PS1 IPP Czar Logs for the week 2019.01.28 - 2019.02.03
(Up to PS1 IPP Czar Logs)
Monday : 2019.01.28
- MEH: manually sent old update chips to cleanup to free up space where possible since some nodes were going red -- also included all WSdiffs up through past few days and WWdiffs only through end of 2018 since the WWdiff update bug still exists
- TdB: IPP backup machines ippb16,ippb17,ipp19 taken down for Haydn to convert to Ubuntu
Tuesday : 2019.01.29
- TdB: Summitcopy for gpc2 has been updated to include compression.
Wednesday : 2019.01.30
- TdB: IPP backup machines ippb16,ippb17,ipp19 set back to up and SSH keys reset for users ippitc,ipp,ipps2
Thursday : 2019.01.31
- TdB: IPP backup machines ippb16-20 had automount problems. Gavin fixed those, and machines are set back to up.
Friday : 2019.02.01
- TdB: IPP processing failed due to an issue related to the permissions set on the recently updated backup machines ippb16-20. Owner of nebulous disks was set to "ipp users" instead of "apache nebulous". Machines set to repair while the issue is being fixed and processing has now continued.
- TdB: ippb16-20 permissions for nebulous drives fixed to apache, and Gene is running chown on all files on these machines. Will be set back to up on monday, due to Friday being a bad day for system-wise changes.
- TdB: machine ipp70 removed from /home/panstarrs/ipp/local/bin/apachedisk_chk.sh since it is down (and had already been removed from nebulous), but still spawning warning emails every 4 hours.
- MEH: something caused >250k warps in various update states saved for MOPS and QUB (ps_ud_MOPS%/QUB%) to get send to cleanup what looks to be ~noon, but didn't notice until nagios sent load warning for ippdb09 ~1523 -- all warps are being retained by default, only broken or test ones should be cleaned -- a large number of cleanups will take a long time to get through, while in that state it the more there are the more chance it will block+fault pstamp requests wanting to update part of that warp
- stop cleanup pantasks and manually try to salvage remaining possible warp_id set state update and label update_recovery for all state goto_cleaned with data_state full
warptool -dbname gpc1 -updaterun -state goto_cleaned -label goto_cleaned -set_state update -set_label update_recovery -warp_id XXXX
- stop cleanup pantasks and manually try to salvage remaining possible warp_id set state update and label update_recovery for all state goto_cleaned with data_state full
Saturday : 2019.02.02
- MEH: salvage of warps finished, cleanup back to run -- some are broken (several cycles of fail in cleanup), change label
warptool -dbname gpc1 -updaterun -state goto_cleaned -label goto_cleaned -set_label goto_cleaned.broken
Sunday : YYYY.MM.DD
Last modified
7 years ago
Last modified on Feb 2, 2019, 7:06:46 AM
Note:
See TracWiki
for help on using the wiki.
