IPP Software Navigation Tools IPP Links Communication Pan-STARRS Links
wiki:PS1_IPP_Czarlog_20190128

PS1 IPP Czar Logs for the week 2019.01.28 - 2019.02.03

(Up to PS1 IPP Czar Logs)

Monday : 2019.01.28

  • MEH: manually sent old update chips to cleanup to free up space where possible since some nodes were going red -- also included all WSdiffs up through past few days and WWdiffs only through end of 2018 since the WWdiff update bug still exists
  • TdB: IPP backup machines ippb16,ippb17,ipp19 taken down for Haydn to convert to Ubuntu

Tuesday : 2019.01.29

  • TdB: Summitcopy for gpc2 has been updated to include compression.

Wednesday : 2019.01.30

  • TdB: IPP backup machines ippb16,ippb17,ipp19 set back to up and SSH keys reset for users ippitc,ipp,ipps2

Thursday : 2019.01.31

  • TdB: IPP backup machines ippb16-20 had automount problems. Gavin fixed those, and machines are set back to up.

Friday : 2019.02.01

  • TdB: IPP processing failed due to an issue related to the permissions set on the recently updated backup machines ippb16-20. Owner of nebulous disks was set to "ipp users" instead of "apache nebulous". Machines set to repair while the issue is being fixed and processing has now continued.
  • TdB: ippb16-20 permissions for nebulous drives fixed to apache, and Gene is running chown on all files on these machines. Will be set back to up on monday, due to Friday being a bad day for system-wise changes.
  • TdB: machine ipp70 removed from /home/panstarrs/ipp/local/bin/apachedisk_chk.sh since it is down (and had already been removed from nebulous), but still spawning warning emails every 4 hours.
  • MEH: something caused >250k warps in various update states saved for MOPS and QUB (ps_ud_MOPS%/QUB%) to get send to cleanup what looks to be ~noon, but didn't notice until nagios sent load warning for ippdb09 ~1523 -- all warps are being retained by default, only broken or test ones should be cleaned -- a large number of cleanups will take a long time to get through, while in that state it the more there are the more chance it will block+fault pstamp requests wanting to update part of that warp
    • stop cleanup pantasks and manually try to salvage remaining possible warp_id set state update and label update_recovery for all state goto_cleaned with data_state full
      warptool -dbname gpc1 -updaterun -state goto_cleaned -label goto_cleaned -set_state update -set_label update_recovery -warp_id XXXX
      

Saturday : 2019.02.02

  • MEH: salvage of warps finished, cleanup back to run -- some are broken (several cycles of fail in cleanup), change label
    warptool -dbname gpc1 -updaterun -state goto_cleaned -label goto_cleaned -set_label goto_cleaned.broken
    

Sunday : YYYY.MM.DD

Last modified 7 years ago Last modified on Feb 2, 2019, 7:06:46 AM
Note: See TracWiki for help on using the wiki.