[[PageOutline]] (Up to [wiki:PS1_IPP_CzarLogs PS1 IPP Czar Logs]) === Monday : 2012.10.22 === Mark is czar * 00:00 MEH: after reboot ipp023, restarted summitcopy, registration, stdscience for clean start. ran ''regtool -revertprocessedimfile -dbname gpc1''. * 01:30 MEH: watching processing summit/registration getting stalled often, looks like ipp018 often not responding for replication stage, putting into repair (too many trying to replicate to ipp018 with so many red disks?). been a problem for a while with {{{ INFO: task mount.nfs:29470 blocked for more than 120 seconds. }}} * seems to have now moved the problem to ipp017.. but handling better. ipp018 may want a reboot (been a whole ~23d since last one). * 07:40 oddly, seems like manually running regtool was needed when this has been automatic in the past (and sometimes still is). has something changed in code related to it, is there a number of tries attempted and then stops? {{{ regtool -updateprocessedimfile -exp_id 536377 -class_id XY11 -set_state pending_burntool -dbname gpc1 }}} * 08:45 Bill restarted pstamp and update pantasks in order to have clear success and failure counts. * 09:15 MEH: ipp023 down again.. same problem as last night. rebooting and taking out of processing [wiki:Ipp023-crash-20121022] * 10:25 MEH: nightlyscience 99% through (finished downloading ~09:20) * 12:15 MEH: Bill changed retention for M31+STS for MPG folks, when nightlyscience finished need to ''make install'' in src/ipp-20120802/ippconfig/recipes/ (shutting down and restarting all pantasks) * this may challenge available disk space on 20TB machines, 8/30 aren't in repair or red. * 12:45 MEH: Chris added check for registration burntool stalling condition to ippScripts. pantasks stopped (mostly distribution, pstamp only active) and rebuilt ippScripts only. * 12:50 MEH: reboot testing ipp046 after BIOS battery replacement hasn't been done yet. does it boot on power cycle properly now -- no. sending email to Hayden, Gavin, Rita * 13:25 MEH: reactivating the LAP label {{{ labeltool -dbname gpc1 -updatelabel -label LAP.ThreePi.20120706 -set_active }}} * 14:40 MEH: Rita reported loss of AC for the ippbXX nodes, so shutting down until repaired tomorrow. * replication pantasks stopped * neb-host ippb00 down etc for all 4 * 14:50 MEH: many LAP faults, turning all reverts off to avoid wasting cycles on the same faults until fixed. * 15:10 MEH: adding deepstack's compute3 into pstamp for 100% increase in nodes to help move through MOPS request. * 17:00 MEH: something is spiking load >150 on stsci nodes, skycell_jpeg is very backlogged * 19:00 MEH: some 43 camera distribution fault 2 from 10/16, reverted and cleared ''disttool -revertrun -dbname gpc1 -fault 2 -label ThreePi.nightlyscience'' * 22:30 MEH: checked on the exposures o6218g0006o--o6218g0054o reported to be missing for DQstats from otis, they all appear to be processed and there on the datastore from 10/18. === Tuesday : YYYY.MM.DD === === Wednesday : YYYY.MM.DD === === Thursday : YYYY.MM.DD === === Friday : YYYY.MM.DD === === Saturday : YYYY.MM.DD === === Sunday : YYYY.MM.DD ===