[[PageOutline]] (Up to [wiki:PS1_IPP_CzarLogs PS1 IPP Czar Logs]) === Monday : 2012.10.22 === Mark is czar * 00:00 MEH: after reboot ipp023, restarted summitcopy, registration, stdscience for clean start. ran ''regtool -revertprocessedimfile -dbname gpc1''. * 01:30 MEH: watching processing summit/registration getting stalled often, looks like ipp018 often not responding for replication stage, putting into repair (too many trying to replicate to ipp018 with so many red disks?). been a problem for a while with {{{ INFO: task mount.nfs:29470 blocked for more than 120 seconds. }}} * seems to have now moved the problem to ipp017.. but handling better. ipp018 may want a reboot (been a whole ~23d since last one). * 07:40 oddly, seems like manually running regtool was needed when this has been automatic in the past (and sometimes still is). has something changed in code related to it, is there a number of tries attempted and then stops? {{{ regtool -updateprocessedimfile -exp_id 536377 -class_id XY11 -set_state pending_burntool -dbname gpc1 }}} * 08:45 Bill restarted pstamp and update pantasks in order to have clear success and failure counts. * 09:15 MEH: ipp023 down again.. same problem as last night. rebooting and taking out of processing [wiki:Ipp023-crash-20121022] * 10:25 MEH: nightlyscience 99% through (finished downloading ~09:20) * 12:15 MEH: Bill changed retention for M31+STS for MPG folks, when nightlyscience finished need to ''make install'' in src/ipp-20120802/ippconfig/recipes/ (shutting down and restarting all pantasks) * this may challenge available disk space on 20TB machines, 8/30 aren't in repair or red. * 12:45 MEH: Chris added check for registration burntool stalling condition to ippScripts. pantasks stopped (mostly distribution, pstamp only active) and rebuilt ippScripts only. * 12:50 MEH: reboot testing ipp046 after BIOS battery replacement hasn't been done yet. does it boot on power cycle properly now -- no. sending email to Hayden, Gavin, Rita * 13:25 MEH: reactivating the LAP label {{{ labeltool -dbname gpc1 -updatelabel -label LAP.ThreePi.20120706 -set_active }}} * 14:40 MEH: Rita reported loss of AC for the ippbXX nodes, so shutting down until repaired tomorrow. * replication pantasks stopped * neb-host ipp00.0 down etc for all === Tuesday : YYYY.MM.DD === === Wednesday : YYYY.MM.DD === === Thursday : YYYY.MM.DD === === Friday : YYYY.MM.DD === === Saturday : YYYY.MM.DD === === Sunday : YYYY.MM.DD ===