== PS1 IPP Czar Logs for the week 2010.01.24 - 2010.01.30 == [[PageOutline]] (Up to [wiki:PS1_IPP_CzarLogs PS1 IPP Czar Logs]) === Monday : 2011.01.24 === * eam : warp seemed to be slow: no progress for about 1 hour. the book was full of DONE runs. I stopped processing, ran 'process_cleanup warpPendingSkyCell' to clear the book, and restarted processing. It ran fine after that * eam : burntool seemed to have gotten stuck. I looked on the new burntool state ippMonitor page and found one of the imfiles did not seem to be making progress. Looking at the pantasks (control status), I noticed that the job for that cell had been running for a very long time (>500 sec). I went to the machine where it was running and noticed that it was hanging on access /data/ipp033.0 (which crashed over the weekend). I used force.umount to clear the mount point, and things moved along from there * bills 15:53 Removed ipp053 from distribution host lists and restarted distribution pantasks. It seemed sluggish anyways. * bills 16:00 cleared some magicDSRun revert faults. I need to automate this! * bills 16:09 updated magic_destreak_cleanup.pl to *not* delete the original uncensored diff stage cmf files. * bills 16:10 set label STS.201009 back to active. * bills 16:25 Executed stacktool -updaterun -set_state drop -stack_id 216256 -set_note 'fails due to problem in ticket 1427' * bills 19:37 several faults have appeared setting STS.201009 back to inactive * bills 21:19 still lots of faults. I suspect that the rsync processes running on this node are related. I did neb-host --host ipp008 --state repair === Tuesday : 2011.01.25 === Bill is czar today * 04:21 Many many faults. ipp012 filesystem is read-only from many nodes. ssh is rejected. Stopping all processing for a few minutes to investigate. === Wednesday : 2011.01.26 === === Thursday : 2011.01.27 === === Friday : 2011.01.28 === === Saturday : 2011.01.29 === === Sunday : 2011.01.30 ===