== PS1 IPP Czar Logs for the week 2010.01.24 - 2010.01.30 == [[PageOutline]] (Up to [wiki:PS1_IPP_CzarLogs PS1 IPP Czar Logs]) === Monday : 2011.01.24 === * eam : warp seemed to be slow: no progress for about 1 hour. the book was full of DONE runs. I stopped processing, ran 'process_cleanup warpPendingSkyCell' to clear the book, and restarted processing. It ran fine after that * eam : burntool seemed to have gotten stuck. I looked on the new burntool state ippMonitor page and found one of the imfiles did not seem to be making progress. Looking at the pantasks (control status), I noticed that the job for that cell had been running for a very long time (>500 sec). I went to the machine where it was running and noticed that it was hanging on access /data/ipp033.0 (which crashed over the weekend). I used force.umount to clear the mount point, and things moved along from there * bills 15:53 Removed ipp053 from distribution host lists and restarted distribution pantasks. It seemed sluggish anyways. * bills 16:00 cleared some magicDSRun revert faults. I need to automate this! * bills 16:09 updated magic_destreak_cleanup.pl to *not* delete the original uncensored diff stage cmf files. * bills 16:10 set label STS.201009 back to active. * bills 16:25 Executed stacktool -updaterun -set_state drop -stack_id 216256 -set_note 'fails due to problem in ticket 1427' * bills 19:37 several faults have appeared setting STS.201009 back to inactive * bills 21:19 still lots of faults. I suspect that the rsync processes running on this node are related. I did neb-host --host ipp008 --state repair === Tuesday : 2011.01.25 === Bill is czar today * 04:21 Many many faults. ipp012 filesystem is read-only from many nodes. ssh is rejected. Stopping all processing for a few minutes to investigate. * 04:25 ipp012 console unresponsive. 'Only output from console: INIT: Id "s0" r[220286.000299] Kernel panic - not syncing: Attempted to kill init!' cycling power * 04:37 processing restarted . Space is running out. * 04:39 There is a very large postage stamp request (> 15000 jobs from MPIA) Setting MPIA labels to inactive for now. * 04:55 distribution pantasks jobs for ipp012 are all failing. setting ipp012 to off. * 07:11 many nodes over 98%. Stopped processing except for summit copy and registration. All of last night's data has been copied but there are 308 still working on registration & burntool. * 07:23 ippdb00:/tmp is full which is causing nebulous errors. Ran: 'sudo mv /tmp/nebulous_server.log /export/ippdb00.0/nebulous_server.log-20110125' Still has zero free. * 08:00 fixed register exp problem being caused by invalid value for $default_host in registration pantasks.Changed it from ipp023 to any * 08:30 burntool is proceeding. * 08:30 since magic is way backlogged and since it doesn't use much space I've turned distribution back to run with destreak.off * 09:26 fixed broken magic_cleanup script. Data is being recovered on ipp053. Turning processing back on. * 09:41 turned on STS label set priority to 1000. Once those exposures are done and distributed we can clean up the data. * 11:13 turned off chip to allow the sts warps to progress faster. * 11:00 ipp053 is < 98% now. Queued magicRuns on ipp050 for cleanup. It's down to 97% as well. * 13:15 chip.on * 13:31 lowered priority of STS.201009 below MD% but above 3pi * 15:24 removed unneeded labels from the list of labels and survey book entries in stdscience. This should lower the latency of getting things queued. To do this execute the following pantasks command "server input del.labels.january" * 15:58 Taking ThreePi.nightlyscience out of survey.publish while I work on enabling publication of muggled data. * 16:54 in stdscience did survey.off to prevent new runs from getting queued for now. * 17:23 I'm testing publishing out of my build. ~ipp/publishing is off but rest assured some data is getting processed. === Wednesday : 2011.01.26 === === Thursday : 2011.01.27 === === Friday : 2011.01.28 === === Saturday : 2011.01.29 === === Sunday : 2011.01.30 ===