== PS1 IPP Czar Logs for the week 2013.10.28 - 2013.11.03 == [[PageOutline]] (Up to [wiki:PS1_IPP_CzarLogs PS1 IPP Czar Logs]) === Monday : 2013.10.28 === * 12:00 heather: restarted stdsci - it had sine waves on ganglia * 19:58 Bill: is running a pantasks_server on ipp064 out of ~bills/sas.30 running sas.30 staticsky. Once one set of compute3 nodes goes to off state in deepsky I will add one set of compute3 to the pantasks. * 21:00 EAM : rebooted ipp041 as it has been having some NFS oddness * 21:04 Bill: added compute3 to my sas.30 pantasks. Set ippc32, 33, 34, 35, 38, 40, 42, 45, 49, 58, and 62 to off since they are still working on stacks * 23:30 MEH: doesn't look like will be much nightly, fixing LAP cam+warps that have languished a week to finish the stacks.. === Tuesday : 2013.10.29 === * 09:45 Bill: queued SAS.20131030 (EXTENDED_SOURCE_FITS_POISSON = FALSE) * 11:20 MEH: regular restart of stdsci * 12:29 Bill: set compute3 to off in stdscience to use for sas.30 * 12:59 switched sas.30 to 2 x wave4 + 1 x compute 2 to avoid memory problems with deepstacks * put compute3 back into stdscience * stack is stopped as I am using the hosts * 16:55 MEH: ipp061 nfs/mounts have been stuck for a while.. can't even access own export disk. restarting nfs and working again and backlogged jobs clearing * 18:07 Bill: SAS.30 processing is done. Restarted stack pantasks * 18:30 MEH: went though a large number of warps, doing a regular restart of stdsci before nightly === Wednesday : 2013.10.30 === Bill is czar today * 04:50 restarted pstamp and update pantasks * 11:45 MEH: stdsci needs regular restart, also turning MD03,04 diffs with new refstacks -- tweak_ssdiff to get the MD03 marginal from last night * 14:50 MEH: need to get through backlog of chip cleanup, other stages off * 18:30 MEH: again >100k in warps, restarting stdsci before nightly starts * 18:36 Bill: restarted registration and summit copy pantasks (their pcontrols were spinning) * reverted faulted chip after {{{ copied valid /data/ippb03.1/nebulous/42/5a/2169045416.gpc1:20120421:o6038g0140o:o6038g0140o.ota64.fits over corrupt /data/ipp043.0/nebulous/42/5a/2021195867.gpc1:20120421:o6038g0140o:o6038g0140o.ota64.fits }}} * 20:20 MEH: looks like ipp061 is unhappy.. * 20:30 Bill killed hung summitcopy and registration processes on ipp061 took it out of the host lists. Sent email to heather about overload by 96GB ippdvo java process * 20:51 java killed but ipp061 is still in a bad way. There are some stack processes there that are probably stuck forever. Tried restarting nfs but it still can't see some nodes. Machine probably needs a power cycle * 01:00 MEH: cleared md09 stack fault 5, stack_id 2881582 === Thursday : 2013.10.31 === Bill is czar today * 05:55 ipp038 has crashed. Nothing on the console. Attempted power cycle, but no response. * 06:25 set ipp038 to down in nebulous to prevent processes from looking for detrend files with instances there * 06:42 many processes are stuck and have been since 05:02 about the time ipp038 died. Setting stdscience to stop to make it clear which processes to kill * 07:24 old stdscience processes killed. stdscience restarted with poll limit 42 === Friday : 2013.11.01 === === Saturday : 2013.11.02 === === Sunday : 2013.11.03 ===