== PS1 IPP Czar Logs for the week 2014.10.13 - 2014.10.20 == [[PageOutline]] (Up to [wiki:PS1_IPP_CzarLogs PS1 IPP Czar Logs]) === Monday : 2014.10.13 === * if pixelservers shutdown described in PS1 shutdown list, will want to be sure all nightly data downloaded by early morning * 06:00 MEH: nightly data downloaded * 09:00 MEH: no nightly data until ~Wednesday -- disable nightly shutoff of nodes in lanl stdlocal * looks like lanl stdlocal needs a regular restart anyways * after restart will see if this works tonight or if there is something else that also needs to be flipped.. -- server input tweak_offstorage * 12:03 Bill: restarted pstamp pantasks and queued ps_ud label for cleanup. * also dropped warp_id 453790 whose corresponding camRun 472138 was in state drop, which prevents updates for that warpRun from succeeding. * 13:41 Bill: turned MPE label off in pstamp. There is a large backlog of requests to be cleaned for that label and the new ones are running so fast that we are falling behind on space. * 14:15 cleanup is done 75% full. Adding MPE label back in. === Tuesday : 2014.10.14 === * 09:20 MEH: Ken has noted lanl stdlocal has been pretty much idle since ~4am -- looks like handful mem fault warps and many fault 2 stacks can try to be cleared. actually only one was mem fault, three are this and don't know what has been done for this case {{{ cannot build growth curve (psf model is invalid everywhere) Backtrace depth: 11 Backtrace 0: p_psAssert Backtrace 1: pmGrowthCurveGenerate Backtrace 2: psphotMakeGrowthCurve Backtrace 3: psphotChoosePSFReadout Backtrace 4: psphotChoosePSF Backtrace 5: psphotReadoutFindPSF Backtrace 6: (unknown) Backtrace 7: (unknown) Backtrace 8: (unknown) Backtrace 9: __libc_start_main Backtrace 10: (unknown) Assertion failed in function pmGrowthCurveGenerate at pmGrowthCurveGenerate.c:89. Error stack: }}} * 10:45 CZW: restarting stdlocal, stopping stdlanl to do LANL side cleanup. === Wednesday : 2014-10-15 === * 8:45 CZW: All ipp/ipplanl pantasks servers shutdown in preparation for MRTC-B power outage. * 14:15 CZW: Most of the nodes are back online, so I've restarted the ipp/ipplanl pantasks. I'm going to allow LANL processing to happen, although I'm still keeping an eye on the disk usage there. === Thursday : YYYY.MM.DD === * 13:00 CZW: I've put ipp082 into repair because it has a very high load. It's responsive, but I want to see if letting it cool off a bit will clear some of the fault 2 issues I'm seeing in processing. === Friday : YYYY.MM.DD === === Saturday : YYYY.MM.DD === === Sunday : YYYY.MM.DD ===