== PS1 IPP Czar Logs for the week 2015.01.12 - 2015.01.18 == [[PageOutline]] (Up to [wiki:PS1_IPP_CzarLogs PS1 IPP Czar Logs]) === Monday : 2015.01.12 === * 00:40 MEH: clearing a fault 5 diff {{{ difftool -updatediffskyfile -fault 0 -set_quality 42 -diff_id 635299 -skycell_id skycell.1607.047 -dbname gpc1 }}} * 00:50 MEH: stdsci poll still low, restarted * 01:00 MEH: stdlocal set.poll 100 in case do catch up, nightly has priority and be finished in time for fiber switchover * 02:00 MEH: ipp095 overloaded a bit again, to repair * 06:57 Bill: two warps were repeatedly getting fault 4 from the "cannot compute curve of growth psf is invalid everywhere" assertion. Set quality for them {{{ warptool -set_quality 42 -fault 0 -updateskyfile -warp_id 1345413 -skycell_id skycell.0509.081 warptool -set_quality 42 -fault 0 -updateskyfile -warp_id 1345472 -skycell_id skycell.1084.070 }}} * 07:00 Bill ipp090 and ipp093 have rather high loads as well. * but processing rate was generally okay * 10:20 MEH: clearing some fault 5 WSdiffs {{{ difftool -updatediffskyfile -fault 0 -set_quality 42 -diff_id 635300 -skycell_id skycell.1607.038 -dbname gpc1 difftool -updatediffskyfile -fault 0 -set_quality 42 -diff_id 635300 -skycell_id skycell.1607.047 -dbname gpc1 difftool -updatediffskyfile -fault 0 -set_quality 42 -diff_id 635450 -skycell_id skycell.1247.076 -dbname gpc1 }}} * 10:30 MEH: clearing some old WSdiff that won't complete due to old and image cleanup.. -- when time may be able to update warps to complete but... {{{ difftool -updaterun -diff_id 634284 -dbname gpc1 -set_label ThreePi.WS.nightlyscience.missed difftool -updaterun -diff_id 621228 -dbname gpc1 -set_label OSS.WS.nightlyscience.missed }}} * did manual chip+warp updates when necessary**, were actually from fault 5 and need to have quality set so now remaining diff should be released ok {{{ difftool -updatediffskyfile -fault 0 -set_quality 42 -diff_id 634284 -skycell_id skycell.0679.021 -dbname gpc1 difftool -updatediffskyfile -fault 0 -set_quality 42 -diff_id 621228 -skycell_id skycell.1529.088 -dbname gpc1 }}} * **chips for 1/10 were cleaned up already, but not warps * 11:50 MEH: starting stdlocal chip only test for 10G link usage, stopping all other processing since will need cleared for fiber changes in couple hours anyways * poll~100 chip only uses ~4Gbits/s both in and out on link * poll~150 uses ~6-7Gbits/s both in and out on link * nightly overall uses ~5-7Gbits/s mostly 6-7G out and 5-6G in -- so even 150 w/ stdlocal was a bit high during nightly * 11:55 MEH: ippc11 screen issue, {{{ There are screens on: 13722.robo (Detached) 12542.RoboCzar (Dead ???) 19893.CzarPoll (Dead ???) 13565.plot (Detached) Remove dead screens with 'screen -wipe'. }}} * wiped, cleared other names and restarted screens * 14:40 MEH: possible changes to network plans -- setting stdlocal back to full poll until nightly starts * 19:40 MEH: nightly starting, stdlocal set.poll 100 -- seems to be just at limit, maybe a slow slippage in nightly chip processing which will eventually turn chip-warp off in stdlocal. ~1300 camera entries so not having stdlocal auto-off would be good.. so set.poll 80 === Tuesday : 2015.01.13 === * 07:45 MEH: seems stdlocal set.poll~80 is a limit, still caused enough of a backup until about 2am when stdlocal triggered auto-stop chip-warp, stdsci processing mostly kept up however, including registration. stdlocal set.poll back to 400 ~0830 * 09:10 MEH: restarting pstamp * 10:40 MEH: processing to stop -- going to start modifying systems * decomission jaws nodes -- move pantasks to ippc01-c09 -- * ippx037-x044 + 4849 switch to n5k * stsci ip to .30. network * redistribute other ippx across multiple 4849 on 6509e * 15:45 MEH: stsci IP moved to .30. net, x037-x044 moved to nk5 and .20. net, all pantasks moved off jaws nodes -- Gene gave all clear for starting nightly pantasks to be ready for data tonight * while processing down, going to rotate the apache ippc0x logs -- * finding stagnant stsci mounts on ippc18,c19,c11 so far.. starting to force umount.. -- cleared ippc18,c19; Gene working on c11 * this why ipp032 high load doing nothing -- ipp032 clear and Serge/MOPS reports all good, all other datanodes seem fine * Gene ippc20-c32-ish also has stsci mounts -- Gene working on -- looks like stdsci and pstamp using some of these mixed LANL compute nodes in c2 group -- stsci and pstamp stop until cleared.. summitcopy+registration can continue * 17:05 MEH: Gene cleared compute mounts, stsci+pstamp back on. revert update faults caused to ps_ud_QUB * 17:25 CZW: stdlanl and stdlocal restarted and running. At Gene's recommendation, the poll on stdlocal has been decreased to 100 as the default setting until we see how things behave. * 19:20 MEH: normal processing * looks like summitcopy ~5-8 behind, registration ~2-3 -- normal * WS diffs running seem more load on ipp093,095,090 and may need to go to repair * waiting for network 10G link to have some regularity -- @1930 -- ~6/5G into/out 6509e w/ poll 100 (roughly that running but warps+camera+stacks also so not clear) * 19:35 MEH: larger fault spike, scrambled summitcopy order some, but catching up * 19:45 MEH: fallout from another massive fault spike -- took about 10 min to recover registration {{{ 150113 19:32:34InnoDB: Warning: cannot find a free slot for an undo log. Do you have too 150113 19:42:30InnoDB: Warning: cannot find a free slot for an undo log. Do you have too }}} * 20:00 MEH: once things pick back up, load on 10G link is ~7/7-8G.. both ipp/jaws networks are ~2G * not-targeted ipp092 repair -- to many want to write to node in WT mode * generally running much better @100 chip in stdlocal * 20:30 MEH: stdlocal camera,warp off so chip poll 200 will have a full 200 chips running at a time * 20:45 MEH: another massive fault wave ~2045 -- {{{ 150113 20:47:52InnoDB: Warning: cannot find a free slot for an undo log. Do you have too }}} * 21:00 MEH: chip poll 200 -- reaching 8-9G on 10G link but... * more rounds of massive faults -- backlog getting close to shutoff.. may not get good measure.. * 21:15 MEH: is there something else running overloading ippdb00? '''stdlanl?''' * try stop cleanup to free up connections? * 21:20 MEH: another huge fault wave -- cleanup still flushing out -- === Wednesday : YYYY.MM.DD === === Thursday : YYYY.MM.DD === === Friday : YYYY.MM.DD === === Saturday : YYYY.MM.DD === === Sunday : YYYY.MM.DD ===