| Version 33 (modified by , 12 years ago) ( diff ) |
|---|
PS1 IPP Czar Logs for the week 2015.01.12 - 2015.01.18
(Up to PS1 IPP Czar Logs)
Monday : 2015.01.12
- 00:40 MEH: clearing a fault 5 diff
difftool -updatediffskyfile -fault 0 -set_quality 42 -diff_id 635299 -skycell_id skycell.1607.047 -dbname gpc1
- 00:50 MEH: stdsci poll still low, restarted
- 01:00 MEH: stdlocal set.poll 100 in case do catch up, nightly has priority and be finished in time for fiber switchover
- 02:00 MEH: ipp095 overloaded a bit again, to repair
- 06:57 Bill: two warps were repeatedly getting fault 4 from the "cannot compute curve of growth psf is invalid everywhere" assertion. Set quality for them
warptool -set_quality 42 -fault 0 -updateskyfile -warp_id 1345413 -skycell_id skycell.0509.081 warptool -set_quality 42 -fault 0 -updateskyfile -warp_id 1345472 -skycell_id skycell.1084.070
- 07:00 Bill ipp090 and ipp093 have rather high loads as well.
- but processing rate was generally okay
- 10:20 MEH: clearing some fault 5 WSdiffs
difftool -updatediffskyfile -fault 0 -set_quality 42 -diff_id 635300 -skycell_id skycell.1607.038 -dbname gpc1 difftool -updatediffskyfile -fault 0 -set_quality 42 -diff_id 635300 -skycell_id skycell.1607.047 -dbname gpc1 difftool -updatediffskyfile -fault 0 -set_quality 42 -diff_id 635450 -skycell_id skycell.1247.076 -dbname gpc1
- 10:30 MEH: clearing some old WSdiff that won't complete due to old and image cleanup.. -- when time may be able to update warps to complete but...
difftool -updaterun -diff_id 634284 -dbname gpc1 -set_label ThreePi.WS.nightlyscience.missed difftool -updaterun -diff_id 621228 -dbname gpc1 -set_label OSS.WS.nightlyscience.missed
- did manual chip+warp updates when necessary, were actually from fault 5 and need to have quality set so now remaining diff should be released ok
difftool -updatediffskyfile -fault 0 -set_quality 42 -diff_id 634284 -skycell_id skycell.0679.021 -dbname gpc1 difftool -updatediffskyfile -fault 0 -set_quality 42 -diff_id 621228 -skycell_id skycell.1529.088 -dbname gpc1
- chips for 1/10 were cleaned up already, but not warps
- did manual chip+warp updates when necessary, were actually from fault 5 and need to have quality set so now remaining diff should be released ok
- 11:50 MEH: starting stdlocal chip only test for 10G link usage, stopping all other processing since will need cleared for fiber changes in couple hours anyways
- poll~100 chip only uses ~4Gbits/s both in and out on link
- poll~150 uses ~6-7Gbits/s both in and out on link
- nightly overall uses ~5-7Gbits/s mostly 6-7G out and 5-6G in -- so even 150 w/ stdlocal was a bit high during nightly
- 11:55 MEH: ippc11 screen issue,
There are screens on: 13722.robo (Detached) 12542.RoboCzar (Dead ???) 19893.CzarPoll (Dead ???) 13565.plot (Detached) Remove dead screens with 'screen -wipe'.
- wiped, cleared other names and restarted screens
- 14:40 MEH: possible changes to network plans -- setting stdlocal back to full poll until nightly starts
- 19:40 MEH: nightly starting, stdlocal set.poll 100 -- seems to be just at limit, maybe a slow slippage in nightly chip processing which will eventually turn chip-warp off in stdlocal. ~1300 camera entries so not having stdlocal auto-off would be good.. so set.poll 80
Tuesday : 2015.01.13
- 07:45 MEH: seems stdlocal set.poll~80 is a limit, still caused enough of a backup until about 2am when stdlocal triggered auto-stop chip-warp, stdsci processing mostly kept up however, including registration. stdlocal set.poll back to 400 ~0830
- 09:10 MEH: restarting pstamp
- 10:40 MEH: processing to stop -- going to start modifying systems
- decomission jaws nodes -- move pantasks to ippc01-c09 --
- ippx037-x044 + 4849 switch to n5k
- stsci ip to .30. network
- redistribute other ippx across multiple 4849 on 6509e
- 15:45 MEH: stsci IP moved to .30. net, x037-x044 moved to nk5 and .20. net, all pantasks moved off jaws nodes -- Gene gave all clear for starting nightly pantasks to be ready for data tonight
- while processing down, going to rotate the apache ippc0x logs --
- finding stagnant stsci mounts on ippc18,c19,c11 so far.. starting to force umount.. -- cleared ippc18,c19; Gene working on c11
- this why ipp032 high load doing nothing -- ipp032 clear and Serge/MOPS reports all good, all other datanodes seem fine
- Gene ippc20-c32-ish also has stsci mounts -- Gene working on -- looks like stdsci and pstamp using some of these mixed LANL compute nodes in c2 group -- stsci and pstamp stop until cleared.. summitcopy+registration can continue
- 17:05 MEH: Gene cleared compute mounts, stsci+pstamp back on. revert update faults caused to ps_ud_QUB
- 17:25 CZW: stdlanl and stdlocal restarted and running. At Gene's recommendation, the poll on stdlocal has been decreased to 100 as the default setting until we see how things behave.
- 19:20 MEH: normal processing
- looks like summitcopy ~5-8 behind, registration ~2-3 -- normal
- WS diffs running seem more load on ipp093,095,090 and may need to go to repair
- waiting for network 10G link to have some regularity -- @1930 -- ~6/5G into/out 6509e w/ poll 100 (roughly that running but warps+camera+stacks also so not clear)
- 19:35 MEH: larger fault spike, scrambled summitcopy order some, but catching up
- 19:45 MEH: fallout from another massive fault spike -- took about 10 min to recover registration
150113 19:32:34InnoDB: Warning: cannot find a free slot for an undo log. Do you have too 150113 19:42:30InnoDB: Warning: cannot find a free slot for an undo log. Do you have too
- 20:00 MEH: once things pick back up, load on 10G link is ~7/7-8G.. both ipp/jaws networks are ~2G
- not-targeted ipp092 repair -- to many want to write to node in WT mode
- generally running much better @100 chip in stdlocal
- 20:30 MEH: stdlocal camera,warp off so chip poll 200 will have a full 200 chips running at a time
- 20:45 MEH: another massive fault wave ~2045 --
150113 20:47:52InnoDB: Warning: cannot find a free slot for an undo log. Do you have too
- 20:50 MEH: ippmd running single deep stack on c29 -- 0 impact
- 21:00 MEH: chip poll 200 -- reaching 8-9G on 10G link but...
- more rounds of massive faults -- backlog getting close to shutoff.. may not get good measure..
- 21:15 MEH: is there something else running overloading ippdb00? stdlanl prep?
- try stop cleanup to free up connections?
- 21:20 MEH: another huge fault wave -- cleanup still flushing out --
- 22:00 MEH: cleanup cleared ~2145 and no faults since ~2130 but was enough overloading to drive stdlocal to auto-off except for stacks.. that's it, no more testing except 10G link load w/ 500 stdlocal stacks maybe... leaving set.poll 200 since seemed okay in case able to catch up in few hours (but likely can go higher)..
Wednesday : YYYY.MM.DD
- 13:50 CZW Stopping ipp/ipplanl processing for Haydn to reboot ipp083/84 for RAID battery issues.
Thursday : YYYY.MM.DD
Friday : YYYY.MM.DD
Saturday : YYYY.MM.DD
Sunday : YYYY.MM.DD
Note:
See TracWiki
for help on using the wiki.
