IPP Software Navigation Tools IPP Links Communication Pan-STARRS Links

Changes between Version 89 and Version 90 of PS1_IPP_Czarlog_20141208


Ignore:
Timestamp:
Dec 12, 2014, 5:54:38 AM (12 years ago)
Author:
Mark Huber
Comment:

--

Legend:

Unmodified
Added
Removed
Modified
  • PS1_IPP_Czarlog_20141208

    v89 v90  
    121121 * 11:20 MEH: ipp084, 087 watch in neb-host up state, if okay then can add to list for larger rwsize for 10G machines
    122122 * 11:25 MEH: restart ippmd -- ipps as normal, adding x0+1b and x0+1 (excluding those for relastro)
    123   * last week tested deep stacks and summitcopy on new storage nodes s4,5 -- now test some load of chip-warp --
     123  * last week tested deep stacks and summitcopy on new storage nodes s4,5 -- now test some load of chip-warp -- no time before nebulous shutdown..
    124124 * 11:40 CZW: At Gene's request, I've sent stop commands to all the ipp/ipplanl pantasks to attempt to catch up the nebulous database replication.
    125125 * 15:xx Gene finished restarting db00
     
    131131  * ippsXX net io in ~40-60MB/s while ippxNNN and c nodes are ~10-20MB/s
    132132 * 16:40 Gene will not need the subset of x0+1 nodes for another week, can use for processing again
    133  * 17:00 MEH: after reconfig on redid clean, turning things back up to see where limit is (if any) before nightly science
     133 * 17:00 MEH: after reconfig on residual cleaning, turning things back up to see where limit is (if any) before nightly science
    134134  * at  stdlocal~363, ippmd~398 chip processing is 2-3x longer than average -- this will be a problem if nightly is the same
    135135  * for nightly science config -- want to avoid any stack processing overlap w/ stdlocal (or ippmd later) --
     
    158158
    159159}}}
    160 
    161160 * 01:00 MEH: registration in odd state again -- stalled o7003g0375o exposure stalling o7003g0376o stuck in pending_burntool. couple ota stuck >3ks and didn't timeout..
    162161  * getting odd error when regtool -revertprocessedimfile -dbname gpc1
     
    170169     database error
    171170}}}
    172 
    173 
    174171 * 01:30 MEH: all pantasks seemed to be hit with 10-20 stalled jobs not clearing..
    175  * 04:50  MEH: registration looks to have seq faulted around 0330..
     172 * 04:50 MEH: registration looks to have seq faulted around 0330..
     173  * Chris also found reason for reg revert problem -- rawImfile had fault but processed chip --
     174{{{
     175| exp_id | class_id | chip_id | fault | fault | data_state | state | state |
     176+--------+----------+---------+-------+-------+------------+-------+-------+
     177| 836124 | XY60     | 1321675 |     2 |     0 | full       | full  | full  |
     178| 836124 | XY73     | 1321675 |     2 |     0 | full       | full  | full  |
     179}}}
    176180 * 05:35 MEH: and another cluster of excess db00 transaction causing all sorts of faults --
    177181{{{