Changes between Version 89 and Version 90 of PS1_IPP_Czarlog_20141208
- Timestamp:
- Dec 12, 2014, 5:54:38 AM (12 years ago)
Legend:
- Unmodified
- Added
- Removed
- Modified
-
PS1_IPP_Czarlog_20141208
v89 v90 121 121 * 11:20 MEH: ipp084, 087 watch in neb-host up state, if okay then can add to list for larger rwsize for 10G machines 122 122 * 11:25 MEH: restart ippmd -- ipps as normal, adding x0+1b and x0+1 (excluding those for relastro) 123 * last week tested deep stacks and summitcopy on new storage nodes s4,5 -- now test some load of chip-warp -- 123 * last week tested deep stacks and summitcopy on new storage nodes s4,5 -- now test some load of chip-warp -- no time before nebulous shutdown.. 124 124 * 11:40 CZW: At Gene's request, I've sent stop commands to all the ipp/ipplanl pantasks to attempt to catch up the nebulous database replication. 125 125 * 15:xx Gene finished restarting db00 … … 131 131 * ippsXX net io in ~40-60MB/s while ippxNNN and c nodes are ~10-20MB/s 132 132 * 16:40 Gene will not need the subset of x0+1 nodes for another week, can use for processing again 133 * 17:00 MEH: after reconfig on re did clean, turning things back up to see where limit is (if any) before nightly science133 * 17:00 MEH: after reconfig on residual cleaning, turning things back up to see where limit is (if any) before nightly science 134 134 * at stdlocal~363, ippmd~398 chip processing is 2-3x longer than average -- this will be a problem if nightly is the same 135 135 * for nightly science config -- want to avoid any stack processing overlap w/ stdlocal (or ippmd later) -- … … 158 158 159 159 }}} 160 161 160 * 01:00 MEH: registration in odd state again -- stalled o7003g0375o exposure stalling o7003g0376o stuck in pending_burntool. couple ota stuck >3ks and didn't timeout.. 162 161 * getting odd error when regtool -revertprocessedimfile -dbname gpc1 … … 170 169 database error 171 170 }}} 172 173 174 171 * 01:30 MEH: all pantasks seemed to be hit with 10-20 stalled jobs not clearing.. 175 * 04:50 MEH: registration looks to have seq faulted around 0330.. 172 * 04:50 MEH: registration looks to have seq faulted around 0330.. 173 * Chris also found reason for reg revert problem -- rawImfile had fault but processed chip -- 174 {{{ 175 | exp_id | class_id | chip_id | fault | fault | data_state | state | state | 176 +--------+----------+---------+-------+-------+------------+-------+-------+ 177 | 836124 | XY60 | 1321675 | 2 | 0 | full | full | full | 178 | 836124 | XY73 | 1321675 | 2 | 0 | full | full | full | 179 }}} 176 180 * 05:35 MEH: and another cluster of excess db00 transaction causing all sorts of faults -- 177 181 {{{
