Changes between Version 41 and Version 42 of PS1_IPP_Czarlog_20131007
- Timestamp:
- Oct 9, 2013, 11:27:06 AM (13 years ago)
Legend:
- Unmodified
- Added
- Removed
- Modified
-
PS1_IPP_Czarlog_20131007
v41 v42 37 37 * 09:50 MEH: ipp049 had long running (>2hr) pswarp and ppImage, killed and reverted but looks like another pair ppImage stalling -- 38 38 * 10:00 MEH: working in nodes that have been recently fixed or out of processing for disk work etc 39 * ipp031,032 40 * ipp034 41 * ipp041 42 * ipp047 43 * ippc12 39 * ipp031,032 -- no indication cannot just use as compute type nodes while disks are down still (should they be down still or put into repair?) 40 * ipp034 -- was in wave2_weak as missing half RAM, fixed so put back into full use 41 * ipp041 -- mobo had lost 1 CPU so was just datanode, fixed so put back into full use 42 * ipp047 -- mobo fried, replaced and was in wave3_weak for testing. seems okay so back into full use -- dvo/ipptopsps seems to be overloading it (mysql ram issue??) so back to wave3_weak 43 * ippc12 -- broken for while, fixed and put back into full compute use 44 44 * 11:15 MEH: still no LAP stacks, putting 2x compute3 from stack into pstamp for MOPS (18.5k jobs) 45 45 * possibly a lot of data for MOPS on ipp017 requested at once? driving load >200.. settling down now. need to check if also was from access by UMD for high throughput tests at that time 46 46 * 12:00 MEH: addressing low disk space on ippc0x machines due to /tmp/nebulous_server.log 47 47 * 12:35 MEH: MOPS stamps finished, restarting pstamp and putting the 2x compute3 back to stack 48 * 13:00 MEH: LAP has worked through the stalled updates and many stacks being loaded. switching STS reprocessing label back to its original 202 priority (''' thisis above LAP''')49 * 13:50 MEH: ipp047 has be on/off being horrible overloaded in RAM and load.. probably also need powercycle -- finally got a login and killed off ppImage, Heather also stopping jobs and seems to be back -- have to watch the RAM use by mysql48 * 13:00 MEH: LAP has worked through the stalled updates and many stacks being loaded. switching STS reprocessing label back to its original 202 priority ('''STS prio is above LAP''') 49 * 13:50 MEH: ipp047 has be on/off being horrible overloaded in RAM and load.. probably also need powercycle -- finally got a login and killed off ppImage, Heather also stopping jobs and seems to be back -- '''have to monitor the RAM use by mysql''' 50 50 * ipp049 is a terrible RAM use state by mysql (>90%) and cannot do much.. all machines with MYSQL need to be monitored and mysql restarted why significantly using RAM... 51 51 * 14:00 MEH: looks like ippdb01 has died.. -- nothing on console, powercycle seems to be rebooting now (will take a terribly long time for RAM check)
