Changes between Version 40 and Version 41 of PS1_IPP_Czarlog_20131007
- Timestamp:
- Oct 9, 2013, 11:16:44 AM (13 years ago)
Legend:
- Unmodified
- Added
- Removed
- Modified
-
PS1_IPP_Czarlog_20131007
v40 v41 34 34 * 09:25 adding LAP label back in now -- 35 35 * 09:40 MEH: continuing general preventative maintenance -- restarting long running pantasks (summit, cleanup, distribution, publishing, stack; did registration last night) -- been month+ since last log archive to clear up space on ippc18.0 36 * ippc18 disk is overused, cannot rsync the logs ippc18.0 to ippc18.1 w/o over-driving load and killing processing.. suspect too many accounts are writing too many logs to home disk. have to set rsync to < 800KPS36 * ippc18 disk is overused, cannot rsync the logs ippc18.0 to ippc18.1 w/o over-driving load and killing processing.. suspect too many accounts are writing too many logs to home disk. have to set rsync to <2000KPS 37 37 * 09:50 MEH: ipp049 had long running (>2hr) pswarp and ppImage, killed and reverted but looks like another pair ppImage stalling -- 38 * 10:00 MEH: working in nodes that have been recently fixed 38 * 10:00 MEH: working in nodes that have been recently fixed or out of processing for disk work etc 39 39 * ipp031,032 40 * ipp034 41 * ipp041 40 42 * ipp047 43 * ippc12 41 44 * 11:15 MEH: still no LAP stacks, putting 2x compute3 from stack into pstamp for MOPS (18.5k jobs) 42 45 * possibly a lot of data for MOPS on ipp017 requested at once? driving load >200.. settling down now. need to check if also was from access by UMD for high throughput tests at that time … … 45 48 * 13:00 MEH: LAP has worked through the stalled updates and many stacks being loaded. switching STS reprocessing label back to its original 202 priority ('''this is above LAP''') 46 49 * 13:50 MEH: ipp047 has be on/off being horrible overloaded in RAM and load.. probably also need powercycle -- finally got a login and killed off ppImage, Heather also stopping jobs and seems to be back -- have to watch the RAM use by mysql 50 * ipp049 is a terrible RAM use state by mysql (>90%) and cannot do much.. all machines with MYSQL need to be monitored and mysql restarted why significantly using RAM... 47 51 * 14:00 MEH: looks like ippdb01 has died.. -- nothing on console, powercycle seems to be rebooting now (will take a terribly long time for RAM check) 48 52 * 16:00 MEH: MD10z stack fault 5 due to all input warps having bad PSF. just setting to quality 42, probably should be 13006 (PPSTACK_ERR_DATA)
