IPP Software Navigation Tools IPP Links Communication Pan-STARRS Links

Changes between Version 40 and Version 41 of PS1_IPP_Czarlog_20131007


Ignore:
Timestamp:
Oct 9, 2013, 11:16:44 AM (13 years ago)
Author:
Mark Huber
Comment:

--

Legend:

Unmodified
Added
Removed
Modified
  • PS1_IPP_Czarlog_20131007

    v40 v41  
    3434 * 09:25 adding LAP label back in now --
    3535 * 09:40 MEH: continuing general preventative maintenance  -- restarting long running pantasks (summit, cleanup, distribution, publishing, stack; did registration last night) -- been month+ since last log archive to clear up space on ippc18.0
    36   * ippc18 disk is overused, cannot rsync the logs ippc18.0 to ippc18.1 w/o over-driving load and killing processing.. suspect too many accounts are writing too many logs to home disk.  have to set rsync to <800KPS
     36  * ippc18 disk is overused, cannot rsync the logs ippc18.0 to ippc18.1 w/o over-driving load and killing processing.. suspect too many accounts are writing too many logs to home disk.  have to set rsync to <2000KPS
    3737 * 09:50 MEH: ipp049 had long running (>2hr) pswarp and ppImage, killed and reverted but looks like another pair ppImage stalling --
    38  * 10:00 MEH: working in nodes that have been recently fixed
     38 * 10:00 MEH: working in nodes that have been recently fixed or out of processing for disk work etc
    3939  * ipp031,032
     40  * ipp034
     41  * ipp041
    4042  * ipp047
     43  * ippc12
    4144 * 11:15 MEH: still no LAP stacks, putting 2x compute3 from stack into pstamp for MOPS (18.5k jobs)
    4245  * possibly a lot of data for MOPS on ipp017 requested at once? driving load >200.. settling down now.  need to check if also was from access by UMD for high throughput tests at that time
     
    4548 * 13:00 MEH: LAP has worked through the stalled updates and many stacks being loaded. switching STS reprocessing label back to its original 202 priority ('''this is above LAP''')
    4649 * 13:50 MEH: ipp047 has be on/off being horrible overloaded in RAM and load.. probably also need powercycle -- finally got a login and killed off ppImage, Heather also stopping jobs and seems to be back -- have to watch the RAM use by mysql
     50  * ipp049 is a terrible RAM use state by mysql (>90%) and cannot do much.. all machines with MYSQL need to be monitored and mysql restarted why significantly using RAM...
    4751 * 14:00 MEH: looks like ippdb01 has died.. -- nothing on console, powercycle seems to be rebooting now (will take a terribly long time for RAM check)
    4852 * 16:00 MEH: MD10z stack fault 5 due to all input warps having bad PSF. just setting to quality 42, probably should be 13006 (PPSTACK_ERR_DATA)