IPP Software Navigation Tools IPP Links Communication Pan-STARRS Links

Changes between Version 41 and Version 42 of PS1_IPP_Czarlog_20131007


Ignore:
Timestamp:
Oct 9, 2013, 11:27:06 AM (13 years ago)
Author:
Mark Huber
Comment:

--

Legend:

Unmodified
Added
Removed
Modified
  • PS1_IPP_Czarlog_20131007

    v41 v42  
    3737 * 09:50 MEH: ipp049 had long running (>2hr) pswarp and ppImage, killed and reverted but looks like another pair ppImage stalling --
    3838 * 10:00 MEH: working in nodes that have been recently fixed or out of processing for disk work etc
    39   * ipp031,032
    40   * ipp034
    41   * ipp041
    42   * ipp047
    43   * ippc12
     39  * ipp031,032 -- no indication cannot just use as compute type nodes while disks are down still (should they be down still or put into repair?)
     40  * ipp034 -- was in wave2_weak as missing half RAM, fixed so put back into full use
     41  * ipp041 -- mobo had lost 1 CPU so was just datanode, fixed so put back into full use
     42  * ipp047 -- mobo fried, replaced and was in wave3_weak for testing. seems okay so back into full use -- dvo/ipptopsps seems to be overloading it (mysql ram issue??) so back to wave3_weak
     43  * ippc12 -- broken for while, fixed and put back into full compute use
    4444 * 11:15 MEH: still no LAP stacks, putting 2x compute3 from stack into pstamp for MOPS (18.5k jobs)
    4545  * possibly a lot of data for MOPS on ipp017 requested at once? driving load >200.. settling down now.  need to check if also was from access by UMD for high throughput tests at that time
    4646 * 12:00 MEH: addressing low disk space on ippc0x machines due to /tmp/nebulous_server.log
    4747 * 12:35 MEH: MOPS stamps finished, restarting pstamp and putting the 2x compute3 back to stack
    48  * 13:00 MEH: LAP has worked through the stalled updates and many stacks being loaded. switching STS reprocessing label back to its original 202 priority ('''this is above LAP''')
    49  * 13:50 MEH: ipp047 has be on/off being horrible overloaded in RAM and load.. probably also need powercycle -- finally got a login and killed off ppImage, Heather also stopping jobs and seems to be back -- have to watch the RAM use by mysql
     48 * 13:00 MEH: LAP has worked through the stalled updates and many stacks being loaded. switching STS reprocessing label back to its original 202 priority ('''STS prio is above LAP''')
     49 * 13:50 MEH: ipp047 has be on/off being horrible overloaded in RAM and load.. probably also need powercycle -- finally got a login and killed off ppImage, Heather also stopping jobs and seems to be back -- '''have to monitor the RAM use by mysql'''
    5050  * ipp049 is a terrible RAM use state by mysql (>90%) and cannot do much.. all machines with MYSQL need to be monitored and mysql restarted why significantly using RAM...
    5151 * 14:00 MEH: looks like ippdb01 has died.. -- nothing on console, powercycle seems to be rebooting now (will take a terribly long time for RAM check)