IPP Software Navigation Tools IPP Links Communication Pan-STARRS Links

Changes between Version 60 and Version 61 of PS1_IPP_Czarlog_20130415


Ignore:
Timestamp:
Apr 19, 2013, 11:08:06 AM (13 years ago)
Author:
Mark Huber
Comment:

--

Legend:

Unmodified
Added
Removed
Modified
  • PS1_IPP_Czarlog_20130415

    v60 v61  
    2929 * 10:50 Haydn replaced BBU for raid in ipp064
    3030 * 10:55 Haydn adding 32G memory to ippc13
    31  * lots to add..
    32  * 17:50 MEH: everything stalled/faulting.. nebulous not working.. restarting apache ippc08,09 seemed to clear up.
     31 * 13:00 Rita/Gavin finished with switch rewiring, ippc01-ippc16 back up and running. starting reboot of stsci machine to load new 3.7.6 kernel. can do a normal reboot procedure from console since won't power down disks. don't have access to power management for power cycle (good thing) so Rita sticking around at MHPCC in case power cycle needed.
     32{{{
     33sudo tcsh
     34su -
     35sync;sync;sync; umount -a; df -h
     36reboot -f
     37uname -a
     38}}}
     39  * stsci00 has an aspensys user logged in, starting with stsci02 and came back up. stsci03,04,05,07 all needed to do local disk scans and are okay.
     40 * 13:30 MEH: Gavin noticed stsci02 network bond config with the 3.7.6 wasn't the same as Aspen configured them, fixed and is rebooting all the stsci nodes again to correct.
     41 * 14:20 MEH: restarted mysql on ipp019, 020
     42 * 14:30 MEH: ippc13 down again, Gavin rebooted, maybe bad memory module (Haydn, Gavin working on memtest later). taking out of processing
     43 * 16:40 MEH: Gavin notes ippc07 at network speed of 100Mbit.. taking out of the nebulous host set and out of processing, someone will look at it tomorrow
     44  * in looking at the load problem noticed that pstamp loads its hosts multiple times in the input file rather than the default of once, having the number of loading defined in the pantasks_hosts.input file
     45 * 17:50 MEH: now everything stalled/faulting..
     46  * nebulous tools and such not working.. stopping processing and restarting apache on ippc01-c09. when did ippc08,09 seemed to clear up
    3347 * 18:30 MEH: MOPS reporting still  missing 3 exposures but nothing showing up. once system gets going again will try and trace back
    34  * 22:30 MEH: processing has been behaving like the parking brake is on.. many red 20TB disks and seen similar behavior before when this happened, but seems more extreme. stsci neb-host up no help.
     48 * 22:30 MEH: processing has been behaving like the parking brake is on.. many red 20TB disks and seen similar behavior before when this happened, but seems more extreme.
     49  * stsci nodes have been fine with new kernel, neb-host up with note of new kernel. no improvement, but now skycell products will go there to ease space usage on other data nodes
     50{{{
     51neb-host stsci00 up --note "upgraded kernel to 3.7.6"
     52}}}
    3553  * even cleanup is painfully slow, so maybe an issue with nebulous or raid speed on a large portion of machines -- allowing a few more commented out nebservers (ippc04,06) but leaving ippc07 out due to network issue
    3654  * cleared old ippc17.0 stalled mounts
    3755  * stsci has dqstats process that uses datastore, had problem with missing dir so set to off and emailed Bill
    3856  * removing ThreePi.WS.nightlyscience label to focus on primary nightly science
     57  * raidstatus?
     58{{{
     59cat ~ipp/raidstatus/*
     60--> ip057 raid pratially degraded but someone put into repair @0923 as it is rebuilt
     61}}}
     62
    3963=== Wednesday : 2013.04.17 ===
    4064Bill is czar today. Oh joy.