Changes between Version 60 and Version 61 of PS1_IPP_Czarlog_20130415
- Timestamp:
- Apr 19, 2013, 11:08:06 AM (13 years ago)
Legend:
- Unmodified
- Added
- Removed
- Modified
-
PS1_IPP_Czarlog_20130415
v60 v61 29 29 * 10:50 Haydn replaced BBU for raid in ipp064 30 30 * 10:55 Haydn adding 32G memory to ippc13 31 * lots to add.. 32 * 17:50 MEH: everything stalled/faulting.. nebulous not working.. restarting apache ippc08,09 seemed to clear up. 31 * 13:00 Rita/Gavin finished with switch rewiring, ippc01-ippc16 back up and running. starting reboot of stsci machine to load new 3.7.6 kernel. can do a normal reboot procedure from console since won't power down disks. don't have access to power management for power cycle (good thing) so Rita sticking around at MHPCC in case power cycle needed. 32 {{{ 33 sudo tcsh 34 su - 35 sync;sync;sync; umount -a; df -h 36 reboot -f 37 uname -a 38 }}} 39 * stsci00 has an aspensys user logged in, starting with stsci02 and came back up. stsci03,04,05,07 all needed to do local disk scans and are okay. 40 * 13:30 MEH: Gavin noticed stsci02 network bond config with the 3.7.6 wasn't the same as Aspen configured them, fixed and is rebooting all the stsci nodes again to correct. 41 * 14:20 MEH: restarted mysql on ipp019, 020 42 * 14:30 MEH: ippc13 down again, Gavin rebooted, maybe bad memory module (Haydn, Gavin working on memtest later). taking out of processing 43 * 16:40 MEH: Gavin notes ippc07 at network speed of 100Mbit.. taking out of the nebulous host set and out of processing, someone will look at it tomorrow 44 * in looking at the load problem noticed that pstamp loads its hosts multiple times in the input file rather than the default of once, having the number of loading defined in the pantasks_hosts.input file 45 * 17:50 MEH: now everything stalled/faulting.. 46 * nebulous tools and such not working.. stopping processing and restarting apache on ippc01-c09. when did ippc08,09 seemed to clear up 33 47 * 18:30 MEH: MOPS reporting still missing 3 exposures but nothing showing up. once system gets going again will try and trace back 34 * 22:30 MEH: processing has been behaving like the parking brake is on.. many red 20TB disks and seen similar behavior before when this happened, but seems more extreme. stsci neb-host up no help. 48 * 22:30 MEH: processing has been behaving like the parking brake is on.. many red 20TB disks and seen similar behavior before when this happened, but seems more extreme. 49 * stsci nodes have been fine with new kernel, neb-host up with note of new kernel. no improvement, but now skycell products will go there to ease space usage on other data nodes 50 {{{ 51 neb-host stsci00 up --note "upgraded kernel to 3.7.6" 52 }}} 35 53 * even cleanup is painfully slow, so maybe an issue with nebulous or raid speed on a large portion of machines -- allowing a few more commented out nebservers (ippc04,06) but leaving ippc07 out due to network issue 36 54 * cleared old ippc17.0 stalled mounts 37 55 * stsci has dqstats process that uses datastore, had problem with missing dir so set to off and emailed Bill 38 56 * removing ThreePi.WS.nightlyscience label to focus on primary nightly science 57 * raidstatus? 58 {{{ 59 cat ~ipp/raidstatus/* 60 --> ip057 raid pratially degraded but someone put into repair @0923 as it is rebuilt 61 }}} 62 39 63 === Wednesday : 2013.04.17 === 40 64 Bill is czar today. Oh joy.
