Changes between Version 11 and Version 12 of PS1_IPP_Czarlog_20120709
- Timestamp:
- Jul 11, 2012, 5:07:19 PM (14 years ago)
Legend:
- Unmodified
- Added
- Removed
- Modified
-
PS1_IPP_Czarlog_20120709
v11 v12 9 9 * 11:07 CZW: Stopped stdscience to prepare for a restart. 10 10 * 11:25 CZW: Running again. 11 === Tuesday : YYYY.MM.DD === 11 12 === Tuesday : 2012.07.10 === 12 13 * 11:00 Serge: Slave on ipp001 is back 13 14 * 11:40 Mark: restarted stdscience, only 40-100 jobs running. back up to >450 jobs running. doing a large volume of MD07.mehtest updates for a few hours 14 15 * 16:18 CZW: I've launched a second "stdscience" like pantasks in ~ipp/ecliptic. I'm using this to process the w-band warps for the ecliptic plane to allow for warp-stack diffs for MOPS. It is using compute3 and stsci nodes, but I may have oversubscribed these nodes somewhat. The data is processing under the label "ecliptic.rp". 15 === Wednesday : YYYY.MM.DD === 16 * 15:58 Serge: set stsci06 to repair in nebulous 16 17 === Wednesday : 2012.07.11 === 18 * 15:58 Serge: set stsci06 to repair in nebulous (EAM: it never actually had crashed, it was failing to respond to ganglia because of nfs hangups) 19 * 17:05 EAM : after serious hang-ups, it looks like the cluster is now unwedged. it seems that the nfs servers on some of the stsci nodes were failing to respond to other stsci nodes and vice versa. I killed off active jobs as much as possible on those machines and then forced the machines to umount the hung partitions. looks like this eventually worked. it is not clear what caused the initial problem, though. 20 17 21 === Thursday : YYYY.MM.DD === 18 22
