| Version 12 (modified by , 14 years ago) ( diff ) |
|---|
(Up to PS1 IPP Czar Logs)
Monday : 2012.07.09
- 09:00 Serge: gpc1 replication broken on ipp001. Dropped the database and recreating it.
- 10:55 Bill: restarted distribution to cause the dropped stages to be added back in
- 11:02 Bill: updated recipes/psphot.config to remove the limit on the number of peaks (previously 50,000)
- 11:07 CZW: Stopped stdscience to prepare for a restart.
- 11:25 CZW: Running again.
Tuesday : 2012.07.10
- 11:00 Serge: Slave on ipp001 is back
- 11:40 Mark: restarted stdscience, only 40-100 jobs running. back up to >450 jobs running. doing a large volume of MD07.mehtest updates for a few hours
- 16:18 CZW: I've launched a second "stdscience" like pantasks in ~ipp/ecliptic. I'm using this to process the w-band warps for the ecliptic plane to allow for warp-stack diffs for MOPS. It is using compute3 and stsci nodes, but I may have oversubscribed these nodes somewhat. The data is processing under the label "ecliptic.rp".
Wednesday : 2012.07.11
- 15:58 Serge: set stsci06 to repair in nebulous (EAM: it never actually had crashed, it was failing to respond to ganglia because of nfs hangups)
- 17:05 EAM : after serious hang-ups, it looks like the cluster is now unwedged. it seems that the nfs servers on some of the stsci nodes were failing to respond to other stsci nodes and vice versa. I killed off active jobs as much as possible on those machines and then forced the machines to umount the hung partitions. looks like this eventually worked. it is not clear what caused the initial problem, though.
Thursday : YYYY.MM.DD
Friday : YYYY.MM.DD
Saturday : YYYY.MM.DD
Sunday : YYYY.MM.DD
Note:
See TracWiki
for help on using the wiki.
