| Version 19 (modified by , 12 years ago) ( diff ) |
|---|
PS1 IPP Czar Logs for the week 2015.01.05 - 2015.01.11
(Up to PS1 IPP Czar Logs)
Monday : 2015.01.05
- 09:15 EAM : stdlocal is sluggish, restarting it now.
- 10:20 EAM : ipp071 seems to be responding to logins again, so i've put it in neb repair (not down)
- 19:20 MEH: w/ ipp071 back available, the half faulted deepstacks can hopefully finish -- ~ippmd/deepstack/ptolemy.rc -- found possible problem, deepstack stop now
Tuesday : 2015.01.06
- 20:20 MEH: try md stacks again -- ~ippmd/deepstack/ptolemy.rc
- 20:50 EAM: summit is getting heavy winds so I'm leaving stdlocal at high load levels. I'll check again after the KP3 telecon (11pm).
- 21:15 EAM: stdlocal was at 120k warps so I'm restarting it.
Wednesday : 2015.01.07
- 10:44 HAF unjammed registration
regtool -updateprocessedimfile -exp_id 848923 -class_id XY55 -set_state pending_burntool -dbname gpc1 regtool -updateprocessedimfile -exp_id 848925 -class_id XY55 -set_state pending_burntool -dbname gpc1
- restarted registration (11:12, haf)
Thursday : 2015.01.08
- 03:45 Bill stdscience not making progress. Jobs in pcontrol queue without associated processes on the target host. No ppImage processes in 'who uses cluster page' Shutting down stdscience for a restart
- unfortunately the who uses cluster page is not working with all of the hosts. There are some ppImage processes still running See for example ippc26
- ganglia shows ipp059 is down (should have loooked there first. Unfortunately the console app won't accept my password when I try to connect with it
- I am hesitant to start up a new stdscience pantasks with things in this state
- ran neb-host ipp059 down
- 05:30 MEH: power cycling ipp059 but then need to go to aas booth and won't be able to check on until later
- only info on console was
<Jan/07 10:48 pm>ipp059 login: [381619.298227] Disabling IRQ #78
- only info on console was
- 06:20 EAM: ipp059 is back up and happy, i've restarted stdscience.
- 07:45 EAM: it looks like cab1 had a hiccup last night. all machines in that cab, ex ipp009 and ipp019, rebooted about 9pm. I stopped all processing to clear out some mount point troubles. then it turned out that ipp018 was having xfs problems on /export/ipp018.0. I tried to repair, but needed to reboot to umount. on boot, /dev/sda4 refused to mount. I ran xfs_check /dev/sda4 and it said the metadata needed to be replayed by mounting. mounting gave an error that "Structure needs cleaning". I ran xfs_repair /dev/sda4 as advised, and it said that the log was unrecoverable. I am now running xfs_repair -L /dev/sda4 as advised, and it is trying to clean up the disk partition. The output is going to /root/sda4.repair.log. I have since restarted the processing with ipp018 in neb 'down'. there will probably be some errors until ipp018 is up if anything put output there (was it in neb up earlier? not sure).
- 09:10 EAM: xfs_repair finally finished and /export/ipp018.0 mounted successfully. nfs exports it just fine now.
Friday : 2015.01.09
Saturday : 2015.01.10
Sunday : 2015.01.11
Note:
See TracWiki
for help on using the wiki.
