Changes between Version 23 and Version 24 of PS1_IPP_Czarlog_20130401
- Timestamp:
- Apr 2, 2013, 11:47:43 AM (13 years ago)
Legend:
- Unmodified
- Added
- Removed
- Modified
-
PS1_IPP_Czarlog_20130401
v23 v24 11 11 * 19:00 MEH: MD09 deepstacks done so compute3 back to stdsci and stack until start MD08 refstack ~tomorrow 12 12 * 22:20 MEH: appear to have lost connection from ifa to production cluster? ganglia also reporting all systems down.. 13 === Tuesday : YYYY.MM.DD===13 === Tuesday : 2013.04.02 === 14 14 mark is czar 15 15 * 07:00 MEH: network issue with summit, but also something odd with production cluster and looks like some systems are down and some were rebooted. … … 20 20 * 08:30 MEH: restarted screens and czarpoll on ippc11, will wait for roboczar until systems back running 21 21 * 08:40 MEH: rebooted ipp036 to fix mount issues, didn't come back up cleanly after whatever happened last night. need to scan mounts on all machines 22 * serge put ipp036 into neb-host repair -- need to take back out23 * ipp028,033,036 disks aren't always mounting22 * serge put ipp036 into neb-host repair -- need to put back up - ok 23 * ipp028,033,036,060 disks aren't always mounting -- Gene noted ypbind gone rogue, need to kill and restart with nfs (sudo tcsh, su - otherwise once kill ypbind no sudo access of course) 24 24 * some wave1 machines trouble mounting wave4 25 25 * 09:30 MEH: ganglia is showing several machines high 1-m load but nothing running, gmond probably needs to be restarted on most/all machines -- done, and login seems fine except for ipp041 date below 26 26 * 09:40: MEH: ippc17 and ippc19 down, ippRequestServer (pss, datastore db and replication) so not possible for PSS to run until fixed 27 27 * 10:00 MEH: ipp041 date is wrong -- 10:25 manually set to w/in 1s. how to fix time sync? 28 * 11:45 MEH: ipp048 and ippc19 still problems, ipp048 to down. 28 29 === Wednesday : YYYY.MM.DD === 29 30
