Changes between Version 30 and Version 31 of PS1_IPP_Czarlog_20170116
- Timestamp:
- Jan 22, 2017, 10:13:19 PM (10 years ago)
Legend:
- Unmodified
- Added
- Removed
- Modified
-
PS1_IPP_Czarlog_20170116
v30 v31 30 30 * MEH: IfA-wide AC failure -- powering down ipp003,001,022 as less critical to ipp002,ippops1,ippops2 31 31 * last gpc1 dump was successful ~0000 last night before ippdb machines went down last night 32 * AC up but leaving machines off until morning and winds lighten in case AC crashes again 32 33 * MEH: was QUB night, will restart pantasks now that things appear to be back online 34 * ipp032 also not booting -- normally neb-host repair, but must set neb-host down 33 35 * ipp121 not booting -- must be set neb-host down until fixed -- hopefully wont bork summitcopy since power glitch interrupted download o7775g0046f-o7775g0049f but only flats 34 * o7775g0046f having regular reg fault -- 35 * ipp032 also not booting -- normally neb-host repair, but must set neb-host down36 * o7775g0046f having regular reg fault -- messed up redownload and ended up with duplicate newExp/Imfiles -- 1190291 original (dropped from gpc1), 1190457 replaced 37 * o7775g0047f handful of files in missing (ip121) and/or corrupted state but sufficient to replicate and repair 36 38 * ippMonitor crashed when systems down, restart on ippc33 37 39 * not sure if nebdiskd is running still or where it runs now (ippdb08 seems to be from latest czarlog entry?) 38 40 * EAM: restarted nebdiskd on ippdb01 (location of nebulous mysql master server) 39 * MEH: Gene starting data shuffle on stare04 started killing ipp088,104 (date nodes critical for priority QUB processing...) -- 140 remote_md5sum.pl jobs and rising -- so set to neb-host repair 41 * MEH: Gene starting data shuffle on stare04 started killing ipp088,104 (date nodes critical for priority QUB processing...) -- 140 remote_md5sum.pl jobs and rising -- so set to neb-host repair -- setting ~ipptest/replication stop -- cannot have datanode minefield for nightly processing... 40 42 * MEH: ipp118, ipp120 appears to be having XFS issues -- neb-host repair until can look into... -- regular faults then started clearing 41 43 * looks like some files on ipp120 is blocking registration.. -- corrupted gpc1/20170123/o7776g0076o/o7776g0076o.ota55.burn.tbl, manually regenerated with fixburntool
