== PS1 IPP Czar Logs for the week 2013.04.01 - 2013.04.07 == [[PageOutline]] (Up to [wiki:PS1_IPP_CzarLogs PS1 IPP Czar Logs]) === Monday : 2013.04.01 === mark is czar * 08:05 MEH: remaining few nightly science exposures finishing up * 09:38 Bill: ippc03 is running out of space in /tmp. The nebulous_server.log file had grown to 22G. Stopped apache, moved the file out of the way, restarted apache, then deleted it. * 11:10 MEH: ipp05,08,09 nebulous_sever.log >22G as well, leaving ~4G on / OS disk and should be cycled out as well -- done. ippc07 will be next, only ~9G left available * 19:00 MEH: MD09 deepstacks done so compute3 back to stdsci and stack until start MD08 refstack ~tomorrow * 22:20 MEH: appear to have lost connection from ifa to production cluster? ganglia also reporting all systems down.. === Tuesday : 2013.04.02 === mark is czar * 07:00 MEH: network issue with summit, but also something odd with production cluster and looks like some systems are down and some were rebooted. * still down: 028,033,036,046,048,c17,c19,db02 and console not reachable -- was ganglia problem for some, ipp046,048,c17,c19 actually down/off network, cannot connect via console * checking if the stsci nodes were power cycled -- no, so probably good thing * ippdb00,01,03 have been so may need to check mysql * various data (non-stsci data) nodes have been rebooted it seems * 08:30 MEH: restarted screens and czarpoll on ippc11, will wait for roboczar until systems back running * 08:40 MEH: rebooted ipp036 to fix mount issues, didn't come back up cleanly after whatever happened last night. need to scan mounts on all machines * serge put ipp036 into neb-host repair -- need to put back up - ok * ipp028,033,036,060 disks aren't always mounting -- Gene noted ypbind gone rogue, need to kill and restart with nfs (sudo tcsh, su - otherwise once kill ypbind no sudo access of course) * some wave1 machines trouble mounting wave4 * 09:30 MEH: ganglia is showing several machines high 1-m load but nothing running, gmond probably needs to be restarted on most/all machines -- done, and login seems fine except for ipp041 date below * 09:40: MEH: ippc17 and ippc19 down, ippRequestServer (pss, datastore db and replication) so not possible for PSS to run until fixed * 10:00 MEH: ipp041 date is wrong -- 10:25 manually set to w/in 1s. how to fix time sync? * 10:20 Bill added PSS MOTD, to remove -- cp /data/ippc17.0/pstamp/web/pstamp_motd.html.empty /data/ippc17.0/pstamp/web/pstamp_motd.html * 11:45 MEH: ipp048 and ippc19 still problems, ipp048 to down. -- Haydn and Rita having trouble logging into ipp048 console, I can log in and disk scan taking the time for reboot -- back up now, neb-host up. * ''ippc19 still being investigated -- burning smell, out for the time being'' * 12:05 MEH: datanodes back online, mounts appear fine. slowly starting pantasks -- summitcopy, registration, stdscience, stack * 12:20 MEH: nebdiskd needs to be started on ippdb00 * 12:35 Serge fixed crashed czardb table processed -- repair table processed. loaded_to_datastore also crashed and repaired. eventually the czarpoll page will be uptodate. loaded_to_ODM also crash and repaired. * 12:50 MEH: distribution, publishing, detrend started. pstamp and update will wait until decide on ippc17 replication * 13:40 MEH: finding other machines with close but wrong time by hours.. more importantly ippdb00,01 are off.. so all processing to stop * 15:00 MEH: system times fixed and /etc/init.d/ntpd restart, ippdb00,01 mysql restarted, processing restarted * 15:05 MEH: wave1 need reboot with testing 3.7.6 kernel or may have disk issues with all up in nebulous -- all into repair and removed from processing, slowly rebooting with 3.7.6 kernel * kernel option may not be a visible option depending on the console state and system logged into, should show up when arrow key up * ipp017 different mobo so keep original kernel and in neb-host repair * ipp011-016 rebooted with new kernel and neb-host up * ipp018-021 rebooted with new kernel and set neb-host up (ipp021 originally okay with original kernel and neb-host up) * 15:15 Bill noticed pstamp confused with ports from detrend. fixed and pstamp running again * 15:18 Bill removed the "pstamp is down" message from postage stamp web pages by emptying the pstamp_motd file. * 17:55 MEH: ippc01-c05, stare, stsci ganglia not reporting right, sudo /etc/init.d/gmond restart * 19:30 MEH: done downloading last nightly data, but stuck and not registering them all. downloading and registering tonights data okay. * 20:30 MEH: cleared offending imfiles and rawExp registration from last night proceeding * 22:20 MEH: normally caught up with last night's and tonight's exposure registration, processing load seems more irregular than normal with more stage faults. * 23:40 MEH: restart roboczar now on ippc11. cleanup stays off (not in roboczar) and replication off as normally runs on ippc19 but not active right now anyways. === Wednesday : 2013.04.03 === Bill is czar today * 08:05 Nightly science has about 40 warps left to process and lots of diffs. Setting chip.off for now. * 09:51 Just 50 diffs to do setting warp to off for a few minutes * 10:21 Letting nightly science diffs finish then will restart stdscience. * 10:36 stdscience restarted * MEH: compute3 out of autoload to stdscience and stack still -- deepstack tests and refstacks soon -- adding manually back into stdscience until i reactivate deepstack after the tuesday mess * 11:00 MEH: appears MD05,06,07 from 4/02 still missing from stacks and diffims, re-adding that date to stdscience needed pick up? will need to tweak_ssdiffs since past default processing window * added and stacks except MD07 picked up @1210. tweak_ssdiff and 4/03 SSdiff running, may need to extend the window depending how long 4/02 stacks take * 13:40 MEH: MD stacks and ssdiff finished, tweak_ssdiff time back to default * 14:45 OTIS reported a problem exposure from Monday night o6384g0201o. One of it's pzDownloadImfiles had fault 110 (HTTP Gone 410 - 300) and the pzDownloadExp had been set to drop. The file is now available so I reverted the imfile and set the state back to run and the exposure copied and registered succesfully. Is now going through chip processing. * 15:00 MEH: md.pv1 notes -- no need to monitor and since low prio, if chip.off etc necessary, then that is not a problem. * MD08.refstack.20130401 chip->warp and prio my move up since we've started observing MD08. deepstack will take compute3 nodes in very near future. * MD08.pv1.20130403 reprocessing necessary of all exposures prior to july 2012 (minimum date of may 15, 2012 so some overlap of "acceptable" processing for comparison pv1 start vs now) * MD09.update2012 almost half the MD09.nightlyscience warps from 2012 were not cleaned up (chips were), moved label MD09.nightlyscience->MD09.update2012 and updating warps that were cleaned up to be available for the MD09 2009-2012 deepstack. * stsci06 has gone down stopping processing * 16:09 set stsci06 to down and set pantasks to run' * 16:57 many stuck jobs waiting for stsci06. Stopping processing and will attempt to kill outstanding processes * 17:18 stdscience, distribution, update, and pstamp pantasks restarted. * 17:20 camera and warp revert.off jobs for now. The faulted jobs need data from stsci06 that is not available. * 17:40 Many Many nebulous errors creating new files. restarted the apache nebulous servers that had high load. This seems to have improved things * 21:25 MEH: with stsci06 down, removing MD08.refstack.20130401 label since remaining fault. chip/camera/warp revert back on === Thursday : 2013.04.04 === Bill is czar today === Friday : 2013.04.05 === === Saturday : 2013.04.06 === === Sunday : 2013.04.07 ===