== PS1 IPP Czar Logs for the week YYYY.MM.DD - YYYY.MM.DD == [[PageOutline]] (Up to [wiki:PS1_IPP_CzarLogs PS1 IPP Czar Logs]) === Monday : 2013.07.29 === mark is czar * 08:30 MEH: doing regular restart of stdsci (turndown in rate) * 08:50 MEH: update appears down, restarting. adding 1x compute3 (taking from stdsci) * might as well restart other main running ones like summitcopy, registration. * 09:15 MEH: and restarting pstamp * 10:10 MEH: in preparation for possible power issues with Flossie, stopping all processing and shutting down compute2+3 nodes (and stare which Gene will specifically take care of for dvo/ipptopsps issues). also ippc30 still PSS pseudo-datanode so leave up as well * 11:15 MEH: compute3 ippc63--ippc31 shutdown started * 11:45 MEH compute2 ippc29--ippc20 shutdown started * 12:15 MEH compute2 ippc10,c14,c15,c16 started * 19:00 Gene powered down all systems === Tuesday : 2013.07.30 === mark is czar * 07:30 MEH: Gavin proceeding with long processes of restarting all machines * 09:50 MEH: Gene has made the shift of the /data/ipp031.0, /data/ipp032.0 locations to stsci03.0, stwci00.0 respectively with lower level symlinks under nebulous across all stsci machines * 10:00 MEH: most machines up, doing secondary check of cpu, memory etc not MIA as we have occasionally found in past (also datetime correct) * ipp061 refound its lost 8G RAM from end of March * ipp060 lost 8G RAM.. doing shutdown and powerdown/up restart -- 8G back * 10:10 Gavin reports {{{ ippc02 - down due to audible alarm (MHPCC are escorting guests and requested we disable it) ippc53 - checking on it, console is blank but I can see outlet is providing 263W. ipp046 - not powering up after power outlet cycle. (suspect BIOS settings not being saved. time for button battery replacement..) }}} * Haydn checking on ippdb04 at ATRC -- up an running now, ganglia thinks it is down still * 10:30 MEH: ippdb00,02 mysql not running so need to start -- see [wiki:Processing] * apache already running on ippc01-09 (taking ippc02 out of list in ~ipp/.tcshrc) * nebdiskd started by ipp@ippdb00 * neb-host ipp046 down * ippc02,c53,046 commented out in ~ippconfig/pantasks_hosts.input * 11:00 MEH: restarted czar screens and czarpoll on ippc11 * 11:10 MEH: still scanning nfs mounts across all system * 11:20 MEH: ipp028,033,036 commonly missing, after a couple nfs restarts and minute they are exporting okay -- ''sudo /etc/init.d/nfs restart'' {{{ -- first restart reported error exportfs: Warning: /export/ipp028.0 does not support NFS export. exportfs: Warning: /export does not support NFS export. }}} * 11:40 MEH: neb-host repair for ipp033, 041 (red and full) so if cleanup frees up space still allows space for dvo work (as with others in that wave) * 12:00 MEH: ipp046 back up, so back to neb-host repair. ippc02 also back up so back into the .tcshrc nebservers. -- both added back into ~ippconfig/pantasks_hosts.input * 12:05 MEH: stdsci pantasks started, will start PSS set next * 12:40 MEH: all pantasks started and running again === Wednesday : YYYY.MM.DD === === Thursday : YYYY.MM.DD === === Friday : YYYY.MM.DD === === Saturday : YYYY.MM.DD === === Sunday : YYYY.MM.DD ===