== PS1 IPP Czar Logs for the week YYYY.MM.DD - YYYY.MM.DD == [[PageOutline]] (Up to [wiki:PS1_IPP_CzarLogs PS1 IPP Czar Logs]) === Monday : YYYY.MM.DD === === Tuesday : YYYY.MM.DD === === Wednesday : YYYY.MM.DD === === Thursday : 2017.09.14 === * MEH: ippc65 down since early AM, Hadyn power cycled * MEH: ipp@ippc19 crontab adding /export/ippdbXX check added now since ippdb08 filled up surprisingly today * MEH: ippdb06,db08 ganglia plots missing since weekend.. -- critical machines, useful to have plots.. -- /etc/init.d/gmond restart * MEH: sending missed and misc nightly processing to cleanup again * 15:30 EAM : ran out of space on ippdb08 /export/ippdb08.0 due to a problem with the binlogs. I cleared out some old files first (a test dump and an old, partial copy of nebulous mysql). I then discovered that the binlogs had been getting very large for a while: this was caused by an error introduced into gzgetexp on July 15 which resulting it in attempting to re-download every failed exposure, instead of only the past night. I fixed the code in pzgetexp to only check for the past 100 exposures. I also cleared the binlogs up to Sept 12. * 16:20 EAM : updated ippMonitor to report replication status for ippdb01 for nebulous (instead of ipp115) * MEH: ippc61 appears to have been crashed since weekend? cannot ssh into, no console login (or messages) -- power cycle twice without booting -- email Haydn to look at if time -- Haydn thinks its already on his list of ones to work on and hasn't been up for a while and maybe IP address mixup -- conflicts w/ what ganglia was reporting and that Weryk has been actively using it as well * MEH: nightly pantasks stop.block.restart seems to get pstamp into weird/broken state if trying to make stamps/update with faults -- happened a couple times now -- manually clearing stalled MOPS jobs... === Friday : 2017.09.15 === * MEH: ippc17 missing from ganglia -- /etc/ganglia/gmond.conf was pointing to ippc18, fixed to be ippc19, restarted and now reporting * ippx107 wasn't reporting either -- gmond needed restart === Saturday : 2017.09.16 === * MEH: pstamp backup of QUB updates -- appears ipps00-s07 having issues stalling updates for chip+warp+diff -- taking out of ippqub:stdscience * IIRC ipps00-s07 are in same rack -- network connection/path? === Sunday : YYYY.MM.DD ===