== PS1 IPP Czar Logs for the week 2018.08.27 - 2018.09.02 == [[PageOutline]] (Up to [wiki:PS1_IPP_CzarLogs PS1 IPP Czar Logs]) === Monday : 2018.09.02 === * 10:15 EAM : ippb04 and ippb05 are up allowing us to have a b-node replication target. I've set them up for now and am re-running the MOPS dailytestset ~~chunk~~ quad, which now seems to be succeeding. * MEH: summitcopy & registration getting allocation of 1x s4 (ipp067-70) to substitute for down compute nodes and test data node use there * MEH: updating stdscience, pstamp, ippqub:stdscience_ws in master ~ippitc/ippconfig/pantasks_hosts.input now from their input files -- no allocation changes other than removing ipps from pstamp now * MEH: ippb06-b15 to repair to allow chips to be updated for stamps -- due to there being replicated products having copies only on the backup nodes still... * MEH: more test reprocessing for MOPS without masks with label/data_group mehtest.nomask * MEH: moving pstamp pantasks from ippc67 to ippc33 -- testing use of lower ippc nodes and network with pantasks (try to trigger network issue seen with lower ippc and ipp123-126) === Tuesday : 2018.08.29 === * MEH: more broken/missing LAP/2014 files flushed out with pstamp test -- also noticing some nfs/cpu wait issues for ipp124 to new ipp127-ipp129 nodes in normal stdscience updates, may be due to over requests of older data on those machines * no apparent ssh interruptions/faults however (extra loadings of ipp123-126 in pstamp) * MEH: Weryk starting nice'd processing again -- ipps nodes to begin with then move on to ippx and ippc30-63 as before * pstamp on ippc33 with Weryk nice'd processing seems fine === Wednesday : 2018.08.30 === * MEH: ipp113 root disk nearly out of space and will bork all kinds of things... -- both apache and messages very very large * manually do log rotation after stopping apache * no nagios warning -- need to verify if that is setup for this node -- we need ping/alive, root+export disk space check on all nodes basically, maybe also load? -- was only for nodes 50 for 5 min warning * MEH: ippx001-x020 root disk also nearly or actually out of space... -- do manual log rotation there like ippx021-x088 a few weeks ago * MEH: ippb07-15 neb-host repair->up now since seem stable after shutdown, ippb16,b17,b19,b20,b23 to neb-host repair * MEH: Weryk adding ippc80-c127 to nice'd processing since PS1 down * ippc102,ippc127 /export/ippc1xx.x permissions drwxr-xr-x and needs to be drwxrwxrwt -- fixed === Thursday : YYYY.MM.DD === === Friday : 2018.08.31 === * MEH: ipp032 with Weryk processing providing the nagios addition for load >50/5 min a good test -- all other nodes except a few of the older ipp054-066 data nodes also triggering * MEH: ippc88 back up for a bit now, put back into nightly processing group * MEH: ippb18 down again, console shows root disk check problem -- unknown if crash/rebooted itself and stalled or if someone powercycled -- leaving in halted state from root disk error until can be looked at === Saturday : YYYY.MM.DD === === Sunday : YYYY.MM.DD ===