PS1 IPP Czar Logs for the week 2018.08.27 - 2018.09.02
(Up to PS1 IPP Czar Logs)
Monday : 2018.09.02
- 10:15 EAM : ippb04 and ippb05 are up allowing us to have a b-node replication target. I've set them up for now and am re-running the MOPS dailytestset
chunkquad, which now seems to be succeeding. - MEH: summitcopy & registration getting allocation of 1x s4 (ipp067-70) to substitute for down compute nodes and test data node use there
- MEH: updating stdscience, pstamp, ippqub:stdscience_ws in master ~ippitc/ippconfig/pantasks_hosts.input now from their input files -- no allocation changes other than removing ipps from pstamp now
- MEH: ippb06-b15 to repair to allow chips to be updated for stamps -- due to there being replicated products having copies only on the backup nodes still...
- MEH: more test reprocessing for MOPS without masks with label/data_group mehtest.nomask
- MEH: moving pstamp pantasks from ippc67 to ippc33 -- testing use of lower ippc nodes and network with pantasks (try to trigger network issue seen with lower ippc and ipp123-126)
Tuesday : 2018.08.29
- MEH: more broken/missing LAP/2014 files flushed out with pstamp test -- also noticing some nfs/cpu wait issues for ipp124 to new ipp127-ipp129 nodes in normal stdscience updates, may be due to over requests of older data on those machines
- no apparent ssh interruptions/faults however (extra loadings of ipp123-126 in pstamp)
- MEH: Weryk starting nice'd processing again -- ipps nodes to begin with then move on to ippx and ippc30-63 as before
- pstamp on ippc33 with Weryk nice'd processing seems fine
Wednesday : 2018.08.30
- MEH: ipp113 root disk nearly out of space and will bork all kinds of things... -- both apache and messages very very large
- manually do log rotation after stopping apache
- no nagios warning -- need to verify if that is setup for this node -- we need ping/alive, root+export disk space check on all nodes basically, maybe also load? -- was only for nodes <ipp080, Gavin setup for all nodes now including a load >50 for 5 min warning
- MEH: ippx001-x020 root disk also nearly or actually out of space... -- do manual log rotation there like ippx021-x088 a few weeks ago
- MEH: ippb07-15 neb-host repair->up now since seem stable after shutdown, ippb16,b17,b19,b20,b23 to neb-host repair
- MEH: Weryk adding ippc80-c127 to nice'd processing since PS1 down
- ippc102,ippc127 /export/ippc1xx.x permissions drwxr-xr-x and needs to be drwxrwxrwt -- fixed
Thursday : YYYY.MM.DD
Friday : 2018.08.31
- MEH: ipp032 with Weryk processing providing the nagios addition for load >50/5 min a good test -- all other nodes except a few of the older ipp054-066 data nodes also triggering
- MEH: ippc88 back up for a bit now, put back into nightly processing group
- MEH: ippb18 down again, console shows root disk check problem -- unknown if crash/rebooted itself and stalled or if someone powercycled -- leaving in halted state from root disk error until can be looked at
Saturday : YYYY.MM.DD
Sunday : YYYY.MM.DD
Last modified
8 years ago
Last modified on Sep 1, 2018, 6:18:52 AM
Note:
See TracWiki
for help on using the wiki.
