IPP Software Navigation Tools IPP Links Communication Pan-STARRS Links
wiki:PS1_IPP_Czarlog_20180827

PS1 IPP Czar Logs for the week 2018.08.27 - 2018.09.02

(Up to PS1 IPP Czar Logs)

Monday : 2018.09.02

  • 10:15 EAM : ippb04 and ippb05 are up allowing us to have a b-node replication target. I've set them up for now and am re-running the MOPS dailytestset chunk quad, which now seems to be succeeding.
  • MEH: summitcopy & registration getting allocation of 1x s4 (ipp067-70) to substitute for down compute nodes and test data node use there
  • MEH: updating stdscience, pstamp, ippqub:stdscience_ws in master ~ippitc/ippconfig/pantasks_hosts.input now from their input files -- no allocation changes other than removing ipps from pstamp now
  • MEH: ippb06-b15 to repair to allow chips to be updated for stamps -- due to there being replicated products having copies only on the backup nodes still...
  • MEH: more test reprocessing for MOPS without masks with label/data_group mehtest.nomask
  • MEH: moving pstamp pantasks from ippc67 to ippc33 -- testing use of lower ippc nodes and network with pantasks (try to trigger network issue seen with lower ippc and ipp123-126)

Tuesday : 2018.08.29

  • MEH: more broken/missing LAP/2014 files flushed out with pstamp test -- also noticing some nfs/cpu wait issues for ipp124 to new ipp127-ipp129 nodes in normal stdscience updates, may be due to over requests of older data on those machines
    • no apparent ssh interruptions/faults however (extra loadings of ipp123-126 in pstamp)
  • MEH: Weryk starting nice'd processing again -- ipps nodes to begin with then move on to ippx and ippc30-63 as before
    • pstamp on ippc33 with Weryk nice'd processing seems fine

Wednesday : 2018.08.30

  • MEH: ipp113 root disk nearly out of space and will bork all kinds of things... -- both apache and messages very very large
    • manually do log rotation after stopping apache
    • no nagios warning -- need to verify if that is setup for this node -- we need ping/alive, root+export disk space check on all nodes basically, maybe also load? -- was only for nodes <ipp080, Gavin setup for all nodes now including a load >50 for 5 min warning
  • MEH: ippx001-x020 root disk also nearly or actually out of space... -- do manual log rotation there like ippx021-x088 a few weeks ago
  • MEH: ippb07-15 neb-host repair->up now since seem stable after shutdown, ippb16,b17,b19,b20,b23 to neb-host repair
  • MEH: Weryk adding ippc80-c127 to nice'd processing since PS1 down
    • ippc102,ippc127 /export/ippc1xx.x permissions drwxr-xr-x and needs to be drwxrwxrwt -- fixed

Thursday : YYYY.MM.DD

Friday : 2018.08.31

  • MEH: ipp032 with Weryk processing providing the nagios addition for load >50/5 min a good test -- all other nodes except a few of the older ipp054-066 data nodes also triggering
  • MEH: ippc88 back up for a bit now, put back into nightly processing group
  • MEH: ippb18 down again, console shows root disk check problem -- unknown if crash/rebooted itself and stalled or if someone powercycled -- leaving in halted state from root disk error until can be looked at

Saturday : YYYY.MM.DD

Sunday : YYYY.MM.DD

Last modified 8 years ago Last modified on Sep 1, 2018, 6:18:52 AM
Note: See TracWiki for help on using the wiki.