IPP Software Navigation Tools IPP Links Communication Pan-STARRS Links
wiki:PS1_IPP_Czarlog_20140811

Version 19 (modified by Mark Huber, 12 years ago) ( diff )

--

PS1 IPP Czar Logs for the week YYYY.MM.DD - YYYY.MM.DD

(Up to PS1 IPP Czar Logs)

Monday : 2014.08.11

  • HAF czar : restarted summitcopy,registration,distribution,stdscience,stack (cluster was off this weekend due to hurricane)
  • HAF czar : ippb00, 01, 03, 04 are set to repair, 02 is set to up
  • HAF czar : restarted czartool/roboczar on c18

Tuesday : 2014.08.12

  • 07:30 MEH: LAP stacks needed to be reverted so some runs now clearing, one had stack label LAP.PV3.20140730.local
  • 08:30 MEH: probably need to do a regular restart of the LAP pantasks as well -- will do tonight if have to reallocate for normal nightly
  • 10:30 MEH: restarted pstamp -- pstamp must run from ipptest for now -- often running out of space, need to change PSTAMP_PRESERVE_DAYS 14->7d or so or doc for doing manually?
    • suspect stamp cleanup not happening because cleanup pantasks not running? starting now -- or does this just cleanup the updated products and not the stamp bundles, need to cleanup stack bundles for space.. and some more doc on this
    • to try and clean up more space for pstamp, as ipptest ran: pstamp_queue_cleanup.pl --preserve-days=12
      • clearing >200G now..
  • 11:00 MEH: LAP pole stdsci poll was set too low for number of nodes available if run out of chips etc. bumped up
  • 12:00 MEH: don't see nebdiskd on ippdb00, needs to be started -- ipp@ippdb00 started nebdiskd
  • 13:10 MEH: unable to log into ipp071.. suspect NIS/ypbind issue again like on 8/4 -- only noticed since nebdiskd couldn't access, has it been broken since reboot the other day or did it recently die? -- Gavin found problem with 10G fiber delay in network startup causing ypbind to timeout. will look into a fix
  • 14:10 MEH: restarted LAP stdlocal pantasks and reloaded nodes from before -- too much, if any pstamp updates happen then will overload datanodes..
  • 15:10 MEH: pstamp having trouble -- seems a reqType unknown is in system and causing problems? -- leaving off until solved because just filling log file
    failure for: request_finish.pl --req_id 405851 --req_type unknown --req_file /data/ippc30.1/pstamp/work/webreq/2014/08/12/web_162610.fits --req_name NULL --product NULL --outdir NULL --redirect-output --dbname ippRequestServer --verbose
    job exit status: 29
    job host: ippc38
    job dtime: 0.604044
    job exit date: Tue Aug 12 15:25:27 2014
    
  • 15:30 MEH: ifaps1 is still down from the weekend storms? -- is it needed to be up?

Wednesday : YYYY.MM.DD

Thursday : YYYY.MM.DD

Friday : YYYY.MM.DD

Saturday : YYYY.MM.DD

Sunday : YYYY.MM.DD

Note: See TracWiki for help on using the wiki.