| Version 19 (modified by , 12 years ago) ( diff ) |
|---|
PS1 IPP Czar Logs for the week YYYY.MM.DD - YYYY.MM.DD
(Up to PS1 IPP Czar Logs)
Monday : 2014.08.11
- HAF czar : restarted summitcopy,registration,distribution,stdscience,stack (cluster was off this weekend due to hurricane)
- HAF czar : ippb00, 01, 03, 04 are set to repair, 02 is set to up
- HAF czar : restarted czartool/roboczar on c18
Tuesday : 2014.08.12
- 07:30 MEH: LAP stacks needed to be reverted so some runs now clearing, one had stack label LAP.PV3.20140730.local
- 08:30 MEH: probably need to do a regular restart of the LAP pantasks as well -- will do tonight if have to reallocate for normal nightly
- 10:30 MEH: restarted pstamp -- pstamp must run from ipptest for now -- often running out of space, need to change PSTAMP_PRESERVE_DAYS 14->7d or so or doc for doing manually?
- suspect stamp cleanup not happening because cleanup pantasks not running? starting now -- or does this just cleanup the updated products and not the stamp bundles, need to cleanup stack bundles for space.. and some more doc on this
- to try and clean up more space for pstamp, as ipptest ran: pstamp_queue_cleanup.pl --preserve-days=12
- clearing >200G now..
- 11:00 MEH: LAP pole stdsci poll was set too low for number of nodes available if run out of chips etc. bumped up
- 12:00 MEH: don't see nebdiskd on ippdb00, needs to be started -- ipp@ippdb00 started nebdiskd
- 13:10 MEH: unable to log into ipp071.. suspect NIS/ypbind issue again like on 8/4 -- only noticed since nebdiskd couldn't access, has it been broken since reboot the other day or did it recently die? -- Gavin found problem with 10G fiber delay in network startup causing ypbind to timeout. will look into a fix
- 14:10 MEH: restarted LAP stdlocal pantasks and reloaded nodes from before -- too much, if any pstamp updates happen then will overload datanodes..
- 15:10 MEH: pstamp having trouble -- seems a reqType unknown is in system and causing problems? -- leaving off until solved because just filling log file
failure for: request_finish.pl --req_id 405851 --req_type unknown --req_file /data/ippc30.1/pstamp/work/webreq/2014/08/12/web_162610.fits --req_name NULL --product NULL --outdir NULL --redirect-output --dbname ippRequestServer --verbose job exit status: 29 job host: ippc38 job dtime: 0.604044 job exit date: Tue Aug 12 15:25:27 2014
- 15:30 MEH: ifaps1 is still down from the weekend storms? -- is it needed to be up?
Wednesday : YYYY.MM.DD
Thursday : YYYY.MM.DD
Friday : YYYY.MM.DD
Saturday : YYYY.MM.DD
Sunday : YYYY.MM.DD
Note:
See TracWiki
for help on using the wiki.
