| Version 28 (modified by , 12 years ago) ( diff ) |
|---|
PS1 IPP Czar Logs for the week YYYY.MM.DD - YYYY.MM.DD
(Up to PS1 IPP Czar Logs)
Monday : 2014.08.11
- HAF czar : restarted summitcopy,registration,distribution,stdscience,stack (cluster was off this weekend due to hurricane)
- HAF czar : ippb00, 01, 03, 04 are set to repair, 02 is set to up
- HAF czar : restarted czartool/roboczar on c18
Tuesday : 2014.08.12
- 07:30 MEH: LAP stacks needed to be reverted so some runs now clearing, one had stack label LAP.PV3.20140730.local
- 08:30 MEH: probably need to do a regular restart of the LAP pantasks as well -- will do tonight if have to reallocate for normal nightly
- 10:30 MEH: restarted pstamp -- pstamp must run from ipptest for now -- often running out of space, need to change PSTAMP_PRESERVE_DAYS 14->7d or so or doc for doing manually?
- suspect stamp cleanup not happening because cleanup pantasks not running? starting now -- or does this just cleanup the updated products and not the stamp bundles, need to cleanup stack bundles for space.. and some more doc on this
- to try and clean up more space for pstamp, as ipptest ran: pstamp_queue_cleanup.pl --preserve-days=12
- clearing >200G now..
- 11:00 MEH: LAP pole stdsci poll was set too low for number of nodes available if run out of chips etc. bumped up
- 12:00 MEH: don't see nebdiskd on ippdb00, needs to be started -- ipp@ippdb00 started nebdiskd
- 13:10 MEH: unable to log into ipp071.. suspect NIS/ypbind issue again like on 8/4 -- only noticed since nebdiskd couldn't access, has it been broken since reboot the other day or did it recently die? -- Gavin found problem with 10G fiber delay in network startup causing ypbind to timeout. will look into a fix
- 14:10 MEH: restarted LAP stdlocal pantasks and reloaded nodes from before -- too much, if any pstamp updates happen then will overload datanodes..
- 15:10 MEH: pstamp having trouble -- seems a reqType unknown is in system and causing problems? -- leaving off until solved because just filling log file
failure for: request_finish.pl --req_id 405851 --req_type unknown --req_file /data/ippc30.1/pstamp/work/webreq/2014/08/12/web_162610.fits --req_name NULL --product NULL --outdir NULL --redirect-output --dbname ippRequestServer --verbose job exit status: 29 job host: ippc38 job dtime: 0.604044 job exit date: Tue Aug 12 15:25:27 2014
- is creating ~ipptest/NULL/reqfinish.405851.log file, seems to be unknown DB? but if ipptest was running before, should be okay?
request 405851 has unknown reqType unknown Running [/home/panstarrs/ipptest/psconfig/ipp-pv3-20140717.lin64/bin/pstamptool -updatereq -req_id 405851 -state stop -fault 5 -dbname ippRequestServer]... Unable to perform /home/panstarrs/ipptest/psconfig/ipp-pv3-20140717.lin64/bin/pstamptool -updatereq -req_id 405851 -state stop -fault 5 -dbname ippRequestServer error code: 768 at /home/panstarrs/ipptest/psconfig/ipp-pv3-20140717.lin64/bin/request_finish.pl line 101. request 405851 has unknown reqType unknown Running [/home/panstarrs/ipptest/psconfig/ipp-pv3-20140717.lin64/bin/pstamptool -updatereq -req_id 405851 -state stop -fault 5 -dbname ippRequestServer]... -> psDBAlloc (psDB.c:166): Database error generated by the server Failed to connect to database. Error: Unknown database 'ippRequestServer' -> pstamptoolConfig (pstamptoolConfig.c:353): unknown psLib error Can't configure database -> main (pstamptool.c:80): (null) failed to configure Unable to perform /home/panstarrs/ipptest/psconfig/ipp-pv3-20140717.lin64/bin/pstamptool -updatereq -req_id 405851 -state stop -fault 5 -dbname ippRequestServer error code: 768 at /home/panstarrs/ipptest/psconfig/ipp-pv3-20140717.lin64/bin/request_finish.pl line 101. - looks like the cmd has left off the -dbserver ippc17
- not only that, but if the cmd did work it would fault because there is no -set_XXXX. something is wrong with this error case then and -state,-fault need to be -set_state,-set_fault?
- 16:35 MEH: so fix the stop/fault cmd and run for the broken ones -- seems to be moving again (problem looks to have happened ~4-6am)
pstamptool -dbname ippRequestServer -dbserver ippc17 -updatereq -req_id 405851 -set_state stop -set_fault 5 pstamptool -dbname ippRequestServer -dbserver ippc17 -updatereq -req_id 405852 -set_state stop -set_fault 5 pstamptool -dbname ippRequestServer -dbserver ippc17 -updatereq -req_id 405853 -set_state stop -set_fault 5 pstamptool -dbname ippRequestServer -dbserver ippc17 -updatereq -req_id 405854 -set_state stop -set_fault 5 pstamptool -dbname ippRequestServer -dbserver ippc17 -updatereq -req_id 405855 -set_state stop -set_fault 5 pstamptool -dbname ippRequestServer -dbserver ippc17 -updatereq -req_id 405856 -set_state stop -set_fault 5 pstamptool -dbname ippRequestServer -dbserver ippc17 -updatereq -req_id 405857 -set_state stop -set_fault 5 pstamptool -dbname ippRequestServer -dbserver ippc17 -updatereq -req_id 405858 -set_state stop -set_fault 5
- does seem to suggest a code bug but unclear the process logistics here so..
- is creating ~ipptest/NULL/reqfinish.405851.log file, seems to be unknown DB? but if ipptest was running before, should be okay?
- 15:30 MEH: ifaps1 is still down from the weekend storms? -- is it needed to be up?
- 16:00 HAF: I haven't talked to Gavin about it, but I'd like it to be up..
- 17:20 MEH: so other PSS jobs have finished leaving a handful in various states, two in fault 4 and four in state new --
Wednesday : 2014.08.13
- 05:54 Bill: reverted registration fault for o6882g0284o XY04. It ran into a problem connecting to the database on first try. Science exposures are all registered now.
- 07:00 MEH: some remaining warps also DB access faulted
Thursday : YYYY.MM.DD
Friday : YYYY.MM.DD
Saturday : YYYY.MM.DD
Sunday : YYYY.MM.DD
Note:
See TracWiki
for help on using the wiki.
