| Version 18 (modified by , 7 years ago) ( diff ) |
|---|
PS1 IPP Czar Logs for the week 2019.04.22 - 2019.04.28
Czar of the week: TdB + EAM (czar-lord), previously CCL
(Up to PS1 IPP Czar Logs)
Monday : 2019.04.22
- CCL: these are full stage on chip, but still show up on red flags
chiptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -chip_id 1182888 chiptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -chip_id 1182889 chiptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -chip_id 1182890 chiptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -chip_id 1182891 chiptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -chip_id 1182892 chiptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -chip_id 1182893 chiptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -chip_id 1182894 chiptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -chip_id 1182895 chiptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -chip_id 1182896 chiptool -dbname gpc1 -setimfiletoupdate -set_label ps_ud_MOPS -chip_id 1182888 chiptool -dbname gpc1 -setimfiletoupdate -set_label ps_ud_MOPS -chip_id 1182889 chiptool -dbname gpc1 -setimfiletoupdate -set_label ps_ud_MOPS -chip_id 1182890 chiptool -dbname gpc1 -setimfiletoupdate -set_label ps_ud_MOPS -chip_id 1182891 chiptool -dbname gpc1 -setimfiletoupdate -set_label ps_ud_MOPS -chip_id 1182892 chiptool -dbname gpc1 -setimfiletoupdate -set_label ps_ud_MOPS -chip_id 1182893 chiptool -dbname gpc1 -setimfiletoupdate -set_label ps_ud_MOPS -chip_id 1182894 chiptool -dbname gpc1 -setimfiletoupdate -set_label ps_ud_MOPS -chip_id 1182895 chiptool -dbname gpc1 -setimfiletoupdate -set_label ps_ud_MOPS -chip_id 1182896 --> can't be updated due to ippb nodes are shutdown.
- CCL: most of them have burntool files on ippb node without copies on ITC nodes.
SELECT rawExp.exp_name,chipProcessedImfile.class_id,rawExp.exp_id,chipRun.chip_id,chipRun.state,chipRun.label,chipRun.data_group,chipRun.dist_group,chipProcessedImfile.fault FROM chipRun, rawExp, chipProcessedImfile WHERE chipRun.exp_id = rawExp.exp_id AND chipProcessedImfile.exp_id = chipRun.exp_id AND chipProcessedImfile.chip_id = chipRun.chip_id AND chipProcessedImfile.fault != 0 AND chipRun.label like 'update.LAP.PV3' group by chip_id; +-------------+----------+--------+---------+--------+----------------+-------------------------------+-------------+-------+ | exp_name | class_id | exp_id | chip_id | state | label | data_group | dist_group | fault | +-------------+----------+--------+---------+--------+----------------+-------------------------------+-------------+-------+ | o5464g0207o | XY75 | 230361 | 1179692 | update | update.LAP.PV3 | LAP.PV3.20140730.20141024 | LAP.ThreePi | 2 | | o5444g0007o | XY76 | 219509 | 1179773 | update | update.LAP.PV3 | LAP.PV3.20140730.20141024 | LAP.ThreePi | 2 | | o5444g0024o | XY75 | 219526 | 1179774 | update | update.LAP.PV3 | LAP.PV3.20140730.20141024 | LAP.ThreePi | 2 | | o5444g0009o | XY75 | 219511 | 1181481 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5444g0026o | XY75 | 219528 | 1181482 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5027g0338o | XY75 | 87852 | 1181498 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5462g0101o | XY26 | 229028 | 1181500 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5343g0321o | XY73 | 173486 | 1181536 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5343g0336o | XY76 | 173502 | 1181539 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5487g0040o | XY76 | 242068 | 1181546 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5487g0052o | XY75 | 242080 | 1181547 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5343g0330o | XY74 | 173494 | 1181573 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5343g0333o | XY73 | 173498 | 1181605 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5343g0372o | XY76 | 173537 | 1181945 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5362g0293o | XY75 | 181552 | 1182008 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5390g0030o | XY75 | 191840 | 1182127 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5390g0067o | XY76 | 191878 | 1182376 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5390g0068o | XY74 | 191879 | 1182678 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5365g0588o | XY73 | 183722 | 1182737 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5365g0590o | XY12 | 183730 | 1182738 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5365g0609o | XY74 | 183764 | 1182742 | update | update.LAP.PV3 | LAP.PV3.20140730.20141025 | LAP.ThreePi | 2 | | o5369g0248o | XY75 | 185907 | 1182788 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | | o5369g0251o | XY27 | 185910 | 1182873 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | | o5371g0256o | XY76 | 186536 | 1183057 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | | o5371g0258o | XY75 | 186538 | 1183093 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | | o5405g0451o | XY75 | 198470 | 1183305 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | | o5405g0470o | XY74 | 198489 | 1183308 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | | o5444g0013o | XY76 | 219515 | 1184003 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | | o5444g0016o | XY76 | 219518 | 1184043 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | | o5343g0353o | XY75 | 173517 | 1184369 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | | o5343g0367o | XY74 | 173532 | 1184371 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | | o5343g0369o | XY74 | 173534 | 1184372 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | | o5404g0151o | XY75 | 197586 | 1184373 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | | o5404g0172o | XY76 | 197607 | 1184374 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | | o6435g0293o | XY30 | 613989 | 1184406 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | | o6525g0133o | XY74 | 646008 | 1184407 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | | o6513g0165o | XY04 | 641174 | 1184566 | update | update.LAP.PV3 | LAP.PV3.20140730.ipp.20141026 | LAP.ThreePi | 2 | | o5343g0351o | XY47 | 173516 | 1184668 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | | o5344g0252o | XY75 | 174158 | 1184675 | update | update.LAP.PV3 | LAP.PV3.20140730.20141026 | LAP.ThreePi | 2 | +-------------+----------+--------+---------+--------+----------------+-------------------------------+-------------+-------+
- MEH: adding QUB.2 to pstamp server for bulk QUB requests (similar to MOPS.2, large/long request not needed ASAP), ps_ud_QUB.2 label and updates done in ippqub:stdscience_ws pantask -- lots of faults, czar not need to worry about at this time and due to the ongoing issue of multiple copies of needed files only on the ippb nodes and none on the actual IPP-ITC cluster...
- TdB: Around 21:00 I noticed that data was no longer downloading from the summit. A quick look at the summitcopy pantask stdout log showed that MySQL was down, which was causing the problem. Gene checked the process list for mysql:
ps aux | grep mys
it was not listed, so it had crashed or had been shutdown. Looking at /var/log/mysql/mysqld.err showed the messages about the crash. Gene restarted MySQL on ippdb09 as follows:restarting MySQL: /etc/init.d/mysql start (it complains about the PID file being left behind, but tries to start anyway. The start also takes too long, and the startup script gives up and claims the start fails. NOTE this is not true: you need to check the process list (ps aux | grep mys) and check for mysql. It will probably be there, it is just taking a long time to start up. You can check the log file (tail -f /var/log/mysql/mysqld.err) and wait for it to say "ready for connections.").
Once mysql was back up, nebulous started working again and downloads proceeded OK.
Also, I put ippb10-11 back to up to deal with the storage there. Temperature in the ARTC IT seems to be holding, as processing picked up.
Tuesday : 2019.04.23
- TdB: Following the fix to the AC system in the ATRC IT centre, the b-nodes that were put to down to reduce the temperature (ippb08-23) are now but pack up.
Following the AC fix, not all machines came back up without problems. ippb08 and ipp17 are reporting to be down in ganglia. For ippb08, a restart of gmond was needed:
sudo /etc/init.d/gmond restartippb17 seems not be mounting its datadirs. Doing less /etc/exports show no dirs are being exported at all, and doing sudo exportfs -s shows nothing is exported either. The exports file needs to be fixed, which required Gavin's help to copy the required data to the file. Finally, the ganglia daemon for the ganglia user account was not running, which was the reason for it still showing up as off on ganglia.
Wednesday : 2019.04.24
- TdB: There are two exposures stuck in the full state. send them to clean and then update again:
chiptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -chip_id 1185108 chiptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -chip_id 1185115
followed by:chiptool -dbname gpc1 -setimfiletoupdate -set_label update.LAP.PV3 -chip_id 1185108 chiptool -dbname gpc1 -setimfiletoupdate -set_label update.LAP.PV3 -chip_id 1185115
That cleared them out. There is also a warp error that need rerunning. A fault 4 with memory block corruption.chiptool -dbname gpc1 -definebyquery -set_label ps_ud_WEB -set_workdir neb://@HOST@.0/gpc1/mops.fixbrokenLAP.20180323 -set_dist_group NULL -set_tess_id RINGS.V3 -set_end_stage warp -set_data_group mops.fixbrokenLAP.20180323 -set_reduction LAP_SCIENCE -exp_name o5250g0237o chiptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -chip_id 1532684 warptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -warp_id 1479726
- MEH: sending more chips finished warp updates to cleanup later tonight --
- MEH: manually restarting ~ippitc/pstamp after a large backlog QUB request finished
Thursday : 2019.04.25
- MEH: queuing up a backlog of OSS.WS for distribution
- TdB: There are some missing burn tables that needed fixing, using:
ipp_apply_burntool_fix.pl --exp_name o6601g0088o --class_id XY61 --verbose --dbname gpc1
and subsequent reversion
There were also some files that cannot load ch.fits files. I sent them to clean and followed that by an update:
warptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -warp_id 1154188followed by:
warptool -dbname gpc1 -setskyfiletoupdate -set_label update.LAP.PV3 -warp_id 1154188There were also some memory block issues, which were sent to be rerun from scratch:
chiptool -dbname gpc1 -definebyquery -set_label ps_ud_WEB -set_workdir neb://@HOST@.0/gpc1/mops.fixbrokenLAP.20180323 -set_dist_group NULL -set_tess_id RINGS.V3 -set_end_stage warp -set_data_group mops.fixbrokenLAP.20180323 -set_reduction LAP_SCIENCE -exp_name o6150g0465o chiptool -dbname gpc1 -definebyquery -set_label ps_ud_WEB -set_workdir neb://@HOST@.0/gpc1/mops.fixbrokenLAP.20180323 -set_dist_group NULL -set_tess_id RINGS.V3 -set_end_stage warp -set_data_group mops.fixbrokenLAP.20180323 -set_reduction LAP_SCIENCE -exp_name o5803g0117o chiptool -dbname gpc1 -definebyquery -set_label ps_ud_WEB -set_workdir neb://@HOST@.0/gpc1/mops.fixbrokenLAP.20180323 -set_dist_group NULL -set_tess_id RINGS.V3 -set_end_stage warp -set_data_group mops.fixbrokenLAP.20180323 -set_reduction LAP_SCIENCE -exp_name o5803g0135o chiptool -dbname gpc1 -definebyquery -set_label ps_ud_WEB -set_workdir neb://@HOST@.0/gpc1/mops.fixbrokenLAP.20180323 -set_dist_group NULL -set_tess_id RINGS.V3 -set_end_stage warp -set_data_group mops.fixbrokenLAP.20180323 -set_reduction LAP_SCIENCE -exp_name o6284g0063o chiptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -chip_id 1202573 chiptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -chip_id 1218742 chiptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -chip_id 1218743 chiptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -chip_id 1227699 warptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -warp_id 1151879 warptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -warp_id 1169382 warptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -warp_id 1169386 warptool -dbname gpc1 -updaterun -set_state goto_cleaned -set_label goto_cleaned -warp_id 1183696
Finally, there were also some files with multiple zero size files, for which non-zero copies where copied over.
Friday : 2019.04.26
- TdB: More burntools again:
ipp_apply_burntool_fix.pl --exp_name o5417g0140o --class_id XY10 --verbose --dbname gpc1 ipp_apply_burntool_fix.pl --exp_name o5408g0398o --class_id XY10 --verbose --dbname gpc1 ipp_apply_burntool_fix.pl --exp_name o5407g0515o --class_id XY10 --verbose --dbname gpc1 ipp_apply_burntool_fix.pl --exp_name o5438g0127o --class_id XY10 --verbose --dbname gpc1 ipp_apply_burntool_fix.pl --exp_name o5512g0091o --class_id XY10 --verbose --dbname gpc1 ipp_apply_burntool_fix.pl --exp_name o5518g0177o --class_id XY10 --verbose --dbname gpc1 ipp_apply_burntool_fix.pl --exp_name o5518g0178o --class_id XY10 --verbose --dbname gpc1
There are two files for which the burntool fix fails:
ipp_apply_burntool_fix.pl --exp_name o5391g0113o --class_id XY10 --verbose --dbname gpc1 ipp_apply_burntool_fix.pl --exp_name o5391g0123o --class_id XY10 --verbose --dbname gpc1It returns an error:
instance location: file:///data/ipp139.1/nebulous/f0/9d/11703263040.gpc1:20100714:o5391g0113o:o5391g0113o.ota10.burn.tbl Running [neb-mv neb://ipp008.0/gpc1/20100714/o5391g0113o/o5391g0113o.ota10.burn.tbl neb://ipp008.0/gpc1/20100714/o5391g0113o/o5391g0113o.ota10.burn.tbl.bad]... neb-rm --force neb://ipp008.0/gpc1/20100714/o5391g0113o/o5391g0113o.ota10.burn.tbl.bad /home/kiawe/eugene/src/psconfig/ipp-gentoo-ipp070.lin64/bin/ipp_apply_burntool_single.pl --dbname gpc1 --camera GPC1 --exp_id 192159 --class_id XY10 --this_uri neb://ipp008.0/gpc1/20100714/o5391g0113o/o5391g0113o.ota10.fits --previous_uri neb://ipp008.0/gpc1/20100714/o5391g0112o/o5391g0112o.ota10.fits --verbose Running [/home/kiawe/eugene/src/psconfig/ipp-gentoo-ipp070.lin64/bin/ipp_apply_burntool_single.pl --dbname gpc1 --camera GPC1 --exp_id 192159 --class_id XY10 --this_uri neb://ipp008.0/gpc1/20100714/o5391g0113o/o5391g0113o.ota10.fits --previous_uri neb://ipp008.0/gpc1/20100714/o5391g0112o/o5391g0112o.ota10.fits --verbose]... Argument "'/home/kiawe/eugene/src/psconfig/ipp-gentoo-ipp070.lin64..." isn't numeric in right bitshift (>>) at /home/kiawe/eugene/src/psconfig/ipp-gentoo-ipp070.lin64/bin/ipp_apply_burntool_fix.pl line 122. Unable to perform /home/kiawe/eugene/src/psconfig/ipp-gentoo-ipp070.lin64/bin/ipp_apply_burntool_single.pl --dbname gpc1 --camera GPC1 --exp_id 192159 --class_id XY10 --this_uri neb://ipp008.0/gpc1/20100714/o5391g0113o/o5391g0113o.ota10.fits --previous_uri neb://ipp008.0/gpc1/20100714/o5391g0112o/o5391g0112o.ota10.fits --verbose: 4 at /home/kiawe/eugene/src/psconfig/ipp-gentoo-ipp070.lin64/bin/ipp_apply_burntool_fix.pl line 123
