PS1 IPP Czar Logs for the week 2018.04.30 - 2018.05.06
(Up to PS1 IPP Czar Logs)
Monday : 2018.04.30
- 14:00 : spent part of the morning clearing pstamp failures. These had three types of problems: 1) chip imfiles in state = error_clean, 2) chip imfiles missing PSF files (stuck only on stsci nodes), 3) warp skyfiles in state error clean. I used these types of commands to clear
chiptool -dbname gpc1 -updaterun -chip_id $chip_id -set_state cleaned chiptool -dbname gpc1 -setimfiletoupdate -chip_id $chip_id -set_label ws_nightly_update -set_update_mode 1 warptool -dbname gpc1 -updaterun -warp_id $warp_id -set_state cleaned warptool -dbname gpc1 -setskyfiletoupdate -warp_id $warp_id -set_label ws_nightly_update neb-mv FOO.psf FOO.psf.bad
- 15:15 CZW: I have updated the ippps2 nightly_science.config to remove all nightly_science targets. This should prevent any training/testing data from being interpreted as real science data. I have also updated the camera format to use the FILTERID header to set the FPA.FILTER concept, which will allow registration to insert full filter names.
- 15:30 CZW: ippb02 rsyncs are stopped so I can transfer data from the .0 partitions of ippb07/8 to the .1 partitions, because I wrote the commands wrong.
- 16:30 CZW: The ingestion of GPC1 stacks into the gpc2 database cannot be completed due to a lack of usable w science data from GPC2. This is needed to construct the reference warp that links the stackRun entry to the rawExp table. Making a note here so I don't forget to fix this when GPC2 w data exists.
- MEH: reprocessed the remaining 2 exposures still broken from before (o7023g0257o, o7023g0317o) so MOPS can get their data as soon as needed without further delays -- looks like the previous fix was still in an odd not-updated state in warp that pstamp could use, a normal cleanup cycle seems to have cleared and now those exposures with the broken LAP label can fully update
- the czars really shouldn't only do fixes partway when there is likelihood of other problems in other chips/skycells and asking them to let us know of problems -- that is an unnecessary delay for MOPS on what is a lower level IPP data problem and things may also then get into odd states for higher levels like the pstamp server, causing even further delays
- MOPS really shouldn't have to be testers by default until all these broken files from the stsci shuffle are fixed -- to reduce as much of the unnecessary delay for MOPS, it would help if the czars would ensure that the entire exposure is clear of problems, the best way to do that is having the entire exposure stage in full state and letting them know after checking that is so -- typically this will involve both chip and warp stages -- reminder that ps_ud% cleanup happens daily 1800-1900 HST
- MEH: restored /export/ippx029.0/ipp dir after Haydn rebuilt the raid -- needs to be tested in processing
- MEH: ganglia not reporting on ippx071, ippc127, ippdb03, ippc17, ippc27, ippx052 for well over a week now -- restarted gmond and seem fine now -- if not reporting then cannot see problems, czars need to be regularly checking this
- MEH: the daily pantasks restart seems to be happening again now
Tuesday : 2018.05.01
- MEH: ippx027.0 disk replaced and needs to be rebooted -- not coming back up, down until further notice
- MEH: apparently the ipp113/web,php for datastore,pstamp all setup to use the old ipp-trunk-20170221 tag, changes for MOPS access and full ROI return requires using ipp-20170121 -- check, modify and monitor for problems
- web/request.php appears very different from trunk version, adding another mod and will need to look into why and check other files...
Wednesday : 2018.05.02
- MEH: ippx027 back online -- Haydn not doing a reboot test since didn't reboot easily on own, should be ok to use but suspect if need to reboot
Thursday : 2018.05.03
- MEH: ippx005 crash/unresponsive this morning, remote power management not setup (There is no outlet associated with this port) -- unclear why or how many other nodes in this state -- Curt powered ippx005 back up ~noon
- MEH: changes still not put into ops tag and rebuild in order to deal with all the missing files when doing updates and want to always avoid doing tag rebuilds on Fridays... -- ran another MOPS daily test quad (MOPS.dailytestset.20180503) to verify the rebuild is similar to that passed by the MOPS check
Friday : 2018.05.04
- MEH: update pstamp files to test updates for old pstamp web request.php to access multiple images in ROI option and increase number of sample points for different images in ROI to help avoid the missing image if position in chip gap with limits for boundary issues
Saturday : YYYY.MM.DD
Sunday : YYYY.MM.DD
Last modified
8 years ago
Last modified on May 6, 2018, 6:33:06 PM
Note:
See TracWiki
for help on using the wiki.
