| Version 25 (modified by , 14 years ago) ( diff ) |
|---|
PS1 IPP Czar Logs for the week YYYY.MM.DD - YYYY.MM.DD
(Up to PS1 IPP Czar Logs)
Monday : 2012.05.07
- 07:25 Mark: registration stalled? had to run to get things moving along
regtool -updateprocessedimfile -exp_id 484476 -class_id XY43 -set_state pending_burntool -dbname gpc1 -- seems to also be burntool taking 20ks but not going to try and fix until get in 0 ipp008 BUSY 21216.52 0.0.3.740b 0 ipp_apply_burntool_single.pl --exp_id 484466 --class_id XY43 --this_uri neb://ipp015.0/gpc1/20120507/o6054g0187o/o6054g0187o.ota43.fits --continue 10 --previous_uri neb://ipp015.0/gpc1/20120507/o6054g0186o/o6054g0186o.ota43.fits --dbname gpc1 --verbose -- nothing in registration log for this. looks like stalled on ipp008 in a neb-replicate for neb://ipp015.0/gpc1/20120507/o6054g0197o/o6054g0197o.ota43.burn.tbl -- which has 3 entries now and one with different md5sum -- culled 0 size one and didnt unstick -- neb-replicate not able to kill. not going to mess with anymore, emailed czar, maybe just needs to restart replication pantasks. otherwise mostly everything registered and processing
- 10:00 heather: messing with addstars...
- 14:05 Bill rebooting ipp008. It isn't playing nicely with others. Also will restart stdscience.
Tuesday : 2012.05.08
Bill is czar today.
- 07:30 restarted summit copy. Set ipp008 to repair. It is bogging things down.
- 09:56 summit copy is proceeding but running into "503 try again" faults from conductor
- set newExp for exp_id 456051 456052 from December 15 to state 'drop' They keep reverting and failing to register. c5972g0035l and c5972g0036l
Wednesday : 2012.05.09
Mark czar today
- 07:30 looks like nightly science finished downloading and processing.
- 11:15 Chris rebuilt ops tag and restarted stdscience to reprocess SAS with changes (details)
Thursday : 2012.05.10
Mark is czar
- 07:40 all nightly data downloaded and processed except for one chip on ipp008 running for 36ks. manually killed ppImage on ipp008, reverted and continued through warp fine. not a particularly crowded field at all, nothing stands out of the normal in the ipp008 log around the time it was originally running.
- 10:30 connections from manoa to production cluster seem to have been interrupted/lost for a bit (as was found earlier in the morning when got in). ippc19 seems to not be coming back or is down (can connect via console however so seems like network issue). when connecting via console noticed 2 pantasks_servers were running as ipp, two replication servers??
- 12:20 looks like Gavin had to reboot ippc19. restarting replication pantasks
- 12:30 Chris doing daily restarting of stdscience and will start LAP processing.
- 12:41 Bill restarted pstamp pantasks. It had many timeout errors in the status output.
- 13:40 upped the stack.poll 100->200 now that there are ~140 hosts. Chris bumped up the unwant 5->7 in stdscience.
Friday : YYYY.MM.DD
Saturday : YYYY.MM.DD
Sunday : YYYY.MM.DD
Note:
See TracWiki
for help on using the wiki.
