| Version 14 (modified by , 10 years ago) ( diff ) |
|---|
PS1 IPP Czar Logs for the week YYYY.MM.DD - YYYY.MM.DD
(Up to PS1 IPP Czar Logs)
Monday : 2016.04.25
- 23:50 MEH: another warp fault 4 that would have stalled until manually cleared -- cannot build growth curve (psf model is invalid everywhere)
warptool -dbname gpc1 -updateskyfile -set_quality 42 -skycell_id skycell.0392.066 -warp_id 1721240 -fault 0
Tuesday : 2016.04.26
- 13:40 EAM : ipp041 fell over, rebooting now.
- 14:20 CZW : starting third batch of wave 3 rsync processes, concurrent with still running second batch. This set is ipp043,44,45,46,47. There are delays built into the jobs to try and stagger the impact.
- 20:15 EAM : stopping and restarting pantasks
Wednesday : YYYY.MM.DD
- 23:24 MEH: looks like registration crashed... restarting
[2016-04-27 23:11:07] pantasks_server[6514]: segfault at 10f8ab8 ip 0000000000408a4e sp 00000000429b9f20 error 4 in pantasks_server[400000+16000]
Thursday : 2016.04.28
- 07:40 MEH: Serge/MOPS request for manual diffims since visit 4 cam quality fault, do v2-3
difftool -dbname gpc1 -definewarpwarp -exp_id 1084724 -template_exp_id 1084735 -backwards -set_workdir neb://@HOST@.0/gpc1/OSS.nt/2016/04/28 -set_dist_group SweetSpot -set_label OSS.nightlyscience -set_data_group OSS.20160428.extra -set_reduction SWEETSPOT -simple -rerun difftool -dbname gpc1 -definewarpwarp -exp_id 1084812 -template_exp_id 1084830 -backwards -set_workdir neb://@HOST@.0/gpc1/OSS.nt/2016/04/28 -set_dist_group SweetSpot -set_label OSS.nightlyscience -set_data_group OSS.20160428.extra -set_reduction SWEETSPOT -simple -rerun
- 08:20 MEH: the .multi diffims still being made, sending a large number to cleanup...
- 08:30 MEH: doing cleanup of ps_ud% to clear the cycling faulting chips and warps...
- 11:48 MEH: manually updating local ~ipp/psconfig/ipp-20141024.lin64/share/pantasks/modules/ipphosts.mhpcc.config to use .20160428avoidfullv2 for data targeting to reduce number of ipp1xx nodes used and some full disks to reduce random allocations
- also reducing s6 group use in stdscience 6x->3x to reduce load there and since stdsci is still overpowered
- restarting nightly pantasks as normally required and make use of these changes
- 16:30 EAM: stopping pantasks to do the pantasks host re-organization (moving pantaskses from ippc01-09 to ippc20-25. I will also double check the pantasks client loading assignments to avoid overloading c20-25.
- 16:40 CZW: Running nebulous restore on rsynced data from ipp033. This is running on ipp100, as it must run directly on the host containing the data. It is incredibly lightweight (doing link() and sql updates), but will likely be running for a few hours.
Friday : YYYY.MM.DD
Saturday : YYYY.MM.DD
Sunday : YYYY.MM.DD
Note:
See TracWiki
for help on using the wiki.
