== PS1 IPP Czar Logs for the week 2014-02-24 - 2014-03-02 == [[PageOutline]] (Up to [wiki:PS1_IPP_CzarLogs PS1 IPP Czar Logs]) === Monday : 2014-02-24 === * 10:37 Bill set stdscience to stop. It is desperately in need of a restart. * serge came by and reported that the postage stamp server is slow. Restarted it and added a set of compute3 nodes. * 10:44 stdscience restarted with standard (low power host set) * restarting staticsky * 14:45 CZW: quick note on how I investigated the issue from last night: * Step 1: check that registration was working correctly: {{{ regpeek.pl }}} * Step 2: investigate the state of diffs: {{{ nightly_science.pl --debug --verbose --queue_diffs --date 2014-02-24 }}} This command does not actually execute any commands, but does the full logic check to see what's up. The interesting line is: {{{ diff_queue: Number of input warps to make diffs is not even for target OSS and object ps1_14_1783! 4 : I should declare an exposure to be qualityy. }}}, which despite the spelling error, points out that object ps1_14_1783 has 5 warps, and so it needs to exclude one (because diffs require an even number of exposures). However, this is done based on exp_id (for historical reasons), which kicks out the visit 1 instance. === Tuesday : 2014-02-25 === * 10:54 Bill: daily restart of pstamp server * 11:27 Bill: stopping staticsky in preparation to rebuild the tag to support distribution of full force runs * back to run I will do this after lunch. * 12:59 set to stop * 16:30 Bill: rebuilt ippsky's ipp and started up a pantasks to distribute the full force runs. This shouldn't take too long to complete. * 16:38 set staticsky to run * 17:30 CZW: In an attempt to keep stdscience from falling behind, I've made the following changes: * stdscince: {{{ hosts add wave2; hosts off ignore_wave2 hosts add wave2; hosts off ignore_wave2 hosts add wave3; hosts off ignore_wave3 hosts add wave3; hosts off ignore_wave3 hosts add compute2; hosts off ignore_compute2 }}} * staticsky: {{{ hosts off wave3 }}} === Wednesday : 2014-02-26 === * 05:10 Bill: processing is pretty far behind again. Set staticsky to stop. Rate may have climbed a little bit. * 06:10 summit copy backlog 67 * 07:30 Bill: added 2 sets of compute3 nodes to stdscience removed LAP label stdscience * pcontrol is spinning I'm not sure that pantasks is keeping things loaded as well as it could. status commands mostly timnig out. * really really need to do a daily restart * 07:53 now we are getting a very large number of database related faults. detselect commands are failing. ipptool -pending queries are timing out. * 07:58 Bill: I'm going to restart stdscience * 08:05 restarted stdscicence adding the deepstack hosts (staticsky is off) * 10:13 Bill: with permission from the czar (gene) I am adding the STS label to stdscience. I will watch for memory problems. (I could not reproduce the result that we got on Friday) * the memory problem is back. It seems that some chips (XY24 and XY31 at least) for some exposures gets really whacky results, generating lots of detections (>300,000), which causes memory to grow very large (>20GB) when writing the cmf file. Unfortunately the targets for these chips (ipp016 and ipp020) have 24 GB of memory so they can't handle multiple ppImage instances. Further debugging is required * 10:55 Gene is going to restart things without the STS label. === Thursday : 2014-02-27 === === Friday : 2014-02-28 === === Saturday : 2014-03-01 === === Sunday : 2014-03-02 ===