IPP Software Navigation Tools IPP Links Communication Pan-STARRS Links
wiki:PS1_IPP_Czarlog_20190311

Version 11 (modified by cclin33, 7 years ago) ( diff )

--

PS1 IPP Czar Logs for the week YYYY.MM.DD - YYYY.MM.DD

(Up to PS1 IPP Czar Logs)

Monday : 2019.03.11

  • MEH: simple oneliner to remove the faulted chip/warps in update.OSS to update.OSS.broke until ipp067 back online -- remember to add the user/password
    mysql -hscidbm -uXXXXX -pXXXXX gpc1 -B -e ' select chip_id,warp_id,warpRun.label,warpRun.data_group,warpRun.state,chipRun.label,chipRun.state from warpRun join warpSkyfile using (warp_id) join fakeRun using (fake_id) join camRun using (cam_id) join chipRun using (chip_id) join chipProcessedImfile using (chip_id) where warpRun.label like "update.OSS" and warpRun.data_group like "OSS%" and warpRun.state like "update" and  ( warpSkyfile.fault>0 or chipProcessedImfile.fault>0 ) group by warp_id limit 500; ' | grep -v data_group | awk '{print "chiptool -dbname gpc1 -updaterun -set_label update.OSS.broke -chip_id  "$1" ;warptool -dbname gpc1  -updaterun -set_label update.OSS.broke -warp_id "$2 }'
    
  • TdB: Following upgrades to Ubuntu and a series of EDAC errors, machines ippb16-23 are now behaving well again. They have been tested rigorously by Gene using rsyncs, and now put back to neb-host up.
  • CCL: set ipp134 down -> repair for testing then -> up, there are ipp134_bck.0, set it up only ATRC is offline
  • CCL: set ipp135 down -> repair for testing (ipp135_bck.o, same as ipp134_bck.0)
  • CCL: restart stdscience at 14:00
  • TdB: The ipp067 related faults on IPP monitor are being cleared out on a regular basis, when the total number of faulted exposures becomes too high (>300). In that case, stop the stdscience pantask and run the batch SQL query/chiptool command Mark gives above. Then, set the pantask back to run.
  • MEH: adding other updates to run in ~ippmops/stdscience targeting ipp123-129 for warps -- stop pantasks if needed

Tuesday : 2019.03.12

  • MEH: ipp130 looks okay, setting neb-host down->repair->up and will just target in ippmops:stdscience for now
    • if it works fine, it could be kept neb-host up when nightly processing starts since it has serial console access in i12 (needs to be added to Consolerc)
  • TdB: Fixed a bunch of red on IPPmonitor by changing the labels to update.OSS.broke. However, a few exposures failed with fault=2 and state=full, which was somewhat unusual. An example is exposure o6982g0381o, XY41 which went through a chip update. The bottom of the update log reads:
          Complete ppImage run: 20.814865 sec
          Memory leaks: 0
    Number of leaks to display: 500
    Running [/home/panstarrs/ippmops/psconfig/ipp-20141024.lin64/bin/ppStatsFromMetadata /data/ipp018.0/nebulous/ee/8d/8954414172.gpc1:OSS.nt:2014:11:21:o6982g0390o.821305:o6982g0390o.821305.ch.1262404.XY41.stats - CHIP_IMFILE]...
    -bg 314.070892 -bg_mean_stdev 1.714188 -bg_stdev 19.923737 -dtime_detrend 9.678738 -dtime_photom 0.000003 -dtime_total 20.751261 -quality 0 
    Running [/home/panstarrs/ippmops/psconfig/ipp-20141024.lin64/bin/funpack -S /data/ipp018.0/nebulous/e2/46/8954413865.gpc1:OSS.nt:2014:11:21:o6982g0390o.821305:o6982g0390o.821305.ch.1262404.XY41.ch.fits > /dev/null]...
    Running [/home/panstarrs/ippmops/psconfig/ipp-20141024.lin64/bin/funpack -S /data/ipp018.0/nebulous/58/d7/8954414001.gpc1:OSS.nt:2014:11:21:o6982g0390o.821305:o6982g0390o.821305.ch.1262404.XY41.ch.mk.fits > /dev/null]...
    Running [/home/panstarrs/ippmops/psconfig/ipp-20141024.lin64/bin/funpack -S /data/ipp018.0/nebulous/3e/19/8954414035.gpc1:OSS.nt:2014:11:21:o6982g0390o.821305:o6982g0390o.821305.ch.1262404.XY41.ch.wt.fits > /dev/null]...
    Running [/home/panstarrs/ippmops/psconfig/ipp-20141024.lin64/bin/funpack -S /data/ipp018.0/nebulous/fc/05/8954414163.gpc1:OSS.nt:2014:11:21:o6982g0390o.821305:o6982g0390o.821305.ch.1262404.XY41.mdl.fits > /dev/null]...
    Running [/home/panstarrs/ippmops/psconfig/ipp-20141024.lin64/bin/chiptool -tofullimfile -chip_id 1262404 -class_id XY41 -dbname gpc1]... 
    
    The files in question cannot be accessed since ipp018 is down on nebulous.
  • TdB: ipp136 and ipp137 put from down to up, to test for files going there. it is currently not targeted for specific products, but will be tomorrow if all works out.
  • CCL: ipp135 and ipp137 put from repair to up, for warp product targetting.
  • CCL: modify the configuration file of PS1 stdscience for warp stage targeting on ipp134 and ipp135 machine by following steps.
    the files are under these two folders
    /data/ippc64.1/ippitc/src/ipp-20170121/ippTasks
    /data/ippc64.1/ippitc/psconfig/ipp-20170121.lin64/share/pantasks/modules
    
    make a copy for backing up and a new file for adding targetting machines from "ipphosts.mhpcc.config"
    cp ipphosts.mhpcc.config ipphosts.mhpcc.config.190312ori
    cp ipphosts.mhpcc.config ipphosts.mhpcc.config.190312_target134_135
    
    add ipp134 & ipp135 to the new file
    ipphosts METADATA
      camera STR skycell
      count S32 48
      sky00 STR  ipp123.1
      sky01 STR  ipp124.1
      sky02 STR  ipp125.1
      sky03 STR  ipp126.1
      sky04 STR  ipp129.0
    ...
      sky30 STR  ipp125.0
      sky31 STR  ipp126.0
      sky32 STR  ipp134.0
      sky33 STR  ipp134.1
      sky34 STR  ipp134.0
      sky35 STR  ipp134.1
      sky36 STR  ipp134.0
      sky37 STR  ipp134.1
      sky38 STR  ipp134.0
      sky39 STR  ipp134.1
      sky40 STR  ipp135.0
      sky41 STR  ipp135.1
      sky42 STR  ipp135.0
      sky43 STR  ipp135.1
      sky44 STR  ipp135.0
      sky45 STR  ipp135.1
      sky46 STR  ipp135.0
      sky47 STR  ipp135.1
     END
    
    overwrite "ipphosts.mhpcc.config" by the new configuration file in the two folders
    cp ipphosts.mhpcc.config.190312_target134_135 ipphosts.mhpcc.config
    
    and then wait for "stdscience" restarting automatically at 3:30pm
    

Wednesday : 2019.03.13

Thursday : 2019.03.14

Friday : 2019.03.15

Saturday : 2019.03.16

Sunday : 2019.03.17

Note: See TracWiki for help on using the wiki.