IPP Software Navigation Tools IPP Links Communication Pan-STARRS Links
wiki:PS1_IPP_Czarlog_20190204

Version 5 (modified by Mark Huber, 7 years ago) ( diff )

--

PS1 IPP Czar Logs for the week 2019.02.04 - 2019.02.11

(Up to PS1 IPP Czar Logs)

Monday : 2019.02.04

JRF: Took over as czar, from the previous night there was a hiccup in the PS2 summitcopy logs, see below. Likely just a disc error (funpack).

*** stderr ***
stderr 13980
funpack returned exit status 26624
Unable to perform dsget: 29 at /data/ippc64.1/ippitc/psconfig/ipp-20170121.lin64/bin/summit_copy.

failure for: summit_copy.pl --uri http://ipp113.ifa.hawaii.edu/ds-gpc1/o8518g0235o/o8518g0235o34.fits --filename neb://ipp093.0/gpc1/20190204/o8518g0235o/o8518g0235o.ota34.fits --summit_id 1445517 --exp_name o8518g0235o --inst gpc1 --telescope ps1 --class chip --class_id ota34 --bytes 49432320 --md5 a86555c5072ea6e317764896c0a0b02c --dbname gpc1 --timeout 600 --verbose --copies 2 --compress --nebulous
job exit status: 29
job host: ippc99
job dtime: 33.97878
job exit date: Sun Feb  3 22:27:45 2019

Tuesday : 2019.02.05

JRF: Ran the night report and saw that there was one exposure that had not gone to the chipRun stage (o8519g0755o). It was intentionally not sent to the next stage, as there is no config for processing this type of exposure yet. It's end stage was registration so all was good. JRF: Will be meeting with Mark later to go through some errors that were reported in the logs (TODO: update this)

JRF: stdscience error log reported too many connections to the database. This had come up previously (in the past month):

 -> psDBAlloc (psDB.c:166): Database error generated by the server
     Failed to connect to database.  Error: Too many connections
 -> faketoolConfig (faketoolConfig.c:373): unknown psLib error
     Can't configure database
 -> main (faketool.c:70): (null)
     failed to configure
Unable to perform faketool -addproces

config error for: fake_imfile.pl --exp_id 1450315 --fake_id 2062988 --class_id XY31 --chiproot=neb://ipp090.0/gpc1/OSS.nt/2019/02/05//o8519g0140o.1450315/o8519g0140o.1450315.ch.2127506 --camroot=neb://any/gpc1/OSS.nt/2019/02/05//o8519g0140o.1450315/o8519g0140o.1450315.cm.2093978 --camera GPC1 --outroot neb://ipp090.0/gpc1/OSS.nt/2019/02/05//o8519g0140o.1450315/o8519g0140o.1450315.fk.2062988 --reduction SWEETSPOT --dbname gpc1 --verbose
job exit status: 3
job host: ippc69
job dtime: 3.755007
job exit date: Mon Feb  4 20:22:07 2019
*** stdout ***
stdout 13038

and another:

 -> psDBAlloc (psDB.c:166): Database error generated by the server
     Failed to connect to database.  Error: Too many connections
 -> faketoolConfig (faketoolConfig.c:373): unknown psLib error
     Can't configure database
 -> main (faketool.c:70): (null)
     failed to configure
Unable to perform faketool -addproces

config error for: fake_imfile.pl --exp_id 1450315 --fake_id 2062988 --class_id XY54 --chiproot=neb://ipp112.0/gpc1/OSS.nt/2019/02/05//o8519g0140o.1450315/o8519g0140o.1450315.ch.2127506 --camroot=neb://any/gpc1/OSS.nt/2019/02/05//o8519g0140o.1450315/o8519g0140o.1450315.cm.2093978 --camera GPC1 --outroot neb://ipp112.0/gpc1/OSS.nt/2019/02/05//o8519g0140o.1450315/o8519g0140o.1450315.fk.2062988 --reduction SWEETSPOT --dbname gpc1 --verbose
job exit status: 3
job host: ippc79
job dtime: 3.872324
job exit date: Mon Feb  4 20:22:07 2019
*** stdout ***
stdout 11212
  • MEH: increased ippdb08 mysql from 256 to 512 for tonight with set global (ephemeral) and not in my.cnf to monitor how things go

JRF: The stdscience log last night on PS1 showed a bunch of errors reported around 20:30. The majority of these were on ippc118. Look at the ganglia logs there was a spike in network activity, up to about 17Mb, but not too large that it should cause an issue. Either way the tasks that failed reverted successfully and continued on after this.

JRF: MEH will look into the reported nightly_science.pl errors. Some are reported for gpc2 in the gpc1 logs!! e.g.

failure for: nightly_science.pl --queue_diffs --date 2019-02-05 --dbname gpc2 --camera GPC2
job exit status: 2
job host: localhost
job dtime: 8.072818
job exit date: Mon Feb  4 21:04:02 2019

Commented out the use of gpc2 dates in the 'stdscience/input' file. Believe this is the reason that there are gpc2 errors being reported in the gpc1 logs above.

JRF: killed a 4 day old job in the database that was run as ippuser to query something. WARNING: very dangerous to go around killing jobs.

Wednesday : 2019.02.06

  • MEH: nebulous apache server (ippc70-c75) log reset --
    stop all pantasks -- email ipp-dev to let know offline for X time 
    
    sudo /etc/init.d/apache2 stop
    sudo /etc/init.d/apache2 status
    
    sudo rm /tmp/nebulous_server.log ; sudo touch /tmp/nebulous_server.log ; sudo chown apache /tmp/nebulous_server.log ; sudo chmod g+w /tmp/nebulous_server.log ; ls -l /tmp/*log
    
    sudo /etc/init.d/apache2 start
    sudo /etc/init.d/apache2 status
    

JRF: ME, MEH, and TdB shall be going though reprocessing some data for Rob without some pixels removed (as an NEO is suspected to be falling into a masked area).

Thursday : 2019.02.07

Friday : 2019.02.08

Saturday : 2019.02.09

Sunday : 2019.02.10

Note: See TracWiki for help on using the wiki.