| Version 16 (modified by , 7 years ago) ( diff ) |
|---|
PS1 IPP Czar Logs for the week 2019.02.04 - 2019.02.11
(Up to PS1 IPP Czar Logs)
Monday : 2019.02.04
JRF: Took over as czar, from the previous night there was a hiccup in the PS2 summitcopy logs, see below. Likely just a disc error (funpack) -- seemed to revert okay and processed
*** stderr *** stderr 13980 funpack returned exit status 26624 Unable to perform dsget: 29 at /data/ippc64.1/ippitc/psconfig/ipp-20170121.lin64/bin/summit_copy. failure for: summit_copy.pl --uri http://ipp113.ifa.hawaii.edu/ds-gpc1/o8518g0235o/o8518g0235o34.fits --filename neb://ipp093.0/gpc1/20190204/o8518g0235o/o8518g0235o.ota34.fits --summit_id 1445517 --exp_name o8518g0235o --inst gpc1 --telescope ps1 --class chip --class_id ota34 --bytes 49432320 --md5 a86555c5072ea6e317764896c0a0b02c --dbname gpc1 --timeout 600 --verbose --copies 2 --compress --nebulous job exit status: 29 job host: ippc99 job dtime: 33.97878 job exit date: Sun Feb 3 22:27:45 2019
Tuesday : 2019.02.05
JRF: Ran the night report and saw that there was one exposure that had not gone to the chipRun stage (o8519g0755o). It was intentionally not sent to the next stage, as there is no config for processing this type of exposure yet. It's end stage was registration so all was good. JRF: Will be meeting with Mark later to go through some errors that were reported in the logs (TODO: update this)
JRF: stdscience error log reported too many connections to the database. This had come up previously (in the past month):
-> psDBAlloc (psDB.c:166): Database error generated by the server
Failed to connect to database. Error: Too many connections
-> faketoolConfig (faketoolConfig.c:373): unknown psLib error
Can't configure database
-> main (faketool.c:70): (null)
failed to configure
Unable to perform faketool -addproces
config error for: fake_imfile.pl --exp_id 1450315 --fake_id 2062988 --class_id XY31 --chiproot=neb://ipp090.0/gpc1/OSS.nt/2019/02/05//o8519g0140o.1450315/o8519g0140o.1450315.ch.2127506 --camroot=neb://any/gpc1/OSS.nt/2019/02/05//o8519g0140o.1450315/o8519g0140o.1450315.cm.2093978 --camera GPC1 --outroot neb://ipp090.0/gpc1/OSS.nt/2019/02/05//o8519g0140o.1450315/o8519g0140o.1450315.fk.2062988 --reduction SWEETSPOT --dbname gpc1 --verbose
job exit status: 3
job host: ippc69
job dtime: 3.755007
job exit date: Mon Feb 4 20:22:07 2019
*** stdout ***
stdout 13038
and another:
-> psDBAlloc (psDB.c:166): Database error generated by the server
Failed to connect to database. Error: Too many connections
-> faketoolConfig (faketoolConfig.c:373): unknown psLib error
Can't configure database
-> main (faketool.c:70): (null)
failed to configure
Unable to perform faketool -addproces
config error for: fake_imfile.pl --exp_id 1450315 --fake_id 2062988 --class_id XY54 --chiproot=neb://ipp112.0/gpc1/OSS.nt/2019/02/05//o8519g0140o.1450315/o8519g0140o.1450315.ch.2127506 --camroot=neb://any/gpc1/OSS.nt/2019/02/05//o8519g0140o.1450315/o8519g0140o.1450315.cm.2093978 --camera GPC1 --outroot neb://ipp112.0/gpc1/OSS.nt/2019/02/05//o8519g0140o.1450315/o8519g0140o.1450315.fk.2062988 --reduction SWEETSPOT --dbname gpc1 --verbose
job exit status: 3
job host: ippc79
job dtime: 3.872324
job exit date: Mon Feb 4 20:22:07 2019
*** stdout ***
stdout 11212
JRF: The stdscience log last night on PS1 showed a bunch of errors reported around 20:30. The majority of these were on ippc118. Look at the ganglia logs there was a spike in network activity, up to about 17Mb, but not too large that it should cause an issue. Either way the tasks that failed reverted successfully and continued on after this.
JRF: MEH will look into the reported nightly_science.pl errors. Some are reported for gpc2 in the gpc1 logs!! e.g.
failure for: nightly_science.pl --queue_diffs --date 2019-02-05 --dbname gpc2 --camera GPC2 job exit status: 2 job host: localhost job dtime: 8.072818 job exit date: Mon Feb 4 21:04:02 2019
Commented out the use of gpc2 dates in the 'stdscience/input' file. Believe this is the reason that there are gpc2 errors being reported in the gpc1 logs above.
MEH: the logs have useful info for tracking down problems and while the ~ippitc/start_server.sh script archives the logs, it might also be good to keep a note of running tasks and taskstats before doing a shutdown -- one way can be like this
echo "status;quit" | pantasks_client -c ~ippitc/stdscience/ptolemy.rc >& stdsci_status.`date +%y%m%d_%H%M%S` echo "status -taskstats;quit" | pantasks_client -c ~ippitc/stdscience/ptolemy.rc >& stdsci_tasksstats.`date +%y%m%d_%H%M%S` echo "controller status;quit" | pantasks_client -c ~ippitc/stdscience/ptolemy.rc >& stdsci_nodes.`date +%y%m%d_%H%M%S`
MEH: increased ippdb08 mysql from 256 to 512 for tonight with set global (ephemeral) and not in my.cnf to monitor how things go
JRF: killed a 4 day old job in the database that was run as ippuser to query something. WARNING: very dangerous to go around killing jobs.
Wednesday : 2019.02.06
MEH: recent [ps-ipp-ops] email
- PROBLEM alert - ippc75/Root Partition is CRITICAL -- nagios on 20190123 -- / nearly full, getting dangerous -- need to plan on resetting the nebulous_server.log (log rotation already in place)
- ippc75 apache nebulous disk warning -- cronjob check on 20190130 -- 98% left, now dangerous since rotated logs can reach this size before rotation, need to deal with ASAP
- RECOVERY alert - ippc75/Root Partition is OK -- nagios on 20190206 -- good
- Fatal | Event occured on: ipp097.ifa.hawaii.edu -- 20190202 -- BBU failed, could consider putting neb-host repair particularly if many of the ipp067-122 nodes start filling up (close)
- Fatal | Event occured on: ipp087.ifa.hawaii.edu -- 20190206 -- BBU failed, already in neb-host repair for this reason
- FATAL Event occurred on: ipp054 -- 20190205 -- BBU failed, already in neb-host repair, nothing to do and just reminder (applies for ipp054-066)
- PROBLEM alert - ippdb05/Root Partition is WARNING -- 20190112 -- / nearly full and very dangerous, log rotate on syslog will help
- RECOVERY alert - ippdb05/Root Partition is OK -- 20190206 -- good
- MySQL replication problem on ippc17 -- 20190206 -- reminder replication for ippRequestServer (pstamp+datastore) DB on ipp113 doesn't have replication setup still
- Cron <ippps2@ippc23> /bin/tcsh /data/ippc18.0/home/ippps2/cron_fix_fault_c23/ipp_ps2_rev4.bat --
- PROBLEM alert - ipp112dev/sda1 is WARNING -- nagios on 20190203 --
JRF: MEH, TdB and myself are going to archive the nebulous logs as they are getting very large, hence nagios complaining about lack of disk space on root. Log sizes:
- ippc71 - 20GB
- ippc72 - 46GB
- ippc73 - 78GB
- ippc74 - 74GB
- ippc75 - 78GB
There are also other logs in the /var/log/ folder that already have log rotation setup and are archived periodically like messages and apache2 that can be useful to find info on problems
- archive the neb log (in this case this is for ippc75, change the location appropriately)
user@ippc75$ rsync -avP /tmp/nebulous_server.log /export/ippc75.0/ cd /export/ippc75.0/ mv nebulous_server.log nebulous_server.log.180206 bzip2 nebulous_server.log.180206
- MEH: nebulous apache server (ippc70-c75) log reset --
stop all pantasks -- email ipp-dev to let know offline for X time sudo /etc/init.d/apache2 stop sudo /etc/init.d/apache2 status sudo rm /tmp/nebulous_server.log ; sudo touch /tmp/nebulous_server.log ; sudo chown apache /tmp/nebulous_server.log ; sudo chmod g+w /tmp/nebulous_server.log ; ls -l /tmp/*log sudo /etc/init.d/apache2 start sudo /etc/init.d/apache2 status
The above makes sure to check that apache is no longer writing to the log before moving it. After moving a new file needs to be created, with the correct permissions, for apache to write to in the future. Then the service can be restarted.
- do a test exposure
After setting everything back up it's a good idea to do a test exposure. So turn on stdscience from the pantasks_client and run the following chiptool command (change the date).
chiptool -dbname gpc1 -definebyquery -set_label MOPS.dailytestset -set_workdir neb://@HOST@.0/gpc1/MOPS.dailytestset.20190206 -set_data_group MOPS.dailytestset.20190206 -set_dist_group NULL -set_tess_id RINGS.V3 -set_end_stage warp -set_reduction SWEETSPOT -simple -exp_name o6370g0351o -pretend
JRF: When checking that we could use the console commands on ipp113, it has some troublesome status. MEH will report this to Hayden.
Outlet Name Status Post-on Delay(s) i05_pdu0[15] i48A_15_BAD OFF(locked) 0.5
N.B. To get out of the console, do: Shift + ~, . That is press shift and tilde together (possibly twice), then release and press period.
MEH: running LAP.PV3 update tests in ~ippmops:stdscience using ippc31-c63, ipps00-14
Thursday : 2019.02.07
JRF: ME, MEH, and TdB shall be going though reprocessing some data for Rob without some pixels removed (as an NEO is suspected to be falling into a masked area).
