IPP Software Navigation Tools IPP Links Communication Pan-STARRS Links
wiki:PS1_IPP_Czarlog_20141208

Version 33 (modified by Mark Huber, 12 years ago) ( diff )

--

PS1 IPP Czar Logs for the week YYYY.MM.DD - YYYY.MM.DD

(Up to PS1 IPP Czar Logs)

Monday : 2014.12.08

  • 08:45 MEH: ippmd using ~280 nodes now that nightly is finished (ippsXX, 2x x2)
  • 09:05 MEH: WS diffs queued fine today -- will be stdlocal ~300, ippmd~300, stdsci~300 -- except 600 diffs will shortly shut off chip-warp in stdlocal
  • 12:55 MEH: the 20T data nodes with space that had put neb-host up late last week to fill and increase the number of data nodes have filled or behaving poorly -- neb-host repair on all 20T nodes now
    • for the 20T nodes w/o processing things seemed mostly okay -- net in was high >50MB/s and seemed to be writing okay, just forced to write constantly and probably could be put back in, but ones doing processing probably shouldn't be

  • 20:35 HAF: various emails floating around, there is a problem with nebulous (mark / gene noticed), and summit copy /registration are faulty. The errors we see are like this:
-> pmConfigConvertFilename (pmConfig.c:1833): System error
     failed to create a new nebulous key: nebclient.c:1012 nebSetServerErr() - SOAP-ENV:Server - error: DBD::mysql::st execute failed: The table 'storage_object' is full at /usr/lib64/perl5/site_perl/5.8.8/Nebulous/Server.pm line 275,  line 12.
 -> pmConfigRead (pmConfig.c:618): System error
     Unable to resolve trace destination: neb://ipp015.0/gpc1/ThreePi.nt/2014/12/09//o7000g0063o.833489/o7000g0063o.833489.ch.1316586.XY23.trace
Unable to perform ppImage: 1 at /home/panstarrs/ipp/psconfig/ipp-20141024.lin64/bin/chip_imfile.pl line 830
	main::my_die('Unable to perform ppImage: 1', 833489, 1316586, 'XY23', 1) called at /home/panstarrs/ipp/psconfig/ipp-20141024.lin64/bin/chip_imfile.pl line 509
  • 20:35 HAF: Serge is investigating, notes that ippdb00 is full and is finding the magical incantations to fix that.
  • 20:48 SC: Magic incantation is this: PURGE BINARY LOGS TO 'mysqld-bin.003942';
  • 22:17 HAF: seeng the same errors again. this time for 'instance' table. Now what? It's jamming up registration and stuff.

Tuesday : 2014.12.09

  • 07:45 EAM : processing has been limping along, with a number of the 'table full' errors. there is enough room on the disk after Serge purged the binary logs, so that is not the cause. the tables are big, but not approaching the InnoDB 64 TB max values. The load on the machine (ippdb00) is modest (3-4). I am guessing that a restart of mysql might clear out something which is cached?
  • 08:10 EAM : more research has revealed the likely cause: the number of concurrent transactions was too large (the full disk probably also triggered this). I turned down the number of nodes doing cleanup and this seems to be making things better (lower error rate). the following mysql bug report is relevant (and points out that 5.5.xx has bumped the transaction limit): http://bugs.mysql.com/bug.php?id=26590
    • PS1_IPP_Czarlog_20141201 -- when looking at the logs on saturday after powerup seems like we have been hitting this harder since mid-october
  • 08:15 EAM : addendum: this is probably not as critical, but we may want to bump the ibdata file. this is from the mysql manual (http://dev.mysql.com/doc/refman/5.0/en/innodb-data-log-reconfiguration.html)
For example, this tablespace has just one auto-extending data file ibdata1:

innodb_data_home_dir =
innodb_data_file_path = /ibdata/ibdata1:10M:autoextend
Suppose that this data file, over time, has grown to 988MB. Here is the configuration line after modifying the original data file to not be auto-extending and adding another auto-extending data file:

innodb_data_home_dir =
innodb_data_file_path = /ibdata/ibdata1:988M;/disk2/ibdata2:50M:autoextend
When you add a new file to the tablespace configuration, make sure that it does not exist. InnoDB will create and initialize the file when you restart the server.
  • 09:00 EAM : put ipp077 to 'repair' now that it has a new 10g card
  • 09:50 MEH: ippmd lite processing back on now that faults seem to be happening less again and stdlocal running -- set stop as needed, just leave note -- nope, md still many faults and so off
  • 11:40 MEH: removing the WS label from stdsci so power is focused on the more timely WW diffs for MOPS
  • 14:50 EAM : ipp034 crashed, power cycling
  • 14:55 CZW: while working on a revert task for registration exposures, I reverted two exposures from XXX (I don't see a chipRun for it, so it might have been engineering) and 2014-08-05. These will likely process as usual and confuse everyone. I was going to reboot ipp034, but I see that Gene has beaten me to it.
  • 19:00 MEH: stdlocal seems to be running w/ 4x x2, not the 3x so turning 1x off to be sure ippmd can run
    • stdsci has a massive loading of 5x c2 and 3x x2, when jobs available hits >550, stdlocal >300 and things are faulting.. -- turning 3x c2 off since stdlocal is also running stacks and that could really hurt nightly processing..
  • 19:45 MEH: was all kind of mess, more than adding few ipps for MD -- was stdsci allocation rebalanced after nightly finished this afternoon? -- taking the 5x:c2 and 2x:x2 out of nightly for now, may want ~100 more like on sunday night if doing WS
    • somehow stdlocal has 4x:x2 again.. massive faults.. turning 1x:x2 off -- getting better
  • 21:00 MEH: looks like balance finally restored, few faults -- stdsci=363, stdlocal~160 (still dropping so may need to turn some back on), ippmd~55 and summit/registration etc for nightly
    • looks like WS are still set to trigger in morning so not going to worry about that
    • massive camera backlog -- raise poll to 30, could probably go higher but don't want to overload
  • 22:30 MEH: stdlocal stabilized ~114, should be able to raise to ~160 so can add x2(+x3) and see
    • seems fine, since stdlocal has stacks and try adding another 60 stdmd to see if another x2(+x3) could go into stdlocal -- suspect not, and stdlocal has switched to only stacks ~170 now so hard to test
  • 22:50 MEH: ippc63 unresponsive --

Wednesday : YYYY.MM.DD

Thursday : YYYY.MM.DD

Friday : YYYY.MM.DD

Saturday : YYYY.MM.DD

Sunday : YYYY.MM.DD

Note: See TracWiki for help on using the wiki.