| Version 22 (modified by , 12 years ago) ( diff ) |
|---|
PS1 IPP Czar Logs for the week YYYY.MM.DD - YYYY.MM.DD
(Up to PS1 IPP Czar Logs)
Monday : 2014.12.08
- 08:45 MEH: ippmd using ~280 nodes now that nightly is finished (ippsXX, 2x x2)
- 09:05 MEH: WS diffs queued fine today -- will be stdlocal ~300, ippmd~300, stdsci~300 -- except 600 diffs will shortly shut off chip-warp in stdlocal
- 12:55 MEH: the 20T data nodes with space that had put neb-host up late last week to fill and increase the number of data nodes have filled or behaving poorly -- neb-host repair on all 20T nodes now
- for the 20T nodes w/o processing things seemed mostly okay -- net in was high >50MB/s and seemed to be writing okay, just forced to write constantly and probably could be put back in, but ones doing processing probably shouldn't be
- 20:35 HAF: various emails floating around, there is a problem with nebulous (mark / gene noticed), and summit copy /registration are faulty. The errors we see are like this:
-> pmConfigConvertFilename (pmConfig.c:1833): System error
failed to create a new nebulous key: nebclient.c:1012 nebSetServerErr() - SOAP-ENV:Server - error: DBD::mysql::st execute failed: The table 'storage_object' is full at /usr/lib64/perl5/site_perl/5.8.8/Nebulous/Server.pm line 275, line 12.
-> pmConfigRead (pmConfig.c:618): System error
Unable to resolve trace destination: neb://ipp015.0/gpc1/ThreePi.nt/2014/12/09//o7000g0063o.833489/o7000g0063o.833489.ch.1316586.XY23.trace
Unable to perform ppImage: 1 at /home/panstarrs/ipp/psconfig/ipp-20141024.lin64/bin/chip_imfile.pl line 830
main::my_die('Unable to perform ppImage: 1', 833489, 1316586, 'XY23', 1) called at /home/panstarrs/ipp/psconfig/ipp-20141024.lin64/bin/chip_imfile.pl line 509
- 20:35 HAF: Serge is investigating, notes that ippdb00 is full and is finding the magical incantations to fix that.
- 20:48 SC: Magic incantation is this: PURGE BINARY LOGS TO 'mysqld-bin.003942';
- 22:17 HAF: seeng the same errors again. this time for 'instance' table. Now what? It's jamming up registration and stuff.
Tuesday : 2014.12.09
- 07:45 EAM : processing has been limping along, with a number of the 'table full' errors. there is enough room on the disk after Serge purged the binary logs, so that is not the cause. the tables are big, but not approaching the InnoDB 64 TB max values. The load on the machine (ippdb00) is modest (3-4). I am guessing that a restart of mysql might clear out something which is cached?
- 08:10 EAM : more research has revealed the likely cause: the number of concurrent transactions was too large (the full disk probably also triggered this). I turned down the number of nodes doing cleanup and this seems to be making things better (lower error rate). the following mysql bug report is relevant (and points out that 5.5.xx has bumped the transaction limit): http://bugs.mysql.com/bug.php?id=26590
- PS1_IPP_Czarlog_20141201 -- when looking at the logs on saturday after powerup seems like we have been hitting this harder since mid-october
- 08:15 EAM : addendum: this is probably not as critical, but we may want to bump the ibdata file. this is from the mysql manual (http://dev.mysql.com/doc/refman/5.0/en/innodb-data-log-reconfiguration.html)
For example, this tablespace has just one auto-extending data file ibdata1: innodb_data_home_dir = innodb_data_file_path = /ibdata/ibdata1:10M:autoextend Suppose that this data file, over time, has grown to 988MB. Here is the configuration line after modifying the original data file to not be auto-extending and adding another auto-extending data file: innodb_data_home_dir = innodb_data_file_path = /ibdata/ibdata1:988M;/disk2/ibdata2:50M:autoextend When you add a new file to the tablespace configuration, make sure that it does not exist. InnoDB will create and initialize the file when you restart the server.
- 09:00 EAM : put ipp077 to 'repair' now that it has a new 10g card
- 09:50 MEH: ippmd lite processing back on now that faults seem to be happening less again and stdlocal running -- set stop as needed, just leave note -- nope, md still many faults and so off
- 11:40 MEH: removing the WS label from stdsci so power is focused on the more timely WW diffs for MOPS
- 14:50 EAM : ipp034 crashed, power cycling
Wednesday : YYYY.MM.DD
Thursday : YYYY.MM.DD
Friday : YYYY.MM.DD
Saturday : YYYY.MM.DD
Sunday : YYYY.MM.DD
Note:
See TracWiki
for help on using the wiki.
