IPP Software Navigation Tools IPP Links Communication Pan-STARRS Links

Changes between Initial Version and Version 1 of Production_Cluster_Status


Ignore:
Timestamp:
Feb 24, 2009, 4:23:57 PM (17 years ago)
Author:
trac
Comment:

--

Legend:

Unmodified
Added
Removed
Modified
  • Production_Cluster_Status

    v1 v1  
     1<pre>
     2
     3Status as of 2008-11-14
     4
     5Nodes down:
     6
     7None.
     8
     9Notes:
     10
     11- ipp015 is rebuilding disk #16
     12
     13Known issues:
     14
     15- random system crashes under heavy load of nodes, occasionally with a printk() of "do_IRQ: X.XXX" which appears to be caused by a hardware interrupt that does not have a handler
     16  (device driver) for it.
     17- ipp004 - disk bay #12 dead (system is usable)
     18- ipp005 - seems to have random instability issues beyond the do_IRQ issue of the other nodes
     19- ipp008 - disk bay #7 dead (system is usable)
     20- ipp013 & ipp016 have dead fans but it is not the fans themselves (tried replacing them).  It appears to be the cabling that leads up to the fan modules that's gone bad.
     21- forcedeth.c max_interrupt_work is still too conservative (requires a cluster restart to change).
     22- ipp025 - suspect bad memory module (system is usable)
     23- Our contact at MHPCC, Brad Thomas, notified me that an alarm from cabinet 4 managed power strip had been set off last evening due to high loads on nodes (ipp017, 030-036).
     24  Please consider distributing your jobs evenly across ipp production nodes.
     25
     26  Below is the layout at MHPCC.
     27  cabinet0: ipp020-21
     28  cabinet1: ipp012-014,008,016,018-019
     29  cabinet2: ipp004-007,015,009-011
     30  cabinet3: ipp023-029
     31  cabinet4: ipp030-036,017
     32
     33http://kiawe.ifa.hawaii.edu/IPPwiki/index.php/Production_Cluster_Status
     34
     35</pre>