IPP Software Navigation Tools IPP Links Communication Pan-STARRS Links

Changes between Version 1 and Version 2 of Production_Cluster_Status


Ignore:
Timestamp:
Oct 8, 2009, 10:03:50 AM (17 years ago)
Author:
eugene
Comment:

--

Legend:

Unmodified
Added
Removed
Modified
  • Production_Cluster_Status

    v1 v2  
    1 <pre>
     1== IPP MHPCC Production Cluster Status ==
    22
    3 Status as of 2008-11-14
     3(Up to [wiki:IPP_for_PS1 IPP for PS1])
    44
    5 Nodes down:
     5'''Status as of 2008-11-14'''
     6
     7=== Nodes down ===
    68
    79None.
    810
    9 Notes:
     11=== Notes ===
    1012
    11 - ipp015 is rebuilding disk #16
     13 * ipp015 is rebuilding disk \#16
    1214
    13 Known issues:
     15=== Known issues ===
    1416
    15 - random system crashes under heavy load of nodes, occasionally with a printk() of "do_IRQ: X.XXX" which appears to be caused by a hardware interrupt that does not have a handler
     17 * random system crashes under heavy load of nodes, occasionally with a printk() of "do_IRQ: X.XXX" which appears to be caused by a hardware interrupt that does not have a handler
    1618  (device driver) for it.
    17 - ipp004 - disk bay #12 dead (system is usable)
    18 - ipp005 - seems to have random instability issues beyond the do_IRQ issue of the other nodes
    19 - ipp008 - disk bay #7 dead (system is usable)
    20 - ipp013 & ipp016 have dead fans but it is not the fans themselves (tried replacing them).  It appears to be the cabling that leads up to the fan modules that's gone bad.
    21 - forcedeth.c max_interrupt_work is still too conservative (requires a cluster restart to change).
    22 - ipp025 - suspect bad memory module (system is usable)
    23 - Our contact at MHPCC, Brad Thomas, notified me that an alarm from cabinet 4 managed power strip had been set off last evening due to high loads on nodes (ipp017, 030-036).
     19 * ipp004 - disk bay #12 dead (system is usable)
     20 * ipp005 - seems to have random instability issues beyond the do_IRQ issue of the other nodes
     21 * ipp008 - disk bay #7 dead (system is usable)
     22 * ipp013 & ipp016 have dead fans but it is not the fans themselves (tried replacing them).  It appears to be the cabling that leads up to the fan modules that's gone bad.
     23 * forcedeth.c max_interrupt_work is still too conservative (requires a cluster restart to change).
     24 * ipp025 - suspect bad memory module (system is usable)
     25 * Our contact at MHPCC, Brad Thomas, notified me that an alarm from cabinet 4 managed power strip had been set off last evening due to high loads on nodes (ipp017, 030-036).
    2426  Please consider distributing your jobs evenly across ipp production nodes.
    2527
    26   Below is the layout at MHPCC.
     28===  Layout at MHPCC ===
     29{{{
    2730  cabinet0: ipp020-21
    2831  cabinet1: ipp012-014,008,016,018-019
     
    3033  cabinet3: ipp023-029
    3134  cabinet4: ipp030-036,017
    32 
    33 http://kiawe.ifa.hawaii.edu/IPPwiki/index.php/Production_Cluster_Status
    34 
    35 </pre>
     35}}}