| | 1 | <pre> |
| | 2 | |
| | 3 | Status as of 2008-11-14 |
| | 4 | |
| | 5 | Nodes down: |
| | 6 | |
| | 7 | None. |
| | 8 | |
| | 9 | Notes: |
| | 10 | |
| | 11 | - ipp015 is rebuilding disk #16 |
| | 12 | |
| | 13 | Known issues: |
| | 14 | |
| | 15 | - random system crashes under heavy load of nodes, occasionally with a printk() of "do_IRQ: X.XXX" which appears to be caused by a hardware interrupt that does not have a handler |
| | 16 | (device driver) for it. |
| | 17 | - ipp004 - disk bay #12 dead (system is usable) |
| | 18 | - ipp005 - seems to have random instability issues beyond the do_IRQ issue of the other nodes |
| | 19 | - ipp008 - disk bay #7 dead (system is usable) |
| | 20 | - ipp013 & ipp016 have dead fans but it is not the fans themselves (tried replacing them). It appears to be the cabling that leads up to the fan modules that's gone bad. |
| | 21 | - forcedeth.c max_interrupt_work is still too conservative (requires a cluster restart to change). |
| | 22 | - ipp025 - suspect bad memory module (system is usable) |
| | 23 | - Our contact at MHPCC, Brad Thomas, notified me that an alarm from cabinet 4 managed power strip had been set off last evening due to high loads on nodes (ipp017, 030-036). |
| | 24 | Please consider distributing your jobs evenly across ipp production nodes. |
| | 25 | |
| | 26 | Below is the layout at MHPCC. |
| | 27 | cabinet0: ipp020-21 |
| | 28 | cabinet1: ipp012-014,008,016,018-019 |
| | 29 | cabinet2: ipp004-007,015,009-011 |
| | 30 | cabinet3: ipp023-029 |
| | 31 | cabinet4: ipp030-036,017 |
| | 32 | |
| | 33 | http://kiawe.ifa.hawaii.edu/IPPwiki/index.php/Production_Cluster_Status |
| | 34 | |
| | 35 | </pre> |