On Early DetectionModel the system65

On Early Detection [/ɔn ˈərli dɪˈtɛkʃən/] n - Few outages in a data center arrive unheralded. They send word before they come: a bearing grows loud, a temperature climbs, a power reading drifts, a line in a log alters its habit. The office of an early detection system is to hear that word and turn it to action before warning ripens into failure.

Of What Such a System Is Made

  • Sensors in the right places. [/ˈsɛnsɒrs ˈɪn ˈthɛ ˈrɪght ˈplækɛs/] n A Temperature sensor on the inlet, the outlet, and the hottest spot within the rack tells more than a single reading taken mid-aisle. Vibration sensors on pumps and fans betray mechanical wear; door and airflow sensors confess containment troubles.
  • Signals that matter. [/ˈsɪgnæls ˈthæt ˈmættɛr/] n Abundance of data is not itself a virtue. A sound system attends to the few readings that foretell real trouble, not to every number that can be gathered.
  • Trending, not just thresholds. [/ˈtrɛndɪng ˈnɒt ˈjʌst ˈthrɛshɒlds/] n A threshold that fires at 30 degrees is blind to the rack that passed from 18 to 28 in ten minutes. The rate of change is oft the truer herald than the absolute limit.
  • Multiple angles. [/ˈmʌltɪplɛ ˈænglɛs/] n One sensor may lie; two sensors, or one sensor joined to a derived signal, are harder to deceive. The same reasoning leads Managing monitoring tools to keep more than one platform.
  • Human checkpoints. [/ˈhʌmæn ˈkhɛkkpɒɪnts/] n Automation tends the routine. The novel belongs to human hands: the strange noise, the unlooked-for pattern, the reading that will not agree with the model.

Of the Places Where Failure Begins

Before failure can be detected early, one must learn where it is born. A few ways of seeking those places:

  • Look at past incidents. [/ˈlʌːk ˈæt ˈpæst ˈɪnkɪdɛnts/] n What has broken before is most likely to break again. The field notes in these pages are full of such precedents.
  • Walk the floor with a checklist. [/ˈwælk ˈthɛ ˈflʌːr ˈwɪth ˈæ ˈkhɛkklɪst/] n What has a single point of failure? What has no sensor at all? What would go unnoticed from the monitoring desk?
  • Ask the equipment. [/ˈæsk ˈthɛ ˈɛqʌɪpmɛnt/] n Many devices already offer their internal counters, wear indicators, and error logs over SNMP or Modbus. To read them costs little, once the art is known.
  • Model the system. [/ˈmɒdɛl ˈthɛ ˈsistɛm/] n A simple energy balance or airflow model will say when a reading is physically implausible, and implausible readings are commonly the most interesting.

Where Custom Devices Find Their Place

Vendors cover the common cases; custom devices fill the gaps between them: the odd sensor, the legacy protocol, the location too costly to wire by commercial means. See Bespoke Devices for the Datacenter and On the Making of Bespoke Devices for how such things are built, and Raspberry Pi in the datacenter for one made concrete.

The discipline of building these systems is kin to IOC operations and to the monitoring work recorded in Managing monitoring tools. For accounts of how the best operators have pressed efficiency and reliability further, see Efficient datacenter case studies.