Operational metrics and reportingAutomate the pull, curate the story23

Operational metrics and reporting [/ˌɑpərˈeɪʃənəl ˈmɛtrɪks ənd rɪˈpɔrtɪŋ/] v - Leadership doesn’t read tickets. They read trends. A large part of my IOC operations work is converting the daily stream of alerts, incidents, and changes into a small set of honest numbers and summaries that tell the reader whether things mend or worsen, and what to do about it.

The Figures Worth Keeping

I keep the reported set short by design; a dashboard of fifty numbers says nothing. The ones that earn their place:

  • MTTD and MTTR [/ˈmttd ˈænd ˈmttr/] n - mean time to detect and to resolve. Detection time is a direct report card on Tuning alert thresholds; resolution time, on staffing and runbooks.
  • Incident volume by severity and domain [/ˈɪnkɪdɛnt ˈvɒlʌmɛ ˈbi ˈsɛvɛrɪti ˈænd ˈdɒmæɪn/] n - GPU vs. network vs. The Engines of the Hall, so it’s plain where the pain lies.
  • Availability and SLA-credit trend [/ˈævæɪlæbɪlɪti ˈænd ˈkælkʌlætɪng ˈslæ ˈkrɛdɪtsslækrɛdɪt ˈtrɛnd/] n - the customer-facing scoreboard, and a leading indicator when credits begin to climb.
  • RCA timeliness and closure [/ˈrkæ ˈtɪmɛlɪnɛss ˈænd ˈklɒsʌrɛ/] n - are On Running the Root-Cause Inquiry delivered on time, and are their action items truly done?
  • Alert noise ratio [/ˈælɛrt ˈnɒɪsɛ ˈrætɪɒ/] n - fired vs. actioned, the health check on the monitoring itself.

Of Incident Summaries

Beside the numbers, leadership wants a readable account of the period: the handful of incidents that mattered, their cause, their cost, and what is being done. I write these as short, plain summaries drawn straight from the incident records in the ITSM ticketing systems and the On Running the Root-Cause Inquiry; no jargon, no blame, only impact and action. An executive should read one in a minute and know whether to worry.

Keeping Reports Honest

Two things ward a report from becoming fiction:

  • One source of truth. [/ˈɒnɛ ˈsæʊrkɛ ˈɒf ˈtrʌth/] n Every number traces back to the ticketing and monitoring systems, never to a hand-kept spreadsheet. If two reports disagree, that is a data problem to mend, not a rounding difference to wave off; it usually points back to Data integrity and log retention.
  • Automate the pull, curate the story. [/ˈæʌtɒmætɛ ˈthɛ ˈpʌll ˈkʌrætɛ ˈthɛ ˈstɒri/] n Let the collection be automated and consistent period to period; my judgment goes into the summary and the recommendations, not into copy-pasting figures. The recurring “why is this number bad” answers are precisely what feeds On Finding the Gaps in One’s Tools.