Onboarding and operational docs [/ˈonboardɪŋ ənd ˌɑpərˈeɪʃənəl ˈdɑks/] n - The most fragile thing in an operations center isn’t a tool; it’s knowledge that lodges in one head alone. If the only way to answer a The Engines of the Hall alarm at 3 a.m. is to wake the one analyst who “just knows,” the operation isn’t resilient, it’s lucky. Writing and keeping the documentation and training that removes that single point of failure is a core part of my IOC operations work.
What I Write and Keep Current
- Runbooks. [/ˈrʌnbʌːks/] n For each recurring alert from Tuning alert thresholds, a short, tested procedure: what it means, how to confirm it, the safe steps, and when to escalate. A runbook never followed in anger is a draft, so I revise mine straight out of On Running the Root-Cause Inquiry.
- Tool guides. [/ˈtʌːl ˈgʌɪdɛs/] n How to actually use the ITSM ticketing systems and the monitoring platforms: not the vendor manual, but our conventions, how we tag, how we route, which dashboards to trust.
- Process docs. [/ˈprɒkɛss ˈdɒks/] n The lifecycles that aren’t optional: incident handling, the RCA workflow, SLA-credit handling, and access requests.
- Onboarding path. [/ˈɒnbɒærdɪŋ ˈpæth/] n A sequenced first-two-weeks for a new analyst, so they reach “safe on-call” by a known road rather than by osmosis.
Of Documents That Get Used
Operational docs rot faster even than dashboards, and a wrong runbook is worse than none. The habits that keep them alive:
- Findable. [/ˈfɪndæblɛ/] n Linked from the alert and the ticket, not buried in a wiki nobody opens. The best runbook is the one attached to the page that fired.
- Owned and dated. [/ˈɒwnɛd ˈænd ˈdætɛd/] n Each doc bears an owner and a last-reviewed date, so staleness shows.
- Born from real incidents. [/ˈbɒrn ˈfrɒm ˈrɪːl ˈɪnkɪdɛnts/] n The richest material already lies in these pages; Field Notes and Problems We Could Only Work Around are, in effect, runbooks and post-mortems for troubles we truly met.
What It Purchases
Good documentation is what lets the operations center scale past the people who built it: swifter onboarding, steadier response, and fewer On Running the Root-Cause Inquiry whose root cause is “the procedure wasn’t written down.” Where I find a process that can’t be documented cleanly because the tool fights it, that is no writing problem; it goes onto On Finding the Gaps in One’s Tools.