Operational Risk Management for COOs: A Practical Playbook

Operational risk is the risk of loss from broken processes, unreliable people, failing systems, or outside events — a payroll run that misses, a supplier that goes dark, a data breach, a key manager who quits mid-quarter. It is not market risk or credit risk. It is the risk that the machine you run stops working the way it should.
As COO, you own that machine, so you own this risk. The goal of managing it is not to eliminate every hazard — that is impossible and would cost more than it saves. The goal is to know your real exposures, put the right controls on the few that could actually hurt you, and be able to prove those controls work when the board, an auditor, or a regulator asks.
This guide walks the loop most operators run in practice: identify the risks, rank them by likelihood and impact, install controls that match, monitor with numbers that move early, and report in a way that drives decisions.
What operational risk actually covers
Basel and most risk frameworks split operational risk into four sources: people, process, systems, and external events. It helps to keep those categories in front of you, because most teams over-index on one (usually cybersecurity) and go blind on the others (usually people and process).
- People: key-person dependence, skill gaps, fraud, human error, turnover in a critical role.
- Process: a manual step with no check, a handoff that drops work, a reconciliation nobody does.
- Systems: an outage, a failed integration, a batch job that silently truncates data.
- External: a supplier collapse, a weather event, a regulatory change, a payment-processor outage you don't control.
Build a register before you build anything else
A risk register is a living list of what could go wrong, how bad it would be, how likely it is, who owns it, and what you are doing about it. Everything else — controls, monitoring, board reports — hangs off this list. Without it, "risk management" is a folder of policies nobody reads.
Strong looks like a short, owned, current register: 30 to 60 real entries, each with a named owner and a next review date, revisited monthly. Weak looks like a 400-line spreadsheet built for an audit two years ago, with no owners and no dates — technically present, operationally dead. A useful register is closer to your P&L than to a compliance binder: something you actually consult when you make decisions.
The practical move: block ninety minutes with each functional lead and ask one question — "what could stop your team from delivering next quarter?" You will get more real risk in that hour than from any generic checklist, because the people doing the work already know where the weak joints are. For a deeper framework on running that identification pass, see the risk assessment framework.
Rank by likelihood and impact — then act on the top corner
You cannot control everything, so you must rank. The standard tool is a likelihood × impact scoring model that produces a heat map. Score each risk on both axes (a 1–5 scale each is enough), multiply for a rough priority score, and sort. The point is not mathematical precision — it is forcing a conversation about which handful of risks deserve real money and attention.
| Score | Likelihood | Impact if it happens | What you do |
|---|---|---|---|
| High × High | Probable this year | Halts a core function or threatens survival | Treat now — assign owner, fund a control, track weekly |
| High × Low | Happens often | Nuisance, absorbed easily | Reduce if cheap; otherwise accept and monitor |
| Low × High | Rare | Severe — outage, breach, key-person loss | Transfer (insurance) or build a tested contingency |
| Low × Low | Rare | Minor | Accept and document; revisit at annual review |
Match the control to the risk
Controls come in three flavours, and a healthy program uses all three rather than piling up one kind.
- Preventive controls stop a bad thing from happening: segregation of duties so no single person can both approve and pay an invoice, access restrictions, approval thresholds, mandatory checklists.
- Detective controls catch it after the fact but before it compounds: reconciliations, monitoring dashboards, audit trails, exception reports, an anomaly alert on unusual transactions.
- Corrective controls limit the damage once something has gone wrong: incident-response runbooks, backups you have actually restored from, insurance, a standby supplier.
Strong control design is proportionate — you do not put a four-eyes approval on a \$50 expense. Weak design either has gaps or is so heavy that people route around it, which is worse than no control because it hides the exposure. The most valuable preventive control in most companies is unglamorous: segregation of duties in finance.
Monitor with indicators that move before the loss does
The difference between managing risk and reacting to incidents is monitoring. You want Key Risk Indicators (KRIs) — leading metrics that rise before the risk materialises — not just a tally of things that already broke.
The useful distinction: a KRI is a signal, a KPI is a result. "Number of overdue vendor payments" is a KRI for supplier-relationship risk — it climbs before a critical supplier walks. "System uptime trending below 99.9%" is a KRI for outage risk. Each should have a threshold that triggers action, not just a number on a dashboard.
Strong monitoring means three or four KRIs per top risk, each with a defined trigger and an owner who acts when it trips. Weak monitoring is a wall of green dials nobody reads, or metrics that only tell you a loss already happened. Set thresholds honestly — an amber that trips constantly gets ignored, and a red that never trips is decoration. Pair these with the operational numbers in your operations metrics so risk indicators sit next to performance ones.
Connect risk to continuity and crisis response
Risk management and business continuity are two halves of the same discipline: one reduces the chance and size of a disruption, the other keeps the business running when one hits anyway. If your risk work stops at the register and never produces tested recovery plans, you have done half the job.
For each high-impact risk, ask: if this happens tomorrow, what do we do in the first hour, and who decides? That answer becomes a continuity plan with defined recovery-time objectives — how long a function can be down before it hurts — and a crisis-communication path for staff, customers, and regulators. The plans that fail on the day are the ones never rehearsed. Build these out with a dedicated business continuity guide and a written crisis communication plan so the response is decided in calm, not invented in panic.
Supply chain deserves its own attention here, because it is the external risk most COOs underweight. Single-source suppliers, thin inventory buffers, and unmapped tier-two dependencies are where a small event upstream becomes your problem. Mapping those dependencies and building alternate sources is the core of supply chain resilience.
Report to the board so they can actually decide
A risk report exists to help the board make decisions, not to demonstrate that you are busy. The strong version is one page: your top five risks, the direction each is trending (better, worse, steady), what you are doing, and the two or three decisions or resources you need from them. The weak version is a 40-slide compliance deck that gets flipped through in silence.
Lead with what changed and what you need. A board does not want the full register — it wants to know which exposures are moving the wrong way and where its help is required (a budget approval, a risk-appetite call, a decision to exit a line of business). Framing the report around decisions rather than status is the single biggest upgrade most operators can make.
Key takeaways
- Operational risk spans people, process, systems, and external events — if your register is mostly IT, you have a blind spot, not a secure company.
- Build a short, owned, current risk register first; every control and report hangs off it.
- Rank by likelihood × impact and give each risk one of four responses: treat, transfer, tolerate, or terminate. Do not treat everything.
- Layer preventive, detective, and corrective controls; a single control is brittle, and segregation of duties in finance is the highest-value one most firms have.
- Monitor with Key Risk Indicators that move before the loss, each with a threshold and an owner who acts.
- Pair risk work with tested continuity and crisis plans — an untested plan fails on the day.
- Report to the board in one page framed around decisions and what you need, not a compliance deck.