Operational Risk Management for COOs: A Practical Playbook

Multiracial colleagues in formal clothes sitting at table with laptop and documents while discussing details of business plan

Operational risk is the risk of loss from broken processes, unreliable people, failing systems, or outside events — a payroll run that misses, a supplier that goes dark, a data breach, a key manager who quits mid-quarter. It is not market risk or credit risk. It is the risk that the machine you run stops working the way it should.

As COO, you own that machine, so you own this risk. The goal of managing it is not to eliminate every hazard — that is impossible and would cost more than it saves. The goal is to know your real exposures, put the right controls on the few that could actually hurt you, and be able to prove those controls work when the board, an auditor, or a regulator asks.

This guide walks the loop most operators run in practice: identify the risks, rank them by likelihood and impact, install controls that match, monitor with numbers that move early, and report in a way that drives decisions.

What operational risk actually covers

Basel and most risk frameworks split operational risk into four sources: people, process, systems, and external events. It helps to keep those categories in front of you, because most teams over-index on one (usually cybersecurity) and go blind on the others (usually people and process).

  • People: key-person dependence, skill gaps, fraud, human error, turnover in a critical role.
  • Process: a manual step with no check, a handoff that drops work, a reconciliation nobody does.
  • Systems: an outage, a failed integration, a batch job that silently truncates data.
  • External: a supplier collapse, a weather event, a regulatory change, a payment-processor outage you don't control.
Strong risk management names exposures in all four. Weak risk management has a thick cyber policy and no idea that one contractor holds the only working knowledge of how the month-end close runs. The tell is simple: if your risk register is 80% IT, you have a coverage gap, not a secure company.

Build a register before you build anything else

A risk register is a living list of what could go wrong, how bad it would be, how likely it is, who owns it, and what you are doing about it. Everything else — controls, monitoring, board reports — hangs off this list. Without it, "risk management" is a folder of policies nobody reads.

Strong looks like a short, owned, current register: 30 to 60 real entries, each with a named owner and a next review date, revisited monthly. Weak looks like a 400-line spreadsheet built for an audit two years ago, with no owners and no dates — technically present, operationally dead. A useful register is closer to your P&L than to a compliance binder: something you actually consult when you make decisions.

The practical move: block ninety minutes with each functional lead and ask one question — "what could stop your team from delivering next quarter?" You will get more real risk in that hour than from any generic checklist, because the people doing the work already know where the weak joints are. For a deeper framework on running that identification pass, see the risk assessment framework.

Rank by likelihood and impact — then act on the top corner

You cannot control everything, so you must rank. The standard tool is a likelihood × impact scoring model that produces a heat map. Score each risk on both axes (a 1–5 scale each is enough), multiply for a rough priority score, and sort. The point is not mathematical precision — it is forcing a conversation about which handful of risks deserve real money and attention.

ScoreLikelihoodImpact if it happensWhat you do
High × HighProbable this yearHalts a core function or threatens survivalTreat now — assign owner, fund a control, track weekly
High × LowHappens oftenNuisance, absorbed easilyReduce if cheap; otherwise accept and monitor
Low × HighRareSevere — outage, breach, key-person lossTransfer (insurance) or build a tested contingency
Low × LowRareMinorAccept and document; revisit at annual review
Every risk gets one of four responses: treat (add a control), transfer (insure or contract it out), tolerate (accept it consciously and write down why), or terminate (stop doing the thing that creates it). The most common failure is treating everything, which spreads your team thin and buries the three risks that could actually sink the quarter. Strong operators can name their top five risks from memory and tell you what is being done about each.

Match the control to the risk

Controls come in three flavours, and a healthy program uses all three rather than piling up one kind.

  • Preventive controls stop a bad thing from happening: segregation of duties so no single person can both approve and pay an invoice, access restrictions, approval thresholds, mandatory checklists.
  • Detective controls catch it after the fact but before it compounds: reconciliations, monitoring dashboards, audit trails, exception reports, an anomaly alert on unusual transactions.
  • Corrective controls limit the damage once something has gone wrong: incident-response runbooks, backups you have actually restored from, insurance, a standby supplier.
Consider a mid-sized firm worried about invoice fraud. Preventive: require two approvers above a set amount and split the "set up a new vendor" and "pay a vendor" permissions. Detective: a weekly report flagging any new vendor bank-account change. Corrective: fraud insurance and a documented clawback process. One control alone is brittle; three layered controls mean a single failure does not become a loss.

Strong control design is proportionate — you do not put a four-eyes approval on a \$50 expense. Weak design either has gaps or is so heavy that people route around it, which is worse than no control because it hides the exposure. The most valuable preventive control in most companies is unglamorous: segregation of duties in finance.

Monitor with indicators that move before the loss does

The difference between managing risk and reacting to incidents is monitoring. You want Key Risk Indicators (KRIs) — leading metrics that rise before the risk materialises — not just a tally of things that already broke.

The useful distinction: a KRI is a signal, a KPI is a result. "Number of overdue vendor payments" is a KRI for supplier-relationship risk — it climbs before a critical supplier walks. "System uptime trending below 99.9%" is a KRI for outage risk. Each should have a threshold that triggers action, not just a number on a dashboard.

Strong monitoring means three or four KRIs per top risk, each with a defined trigger and an owner who acts when it trips. Weak monitoring is a wall of green dials nobody reads, or metrics that only tell you a loss already happened. Set thresholds honestly — an amber that trips constantly gets ignored, and a red that never trips is decoration. Pair these with the operational numbers in your operations metrics so risk indicators sit next to performance ones.

Connect risk to continuity and crisis response

Risk management and business continuity are two halves of the same discipline: one reduces the chance and size of a disruption, the other keeps the business running when one hits anyway. If your risk work stops at the register and never produces tested recovery plans, you have done half the job.

For each high-impact risk, ask: if this happens tomorrow, what do we do in the first hour, and who decides? That answer becomes a continuity plan with defined recovery-time objectives — how long a function can be down before it hurts — and a crisis-communication path for staff, customers, and regulators. The plans that fail on the day are the ones never rehearsed. Build these out with a dedicated business continuity guide and a written crisis communication plan so the response is decided in calm, not invented in panic.

Supply chain deserves its own attention here, because it is the external risk most COOs underweight. Single-source suppliers, thin inventory buffers, and unmapped tier-two dependencies are where a small event upstream becomes your problem. Mapping those dependencies and building alternate sources is the core of supply chain resilience.

Report to the board so they can actually decide

A risk report exists to help the board make decisions, not to demonstrate that you are busy. The strong version is one page: your top five risks, the direction each is trending (better, worse, steady), what you are doing, and the two or three decisions or resources you need from them. The weak version is a 40-slide compliance deck that gets flipped through in silence.

Lead with what changed and what you need. A board does not want the full register — it wants to know which exposures are moving the wrong way and where its help is required (a budget approval, a risk-appetite call, a decision to exit a line of business). Framing the report around decisions rather than status is the single biggest upgrade most operators can make.

Key takeaways

  • Operational risk spans people, process, systems, and external events — if your register is mostly IT, you have a blind spot, not a secure company.
  • Build a short, owned, current risk register first; every control and report hangs off it.
  • Rank by likelihood × impact and give each risk one of four responses: treat, transfer, tolerate, or terminate. Do not treat everything.
  • Layer preventive, detective, and corrective controls; a single control is brittle, and segregation of duties in finance is the highest-value one most firms have.
  • Monitor with Key Risk Indicators that move before the loss, each with a threshold and an owner who acts.
  • Pair risk work with tested continuity and crisis plans — an untested plan fails on the day.
  • Report to the board in one page framed around decisions and what you need, not a compliance deck.

Frequently asked questions

What is the difference between operational risk and enterprise risk management (ERM)? Operational risk is one category — loss from broken processes, people, systems, or external events. ERM is the umbrella that also covers strategic, financial, and compliance risk across the whole organisation. As COO you usually own operational risk directly and feed it into the wider ERM picture the CEO and board oversee. Treating operational risk as if it were the entire ERM program is a common scoping error. How often should a COO review the risk register? Review it monthly with functional owners, reassess more fully each quarter, and run a deep review annually. High × high risks warrant weekly tracking until the control is proven. The failure mode is a register built once for an audit and left to rot; one nobody has opened in six months is not managing anything. Which operational risks do COOs most often underestimate? People risk and single-supplier dependence. Teams pour attention into cybersecurity, then discover that one undocumented person holds the only working knowledge of a critical process, or that a sole supplier has no backup. Both are cheap to reduce once named — cross-training and a runbook for key-person risk, an alternate source for supplier risk — but they rarely get named without a deliberate identification pass. How do you measure whether a risk program is actually working? Look at leading indicators, not just the absence of disasters. Track whether KRIs trend the right way, whether controls pass their tests, how fast incidents are detected and resolved, and whether near-misses are reported at all — a rise in near-miss reports usually means a healthier culture. A program with zero reported incidents is more likely blind than safe. Should smaller companies bother with formal operational risk management? Yes, but proportionately. A smaller firm does not need a GRC platform or a dedicated risk officer — it needs a one-page register of its ten real exposures, a named owner for each, a handful of KRIs, and a tested plan for the two or three things that could stop the business. The discipline scales down; the paperwork should too. How do controls and compliance relate to operational risk? Compliance is a subset of your controls — the ones required by law or regulation rather than by your own risk appetite. Meeting a regulation reduces certain risks but never covers all of them, so a "compliant" company can still be badly exposed. Build controls around your actual risks first, then confirm they also satisfy the compliance management requirements you are bound by.