The cost of acting. The cost of waiting.
Cyber containment decisions, business interruption and the value of better evidence.
In this perspective
The security team wants to disconnect a supplier. Operations wants to keep the connection open. Both have a good reason.
One sees a route an attacker could use to reach the plant. The other sees the maintenance access that keeps it running. Close the connection and the interruption may start immediately. Leave it open and the business carries an uncertain loss that could be much larger.
Which decision would you sign off?
It is an uncomfortable question because the costs arrive on different terms. The production manager can point to the work that will stop. The security team has to explain what might happen, how likely it is and how much confidence to place in that assessment.
A faster answer is useful only if we understand what we are authorising. Much of the work needed to reach that answer should happen before the alert arrives.
Closing the connection is only one option
In the Cyber Capital Lab, a fictional manufacturer discovers that a supplier-access gateway is affected by a newly disclosed vulnerability. Further investigation confirms an attack path and identifies a maintenance dependency.
Disconnecting the gateway reduces the modelled cyber exposure, but interrupts production. A narrower restriction preserves an approved maintenance route, provided the team can establish and enforce the right boundaries.
The first article, “The cyber budget is agreed. The risk has changed. Now what?”, explored how new evidence changes the financial picture. This time, I want to follow that through to the decision: when should we act, what is worth investigating first, and how long can we reasonably wait?
The choice is rarely just “disconnect” or “do nothing”. Restricting an identity, revoking a session or applying a temporary network rule may buy time for a controlled change. Extra monitoring may help the team spot activity during that period, although it does not remove the exposure. Each option needs its own assessment.
The options also have different dependencies. A proposed restriction may rely on accurate identity ownership. A recovery plan may rely on the same remote access that containment would remove. You may be able to reverse a configuration change in minutes. You cannot necessarily recover the production lost while it was in place.
Those details belong in the decision, even when they make the meeting less comfortable.
What does “wait” actually mean?
The cost of acting now might include engineering effort, a production interruption and the work required to restore normal operation. The cost of waiting includes the exposure retained during the delay and any consequences of responding later.
“Wait for a patch” leaves a lot unsaid. How long do we expect to wait? What protection remains in place? What happens if the fix is delayed or fails testing? At what point do we choose another response?
A 72-hour risk estimate cannot simply be divided by 72 to produce a reliable hourly cost of delay. That would assume the risk is distributed uniformly through the period. Exploit availability, attacker activity and the state of the environment may change while we wait.
The same applies to the intervention. A control that takes a day to deploy does not provide its full benefit from the moment someone approves it.
Before comparing the figures, I would want answers to five questions:
| Decision input | What needs to be established |
|---|---|
| Exposure during the delay | What could happen before the intervention becomes effective, and under which assumptions? |
| Time to effective action | How long until the control is deployed and its effect verified? |
| Operational consequence | Which services stop or degrade, and what does that mean for the business? |
| Recovery | What is needed to return to an acceptable operating state? |
| Review trigger | What evidence, deadline or business change causes us to reconsider? |
We may still disagree about the estimates. But we can now see what each option involves, over the same period.
What would another hour of investigation buy us?
“We need more evidence” sounds reasonable. Sometimes it is exactly right. Sometimes it is a difficult decision being passed to the next meeting.
I would want to know what the missing information could change.
For the supplier-access scenario, the first question is whether the route described in the threat intelligence actually exists here. The affected software version matters, but so do the permissions, configuration and active sessions that make the route usable. A line between two assets is only the beginning of that investigation.
The next question is what depends on the route. The supplier identity might reach an engineering workstation through a gateway and jump host. Closing the gateway could remove that access. It could also disconnect the engineer or recovery service we need to keep production running.
Attack-path mapping and digital-twin simulation can help answer those questions. I explored the technical foundations in Before the SOC acts: attack-path modelling and digital twins across IT and OT. Here, the question is how those findings change the financial comparison.
Test the intervention, then price the consequences
Imagine comparing two responses: closing the supplier connection completely, or restricting it to an approved maintenance session with narrower permissions.
An attack-path model can help us test whether each change removes the relevant route. A simulation with sufficient operational detail can help us investigate which services would be affected, how long an interruption might last and what recovery would require.
A firewall change might break the attack path and still stop a legitimate service. To assess that consequence, the simulation needs the relevant operational dependencies, with a clear account of how current and complete the information is.
We can then carry the findings into the financial assessment:
| What we investigate | What it could change in the financial comparison |
|---|---|
| Whether the attack prerequisites exist | The likelihood assumptions and their uncertainty range |
| Whether a proposed restriction blocks the relevant route | The estimated exposure remaining after intervention |
| Which business services depend on the affected access | The scope and cost of interruption |
| Whether recovery remains possible after containment | The restoration time, effort and potential additional loss |
| Where the evidence is stale or incomplete | The range of outcomes we should test before relying on the recommendation |
Suppose the initial estimate allows two hours for the interruption. The dependency review suggests eight: the supplier whose access we plan to remove is also needed to restore the service.
The attack path has not changed. The cost of closing it has. That could be enough to change the decision.
Or the investigation might show that a narrower restriction preserves maintenance access while removing the route under consideration. The financial comparison can then reflect both the reduced exposure and the lower interruption cost, subject to the restriction working as intended.
Before accepting the preferred option, I would push it a little. Let restoration take longer. Reduce the assumed effectiveness of the restriction. Does the recommendation still hold, or was one optimistic assumption doing all the work?
The figures remain estimates, conditional on the evidence and assumptions behind them. A digital twin can help test particular scenarios; a reliable probability of attack and a complete estimate of business impact require further justification. Each financial assumption needs that connection to the evidence. Costs need checking too, so the same interruption does not appear twice in the total.
That brings us back to the value of another hour. If checking a maintenance dependency could change the response, it may be worth investigating while a temporary restriction provides protection. If a further report merely repeats the vulnerability’s severity, it may add little.
What decision could this evidence change, how soon can we obtain it, and what protection do we have while we look?
More information is not automatically more valuable. It may also reveal that the exposure is worse than we thought.
The current Cyber Capital Lab illustrates how a financial estimate changes as evidence arrives. Its playback is not an elapsed-time model of an attacker progressing through the environment. Estimating the cost of a real delay would require additional assumptions and evidence.
The approval process starts before the alert
Automatic containment makes sense where the action, its scope and its consequences are sufficiently understood. That confidence has to come from preparation: knowing which assets and identities are eligible, what evidence is required, which dependencies need checking and what happens if the action fails.
A wider or less familiar intervention may need someone to weigh the consequences and authorise it. That person needs both the information and the authority to make the call.
“Human approval required” is not a complete operating model. Who is the person? Can they be reached? What are they authorised to decide? What happens if the identity service or communications channel they need is unavailable?
A team can detect an attack in seconds and still wait a long time for an answer to those questions.
NIST’s incident response guidance places response within wider cybersecurity risk management. That is a useful starting point for treating preparation, business context and recovery as part of the response capability.
For me, the strongest machine-speed defence proposition is one that can show which decisions have already been made, which conditions make them valid, and how it recognises when those conditions no longer hold.
Who pays, and when?
Insurance adds another question to the decision: how much of the loss will the business ultimately bear?
In the Lab, assumed insurance recoveries apply to the modelled cyber losses. The separate costs of the chosen containment action and its planned production interruption remain with the business. That is an illustrative modelling choice, not a conclusion about how a particular policy would respond.
For an actual decision, I would want the relevant insurance assumptions reviewed in advance with the people responsible for the policy. Which response and interruption costs might be recoverable? What conditions or approvals need to be considered? What remains uncertain?
Then there is timing. The business may need cash, engineers and supplier capacity today, while any reimbursement arrives later. Moving a cost to an insurer does not bring the production line back into service.
During an incident, a model should keep an unresolved coverage question visible. It should not convert a policy summary into guaranteed reimbursement. Nor should the team have to improvise its insurance escalation process while deciding how to contain an attack.
The clock should follow the business outcome
Detection and response times remain useful measures, provided we define their start and end points. A single average can hide the part of the process that needs attention.
I would separate the time taken to:
- Detect and recognise the relevant activity.
- Assemble enough evidence to make the decision.
- Obtain any required authorisation.
- Implement and verify containment.
- Restore the affected business service.
That makes it easier to see whether the constraint is detection, investigation, ownership, execution or recovery.
Speed belongs alongside unnecessary interruption, failed containment and recovery consequences. A response that is five minutes faster but interrupts the wrong service has left the business with a problem the headline metric will not explain.
Give the agent something useful to challenge
An AI agent could help with the work between receiving new evidence and deciding what it means for the model.
It could bring together threat intelligence, validated attack-path findings, dependency records and simulation results, then identify which financial assumptions need reviewing. Each proposed change should show its source, the limits of the evidence and the scenarios it affects. If two sources disagree, it should make that disagreement visible and identify who can resolve it.
A useful output might say: “The proposed isolation affects a maintenance dependency. The current cost estimate assumes two hours of interruption, but the operations record specifies eight. Here is the source, the proposed model change and its effect on the comparison.”
It could also propose a sensitivity check: would the decision change if recovery took longer, or if the restriction left more exposure than expected? That gives the decision-maker something concrete to review before accepting a revised model.
The agent should not turn missing telemetry into evidence of safety, invent a probability from a vulnerability description or treat its own confidence as a measure of real-world attack likelihood. Mathematical consistency and well-supported inputs are separate requirements.
I would keep the calculations in a deterministic tool and require approval for changes to the accepted risk assumptions. Authority to update a working model would also remain separate from authority to execute containment or accept the residual risk.
That is a capability we would need to test. I would want to know which dependencies the agent missed, which changes it could substantiate and when it recognised that the evidence was insufficient. A persuasive recommendation is easy to overvalue when you are under pressure.
Make waiting a decision with an expiry
Sometimes the evidence justifies immediate action. Sometimes a narrower restriction buys useful time. Either way, “monitor the situation” leaves too much open.
If we wait, someone should own that decision. The team should know what protection remains, what it expects to learn and when it will reconsider. A deadline or a specific change in the evidence gives that review a purpose.
The same discipline applies after acting. Was the control effective? Did the expected operational consequence occur? Has the intervention removed access that recovery now needs?
You can explore both sides of the decision in the Playground.
Explore the attack-path and digital-twin simulation to consider the technical routes and operational dependencies behind a containment decision.
Then explore the Cyber Capital Lab to compare the financial trade-offs. Start at “Dependency found”, compare broad isolation with restricting supplier access, then switch to peak production.
The demonstrations are separate and simplified, with no data exchanged between them. Use them to explore how the technical and financial questions connect: what can we interrupt, what else depends on it, and how does that change the decision?
Before the next alert arrives, it is worth asking: who could authorise that supplier disconnection, and what would they need to know?