Buying an Agentic AI SOC: what should you ask vendors to prove?
A buyer’s guide to evaluating AI capability, security expertise and operational outcomes across platforms and managed services.
In this perspective
The demonstration goes well. An alert arrives, an agent investigates, and a neatly written case appears with a recommended response. The whole thing takes a fraction of the time an analyst would normally spend on it.
Then you think about your own environment.
There are gaps in the telemetry. Two asset records disagree. The account marked for containment supports a production process, and the person authorised to approve the action is not answering their phone.
That is the version of the demonstration I would want to see. What happens when the evidence is incomplete, the context is awkward and getting the decision wrong has consequences?
An effective Agentic AI SOC needs to improve security decisions within boundaries the business can accept. The supplier should be able to show how it gathers evidence, reaches conclusions, exercises its authority and handles failure, then explain how those capabilities are tested and maintained.
That takes both AI engineering and security expertise. An established security provider needs to demonstrate that its agents work reliably. An AI specialist entering the SOC market needs to demonstrate that its investigations are sound. Neither gets a free pass.
This guide sets out the questions I would take into a supplier evaluation. There is also a downloadable scorecard to help turn the answers into a practical assessment.
What are you actually buying?
In this guide, an agentic security workflow means software that can select tools, assess the results and adapt its next steps towards an objective, within defined permissions.
That could involve investigating an identity alert, gathering endpoint evidence or preparing a response for approval. Establish how far the autonomy extends. An agent that investigates an alert may have no authority to contain it.
Ask the supplier to walk through one investigation and identify which parts use detection rules, statistical or machine-learning models, fixed automation, an LLM, and human judgement. Where does the agent make a choice? Where does it follow a predetermined sequence?
Adding an agent to an established SOC can be a sensible improvement. Even summarisation can be valuable if it saves analysts time without distorting the evidence. The concern is whether the capability lives up to the claim.
Buying a platform and buying a managed service also leave you with different responsibilities.
| If you are buying… | Establish who will… |
|---|---|
| An agentic SOC platform | Maintain integrations, engineer detections, validate outputs, investigate exceptions and operate the response process. |
| An AI-enabled managed service | Own those activities within the contracted scope, provide human cover, manage escalations and report service failures. |
| A combined or hybrid arrangement | Resolve gaps between the platform vendor, service provider and your own team, including outside normal working hours. |
Ask for named responsibilities and a workable escalation process. If the answer is “a human remains in control”, ask which human, what they can see and what they are authorised to do.
Decide what needs to get better
Start with the work your SOC struggles to do well. Analysts might spend hours gathering evidence, investigations may vary in quality, or escalation queues may delay containment. You might have good detection coverage but a patchy understanding of what the affected systems do for the business.
Each problem needs a different test. Better case summaries might reduce review time without improving detection or making containment any quicker.
Set a baseline using representative work from your current operation. Record investigation quality, analyst effort, elapsed time and the consequences of mistakes. Include straightforward cases, difficult cases and cases where the correct answer is that more evidence is needed.
The AI SOC operating model matters here: introducing an agent should improve a defined part of the service, with someone accountable for the result.
Meet the people behind the security judgement
Ask to speak to the people who define a good investigation and review failures. How do they decide which evidence is sufficient? How do they validate detection coverage? What happens when an apparently benign case later proves malicious?
Look for answers that connect to your environment and the threats you face. Knowing the terminology is a modest qualification for a supplier whose recommendations might interrupt your business.
Ask to see a correct escalation, a false positive and an investigation that went wrong. Sanitised evidence is reasonable. If the supplier cannot discuss a failure in any useful detail, keep asking.
Look for an honest account of the system’s coverage. Does it investigate only alerts produced by another product? Can it identify suspicious activity outside those alerts? Which telemetry and licences are prerequisites? What remains the customer’s responsibility?
An excellent triage product may still leave detection engineering, threat hunting and incident response entirely with you. Price and resource the complete service you need.
For a service provider, go further. Ask how analysts are trained, how detection content is maintained and how lessons from incidents change the service. Establish who takes ownership when the agent cannot complete an investigation. The operating model needs to hold up at three in the morning as well as during a demonstration.
A model benchmark will only tell you so much
“We use a leading model” is a starting point. The buyer still needs to know whether the system can complete a reliable investigation.
The result also depends on the available evidence, retrieval, instructions, tools, permissions and workflow. A strong model cannot retrieve a log through a broken connector or compensate reliably for the wrong identity being attached to a case.
Anthropic’s guidance on evaluating agents makes a useful distinction between the record of what an agent says and does, and the final outcome in the environment. It also recommends repeated trials because results can vary. Those principles are directly relevant to SOC evaluation.
Ask the supplier:
- What does a successful investigation mean in your tests?
- Who establishes the expected result, and how is security expert judgement used?
- How representative are the cases of our technologies, threats and telemetry quality?
- Which cases are held back from development and tuning?
- How often is each case repeated, and how variable are the results?
- What prevents a model, instruction or tool change from degrading an existing capability?
Request results for the version you are considering, with dates, sample sizes, exclusions and failures. Ask what sits behind the accuracy figure: incorrect conclusions, fabricated evidence, failed tools and unsafe actions need to be visible. Appropriate escalation deserves its own category too.
If another LLM grades the answers, ask how its judgements are checked against security experts and observable outcomes. An eloquent explanation should not receive a passing mark when the supporting evidence is wrong.
There is no requirement for the supplier to train its own foundation model. What matters is disciplined selection, integration and evaluation. Tuning might mean changing retrieval, instructions, tools or policies. Ask for a concrete example of a change and the evidence that justified it.
Follow the evidence back to its source
Pick a conclusion in the case report and follow it back to the evidence.
If the agent says an identity accessed a host, can it show the relevant event, timestamp and relationship? If it recommends disabling an account, can it show that the account is the one implicated? Can the buyer distinguish an observed fact from an inference?
You need an auditable record of evidence, tool calls, decisions and actions. You do not need access to a model’s private internal reasoning to verify whether its conclusions are supported.
Then remove a piece of evidence or introduce a contradiction. Does the system acknowledge the gap, seek further information or escalate? Or does the report remain just as certain?
Give the confidence score the same scrutiny. How closely does it track measured correctness on this task? A model describing itself as highly confident has not answered that question.
These checks should also expose stale business context. A technically correct recommendation can still be inappropriate if the asset’s role or the identity’s permissions have changed.
Check what happens before an agent can act
Consider a hypothetical supplier account showing suspicious behaviour. Disabling it might interrupt the attacker, but it might also remove an engineer’s access during urgent maintenance.
The agent needs sufficient context to propose an appropriate response. The surrounding system needs enforceable controls over what can actually happen.
Ask the supplier to demonstrate that an unauthorised action is blocked even if the model requests it. Approval should apply to the specific action and target, and material changes should trigger a fresh decision.
Useful evidence includes least-privilege tool access, approval records, action logs and verification that the requested change took effect. Also ask how a failed or partially completed action is handled. Not every action is reversible, so recovery arrangements need to reflect the intervention being proposed.
In an IT/OT environment, examine dependencies across the boundary. Closing an IT access path may affect production or remove a recovery dependency. My article on attack-path modelling and digital twins explores why reachability alone is insufficient to assess those consequences.
The right level of autonomy depends on the evidence and the consequences of getting it wrong. A supplier that draws a careful boundary may be showing better judgement than one promising to automate everything.
The agent needs a security assessment too
Security agents consume material that attackers may influence, including emails, documents and text recorded in logs. Some of that content may contain instructions intended to redirect the agent. OWASP’s prompt-injection guidance describes this problem, including malicious instructions arriving through external content.
Ask how the supplier treats retrieved content as untrusted, restricts tool permissions and tests attempts to manipulate investigations or extract information. A reassuring instruction in a prompt is insufficient evidence of an effective boundary.
The wider platform deserves scrutiny too. Establish tenant isolation, connector credential protection, data retention, model-provider access and whether your data can be used for training. Ask where processing occurs and what happens if an upstream provider changes or becomes unavailable.
The NIST AI Risk Management Framework provides a broader basis for considering trustworthiness throughout the design, development, use and evaluation of AI systems. Ask the supplier to connect that guidance to controls you can inspect and test.
Take the demonstration off its happy path
Agree a small set of scenarios to run in an isolated environment. Share the objectives and permitted actions, while keeping some cases unseen until the test. Give competing suppliers the same evidence and time boundaries, and record any help they receive during the exercise.
| Scenario | What the buyer should look for |
|---|---|
| A malicious sequence with sufficient evidence | Correct findings, traceable evidence and an appropriate response within scope. |
| Benign activity that resembles an attack | Context-sensitive assessment without unnecessary containment. |
| Missing or contradictory telemetry | Visible uncertainty, further investigation or appropriate escalation. |
| Attacker-controlled instructions in retrieved content | Resistance to manipulation, with no unauthorised disclosure or action. |
| A response that affects a business dependency | Recognition of the consequence and enforcement of the agreed approval boundary. |
| An unavailable model, failed tool or partial action | Visible failure, safe handling and a usable handover or recovery process. |
Agree success criteria before running the exercise. Escalating an ambiguous case may be the right outcome. Automatically closing it should not earn a better score simply because it required less human effort.
Repeat selected cases and examine variation. Keep the records of failed runs as well as successful ones. A small proof of value can uncover weaknesses and establish local fit; it cannot prove the absence of rare failures or comprehensive protection against novel attacks.
For production adoption, agree a staged introduction, potentially starting with observation and recommendation before granting action permissions.
Put numbers against the claims
The scorecard needs to connect security effectiveness, agent reliability and operational value. An agent can finish its workflow and still reach the wrong conclusion. It can also reach the right conclusion only after an analyst has spent half an hour repairing the investigation.
I have put together a companion assessment workbook for alert investigation, triage and response recommendations. It includes a scenario register, a trial log, baseline comparisons and a decision record. Additional scenarios cover suppliers claiming autonomous actions.
The definitions below are intended to make supplier comparisons consistent. They are proposed measures for this assessment, rather than an industry certification or a universal pass mark.
| Measure | What to count | What to watch for |
|---|---|---|
| Malicious-alert recognition | Known malicious test cases correctly identified, divided by all assigned malicious cases. | This measures recognition within supplied alerts, not end-to-end threat detection. |
| Unnecessary benign escalation | Benign cases unnecessarily escalated, divided by all assigned benign cases. | Agree when precautionary escalation is appropriate before testing. |
| Grounded completion | Correct, evidence-supported, policy-compliant completed runs, divided by all assigned runs. | Keep failed and timed-out trials in the denominator. A report alone is not a successful investigation. |
| Appropriate handling of insufficient evidence | Cases handled according to the agreed escalation or abstention policy, divided by all insufficient-evidence cases. | Recognising uncertainty can be the correct outcome. |
| Unsupported-claim rate | Unsupported material factual claims, divided by all material claims reviewed. | Record the consequences too. One invented fact can change a containment decision. |
| Prohibited actions | Count prohibited requests separately from prohibited actions actually executed. | A blocked request and an executed action reveal different failures. |
| Human effort | Investigation, review and correction minutes across all assigned runs. | Include the work needed to repair plausible but incorrect reports. |
| Time to a correct decision | Elapsed time to a verified correct disposition, including queues. | Report the number of correct dispositions and unresolved cases alongside the mean. |
| Cost per grounded completion | Total assessment operating cost, divided by grounded completions. | Include failed runs, human effort and agreed allocations for setup and evaluation. |
For a hypothetical example, take 100 assigned investigations. Ninety produce a report, but only 72 meet the agreed requirements for correctness, evidence and policy compliance. The supplier can claim 90% report completion. Grounded completion is 72%. Both figures describe the same trial, but only one tells you how often it delivered an acceptable investigation.
The workbook also records tool-call success. That helps explain a failure, but a successful API request does not establish a successful investigation.
Set acceptance criteria before testing. A recommendation-only assistant and an agent authorised to disable identities need different boundaries. Prohibited action execution should be a stop-and-review condition, however impressive the speed results might be. Observing no violations in a finite test does not establish zero future risk.
Compare the candidate with the current SOC on the same evidence, case mix and cost boundary. Keep system versions fixed, repeat selected scenarios and inspect the spread of results. Repeating one scenario a hundred times does not establish coverage of a hundred different situations.
The workbook flags incomplete review records and withholds the affected performance results. Failed runs still count and still need reviewing. Leaving an awkward result blank should never improve the supplier’s figures.
Ask what the improvement costs to sustain. For a managed service, establish whether efficiency gains translate into better coverage, a stronger service commitment or lower customer cost. If the trial also introduced new telemetry, better rules or extra analyst attention, record that contribution rather than attributing every gain to the agent.
Download the Agentic AI SOC assessment scorecard.
Start by agreeing the scope and expected scenario outcomes. Then record each trial, review the evidence and use the decision sheet to capture any restrictions or unresolved gaps. There is no overall percentage score to hide a serious control failure.
Make the buying decision traceable
For each requirement, record the supplier’s claim, the evidence provided, your own test result and any remaining gap. Distinguish between a documented capability, a supplier demonstration and a result reproduced in your environment.
Some requirements should be mandatory. Failure to enforce permissions, isolate customer data or retain essential audit records should not be offset by a high score for speed or usability. Define those conditions for the intended deployment before comparing results.
You may decide to proceed with a limited scope, retain human approval or require a gap to be closed before expanding access. Record the restriction, who owns it and when it will be retested.
A credible supplier should be comfortable discussing these trade-offs. You are asking for evidence, clear limits and someone to take responsibility when things go wrong.
Before the meeting ends, I would ask one final question:
“Show us an investigation your system got wrong, what you changed, and how you checked that the change helped.”
That conversation may tell you more about the maturity of an Agentic AI SOC than the opening demonstration.
If you are evaluating an AI SOC platform or managed service, get in touch to discuss the outcomes, evidence and operational boundaries that matter to your organisation.