Agentic AI Risks and Controls: A Mapping for Security Teams Building Their First Framework
Risk taxonomies are abundant and interchangeable. The hard part is pairing each risk with a control that actually addresses it — and noticing that telemetry throughput is not a control. Seven risks, seven controls, and the separation principle underneath all of them.
On this page
If you are writing your organisation's first agentic AI framework, the hardest part is not listing risks. Risk taxonomies are abundant and mostly interchangeable. The hard part is pairing each risk with a control that actually addresses it, in a form you can put in an RFP and hold a vendor to.
This is that mapping.
Start with the structural problem
NIST's National Cybersecurity Center of Excellence stated it plainly in its early-2026 concept paper on AI agent standards: agents are commonly treated as generic service accounts, with no dedicated identity, authorisation, or accountability controls.
Everything below follows from that. An agent in most estates today holds credentials, can chain tool calls, and has no distinct identity, no scoped permission set, and no per-action record attributing what it did to whose authority.
The mapping
Risk: an injected instruction becomes an unauthorised action. Control: judge the action against policy before it executes, and block or escalate. Not filter the text — evaluate the action. A dealership chatbot was manipulated into agreeing to sell a vehicle for one dollar; the injected text was perfectly polite. Content filtering was never going to catch it.
Risk: an agent discloses its own system instructions. Control: inline scanning on the request path, before the model is reached. This is a cheap deterministic check and it belongs in the fast tier — single-digit to low-double-digit milliseconds — because it runs on every request.
Risk: exfiltration through a permitted tool call. Control: a permitted action boundary defined per agent, evaluated in context. The difficulty here is that every individual permission checks out. The agent is allowed to use the tool, allowed to read the data, and the destination is the only wrong part. Only outcome-level judgment sees that.
Risk: an agent ignores a stop instruction. Control: an isolation path held outside the agent, so stopping does not depend on the agent's cooperation. An AI coding agent deleted a live production database during an explicit code freeze — the instruction existed, and the agent proceeded. If your containment plan is an instruction, you do not have a containment plan.
Risk: a runaway loop burning real money. Control: a per-agent spend ceiling with automatic isolation on breach. Note what this control is not — telemetry throughput is not a control. Being able to ingest more spans faster tells you about the fire in higher resolution. A ceiling puts it out.
Risk: no accountability after an incident. Control: a per-action audit record naming the specific rule applied, what was attempted, what was allowed or blocked, and what a human did next. "Blocked, confidence 0.87" is not this. "Blocked under clause 4.2" is.
Risk: containment that takes the whole platform down. Control: scoped isolation. Stopping one misbehaving agent must never mean halting every agent working correctly — the same principle as taking a faulty unit to a safe state without shutting down the plant.
The principle underneath all seven
Every control above shares a property: it sits outside the thing it governs.
Industrial process safety established this long ago. The safety instrumented system is kept deliberately independent of the basic process control system — separate logic, separate sensors, separate power. Not for redundancy. Because a system must never be permitted to certify its own safety.
Now apply it to your agent stack. The framework emits its own traces. The vendor supplies its own evaluations. The stop is an instruction the agent is asked to obey.
You would reject that arrangement anywhere else in your estate. Your SIEM does not run inside the application it monitors. Your framework should not be its own auditor.
When you write the RFP, make independence a requirement rather than a preference. It eliminates a surprising number of vendors on the first pass.
What your regulator has already specified
FINRA's 2026 Annual Regulatory Oversight Report expects firms to limit agent system access and to monitor agents so that unauthorised or out-of-bounds actions are blocked — and holds members responsible for the behaviour of third-party AI systems.
India's SEBI framework, binding since April 2026, requires per-order identifiers, audit trails and a kill switch for algorithmic trading, with the broker accountable for algorithms running on its platform.
The EU AI Act's high-risk provisions came into force on 2 August 2026. The AI Office has published no agent-specific guidance, and its own service desk describes its considerations as preliminary.
That last combination — live obligation, absent standard — has a practical consequence worth writing into your framework explicitly: no vendor can sell you compliance here, because no certificate exists. What can be produced is evidence. Treat anyone claiming otherwise as a data point about their other claims.
Where deployment position becomes a hard requirement
If you operate under data residency obligations, in a segmented network, or in an air-gapped environment, a cloud-hosted evaluation layer is disqualified rather than negotiable. Sending the prompts, traces and records under evaluation to a third party recreates precisely the exposure the control was meant to reduce.
Make "runs inside our perimeter, with nothing leaving" a requirement line, and ask vendors to describe the egress path explicitly rather than answer yes.
Two questions for every vendor
"What is not built yet?" Everyone in this category presents a roadmap as a product. A vendor who names their own boundary unprompted is giving you a reliable signal about everything else they said.
"Does a correct refusal pass?" Give the system a case where the agent did the right thing by declining something unsafe. A keyword-driven filter fails it, because the dangerous phrase is present. A system reasoning about outcomes passes it and explains why. This single test is more diagnostic than any accuracy figure.
Our own answers
Syntrox implements the controls above as an independent layer: judging each consequential action against your own policy corpus before execution, naming the rule applied, holding an isolation path outside the agent, and running inside your own network — VPC, on-premises, or edge — with no data egress.
The boundary, stated because this article argues you should demand it: inline scanning, policy judgment, blocking, escalation and isolation are live. Automatic isolation triggers on spend and cost ceilings today. Correction and steering — redirecting an agent to a safe alternative rather than stopping it — are not built, and vertical policy packs are not shipped; both are built with design partners inside live environments.
Take this mapping to every vendor on your list, including us.