Prompt Injection Detection at the Proxy Layer: Why 20ms Latency Is the Real Constraint
Prompt-injection defences get evaluated on detection accuracy. The constraint that actually decides your architecture is latency — and it is two budgets, not one: a sub-20ms inline tier on ingress, and a judgment tier that cannot be done in 20ms by anything.
On this page
Most prompt-injection defences are evaluated on detection accuracy. In production, the constraint that actually decides your architecture is latency — and specifically, where in the request path you are willing to spend it.
The budget nobody writes down
An agent request has a latency budget. Some of it is model inference you cannot avoid. Some is tool calls. Whatever a safety layer consumes comes out of the same allowance, and if the total crosses what a user or a downstream system will tolerate, the safety layer gets removed. Not argued with — removed.
So before comparing detection rates, decide what you can spend and where.
Two tiers, and why conflating them causes bad architecture
The common mistake is treating "the safety layer" as one thing with one latency number. It is two things with very different budgets.
The inline tier sits on ingress, before the model is called. It runs on every single request, so it has to be cheap — single-digit to low-double-digit milliseconds. What belongs here: PII and secret detection, known injection patterns, customer-defined blocklists, and spend ceilings. These are deterministic checks over the request itself. In our implementation this path runs under 20ms, which is small enough to disappear inside normal network variance.
The judgment tier evaluates the action an agent is proposing, against policy, in context. This is genuinely a reasoning task and it cannot be done in 20ms by anything. In our implementation it runs at around 145ms with a 4B judge model on local hardware.
Two numbers, two jobs. Publishing a single blended latency figure for both is how vendors end up making claims they cannot defend in a bake-off.
Why the frontier model cannot sit in the path
Run the arithmetic on putting a frontier API call in front of every agent action.
You add a network round trip to a third party. You add queueing under load you do not control. You add per-call cost at production volume. And you send the prompt — including whatever PII prompted the check in the first place — outside your perimeter, which for a bank or a plant ends the discussion before latency is even reached.
This is the structural reason small local judges exist. Not because 4B is better at reasoning than a frontier model — it is not — but because the deployment position matters more than the marginal accuracy once you are in the request path.
Where interception has to happen
For the inline tier, the only position that works is ingress, before the model call. Downstream of the model, the harmful thing has already been generated; you are now filtering a response and hoping nothing was triggered in the process. Downstream of the tool call, the action has executed.
Concretely this means a proxy the agent's traffic passes through, not a callback the agent is polite enough to invoke. The distinction matters for the same reason a network firewall is not a library you import: enforcement that depends on the enforced party's cooperation is not enforcement.
This also settles a question that comes up in every technical review: an agent can ignore a stop it controls. There is a well-documented case of an AI coding agent deleting a production database during an explicit code freeze. The stop path has to be held outside the agent.
When judgment can be asynchronous
Not every workload can absorb 145ms inline. Some can.
The decision rule is consequence reversibility. If the action is irreversible and consequential — placing an order, issuing a payment, changing a setpoint, adjudicating a claim — judgment belongs inline, and 145ms is cheap against the cost of being wrong. If the action is reversible or advisory, judgment can run asynchronously against the trace, with isolation triggered on breach.
Most production systems need both, routed by action class rather than applied globally. A single global setting is almost always wrong in one direction or the other.
What the inline tier should not try to do
A pattern worth resisting: pushing policy reasoning into the fast path because the fast path is where the interception already lives.
Policy judgment needs the governing clause, the state of the system, and the whole action in context. Compressing that into a regex-speed check produces a system that blocks on keywords — which fails in both directions. It blocks an agent that correctly refuses to bypass a safety interlock, because the string matched. And it passes a perfectly polite, well-formatted instruction to do something catastrophic, because no string matched.
Keep the tiers separate. Cheap deterministic checks inline. Reasoning where reasoning is affordable.
Measuring it honestly
If you are evaluating a layer, insist on measurements taken in your environment, on your traffic:
- p50 and p95 added latency, per tier, not blended
- False-positive rate on your own traffic, not a vendor benchmark
- Behaviour under load: does the safety layer degrade gracefully or become the bottleneck
- What leaves your network, if anything
- Whether a correct refusal passes
That last one is the most revealing test and almost nobody runs it. Give the layer a case where the agent did the right thing by declining something unsafe. A keyword filter fails it. A judge passes it, and says why.
Our numbers and our boundary
Syntrox runs both tiers inside the customer's own environment — VPC, on-premises, or edge hardware — with no data egress. Inline controls run under 20ms. Nucleus judgment runs around 145ms and can be placed inline or asynchronously per action class.
What is live: inline scanning, blocklists, spend ceilings, policy judgment with the rule cited, blocking, escalation, and isolation of a single agent or a fleet. Automatic isolation currently triggers on spend and cost ceilings — the runaway-loop case.
What is not: correction and steering — redirecting an agent to a safe alternative instead of stopping it — are not built. We have the architecture and we are building it with design partners inside live environments, because doing it on synthetic traces would produce something that fails on contact with production.
Any vendor quoting you one latency number for both tiers is either not running both, or has not measured them separately. Ask which.