AGENT SECURITY APPROACHES

Where should AI agent security actually sit?

Teams usually put AI agent security in one of 4 places: inside the model or harness, at an AI gateway, around the agent in a sandbox, or beside the agent in its execution path. The table compares them on 4 criteria.

Four places to put AI agent security, compared
ApproachContext depthDetectionIndependence from the agentManageability
Harness and model controlsSees the conversation and the harness's rules inside one agent.Strong on unsafe content, PII in text and jailbreaks; not built to correlate actions across steps.Low: the controls run inside the agent they govern.One policy per harness, so many frameworks means many rulebooks.
AI gateways (MCP, LLM, API)Sees each request that passes through, with little of the user, goal or history behind it.Per-request rules on 3 of 11 action lanes; the other 8 never pass through.High: runs outside the agent, on the network path.Strong: central policy, auth, rate limits and logs in one place.
Agent sandboxes (for example NVIDIA OpenShell)Sees the agent's outbound traffic and its policy; deeper context can come from add-ons via supervisor middleware.The OpenShell supervisor inspects outbound HTTP, GraphQL and MCP traffic against policy.High: the supervisor beside each sandbox is the agent's only way out.Policies set what each agent may touch, enforced by the supervisor.
Ackuity (Agent Execution Control Switch)29 signals across 6 dimensions per action, from user and intent to target and history.60+ threat models in 14 categories, correlated across steps, with a decision in 40 to 100 ms.High: runs beside the agent, outside its control, on context the agent cannot edit.One policy across agents and environments, with a 5-step response ladder.

PLACEMENT

Why does it matter where agent security sits?

Where a control sits decides what it can see and whether the agent can tamper with it. A guard at the front door sees everyone who walks in and nothing that happens upstairs. A guard inside each office sees everything, but reports to the person being watched.

A map of where 4 security approaches sit relative to one agent. 1: Harness and model controls sit inside the agent and are best at content safety, PII in text and jailbreak resistance. 2: An agent sandbox, for example NVIDIA OpenShell, draws a boundary around the agent's process and is best at limiting what the agent can reach. 3: Ackuity sits beside the agent, in its execution path, in the Goldilocks zone: not inside the agent, not at the gateway. It is best at deciding whether each action, right now, should run. 4: AI gateways sit at the perimeter on the LLM, MCP and API calls routed through them and are best at auth, rate limits, central policy and logging. Other actions cross the perimeter without a gateway. Most teams run more than one layer.2AGENT SANDBOXAgentMODEL + HARNESS1Harness andmodel controlsGOLDILOCKS ZONE3Ackuitybeside the agent,in its execution pathnot inside the agent,not at the gatewayPERIMETER4AI gatewayLLM·MCP·APIOTHER ACTIONSTARGETS1INSIDE THE AGENTHarness andmodel controlsBEST ATcontent safety, PIIin text, jailbreaks2AROUND THE PROCESSAgent sandboxfor exampleNVIDIA OpenShellBEST ATlimiting what anagent can reach3BESIDE THE AGENTAckuityexecution control,Goldilocks zoneBEST ATwhether this action,right now, should run4AT THE PERIMETERAI gatewayson calls routedthrough themBEST ATauth, rate limits,policy and logging
Four places to put AI agent security: harness and model controls inside the agent, a sandbox around its process, Ackuity beside it in the execution path, and AI gateways at the perimeter. Each layer is best at a different job, and most teams run more than one.

Agent security has the same trade-off, and most teams run more than one layer.

CRITERION 1

What is context depth?

Context depth is how much a control knows about an action at the moment it decides. A shallow view sees the request: a tool name, a URL, a query. A deep view also knows which user asked, what the agent was trying to do, what it planned, what it did earlier in the session and what the target system holds.

Ackuity's Agent Security Context Graph is built for that deep view. The Agent Security Context Graph is a security context graph for AI agents: everything relevant to a single agent action, across six dimensions and 29 signals, assembled outside the agent before the action runs, so the agent cannot edit it.

CRITERION 2

What does detection mean for AI agents?

Detection is how well a control spots a bad action, including one that looks harmless on its own. Many agent attacks are spread across steps: read a secret, encode it, send it out. Each step passes a per-request rule. Catching the pattern takes correlation across steps and a catalogue of known threats, such as Ackuity's 60+ threat models in 14 categories.

Coverage counts too. Agents act through 11 kinds of action, and only 3 of them (LLM, MCP and API calls) can pass a gateway. Hugging Face counted about 17,600 actions by rogue agents. They ran through shell commands, raw HTTP and cloud metadata calls, paths an LLM, MCP or API gateway never sees. Read the Hugging Face incident timeline.

CRITERION 3

Why should security be independent of the agent?

Independence is whether the agent can influence the control that judges it. A rule inside the agent is read by the same model that decides what to do next, so a clever plan or an injected instruction can route around it. Controls outside the agent, whether at a gateway, in a sandbox supervisor or in a sidecar beside it, don't share the agent's reasoning, so the agent has no say in the decision.

CRITERION 4

What makes agent security manageable?

Manageability is how much work it takes to write, change and audit policy across every agent you run. Gateways score well because one chokepoint means one place to set rules. Harness controls score lower, since each framework keeps its own format.

Proportion matters as well. Ackuity's response ladder can allow an action, constrain it, pause it for human approval, block it, or terminate the container in catastrophic cases. A risky action doesn't have to mean a stopped agent.

LAYERS

How do these approaches fit together?

Guardrails keep model output clean. Gateways handle auth, quotas and central policy at the chokepoint. Sandboxes limit what an agent can reach. Ackuity adds a context-aware decision on each action, made outside the agent, before the action runs.

Ackuity also feeds the other layers. The context it assembles for each action gives gateways, IAM and SIEM tools the who, what and why they lack, and alerts route to the SOC.

Sandboxes and Ackuity answer different questions. OpenShell decides what an agent may touch. Ackuity decides whether this action, right now, should run. Ackuity is building an integration on the OpenShell supervisor's middleware hook, which lets an external service allow, deny or modify agent traffic. The integration is in development; see Ackuity for NVIDIA OpenShell and what NVIDIA OpenShell is. NVIDIA's Sentry can quarantine an agent in milliseconds on BlueField-4. Quarantine is the right last resort and the wrong first move, as we explain in kill switch vs control switch.

WHERE ACKUITY SITS

So where does Ackuity sit?

Ackuity, the Agent Execution Control Switch, sits beside each agent, in its execution path: not inside the agent, not at the gateway. We call that placement the Goldilocks zone. Beside the agent, Ackuity is close enough to see intent, plan and history, and still outside the agent's control.

NVIDIA, OpenShell and BlueField are trademarks of NVIDIA Corporation.

QUESTIONS PEOPLE ASK

Frequently asked questions.

What is the best approach to AI agent security?

A layered one: model guardrails for content, gateways for central policy, sandboxes for containment, and independent execution control for the actions themselves. Ackuity adds the context-aware decision on each action and feeds the other layers the context they lack.

What is an agent sandbox?

An agent sandbox runs an AI agent in an isolated environment and controls what it can reach. In NVIDIA OpenShell, a supervisor beside each sandbox is the agent's only way out and inspects outbound HTTP, GraphQL and MCP traffic against policy. Ackuity's OpenShell integration is in development.

Is Ackuity an AI gateway?

No. Ackuity sits beside each agent in its execution path, away from the network chokepoint where gateways live. From there it sees actions that never cross a gateway, such as CLI commands and SQL queries, along with the agent's goal and history.

What is the Goldilocks zone in AI agent security?

It is the spot beside the agent, in its execution path: not inside the agent, not at the gateway. Ackuity calls that placement the Goldilocks zone. It is close enough to see intent, plan and history, and outside the agent's control, so the agent can't reason around the check.

NEXT STEP

Agents are going to act. Decide which actions run.

Show us the actions you need to control, and we’ll show you where Ackuity fits.

Request access