An AI agent kill switch is a control outside an agent that stops the agent when it misbehaves, and the agent cannot override it. Most kill switches stop the whole agent, which makes them the right last resort and a poor first response to most bad actions.
What does NVIDIA's agent safety launch get right?
On September 28, 2026, NVIDIA announced the Open Agent Safety Platform. Two of its design choices deserve credit.
The first is that enforcement lives outside the agent. NVIDIA OpenShell runs each agent in a sandbox with kernel-level isolation, and a supervisor on the trusted side of the boundary is the agent's only allowed egress path. The agent cannot rewrite the policy that governs it.
The second is an out-of-band watchdog. Sentry runs on NVIDIA BlueField-4 DPUs, and NVIDIA says it can quarantine an agent that tries to move outside its boundaries in milliseconds. Because it runs on separate hardware on the node's path to the model, the agent has no way to argue with it.
We agree with both choices. A control the agent can influence is a suggestion. Our explainer on what NVIDIA OpenShell is covers the platform in detail.
Why is a kill switch the wrong first move?
A kill switch answers one question: should this agent keep running? For most bad actions, that is too blunt a question.
Take a hypothetical customer service agent that has handled 400 refund requests correctly today. Request 401 asks for a refund to an account that does not belong to the customer. Quarantining the agent stops the bad refund. It also stops the good requests queued behind it, interrupts work in progress and leaves someone to investigate and restart the agent.
The problem was one action, so the response should be sized to that action. When I talk this through with our team, I put it this way: "an agentic kill switch is overkill" for most of what goes wrong.
What is a control switch for AI agents?
A control switch decides action by action. Before each action runs, it asks whether this action, from this agent, for this user, makes sense now. Most work continues while the one bad step is constrained, sent to a human or blocked.
We call this category the Agent Execution Control Switch. Ackuity sits beside the agent in the execution path and outside the agent's control. It checks each action against the security context graph for AI agents, which holds 29 signals across six dimensions: user, agent, intent and goal, target system and data, tools and supply chain, and history. It also compares the action with behavioral baselines and checks the sequence of actions against 60+ threat models in 14 categories, mapped to NIST, OWASP and MITRE ATLAS. A decision takes 40 to 100 ms, measured as decision time rather than end-to-end latency.
How does the response ladder work?
Ackuity's response ladder has five decisions, chosen in proportion to severity and confidence:
- Allow: the action is signed, logged and run.
- Constrain: the action is masked or scoped, then run.
- Human in the loop: the action pauses until a person approves it.
- Block: the action is dropped before it reaches the target.
- Terminate: the container comes down, for catastrophic cases only.
Below the block threshold, Ackuity can also alert the SOC, and it logs each action with its query, chain of thought and the data it touched. In the refund example, Ackuity blocks request 401 or sends it to a person, and the agent keeps working through the queue.
Why does a control switch need to see every action?
Per-action decisions only help if the control sees the actions. In the incident covered in our OpenAI and Hugging Face incident timeline, Hugging Face counted about 17,600 actions by rogue agents. They ran through shell commands, raw HTTP and cloud metadata calls, paths an LLM, MCP or API gateway never sees.
That is why placement matters as much as the decision. The control belongs in the Goldilocks zone: not inside the agent, not at the gateway. Beside the agent, in the execution path, it can see intent, plan and history while staying out of the agent's reach. OpenShell's supervisor sits in that kind of position for network traffic, and we are building an integration with its supervisor middleware, described in Ackuity and NVIDIA OpenShell. The integration is in development. Outside OpenShell, Ackuity's own sidecar takes the same position on any Kubernetes cluster.
When should you terminate an agent?
Terminate when the agent itself can no longer be trusted, as opposed to one action being wrong. Signs include poisoned memory, a goal that has been hijacked across many steps, a burst of blocked actions in a short window, or attempts to reach past the agent's boundaries. In those cases a whole-agent stop, whether Ackuity's Terminate or a hardware quarantine such as Sentry, is the proportionate response.
A sound design keeps both. Per-action control handles the everyday mistakes and attacks, and the kill switch stays ready for the rare case that needs it.
NVIDIA, OpenShell and BlueField are trademarks of NVIDIA Corporation.
For the full picture, including where a switch should sit and what to look for, read our guide: What is an AI agent kill switch?