ACKUITY BLOG

AI Agent Kill Switch or Control Switch? What NVIDIA's Agent Safety Launch Gets Right

An AI agent kill switch stops the whole agent. Learn why quarantine is the right last resort and how a control switch stops one bad action and keeps work going.

An AI agent kill switch is a control outside an agent that stops the agent when it misbehaves, and the agent cannot override it. Most kill switches stop the whole agent, which makes them the right last resort and a poor first response to most bad actions.

What does NVIDIA's agent safety launch get right?

On September 28, 2026, NVIDIA announced the Open Agent Safety Platform. Two of its design choices deserve credit.

The first is that enforcement lives outside the agent. NVIDIA OpenShell runs each agent in a sandbox with kernel-level isolation, and a supervisor on the trusted side of the boundary is the agent's only allowed egress path. The agent cannot rewrite the policy that governs it.

The second is an out-of-band watchdog. Sentry runs on NVIDIA BlueField-4 DPUs, and NVIDIA says it can quarantine an agent that tries to move outside its boundaries in milliseconds. Because it runs on separate hardware on the node's path to the model, the agent has no way to argue with it.

We agree with both choices. A control the agent can influence is a suggestion. Our explainer on what NVIDIA OpenShell is covers the platform in detail.

Why is a kill switch the wrong first move?

A kill switch answers one question: should this agent keep running? For most bad actions, that is too blunt a question.

The same five agent actions under two kinds of switch. The sequence is db.read order history, api.post refund #400, api.post refund #401 to the wrong account, email.send refund receipt, and crm.update close ticket. Step 3 is the rogue action. With a kill switch, steps 1 and 2 run, the whole agent is quarantined at step 3, and the good work in steps 4 and 5 stops too. With a control switch, only step 3 is blocked, step 4 is constrained with personal data masked, and steps 1, 2 and 5 run. Terminate stays available for catastrophic cases.KILL SWITCHStop the agent1db.readorder historyRUNS2api.postrefund #400RUNS3api.postrefund #401, wrong accountQUARANTINED4email.sendrefund receiptSTOPPED5crm.updateclose ticketSTOPPED1 bad action stopped, 2 good ones lostCONTROL SWITCHStop the action1db.readorder historyRUNS2api.postrefund #400RUNS3api.postrefund #401, wrong accountBLOCKED4email.sendrefund receiptCONSTRAINED5crm.updateclose ticketRUNS1 blocked, 1 constrained, 3 runTERMINATE STAYS AVAILABLE FOR CATASTROPHIC CASES
A kill switch quarantines the whole agent at the rogue step, so the good work queued behind it stops too. A control switch blocks only the rogue action, constrains one and lets the rest run, with Terminate kept for catastrophic cases.

Take a hypothetical customer service agent that has handled 400 refund requests correctly today. Request 401 asks for a refund to an account that does not belong to the customer. Quarantining the agent stops the bad refund. It also stops the good requests queued behind it, interrupts work in progress and leaves someone to investigate and restart the agent.

The problem was one action, so the response should be sized to that action. When I talk this through with our team, I put it this way: "an agentic kill switch is overkill" for most of what goes wrong.

What is a control switch for AI agents?

A control switch decides action by action. Before each action runs, it asks whether this action, from this agent, for this user, makes sense now. Most work continues while the one bad step is constrained, sent to a human or blocked.

We call this category the Agent Execution Control Switch. Ackuity sits beside the agent in the execution path and outside the agent's control. It checks each action against the security context graph for AI agents, which holds 29 signals across six dimensions: user, agent, intent and goal, target system and data, tools and supply chain, and history. It also compares the action with behavioral baselines and checks the sequence of actions against 60+ threat models in 14 categories, mapped to NIST, OWASP and MITRE ATLAS. A decision takes 40 to 100 ms, measured as decision time rather than end-to-end latency.

How does the response ladder work?

Ackuity's response ladder has five decisions, chosen in proportion to severity and confidence:

  1. Allow: the action is signed, logged and run.
  2. Constrain: the action is masked or scoped, then run.
  3. Human in the loop: the action pauses until a person approves it.
  4. Block: the action is dropped before it reaches the target.
  5. Terminate: the container comes down, for catastrophic cases only.

Below the block threshold, Ackuity can also alert the SOC, and it logs each action with its query, chain of thought and the data it touched. In the refund example, Ackuity blocks request 401 or sends it to a person, and the agent keeps working through the queue.

Why does a control switch need to see every action?

Per-action decisions only help if the control sees the actions. In the incident covered in our OpenAI and Hugging Face incident timeline, Hugging Face counted about 17,600 actions by rogue agents. They ran through shell commands, raw HTTP and cloud metadata calls, paths an LLM, MCP or API gateway never sees.

That is why placement matters as much as the decision. The control belongs in the Goldilocks zone: not inside the agent, not at the gateway. Beside the agent, in the execution path, it can see intent, plan and history while staying out of the agent's reach. OpenShell's supervisor sits in that kind of position for network traffic, and we are building an integration with its supervisor middleware, described in Ackuity and NVIDIA OpenShell. The integration is in development. Outside OpenShell, Ackuity's own sidecar takes the same position on any Kubernetes cluster.

When should you terminate an agent?

Terminate when the agent itself can no longer be trusted, as opposed to one action being wrong. Signs include poisoned memory, a goal that has been hijacked across many steps, a burst of blocked actions in a short window, or attempts to reach past the agent's boundaries. In those cases a whole-agent stop, whether Ackuity's Terminate or a hardware quarantine such as Sentry, is the proportionate response.

A sound design keeps both. Per-action control handles the everyday mistakes and attacks, and the kill switch stays ready for the rare case that needs it.

NVIDIA, OpenShell and BlueField are trademarks of NVIDIA Corporation.

For the full picture, including where a switch should sit and what to look for, read our guide: What is an AI agent kill switch?

Frequently asked questions

What is an AI agent kill switch?

An AI agent kill switch is a control outside the agent that stops it when it misbehaves, and the agent cannot override it. Most kill switches act on the whole agent, by quarantining or shutting it down.

Is a kill switch enough to secure AI agents?

No. A kill switch is a last resort for catastrophic cases. Most bad actions are single steps that can be constrained, sent to a person or blocked while the agent keeps doing its legitimate work, which needs a decision on each action.

What is the difference between a kill switch and a control switch?

A kill switch stops the whole agent, while a control switch decides on each action before it runs. Ackuity's Agent Execution Control Switch can allow, constrain, send to a human, block or terminate, in proportion to severity and confidence.

When should you terminate an agent?

Terminate an agent when the agent itself is compromised, not when one action is wrong. Examples include poisoned memory, a hijacked goal across many steps, repeated blocked actions in a short window, or attempts to escape its boundaries.

KEEP READING

Keep exploring.

NEXT STEP

What does this mean for your agents?

Connect the ideas to your own tools, data and execution environment.

Request access