AI AGENT KILL SWITCH

What is an AI agent kill switch?

An AI agent kill switch is an independent control that stops a harmful agent action before it executes. The agent has no influence over the decision, because the switch sits outside its reach. Where you put that switch decides what it can see, and what it can stop.

WHY NOW

Why do AI agents need a kill switch now?

AI agents need a kill switch because they now take actions on real systems, and some of those actions go wrong faster than a person can step in. A chatbot that says something wrong produces a bad sentence. An agent that does something wrong can delete a table, email a customer list or pull credentials from a cloud metadata endpoint, all in the time it takes you to read this line.

Agents go rogue for ordinary reasons. A web page carries hidden instructions. A tool description has been tampered with. A goal drifts over 30 steps until the agent is solving a problem nobody asked it to solve. All it takes is an agent with access and a plan that went sideways.

The clearest public example so far is the incident covered in our OpenAI and Hugging Face incident timeline. Hugging Face counted about 17,600 actions by rogue agents. They ran through shell commands, raw HTTP and cloud metadata calls, paths an LLM, MCP or API gateway never sees. That detail matters for everything that follows, because a switch can only stop what it can see.

OPTION 1

Can't the agent police itself?

Controls built into the model or the agent harness help, but they cannot act as a kill switch, because the controlled party is also the control. System prompts, model guardrails and harness policies all live inside the thing they are meant to restrain.

Picture a teenager handed the house rules and the only key to the front door. A clever one reads the rules for what they leave out. Agents do something similar: they reason, and an agent that can read its own rules can also reason its way around them, especially when a prompt injection has given it a new goal.

There is a practical limit too. Each harness carries its own policy, so a company with agents on three frameworks has three rulebooks and no single view across systems or agents. These tools still earn their place. Ackuity, the Agent Execution Control Switch, runs alongside them, and our comparison of Ackuity and model guardrails covers how the two fit.

OPTION 2

Why not put the switch at the gateway?

A gateway can stop the agent traffic that passes through it, but most agent actions never pass through one. LLM, MCP and API gateways sit on a few well-marked roads, and agents also use the back roads.

An agent on the left takes 11 kinds of action, drawn as lanes flowing right. Ackuity sits beside the agent and every lane passes through it, so it checks all 11. A gateway further along crosses only 3 lanes: LLM calls, MCP tool calls, API calls. The other 8 never meet a gateway: CLI commands, HTTP probes, SQL queries, IMDS metadata calls, Credential searches, Memory writes, RAG retrieval, A2A hand-offs.LLM callsMCP tool callsAPI callsCLI commandsHTTP probesSQL queriesIMDS metadata callsCredential searchesMemory writesRAG retrievalA2A hand-offsAGENTAgentACKUITY11 of 11GATEWAY3 actionspass a gateway8 actionsnever meet oneA GATEWAY SEES 3 OF 11. ACKUITY, BESIDE THE AGENT, CHECKS ALL 11.
Of the 11 kinds of action an agent can take, only LLM, MCP and API calls pass a gateway, and the other 8 never meet one. Ackuity sits beside the agent, so it checks all 11.

Think of a toll booth on a highway. It sees every car that drives through it and none of the cars that take the side streets. Of the 11 kinds of action an agent can take, 3 pass a gateway: LLM calls, MCP tool calls and API calls. Eight never meet one:

  • CLI commands
  • HTTP probes
  • SQL queries
  • IMDS metadata calls
  • Credential searches
  • Memory writes
  • RAG retrieval
  • A2A hand-offs

Even for traffic it does see, a gateway sees the call, not the context. It knows a request went to a database. It does not know which user asked, what the agent was told to do or what it did five steps earlier. Its rules look at one request at a time, so an attack spread across several harmless-looking steps slips past. Gateways, identity tools and SIEMs remain useful, and Ackuity feeds them the context they are missing. Our page on Ackuity and AI gateways goes into detail.

PLACEMENT

So where should the switch sit?

An AI agent kill switch should sit beside the agent, in the execution path: not inside the agent, not at the gateway. Close enough to see the agent's intent, plan and history, and far enough away that the agent has no control over it.

Three possible positions for a kill switch around one agent. Position A, inside the agent, covers harness and model guardrails such as system prompts and policies. It sees the agent's reasoning, but the agent can reason around its own rules, so it is too close. Position B, beside the agent in the execution path, is where Ackuity sits: the Goldilocks zone, not inside the agent and not at the gateway. It sees intent, plan and history, sits outside the agent's control, and every action passes through it. Position C, at the gateway, sees only the LLM, MCP and API calls that cross it, without context, while shell, SQL and cloud metadata calls take other paths around it, so it is too far.AGENTAgentHarness guardrailsAACKUITYControl switchBLLM, MCP, APIShell, SQL, IMDSCGATEWAYAINSIDE THE AGENTToo closeSEESPrompts, policies andthe agent's reasoningLIMITThe agent can reasonaround its own rulesBBESIDE THE AGENTGoldilocks zoneSEESIntent, plan and historyfor every actionPOSITIONOutside the agent's control,in the execution pathCAT THE GATEWAYToo farSEESOnly the calls thatcross it, without contextLIMITMost actions takeother paths
A kill switch inside the agent sees its reasoning, but the agent can reason around it, and a gateway sees only the calls that cross it, without context. Ackuity sits beside the agent in the execution path, where it sees intent, plan and history and stays outside the agent's control.

We call that spot, not inside the agent and not at the gateway, the Goldilocks zone. Inside the agent is too close, because the agent can reason around its own rules. Out at the gateway is too far, because most actions never pass and the context is missing. Beside the agent, every action has to go through the switch before it reaches a shell, a database or another agent, and the switch has the full story when it decides.

KILL OR CONTROL

Stop the agent, or stop the action?

A kill switch stops the whole agent. A control switch stops the one action that is wrong and lets the rest of the work continue. Most incidents call for the second.

The same five agent actions under two kinds of switch. The sequence is db.read order history, api.post refund #400, api.post refund #401 to the wrong account, email.send refund receipt, and crm.update close ticket. Step 3 is the rogue action. With a kill switch, steps 1 and 2 run, the whole agent is quarantined at step 3, and the good work in steps 4 and 5 stops too. With a control switch, only step 3 is blocked, step 4 is constrained with personal data masked, and steps 1, 2 and 5 run. Terminate stays available for catastrophic cases.KILL SWITCHStop the agent1db.readorder historyRUNS2api.postrefund #400RUNS3api.postrefund #401, wrong accountQUARANTINED4email.sendrefund receiptSTOPPED5crm.updateclose ticketSTOPPED1 bad action stopped, 2 good ones lostCONTROL SWITCHStop the action1db.readorder historyRUNS2api.postrefund #400RUNS3api.postrefund #401, wrong accountBLOCKED4email.sendrefund receiptCONSTRAINED5crm.updateclose ticketRUNS1 blocked, 1 constrained, 3 runTERMINATE STAYS AVAILABLE FOR CATASTROPHIC CASES
A kill switch quarantines the whole agent at the rogue step, so the good work queued behind it stops too. A control switch blocks only the rogue action, constrains one and lets the rest run, with Terminate kept for catastrophic cases.

Take a hypothetical support agent that has handled 400 refunds correctly today. Refund 401 goes to an account that does not belong to the customer. Shutting the agent down stops refund 401. It also stops every good request queued behind it and leaves someone to investigate and restart. Blocking refund 401, or sending it to a person, fixes the actual problem.

Quarantine still has a job. It is the right last resort and the wrong first move. NVIDIA's Open Agent Safety Platform, launched on September 28, 2026, is a good market example. Its Sentry component runs on BlueField-4 hardware, and NVIDIA says it can quarantine an agent that tries to move outside its boundaries in milliseconds. NVIDIA OpenShell, part of the same platform, shows the placement idea: each agent runs in a sandbox, and a supervisor beside each sandbox is the agent's only way out, inspecting outbound HTTP, GraphQL and MCP traffic against policy before it leaves. Our explainer on what NVIDIA OpenShell is covers the platform, and kill switch vs control switch covers the trade-off in more depth.

Ackuity is an NVIDIA Inception member and is building an integration with OpenShell's supervisor middleware. The integration is in development; details are on our NVIDIA OpenShell integration page.

HOW IT DECIDES

How does an Agent Execution Control Switch decide?

An Agent Execution Control Switch checks each agent action against its full context and a library of known attack patterns, then picks a response sized to the risk, all before the action runs. Ackuity is the Agent Execution Control Switch AI builders add to their agents, verifying every action before it runs.

The context comes from the Agent Security Context Graph, which assembles 29 signals across six dimensions for every action: the user, the agent, its intent and goal, the target system and data, its tools and supply chain, and its history. The context graph is built outside the agent, so the agent cannot edit it.

Detection is neurosymbolic. Rules, behavioral baselines and small language models each look at the action, and their findings are correlated across steps in real time. The checks cover 60+ threat models in 14 categories, mapped to NIST, OWASP and MITRE ATLAS, from memory poisoning to MCP threats.

Then the response ladder picks one of five decisions, in proportion to severity and confidence:

  1. Allow: signed, logged and run.
  2. Constrain: data masked or scope narrowed, then run.
  3. Human in the loop: paused until a person approves it.
  4. Block: dropped before it reaches the target.
  5. Terminate: container shut down, kept for catastrophic cases.

Below the block threshold, Ackuity can alert your SOC, and it logs each action with its query, chain of thought and the data it touched. A decision takes 40 to 100 ms. That figure is decision time, measured separately from end-to-end latency.

DEPLOYMENT

How do you add one without rewriting your agent?

You add Ackuity without changing agent code, through one of three ways in:

  • Sidecar: an open source sidecar runs in the agent's pod, and an init container reroutes its traffic. The agent never knows. This gives the deepest coverage.
  • API injection: for Copilot Studio and similar platforms where you cannot run a sidecar.
  • Event pull: Ackuity reads events from OpenTelemetry, Langfuse or LangSmith. It is observe-only and the fastest way to start.

You choose per policy whether the sidecar fails open or fails closed, and a sidecar failure affects one pod, not your whole estate. Your data stays in your own cloud account. For stricter requirements, Ackuity can hold the agent's credentials and run the action on its behalf, handling mTLS. The full list of connectors is on our integrations page.

CHECKLIST

What to look for in an AI agent kill switch

  1. It runs outside the agent, so the agent cannot edit, disable or argue with it.
  2. It sees every kind of action, including shell commands, SQL queries, metadata calls and A2A hand-offs that never reach a gateway.
  3. It decides with context: who the user is, what the agent was asked to do and what it has already done.
  4. It decides before the action runs, with a decision time you can measure.
  5. It responds in proportion, able to constrain, pause or block one action, with whole-agent shutdown kept for catastrophic cases.
  6. It has a failure mode you choose, open or closed, and a failure stays contained.
  7. It keeps a record of each decision and its evidence, and that record stays in your environment.

NVIDIA, OpenShell and BlueField are trademarks of NVIDIA Corporation.

QUESTIONS PEOPLE ASK

Frequently asked questions.

What is an AI agent kill switch?

An AI agent kill switch is an independent control that stops a harmful agent action before it executes, and the agent has no influence over the decision. Traditional kill switches stop the whole agent. A control switch, such as Ackuity's Agent Execution Control Switch, acts on each action and keeps whole-agent shutdown for catastrophic cases.

How do you stop an AI agent from taking a harmful action?

You stop a harmful agent action by checking it before it runs, from a control the agent cannot influence. That control needs to sit beside the agent in the execution path, see every kind of action, weigh the action against context such as the user, the goal and recent history, and then allow, constrain, pause or block it.

Is a kill switch enough?

A kill switch on its own is too blunt for most incidents, because stopping the whole agent also stops all the good work it was doing. Most problems come down to one bad action, so you need per-action control first, with whole-agent shutdown or quarantine kept as the last resort.

Where should an agent kill switch sit?

An agent kill switch should sit beside the agent, in the execution path: not inside the agent, not at the gateway. Ackuity calls this the Goldilocks zone. Inside the agent, the agent can reason around its own rules. At a gateway, most actions never pass and the context is missing.

Can an MCP gateway act as a kill switch?

An MCP gateway can block MCP tool calls that pass through it, but it cannot act as a full kill switch because most agent actions never cross it. Of 11 kinds of agent action, 3 pass a gateway. CLI commands, HTTP probes, SQL queries, IMDS metadata calls, credential searches, memory writes, RAG retrieval and A2A hand-offs do not. Gateways stay useful, and Ackuity feeds them context.

Does adding a control switch slow agents down?

Ackuity makes a decision in 40 to 100 ms per action. That is decision time, measured separately from end-to-end latency. Most agent steps already wait far longer on a model response, and the sidecar runs in the agent's own pod, so the check happens close to the action.

NEXT STEP

Agents are going to act. Decide which actions run.

Show us the actions you need to control, and we’ll show you where Ackuity fits.

Request access