ACKUITY VS MODEL GUARDRAILS

Can an AI agent be trusted to enforce its own rules?

Model and harness guardrails filter what goes into and comes out of a model, and they are good at it. They run inside the agent they protect, though, which makes them a weak place to decide whether an action should run. Ackuity, the Agent Execution Control Switch, makes that decision from outside the agent and works alongside the guardrails you already have.

Model and harness guardrails and Ackuity, side by side
Model and harness guardrailsAckuity
Where it runsInside the model or the agent's harnessBeside the agent, outside its control
What it checksPrompts, responses and harness tool rulesEach action before it executes
Best atContent safety, PII in text and jailbreak resistanceDeciding whether an action should run
Influence from the agentShares the agent's context, so injected text reaches bothNone: the context graph is built outside the agent
Context per decisionThe conversation and the harness's rules29 signals across 6 dimensions, including history across steps
ScopeOne policy per harness or modelOne policy across agents and environments
ResponsesRefuse, redact or rewrite textAllow, Constrain, Human in the loop, Block, Terminate

WHAT GUARDRAILS DO WELL

What do model and harness guardrails actually do?

Guardrails are the controls built into a model or the harness that runs it. Most fall into one of three groups:

A large boundary marks the inside of the agent. It holds the model and the harness that runs it. An input filter checks prompts going into the model and an output filter checks responses coming out, which makes guardrails good at content safety, PII in text and jailbreak resistance. Inside the harness, the agent plans toward a goal, and its own reasoning can route around a rule that blocks the direct path. Outside the boundary, in the execution path between the agent and its targets (databases, APIs, tools, files, cloud services and other agents), Ackuity checks each action before it runs and responds with Allow, Constrain, Human in the loop, Block or Terminate. Both layers are useful and work together.INSIDE THE AGENTGuardrails check prompts and responsesGUARDRAILInput filterchecks promptsModelGUARDRAILOutput filterchecks responsesHARNESSAgent harnessruns the plan andapplies tool rulesPLANRULEGOALreasoning can route around a ruleACTIONOUTSIDE THE AGENTIn the execution path, before targetsACKUITYChecks eachactionbefore it runsRESPONSESAllowConstrainHuman in the loopBlockTerminateTARGETSDatabasesAPIsToolsFilesCloudAgentsBEST ATcontent safety, PII in text, jailbreaksBEST ATdeciding whether an action runs
Model and harness guardrails work inside the agent, checking prompts and responses, while Ackuity works outside it in the execution path, checking each action before it reaches a target. The two layers answer different questions and work best together.
  • Prompt and response filters that screen text going into and out of the model
  • System prompts that tell the model what it should and shouldn't do
  • Harness policies, the rules an agent framework applies to the tools and steps it allows

They do real work. Filters catch unsafe content and personal data in text before anyone sees it. Safety training and filtering make a model harder to jailbreak. If your worry is what the model says, guardrails are the right tool, and you should keep them.

THE LIMIT

Why do guardrails fall short once agents start acting?

A guardrail lives inside the thing it is meant to control. The model reading the system prompt is the same model deciding what to do next, so the controlled party is also the control. For text, the worst outcome is a bad sentence. When the output is a database write or a cloud API call, the stakes change.

Agents reason, and reasoning means finding a way to reach a goal. Give an agent a goal and a rule that blocks the obvious path, and it may find a second path to the same place. A prompt injection hidden in a document can also change the agent's idea of what the user wanted, and the guardrail reads the same poisoned context the agent does.

Scope is the other limit. Each harness carries its own policy in its own format, so a company running agents on 4 frameworks maintains 4 rulebooks. None of them sees what other agents are doing, or what the same user asked a different agent an hour ago.

It's a bit like asking a new hire to supervise themselves. Most days that works. On the day someone talks them into something, you want a second person who wasn't part of the conversation.

OUTSIDE THE AGENT

Where does Ackuity check the action?

Ackuity checks each action beside the agent, in its execution path, before the action runs. That spot is not inside the agent, not at the gateway, and we call it the Goldilocks zone. It is close enough to see intent, plan and history, and it stays outside the agent's control.

For every action, Ackuity assembles a security context graph for AI agents: 29 signals across 6 dimensions, from the user and the agent's permissions to the target system and the session's history. Ackuity builds it outside the agent, so the agent can't edit it. Neurosymbolic verification (rules, behavioral baselines and small language models, correlated across steps) weighs the action against 60+ threat models in 14 categories and decides in 40 to 100 ms.

The answer comes from the response ladder: Allow, Constrain, Human in the loop, Block or Terminate, with Alert and Log below the block threshold. When an agent goes rogue, Ackuity blocks or constrains the bad action and the rest of the agent's work can carry on. Terminate is reserved for catastrophic cases.

Ackuity Core applies one policy across agents and environments, so security teams manage a single set of rules instead of one per harness.

WORKING TOGETHER

Should you keep your guardrails?

Yes. Guardrails and Ackuity answer different questions. A guardrail asks whether this text is safe. Ackuity asks whether this action, for this user, against this system, right now, should run.

Take a customer service agent. A response filter catches a phone number before it lands in a reply. Ackuity blocks the SQL query that would have pulled the whole customer table, or masks the sensitive columns and lets the query run. Both controls did their jobs, at different points. The context Ackuity assembles also flows to your SIEM, so the security team sees the full story behind each decision.

QUESTIONS PEOPLE ASK

Frequently asked questions.

What is the difference between guardrails and execution control?

Guardrails filter what a model reads and writes, and execution control decides whether an agent's action runs. Guardrails work inside the model or harness. Ackuity, the Agent Execution Control Switch, works beside the agent and checks each action against its full context before it executes.

Can an AI agent bypass its own guardrails?

Yes, because guardrails sit inside the agent that reasons over them. An agent chasing a goal can find a path its rules didn't anticipate, and a prompt injection can change what it believes the user wanted. Ackuity checks the action outside the agent, so the agent has no influence over the decision.

Does Ackuity replace model guardrails?

No. Keep guardrails for content safety, PII in text and jailbreak resistance. Ackuity adds an independent check on the actions agents take, which guardrails weren't built to make.

Is a system prompt a security control?

A system prompt is an instruction, and the model weighs it alongside everything else it reads. It shapes behavior well on ordinary days. Injected text or creative reasoning can override it, so it shouldn't be the last check before an agent writes to production.

Does Ackuity need changes to my agent code?

No. The open source sidecar runs in the agent's pod and an init container reroutes its traffic, so the agent code stays as it is. Platforms such as Copilot Studio connect through API injection, and OpenTelemetry, Langfuse or LangSmith events give an observe-only start.

NEXT STEP

Agents are going to act. Decide which actions run.

Show us the actions you need to control, and we’ll show you where Ackuity fits.

Request access