Guardrails are the controls built into a model or the harness that runs it. Most fall into one of three groups:
- Prompt and response filters that screen text going into and out of the model
- System prompts that tell the model what it should and shouldn't do
- Harness policies, the rules an agent framework applies to the tools and steps it allows
They do real work. Filters catch unsafe content and personal data in text before anyone sees it. Safety training and filtering make a model harder to jailbreak. If your worry is what the model says, guardrails are the right tool, and you should keep them.