All posts
AI Agents

AI Agent Guardrails: Hard Checks Beat Prompts

AI agent guardrails in layers: hard checks that fail the run, soft rules in prompts, a human approval gate, and the linter rules I run on my own blog.

October 3, 2026·7 min read·by Olexander Cheberko
Table of contentstap to expand

AI agent guardrails are the checks that keep an agent's work inside your rules, and they come in two strengths: hard checks that fail the run or block the action, and soft rules that only ask the agent to comply. Hard checks beat prompts because an agent can forget a prompt, but it cannot get past a check that fails. Below is how I layer both on this blog, which agents write, with the rules my linter enforces and why.

What are AI agent guardrails?

Limits on what an agent may read, do and produce, enforced before it acts, while it acts and before its output takes effect. For an agent that produces work, such as code or a published page, the guardrail that matters most checks the output before it ships, plus the person who approves it.

What is the difference between hard and soft guardrails?

Lauren Tan, an engineer at Cursor, draws this line in a recorded talk on how she came to trust coding agents. In her layers, a strict codebase convention is the strongest, and CI checks, lint rules for bad patterns she has observed and compiler diagnostics turn CI red. Rules, skills and a review bot are soft: an agent can forget them or apply them unevenly. She layers both but will not rely on the soft layer alone.

Claude Code's docs make the same split: Claude treats CLAUDE.md and auto memory as "context, not enforced configuration", and to block an action regardless of what Claude decides, you use a PreToolUse hook (memory).

LayerExamplesWhen a rule breaks
HardCompiler, CI job, lint script, a Claude Code hook or deny ruleThe run fails or the action is blocked
SoftPrompts, CLAUDE.md, memory, skills, review bot commentsThe agent may comply, or may not
HumanApproval before a commit, a push or a sendNothing takes effect until a person says yes

Which layer should a rule live in?

Tan's test is the review comment: each time a person enforces a constraint by commenting on a pull request, treat it as a smell and ask "how do I turn this into a hard rule" (talk). Her answers: a lint rule, a CI failure, or eliminating the problem entirely. I sort rules by what it takes to check them:

  • A script can check it with no judgment: hard. A dash, a link, a length, a banned phrase.
  • It needs judgment, such as whether a claim matches a vendor's help page: soft, but checked by an agent that did not write the text.
  • It goes public or is hard to undo: a person approves it.

How did I apply this to an AI-written blog?

Each article gets one writer agent, an independent fact-checker that re-reads the vendor's help pages and fixes or cuts claims, editor passes, and one agent that renders every page in a real browser at desktop and phone width. One run on October 2, 2026 used 32 agents to write 13 manuals, tighten 5 drafts and render-check 20 pages.

  • Soft: prompts and memory rules. I rejected early drafts as padded, and the fix became a fixed manual format (numbered checks with Where, Do, You should see, If not) plus a memory rule. Third-party SEO skills here run under my rules block: no paid API call without my approval, no indexing submissions from the skills, my positioning rules win. A rules block is still a prompt.
  • Hard: scripts/content_lint.py, which exits with an error when a rule breaks. Each writer agent must reach zero errors on its file before handing it back.
  • Human: I approve every commit and every push. Agents never push on their own.

What does my content linter check, and why?

Check (fails the run)Why it exists
An em dash, or a spaced en dash used as a separatorHouse style bans them, models write them freely, and a style rule in a prompt is the kind an agent forgets
A link to a removed page or to a page that does not existPages I removed answer 410 Gone, and a draft can link a sibling draft that never ships. The reader hits a dead end
A title over 60 characters or a description over 155Google truncates title links "as needed, typically to fit the device width" (Google), so the query goes first and the text stays short
A category outside the fixed listMy blog code quietly files an unknown category under Strategy, so a typo misfiles a post with no error
An FAQ question that repeats one on another pageOne question should have one answer page on the site. Two live pages already repeated a question from another page when the check arrived
A manual whose declared step count differs from its checks, or a check without its expected result and fixThe manual format from the rejected drafts, as code. A check with no expected result and no fix is padding with a number
Banned phrasing: "HIPAA-compliant" about my work, "NYC-based", invented frequency such as "the case I see most"I call my work HIPAA-conscious, I work remotely, and a fact-checker caught an agent inventing frequency with nothing behind it

Softer rules only warn, such as "we" or "our" in prose: this site speaks for one person. Known errors in pages I am not editing yet sit in a baseline file until their dated fix, so the check could be strict from day one.

Each error names the rule it broke, and the banned-phrase errors name the fix, so the agent can repair the file and rerun the check. An excerpt, shortened:

BANNED = [
    (r"(?i)\bHIPAA[- ]compliant\b", "say HIPAA-conscious, never HIPAA-compliant about the work"),
    (r"(?i)\bNYC-based\b", "say NYC-focused, run remotely"),
    (r"(?i)\bthe case I see most\b|\bclients often\b",
     "invented frequency claim: state one real case or nothing"),
]

Claude Code hooks work the same way: a PreToolUse hook script that exits with code 2 blocks the action, and in the docs' file-protection example its Blocked: message reaches Claude as feedback (hooks). The docs call this "deterministic control": the check runs because the hook fires, not because the agent remembered.

What can hard checks not catch?

Judgment. A fact-checker caught a claim about Google Ads "Calls from ads" that left out what happens with call recording on: AI analyzes the recording to determine lead quality (Google). Agents also audited public pages of practices I researched for outreach, read-only, with headless Chrome and network capture. One finding said a site set consent to "denied" with no banner; it held only because the check ran from Bulgaria, inside the EU, and I withdrew it. No regex finds either mistake.

Where does the human stay in the loop?

At the commit and the push: I approve every one, and agents never push on their own. Tan describes the far end of the trust curve in the same talk: her agents now merge their own pull requests and she reviews them on main afterward, and the talk is about what it took to get there. My pages are public and carry my name, so the gate stays.

To enforce the gate rather than promise it, the permissions.ask setting makes Claude Code "always prompt before listed tool uses" (settings), and rules are evaluated deny, then ask, then allow, so a narrower allow rule cannot skip the prompt (permissions).

{
  "permissions": {
    "ask": ["Bash(git push *)"]
  }
}

The permissions page notes the limit: a push written another way, such as git -C . push origin main, does not match, so this gates the usual command and is not a security boundary.

What does this mean for tracking work?

The same split applies to conversion tracking, my actual job. A script can confirm that a tag loads and a conversion request leaves the page; a person judges whether that conversion is the one the business should bid on. If you want both layers built into your measurement, see marketing analytics.

Tags

ai-agent-guardrailshuman-in-the-loopclaude-codecontent-lintingai-agents

Frequently asked questions

What is the difference between hard and soft guardrails for AI agents?

A hard guardrail fails the run or blocks the action when a rule breaks: a compiler error, a failing CI job, a lint script that exits with an error, a Claude Code hook that exits with code 2 or a deny rule. A soft guardrail asks the agent to comply: a prompt, a CLAUDE.md file, a skill, a review bot's comment. An agent can forget a soft rule. It cannot get past a hard check that fails.

Is a rule in CLAUDE.md a guardrail?

A soft one. Claude Code's documentation says Claude treats CLAUDE.md as context, not enforced configuration, and points to a PreToolUse hook when an action must be blocked regardless of what Claude decides. Keep CLAUDE.md for guidance and move anything a script can check into a hook, a permission rule or a lint step.

What does human in the loop mean for AI agents?

A person approves an agent's output before it takes effect. On my blog that point is the commit and the push: agents write, fact-check and render the pages, and nothing reaches the site until I approve it. The gate belongs wherever an action is public or hard to undo.

Can an AI agent fix the errors a linter reports?

Usually, if the error names the fix. My linter's banned-phrase messages do, for example 'say HIPAA-conscious, never HIPAA-compliant about the work', so the writer agent can correct the file and run the check again. What a linter cannot judge, such as whether a claim matches a vendor's help page, goes to a separate fact-checker.

Related posts