A prompt is not a guardrail

Instructions in a prompt are a preference. A guardrail is something that can refuse, that leaves a record when it refuses, and that you can write a test against.

Vishal Chiniwar Co-founder and CTO 19 May 2026

Share
A gate of deterministic checks between a model and a live pageModel output passes through four enforced checks before it is allowed to reach a page.ModelGROUNDINGSCHEMA VALIDSCOPE CHECKHUMAN APPROVALA PROMPT IS A REQUESTA GUARDRAIL REFUSES

What the model can actually see

A model asked to update a page does not have the page. It has whatever you put in the request, which is a few thousand tokens of text that somebody chose. It does not have the other forty pages that mention the same product. It does not know that the sentence it is rewriting is the one place your pricing is authoritative. It does not know that the phrase it finds clumsy is legally reviewed wording.

This is the root of most of the damage, and it is not a model quality problem. A better model with the same context makes a better version of the same mistake. It will produce a more fluent sentence that still contradicts the pricing page, because contradicting the pricing page was never something it was in a position to notice.

So the design question is not how do we stop the model being wrong. It is what has to be true about the system around the model so that a wrong output cannot reach a live page.

What breaks, specifically

It is worth being concrete, because the abstract version of this argument produces abstract safeguards.

The table below is the set we design against. Each row is a thing that happens when a model edits a site with insufficient context, and each one has a guardrail that is mechanical rather than linguistic. Notice that none of the guardrails is an instruction. Every one of them is a check that either passes or does not.

Failure modes when a model edits a site, and the guardrail for each
What breaksHow it happensThe guardrailWhere it lives
ContradictionThe edited page now disagrees with another page about a factCompare the new value against the canonical fact set before acceptingVerification, before write
Silent scope escapeAn edit to one field also rewrites the heading above itAccept a diff against named fields only, reject anything outsideVerification, before write
Broken referenceA link or component reference is rewritten into something that does not resolveResolve every reference in the output against the live route tableVerification, before write
Tone driftLegally reviewed or brand fixed wording gets improvedMark spans as immutable in the source, reject any diff that touches themGrounding and verification
Structure lossA table becomes prose, or a heading level disappearsCompare the structural outline before and after, reject changes not requestedVerification, before write
Stale basisThe edit is correct against a version of the page that has since changedCarry the observation time, refuse if the page changed after itObservation and verification
Schema divergenceThe prose changes and the structured data still describes the old claimRe-derive schema from the new content and comparePost write, blocking

Why the prompt cannot carry this

Everyone building these systems starts by putting the rules in the prompt. Do not change the heading. Do not alter pricing. Keep the tone. It works in testing, which is the problem.

Instructions in a prompt fail in three ways that a check in code does not. They degrade under unusual input, and unusual input is exactly the case you are trying to survive. They cannot be tested in isolation, because the only way to know whether the instruction held is to run the whole generation and read the output. And when they fail they leave nothing behind: there is no artifact that says the rule was violated, only an output that happens to be wrong.

A check in code has the opposite properties on all three. It behaves the same on unusual input, it has a unit test, and a failure produces a record with a reason attached.

This is not an argument against putting instructions in the prompt. Do that too, because it improves the average output and reduces how often the checks have to fire. It is an argument against counting the instruction as the safety mechanism. The prompt makes good output more likely. The guardrail makes bad output impossible to publish. Those are different jobs and only one of them can be audited.

What a guardrail has to be able to do

The working definition we use has three parts, and a mechanism that fails any of them is not a guardrail regardless of what it is called.

It can refuse. Not warn, not flag, not lower a confidence score. Refuse, so that the operation does not happen. If the only outcome is a warning on a dashboard, you have built monitoring, which is useful and is not a control.

It leaves a record when it refuses, with the reason, the input, and the rule that fired. Without this you cannot tell the difference between a guardrail that is working and a guardrail that never runs.

It has a test. If you cannot write a case that the guardrail rejects and a case it accepts, it is not specified tightly enough to rely on. This is the requirement that kills most vaguely worded rules, and killing them early is the point.

Scope has to be data, not description

The most common architectural mistake I see is expressing scope in language. The system is told it may edit marketing copy but not legal text, or that it should only change things the user asked about. Both of those are judgements, and handing a judgement to the component you are trying to constrain defeats the exercise.

Scope has to be a structure the checking layer can evaluate without interpreting anything. Which pages, by id. Which fields on those pages, by name. Which spans within those fields are immutable. What kinds of change are permitted: replace text, yes; change a link target, no; alter heading level, no.

Once scope is data, the check is trivial and unarguable. The proposed diff either touches only permitted fields on permitted pages or it does not. There is no threshold to tune and no confidence score to interpret, which is exactly what you want in the component that decides whether something reaches a customer's website.

  • Pages, by stable entity id rather than by URL
  • Fields on those pages, by name, with an explicit allow list
  • Immutable spans within permitted fields, marked in the source
  • Permitted operation types, with everything unlisted denied
  • A maximum size of change, so a rewrite cannot arrive disguised as an edit

Where the guardrails have to sit

Position matters as much as content. A check in the wrong place is decoration.

Anything that decides whether an operation is permitted must run after generation and before the write, in the same process as the write, with no path around it. If there is a code path that reaches the write without passing the check, that path will be taken eventually, usually by a retry handler somebody added under time pressure.

Checks that run in the model's own reasoning are not guardrails. Checks that run in the client are not guardrails. Checks that run asynchronously after publication are monitoring, which is valuable and different: they tell you something already went wrong.

The practical rule we work to is that the write function itself takes a validated proposal type, and there is no constructor for that type except the one that runs the checks. Then the guardrail is not something you remember to call. It is the only way to obtain the thing the write needs.

Human review is not the safety layer

There is a common shape where generated changes queue for human approval and that queue is described as the safety mechanism. It is not one, and the reason is well understood outside this field.

A reviewer looking at a stream of plausible changes is performing a vigilance task. Vigilance degrades with volume and with base rate. If thirty nine of the last forty items were fine, attention on the fortieth is not what it was on the first, and the fortieth is where the interesting failure is. Nothing about the reviewer being conscientious changes this.

Human review is genuinely valuable for judgement: is this the right change to make, does it say what we want to say, is the argument correct. Those are questions a machine cannot answer and a person can. Correctness questions, does this contradict another page, does this link resolve, does this touch a field it should not, belong to the machine, because the machine is not bored on the fortieth item.

Refusals are the most useful output you have

A system with guardrails produces refusals, and the instinct is to treat them as friction to be reduced. That is backwards.

Every refusal is a specific statement about a boundary the system met. A refusal because a proposed change would have contradicted the pricing page tells you the model was working from a stale or incomplete fact set. A cluster of refusals on the same field tells you the scope is drawn wrong, or that people keep wanting to do something the system does not allow.

So refusals need somewhere to go beyond a log. They need a queue that a person reviews periodically, grouped by rule, because the shape of the cluster is the signal. If one rule accounts for most refusals, either that rule is too tight or the capability behind it is missing.

This is also the part that keeps the guardrails honest over time. A rule nobody ever looks at gets loosened during the first incident where it is inconvenient. A rule with a visible refusal queue gets argued about with evidence.

Where teams get this wrong

Four patterns worth naming, all of which are recoverable if caught before anything is live.

Confidence thresholds as a control. A model's own confidence is not calibrated against your business rules, and tuning a threshold is guessing at a number that will drift with every model change. Deterministic checks do not need tuning.

Guardrails as a post launch phase. The check has to exist before the first generated output, because the first generated output is when everyone learns how much they trust the system. Building the check afterwards means building it against outputs you have already accepted.

Guardrails that need a network call. A check that fetches something will be slow, and a slow check in the write path gets made optional under load. Every check should be answerable from data already in hand, which is why nothing downstream of observation is allowed to fetch.

One giant validator. A single function that decides everything cannot tell you which rule fired, cannot be tested per rule, and becomes the thing nobody wants to modify. Small named rules, each with its own test and its own refusal reason.

How to evaluate a guardrail layer

These are the questions I would want asked about ours, and they are more informative than looking at a list of rules.

The last one is the one that separates a real control layer from a described one. If nobody can produce a refusal on demand, the checks may never have run.

Questions that tell you whether guardrails are real

  • Can you name the code path that reaches a write, and show that every branch of it passes the checks?
  • For each rule, is there a test case that it rejects and one that it accepts?
  • What happens to a refusal? Where does the record go, and who reads it?
  • Is scope expressed as data, or described in language somewhere?
  • Which checks require a network call, and what happens to them under load?
  • How does a change know which version of the page it was written against?
  • Can you trigger a refusal right now, deliberately, and show me the record it produced?

What good looks like in practice

A system with real guardrails is noticeably less impressive in a demo. It declines things. It asks for scope it has not been given. It refuses an edit because the page changed four minutes ago and asks to re-observe first.

In production those same behaviours read as trustworthy rather than as limitations. The people who have to live with the system stop watching it, which is the actual goal. A tool that requires supervision has not saved anyone any work.

The clearest signal that it is working is boring: over a few weeks, the refusal queue shifts from correctness refusals to scope refusals. Correctness refusals mean the system is generating things that contradict reality, which is a grounding problem. Scope refusals mean it is generating reasonable things it has not been permitted to do, which is a permissions conversation. The second is a much better problem to have.

Where this sits in what we are building

Creogen is in private development, so read this as the reasoning that shaped it rather than a description of a shipped product. The constraint we set early was that the write path takes a validated proposal and there is no other way to construct one. Everything above follows from having taken that seriously rather than from any insight about models.

Most of it came from building four Webflow Designer apps and watching what happens when an automated change meets a site somebody else is also editing. The stale basis row in the table is there because we hit it, not because we predicted it.

None of this requires our tooling. Express scope as data, put the checks after generation and before the write with no path around them, give refusals a queue, and write a test per rule. That is available with whatever you are already using, and it is most of the distance.

Written by

Vishal Chiniwar

Co-founder and CTO

Vishal builds the systems behind Creoglyph. He wrote the four Webflow Designer apps the company ships, and now works on the architecture behind Creogen and Creobot: retrieval, grounding, CMS data models, install flows and the validation that has to sit between a model and a live marketing site.

Related reading

Questions this raises

Put them there as well. Instructions improve the average output and reduce how often the checks fire. The mistake is counting the instruction as the safety mechanism, because it cannot refuse, cannot be tested in isolation, and leaves no record when it fails.

No. Reviewing a stream of plausible changes is a vigilance task, and vigilance degrades with volume. Humans are good at judgement questions and poor at correctness checking on the fortieth similar item. Give the machine the correctness checks and the human the judgement.

Stricter than feels comfortable, then loosen with evidence from the refusal queue. Loosening a rule because refusals cluster on it is a decision with data behind it. Tightening a rule after something reached a live page is not.

Three. A field level allow list so scope is data rather than description, a check that every factual claim traces to a retrieved chunk, and a reference resolver so nothing links to something that does not exist. Those three cover the failures that reach a customer's page, and each is answerable from data already in hand.

Give refusals a visible queue that somebody reviews, so loosening a rule becomes an argument with evidence rather than a decision made under pressure. A rule nobody looks at gets relaxed the first time it is inconvenient. A rule with a refusal history attached is much harder to wave through.

It can, and it should not be counted on. Putting the rules in the prompt improves average output and reduces how often the checks fire, which is worth having. It is not a control, because it cannot refuse, cannot be tested alone, and leaves nothing behind when it fails.

Not zero and not most of them. Zero means the checks are not running or the scope is drawn so wide that nothing can violate it. A majority means the context is wrong rather than the model. What matters more than the rate is the trend in reasons: correctness refusals shifting toward scope refusals over a few weeks is the system working.

What stops a model writing to your live site?

If the answer is a prompt, that is not a guardrail. We will go through what validation actually has to sit in that path.

Creogen validates before it writes. That constraint shaped the whole system.

A visitor question becoming a change on a page A question enters on the left, Creobot captures it, Creogen turns it into an operation, and one block on the page is marked as changed. Creobot Creogen QUESTION TO CHANGE