The model is editing a page it has never seen the rest of
Give a model a paragraph and ask it to improve the paragraph and it will. The damage happens in the forty other places that paragraph was quietly load bearing.
Vishal Chiniwar Co-founder and CTO 17 June 2026
A small edit with a wide blast radius
Take a request that sounds harmless. Tighten the third paragraph on the security page.
A model does that well. The paragraph gets shorter, clearer, better. In the process the phrase data is encrypted at rest and in transit becomes data is encrypted end to end, which is a different claim, and is the phrase your enterprise questionnaire template quotes verbatim, and is referenced by the compliance section on the pricing page which now says something the security page no longer says.
Nothing about the edit was wrong on its own terms. The paragraph is genuinely better. The failure is that the paragraph was load bearing in four places the model could not see, because none of them were in the request.
This is the whole category. The damage is almost never in the page that was edited. It is in the relationships that page was holding up.
What is actually missing
It helps to be precise about the missing information rather than saying the model lacks context, which is true and unactionable.
There are four distinct kinds, they fail differently, and they are supplied by different mechanisms. Conflating them is why teams add more context and get the same failures.
What else exists. The other pages that mention this subject, and specifically the ones that would contradict a change here. A model that cannot see the pricing page cannot avoid contradicting it.
What is fixed. Legally reviewed wording, brand fixed phrasing, product names that were decided rather than chosen, numbers that came from somewhere authoritative. These look like ordinary prose and are not.
What depends on this. Which pages link here with anchor text that describes this page. Which components render this field. Which schema block derives from this text. A change to a heading can silently invalidate the anchor text on nine other pages.
When this was true. The observation time of everything supplied. A model reasoning over a snapshot taken an hour ago will produce an edit that was correct an hour ago.
The failure modes, and what actually stops each one
This is the table we work from. Each row is a real failure with a specific missing context type and a specific mechanism, because generic remedies do not map onto specific failures.
| Failure | Missing context | What stops it | Cost of not having it |
|---|---|---|---|
| Claim drift | What else exists | Retrieve the pages that state the same fact, not just the target | Two pages disagree, a buyer finds one |
| Fixed wording rewritten | What is fixed | Immutable spans marked in the source, diff rejected if touched | Legal or brand wording silently changed |
| Anchor text invalidated | What depends on this | Manifest of inbound links with their anchor text | Nine pages describe a page that no longer says that |
| Component break | What depends on this | Field level map of which components render which fields | A rendering error, or worse, an empty section |
| Schema divergence | What depends on this | Re-derive markup from final content and compare | Machine readable page contradicts the visible one |
| Stale basis edit | When this was true | Observation timestamp carried with every supplied chunk | An edit correct against a version that no longer exists |
| Terminology split | What is fixed | Canonical name list supplied as data, not as instruction | The product has two names, search treats them as two things |
| Silent scope creep | What is fixed | Field level allow list, everything unlisted denied | A heading changes when only a paragraph was requested |
Why more context is not the answer
The obvious response to a context problem is to supply more context, and the obvious implementation is to give the model the whole site. This makes things worse and it is worth understanding why, because the instinct is strong.
Relevance dilutes. A model given forty pages and asked about one has to find the relevant material inside a much larger volume, and the material that matters competes with material that merely looks similar. Output gets vaguer rather than better grounded.
Cost and latency both go up in a way that gets optimised away later, usually by whoever is on call when the bill arrives, usually without a record of which context was dropped.
And most importantly, volume does not supply the kinds of context that are actually missing. What is fixed and what depends on this are not properties you can read off the page text at all. A thousand pages of context still does not tell a model that a particular sentence was written by a lawyer, or that nine pages link here with anchor text describing this heading.
The answer is not more context. It is structured context: a small amount of the right kind, supplied as data rather than as prose.
The context that is hardest to supply
Of the four types, what is fixed is the one nobody has written down, and it is the one that causes the most expensive failures.
Every site has sentences that look like ordinary prose and are not. Wording a lawyer approved. A phrase that came out of a positioning exercise and is used identically across sales decks. A number that was signed off by someone in finance. A product name that was chosen after an argument and is not up for revision.
None of that is marked anywhere. It is carried by the people who were in the room, and when those people leave, the site becomes a document where every sentence has equal standing. A model treats it that way too, correctly, because nothing told it otherwise.
The remedy is unglamorous and is mostly not a technical task. Somebody has to go through the pages that matter and mark the spans that are fixed, with a note on why and who to ask. It takes a day for a normal marketing site. The reason it does not get done is that it produces nothing visible, and the reason it is worth doing is that it is the only thing that makes those spans defensible against an automated editor, a new hire, or a rushed campaign.
The manifest is the cheap part
If you build one thing before letting anything automated near your site, build a manifest of what each page is and what points at it.
It is not sophisticated. For each page: a stable id, the current URL and its history, which fields exist, which of those are immutable, which other pages link here and with what anchor text, which components render which field, and which schema block derives from which content.
Almost all of that is derivable from a site you already have, in an afternoon, with no model involved. It is a crawl plus a parse plus a join.
And it is what every remedy in the table above depends on. Without it, checking whether an edit invalidates an anchor text somewhere is an unbounded search. With it, it is a lookup. The reason this is the first thing to build is that it converts every downstream check from a research problem into an index problem.
You cannot check what depends on a page unless something already knows what depends on that page.
Source mapping: which words came from where
The second structure worth building is a map from rendered output back to the field it came from.
This sounds like a rendering detail and it is the thing that makes an edit safe to apply. When a model returns improved text, you need to know which field that text belongs to, whether the model changed anything outside that field, and whether any of what it changed was marked immutable.
Without a source map, applying a model's output means replacing a block of rendered content, and you have no way to tell whether the replacement respected boundaries you care about. With one, applying an edit is a field level operation and out of scope changes are detectable rather than plausible.
It also makes the failure diagnosable. When something goes wrong, the question is which field and which rule, and the source map is what makes that question answerable in seconds instead of an afternoon.
Preview has to be against the site
Almost every tool in this space previews the page that changed. That is the wrong unit, and it is wrong in exactly the way the whole article describes.
The failure is a relationship. Previewing the edited page shows you a better paragraph, because the paragraph is better. It does not show you that the pricing page now disagrees with it, or that nine anchor texts now describe something that is no longer there.
A useful preview answers a different question: what else changes, or becomes wrong, if this is applied. Concretely that means resolving the manifest, running the contradiction check against every page that states the same fact, re-deriving schema, and rechecking the anchor text of every inbound link.
None of that is expensive once the manifest exists. All of it is impossible without it, which is the argument for the ordering.
A worked example, start to finish
Follow the security page edit through a system that has the structures above, to see where it stops.
The request arrives: tighten the third paragraph. The source map resolves that paragraph to a named field on a named entity. The field allow list permits text replacement on that field, so scope passes.
Grounding retrieves the page, the three other pages that mention encryption, and the canonical fact set. It attaches an observation time to each. The immutable span register shows that the encryption sentence is marked fixed, with a note naming who approved it.
The model returns a better paragraph. Verification runs. The immutable span check fires: the returned text differs inside a fixed span. The edit is refused before anything is written, with a reason code naming the span and the note attached to it.
What the operator sees is not a failure. It is a specific message: this paragraph contains wording that was approved by somebody, here is who, ask them if you want it changed. The paragraph outside that sentence can still be tightened, and the second attempt does exactly that and passes.
The whole sequence takes a few hundred milliseconds and none of it involved a model deciding what it was allowed to do.
Where teams get this wrong
Four patterns, all of which are recoverable if caught before anything automated touches a live site.
Treating context as a prompt engineering problem. Supplying what is fixed and what depends on this cannot be done in prose, because those are relationships in a graph and prose is not a graph. Attempting it produces long prompts and the same failures.
Building the editor before the manifest. The editor demos well and the manifest does not, so the ordering is almost always wrong. Then the first bad edit arrives and the manifest has to be built anyway, under pressure, against a site that now has an incident attached to it.
Assuming the CMS knows the relationships. It knows the fields. It usually does not know which pages link where with what anchor text, which components consume which fields, or which text is legally fixed. Those live in people's heads until somebody writes them down.
Reviewing the diff rather than the consequences. A reviewer shown a before and after of one paragraph will approve it, because the paragraph is better. The relevant information is what else this breaks, and that is not visible in the diff.
How to evaluate a system that edits your site
These are the questions I would ask of any tool proposing to edit a website, including ours. They are answerable in a few minutes by anyone who built one.
The second question is the one that separates a text editor with a model attached from something that understands a site.
- What does the system know about a page beyond its own content?
- Before applying an edit, what does it check that is not on the edited page?
- How does it know which text on a page may not be changed?
- What does the preview show: the page, or everything the change affects?
- How does an edit know which version of the page it was written against?
- When it refuses, what does the refusal say, and where does that record go?
What good looks like
A system with adequate context is noticeably more conservative and much less interesting to watch.
It declines edits it cannot ground. It says the page changed since it looked and asks to re-observe. It reports that a proposed change would contradict two other pages, and names them. It refuses to touch a span somebody marked fixed, without being asked to remember that.
The measurable version: over a few weeks, the proportion of edits that are reverted after publication approaches zero, and the proportion refused before publication is non trivial and stable. If nothing is ever refused, the checks are not running. If everything is refused, the context is wrong rather than the model.
The behavioural version is simpler. The people who own the site stop reading every diff, because the class of failure they were watching for is being caught by something that does not get bored.
Where we actually are with this
Creogen is in private development and I would rather state that than let a present tense imply a shipped product. The manifest and the source map are the parts that exist and that everything else was built against, in that order, deliberately.
The reason the ordering is not an opinion is that we tried it the other way first while building the Webflow Designer apps. The editing capability came first, it worked, and then a change to a field broke something that consumed that field in a place nobody had recorded. There was no manifest to consult, so finding it was a search rather than a lookup. That is why manifest before editor is stated as strongly as it is here.
Nothing in this article requires our tooling. The manifest is a crawl, a parse and a join. Marking immutable spans is a field in your CMS. Previewing against the site rather than the page is a decision about what your preview resolves. All three are available with whatever you are already running.
The short version
A model editing your site is not looking at your site. It is looking at a request, and everything your site knows that is not in that request is invisible to it.
Four kinds of context close the gap: what else exists, what is fixed, what depends on this, and when this was true. None of them is supplied by giving the model more text, and all of them are supplied by building a small amount of structure first.
If you take one thing: build the manifest before the editor. Everything in the failure table becomes a lookup instead of a search, and the ordering costs nothing if you do it first and a great deal if you do not.
Related reading
A prompt is not a guardrail
Instructions in a prompt are a preference. A guardrail is something that can refuse, that leaves a record when it refuses, and that you can write a test against.Vishal Chiniwar19 May 2026Write the pass conditions before you write the generator
If you cannot say what has to be true for a generated page to be publishable, you do not have a validation problem yet. You have a specification problem.Vishal Chiniwar1 June 2026Your CMS model decides what you will be able to fix later
Most collections are modelled for the page that exists today. Then the product gets renamed and you find out what the model made impossible.Vishal Chiniwar19 June 2026
Questions this raises
No. Two of the four missing context types, what is fixed and what depends on this, are not present in the page text at any length. They are relationships and permissions that live outside the content, so no amount of text supplies them.
A diff shows what changed on the edited page. The failures here are on other pages, and a reviewer looking at a better paragraph will correctly conclude it is a better paragraph. The information needed is what else this breaks, which is not in the diff.
For a site of a few hundred pages it is a crawl, a parse and a join, and it is an afternoon of work rather than a project. The expensive part is the marking of immutable spans, because that requires somebody to decide what is fixed, which is a human judgement nobody has written down before.
For a site of a few hundred pages it is a crawl, a parse and a join, which is an afternoon rather than a project. The expensive part is marking the immutable spans, because that needs somebody to decide which sentences are fixed and who approved them, and nobody has written that down before.
What else exists, because contradiction is the most common failure and the most visible to a buyer. Then what is fixed, because the damage from rewriting legally reviewed wording is the least recoverable. The other two are cheaper to add once the manifest exists.
A well structured CMS gives you fields, which is most of the way there. What it usually does not give you is the mapping from rendered output back to the field it came from, which is what makes an edit checkable rather than a block replacement. If your renderer can tell you that, you already have a source map.
Whoever knows why the wording is fixed, which is usually not the person who will edit it. In practice this means a session with legal, marketing and whoever owns positioning, going through the pages that matter. It takes a day for a normal site and it is the only part of this that is not a technical task.
A model editing your site cannot see the site.
It sees the request. That gap is where the damage happens. We will go through what context has to be supplied and what has to be forbidden.
Creogen supplies page context and refuses edits it cannot ground. In private development.