Architecting a system that has to be right about someone else's website

Every hard problem in a website operations platform is the same problem: the system is reasoning about state it does not own and cannot lock.

Vishal Chiniwar Co-founder and CTO 13 May 2026

Share
Four layers and a stable content identityIngest, analysis, proposal and application as separate layers, with content identity rather than URL as the primary key.Ingestread onlyAnalysispureProposalno writesApplicationwritesContent identityURL IS AN ATTRIBUTENOT THE KEY

The constraint that shapes everything else

A website operations platform sits in an unusual position. It has to hold opinions about a system it does not control. The site can change underneath you at any moment, by a person you have never met, through an interface you do not own, with no notification. There is no transaction, no lock, no event stream you can subscribe to that tells the whole truth.

Every architectural decision below follows from that. If you design as though you own the state, you will build something that is confidently wrong, which on a live marketing site is worse than being unavailable. An operations tool that reports a problem which was fixed last Tuesday burns its credibility faster than one that says it is not sure.

So the first design rule is that every claim the system makes carries the moment it was observed and the thing it was observed on. Not a timestamp in a log, a timestamp attached to the claim itself, surfaced wherever the claim is surfaced. It is the cheapest thing in the whole architecture and it is the one that keeps the rest honest.

You are not building a source of truth. You are building a system that is careful about how stale it might be.

Identity, and why the URL will not do

The obvious primary key is the URL. It is wrong, and it is wrong in a way that does not show up until about month four.

URLs move. A page gets a better slug. A section gets reorganised and everything under it shifts by one path segment. A CMS collection gets renamed and every item URL changes at once. If the URL is your identity, all of that reads as: the old page was deleted and a new page appeared. Every finding you had attached to it is now orphaned, and every trend you were building is now two half trends.

The identity has to be a stable internal id, with the URL as an attribute that has a history. Then a rename is an event on an existing entity rather than a death and a birth. This sounds like an obvious data modelling point and it is, but it is the single decision I have seen cost the most rework when it is deferred, because retrofitting identity means rewriting every table that referenced the old key.

What the identity has to survive, and what breaks if the URL is the key
EventWhat actually happenedWhat a URL keyed system records
Slug changeOne page, new addressOne deletion and one creation, history lost
Section reorganisedTwenty pages, new prefixTwenty deletions and twenty creations
Collection renamedEvery item URL changes at onceThe entire collection appears to churn overnight
Locale addedSame page, additional addressA duplicate page with no relationship to the original
Page genuinely deletedOne page goneIndistinguishable from a slug change

The five stages, and why they are separate

It is tempting to build this as one pipeline: look at the site, find problems, fix them. That collapses five stages that fail in five different ways, and when the collapsed thing goes wrong you cannot tell which part went wrong.

Observation gathers what is there. Grounding turns raw observation into things the system can reason about. Proposal decides what should change. Verification decides whether the proposal is allowed. Evaluation decides, later, whether the change did what it was supposed to.

Keeping them separate has a specific practical payoff: each stage can be wrong on its own and be diagnosed on its own. If the system proposes something absurd, you can ask whether observation missed something, whether grounding lost context, or whether the proposal logic is at fault, and the answer is in one place.

The five stages, what each one owns, and how each one fails
StageOwnsCharacteristic failureWhat the failure looks like
ObservationFetching and recording current statePartial readsA page is reported missing because a fetch timed out
GroundingTurning pages into structured, addressable factsLost contextA claim is attached to the wrong page or the wrong element
ProposalDeciding what should change and whyConfident nonsenseA suggestion that would be correct on a different site
VerificationDeciding whether a proposal may proceedPermissive checksSomething invalid reaches a live page
EvaluationDeciding whether the change workedNo recorded expectationNobody can say whether it helped, so nothing is learned

Observation has to assume it is missing something

The naive version of observation is a crawl. Crawls are fine and you need one, but a crawl tells you about pages that are linked from other pages, which is not the same as the set of pages that exist. Orphan pages, pages reachable only from an email campaign, pages behind a query parameter, and pages that exist in the CMS but are not published all sit outside it.

So observation needs at least two sources that disagree: what the crawl found, and what the content system says should exist. The disagreement is not a bug in the design, it is the most useful output. A page the CMS knows about and the crawl never reached is either orphaned or broken, and both of those are worth telling someone about.

The other thing observation has to do is record failure as data rather than as an exception. If a fetch times out, that is not nothing, it is an observation with a specific character. Treating it as an error and retrying silently is how you end up reporting that a page vanished when the truth is that a CDN had a bad ninety seconds.

Grounding is where most of the value and most of the risk live

Grounding is the step that turns a page into things you can say something about. A page is not a useful unit for reasoning. A claim on a page is. Pricing stated in three places, a product name used in a heading, a link with anchor text pointing at a document that no longer exists.

This is where retrieval belongs, and it has to be scoped. A model reasoning about a page needs the page, the pages that reference it, and the canonical facts the site has agreed on. It does not need the whole site, and giving it the whole site makes the output worse rather than better, because the relevant context gets diluted by volume.

The rule we work to is that every generated statement has to be traceable to a specific retrieved chunk, and if it cannot be, it does not get made. That is a constraint on the architecture, not a prompt instruction. Prompts are not a control surface. The check has to live in code, downstream, where it can refuse.

  • Retrieve the page under discussion in full, not summarised
  • Retrieve pages that link to it, because they are the ones that will contradict it
  • Retrieve the canonical fact set: product names, prices, plan limits
  • Attach the observation timestamp to every retrieved item
  • Refuse to emit a claim that cannot be traced to a retrieved chunk

Proposal, and the discipline of not acting

Proposal is the stage everyone wants to build first and it should be built last. It is also the stage where the temptation to write directly to the live site is strongest, and it should be resisted absolutely.

A proposal is a described change: this URL, this element, this new value, this reason, this expected effect. It is data. It does not touch anything. That separation is what makes the whole thing reviewable, reversible and auditable, and it is what lets a human be in the loop without the human becoming a bottleneck on every trivial edit.

It also means the proposal can be wrong safely. A proposal that would have been damaging is a rejected record, not an incident.

Verification is not a review step

There is a common shape where generated output goes to a human for approval and that is called the safety layer. It is not one. Humans approving a stream of plausible looking changes will approve almost all of them by the fortieth, and the fortieth is where the bad one is.

Verification has to be mechanical and it has to be written before the generation, not after. The question it answers is: what has to be true for this change to be allowed to exist. Does the URL still resolve. Does the element still exist at that selector. Does the new value contradict a canonical fact. Does the page still validate. Is the change within the scope the site owner permitted.

Every one of those is a check that can fail without a human present, and every one of them is a check a human would not reliably perform on the fortieth item. The human review sits on top of that, for judgement, not for correctness.

What has to be true before generated output reaches a live page

  • The target URL resolves and is the entity the proposal was written against
  • The element still exists, and the observation it was based on is recent enough to trust
  • Every factual claim traces to a retrieved chunk, with that chunk recorded alongside
  • The value does not contradict a canonical fact the site has already agreed
  • The change is inside the scope the site owner granted, by page and by field
  • The page still validates structurally, and the schema still matches the rendered content
  • The expectation is recorded, so evaluation is possible afterwards
  • The change is reversible, and the reversal is stored with it

Evaluation is the stage everyone skips

The change ships. The ticket closes. Nobody ever asks whether it worked. This is not laziness, it is a missing data structure: there is nowhere to write down what the change was supposed to do, so afterwards there is nothing to compare against.

The fix is small and it has to happen before the change rather than after. When a proposal is created, it records an expectation. Not a metric target, which invites gaming, but a statement: this page should stop producing the same support question, or this page should start being able to answer this query. Then evaluation is a comparison, and the honest outcome that it made no difference is available.

Without a recorded expectation, every result can be read as a success, and a system where every result is a success is a system that is not learning anything.

Where teams building this go wrong

Four mistakes, all of which we have either made or come close to making.

Treating the model as the system. The model is one component in the proposal stage. If the architecture diagram has a model in the middle with arrows coming out of it, the interesting parts have not been designed yet. The interesting parts are what constrains the model, what checks it, and what records whether it was right.

Putting the guardrail in the prompt. Instructions in a prompt are a preference, not a control. They degrade under unusual input, they cannot be tested in isolation, and they leave no artifact when they fail. A check in code either passes or does not, and it can have a test written against it.

Caching without an invalidation story. Fetching a site repeatedly is slow and rude, so everyone caches. Very few decide what makes a cached observation too old to act on. Staleness tolerance is not the same for every decision: reading a page to summarise it can tolerate hours, writing to it cannot tolerate minutes.

Building the dashboard first. The dashboard is the part a stakeholder can see, so it gets built early, and then the data model gets shaped by what the dashboard needs rather than by what is true. Every one of those shortcuts is paid for later at a bad exchange rate.

Performance is a design constraint, not a tuning phase

The cost profile of this kind of system is unusual. Observation is network bound and bursty. Grounding is memory bound and embarrassingly parallel. Proposal is latency bound in a way users feel directly. Verification is cheap and must never be skipped for speed.

That last one is the important one. There will be pressure to make verification optional under load, because it sits in the path between a user pressing a button and something happening. Making it optional under load means it is absent exactly when the system is busiest and most likely to be handling something unusual.

The way out is to make verification cheap rather than optional. Every check should be answerable from data already in hand, which is another reason no stage downstream of observation is allowed to fetch. A check that needs a network call is a check that will eventually be skipped.

Interfaces between the stages

Since the stages are separate, the contracts between them matter more than the internals. Two rules have saved us the most trouble.

First, every message between stages carries the entity id, the observation time, and the provenance of anything derived. A proposal that arrives without knowing when its observation was taken cannot be verified safely, and by the time it reaches verification it is too late to go and ask.

Second, no stage is allowed to fetch. Observation fetches. Everything downstream works from what observation recorded. This sounds restrictive and it is deliberately so, because a grounding step that can go and fetch the page again will do it under time pressure, and then you have two different versions of reality inside one decision.

What this means for build order

If you are building something in this shape, the order that has worked for us is the opposite of the order that feels satisfying.

Identity first, because retrofitting it is the most expensive mistake available. Then observation, including the failure recording, because everything downstream is only as good as the record it reasons over. Then verification, before proposal, so that the first proposal you generate meets a check that already exists. Then proposal. Then evaluation, which needs real proposals to have something to evaluate.

Building proposal first is the common path and it produces an impressive demo and a system with no floor under it.

  • Identity and the entity model, with URL history
  • Observation, with failures recorded as data
  • Verification checks, written before anything generates
  • Proposal, as described data that touches nothing
  • Evaluation, comparing outcome against a recorded expectation

How to evaluate an architecture like this

If someone shows you a system in this space, including ours, these are the questions that separate a real architecture from a demo. I would want to be asked all of them.

The answers matter less than whether the person can answer at all. A system that has thought about staleness will have an opinion about staleness.

  • What is the primary key for a page, and what happens to history when a slug changes?
  • How does a claim carry the time it was observed, and where does a user see that?
  • What check would stop an invalid change reaching a live page without a human?
  • Where is the expectation for a change recorded, and when?
  • What happens when observation partially fails, and how is that different from a page being gone?

Where we actually are with this

Creogen is the platform this architecture describes and it is in private development. I would rather state that than let the present tense imply more than it should. The identity model, the observation layer and the verification stage are the parts that exist and that shaped everything else. Creobot, the conversation engine, feeds the observation side with what visitors actually asked, which is the input the site otherwise does not have.

The reason we are building it this way rather than faster is that the failure mode of getting it wrong is publishing something incorrect on somebody's live marketing site. That is not a recoverable class of error in the way a broken internal dashboard is. It is the customer's front door.

Most of what is above came from building four Webflow Designer apps and getting the identity question wrong in one of them first. That rework is the reason identity is at the top of the build order rather than somewhere in the middle.

Written by

Vishal Chiniwar

Co-founder and CTO

Vishal builds the systems behind Creoglyph. He wrote the four Webflow Designer apps the company ships, and now works on the architecture behind Creogen and Creobot: retrieval, grounding, CMS data models, install flows and the validation that has to sit between a model and a live marketing site.

Related reading

Questions this raises

Because a described change is reviewable, reversible and auditable, and a direct write is none of those. The separation also means a wrong proposal is a rejected record rather than an incident on a live page.

No. Humans approving a stream of plausible looking changes approve almost all of them after the first few dozen. Mechanical checks written before generation catch the things attention does not. Human review sits on top of that for judgement, not for correctness.

It changes. Slug edits, section reorganisation, collection renames and locale additions all change URLs without changing the page. If the URL is the key, each of those reads as a deletion plus a creation and the history is lost.

Yes, and it is the most expensive retrofit in this list, which is why the article puts it first in the build order. The path is additive: introduce an entity table, backfill it from current URLs, add the URL as an attribute with a history, then migrate each table that referenced the old key one at a time. Expect it to take longer than the original implementation did.

It depends on the action rather than on a single threshold. Reading a page to summarise it can tolerate hours. Writing to it cannot tolerate minutes, because somebody else may be editing. Set the tolerance per operation type and record it, rather than picking one number for the whole system.

Identity and observation. Those two produce a record of what the site is and what it looked like, which is useful on its own and is the thing everything else reasons over. Proposal without them is a system making confident claims about state it has not established.

No. Every constraint in it comes from reasoning about a content system you do not own and cannot lock, which is true of any hosted CMS. The Webflow specifics live in the observation layer, in how you fetch and what the platform tells you, and nothing above that stage depends on which platform is underneath.

Where would your stack fail without telling anyone?

That is the question the architecture has to answer. Bring the part you are least sure about and we will draw it out properly.

Creogen is in private development. We would rather show the architecture than the marketing.

A visitor question becoming a change on a page A question enters on the left, Creobot captures it, Creogen turns it into an operation, and one block on the page is marked as changed. Creobot Creogen QUESTION TO CHANGE