Write the pass conditions before you write the generator

If you cannot say what has to be true for a generated page to be publishable, you do not have a validation problem yet. You have a specification problem.

Vishal Chiniwar Co-founder and CTO 1 June 2026

Share
Six deterministic checks before human reviewGenerated output passes six mechanical checks, then a human review, before it is allowed to become a page.GenerateduntrustedSTRUCTURELINKSSCHEMAFACTSVOCABULARYDUPLICATIONHuman review

The question that comes before the pipeline

Teams usually arrive at validation after they have a generator. The generator produces pages, somebody notices a bad one, and validation gets added to catch bad ones. That ordering guarantees you build checks shaped around the failures you already saw.

The better starting point is a question with nothing to do with generation: what has to be true for a page on this site to be publishable at all. Most teams have never written that down, including for pages a human wrote. Doing it produces a list that applies to everything and that a generator then has to satisfy, rather than a list of patches.

It is also the fastest way to discover that your standards are not agreed internally. If two people write the list and they differ, that disagreement was already there and was previously resolved by whoever happened to be publishing.

Validation is not review

These get conflated constantly and they answer different questions.

Validation asks whether the page is allowed to exist: does it resolve, does it contradict a known fact, does it have the required structure, does the schema match the content. Every one of those has a deterministic answer and none of them requires taste.

Review asks whether the page is any good: is the argument right, is this the thing we want to say, does it sound like us. None of those has a deterministic answer and all of them require somebody who knows the company.

Keeping them separate matters because they scale differently. Validation scales to any volume. Review does not, and a system that requires review of every generated page has not saved anyone work. The design goal is to make validation strict enough that review can be about judgement rather than about correctness.

The five stages

Order matters, and the ordering principle is cheapest and most certain first. Later stages are expensive, so they should only ever run on candidates that survived the cheap ones.

Structural comes first because it needs nothing beyond the output itself. Referential next, because it needs the site's route table but not its content. Factual next, needing the canonical fact set. Then consistency against the rest of the site. Then policy, which is the smallest and the one most specific to your company.

The five validation stages, in the order they should run
StageQuestion it answersNeedsTypical failure
StructuralIs this a well formed page of the required shape?Only the outputMissing h1, heading levels skipped, empty required field
ReferentialDoes everything it points at exist?Route table, asset listLink to a deleted page, image reference that does not resolve
FactualDoes every claim trace to a source?Canonical fact set, retrieved chunksA price with no supporting chunk, an invented plan name
ConsistencyDoes it agree with the rest of the site?Existing page contentContradicts the pricing page, uses a retired product name
PolicyIs it allowed by our own rules?Policy configBanned phrasing, missing disclosure, wrong locale conventions

Worked example: a page that passes

Take a generated integration page. It has an h1 naming the integration, three h2 sections, a short paragraph under each, a specification table, and a link to the relevant documentation.

Structural passes: one h1, no skipped heading levels, all required fields present, the table has a header row and consistent column counts. Referential passes: the documentation link resolves against the current route table, and the logo asset exists at the referenced path. Factual passes: the two claims that matter, that the integration supports two way sync and is available on Business and above, each trace to a retrieved chunk from the product data, recorded alongside. Consistency passes: the plan name matches the canonical plan list, and no existing page says anything different about this integration. Policy passes: no banned phrasing, the availability disclosure is present.

Note what validation did not decide. It did not decide whether the three sections were the right three, whether the writing is good, or whether this integration deserves a page. Those are review questions and they are still open. What validation established is that publishing it would not put anything false or broken on the site.

Worked example: a page that fails, and how

Same generator, different input. The page describes an integration that was renamed six months ago.

Structural passes, because the shape is fine. This is the important part: the page looks completely correct at the structural level, which is why structure alone is not validation.

Referential fails. The documentation link points at the old path, which now returns a redirect rather than a page. Reason code: reference resolves to a redirect rather than a target. That is a real distinction, because a redirect might be acceptable policy and a missing page is not.

Factual fails separately. The page states availability on a plan tier that was merged into another tier in the last pricing change. The claim traces to a retrieved chunk, but the chunk carries an observation timestamp older than the pricing update, so recency ranking should have preferred the current page and did not. Reason code: claim supported only by a chunk older than the canonical fact it contradicts.

Consistency fails third, on the product name, against four existing pages that use the new one.

Three failures, three reason codes, one root cause: the index was stale. That is the value of separate reason codes. A single failed boolean would have told you the page was bad. The codes tell you the indexing pipeline missed an update, which is where the actual fix belongs.

Where the pipeline sits, and what happens on failure

Validation runs after generation and before publication, in the same process as the write, with no branch around it. That much follows from the general guardrail argument and I will not repeat it.

What is specific here is what happens next, because a validation layer that only blocks is half a system. Each reason code carries a disposition, and the dispositions are what turn failures into either a fix or a retry rather than into a queue nobody drains.

A staleness failure re-observes and retries once, automatically, because the correct response to the basis having changed is to look again. If it fails a second time on the same code, it stops and routes to a person, because two staleness failures in a row usually means something is editing the page concurrently and retrying forever is rude.

A factual failure never retries. Regenerating with the same context produces the same claim with different wording, which is worse, because now the failure looks like a different failure. It routes to grounding: the fact set is wrong, or the retrieval missed, and both of those are upstream problems.

A policy failure routes straight to a person with the rule quoted, because policy is the one category where the right answer is sometimes to change the rule.

Every failure needs a reason code

This is the design decision I would defend hardest. A validator that returns pass or fail is nearly useless at any volume.

Reason codes let failures be aggregated, and the aggregate is where the information is. Forty failures spread across twelve codes is a generator producing generally weak output. Forty failures all carrying one code is a specific broken thing, usually upstream, usually fixable in an afternoon.

They also let you set different policies per code. Some failures should block absolutely. Some should route to a human. Some should trigger a re-observation and a retry, which is the correct response to a staleness failure and the wrong response to a policy failure.

The rule we work to is that a check may not fail without naming which rule fired and what it saw. If a rule cannot articulate that, it is not specified tightly enough to be relied on.

  • Block: factual claims with no supporting chunk, references that do not resolve
  • Re-observe and retry: staleness failures where the basis has changed since observation
  • Route to a human: policy and tone failures, which need judgement
  • Warn and allow: structural preferences that are not correctness issues

The checks that are worth writing first

If you are starting from nothing, this is the order that has given us the most protection per hour of work. None of these needs a model and all of them are testable in isolation.

The last two are the ones people leave out and they are the ones that catch the failures that reach customers.

The first validation checks worth writing

  • Exactly one h1, and no skipped heading levels
  • Every internal link resolves against the current route table, and redirects are reported separately from misses
  • Every image reference resolves and carries alt text
  • Every number, price and plan name appears in the canonical fact set
  • Every factual claim carries the id of the chunk that supports it
  • No claim is supported only by a chunk older than a contradicting canonical fact
  • Structured data is re-derived from the final content and compared with what is emitted
  • The page does not contradict any existing page on a fact both of them state

The check that catches the most: schema against content

Re-deriving the structured data from the final rendered content and comparing it against what the page emits is the single highest yield check we have.

It catches a whole class of failures that nothing else does. Content gets edited after the schema was generated. A price changes in the prose and not in the markup. A FAQ answer is rewritten and the FAQ schema still carries the old text. Every one of those produces a page that is internally inconsistent in a way a human reader will never see and an automated consumer absolutely will.

It is also the failure with the worst consequences relative to how boring it is. Markup that contradicts the page is a machine readable statement that you are unreliable, made to exactly the systems you were hoping would quote you.

Where teams get this wrong

Checks that cannot fail. A surprising number of validation suites contain rules that no real input could violate, usually because the rule was written against a hypothetical rather than an observed failure. The test for this is simple and uncomfortable: for each rule, can you produce an input that fails it? If not, the rule is reassurance.

Validation as a scoring function. Combining checks into a quality score and setting a threshold reintroduces exactly the ambiguity the checks removed. A page either may be published or may not. Scores are useful for ranking candidates, not for deciding.

Validating only generated pages. If the standard is real, it applies to human written pages too. Running it across the existing site is also the fastest way to find out whether the standard is reasonable, and it usually finds things.

Running expensive checks first. If the consistency check queries the whole site, it should not run on a page that already failed structurally. Cheap and certain first is not just about compute, it is about which failure gets reported, and the first reported failure should be the most actionable one.

How to evaluate a validation layer

These are the questions I would want asked about ours. Most of them can be answered in a few minutes by someone who built it, and the ones that cannot are the interesting ones.

The first question is the one that decides everything else. If validation runs after publication, it is monitoring.

  • Does validation run before publication, in the write path, with no branch around it?
  • For each rule, can you produce an input that fails it?
  • Does every failure carry a reason code, and are failures aggregated by code anywhere?
  • Which failures block, which retry, and which route to a person?
  • Does the same standard run against human written pages?
  • Is the structured data re-derived from final content, or generated once and trusted?

What validation cannot catch

Worth being explicit about the ceiling, because a validation layer that is trusted beyond its reach is more dangerous than one nobody trusts.

It cannot tell you the page is pointless. A structurally perfect, factually grounded, internally consistent page about something nobody asked about will pass everything. Whether a page should exist is a content strategy question and no check answers it.

It cannot tell you the argument is wrong. Claims can each trace to a source and still add up to a conclusion the company would not stand behind. That is review, and it is why review still exists.

It cannot tell you the tone is off in a way that matters. Policy checks catch banned phrasing and required disclosures. They do not catch a page that is technically compliant and reads like it was written by someone who has never spoken to a customer.

And it cannot make a stale fact set current. Every factual check is only as good as the canonical data behind it. Validation catches disagreement with what you believe. It cannot catch what you believe being out of date, which is why observation and indexing are the layers that actually determine correctness.

What good looks like

A working validation layer is unexciting. Pages either publish or come back with a specific reason. Nobody argues about whether a page is acceptable, because acceptable is defined and the definition is in code.

The observable signal that it is healthy is the distribution of failures over time. Early on, most failures are structural, because the generator has not learned the shape. Then they shift to factual and consistency, which is a grounding problem. Eventually most failures are policy, which means the system is producing correct pages that break a rule you chose. That is the good end state, and it is a conversation about policy rather than about correctness.

If your failures are still mostly structural after a few months, the generator is not the problem. The specification is.

Where we are with this

Creogen is in private development and I would rather say that than let a present tense imply a shipped product. The validation stage is one of the parts that exists, because we built it before the proposal stage rather than after, on the reasoning above.

The schema against content check is there because of a specific thing we hit while building the Webflow apps: content edited after the markup was produced, with nothing to notice the divergence. It was invisible in the interface and completely visible to anything reading the page programmatically.

None of this needs our tooling. Write the pass conditions down, order the checks cheapest first, give every failure a reason code, and run the whole thing against your existing human written pages to find out whether the standard is honest.

Written by

Vishal Chiniwar

Co-founder and CTO

Vishal builds the systems behind Creoglyph. He wrote the four Webflow Designer apps the company ships, and now works on the architecture behind Creogen and Creobot: retrieval, grounding, CMS data models, install flows and the validation that has to sit between a model and a live marketing site.

Related reading

Questions this raises

Different failures deserve different policies, which is why reason codes matter. Unsupported factual claims and broken references should block. Staleness should trigger re-observation and a retry. Policy and tone should route to a person. A single blanket answer wastes the information the codes carry.

Yes, and for the same reason as guardrails generally. Review is a vigilance task that degrades with volume, and it is poor at correctness checking on the fortieth similar item. Let the machine check correctness and the person judge quality.

It should, and running it is the fastest honesty test available. If the standard is real it applies to everything. If it fails on half your existing site, either the standard is wrong or you have just found a lot of work worth doing.

You mostly cannot, and that is the correct finding rather than a gap to fill. Tone belongs to review, not validation. What you can encode is the mechanical part: banned phrasing, required disclosures, locale conventions. Anything left over is a judgement and needs a person, which is why the two stages stay separate.

Look at the reason codes before touching the thresholds. A cluster on one code is almost always an upstream problem, usually a stale index or a fact set that is missing entries. Loosening the check in that situation ships the underlying problem instead of fixing it.

Consistently, in my experience, and it is the fastest honesty test for the standard itself. If your rules fail on half the existing site, either the rules are wrong or you have discovered real work. Both outcomes are worth having before you point the validator at generated output.

What has to be true before generated output goes live?

Write that list down and you have most of a validation layer. We will help you work out what belongs on it.

Creogen will not publish output that fails its own checks. That is the whole design.

A visitor question becoming a change on a page A question enters on the left, Creobot captures it, Creogen turns it into an operation, and one block on the page is marked as changed. Creobot Creogen QUESTION TO CHANGE