Structured data is a promise, and the page has to keep it

Markup that disagrees with the page is worse than no markup, because it is checkable. You have told an automated reader you are unreliable.

Vishal Chiniwar Co-founder and CTO 10 June 2026

Share
Visible page content mapped to structured entitiesEach visible block of a page maps to a structured data entity, with the highlighted pair showing that schema must match what is shown.MATCHES WHAT IS VISIBLEPAGESTRUCTURED ENTITIES

The failure that costs the most

Two representations of the same document, emitted from the same request, disagreeing with each other. That is the shape of the most common structured data failure, and it is worth starting there rather than with schema types, because it explains why more markup is not the remedy.

A page says a plan costs one figure. The JSON-LD on the same page says another. Both are emitted to every automated reader that requests the URL, and one of them is wrong.

This happens constantly and it is almost never noticed, because no human sees both. The prose was edited during a pricing change. The markup was generated when the page was built and has been carried forward untouched ever since. The visual page is correct. The machine readable page is two quarters out of date.

This is worse than having no structured data at all, and the reason is specific. Absence of markup is neutral: a reader falls back to parsing the page and gets the right answer. Contradictory markup is a claim, made in a format designed to be trusted, that turns out to be false when checked against the same document. You have handed a checkable inconsistency to exactly the systems whose job is to decide whether to rely on you.

The check for this takes about a minute per page and almost nobody runs it: open the page, open the emitted JSON-LD, and read the two side by side looking only at numbers and names. On most sites of any age you will find at least one disagreement, and it will be in a field that changed at some point without the markup following.

Schema is the last step, not the first

Most structured data work starts by choosing types and filling in properties. That ordering produces markup that describes a page rather than describing the thing the page is about, and the difference matters more than it sounds.

The prior questions are: what entities does this site actually contain, what relationships hold between them, and which URL is authoritative for each. Almost no site has these written down. They exist implicitly in the information architecture, in the CMS collections, and in the heads of two or three people.

Once they are explicit, the markup is close to mechanical, because schema types are a vocabulary for expressing exactly those things. If the entities are unclear, no amount of care with property names fixes it, and you end up with technically valid markup that asserts very little.

What an answer engine is actually doing with this

It helps to be concrete about the consumer, because a lot of structured data work is done with a vague idea of who is reading it.

A system trying to answer a question from your site has to do three things. Find the passage that contains the claim. Decide the claim is about the entity in the question rather than a similar one. Decide your page is a source worth attributing.

Structured data helps mainly with the second and third. It says explicitly that this page is about this entity, which resolves the ambiguity between your product and a similarly named one, and between your Business plan and somebody else's. And relationships let a reader answer a question that no single passage on your page states, such as which plan a feature belongs to, by combining two things you did state.

It helps very little with the first, which is a property of how the page is written. That is why markup on an unstructured page changes almost nothing, and why the entity work has to come with page structure work rather than instead of it.

Entities: what your site is actually about

An entity is a thing that has an identity independent of the page describing it. Your product is an entity. Each pricing plan is an entity. Each integration, each person on the team, each article, each glossary term.

A page is not an entity. A page is a document that describes one or more entities, and conflating the two is the most common modelling error. When you treat the page as the thing, you cannot express that two pages describe the same product, and you cannot express that a product exists whether or not it currently has a page.

The practical exercise takes about an hour. List the nouns your company uses in sales conversations. Product. Plan. Integration. Use case. Customer type. Feature. Then for each, ask whether it has an identity separate from any page, and whether more than one page mentions it. The ones that pass both are your entities, and they are usually already sitting in CMS collections without having been named as such.

Entities, relationships and where each should live
EntityAuthoritative URLKey relationshipsSchema type
ProductOne product pageoffers plans, has features, integrates withSoftwareApplication or Product
PlanA section of the pricing pagebelongs to product, includes featuresOffer
IntegrationOne page per integrationconnects product to third partyProduct or a defined term, by case
ArticleOne article pagewritten by person, about topicArticle
PersonOne team or author pageauthor of, works forPerson
Glossary termOne anchor on one pagerelated to other termsDefinedTerm
OrganisationHomepagepublisher of, employsOrganization

Relationships are where the value is

Entities alone give an automated reader a list of nouns. Relationships give it something it can reason with, and reasoning is what an answer engine is doing when it decides whether your page answers a question.

The relationships that matter most for a marketing site are unglamorous. Which plans belong to which product. Which features belong to which plan. Which integrations connect to which product. Who wrote which article. Which article is about which topic.

Expressing a relationship in schema is done by reference rather than by repetition: the plan carries an identifier pointing at the product, not a copy of the product's description. This is the same principle as canonical facts in the prose, and for the same reason. Repetition creates opportunities to disagree.

Most sites express none of these. They emit an Organization block on every page and nothing else, which asserts that a company exists and stops there.

Modelling the relationships you actually have

Once entities are named, the relationships are usually obvious and it is worth writing them as sentences before writing them as markup. Product offers plan. Plan includes feature. Product integrates with third party. Person writes article. Article is about topic.

Two rules keep this from sprawling. Model only relationships that a buyer would ask about, because those are the questions an answer engine is trying to satisfy on your behalf. And model in one direction, from the more specific to the more general, so a plan points at its product rather than a product listing all its plans inline. Bidirectional description doubles the opportunities to disagree.

The relationship that gets missed most often is the one between a feature and a plan. Sites list features on the product page and plans on the pricing page, and never state which features belong to which plan anywhere a machine can read. That is precisely the question buyers ask most, and it is why so many assistants get plan availability wrong on so many sites, including for products whose pricing pages are perfectly clear to a human.

One entity, one canonical URL

If four pages describe the same integration, an automated reader has to decide which one to believe. Usually they will not agree, because they were written at different times.

The rule is one entity, one authoritative URL, and every other mention references it rather than restating it. This is exactly the same decision as deciding which page canonically states your pricing, and it pays in both directions: the prose stops contradicting itself and the markup stops splitting the signal.

The failure is easy to detect. Search your own site for a plan name and count the pages that state its price. If the answer is more than one, you have a canonicalisation problem before you have a schema problem, and adding markup to all of them makes it worse rather than better because now the disagreement is machine readable.

Derive the markup, never author it

The single change that prevents most structured data rot: generate the markup from the same source as the rendered content, in the same build, and never let a person edit it independently.

If the price in the markup comes from the same field as the price in the prose, they cannot disagree. If somebody types the price into a schema block by hand, they will disagree, not immediately but eventually, and eventually is about two quarters.

Then verify the derivation rather than trusting it. Re-derive the structured data from the final rendered content and compare it with what the page emits. That comparison catches edits that happened after generation, which is the case hand written markup cannot survive and derived markup can still be caught out by if the pipeline has a gap.

This is the check that has caught the most on our own site, and it is entirely boring.

The page underneath still has to be structured

Markup describes structure. It does not create it. A page whose claims are buried inside persuasive paragraphs is not made extractable by wrapping it in JSON-LD, and this is where a lot of structured data effort is wasted.

If a plan's feature list exists only as prose in a paragraph, emitting an Offer with an itemised feature list is an assertion the page does not visibly support. It may pass validation. It is still describing something a reader cannot see, and the gap between the two is exactly the kind of thing automated readers are increasingly checking.

The order that works is: make the claim exist somewhere with a boundary on the visible page, in a table row, a list item, or a short paragraph under a plain heading. Then describe that structure in markup. The markup is a description of a real thing rather than a claim about an absent one.

Where teams get this wrong

Marking up everything. Every type in the vocabulary applied to every page produces a lot of assertions, most of them thin, and a reader has no way to tell which ones you meant. Fewer, accurate, well related blocks are worth more than comprehensive coverage.

Treating validation as correctness. A schema validator checks that your markup is well formed and uses properties correctly. It cannot tell you the price is wrong, that the entity does not exist, or that four pages claim to be authoritative for the same thing. Passing validation is the floor.

Authoring markup separately from content, which is the rot mechanism described above and worth repeating because it is the most common.

Chasing rich result formats. Building the entity model around whatever visual treatment a search engine currently rewards produces a model that is wrong the next time the treatment changes. Model the things your business actually contains. Those are stable.

How to evaluate what you have

This is the pass to run on an existing site. It is mostly reading rather than tooling, and it takes an afternoon.

The first two items catch the majority of real problems on most sites.

A structured data audit that finds real problems

  • For each page with markup, compare every factual value in the markup against the same value in the visible content
  • Search the site for each plan name and count the pages that state its price
  • List your entities and check each has exactly one authoritative URL
  • Check whether any relationships are expressed, or only isolated entities
  • Confirm the markup is derived from the same source as the content, not typed separately
  • Find one claim in the markup and check the visible page states it with a boundary around it
  • Check that removing a page also removes its entity assertions from anywhere else

What good looks like

A site doing this well emits less markup than you might expect, and it is all load bearing.

There is one authoritative page per entity. Relationships are expressed by reference rather than by repeated description. Every value in the markup appears identically on the visible page. The markup is generated in the build from the same fields as the content, and a check re-derives and compares it before publication.

The behavioural test is that changing a price is one edit, in one place, and both the prose and the markup on every affected page change with it. If that takes more than one edit, the structure has not been done, whatever the markup looks like.

Where this shows up in what we build

This site emits schema derived from the same JSON that renders the content, and the build compares the two. That is not a claim about a product, it is a property of this site you can check by reading the source.

Creogen is being built around the entity question rather than the markup question: which pages describe the same thing, which of them disagree, and which is supposed to be authoritative. Markup follows from that and is the easy part. It is in private development.

Most of what is above needs no tooling at all. Write down your entities, give each one an authoritative URL, derive the markup from the content, and compare the two before publishing. That is an afternoon and most of the value.

Written by

Vishal Chiniwar

Co-founder and CTO

Vishal builds the systems behind Creoglyph. He wrote the four Webflow Designer apps the company ships, and now works on the architecture behind Creogen and Creobot: retrieval, grounding, CMS data models, install flows and the validation that has to sit between a model and a live marketing site.

Questions this raises

No. Fewer accurate blocks with real relationships are worth more than comprehensive coverage of thin assertions. A reader cannot tell which of a hundred assertions you actually meant.

Validation checks that the markup is well formed and uses properties correctly. It cannot tell you a price is wrong, that four pages claim to be authoritative for the same entity, or that the visible page does not support the claim. Passing is the floor.

Only for what the page genuinely contains, usually the organisation and the breadcrumb. Asserting itemised structure that the visible page does not show is a claim you cannot support, and it is checkable.

List the nouns your company uses in sales conversations. Product, plan, integration, use case, feature, customer type. For each one, ask whether it has an identity separate from any page and whether more than one page mentions it. The ones that pass both are your entities, and they are usually already sitting in CMS collections.

No, and it is actively harmful. Markup that asserts structure the visible page does not show is a checkable claim you cannot support. Fix the page first so the claim exists with a boundary around it, then describe that structure in markup.

Open the page, open the emitted JSON-LD, and read them side by side looking only at numbers and names. It takes about a minute per page and on most sites of any age you will find at least one disagreement in a field that changed without the markup following.

No. Model from the more specific to the more general, so a plan points at its product rather than a product listing all its plans inline. Bidirectional description doubles the number of places that can disagree, which is the same reason facts get one canonical home in the prose.

Structured data is a promise you have to keep.

Markup that disagrees with the page is worse than no markup. Bring a page and its schema and we will check they still match.

Every page on this site emits schema that matches what is rendered. It is validated in the build.

A visitor question becoming a change on a page A question enters on the left, Creobot captures it, Creogen turns it into an operation, and one block on the page is marked as changed. Creobot Creogen QUESTION TO CHANGE