Nobody wrote down what the change was supposed to do
Which is why nobody can say whether it worked. The measurement problem on most sites is not instrumentation. It is that no expectation was recorded before the change.
Vishal Chiniwar Co-founder and CTO 30 June 2026
The missing artifact
There is no field for it. Look through whatever your team uses to track website work, and find the place where somebody records what a change is supposed to cause. On most setups that place does not exist, which means the question of whether a change worked is being asked of a system that was never built to answer it.
A change ships. The ticket closes. Six weeks later somebody asks whether it worked, and the honest answer is that nobody can tell, because nowhere in the process was there a place to write down what it was supposed to do.
This is usually described as a measurement problem and treated with more analytics. It is not a measurement problem. It is a missing data structure. There is no field, anywhere in most teams' workflow, that holds the sentence this change should cause X. So afterwards there is nothing to compare against, and any number that moved can be narrated as the result.
Adding that field costs nothing and is the single highest yield change available. Everything below assumes it exists, because without it the rest is decoration.
Four layers, four different questions
Measurement on a website is not one thing, and the most common failure is owning one layer and trying to make it answer questions that belong to another.
Each layer has a characteristic question, a characteristic latency, and a characteristic way of being wrong. The table is the version I sketch when somebody says their analytics does not tell them anything useful, because usually the analytics is fine and is being asked a question from a different row.
| Layer | Question it answers | Latency | How it misleads |
|---|---|---|---|
| Delivery | Did the change actually ship, everywhere, and is it still there? | Minutes | Assumed rather than checked, so a reverted change looks like a failed change |
| Behaviour | Did people interact with the thing differently? | Days | Sensitive to traffic mix, so a campaign change looks like a page change |
| Outcome | Did the thing we care about commercially move? | Weeks to months | Too far downstream to attribute to one page with any confidence |
| Qualitative | Did the friction we were targeting stop happening? | Days | Anecdotal, easy to over read, and the only layer that explains why |
Start with delivery, because it is free and it is wrong more often than you expect
The cheapest layer is the one almost nobody instruments. Did the change actually ship, is it live on every page it should be on, and is it still there today.
This sounds trivial. In practice a meaningful proportion of changes that appear not to have worked were reverted by a subsequent deploy, never applied on a template variant, or applied on desktop and not on the responsive layout. Anyone who has looked has found at least one.
The check is a snapshot comparison: assert the specific change is present on the specific set of URLs, on a schedule. It takes an hour to build, it is deterministic, and it converts a whole class of confusing results into a boring answer.
Do this before adding a single analytics event. Sophisticated measurement of a change that is not live is an expensive way to be confused.
Event quality is the ceiling on everything above it
The behaviour layer is only as good as the events under it, and event quality on most marketing sites is poor in ways nobody has audited.
The recurring problems are consistent enough to list. Events that fire on render rather than on interaction, so they count impressions and are named as clicks. Duplicate fires from a component that mounts twice. Events lost to consent tooling in some regions, silently, so a geography looks like it stopped converting. Names that changed meaning when somebody reused an event for a new element rather than adding one.
Any of those makes the analysis above it confidently wrong rather than uncertain, which is the worse failure. Uncertainty prompts caution. Confident wrongness prompts a decision.
The audit is unglamorous: for each event you intend to rely on, trigger it manually and confirm it fires once, at the right moment, with the right properties, on both breakpoints, with consent granted and denied. Half a day for the events that matter.
- Fires on the interaction, not on the element rendering
- Fires exactly once per interaction, verified on a component that mounts twice
- Carries the identifier of the page and the specific element, not just a page path
- Behaves predictably when consent is denied, and that behaviour is known rather than assumed
- Has not been reused for a second purpose since it was named
The baseline problem, and the failure that looks like a result
Comparing after against before requires a before, and on most sites the before is not what people think it is.
Traffic mix moves constantly. A campaign starts, a post gets shared, a competitor changes their bidding, a seasonal pattern arrives. Any of those changes who is arriving, and the same page converts differently for different audiences. So a comparison of last month against this month is comparing a page change and an audience change at once, attributed entirely to the page.
Two practical mitigations, neither of which is a proper experiment. Take a longer baseline than feels necessary, so ordinary variation is visible and you can see whether your effect is inside it. And check a control surface, some part of the site that did not change, over the same window. If the unchanged part moved similarly, you measured the week rather than the change.
Neither of these gives you causality. Both stop you claiming it.
A change ships on a Tuesday. The number moves that week. The change gets recorded as successful and the approach gets repeated. What actually happened is that a campaign started on the Wednesday, or a mention landed somewhere, or the previous week was a public holiday in the largest market.
The control surface is the cheapest defence: some part of the site that did not change, checked over the same window. If it moved by a similar proportion, you measured the week.
The more uncomfortable defence is to look at the same window in the previous two periods before concluding anything. Most weekly movements on a marketing site sit comfortably inside ordinary variation, and seeing that plotted once tends to permanently change how a team talks about results.
What attribution cannot do, stated plainly
This is the part where most measurement conversations go wrong, and it is worth being blunt because the pressure to overclaim is considerable.
On a marketing site with typical traffic, a single page change will usually not produce an effect large enough to resolve from noise within a reasonable window. That is not a failure of your tooling. It is arithmetic. Small effects on modest sample sizes take a long time to become visible, and by then several other things have changed.
Which means the honest answer for most changes is that we cannot tell, and the useful response is not to buy better analytics. It is to stop pretending the outcome layer will answer, and use the layers that will.
The pages where outcome measurement genuinely works are the few with enough volume to resolve anything, usually a pricing page or a primary signup path. For those, run a real test. For everything else, measure delivery, behaviour and qualitative, and be clear that you are measuring whether the change did the thing it was designed to do rather than whether it made money.
The layer everyone dismisses is often the only one that works
The qualitative layer gets treated as anecdote, and it is, and on most changes it is the only layer that will give you an answer inside a quarter.
If you added a section because sales answered the same question every week, the measurement is whether sales still answers it every week. Ask them. That is a real observation about a real behaviour, it arrives in days rather than months, and it is not confounded by traffic mix.
The same applies to support tickets on a topic, to the workaround links people send, and to whether a specific objection still comes up on discovery calls. These are countable, they are attributable to the change because the causal chain is short and visible, and no dashboard will ever show them.
The reason this feels unsatisfying is that it is not scalable and it does not produce a chart. The reason it is worth doing anyway is that it answers the question, which the chart usually does not.
Instrumentation that is worth building, in order
If you are starting from an ordinary analytics install, this is the order that has given us the most answer per hour of work.
First, the delivery assertion. A scheduled check that named changes are present on named URLs. It is deterministic, it needs no statistics, and it removes the most common cause of confusing results.
Second, a change log with the expectation field. Not a tool, a table: date, URL, what changed, what it should cause, decision rule, review date, outcome. Six columns and a discipline.
Third, an event audit of the ten events you actually use. Not all of them, the ten. Most sites have a few hundred events and rely on about ten, and the other few hundred are noise that makes the ten harder to find.
Fourth, a qualitative route: an agreed way for sales and support to say this question is still coming up, that lands somewhere a person reads. This is a process rather than software and it is the one that produces answers fastest.
Proper experimentation comes fifth, and only for the two or three surfaces with enough volume to support it. Building it first is the common error and it produces a well instrumented system that still cannot tell you whether last week's change helped.
Write the decision rule before you look
The last piece is the one that makes the rest honest. Before shipping, write what you will do at each outcome, including the outcome where nothing happens.
A decision rule looks like: if the question stops appearing in sales calls within six weeks, keep it and do the same for the next two questions. If it still appears at the same rate, the page is not the problem and we stop working on the page. If we cannot tell, we treat that as no effect and move on rather than extending the window.
That last clause matters more than it looks. Without it, the default behaviour when a result is ambiguous is to keep looking until it resolves in the direction somebody wanted. Writing down in advance that ambiguous means no effect removes an entire genre of self deception.
Write it in the same place as the expectation, at the same time, before the change ships.
A measurement checklist that is actually usable
This is the version that survives contact with a normal team. It is deliberately short, because a long measurement checklist gets skipped entirely and a short one gets done.
The first two items are worth more than the rest combined.
Before a change ships
- Write the expectation as a sentence a person could check, not a metric target
- Write the decision rule, including what you will do if the result is ambiguous
- Name which of the four layers will answer this, and accept that it may be qualitative
- Assert that the change is live on every URL it should be on, on a schedule
- Confirm the events you will rely on fire once, on interaction, on both breakpoints
- Choose a control surface that did not change, and check it over the same window
- Set the review date now, in somebody's calendar, and name who does the review
Where teams get this wrong
Retrofitting the expectation. Writing what the change was meant to do after seeing the result is not measurement, it is narrative construction, and it will always succeed.
Metric targets instead of expectations. As soon as the record says conversion will improve by some figure, it becomes something to defend rather than something to examine, and the honest outcomes disappear from the record.
Measuring the site rather than the change. Site level dashboards move for a hundred reasons. Attributing that movement to whichever change happened to ship that week is guessing with a chart attached.
Never recording a negative. A record where every change worked is not believable and, worse, is not useful, because it contains no information about what to stop doing. The value of the record is entirely in the changes that did not help.
What good looks like
A team measuring well does not have a better dashboard. It has a short written record where each entry has a date, a URL, an expectation, an outcome, and a decision.
Some of those outcomes say it worked. Some say it did not. A meaningful number say we could not tell, and those are recorded as no effect rather than argued about. Nobody is defending a number.
The observable difference over a couple of quarters is that the team stops repeating changes that did not work, which sounds obvious and is genuinely rare. Without a record, the same idea gets proposed again in eighteen months by somebody new, and there is nothing to point at.
How this connects to what we are building
Creogen records the expectation at the point a change is proposed rather than after, and ties it to a URL and a decision. It is in private development and I will not describe it as more finished than that. The design choice worth naming is that the expectation field is not optional, because an optional field for something nobody enjoys writing is an empty field.
Creobot feeds the qualitative layer, which is the layer this article argues is underused. A question asked repeatedly with no page behind it is a measurable thing that stops being asked, and that is the shortest causal chain available on a website.
None of this needs tooling. The expectation, the decision rule and the review date are three lines in whatever you already use. The delivery check is an hour of scripting. Those four things are most of the value in this article.
Related reading
The retainer conversation goes badly because the report is about you
Hours and tickets tell a client you were busy. They do not tell a client the site is better. Those are different questions and only one of them is theirs.Sachin Aathreyaa K M22 May 2026Websites do not need another redesign. They need an owner.
The three year redesign cycle is not a design problem. It is an ownership vacuum, and it produces the same site twice for the same reason.Sachin Aathreyaa K M11 May 2026Your buyers already told you what the site is missing
The most valuable data on your website is not in the dashboard. It is in the questions people ask before they buy, and almost nobody collects them.Sachin Aathreyaa K M4 June 2026
Questions this raises
Most sites do not, for most changes, and that is worth accepting rather than working around. Use the delivery and qualitative layers, which work at any volume, and reserve outcome measurement for the two or three pages with enough traffic to resolve an effect.
It is a hypothesis written in terms somebody could check without a dashboard. The distinction matters because a hypothesis stated as a metric target becomes something to defend, and one stated as an observable behaviour stays something to examine.
Decide that before shipping and write it down, along with what you will do if the answer is still ambiguous when the window closes. Choosing the window afterwards is how ambiguous results become positive ones.
Record that as the outcome. We could not tell is a real result and treating it as one is what stops ambiguous results being read as successes. Write the decision rule in advance so that unresolvable means no effect, and move on rather than extending the window until the answer arrives.
It is weaker than a real experiment and considerably better than nothing. If an unchanged part of the site moved by a similar proportion over the same window, you measured the week rather than the change. It cannot establish causality and it can stop you claiming it.
Only if somebody has checked them by hand. The recurring problems are events firing on render rather than interaction, duplicate fires from a component that mounts twice, and losses to consent tooling in some regions. Each makes the analysis above it confidently wrong rather than uncertain, which is the worse failure.
Nobody wrote down what the change was supposed to do.
Which is why nobody can say whether it worked. Bring a recent change and we will reconstruct the expectation you should have recorded.
Creogen records the expectation before the change, not after. It is the cheapest part and the one always skipped.