An assistant that invents your pricing is worse than no assistant
On a marketing site the cost of a wrong answer is not a bad experience. It is a commitment your company did not make, in writing, to a buyer.
Vishal Chiniwar Co-founder and CTO 26 May 2026
Start from the failure, not the feature
Consider a specific failure. A visitor asks a site assistant whether the plan includes single sign on. The assistant says yes. The plan does not include single sign on, it is an add on, and it has been an add on since a pricing change two quarters ago.
Nothing about that answer looked wrong. It was fluent, confident, on topic and phrased exactly like the correct answer would have been. The visitor screenshots it. Two weeks later that screenshot is in a procurement thread, and somebody at your company has to explain that the website said something the company did not mean.
That is the failure mode worth designing around, and notice what it is not. It is not the model producing obvious nonsense, which is embarrassing and harmless because nobody believes it. It is the model producing a plausible answer that is slightly out of date. Plausibility is what makes it dangerous.
A wrong answer on a marketing site is not a bad experience. It is a commitment nobody authorised, in writing, to a buyer.
What a wrong answer actually costs
The costs are not evenly distributed and it is worth separating them, because they call for different responses.
There is the direct commercial cost, which is a deal that proceeds on a false premise and either falls apart later or closes on terms nobody agreed. There is the support cost, which is the same wrong answer being given repeatedly to people who never contact you, so you cannot even count it. There is the trust cost, which is the visitor who catches the error and now discounts everything else on the page. And there is the internal cost, which is that the assistant gets switched off, and the switching off is permanent because nobody wants to be the person who turns it back on.
That last one is the reason to take grounding seriously up front. A site assistant does not usually get a second attempt inside the same company.
Why the model's own knowledge cannot be the source
A model trained on the general web knows something about your company, and that knowledge is a snapshot of an unspecified moment. It does not know your current pricing, your current plan names, or that you renamed the product last spring. Worse, it does not know that it does not know, so it will answer anyway with the same fluency it uses for things it is right about.
Fine tuning is sometimes offered as the answer to this and it is a poor one for a marketing site. It bakes facts into weights, so every pricing change becomes a retraining event. A pricing page changes far more often than anyone plans to retrain, and the gap between the two is exactly the window where the assistant is confidently out of date.
Retrieval separates the volatile part from the stable part. The model supplies language ability, which changes slowly. The retrieved documents supply the facts, which change constantly. When you change your pricing page, the assistant is correct within one indexing cycle rather than one training cycle.
What retrieval has to get right
Retrieval sounds simple and the naive version fails in specific ways. Three of them do most of the damage.
The first is chunking on character count. Splitting a page every eight hundred characters cuts a pricing table in half, so the retrieved chunk carries plan names with no prices or prices with no plan names, and the model helpfully fills in the gap. Chunk on document structure instead: a heading and the content under it, a table kept whole, a list of features kept with the plan it belongs to.
The second is indexing without recency. If the index holds both the current pricing page and a cached copy from before the change, similarity search will happily return the old one, because the old one is often more similar to the question. Every chunk needs the time it was indexed and the retrieval has to prefer current over similar.
The third is retrieving the page and not its contradictions. If three pages mention plan limits and two of them are stale, retrieving only the best match hides the disagreement. Retrieving all three surfaces it, and a system that can see a contradiction can decline to answer rather than pick.
| Naive approach | What goes wrong | What to do instead |
|---|---|---|
| Fixed character chunks | Tables and lists split mid structure, prices separated from plan names | Chunk on headings, keep tables and feature lists whole |
| Similarity only | Stale copies win because they are often more similar to the question | Rank on recency alongside similarity, and record index time per chunk |
| Single best match | Contradictions between pages stay hidden | Retrieve the competing pages too, and treat disagreement as a signal |
| Whole site in context | Relevant material diluted by volume, answers get vaguer | Scope retrieval to the page, its references, and the canonical fact set |
| No citation | Nothing is checkable, by the visitor or by you | Cite the source page on every factual claim |
Grounding is a refusal mechanism
Grounding is often described as giving the model context. That undersells it. The useful definition is narrower: a grounded system is one that will not make a claim it cannot attribute to a retrieved source.
The important word is will not. This is not an instruction in a prompt, because a prompt is a preference and it degrades exactly when the input is unusual, which is when you need it most. It is a check in code that sits after generation and before the answer reaches a person. If a factual claim has no supporting chunk, the answer does not go out in that form.
The behaviour this produces feels worse in a demo and is much better in production. The assistant says it does not have that information and points at the contact page. Everyone building these systems has to decide whether they are optimising for the demo or for the screenshot in the procurement thread.
Worth being precise about what gets checked. Not every sentence needs a source. Connective language, restatement of the question and general phrasing are fine ungrounded. What needs a source is anything a reader could act on: a price, a plan limit, a feature, an availability claim, a date. The check is on claims, not on words, and conflating the two produces a system that refuses to speak at all.
Citation is not decoration
Every factual answer should name the page it came from, visibly, as a link the visitor can follow.
The obvious reason is that the visitor can check. The better reason is that citation is a forcing function on the architecture. A system that has to cite cannot answer from the model's own memory, because there is nothing to point at. Requiring the citation makes ungrounded answers structurally impossible rather than merely discouraged.
It is also the cheapest debugging tool you will have. When someone reports a wrong answer, the citation tells you immediately whether retrieval fetched the wrong page or the right page said the wrong thing. Those have completely different fixes and without the citation you are guessing.
The latency tradeoff nobody warns you about
Retrieval costs time, and on a marketing site the time budget is not generous. A visitor who asks a question and waits four seconds has already decided the thing is slow, and the interaction quality drops from there regardless of how good the answer turns out to be.
The naive fix is to retrieve less, which trades correctness for speed and is the wrong direction on a page where correctness is the whole point. The better fix is to move work out of the request path. Embed at index time rather than query time. Keep the canonical fact set, the prices and plan names and product names, in a small hot structure that is looked up rather than searched. Stream the answer so the first tokens arrive while the rest is still being produced.
There is also a decision about where the assistant sits in the page's own performance budget. A site assistant is not the reason someone visited. If it costs a meaningful amount of main thread time before it is opened, it is taxing every visitor to serve the few who use it. Load it on interaction, not on page load, and keep the idle cost close to nothing.
What good looks like from the outside
A grounded assistant does not look impressive, which is worth preparing people for internally.
It answers the questions the site answers, quickly, and it names the page it took the answer from. When you ask something the site does not cover, it says so and offers a person rather than producing a paragraph. When two pages disagree, it says the site is inconsistent here rather than choosing one. It never mentions a plan name that does not exist, and it never describes a feature the product does not have.
The last two are the whole game. Everything else is polish. A system that never invents a plan name is a system a sales team can stop worrying about, and a system a sales team has stopped worrying about is one that stays switched on.
Where teams get this wrong
Four patterns, all common, all survivable if caught early.
Treating retrieval as an optimisation. It gets scheduled after launch because the demo works without it. The demo works because the demo asks questions the model happens to know. Real visitors ask about your plan limits.
Indexing everything. Blog posts from three years ago, old release notes, careers pages. All of it competes with the pricing page in similarity space, and some of it is confidently obsolete. A smaller index of current pages beats a large index of everything.
Measuring on answer quality alone. An assistant can be rated highly by users while being wrong about pricing, because users cannot tell. The metric that matters is what proportion of factual claims carry a valid citation, and that one you can compute.
Letting it answer everything. Refusal has to be a first class outcome with its own path, not a fallback that feels like failure. Legal questions, commitments about roadmap, anything about a customer's specific contract: these should route to a person by design.
How to evaluate an assistant before it goes live
This is the set I would run against any site assistant, ours included. It takes an afternoon and it is more informative than a general quality score.
The last one is the one that catches the most systems. Almost anything will handle a question whose answer is on one page. The interesting behaviour is what happens when two pages disagree.
- Ask about something that changed recently. Does it use the current page or a stale copy?
- Ask something the site genuinely does not answer. Does it refuse, or does it improvise?
- Check that every factual claim carries a citation you can click and verify.
- Ask the same question five ways. Do the answers stay consistent with each other?
- Point it at two pages that contradict each other. Does it notice, or does it pick one silently?
Where this sits in what we are building
Creobot is the conversation engine we are building and it is in private development, so treat everything here as the reasoning behind it rather than a description of a shipped product.
The constraint we set at the start was that it answers from the site and says which page it answered from, and that when it cannot ground an answer it says so and offers a person. That constraint is what shaped the retrieval design rather than the other way round. The other half is that the questions it cannot answer are the most valuable output, because a question asked repeatedly with no page behind it is a content gap with evidence attached. That is what feeds Creogen.
None of this requires our tools. Structure the index on document structure, prefer current over similar, cite on every factual claim, and make refusal a real path. That is most of the distance, and it is available with commodity components.
Related reading
A prompt is not a guardrail
Instructions in a prompt are a preference. A guardrail is something that can refuse, that leaves a record when it refuses, and that you can write a test against.Vishal Chiniwar19 May 2026A question is unstructured until you give it a shape
Free text does not aggregate. Ten thousand visitor questions in a table is not data, it is a transcript, and nobody reads transcripts.Vishal Chiniwar6 July 2026Architecting a system that has to be right about someone else's website
Every hard problem in a website operations platform is the same problem: the system is reasoning about state it does not own and cannot lock.Vishal Chiniwar13 May 2026
Questions this raises
Not for facts that change. Fine tuning bakes information into weights, so every pricing change becomes a retraining event. Retrieval keeps the volatile facts in documents you can update today and have correct within one indexing cycle.
That usually means the index is missing pages rather than that grounding is too strict. Look at the refusals as a list. If people are asking things your site does not answer, the fix is a page, not a looser threshold.
A single source link per factual claim is not clutter, it is the thing that makes the answer checkable. It is also the fastest way to diagnose a wrong answer, because it tells you whether retrieval or the source page was at fault.
On publication rather than on a schedule, if you can hook it. A time based rebuild guarantees a window where the assistant is confidently out of date, and the length of that window is whatever interval you chose. If you can only do scheduled rebuilds, make the interval shorter than the rate at which your pricing changes.
Say the site is inconsistent on this point and offer both, or decline and route to a person. Picking one silently is the worst option, because it hides a real problem from the only party who could fix it. A contradiction surfaced is a content bug report arriving for free.
Index what is current and authoritative, not everything you have published. Old posts compete with the pricing page in similarity space and some of them are confidently obsolete. A smaller index of pages you would stand behind today beats a large index of everything you have ever written.
Ask about something that changed recently and check whether the answer uses the current page. Ask something the site genuinely does not cover and see whether it refuses or improvises. Then point it at two pages that contradict each other, which is the test almost every system fails.
An assistant that invents your pricing is worse than no assistant.
Grounding is not a refinement, it is the product. Tell us what your assistant is answering from and we will look at where it can drift.
Creobot answers from your site and says so. Retrieval and citation are not optional there.