A question is unstructured until you give it a shape
Free text does not aggregate. Ten thousand visitor questions in a table is not data, it is a transcript, and nobody reads transcripts.
Vishal Chiniwar Co-founder and CTO 6 July 2026
Why the transcript is not the deliverable
Two records arrive in the same table. One is a string with a timestamp. The other is a normalised phrase, a group id, an intent, a source entity, a confidence and a review state. Both came from the same visitor typing the same sentence. Only the second one can be counted, sorted or acted on.
Collecting visitor questions is easy. Any chat widget, search box or contact form produces them. What comes out is a chronological list of free text, and a chronological list of free text is not something anybody can act on.
You cannot count how many people asked the same thing, because they asked it in different words. You cannot tell which page they were on when they asked, unless something recorded it. You cannot tell whether the question is still relevant, because a question about a feature that shipped last quarter looks identical to one about a feature that does not exist.
So the raw collection is a transcript. The useful artifact is a set of signals: normalised, grouped, attributed to a page, with a state. Getting from one to the other is the entire engineering problem, and it is a data modelling problem rather than a model problem.
The distinction that matters is between a record of an event and a statement about the world. A transcript row is an event: this string arrived at this time. A signal is a statement: some number of people want to know this thing, on this page, and it is currently unanswered. The second is what a person can act on, and everything below is about the transformation between them.
The schema is the product decision
Everything downstream inherits the schema, and anything it fails to record is permanently unavailable. So this is worth getting close to right before collecting at volume, because retrofitting a field means the historical records do not have it.
The shape below is the working version. It is deliberately small. Every field earns its place by being something a downstream decision depends on, and I have removed several that seemed obviously useful and turned out never to be read.
| Field | Type | What it is for | What breaks without it |
|---|---|---|---|
| id | stable identifier | Referring to this signal after it merges or splits | Merged signals lose their history |
| raw | original text, verbatim | Re-deriving everything else when classification changes | You can never reprocess, only reclassify going forward |
| normalised | canonical phrasing | Grouping and display | Duplicates that differ by punctuation |
| group_id | reference | Counting how many people asked the same thing | No aggregation, so no priority |
| intent | enum | Separating buying questions from support questions | Support noise drowns pre sales signal |
| category | enum | Routing to the person who owns that area | Everything lands in one queue |
| source_url | string | Knowing which page failed to answer | You know the gap, not where to fix it |
| answered | boolean and reference | Whether a page covered it at the time | Cannot distinguish a gap from a findability problem |
| confidence | float | How sure the classification is, not the answer | Low confidence rows treated as fact |
| observed_at | timestamp | Staleness and trend | A question from last year ranks with one from today |
| review_state | enum | Whether a person has looked at this group | The list becomes untrusted and then ignored |
Normalisation, and what it must not throw away
Normalisation makes two phrasings of one question comparable. Lowercasing, stripping punctuation, expanding contractions, resolving obvious synonyms against your own product vocabulary.
The rule that matters is that the raw text is kept forever and normalisation is derived. Teams routinely normalise in place because storage feels wasteful, and then the classification scheme changes six months later and there is nothing to reprocess. Keeping raw means every downstream decision is revisable. Losing raw means every downstream decision is permanent.
The other constraint is that normalisation must not resolve things that are genuinely different. Do you support single sign on and do you support single sign on on the starter plan are not the same question, and an aggressive synonym pass will merge them. When in doubt, do less, because grouping downstream can merge and cannot unmerge.
Grouping is where the risk lives
Grouping decides that several questions are the same question. It is where the value is, because a count is what turns forty transcripts into one priority, and it is where the damage is, because a wrong merge is invisible afterwards.
A wrong split is survivable. Two groups that should be one produce two smaller counts, somebody notices the duplicate eventually, and merging them is a cheap operation.
A wrong merge is not. Two genuinely different questions in one group produce a count that looks important and an answer that satisfies neither. Worse, once merged, the distinction is not visible in the interface, so nobody can see there was ever a question there.
The design consequence is to bias toward splitting. Set the similarity threshold conservatively, make merging a cheap explicit action a person can take, and never merge automatically above a certain group size without review. Grouping should be a suggestion the system makes, not a fact it asserts.
This is also why group_id is a reference rather than a rewritten field. The individual signals keep their identity, so an incorrect grouping decision can be reversed without losing anything.
There is a measurement worth taking on your own grouping, and it is cheap. Take fifty grouped signals, have a person read each group, and count how many contain a question that does not belong. That number is your merge error rate and it is the only honest thing you can say about grouping quality. Anything derived from similarity scores is measuring the algorithm against itself.
Confidence about what, exactly
Confidence is the field most often misused, because there are two different things it could mean and systems routinely record one and display the other.
It should mean: how sure are we that this classification is right. This question is a pricing question with confidence 0.6 is a useful statement about a routing decision.
It should not mean: how sure are we that the answer given was correct. That is a separate property, it belongs to the answer rather than to the signal, and conflating them produces an interface where a low number could mean either we are not sure what they were asking or we are not sure what we told them.
The practical rule is that confidence gates routing and nothing else. Below a threshold, a signal goes to a human queue rather than into a category. It never gates whether the signal is recorded, because a question we could not classify is still a question somebody asked, and the unclassifiable ones are often the interesting ones.
Intent categories that survive contact with real questions
The intent enum is the field teams get wrong most often, usually by making it too fine grained on day one. A taxonomy with twenty intents looks thorough and produces classification confidence in the low fives, because the boundaries between neighbouring categories are not real.
Five intents have held up for us. Pre purchase, meaning somebody deciding whether to buy. Capability, meaning does the product do this specific thing. Commercial, meaning price, plans, contracts and terms. Operational, meaning an existing user with a task. And unclassified, which is not a failure state but a genuine category with its own queue.
The value of keeping it coarse is that the boundaries are defensible, so classification confidence stays high and routing is reliable. Sub categories can be derived later from the group, once you have enough of them to see the real clusters, and derived sub categories are revisable in a way that a day one enum is not.
The one distinction worth adding early, because it changes who reads the signal, is whether the asker is a prospect or a customer. That is often inferrable from the page they were on, which is another reason source URL earns its place in the schema.
Source URL is the field that makes it actionable
A signal without a source URL tells you a question exists. A signal with one tells you which page failed to answer it, which is the difference between a research finding and a task.
It also lets you distinguish two situations that look identical in aggregate. Fifty people asking about pricing from the pricing page is a clarity problem on a page that exists. Fifty people asking about pricing from a product page is a linking problem, or a sign that the product page raises the question and does not route to the answer. The remedies are completely different and only the source URL separates them.
Record the full path including the fragment where you have it, and record it as the URL at the time of observation, resolved against the entity id if you have one. Pages move, and a signal attributed to a URL that no longer exists is a signal you cannot act on.
Staleness, and why signals have to expire
A signal is a statement about a moment. Somebody asked this, on this page, when the site said what it said then.
If you ship a page that answers a question, the historical signals about that question do not disappear, and if nothing distinguishes them, your list still shows it as a top gap forever. Teams then either stop trusting the list or manually delete history, and both of those are worse than the original problem.
The mechanism that works is not deletion. It is a decay on ranking plus an event that marks a group as addressed at a point in time. After that, the count that matters is the count since it was addressed. The historical count is still there and is now the thing that tells you whether the fix worked, which is a genuinely useful second life for the same data.
A group that keeps accruing signals after being marked addressed is the most valuable row in the whole system. It means somebody wrote a page and the question is still being asked, which is the findability or credibility distinction arriving as data rather than as opinion.
Review state, or the list nobody trusts
The last field is the one that seems administrative and determines whether the whole thing survives contact with a team.
Without a review state, every visit to the list starts from scratch. You cannot tell which groups somebody has already looked at and decided not to act on, so the same low value items get re evaluated every time, and after about three visits people stop opening it.
Four states are enough: new, triaged, actioned, dismissed. Dismissed needs a reason, because a dismissal without one is indistinguishable from an oversight and will be re litigated.
This is not sophisticated and it is the difference between a signal pipeline that gets used and one that becomes a dashboard nobody has opened since the demo.
Where teams get this wrong
Classifying at collection time and discarding raw. This forecloses every future improvement, because the classification scheme will change and there is nothing to reprocess against.
Automatic merging with a permissive threshold. It produces impressive looking aggregation and quietly destroys distinctions. Wrong merges do not announce themselves.
Treating unclassifiable questions as errors. A question the system could not categorise is frequently the most valuable one, because it is about something nobody anticipated. Route it to a person rather than dropping it.
Building the dashboard before the schema. The dashboard needs a shape, so the shape gets designed to fit the dashboard rather than to record what happened, and the fields that turn out to matter later were never captured.
No expiry mechanism. The list is correct for a quarter and then permanently shows a question that was answered in month two.
How to evaluate a signal pipeline
These are answerable in a few minutes by anyone who built one, including us.
The second one is the question that separates a pipeline from a transcript viewer.
Questions that tell you whether the signals are real
- Is the raw text kept verbatim, and can you reprocess history against a new classification?
- Can you tell how many distinct people asked the same thing, and how sure you are of that grouping?
- Does confidence describe the classification or the answer, and is that consistent everywhere?
- Does every signal carry the page it came from, resolved to an entity rather than a raw URL?
- What happens to a group after it is addressed, and can it accrue new signals afterwards?
- Can a person see which groups have already been triaged and dismissed, with reasons?
- What happens to a question the system cannot classify?
Where Creobot actually is
Creobot is in private development and it is not magic. It is a conversation engine that answers from the site, and the schema above is the thing behind it that matters more than the answering does.
The design commitment worth naming is that a question the assistant could not answer is treated as a first class output rather than as a failure to be minimised. An assistant optimised purely for answer rate has an incentive to produce something plausible for questions it cannot ground, which is the failure mode described elsewhere on this site. Recording the miss instead is the more useful behaviour and it is what makes the signal pipeline worth having.
The part that is genuinely hard is grouping, and I would not claim we have solved it. The current position is conservative thresholds and cheap human merging, which produces more groups than are strictly correct and no invisible wrong merges. That is a deliberate trade rather than a limitation we are hiding.
None of this requires a model to start with. The schema above works with a spreadsheet and a person doing the grouping by hand for the first few hundred, and doing it by hand first is the fastest way to learn whether your intent categories are the right ones.
Related reading
An assistant that invents your pricing is worse than no assistant
On a marketing site the cost of a wrong answer is not a bad experience. It is a commitment your company did not make, in writing, to a buyer.Vishal Chiniwar26 May 2026Your buyers already told you what the site is missing
The most valuable data on your website is not in the dashboard. It is in the questions people ask before they buy, and almost nobody collects them.Sachin Aathreyaa K M4 June 2026Your CMS model decides what you will be able to fix later
Most collections are modelled for the page that exists today. Then the product gets renamed and you find out what the model made impossible.Vishal Chiniwar19 June 2026
Questions this raises
Because the classification scheme will change. If raw is discarded, every improvement applies only to signals collected afterwards, and the historical data becomes permanently shaped by decisions made on day one.
It should be suggested automatically and confirmed cheaply. Wrong splits are visible and easy to fix. Wrong merges are invisible and destroy the distinction, so the threshold should be conservative and merging should stay an explicit action.
The classification, not the answer. Those are separate properties and conflating them produces a number nobody can interpret. Confidence should gate routing to a human queue and nothing else.
Fewer than feels thorough. Five coarse intents with defensible boundaries produce reliable classification. Twenty fine ones produce confidence in the low fives because the boundaries between neighbours are not real. Sub categories can be derived later from the groups once you can see the actual clusters.
Take fifty grouped signals, have a person read each group, and count how many contain a question that does not belong. That is your merge error rate and it is the only honest measure. Anything derived from similarity scores is measuring the algorithm against itself.
Mark it addressed at a point in time and keep the history. The count that matters afterwards is the count since it was addressed, and the historical count becomes the thing that tells you whether the fix worked. A group still accruing signals after being addressed is the most valuable row in the system.
A question is unstructured until you give it a shape.
Free text does not aggregate. If you are collecting visitor questions and cannot do anything with them, the schema is the missing piece.
Creobot turns questions into structured records. In private development.