The assistant is not the reason anybody came to the page

A model on a marketing site introduces one new cost and amplifies about six old ones. Most teams instrument the new one and ignore the rest.

Vishal Chiniwar Co-founder and CTO 21 July 2026

Share
A breakdown of where time is spent on an AI featureFour bars showing bundle, hydration, inference and render time, with a budget line marked across them.WHERE THE TIME GOESBUNDLEHYDRATEINFERENCERENDERBUDGET

Two different clocks

When a page with an assistant feels slow, the first thing to establish is which of two unrelated things is happening, because they have nothing in common except the symptom.

The first is model latency: the time between a request and a useful response. It is measured in hundreds of milliseconds to seconds, it happens on somebody else's infrastructure, and the user is expecting to wait because they just asked a question.

The second is main thread cost: the JavaScript your page runs to render an interface, hydrate a component, animate something, or process a stream. It is measured in tens of milliseconds and the user is not expecting to wait at all, because they did not ask for anything.

These get conflated constantly and the remedies are opposite. Model latency is improved by streaming, caching and smaller prompts. Main thread cost is improved by shipping less JavaScript. Applying the first set of fixes to the second problem produces a page that starts responding sooner and still stutters when you scroll.

The cost of a feature nobody opened

The single most common finding on a site with an assistant is that the assistant costs its full weight on every page load, including for the large majority of visitors who never open it.

That happens because the widget is added the way third party widgets have always been added: a script tag, loaded on every page, initialising on load. It parses, it evaluates, it sets up listeners, it may fetch a config, and none of that was requested by the person reading your pricing page.

The correct shape is that an unopened feature costs a button. Render the affordance in HTML, attach one listener, and load everything else on the first interaction. The person who wants it waits a couple of hundred milliseconds once. The person who does not pays nothing.

This is not a sophisticated technique and it is skipped because the default integration path for most tools does not do it. Which means the work is deciding to do it rather than working out how.

A budget that can fail a build

A performance budget that lives in a document is a preference. One that fails a build is a constraint, and the difference is whether anybody is arguing at three in the afternoon on a release day.

The numbers below are the ones we work to on this site. They are not universal and anybody presenting a single set of thresholds as correct for all sites is guessing about your audience and your market. What matters is that a number exists, it is per route rather than site wide, and something enforces it.

Per route matters because averages hide everything. A site where the homepage is light and the one page carrying the assistant is heavy has a good average and a bad experience for the people closest to buying.

Budgets we work to, and what each one is protecting
WhatOur budgetWhat it protectsWhere it is enforced
JavaScript per route, compressedUnder 100 KBParse and evaluate time on mid range phonesBuild step, fails the build
JavaScript before first interactionAs close to zero as the page allowsEverything, this is the one that matters mostBuild step
Interaction handler workUnder 50 ms per handlerInput responsiveness, the thing people feelManual profiling, not automated
Animation propertiesTransform and opacity onlyCompositing rather than layout on the main threadCode review and linting
Third party scriptsCounted, named, and owned by a personUnbounded growth from tags nobody removesA list in the repository
Layout shift on loadNo unreserved space for anything asyncThe page moving under somebody's thumbReserved dimensions in markup

Hydration is where the framework decision shows up

If the site is built with a framework that hydrates, the assistant is a good example of the cost being paid in the wrong place.

Hydration walks the rendered markup and reattaches behaviour, and its cost scales with how much interactive surface exists rather than with how much of it anybody uses. An assistant panel that is closed still has its component tree, its state, its event handlers and its dependencies, all of which get hydrated so that it is ready in case.

There are established ways around this and they are all versions of the same idea: do not hydrate what is not visible, and treat interactivity as something you opt into per component rather than something the page has by default.

This site is static HTML with a small amount of vanilla JavaScript, which sidesteps the problem rather than solving it. That is a legitimate choice for a marketing site and it is not available to everybody, so the general point stands: the framework decision determines whether a closed panel is free or expensive, and that is worth knowing before it is load bearing.

Streaming helps with one clock and not the other

Streaming a response token by token is the standard answer to model latency and it is a good one. First useful content arrives in a few hundred milliseconds instead of after several seconds, and perceived responsiveness improves considerably.

It also creates a main thread problem that did not exist before. Every chunk that arrives triggers a DOM update. Naively implemented, that is dozens of updates per second, each one potentially causing layout and paint, on a thread that is also handling scrolling.

The remedies are ordinary and are usually skipped because streaming feels like it is already the optimisation. Batch incoming chunks into animation frames rather than rendering each one. Append text to a single node rather than rebuilding a tree. Keep the streaming container out of the layout flow of the rest of the page, so an update inside it cannot force layout on everything around it.

The observable version of getting this wrong: the answer appears quickly and the page stutters while it does. Users report that as slow, which is exactly the wrong lesson for the team to take, because it points them back at latency.

Third party scripts and the thing that always happens

Every marketing site accumulates tags. Analytics, a chat widget, a heatmap tool, a consent manager, an ad pixel from a campaign two years ago that nobody has authority to remove.

The specific problem is that this is the one part of the page nobody owns. Each tag was added by a different team for a different reason with a different justification, and the aggregate is not anybody's responsibility. So it grows monotonically.

The mechanism that works is a list in the repository: every third party script, why it is there, and a named person. Anything not on the list is removed. Anything whose named person has left the company is reviewed. That converts an unbounded accumulation into a thing with an owner, which is the same move as everything else in this project.

The second mechanism is loading them properly. Most tags do not need to run before the page is interactive, and most tag documentation suggests otherwise because the vendor optimises for their data completeness rather than your page.

Canvas and animation cost

This site runs a canvas background, so it is worth being honest about what that costs and what keeps it acceptable.

A canvas animation runs continuously and consumes main thread time on every frame unless something stops it. Three constraints keep it from being a problem. It stops entirely when the tab is not visible, because there is no reason to animate for nobody. It respects reduced motion by not moving. And its per frame work is bounded rather than proportional to anything that grows.

The general rule for animation on a marketing site is narrow: transform and opacity only, because those can be handled by the compositor without touching layout. Animating height, width, top or left forces layout on every frame and is the most common cause of an interface that feels heavy.

The sanctioned exception on this site is the accordion, which animates height because there is no honest way to animate a disclosure without it. That is a deliberate exception written down rather than an oversight, and one exception with a reason is a different thing from a codebase where nobody applied the rule.

Reduced motion and keyboard access are performance work

These usually appear under accessibility and belong in a performance conversation too, for a reason that is easy to state.

A feature somebody cannot use is pure cost. It has weight, it consumes main thread time, and it delivers nothing to that person. An assistant that cannot be opened from a keyboard costs its full JavaScript budget for a user who will never get a single answer out of it.

The reduced motion contract is worth stating rather than assuming. Our version keeps colour and opacity transitions and drops movement, rather than disabling animation entirely, because removing all transitions makes an interface feel abrupt rather than calm. That is a decision with a defensible reason, which is better than either extreme applied without one.

The specific thing to check on an AI feature: does the streaming response announce itself to a screen reader without flooding it. A live region set to assertive on a token stream is unusable. Polite, on a container that updates when a chunk boundary is reached rather than per token, is the shape that works.

What to measure, and on what

Field data over lab data wherever you can get it, because your users are not on your laptop.

The metric that correlates best with the complaint people actually make is interaction responsiveness rather than load time. A page that loads in two seconds and responds instantly feels fast. A page that loads in one second and takes 300 ms to respond to a tap feels broken.

Test on a mid range Android device on a throttled connection, not on a recent phone on office wifi. The gap between those two is where most performance regressions live undetected, and it is larger than most teams expect.

And measure the page with the feature closed, which is the state almost every visitor experiences. Measuring an assistant by opening it and timing the response measures the experience of a small minority.

  • Field interaction responsiveness, not lab load time
  • The page in its closed state, which is what most visitors get
  • A mid range device on a throttled connection
  • JavaScript per route, compressed, as a build output
  • Layout shift caused by anything that arrives asynchronously

Where teams get this wrong

Treating performance as a phase. A pass at the end produces a number that regresses within two releases, because nothing prevents the next addition.

Optimising latency when the problem is the main thread. The two clocks again. If the page stutters while the answer streams, faster tokens will not help.

One budget for the whole site. Averages hide the page that matters, which is usually the one with the feature on it.

Loading everything in case. The overwhelming majority of visitors never open the assistant, and shaping the page around the minority who do is the wrong default.

Assuming a framework handles it. Frameworks make choices about when interactivity is paid for, and those choices are defaults rather than guarantees.

What good looks like

A page carrying an AI feature and doing this well is indistinguishable, in its closed state, from the same page without the feature. That is the target and it is achievable.

Concretely: nothing loads for the assistant until somebody touches it. The build fails if a route exceeds its JavaScript budget. Animations use transform and opacity, with the exceptions written down. Third party scripts are a list with owners. The streaming response updates on animation frames rather than on tokens. Reduced motion drops movement and keeps colour. And the feature is fully usable from a keyboard.

None of that is one trick, which is the point. Performance on a page like this is six or seven ordinary disciplines held at once, and the AI part introduces exactly one new problem while making the existing ones more visible.

Where this shows up in what we build

Creobot is a conversation engine in private development, and the constraint we set for it is the one in this article: a site that has it should not be measurably slower for the people who never open it.

That constraint has shaped it more than any feature decision. It is why the affordance is markup rather than a mounted component, why nothing is fetched until the first interaction, and why the streaming path batches into frames rather than rendering per token.

This site is the honest test case. It carries a canvas background, scroll driven animation, an interactive glossary and a filterable blog, and it is static HTML with vanilla JavaScript and a build that enforces the budget. That is not a claim about a product. It is a property of the site you are reading, and the source is there.

Written by

Vishal Chiniwar

Co-founder and CTO

Vishal builds the systems behind Creoglyph. He wrote the four Webflow Designer apps the company ships, and now works on the architecture behind Creogen and Creobot: retrieval, grounding, CMS data models, install flows and the validation that has to sit between a model and a live marketing site.

Related reading

Questions this raises

Usually not, for the majority of visitors. Model latency affects people who asked a question and are expecting to wait. Main thread cost affects everybody, including the large majority who never open the feature, and they were not expecting to wait for anything.

It fixes the wait for the first useful content and can create a main thread problem, because each chunk triggers a DOM update. Batch chunks into animation frames and append to a single node, or the answer arrives quickly on a page that stutters.

There is no correct universal number and anybody who gives you one is guessing about your audience. What matters more than the number is that it is per route rather than site wide, and that something fails when it is exceeded.

It fixes the wait for first useful content and can create a main thread problem, because each chunk triggers a DOM update. Batch chunks into animation frames and append to a single node, or you get an answer that arrives quickly on a page that stutters while it does.

On a marketing site most visitors do not, so measure before assuming. Even where usage is high, rendering the affordance in markup and loading the rest on first interaction costs the user a couple of hundred milliseconds once and costs everyone else nothing.

On a mid range Android device on a throttled connection, measuring interaction responsiveness rather than load time, with the assistant closed. That closed state is what almost every visitor experiences, and opening the feature to time a response measures the experience of a small minority.

An assistant should not cost you the page it sits on.

Model latency and main thread cost are separate problems and get confused constantly. Bring a page with something bolted onto it.

Every performance term in this article is defined plainly in the glossary.

A visitor question becoming a change on a page A question enters on the left, Creobot captures it, Creogen turns it into an operation, and one block on the page is marked as changed. Creobot Creogen QUESTION TO CHANGE