The workflow is fine until somebody else is editing at the same time
A Webflow app runs inside a Designer session it does not own, against a CMS somebody else may be changing, through an API with limits. All three are normal and all three break things.
Vishal Chiniwar Co-founder and CTO 15 July 2026
Three things you do not control
A Designer app runs in an unusual position and it is worth naming the constraints before the workflow design, because every recommendation below follows from them.
You do not own the session. The user can navigate to another page, switch to a different site, close the panel, or lose their connection, at any point, including in the middle of your operation. None of that is misuse.
You do not own the data. Another person may be editing the same collection while your bulk operation runs. There is no lock you can take and no transaction spanning your changes.
You do not own the pace. The API has limits, they are enforced, and hitting them is a normal condition rather than an exception.
An app designed as though it owns any of those three works perfectly in testing, where one person is doing one thing on a small site with a good connection.
The unit of work has to be resumable
The naive shape for a bulk operation is a loop: fetch the items, apply the change to each, report done. It works up to a size and then stops working, and the size is smaller than most people expect.
At ninety items with rate limiting and normal network variance, a run takes long enough that the probability of an interruption stops being negligible. The user switches page, the connection drops, the panel closes. Whatever happened up to that point happened, and the rest did not.
So the operation cannot be one atomic thing, because you have no mechanism to make it atomic. It has to be a sequence of individually completed steps with a recorded position, so that resuming is picking up from item forty seven rather than starting again.
That has a design consequence people resist: the app needs somewhere to keep progress that survives the panel closing. A run record with a list of item ids, each marked pending, done or failed. It feels heavy for what looks like a loop, and it is the difference between an app that works on ninety items and one that works on nine.
The workflow states
These are the states a bulk operation actually occupies, as opposed to the three that appear in a first implementation. The last column is the one that separates a workflow from a spinner.
The states worth dwelling on are partially applied and stale basis, because both are common, both are invisible if you do not design for them, and neither is an error in any meaningful sense.
| State | Cause | What the user is told | Next action |
|---|---|---|---|
| Idle | Nothing running | What the operation will change, and to how many items | Start, after a preview |
| Running | Normal execution | Item n of m, and which item is current | Pause, which must actually pause |
| Paced | Rate limit reached | Waiting, not failing. Expected to resume shortly | Nothing, and no error styling |
| Interrupted | Session ended, panel closed, connection lost | How far it got, precisely | Resume from the recorded position |
| Partially applied | Some items succeeded, some failed | The exact split, with the failed items listed | Retry only the failures |
| Stale basis | An item changed since it was read | Which items changed, and what they say now | Re read and re decide, per item |
| Conflict | Another editor is writing the same field | Somebody else changed this, here is what they set | Keep theirs, keep ours, or open the item |
| Complete | Everything succeeded | What changed, and how to undo it | Undo, or close |
| Failed | The operation could not proceed at all | The specific reason, not an error code alone | Retry, or a support path carrying the run id |
Where the run state actually lives
Saying the operation needs a recorded position raises an immediate question that the resumability argument skips: recorded where.
Panel memory does not survive the panel closing, which is the exact case you are designing for, so it is not an option on its own. Browser storage survives a panel close and a navigation, and does not survive a different browser, a different machine, or a cleared profile. Your own server survives all of that and requires you to have one, and to have a way of associating a run with a user and a site.
The choice is a real tradeoff rather than a best practice. Browser storage is enough for an app whose operations are minutes long and whose users work from one machine, and it costs nothing. Server side run records are the right answer once operations are long enough that somebody might come back tomorrow, or once you want to be able to answer a support question about what happened.
The thing to avoid is the middle position where progress is held in memory and written to storage only at the end, which is the shape that looks like it has persistence and has none where it matters.
Rate limits are a state, not an error
The most common implementation mistake is treating a rate limit response as a failure, showing an error, and stopping. It is neither a failure nor unexpected.
A rate limit is the platform telling you the pace, which is information you asked for by making requests. The correct response is to slow down and continue, and to tell the user that is what is happening in language that does not look like something went wrong.
Practically that means honouring the retry interval the response gives you rather than guessing, backing off with jitter so that many clients do not resynchronise into a wave, and pacing proactively rather than sprinting until you are told to stop. An app that runs at a steady sustainable rate finishes a large operation sooner than one that bursts and gets throttled, which is counterintuitive enough that it is worth measuring once to believe.
The user facing part matters too. Waiting on the platform, resuming shortly is a true statement and reads as normal. Error: 429 is also true and causes people to close the panel, which turns a pause into an interruption.
Partial writes are the normal outcome
Any operation long enough to be interrupted will sometimes be interrupted, and a partial result is what you have. The question is only whether the app knows precisely what happened.
The requirement is exactness. Forty seven of ninety succeeded is a usable statement. Something went wrong is not, because it leaves the user with a collection in an unknown state and no way to reason about it except by checking every item.
Which means every item level result is recorded as it happens rather than aggregated at the end. If the process dies before the end, the record still holds everything up to that point, and the user opening the panel again sees the truth rather than a reset.
The retry then operates on the failed set only. Retrying the whole operation is the behaviour that produces duplicates and second guesses, and it is what happens by default when the only recorded state is started and finished.
Stale basis and concurrent editors
Your app reads ninety items, computes changes, and starts writing. Between the read and the write to item sixty, somebody in another tab edits item sixty.
Writing anyway destroys their edit silently, and silently is the important word. They will not know. They will find it weeks later and will not connect it to your app.
The mitigation is to carry whatever version indicator the platform gives you from the read into the write, and to treat a mismatch as a state rather than an error. Somebody changed this since we looked, here is what it says now, here is what we were going to set. Then the user decides, per item, and the decision is theirs rather than implied by your write order.
This is the failure I would check first in any Designer app, including the ones we shipped, because it never appears in testing. Testing is one person in one tab, and this failure requires two.
Idempotency, so a retry cannot duplicate
Everything above produces retries. Users retry, your error handling retries, and a resumed run replays whatever was in flight when it stopped. So the write has to be safe to repeat.
The mechanism is a key you control rather than one you infer. Generate a run id at the start and an operation key per item, carry both through, and make the write recognise a key it has already applied rather than applying it twice. Inferring from content does not work, because two items can legitimately end up with the same value.
Where the platform does not give you a way to make a write conditional, the fallback is to record the intent before the write and check the record before repeating. That is weaker, because the process can die between the write and the record, and it is considerably better than nothing.
The failure this prevents is specific and unpleasant: a user resumes an interrupted run and the items that had already succeeded get written again. Usually harmless, occasionally not, and always confusing when the change log shows two edits from one operation.
Preview before, undo after
Two things bracket a bulk operation and both are frequently missing.
Preview means showing the exact set of items that will change and what each one will become, before anything happens. Not a count, the list. A count is a promise the user has to trust. A list is something they can scan, and people reliably spot the two items that should not have matched their filter.
Undo means recording the previous value of every field you write, so the operation can be reversed. This is not complicated and it converts the emotional weight of pressing the button entirely. Users approach a destructive bulk edit differently when the panel says this can be undone, and the throughput difference on adoption is larger than any interface improvement.
If undo is genuinely impossible for a class of operation, say so in the preview, at the point of decision. Discovering it afterwards is the worst version.
What to log against what to show
These are different audiences and conflating them produces either a useless interface or an unhelpful log.
The user needs to know what happened to their data. Which items changed, which did not, why not, and what they can do about it. They do not need HTTP status codes, request identifiers, or the shape of the response, and putting those in the interface makes an ordinary pause look like a system failure.
You need to know why. The request, the response, the timing, the retry count, the rate limit headers, the version indicators that were compared. That belongs in a log with the run id attached, and none of it belongs on screen.
The connection between the two is the run id, which appears in both. That single value is what lets a user say the operation did not work and gives you everything without them having to describe anything.
The error message principle we work to is that the screen says what happened to their data and offers a next action, and never asks the user to interpret something that is our problem to interpret.
The support path is part of the workflow
This is the same argument as for install flows and it applies with more force here, because a bulk operation that half completed leaves the user in a state they cannot describe.
Every terminal failure should carry a run id, the operation type, the counts, and the specific error, in a form the user can send without explaining anything. A support conversation that starts with a run id is a different conversation from one that starts with your app broke my collection.
It also converts your support volume into data. Failures that people actually report are a different set from failures in your logs, because logs record everything and reports record what mattered enough to write in about. The gap between those two sets is worth looking at.
The related surface is the success and error pages that sit outside the Designer panel. Those should carry the same run id, because a user who was redirected out of the panel has lost all the context the panel held.
Where teams get this wrong
Designing for a small site. Everything works at ten items. The problems begin somewhere around fifty and the design decisions that survive them have to be made before you have a user with ninety.
Treating rate limits as failures. It converts a normal pause into an interruption, and it teaches users your app is unreliable when the platform was behaving correctly.
Aggregating results at the end. If nothing is recorded per item as it happens, an interruption erases everything you knew about what already succeeded.
No undo, and no statement about it. The absence is felt as risk, and users respond by not using bulk features on anything that matters, which is exactly the case bulk features exist for.
Testing with one tab. Concurrent editing is the most common real world failure and the least likely to be caught, because catching it requires deliberately doing something a solo tester never does.
How to check an app you already shipped
This is the pass I would run on ours and on anybody else's. Most of it can be done in an afternoon against a test site.
The concurrency item is the one most likely to fail, and it is worth doing first because the fix is a version check rather than a redesign.
An afternoon of deliberate abuse
- Run a bulk operation on ninety items and close the panel halfway. Reopen it. Does it know where it got to?
- Run one and let the connection drop. Does it resume, restart, or forget?
- Trigger a rate limit deliberately. Does the user see a pause or an error?
- Edit an item in a second tab while the operation is running. Does the app notice or overwrite?
- Force a failure on item forty. Does the retry cover only the failures?
- Check whether the preview lists items or only counts them
- Check whether the operation can be undone, and whether the user is told before they commit
- Check that every failure screen carries a run id somebody could send you
What good looks like
A reliable Designer app is one that is boring under conditions that are not boring.
It tells you what it is about to change, item by item. It goes at a steady pace and says when it is waiting on the platform. If you close the panel it picks up where it stopped. If somebody else edited something it asks rather than deciding. When it finishes it tells you exactly what changed and offers to undo it. When it fails it tells you which items and gives you something to send us.
None of that is technically hard. All of it is a consequence of having decided that interruption, concurrency and pacing are the normal case rather than the exception.
Where this came from
We have built four Webflow Designer apps and three are public. Most of the states in that table are not from a design exercise. They are cases that happened to real users and arrived as support conversations, which is why the list includes things like stale basis that would not occur to anybody drawing a flow.
The resumable unit of work is the change I would make first if I were starting again. Our early version of a bulk operation was a loop with a progress bar, which is the obvious implementation and is fine until somebody closes the panel. The rewrite to a recorded run with per item state was not large and it should have been the first design rather than the second.
I am not putting usage numbers on any of this. The states are the transferable part and they hold whether an app has ten users or ten thousand. What the volume gave us was the frequency of the unusual cases, and the lesson from that is only that the unusual cases are more common than a solo tester will ever believe.
Related reading
An install flow is a trust flow that happens to use OAuth
The person installing your app is deciding whether to give you access to something they care about, on a screen you did not design, in about ninety seconds.Vishal Chiniwar13 July 2026We built four apps before we noticed they were all the same app
Every one of them fixed something at the scale of a whole site rather than a page. It took us longer than it should have to see what that meant.Sachin Aathreyaa K M24 July 2026A prompt is not a guardrail
Instructions in a prompt are a preference. A guardrail is something that can refuse, that leaves a record when it refuses, and that you can write a test against.Vishal Chiniwar19 May 2026
Questions this raises
You cannot make it atomic, because there is no transaction spanning the writes and no lock available. The achievable property is resumability: individually completed steps with a recorded position, so an interruption costs you the current item rather than the run.
As a state rather than an error. Honour the interval the response gives you, back off with jitter, and pace proactively rather than sprinting until throttled. Tell the user you are waiting on the platform, because an error message at that moment makes them close the panel.
Concurrent editing. Your app reads an item, somebody edits it in another tab, and your write destroys their change silently. It never shows up in testing because testing is one person in one tab, and the fix is carrying a version indicator from the read into the write.
It depends on how long your operations are. Browser storage survives a panel close and a navigation and is enough for operations measured in minutes. Server side records survive a different machine and let you answer a support question about what happened. The shape to avoid is progress held in memory and written only at the end.
Record the intent before the write and check the record before repeating. It is weaker than a real conditional write, because the process can die between the write and the record, and it is considerably better than nothing. Carry a run id and a per item operation key regardless.
No. It is the platform telling you the pace, which is information you asked for by making requests. Show a pause with no error styling and resume. An error at that moment makes people close the panel, which converts a pause into an interruption.
Webflow app work fails in ways the docs do not cover.
Rate limits, partial writes, stale Designer state. If you are building on the platform we will compare scar tissue.
Four Webflow Designer apps built. Most of this came from breaking them first.