Design the API for whoever is debugging it at midnight

Clever is a tax somebody else pays, usually eighteen months later, usually somebody who was not in the room when the cleverness was decided.

Vishal Chiniwar Co-founder and CTO 3 August 2026

Share
An API judged at the moment it failsA stopped integration meets an actionable error, above three principles that make an API debuggable.MIDNIGHTIntegration stoppedError you can act onCONSISTENT NAMING AND SHAPESVERSION IN THE PATH FROM DAY ONEIDEMPOTENCY KEY ON EVERY WRITENOBODY EVALUATES AN API WHEN IT WORKS

Who the reader is

Every API has a real primary user and it is rarely the one in the design discussion. It is not the person integrating on a good day with documentation open. It is somebody eighteen months later, at an unreasonable hour, with a failing job and no context, trying to work out why a write that succeeded last week is now returning something they have not seen before.

Design for that person and most decisions resolve themselves. Names are boring and unambiguous. Errors say what happened and what to do. Behaviour is predictable rather than convenient. Nothing is inferred that could be stated.

Design for the demo instead and you get an API that reads beautifully in a tutorial and cannot be debugged, which is the failure mode that costs a company years.

Identifiers, and why slugs are not one

The first decision, and the one everything inherits.

A slug is an address. It is chosen for humans, it appears in URLs, and it changes when somebody decides a better one exists. An identifier is what a thing is, and it must never change for as long as the thing exists.

Using the slug as the identifier is the most common mistake in content APIs and it is invisible for about four months. Then somebody renames a product, the slug changes, and every consumer that stored a reference is now pointing at nothing. From the API's side it looks like a deletion and a creation. From the consumer's side their data is corrupt.

So: an opaque, immutable identifier as the primary key, the slug as a mutable attribute, and a way to look up by slug that returns the identifier. Ideally also a slug history, so an old slug resolves to the current entity rather than a 404, because somebody will have bookmarked it and something will have cached it.

The one further constraint is that the identifier should be opaque. If it encodes a type or a date or a sequence, somebody will parse it, and you will have accidentally made your internal structure part of your contract.

Errors are a product surface

Error messages get written last, by whoever is closest to the code, at the moment they are least interested. They are then read at the worst possible moment by somebody with the least context. That asymmetry is worth designing against deliberately.

A useful error has four parts: a stable machine readable code, a human sentence saying what happened, the specific thing that caused it, and what to do next. Validation failed has none of them. Field publishedAt must be an ISO 8601 date, received 12 March 2024 has all four.

The code has to be stable, because consumers will branch on it and a message string will be edited. If the code changes, integrations break silently in the class of way nobody notices until a customer does.

And errors should name the field. An API that says a request is invalid without saying which of thirty fields is invalid has converted a five second fix into an afternoon of bisection.

The test I use: could somebody who has never seen this system resolve the problem from the error alone. If not, the error is a notification rather than a message.

  • A stable code that consumers can branch on and you will not edit
  • A sentence saying what happened in ordinary language
  • The specific field or value that caused it, quoted back
  • What to do next, when there is a next thing to do
  • A request identifier, so a support conversation starts with something

Validation belongs at the boundary

A content API is a boundary between systems that trust each other differently, and the boundary is where correctness has to be enforced, because it is the last point at which you can refuse.

Validate on write, completely, and reject rather than coerce. Coercion is well intentioned and it is how a field ends up containing four representations of the same value. An empty string becoming null, a string becoming a number, a date being parsed leniently. Each is a small kindness that costs somebody a debugging session later.

Validate the whole request rather than failing on the first problem. Returning one error at a time turns a ten field form into ten round trips, and the person on the other end learns to hate you at a rate of one error per attempt.

And keep validation rules in one place rather than duplicating them in the client. Duplicated rules diverge, and when they diverge the client's version is the one that is wrong, because the server's is the one that is enforced.

Schema versions, decided early or migrated later

Content schemas change. Fields are added, meanings shift, something that was a string becomes a reference. The question is not whether but how consumers find out.

The decision has to be made before the first external consumer, because after that any versioning approach is a migration with people's integrations attached.

Additive changes should be safe by contract: consumers must ignore unknown fields, and you must never repurpose an existing field name for a different meaning. Repurposing is the change that breaks consumers most quietly, because nothing errors, the data is simply now wrong.

For breaking changes, pick one approach and be consistent. A version in the path is unfashionable and it is unambiguous, which for a content API used by people integrating occasionally is worth more than elegance. Whatever you choose, publish what constitutes a breaking change, because the useful part of a versioning policy is the definition rather than the mechanism.

Idempotency, because retries are guaranteed

Networks time out. Clients retry. Your own job runner retries. None of that is exceptional, so any write that can be issued twice has to be safe to issue twice.

The mechanism is a key supplied by the client on the request, stored by you against the result. A repeat with the same key returns the original result rather than performing the operation again. This is well established and it is skipped constantly in internal APIs on the reasoning that our client will not retry.

The subtlety worth stating is what to do when the same key arrives with a different body. That is a client bug, and returning the original result silently hides it. Returning a conflict with a clear message surfaces it at the point where somebody can fix it.

For content specifically, the write that most needs this is publication. Publishing twice usually looks harmless, and it produces two entries in a change history, two webhook deliveries, and two notifications to whoever is watching.

Permissions per field, not just per endpoint

Endpoint level permissions are the default and are too coarse for content, because a content record is not uniformly sensitive.

A marketing tool may need to update a page's body and must not be able to change its price. An automated process may write a draft and must not publish. A translator may edit copy and must not touch a reference to another entity. Endpoint level access cannot express any of that, so teams either grant too much or build parallel endpoints.

Field level scopes handle it directly. The API knows which fields a credential may write, and a request touching anything outside that set is rejected with a message naming the field.

This becomes necessary rather than nice the moment anything automated writes to a content system, because a token with broad write access held by an automated process is the arrangement that produces the worst incidents in this category.

Endpoints, and what each one has to say

The shape below is what a content API needs at minimum, with the property of each that is most often missing. None of it is unusual and the right hand column is where the debugging pain comes from.

The row worth noting is the last one. A change history endpoint is the one most often left out and the one most valuable during an incident, because the first question is always what changed.

Content API endpoints and the property most often missing from each
EndpointPurposeMost often missingCost of missing it
List collectionPaginated set of entitiesA stable cursor, not an offsetItems skipped or repeated while paging
Get by idOne entity by immutable idAn observed at timestamp on the responseConsumers cannot tell how stale it is
Resolve by slugSlug to identifier, including old slugsSlug historyRenames become 404s for everything cached
CreateNew entityIdempotency key handlingDuplicates on retry
UpdateChange fieldsField level permission and a version preconditionConcurrent edits silently overwritten
PublishMove to liveSeparation from update, and idempotencyAccidental publication, doubled webhooks
ValidateCheck without writingExisting at allConsumers discover rules by failing
Change historyWhat changed, when, by whomExisting at allNo way to answer what happened during an incident

Pagination that survives a moving collection

Content collections change while somebody is paging through them, which is what makes offset pagination wrong rather than merely unfashionable.

Page two of an offset paginated list, after an item was inserted at the top, contains one item from page one. After a deletion, one item is skipped entirely. A consumer syncing a collection this way ends up with a set that is quietly incomplete, and nothing errors.

Cursor pagination fixes it because the cursor encodes a position in a stable ordering rather than a count. The requirement is that the ordering is stable, which means ordering by something immutable and unique rather than by a timestamp that can collide.

Also return the page size you actually applied rather than assuming the requested one was honoured. A consumer that asked for 500 and received 100 without being told will silently under fetch.

Stale data, and saying when something was true

This is the property most content APIs omit and it is the one that matters most for anything operational reading them.

Every response should carry the time the data was observed or last modified. Without it, a consumer holding a record has no way to reason about whether to trust it, and will either re fetch constantly or act on something arbitrarily old.

For anything cached or derived, say so explicitly, and say how stale it may be. A cached response that does not identify itself as cached is indistinguishable from a current one, which means a consumer cannot make a decision about whether the staleness matters for what they are doing.

The related mechanism is a version precondition on writes: the client sends the version it read, and the write fails if the entity has changed since. That converts a silent overwrite of somebody else's edit into a conflict the client can handle, which is the difference between losing data and having a decision to make.

Logs, and the identifier that connects everything

A request identifier returned on every response, including successful ones, is the cheapest debugging affordance available and it is frequently absent.

It has to appear in three places to do its work: the response, your logs, and any error surfaced to a human. Then a support conversation begins with a value rather than a description, and the timeline of one operation can be reconstructed without correlating timestamps.

For content APIs specifically, the change history is the log that matters most. Who changed what, when, through which credential, and what the previous value was. During any incident involving a website, the first question is what changed, and a content system that cannot answer it turns a ten minute investigation into a day.

That history should be queryable by entity, because the question is almost always about a specific page rather than about a time window.

A design checklist

The set worth checking before anybody external integrates. Most of these are cheap at the start and expensive after the first consumer exists.

The first two are the ones that cannot be retrofitted without a migration.

Before the first external consumer

  • Identifiers are opaque and immutable, and slugs are a separate mutable attribute
  • Old slugs resolve rather than 404
  • Versioning approach chosen, and what counts as a breaking change is written down
  • Every error has a stable code, a sentence, the offending field, and a next action
  • Validation rejects rather than coerces, and reports all failures at once
  • Writes accept an idempotency key, and a repeat with a different body is a conflict
  • Permissions are per field, not only per endpoint
  • Pagination is cursor based over a stable, unique ordering
  • Every response carries the time the data was true
  • Every response carries a request id that also appears in your logs
  • A change history is queryable by entity

Where teams get this wrong

Optimising for the first integration. The first consumer is written by somebody who can ask you questions. Every subsequent one cannot.

Being clever with names. An endpoint that does something slightly different from what its name suggests will be misused forever, and the documentation correcting it will not be read.

Coercing input to be helpful. Every coercion is a decision made on somebody's behalf that they will later have to reverse engineer.

Treating errors as an implementation detail. They are the part of an API most read under pressure, and the part written with the least attention.

Deferring versioning. It is a five minute decision before the first consumer and a project afterwards.

Assuming internal means safe. Internal APIs acquire external consumers, always, usually without a conversation.

What good looks like

A well designed content API is boring to use and boring to debug, and both of those are the goal.

Concretely: you can rename anything without breaking a consumer. An error tells you which field and what to do. A retry cannot duplicate. Two people editing the same entity produces a conflict rather than a loss. You can tell how old any response is. And when something goes wrong at an unreasonable hour, the change history answers what happened in about a minute.

The internal signal is what your support requests look like. If they arrive with a request id and a specific error code, the API is telling people what happened. If they arrive as narrative descriptions of what somebody was trying to do, it is not.

Where this shows up in what we build

Creogen reads and writes content systems on behalf of people who are not watching at the time, which is the condition that makes every property above load bearing rather than tidy. It is in private development and I would rather say that than describe a finished thing.

Two of these came directly from getting them wrong first while building the Webflow Designer apps. The identifier decision is one, and it is why identity sits at the top of our build order rather than in the middle. Idempotency on publication is the other, and we found it the way most people do, by generating two of something and having to explain it.

Nothing here needs our tooling and none of it is novel. It is a list of ordinary decisions, most of which are cheap before the first consumer and expensive afterwards, which is the only reason it is worth writing down in advance rather than learning in order.

Written by

Vishal Chiniwar

Co-founder and CTO

Vishal builds the systems behind Creoglyph. He wrote the four Webflow Designer apps the company ships, and now works on the architecture behind Creogen and Creobot: retrieval, grounding, CMS data models, install flows and the validation that has to sit between a model and a live marketing site.

Related reading

Questions this raises

Because slugs change. When one does, every consumer holding a reference is pointing at nothing, and from the API side it looks like a deletion plus a creation. Use an opaque immutable id, keep the slug as a mutable attribute, and resolve old slugs rather than returning 404.

It is until something automated writes to your content system, at which point endpoint level access means granting a process the ability to change a price when it only needed to draft a paragraph. That is the arrangement behind most bad incidents in this category.

A change history queryable by entity. During any incident the first question is what changed, and an API that cannot answer it turns a ten minute investigation into a day.

Survivable and expensive to fix later. The path is additive: introduce an opaque id, backfill it, keep the slug as a mutable attribute, add slug history so old addresses resolve, then migrate consumers. Doing it before you have external consumers is a morning. Afterwards it is a coordinated migration.

Because it is a decision made on somebody's behalf that they will later have to reverse engineer. An empty string becoming null, a lenient date parse, a string quietly becoming a number: each is a small kindness that produces a field containing four representations of the same value.

Decide before the first external consumer, which makes it a five minute choice rather than a migration. Make additive changes safe by contract, so consumers ignore unknown fields and you never repurpose a field name for a new meaning. Publish what counts as a breaking change, because the definition is the useful part.

For a collection that changes while somebody is paging through it, yes. An insertion repeats an item across pages and a deletion skips one entirely, and nothing errors, so a consumer syncing your collection ends up quietly incomplete. Cursor pagination over a stable unique ordering removes the whole class.

Design the API for whoever is debugging it at midnight.

Clever is a tax someone else pays. Bring an integration that keeps surprising you and we will look at where the surprise comes from.

Three public Webflow apps and the internal services behind them. All of it debugged at unreasonable hours.

A visitor question becoming a change on a page A question enters on the left, Creobot captures it, Creogen turns it into an operation, and one block on the page is marked as changed. Creobot Creogen QUESTION TO CHANGE