Skip to content

Errors, retries & limits

The conventions every call shares, whichever surface made it: what a response looks like, how a failure is described so a program can branch on it, how to retry without doing the work twice, and where the ceilings are. The SDKs raise these as ImageStepError, the CLI maps them to exit codes, MCP returns them as tool errors — same fields everywhere.

The envelope

Base URL https://api.imagestep.dev. Every /api/v1 response is JSON, and success and failure are disjoint: data is present exactly on success, error exactly on failure. A list adds meta.

A success — here the answer to a submit, with data cut down to a few of the job's fields:

{
  "success": true,
  "message": "Job created",
  "data": {
    "id": "job_61ed7d70cc1e4e748ca2e49abb4fa2a4",
    "type": "process",
    "status": "PENDING",
    "op": "grayscale",
    "totalItems": 1
  },
  "timestamp": "2026-09-17T16:46:30.548Z"
}

A list: data is the page, meta says where it sits.

{
  "success": true,
  "message": "Success",
  "data": [
    { "id": "job_f62c2a6f81534cd8b296ea2b7f8d2521", "type": "process", "status": "COMPLETED", "totalItems": 1 },
    { "id": "job_d59ca98c1d6d4b5fbc8eef2477dc5681", "type": "ai-edit", "status": "FAILED", "totalItems": 3 }
  ],
  "meta": {
    "total": 412,
    "page": 0,
    "perPage": 100,
    "hasMore": true,
    "nextCursor": "am9iOjE3ODk3NDUxMTMxMTgyMTQ6am9iX2Q1OWNhOThjMWQ2ZDRiNWZiYzhlZWYyNDc3ZGM1Njgx"
  },
  "timestamp": "2026-09-17T15:41:53.118Z"
}

A failure: no data, and everything a program needs is in error.

{
  "success": false,
  "error": {
    "code": "invalid_param",
    "message": "format must be one of [jpeg, tiff, gif, webp, avif, png]",
    "retryable": false,
    "param": "format",
    "requestId": "fd62714a-193e-4628-88ed-295cc70a9209"
  },
  "timestamp": "2026-09-17T16:46:57.073Z"
}

Authentication

Programs authenticate with an API key, created on /keys and sent as Authorization: ApiKey is_sk_…. The SDKs, the CLI and the MCP server read it from IMAGESTEP_API_KEY. A missing, mistyped or revoked key is 401 unauthorized; a key on an endpoint only the console may use is 403 forbidden. A key cannot manage keys — it can neither list, mint nor revoke another one — with exactly one exception, pointing the other way: DELETE /api/v1/api-keys/self revokes the key that made the call (that is what imagestep logout does). Three things need no key at all: the op catalogue, the public price list and the agent guidelines.

The error object

fieldtypemeaning
codestringone of the closed set below. Branch on it, never on message, which is for people and changes
retryablebooleanwhether the same call can be made again. The status code cannot tell you: 429 and 503 are retryable, 400 and 402 are not, and request_in_progress is a 409 that is
paramstringthe request parameter the error is about, when it is about one — the field to fix; body when it is the body itself, a bare list that is empty or too long
detailsobjectoptional machine-readable context: received, missing, fields, limit, reason
messagestringone sentence for a person
requestIdstringthis request's id, the same value as the X-Request-Id header; quote it when reporting

The rule an agent follows: retryable: true — wait, then make the same call, with the same Idempotency-Key. retryable: false — do not resend the same body: fix what param names, or, when there is nothing in the request to fix (insufficient_credit, asset_count_exceeded), the account is the problem and a person has to act.

How long to wait: what Retry-After says, when the answer carries it — every rate_limited does, and so do the synchronous lane's and the job waits' provider_unavailable — and otherwise a backoff that doubles, the SDKs' 0.5 s then 1 s, given up after a few tries. One retryable answer is not a queue: an internal_error whose details.reason is query_timeout times out again unless the request is narrowed. The SDKs and the CLI have already retried a retryable answer twice by the time it reaches your code, so there it is one to come back to later:

import { ImageStepError, JobFailedError } from "imagestep";

try {
  await client.ops.upscale(ids, { wait: true });
} catch (err) {
  // retryable: the SDK already retried, so come back later — after what the service asked for
  if (err instanceof ImageStepError && err.retryable) queueForLater(ids, err.retryAfter);
  else if (err instanceof JobFailedError) console.error(err.job.status, err.job.id);
  else throw err; // not retryable: fix what err.param names
}

The same fields on every surface: an ImageStepError carries them as properties, the CLI turns retryable into its exit code (4 retryable, 3 not), and an MCP tool error is the error object itself.

Error codes

The closed set, by status. Adding a code is a minor change; renaming or removing one is breaking. A sub-reason goes in details, never in a new code.

codeHTTPretryablewhen it happens
invalid_param400noa parameter is missing, malformed or out of range; param names it (details.fields, when several are). Fix it: the same body fails again
invalid_state400nowell-formed, but what it acts on is not in a state to honour it — cancelling a job that has finished, resuming one that has not failed
unsupported_format400nothe bytes are not an image format this service reads
unauthorized401nono credential, or one this service does not accept — a revoked or mistyped key. Fix the key; retrying it gets an address blocked (below)
insufficient_credit402nothe balance does not cover the job; top up, then send it again. The dry run's sufficientCredit says so first
forbidden403noauthenticated, but not allowed — a built-in you tried to edit, a key on an endpoint only the console may use
asset_not_found · job_not_found · preset_not_found · not_found404nono such thing, or it is not yours — the two are never told apart
method_not_allowed405nothe path exists, not for this method; the Allow header and details.allowed list the ones it takes
not_acceptable406nonothing this endpoint produces satisfies your Accept; JSON is the only representation
idempotency_key_reuse409nothis Idempotency-Key already means a different request — or the first one's answer could not be kept (idempotency keys)
request_in_progress409yesthe first request under this key has not answered yet; send it again shortly, with the same key
subscription_not_actionable409nothe plan action names no subscription this service can act on
payload_too_large413nothe body is over the ceiling for its kind (below); details.limit is the ceiling in bytes. Split it into two requests
unsupported_media_type415nothe body's Content-Type is not one this endpoint reads; details.supported lists the ones it does
asset_count_exceeded422nothe outputs would take the account past its plan's stored-asset ceiling; delete assets or upgrade. The dry run's assetCountLeft says so first
resource_limit_exceeded422noa per-account hard limit — webhook endpoints, for one — would be exceeded
job_not_resumable422noa resume of an attempt that is not the latest, or past the resume budget; details.reason says which
provider_rejected422noan upstream provider refused this input; the same input fails again (on job items)
rate_limited429yestoo many requests, or too many at once (below); wait the Retry-After seconds it carries
service_misconfigured500noour side, and retrying will not help; tell us
internal_error500yesunclassified, on our side; may be transient. With details.reason "query_timeout" the same request can time out again — narrow it before a second retry
provider_unavailable503yesan upstream AI provider, the image worker or the host of a url you sent failed, timed out or is full; details.reason says which, and Retry-After, when present, how long to wait

Failures inside a job

A job that was accepted and then failed carries no error envelope — the request was fine. The verdict is on the items: each failed one has an errorCode from the table above and that code's retryable — or neither, when a worker could not name the cause, which is to be read as not retryable — and the job repeats the verdict at the top. Which codes appear on items, and how resume uses them, is on jobs.

Idempotency keys

Every write (POST / PUT / PATCH / DELETE) accepts an optional Idempotency-Key header — one per logical operation, a UUID is the usual choice. It is what makes a retry safe:

you sendyou get
no keynormal execution, no deduplication
a new keynormal execution; the status and body are stored under (account, key)
same key, same bodythe stored response, verbatim, plus Idempotency-Replayed: true — nothing runs again
same key, different body409 idempotency_key_reuse
same key while the first is still running409 request_in_progress, retryable
  • “The same body” is the same bytes: the method, the path, the query string and the raw body, hashed. A retry that serialises its JSON again — keys in another order, other whitespace — is a different body and a 409; resend the bytes you sent.
  • Keys belong to the account, not to one API key, and are kept 24 hours; after that the same key is new again.
  • A 5xx releases the key, so a retry after one runs. A first request that died without answering holds its key for 5 minutes — request_in_progress until then — and a retry after that runs.
  • Two limits, both of which mean the first request ran: a body over 1 MB is not fingerprinted, and a response over 512 KB is not stored. A repeat of either is refused with 409 idempotency_key_reuse rather than replayed — read what the first one made instead of sending it again.
  • Ignored rather than honoured on the synchronous image endpoints, which create nothing a replay could protect; on the two batch reads that are a POST only because a page of ids does not fit a query string; and on POST /api/v1/assets/upload, whose own rule is that the same bytes are the same asset (existing: true) — send it again. A job's dry run ignores it too: it creates nothing, so pricing a body and then submitting it under one key works.
  • A job submit that asked to wait replays with the job as it is now — waiting again while it runs — rather than the snapshot it stored (jobs).

The SDKs and the CLI send a key on every write and reuse it across their own retries. By hand, and from an agent:

KEY=$(uuidgen)   # one per logical operation — keep it for every retry of this submit
curl -si https://api.imagestep.dev/api/v1/jobs -H "Authorization: ApiKey $IMAGESTEP_API_KEY" -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" -d '{"op":"grayscale","assetIds":["<asset-id>"]}'

Sent twice, the second answer is the first one, byte for byte, and says so. Its X-Request-Id is this request's; a stored failure's error.requestId is the original's.

HTTP/2 201
content-type: application/json
idempotency-replayed: true
x-request-id: ca3c0512-775e-42f4-83e8-81dce29bb702

Rate limits and quotas

Every response that got past authentication carries the current request budget:

RateLimit-Limit: 600
RateLimit-Remaining: 587
RateLimit-Reset: 43

RateLimit-Limit is the requests the window allows, RateLimit-Remaining what is left of it, and RateLimit-Reset the seconds until the window turns. The budget is counted per account (per client IP when anonymous) on each server, so it is a ceiling to pace by, not a quota to spend down — the plan's limits are other codes, so a program can tell “slow down” from “you have run out”:

ceilingcodewhat to do
600 requests per 60 s, per account (per IP when anonymous)429 rate_limited + Retry-Afterslow down; pace by RateLimit-Remaining, and retry after the header says
60 failed authentications from one IP within 60 s429 rate_limited + Retry-Afterfix the key: every request from that address is refused until the window turns, a good key included
4 synchronous image calls in flight per account429 rate_limited + Retry-Afterthe same: it is the sync lane's own concurrency
8 job waits open per account429 rate_limited + Retry-Afterread the job without wait, or wait again shortly; a waited submit is never refused for it
a request body of 8 MB of JSON, 34 MB of image or multipart413 payload_too_largesplit it into two requests; details.limit is the ceiling
credit — for AI ops, and for deterministic runs past the Free plan's monthly allowance, which are paid rather than refused402 insufficient_credittop up at details.topUpUrl; details says what was needed and what could be spent, and the dry run's sufficientCredit warns first
stored assets422 asset_count_exceededdelete, or upgrade at details.upgradeUrl; details.left says how many fit, and the dry run's assetCountLeft warns first
2 running jobs per type per accountnone — the next one waits in PENDINGnothing; it starts when one finishes

Pagination

Every list takes page (from 0) and perPage (100 by default and at most); meta.hasMore says whether to ask again, and meta.nextCursor where the next page starts. One shape, every list — assets, collections, jobs, a job’s items, an endpoint’s deliveries, your reports and the templates. A page or perPage out of range is clamped rather than refused, so perPage=500 is not an error, it is 100; meta always reports what was actually served:

{
  "total": 412,
  "page": 0,
  "perPage": 100,
  "hasMore": true,
  "nextCursor": "YXN0OjE3OTAwMDAwMDAwMDA6YXN0XzRjMWYwZThhOWIyZDRlNmY4YTFjM2I1ZDdlOWYwYTJi"
}

To read on, send nextCursor back as cursor instead of a page number. The service then reads the rows after it and counts nothing, so the answer carries no total and no page, and nextCursor is null on the last page:

{
  "perPage": 100,
  "hasMore": false,
  "nextCursor": null
}

That is the way to walk a long list. Page k by number makes the service read and throw away every row before it and count the whole filter again, so walking a hundred thousand assets by page number is a thousand counts; by cursor every page costs what the first did. It also holds still: rows that arrive while you walk land on top, and a cursor carries on below them. Send the cursor unchanged and with the same filters — it is opaque, and it names a position in one listing’s order, not a query. With page beside it, or from another listing, it is 400 invalid_param on cursor.

The one loop worth getting right is the one that walks to the end — and the SDKs, the CLI and the n8n node have it, so you do not write it. The next page is wherever the answer says it starts:

for await (const asset of client.assets.iterate({ collection: "shoot-01" })) {
  console.log(asset.id);
}

Ask for less when a page has somewhere small to go — the console’s grid asks for 48, the MCP search_assets tool for 20, because a page lands in a model’s context and a context is a viewport. Anything walking the list itself should take the 100: it is the fewest round trips.

A document that would otherwise carry a list inside it does the same: a job comes back with its first 100 items and itemsTruncated, and the rest are GET /api/v1/jobs/{id}/items.

Some listings are not paged at all, because they are catalogues or small bounded sets a caller wants whole: GET /api/v1/ops, /api/v1/ai-models, /api/v1/presets, /api/v1/webhook-endpoints, /api/v1/jobs/counts and a template’s /versions answer with the whole list and no meta.

Request ids

Every response carries X-Request-Id, and every error repeats it as error.requestId. Send your own — up to 128 characters of letters, digits, ._-: — and it is used as is, so one string joins your trace to ours; anything else is replaced, not trimmed, and without one you get a UUID. It is set before authentication, so a 401 and a 429 carry it too. Quote it when you write to support, or when an agent files a report with POST /api/v1/feedback.