Errors, retries & limits
The conventions every call shares, whichever surface made it: what a response looks like, how a failure is described so a program can branch on it, how to retry without doing the work twice, and where the ceilings are. The SDKs raise these as ImageStepError, the CLI maps them to exit codes, MCP returns them as tool errors — same fields everywhere.
The envelope
Base URL https://api.imagestep.dev. Every /api/v1 response is JSON, and success and failure are disjoint: data is present exactly on success, error exactly on failure. A list adds meta.
A success — here the answer to a submit, with data cut down to a few of the job's fields:
{
"success": true,
"message": "Job created",
"data": {
"id": "job_61ed7d70cc1e4e748ca2e49abb4fa2a4",
"type": "process",
"status": "PENDING",
"op": "grayscale",
"totalItems": 1
},
"timestamp": "2026-09-17T16:46:30.548Z"
}A list: data is the page, meta says where it sits.
{
"success": true,
"message": "Success",
"data": [
{ "id": "job_f62c2a6f81534cd8b296ea2b7f8d2521", "type": "process", "status": "COMPLETED", "totalItems": 1 },
{ "id": "job_d59ca98c1d6d4b5fbc8eef2477dc5681", "type": "ai-edit", "status": "FAILED", "totalItems": 3 }
],
"meta": {
"total": 412,
"page": 0,
"perPage": 100,
"hasMore": true,
"nextCursor": "am9iOjE3ODk3NDUxMTMxMTgyMTQ6am9iX2Q1OWNhOThjMWQ2ZDRiNWZiYzhlZWYyNDc3ZGM1Njgx"
},
"timestamp": "2026-09-17T15:41:53.118Z"
}A failure: no data, and everything a program needs is in error.
{
"success": false,
"error": {
"code": "invalid_param",
"message": "format must be one of [jpeg, tiff, gif, webp, avif, png]",
"retryable": false,
"param": "format",
"requestId": "fd62714a-193e-4628-88ed-295cc70a9209"
},
"timestamp": "2026-09-17T16:46:57.073Z"
}Authentication
Programs authenticate with an API key, created on /keys and sent as Authorization: ApiKey is_sk_…. The SDKs, the CLI and the MCP server read it from IMAGESTEP_API_KEY. A missing, mistyped or revoked key is 401 unauthorized; a key on an endpoint only the console may use is 403 forbidden. A key cannot manage keys — it can neither list, mint nor revoke another one — with exactly one exception, pointing the other way: DELETE /api/v1/api-keys/self revokes the key that made the call (that is what imagestep logout does). Three things need no key at all: the op catalogue, the public price list and the agent guidelines.
The error object
| field | type | meaning |
|---|---|---|
| code | string | one of the closed set below. Branch on it, never on message, which is for people and changes |
| retryable | boolean | whether the same call can be made again. The status code cannot tell you: 429 and 503 are retryable, 400 and 402 are not, and request_in_progress is a 409 that is |
| param | string | the request parameter the error is about, when it is about one — the field to fix; body when it is the body itself, a bare list that is empty or too long |
| details | object | optional machine-readable context: received, missing, fields, limit, reason |
| message | string | one sentence for a person |
| requestId | string | this request's id, the same value as the X-Request-Id header; quote it when reporting |
The rule an agent follows: retryable: true — wait, then make the same call, with the same Idempotency-Key. retryable: false — do not resend the same body: fix what param names, or, when there is nothing in the request to fix (insufficient_credit, asset_count_exceeded), the account is the problem and a person has to act.
How long to wait: what Retry-After says, when the answer carries it — every rate_limited does, and so do the synchronous lane's and the job waits' provider_unavailable — and otherwise a backoff that doubles, the SDKs' 0.5 s then 1 s, given up after a few tries. One retryable answer is not a queue: an internal_error whose details.reason is query_timeout times out again unless the request is narrowed. The SDKs and the CLI have already retried a retryable answer twice by the time it reaches your code, so there it is one to come back to later:
JavaScript
import { ImageStepError, JobFailedError } from "imagestep";
try {
await client.ops.upscale(ids, { wait: true });
} catch (err) {
// retryable: the SDK already retried, so come back later — after what the service asked for
if (err instanceof ImageStepError && err.retryable) queueForLater(ids, err.retryAfter);
else if (err instanceof JobFailedError) console.error(err.job.status, err.job.id);
else throw err; // not retryable: fix what err.param names
}Python
from imagestep import ImageStepError, JobFailedError
try:
client.ops.upscale(ids, wait=True)
except JobFailedError as err:
print(err.job["status"], err.job["id"])
except ImageStepError as err:
if not err.retryable:
raise # fix what err.param names
# retryable: the SDK already retried, so come back later — after what the service asked for
queue_for_later(ids, err.retry_after)CLI
imagestep jobs submit --op upscale --asset-ids "$ID" --wait -o json > job.json
case $? in
0) ;; # done: job.json is the finished job
4) sleep 30 && exec "$0" "$@" ;; # retryable — the same call again, later (the CLI already retried)
3) jq -r '.error | "\(.code): fix \(.param)"' job.json >&2; exit 3 ;; # refused — with -o json the error object is on stdout
2) echo "still running" >&2 ;; # --wait ran out of patience; the job did not — read it, do not resubmit
1) exit 1 ;; # the job ended FAILED or CANCELLED (job.json has it, verdict and all) — or nothing was sent: a bad flag, no key
esaccurl
body=$(curl -s https://api.imagestep.dev/api/v1/jobs \
-H "Authorization: ApiKey $IMAGESTEP_API_KEY" -H "Content-Type: application/json" \
-H "Idempotency-Key: $KEY" -d "$REQUEST")
if jq -e '.success' <<<"$body" > /dev/null; then
jq -r '.data.id' <<<"$body"
elif jq -e '.error.retryable' <<<"$body" > /dev/null; then
sleep 30 # then send the same request again, with the same Idempotency-Key
else
jq -r '.error | "\(.code): fix \(.param // "the account")"' <<<"$body" >&2; exit 3
fiMCP
# a failed tool call: isError is true and the result is ONE object — as text and as structuredContent — in the API's shape
{"error": {"code": "rate_limited", "message": "Rate limit exceeded", "retryable": true, "status": 429, "requestId": "5b1f0c9e-2a47-4d0e-9a63-7f1c2e8d4b90"}}
# retryable: true → call the same tool again, with the same idempotency_key
# retryable: false → change what "param" names; the same arguments will fail the same wayThe same fields on every surface: an ImageStepError carries them as properties, the CLI turns retryable into its exit code (4 retryable, 3 not), and an MCP tool error is the error object itself.
Error codes
The closed set, by status. Adding a code is a minor change; renaming or removing one is breaking. A sub-reason goes in details, never in a new code.
| code | HTTP | retryable | when it happens |
|---|---|---|---|
| invalid_param | 400 | no | a parameter is missing, malformed or out of range; param names it (details.fields, when several are). Fix it: the same body fails again |
| invalid_state | 400 | no | well-formed, but what it acts on is not in a state to honour it — cancelling a job that has finished, resuming one that has not failed |
| unsupported_format | 400 | no | the bytes are not an image format this service reads |
| unauthorized | 401 | no | no credential, or one this service does not accept — a revoked or mistyped key. Fix the key; retrying it gets an address blocked (below) |
| insufficient_credit | 402 | no | the balance does not cover the job; top up, then send it again. The dry run's sufficientCredit says so first |
| forbidden | 403 | no | authenticated, but not allowed — a built-in you tried to edit, a key on an endpoint only the console may use |
| asset_not_found · job_not_found · preset_not_found · not_found | 404 | no | no such thing, or it is not yours — the two are never told apart |
| method_not_allowed | 405 | no | the path exists, not for this method; the Allow header and details.allowed list the ones it takes |
| not_acceptable | 406 | no | nothing this endpoint produces satisfies your Accept; JSON is the only representation |
| idempotency_key_reuse | 409 | no | this Idempotency-Key already means a different request — or the first one's answer could not be kept (idempotency keys) |
| request_in_progress | 409 | yes | the first request under this key has not answered yet; send it again shortly, with the same key |
| subscription_not_actionable | 409 | no | the plan action names no subscription this service can act on |
| payload_too_large | 413 | no | the body is over the ceiling for its kind (below); details.limit is the ceiling in bytes. Split it into two requests |
| unsupported_media_type | 415 | no | the body's Content-Type is not one this endpoint reads; details.supported lists the ones it does |
| asset_count_exceeded | 422 | no | the outputs would take the account past its plan's stored-asset ceiling; delete assets or upgrade. The dry run's assetCountLeft says so first |
| resource_limit_exceeded | 422 | no | a per-account hard limit — webhook endpoints, for one — would be exceeded |
| job_not_resumable | 422 | no | a resume of an attempt that is not the latest, or past the resume budget; details.reason says which |
| provider_rejected | 422 | no | an upstream provider refused this input; the same input fails again (on job items) |
| rate_limited | 429 | yes | too many requests, or too many at once (below); wait the Retry-After seconds it carries |
| service_misconfigured | 500 | no | our side, and retrying will not help; tell us |
| internal_error | 500 | yes | unclassified, on our side; may be transient. With details.reason "query_timeout" the same request can time out again — narrow it before a second retry |
| provider_unavailable | 503 | yes | an upstream AI provider, the image worker or the host of a url you sent failed, timed out or is full; details.reason says which, and Retry-After, when present, how long to wait |
Failures inside a job
A job that was accepted and then failed carries no error envelope — the request was fine. The verdict is on the items: each failed one has an errorCode from the table above and that code's retryable — or neither, when a worker could not name the cause, which is to be read as not retryable — and the job repeats the verdict at the top. Which codes appear on items, and how resume uses them, is on jobs.
Idempotency keys
Every write (POST / PUT / PATCH / DELETE) accepts an optional Idempotency-Key header — one per logical operation, a UUID is the usual choice. It is what makes a retry safe:
| you send | you get |
|---|---|
| no key | normal execution, no deduplication |
| a new key | normal execution; the status and body are stored under (account, key) |
| same key, same body | the stored response, verbatim, plus Idempotency-Replayed: true — nothing runs again |
| same key, different body | 409 idempotency_key_reuse |
| same key while the first is still running | 409 request_in_progress, retryable |
- “The same body” is the same bytes: the method, the path, the query string and the raw body, hashed. A retry that serialises its JSON again — keys in another order, other whitespace — is a different body and a
409; resend the bytes you sent. - Keys belong to the account, not to one API key, and are kept 24 hours; after that the same key is new again.
- A 5xx releases the key, so a retry after one runs. A first request that died without answering holds its key for 5 minutes —
request_in_progressuntil then — and a retry after that runs. - Two limits, both of which mean the first request ran: a body over 1 MB is not fingerprinted, and a response over 512 KB is not stored. A repeat of either is refused with
409 idempotency_key_reuserather than replayed — read what the first one made instead of sending it again. - Ignored rather than honoured on the synchronous image endpoints, which create nothing a replay could protect; on the two batch reads that are a
POSTonly because a page of ids does not fit a query string; and onPOST /api/v1/assets/upload, whose own rule is that the same bytes are the same asset (existing: true) — send it again. A job's dry run ignores it too: it creates nothing, so pricing a body and then submitting it under one key works. - A job submit that asked to
waitreplays with the job as it is now — waiting again while it runs — rather than the snapshot it stored (jobs).
The SDKs and the CLI send a key on every write and reuse it across their own retries. By hand, and from an agent:
curl
KEY=$(uuidgen) # one per logical operation — keep it for every retry of this submit
curl -si https://api.imagestep.dev/api/v1/jobs -H "Authorization: ApiKey $IMAGESTEP_API_KEY" -H "Content-Type: application/json" \
-H "Idempotency-Key: $KEY" -d '{"op":"grayscale","assetIds":["<asset-id>"]}'MCP
# tool call — the tools take the key as an argument; repeat the call with the same one
transform {"op": "grayscale", "asset_ids": ["<asset-id>"], "idempotency_key": "3f0e7c1a-9b2d-4e55-8a61-0c9d5e2f7b14"}Sent twice, the second answer is the first one, byte for byte, and says so. Its X-Request-Id is this request's; a stored failure's error.requestId is the original's.
HTTP/2 201
content-type: application/json
idempotency-replayed: true
x-request-id: ca3c0512-775e-42f4-83e8-81dce29bb702Rate limits and quotas
Every response that got past authentication carries the current request budget:
RateLimit-Limit: 600
RateLimit-Remaining: 587
RateLimit-Reset: 43RateLimit-Limit is the requests the window allows, RateLimit-Remaining what is left of it, and RateLimit-Reset the seconds until the window turns. The budget is counted per account (per client IP when anonymous) on each server, so it is a ceiling to pace by, not a quota to spend down — the plan's limits are other codes, so a program can tell “slow down” from “you have run out”:
| ceiling | code | what to do |
|---|---|---|
| 600 requests per 60 s, per account (per IP when anonymous) | 429 rate_limited + Retry-After | slow down; pace by RateLimit-Remaining, and retry after the header says |
| 60 failed authentications from one IP within 60 s | 429 rate_limited + Retry-After | fix the key: every request from that address is refused until the window turns, a good key included |
| 4 synchronous image calls in flight per account | 429 rate_limited + Retry-After | the same: it is the sync lane's own concurrency |
| 8 job waits open per account | 429 rate_limited + Retry-After | read the job without wait, or wait again shortly; a waited submit is never refused for it |
| a request body of 8 MB of JSON, 34 MB of image or multipart | 413 payload_too_large | split it into two requests; details.limit is the ceiling |
| credit — for AI ops, and for deterministic runs past the Free plan's monthly allowance, which are paid rather than refused | 402 insufficient_credit | top up at details.topUpUrl; details says what was needed and what could be spent, and the dry run's sufficientCredit warns first |
| stored assets | 422 asset_count_exceeded | delete, or upgrade at details.upgradeUrl; details.left says how many fit, and the dry run's assetCountLeft warns first |
| 2 running jobs per type per account | none — the next one waits in PENDING | nothing; it starts when one finishes |
Pagination
Every list takes page (from 0) and perPage (100 by default and at most); meta.hasMore says whether to ask again, and meta.nextCursor where the next page starts. One shape, every list — assets, collections, jobs, a job’s items, an endpoint’s deliveries, your reports and the templates. A page or perPage out of range is clamped rather than refused, so perPage=500 is not an error, it is 100; meta always reports what was actually served:
{
"total": 412,
"page": 0,
"perPage": 100,
"hasMore": true,
"nextCursor": "YXN0OjE3OTAwMDAwMDAwMDA6YXN0XzRjMWYwZThhOWIyZDRlNmY4YTFjM2I1ZDdlOWYwYTJi"
}To read on, send nextCursor back as cursor instead of a page number. The service then reads the rows after it and counts nothing, so the answer carries no total and no page, and nextCursor is null on the last page:
{
"perPage": 100,
"hasMore": false,
"nextCursor": null
}That is the way to walk a long list. Page k by number makes the service read and throw away every row before it and count the whole filter again, so walking a hundred thousand assets by page number is a thousand counts; by cursor every page costs what the first did. It also holds still: rows that arrive while you walk land on top, and a cursor carries on below them. Send the cursor unchanged and with the same filters — it is opaque, and it names a position in one listing’s order, not a query. With page beside it, or from another listing, it is 400 invalid_param on cursor.
The one loop worth getting right is the one that walks to the end — and the SDKs, the CLI and the n8n node have it, so you do not write it. The next page is wherever the answer says it starts:
JavaScript
for await (const asset of client.assets.iterate({ collection: "shoot-01" })) {
console.log(asset.id);
}Python
for asset in client.assets.iterate(collection="shoot-01"):
print(asset["id"])CLI
imagestep asset list -c shoot-01 --all -o json | jq -r '.[].id'curl
cursor=""
while :; do
body=$(curl -s "https://api.imagestep.dev/api/v1/assets?collection=shoot-01${cursor:+&cursor=$cursor}" -H "Authorization: ApiKey $IMAGESTEP_API_KEY")
echo "$body" | jq -r '.data[].id'
[ "$(echo "$body" | jq -r .meta.hasMore)" = true ] || break
cursor=$(echo "$body" | jq -r .meta.nextCursor)
doneMCP
# one page per call; the answer carries hasMore and nextCursor — send that back as cursor for the next
search_assets {"collection": "shoot-01"}Ask for less when a page has somewhere small to go — the console’s grid asks for 48, the MCP search_assets tool for 20, because a page lands in a model’s context and a context is a viewport. Anything walking the list itself should take the 100: it is the fewest round trips.
A document that would otherwise carry a list inside it does the same: a job comes back with its first 100 items and itemsTruncated, and the rest are GET /api/v1/jobs/{id}/items.
Some listings are not paged at all, because they are catalogues or small bounded sets a caller wants whole: GET /api/v1/ops, /api/v1/ai-models, /api/v1/presets, /api/v1/webhook-endpoints, /api/v1/jobs/counts and a template’s /versions answer with the whole list and no meta.
Request ids
Every response carries X-Request-Id, and every error repeats it as error.requestId. Send your own — up to 128 characters of letters, digits, ._-: — and it is used as is, so one string joins your trace to ours; anything else is replaced, not trimmed, and without one you get a UUID. It is set before authentication, so a 401 and a 429 carry it too. Quote it when you write to support, or when an agent files a report with POST /api/v1/feedback.