Skip to content

MCP server

Not "image generation via MCP" — that is free everywhere. This is the pipeline: assets, atomic ops, versioned presets and jobs as tools an agent can plan with. Every tool takes and returns references — asset ids, sizes, public URLs — so image bytes never enter the context window, every call that spends money can be priced first, and every failure says whether it is worth retrying.

# Claude Code, local: reads files from your disk
claude mcp add imagestep -e IMAGESTEP_API_KEY=is_sk_… -- npx -y @imagestep/mcp

# Claude Code, hosted: nothing to install
claude mcp add --transport http imagestep https://mcp.imagestep.dev/mcp --header "Authorization: Bearer is_sk_…"

MCP or the CLI?

Both reach the same API with the same key; the question is whether the agent has a shell. Claude Desktop, a hosted agent or any MCP client without one uses this server. Claude Code, Codex or Cursor working inside a repository is usually better served by the CLI and its agent skill — local files, batches, and a script a person can read and run again. The decision table and the operating contract behind both are on the agents page.

Two servers, one key

The same package runs in two places, and which one you connect decides what "a file" and "a wait" mean. Both need an ImageStep API key, created on /keys; both reach the same API as the SDKs and the CLI.

ComparedLocal (stdio)Hosted (Streamable HTTP)
What runsnpx -y @imagestep/mcphttps://mcp.imagestep.dev/mcp
InstallNode 20+ on the machine the client runs onnothing — the client connects over HTTPS
KeyIMAGESTEP_API_KEYAuthorization: Bearer is_sk_…
Local filesyes: file_paths uploads from diskno: pass asset_ids or public urls
A synchronous resulta temp file on your machine (path), gone when the server exitsa signed URL, valid about five minutes
Wait per calldefault 180 s, up to 600 s (wait_seconds)90 s, the default and the cap — then the job handle with timedOut: true
Good forClaude Desktop, Cursor, Claude Code on your laptophosted agents, shared workspaces, anything without a shell

The hosted server is stateless: every request builds a server bound to the key in its header and tears it down after the response. There are no sessions to leak and nothing to warm up. The op catalogue behind it is cached for five minutes, so a tool call does not pay for a catalogue read.

Connect a client

Claude Code

claude mcp add imagestep -e IMAGESTEP_API_KEY=is_sk_… -- npx -y @imagestep/mcp                                   # local
claude mcp add --transport http imagestep https://mcp.imagestep.dev/mcp --header "Authorization: Bearer is_sk_…"   # hosted
claude mcp list                                                                                                # confirm it is connected

Add --scope project to share the entry with a repository. It lands in .mcp.json, which is committed, so keep the key out of it: Claude Code expands ${IMAGESTEP_API_KEY} from each person's environment.

{ "mcpServers": { "imagestep": { "command": "npx", "args": ["-y", "@imagestep/mcp"], "env": { "IMAGESTEP_API_KEY": "${IMAGESTEP_API_KEY}" } } } }

Claude Desktop

Add this to claude_desktop_config.json (Settings → Developer → Edit Config) and restart Claude Desktop. Local, so the agent can upload files from your disk.

{
  "mcpServers": {
    "imagestep": {
      "command": "npx",
      "args": ["-y", "@imagestep/mcp"],
      "env": { "IMAGESTEP_API_KEY": "is_sk_…" }
    }
  }
}

Cursor

The same shape in .cursor/mcp.json (per project) or ~/.cursor/mcp.json (everywhere) — local:

{ "mcpServers": { "imagestep": { "command": "npx", "args": ["-y", "@imagestep/mcp"], "env": { "IMAGESTEP_API_KEY": "is_sk_…" } } } }

or hosted:

{ "mcpServers": { "imagestep": { "url": "https://mcp.imagestep.dev/mcp", "headers": { "Authorization": "Bearer is_sk_…" } } } }

Any client that speaks Streamable HTTP

Point it at https://mcp.imagestep.dev/mcp and send the key as Authorization: Bearer is_sk_… (the API's own ApiKey is_sk_… scheme is accepted too). A request without a key is a 401 whose body is the usual error object. There is no OAuth flow and no session id; every request is authenticated on its own. The answer is one JSON body, but the transport still wants both types in Accept, and answers 406 without them.

curl -s https://mcp.imagestep.dev/mcp \
  -H "Authorization: Bearer $IMAGESTEP_API_KEY" -H "Content-Type: application/json" -H "Accept: application/json, text/event-stream" \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'

Check the key, not just the server

Ask the agent to run search_assets with per_page: 1. It needs the key, so a working setup answers with a page of assets (empty on a new account) and a wrong or revoked key with unauthorized. Listing the tools or reading imagestep://ops proves only that the server started: the catalogue is public and answers whatever key is sent.

What a conversation looks like

Each request below is a few tool calls, and none of them puts an image into the context.

"Upload ~/shoot/product.jpg, remove the background, upscale it 2× and give me the URL"

  1. transform — { op: "remove_bg", file_paths: ["/Users/me/shoot/product.jpg"] }: an AI op, so a job; the answer carries outputs[0].assetId.
  2. transform — { op: "upscale", asset_ids: ["<that id>"], parameters: { scaleFactor: 2 } }.
  3. Read outputs[0].publicUrl from the second answer. Two calls, zero image bytes in context.

"What would it cost to upscale everything in the shoot-01 collection?"

  1. search_assets — { collection: "shoot-01", per_page: 100 }; collect the ids.
  2. transform — { op: "upscale", asset_ids: [...], parameters: { scaleFactor: 2 }, dry_run: true }; quote estimate.estimatedCredits and sufficientCredit. Nothing was created.

"Make the OG image for these three blog posts"

  1. Read imagestep://ops/render_template for the parameter names, if not already in context.
  2. transform — { op: "render_template", parameters: { templateId: "builtin-template-og-image", items: [ {...}, {...}, {...} ] } }: one job, three PNGs, three public URLs.

The tools

generate, transform and run_preset submit work and spend credits, and share the arguments in the table below on top of their own. save_preset and send_feedback write for free, and job_status and search_assets only read. No tool cancels or resumes a job, or deletes anything.

ArgumentTypeMeaning
collectionstring ≤ 200Collection the new assets go in — uploads, fetched URLs and outputs; filter on it later with search_assets. Without one, an output is in its input's collection.
retention_daysinteger ≥ 1Keep what this call stores — the images you hand it and the ones it makes — this many days instead of the account's plan retention. Shorter only: more is kept for the plan's time. Use it when the outputs are consumed at once (posted, sent on) and need not be kept.
idempotency_keystring 1–255Your key for this submission. The same key with the same arguments returns the job the first call created instead of a second one.
waitboolean, default trueWait for the job to finish and return its outputs. false returns the job handle at once; poll job_status.
wait_secondsinteger 5–600, default 180How long to wait when wait is true. The hosted server waits at most 90 s per call (90 s when you name none) and clamps a larger value.
publishboolean, default truePublish the outputs so each has a stable publicUrl on the CDN — the output at full size, never a thumbnail.

transform and run_preset also take their images the same way — any mix of these (more below):

ArgumentTypeMeaning
asset_idsstring[] ≤ 500Assets already in the account.
file_pathsstring[] ≤ 100Local files, uploaded first. Only when the server runs on your machine over stdio; the hosted server refuses them.
urlsstring[] ≤ 100Public http(s) image URLs. The service fetches them — never this server — and each becomes an asset; a private or loopback address is refused.
dry_runboolean, default falsePrice only: returns the estimate for exactly this call, creates nothing, uploads nothing and charges nothing. asset_ids are priced as themselves, file_paths and urls by how many there are.

generate — images from a prompt

Text to image. Returns asset ids and public URLs, never bytes.

ArgumentTypeMeaning
promptstring 1–4000, requiredWhat to draw.
countinteger 1–10, default 1How many images.
modelstringAn image model id. Default: the op's defaultModel (the resource imagestep://ops/generate names it); every model with its price is imagestep://models/ai_image.
parametersobjectModel-specific parameters, for example {"aspectRatio": "16:9"}. The model catalogue says what each accepts.
presetstringA saved preset — slug, id, or slug@version to run exactly that version — whose generate step supplies the model, default parameters and consistency subjects, so a batch shows the same character or product. Save one with save_preset; your prompt, model and parameters still override it.
dry_runboolean, default falsePrice only: returns the estimate for exactly this call, creates nothing, uploads nothing and charges nothing. asset_ids are priced as themselves, file_paths and urls by how many there are.
{ "prompt": "a red bicycle on a white background, studio light", "count": 2, "collection": "bikes", "idempotency_key": "bikes-2026-09-15-a" }

transform — one atomic op on images, or render a template

Apply one op from the catalogue — every op there but generate, which has its own tool. The list the tool offers is read from the live catalogue when the server starts (hosted: per request, from the five-minute cache), so an op added to the API is callable without a new package version; when the catalogue cannot be read the tool falls back to the list it shipped with and its description says so.

ArgumentTypeMeaning
opone of 24 values, requiredThe op to run — the list above. One op does one thing: resize resizes and does not re-encode; a parameter it does not take is refused.
promptstring ≤ 4000Required by edit, optional for ops that have a default prompt, ignored by the rest — the tool description says which, from the catalogue.
modelstringOverride the op's default model (AI ops). Models with prices: imagestep://models/ai_image, or imagestep://models/analyze for analyze.
parametersobjectThe op's parameters, by name. The tool description lists every op's parameters as name(type=default) from the live catalogue, with the values a parameter takes, what a required one means, and for an AI op its model's own parameters object; the full contract with descriptions is the resource imagestep://ops/{op}. Never invent one.
variantsarray 1–20 of { name?, parameters }Several outputs from one call, as one job: the op runs once per variant for every input, each variant merged over the shared parameters. Every social size from one call: parameters {"fit":"cover"}, variants [{"name":"ig","parameters":{"width":1080,"height":1350}}, {"name":"x","parameters":{"width":1600,"height":900}}]. Each result item carries its variant name.

Two modes, chosen by the input — the agent does not pick. When all of these hold, the op runs synchronously and nothing is stored in the account: exactly one file_paths or urls entry, an op the catalogue gives a synchronous form (its syncEndpoint — most deterministic ops, never an AI op), wait: true, no variants, no dry_run. There is no asset id afterwards because none was created. Where the result goes depends on the server: local, a temp file and the answer carries its path; hosted, a signed url that expires in about five minutes. Everything else — several images, any asset_ids, any AI op, wait: false — runs as a job and returns asset ids and public URLs, with progress and webhooks.

Two ops are special. render_template takes no images: pass parameters: {templateId: "<id or id@version>", items: [one object of template variables per PNG]}, and any image handed to it is refused rather than silently uploaded. read_metadata costs nothing and makes no job: on one local file it stores nothing either, while urls and asset_ids are read as assets — fetched URLs become assets that count toward the asset quota.

One local file, an op with a synchronous form, and a wait — this one runs synchronously and stores nothing:

{ "op": "resize", "file_paths": ["/Users/me/shoot/hero.jpg"], "parameters": { "width": 1600 } }

An AI op on a stored asset is a job; the answer is a job handle with its outputs:

{ "op": "upscale", "asset_ids": ["ast_1a2b"], "parameters": { "scaleFactor": 2 }, "idempotency_key": "hero-upscale-1" }

Every social size from one call, as one job:

{
  "op": "resize",
  "asset_ids": ["ast_1a2b"],
  "parameters": { "fit": "cover" },
  "variants": [
    { "name": "ig", "parameters": { "width": 1080, "height": 1350 } },
    { "name": "og", "parameters": { "width": 1200, "height": 630 } }
  ]
}

One PNG per row, and no images in:

{
  "op": "render_template",
  "parameters": { "templateId": "builtin-template-og-image", "items": [{ "title": "Hello" }, { "title": "World" }] }
}

run_preset — a saved, versioned list of steps

Presets are how you chain ops (resize → convert → sharpen) and how a batch stays consistent: the same steps, the same version, every run. Pass a slug or id — a built-in like builtin-util-to-webp, or one of your own. No tool lists presets, so a slug that does not exist is answered with preset_not_found and a message naming every preset on the account. For a recurring character or product, a preset's subjects carry both halves of consistency (reference images and a locked descriptor), so the agent names the subject as {{subject.<name>}} and does not re-describe it. A preset whose first step is generate starts from its prompt and takes no images: call it with preset alone, or give this run its own prompt — a new scene for the same subjects, without saving a version per scene — and a count. Any other preset with no image is invalid_param on asset_ids. More on presets: presets.

ArgumentTypeMeaning
presetstring, requiredPreset slug or id, or slug@version to run exactly that version — what save_preset answers with.
promptstring 1–4000This run's prompt, in place of the one on the preset's AI step — a new scene for the same subjects, with each named as {{subject.<name>}}. A preset of several steps, or with no AI step, has no one step for it and answers invalid_param on prompt.
countinteger 1–10How many images a preset that starts from a prompt makes; default 1. With images there is one output per input.
{ "preset": "builtin-util-web-optimize", "urls": ["https://example.com/a.jpg", "https://example.com/b.jpg"], "collection": "site" }
{ "preset": "bottle-shots@2", "prompt": "{{subject.bottle}} on a beach at dusk", "count": 3 }

save_preset — keep a chain as a versioned preset

A chain the agent has run twice is a preset: save it once, and every later run is one call to run_preset with the same steps at the same version. Each step is one op from the catalogue with the parameters transform takes for it. The service checks every step before storing anything and refuses a bad one — or an order that cannot run — with invalid_param naming steps[i]; steps that are legal but probably not what was meant are saved and answered with warnings. Steps may mix deterministic ops and AI ops: a preset with an AI step beside other steps runs as one chain job, and the dry run prices it step by step. Processing-registry steps are written in the console, the SDK or the CLI, where a person reads them, never from a tool. Free: no credits, no job.

ArgumentTypeMeaning
namestring 1–200, requiredWhat a person sees in the console.
slugstring 1–100URL-safe identifier, derived from name when omitted. builtin- is reserved, and a slug already in use is refused.
descriptionstring ≤ 2000One sentence on what the preset is for.
stepsarray 1–20 of { op, model?, prompt?, parameters? }, requiredRun in order. model and prompt matter on an AI step only; {{subject.<name>}} in a prompt expands to that subject's descriptor.
subjectsarray ≤ 4 of { name, referenceAssetIds, descriptor }Consistency for a recurring character or product: up to 4 of your finished assets pin the geometry, the descriptor locks colour and material. Needs a generate or edit step.
idempotency_keystring 1–255The same key with the same arguments returns the preset the first call saved instead of a second one.
{
  "name": "Web hero, 1600 WebP",
  "slug": "web-hero",
  "steps": [
    { "op": "resize", "parameters": { "width": 1600 } },
    { "op": "convert", "parameters": { "format": "webp", "quality": 82 } }
  ]
}

The answer names the version that was saved — hand preset to run_preset as it is, so a later run is exactly this version:

{
  "preset": "web-hero@1",
  "id": "pre_0b9f3c1e5d2a4e779c416a2f8d1e7b30",
  "slug": "web-hero",
  "version": 1,
  "steps": 2,
  "note": "Run it with run_preset {preset: \"web-hero@1\"}. Changing its steps later makes a new version; this one stays runnable as web-hero@1."
}

job_status — progress and outputs

Per-item states of a job by id and, once it is finished, its outputs — the same job handle a write tool answers with. Free, and read-only by default: checking on a job does not make its results public.

Items come a page at a time, 100 to a page. When a job has more, the answer carries itemsCursor — send it back as items_cursor to read the next page, until an answer has none — and items_status reads only the items in one state: FAILED is what a resume would run again.

It reads the job an agent was handed, and there is no tool for the history — listing past jobs, or counting them by status. An agent follows its own work by id; reading back what ran is a person's check, on the console's Jobs page, or a program's, with GET /api/v1/jobs and GET /api/v1/jobs/counts.

ArgumentTypeMeaning
job_idstring, requiredThe jobId a write tool returned.
publishboolean, default falseAlso publish the outputs and return their public URLs. Off by default — pass true when you want the URLs.
items_statusPENDING | PROCESSING | COMPLETED | FAILED | CANCELLEDOnly the items in this state, a page at a time — FAILED is what a resume would run again. The answer's items are that page, and its itemsCursor reads on.
items_cursorstring ≤ 500The itemsCursor of an earlier job_status answer: the page of items after it. Send the same items_status as that call, if it had one.

search_assets — find what is already there

Find assets by collection, keyword, tag, MIME type, size, ingest state, when they were made, or the job that made them. Returns references — id, dimensions, tags, public URL, expiry — never bytes, a page at a time: follow nextCursor to read on. With group_by: "collection" it lists your collections instead, which is how an agent checks a name before it filters or submits into it: a misspelt collection is simply a new one. Free.

ArgumentTypeMeaning
group_bycollectionList collections — collection, count, lastAddedAt — most recently added to first, instead of assets. Only q (part of the name), page, cursor and per_page apply; any other filter is invalid_param.
qstring ≤ 200Keyword in the name or metadata; with group_by, part of a collection's name.
collectionstring ≤ 200Collection name, matched exactly.
tagstring ≤ 100One of the asset's tags — your own labels — matched exactly.
mimestring ≤ 100Exact MIME type, for example image/png.
viewALL | PUBLISHED, default ALLOnly published assets, or all.
min_widthintegerPixel lower bound on width.
min_heightintegerPixel lower bound on height.
statusPROCESSING | DONE | FAILEDIngest state. FAILED is an upload whose ingest never finished — check this after uploading a batch.
created_fromstring ≤ 40Only assets made since: an ISO-8601 date or instant in UTC (2026-09-14), or epoch millis.
created_tostring ≤ 40Only assets made before; a bare date covers the whole of that day.
has_collectionbooleanfalse is everything not in a collection — what to tidy up. Cannot be combined with collection.
job_idstring ≤ 100Everything one job produced — the id job_status returns. An upload has no job and never matches.
pageinteger ≥ 0A numbered page, from 0; that answer carries total and hasMore.
cursorstring ≤ 500The nextCursor of the previous answer — the page after it, read without counting. How to read on.
per_pageinteger 1–100, default 20Page size — 20 here, where the REST default is 100: a page lands in your context, and a context is a viewport. Ask for more when you are collecting ids rather than reading rows.

send_feedback — report what ImageStep could not do

Whenever the agent concludes ImageStep cannot do what the task asks — an input it cannot handle, an action no tool performs, an op that does not exist, a parameter that is missing, a result that is wrong — it reports it here before it tells the person, instead of working around it. Reports reach the people who build ImageStep and are kept with the account's own data. Free: no credits, no job.

ArgumentTypeMeaning
kindcapability_gap | bug | other, requiredcapability_gap: something ImageStep cannot do. bug: something it does wrong.
messagestring 1–4000, requiredWhat was needed, what was tried, what happened.
opstring ≤ 64The op it concerns. It is not checked against the catalogue: an op that does not exist yet is the most useful report.
contextobjectWhat helps reproduce it — ids, parameters, the requestId of a failed call. At most 4000 characters once encoded; never image bytes or keys.
idempotency_keystring 1–255The same key with the same arguments files one report, not two.
{ "kind": "capability_gap", "op": "detect_faces", "message": "Needed face boxes to crop portraits; no op returns them.", "context": { "tried": "analyze" } }

Resources: the catalogue in the session

An MCP-only client cannot curl the API or open the console, and "never invent a parameter" is a rule it can only follow if it can read the contract. The server exposes it as resources — attach one to the session the way you attach a file. Each is read from the service when it is asked for (the op catalogue through the same five-minute cache the tools use); a read that cannot reach the service fails rather than answering from a stale copy.

URITypeWhat it holds
imagestep://agent-guidelinestext/markdownThe operating contract: read the catalogue, price before spending, branch on retryable, keep batches consistent with a preset, report a gap instead of working around it.
imagestep://opsapplication/jsonThe whole op catalogue: every op with its parameter contract (type, default, description), prompt and asset requirements, default model, pricing and whether it has a synchronous form.
imagestep://usageapplication/jsonWhat the account has spent over the last 30 days, by op: credits charged, jobs created, items settled and synchronous calls. Read it before a large batch.
imagestep://ops/{op}application/jsonOne op's entry — listed per op in the client's resource picker when the catalogue was readable at startup.
imagestep://models/{mode}application/jsonThe model catalogue with prices: ai_image (generate, edit and the AI image ops) or analyze.

Inputs: ids, files, URLs

transform and run_preset take any mix of asset_ids, file_paths and urls, and turn them into one list of asset ids before the job is submitted. Files and URLs become assets in your account (in collection, if you name one), which is what makes the result of one step the input of the next: carry the resultAssetId or the outputs[].assetId forward, never a re-uploaded copy.

  • file_paths work only over stdio, where the agent and the server share a filesystem; the hosted server answers invalid_param on file_paths. The server process reads them, so pass absolute paths — a relative one resolves against wherever the client started the server.
  • urls must be public. A loopback, private or link-local address is refused by this server before anything leaves the process, and again by the service that fetches; the service does not follow redirects. Its size and count limits are on getting images in.
  • The one exception is the synchronous path: one file or URL with an op that has a synchronous form is sent straight to the transform endpoint and stores nothing, so there is no asset id to carry — ask for a job (wait: false, or several inputs) when you need one.

Waiting, timeouts and job_status

A write tool with wait: true (the default) holds the call open until the job is finished, up to wait_seconds, then answers with the outputs. If the wait runs out, the answer is not an error: it is the job handle with timedOut: true and a note. The job keeps running and is charged for what completes, so the agent polls job_status with that jobId — with publish: true once it wants the URLs — and does not submit the same work again. How long to ask for is in the tool's own description: the server reads each op's typicalSeconds off the catalogue — the time one item usually takes on the default model, measured on production — so the figures an agent sees are the service's, not a list kept here.

The hosted server waits at most 90 seconds per call, and clamps a larger wait_seconds to that. The reason is the edge in front of it: a response that has not started after 100 seconds is dropped with an HTML error that carries no job id, while the job keeps running; 90 seconds leaves room for the submit and the last poll. Over stdio the wait is what you ask for: default 180 s, maximum 600 s.

wait: false returns the handle immediately — the right shape for a batch the agent wants to fire and check later, or for a client whose own tool timeout is shorter than the wait.

What comes back

Every answer is one JSON object, in the tool result's text and in structuredContent. A job — a write tool's answer, or job_status — answers with a job handle: the trace an agent needs to say whether a step worked and what it cost, without a dashboard.

{
  "jobId": "job_d59ca98c1d6d4b5fbc8eef2477dc5681",
  "type": "ai-edit",
  "status": "COMPLETED",
  "op": "upscale",
  "totalItems": 1,
  "completedItems": 1,
  "failedItems": 0,
  "creditsCharged": 1512,
  "items": [
    { "index": 0, "status": "COMPLETED", "sourceAssetId": "ast_1a2b", "resultAssetId": "ast_9x8y", "provider": "fal", "model": "fal-ai/clarity-upscaler", "durationMs": 8412, "credits": 1512 }
  ],
  "outputs": [
    { "assetId": "ast_9x8y", "name": "hero.png", "status": "DONE", "mimeType": "image/png", "width": 3200, "height": 2400, "size": 1834211, "publicUrl": "https://cdn.imagestep.dev/ast_9x8y" }
  ]
}
FieldWhat the agent does with it
statusCOMPLETED, FAILED or CANCELLED once it has stopped; anything else means it is still moving — poll job_status
timedOut · notethe wait ran out, not the job: poll job_status with the jobId, never submit again
failedItems · errorCode · retryablea job that ran and failed: the verdict on its items — see Errors below
creditsChargedwhat the job cost; an item that failed was not charged
items[]one per input (per variant): status, the ids in and out, error, errorCode, retryable, provider, model, durationMs, credits, and for a chain step and failedStep (the segment a resume starts again from) — the first 100; past that itemsNote names the job_status call (items_cursor) that reads on
outputs[]the assets it made, with publicUrl when published — the first 20; past that outputsTruncated is true and outputsNote names the search_assets call (job_id and cursor) that reads the rest

Four other shapes, each named by what you asked for:

  • A synchronous transform answers { mode: "sync", stored: false, op, path | url, mimeType, width, height, bytes, durationMs } — with expiresInSeconds on the hosted server — and a note saying no asset id exists.
  • A dry run answers { estimate: { estimatedCredits, creditBalance, sufficientCredit, totalItems, costPerItem, … } }; run_preset adds the resolved preset's id, slug, version and step count.
  • An analyze job answers with the job handle plus analyses: one { assetId, output } per completed item, the answer in output. It creates no output asset and publishes nothing.
  • read_metadata answers { mode: "sync", stored: false, … } with the metadata of one local file, or { assets: [ … ] } for URLs and asset ids.

Price before you spend

AI ops charge credits per item; the tool descriptions quote each op's price on its default model, read from the catalogue. Deterministic ops are free on paid plans; on Free they count against the monthly allowance, and each run past it is paid from the balance. Every write tool takes dry_run: true, which runs the same resolution a real submit runs — preset lookup, model validation, item count, per-item price — and stops before the job; how the estimate is computed is on jobs. A dry run doubles as validation: a bad parameter is refused here, without spending a job to find out. It creates nothing, and uploads nothing: the estimate never depends on the image itself — only on the op, model, parameters and how many images — so file_paths and urls are priced by count (imageCount) and asset_ids as themselves.

sufficientCredit: false in an estimate is not an error. It says a real submit would answer insufficient_credit, and no change to the request fixes that — the agent should stop and tell the person. Prices for every model are on /pricing and in the resource imagestep://models/ai_image; what the account has spent lately is imagestep://usage.

Retrying safely

An MCP client that times out makes the agent call the tool again. Without a key, that is a second job and a second charge. Pass the same idempotency_key on generate, transform or run_preset: the same key with the same arguments returns the job the first call created, and the same key with different arguments is refused with idempotency_key_reuse (how the service keeps keys). Use one whenever a retry is possible — which is whenever wait is on. save_preset and send_feedback take one too: a retried save returns the preset it saved, a retried report files one.

The tools also carry MCP annotations, the hints a client reads to decide what needs a person's confirmation: job_status and search_assets are read-only; the write tools are not idempotent on their own; and none is destructive, because every job writes new assets and never overwrites the one it read.

Errors

A failed call is a tool result with isError: true and one object, in the API's own shape, so the agent branches on the same fields it would from the REST API or the CLI:

{ "error": { "code": "invalid_param", "message": "width must be between 1 and 16384", "retryable": false, "param": "width", "status": 400, "requestId": "3c9e1f40-7b2d-4a8e-9f61-0d5a2b7c8e14" } }
  • Branch on retryable, not on the message. true: wait, then make the same call again. false: fix what param names; when param is null (for example insufficient_credit) the account is the problem, not the request. The codes and what each means are on errors, retries & limits.
  • details carries the sub-reason when there is one; requestId is what to quote when reporting the failure to a person. A refusal this server makes itself, before any request (a file_paths on the hosted server, a private URL), has code, message, retryable and param only.
  • A job that was accepted and then failed is a different case, and not an error result: the answer is the job handle with status: "FAILED", failedItems, and each item's error, errorCode and retryable. What did not complete was not charged; creditsCharged is what was. No tool resumes it — that is a resume from the CLI, the SDKs or the API, worth its price only when retryable is true.

Run it yourself

IMAGESTEP_API_KEY=is_sk_… npx -y @imagestep/mcp                                            # stdio, one key for the process
npx -y @imagestep/mcp --http --port 8787                                                   # Streamable HTTP on /mcp, key per request; GET /healthz
IMAGESTEP_API_KEY=is_sk_… IMAGESTEP_BASE_URL=http://localhost:5001 npx -y @imagestep/mcp   # against another API host

The HTTP mode is the same code the hosted server runs: stateless, one server per request, file_paths disabled, the 90-second cap on waits. --port falls back to PORT, then 8787. It listens on 127.0.0.1 and answers only to its own loopback names, so a web page that points a hostname at your machine cannot drive it; --host 0.0.0.0 opens it to the network, and --allowed-hosts names the hosts it answers to there. Package: @imagestep/mcp on npm, MIT, source in the repository. The server card at /.well-known/mcp/server-card.json describes the hosted server for clients that discover servers by URL.