Skip to content

Edit an image with a prompt

edit takes an image and an instruction and gives back a changed image: an object removed, a background replaced, a style applied. It is the catalogue op for changes you can describe but cannot express as a number.

What it does

Image in, sentence in, image out. The instruction covers three jobs that used to need three different programs: an instruction edit ("make the sky overcast"), a style transfer ("as a pencil sketch") and inpainting by description ("remove the cup on the left") — no mask, because the description is the mask. The output is a new asset; the input stays as it was.

Reach for it when the change is a judgement, and reach for something else when it is arithmetic. Cropping to 1200×630, converting to WebP, rotating 90°, writing a price in the corner: those are deterministic ops — exact, repeatable, free on a paid plan, and available synchronously; a logo in the corner is one too (overlay), as a job. Cutting the subject out is remove_bg, which is one op and one price instead of a prompt that may or may not land.

Call it

One op, six surfaces, one vocabulary: the op name below is the same string everywhere — the SDKs' named helpers only spell it in their language's case, ops.edit in JavaScript — because every surface enumerates GET /api/v1/ops instead of carrying its own list. The request is the catalogue's own example for this op, so it is the one the service has checked.

// pnpm add imagestep
const job = await client.ops.edit("ast_1a2b", "replace the background with a plain warm-grey studio wall", { wait: true });
const [out] = await client.jobs.outputs(job);

The n8n tab is a workflow to paste onto the canvas: the edit entry of the node's Op dropdown, which is filled from the same catalogue. Full references: SDKs, CLI, MCP, REST.

Parameters, model and price

Read live from GET https://api.imagestep.dev/api/v1/ops/edit — the same entry the SDKs, the MCP server and the n8n node enumerate, so this table cannot fall behind the API. The rows are fields of the request body. parameters goes to the model: its keys are the model's own, listed under parameters for each model in GET /api/v1/ai-models, and a key the model does not declare, or a value outside its options or bounds, is 400 invalid_param naming it.

edit

ai job type ai-edit · input scaled to ≤ 1536 px on the long edge

Image + prompt → image (instruction edit, style transfer, inpaint by description).

ParameterTypeWhat it does
promptstringThe edit instruction.
modelstringA model whose categories include 'image_edit' — GET /api/v1/ai-models lists them. Any other model is 400 invalid_param, with the ones that fit in details.allowed. default google/gemini-3.1-flash-image-preview
parametersobjectPassed to the model; GET /api/v1/ai-models lists each model's own, and a key or value it does not declare is 400 invalid_param. The default model reads aspectRatio and imageSize — the output size, whatever the input's, and what its price depends on.

AI op: per-item USD from the model's price; credits are charged per item. Price a batch with POST /api/v1/jobs?dryRun=true before spending. Default model google/gemini-3.1-flash-image-preview: $0.0567 - $0.1903 per item.

That is how it is priced. What one request costs — another model, a bigger scale factor, forty images — is the dry run, which prices the exact body you are about to send without creating anything.

In a workflow

An edit is a judgement, so the useful pattern is to make it repeatable: save the instruction as a preset and version it. Then a wording change is version + 1 rather than a diff somewhere in your code, and slug@3 pins the version that produced the batch you already shipped.

{
  "name": "Studio wall",
  "slug": "studio-wall",
  "description": "One saved instruction, run over a shoot, then tidied up deterministically.",
  "steps": [
    {
      "op": "edit",
      "prompt": "replace the background with a plain warm-grey studio wall"
    },
    {
      "op": "resize",
      "parameters": {
        "width": 2000,
        "fit": "inside"
      }
    },
    {
      "op": "convert",
      "parameters": {
        "format": "webp",
        "quality": 82
      }
    }
  ]
}
import { readFile } from "node:fs/promises";

const preset = await client.presets.create(JSON.parse(await readFile("studio-wall.json", "utf8")));
console.log(preset.slug, preset.version); // "studio-wall" 1

Run it pinned — studio-wall@1 keeps a whole shoot on the instruction it started with, whatever you change afterwards:

const price = await client.presets.run("studio-wall@1", ["<asset-id>","<another-id>"], { dryRun: true }); // nothing is created
console.log(price.estimatedCredits, price.steps);

const job = await client.presets.run("studio-wall@1", ["<asset-id>","<another-id>"], { wait: true });
const outputs = await client.jobs.outputs(job);

From an agent, the loop that works is: dry run to price it, run it, then analyze or a human in the console playground to check the result before it goes anywhere. An edit is the op where "it succeeded" and "it is right" are different questions.

Limits

The model re-renders the whole image, so faces, logos and text can drift even when you asked for something else; say what must stay the same, and use subjects when identity matters. It also writes at its own size, not your photo's: the default model is handed at most 1536 px on the long edge and answers at the imageSize in parameters (priced at 1K when you name none), so a 6000 px original comes back at the size asked for, not at 6000 px — ask for a larger one, or keep the original for print. A prompt is required — edit with no instruction is 400 invalid_param naming prompt. Content the model refuses comes back as a failed item with a code and a refund, not as a silent no-op.

FAQ

Do I need a mask?
No. You describe the region instead — "remove the cup on the left" — which is why this is inpainting by description. There is no mask parameter on the op.
Does it edit the original?
No. The result is always a new asset and the input is untouched — no job overwrites the asset it read, so the input id stays valid for the next step.
Can I edit several images with the same instruction?
Yes — pass every id in assetIds and the prompt applies to each. They are items of one job, settled one by one; one that fails is not charged.
When should I use a deterministic op instead?
Whenever the change has a number in it. Crop, resize, rotate, a colour adjustment, a watermark — those are deterministic ops: exact, repeatable, free on a paid plan, and every one but the watermark (overlay, which reads a stored layer) also runs synchronously, bytes in and bytes out.
Why did the whole image change when I asked for one small thing?
An instruction edit re-renders the picture. If identity has to survive, say what must not change in the prompt, or keep the subject fixed with preset subjects and reference images.