← Blog · Download & Install · GPU acceleration

Edit images with plain English on your own box

August 2026 · SearchAI Inference Server team

/v1/images/edits20B instruction editor AWS g5.2xlarge · A10Gfits 24 GB VRAM private

"Make the background a bright sunny sky." "Add a red banner across the top." "Remove the person on the left." Instruction-based image editing — type what you want changed, get the edited image back — is one of the most requested capabilities in private AI, and one of the most sensitive: product shots, ID documents, medical imagery, and internal screenshots are exactly the pixels you don't want leaving your network.

The inference server ships this behind an OpenAI-compatible endpoint: /v1/images/edits, powered by a 20-billion-parameter Qwen-Image-Edit-class editor running through a stable-diffusion.cpp adapter, on your own GPU box. This post is the measured account: the exact AWS instance, the one-line install, real timings, and what "minutes-class editing" honestly means.

The setup — same $1.21/hour box as the GPU benchmark

Everything runs on the instance from our GPU acceleration benchmark: an AWS g5.2xlarge (NVIDIA A10G 24 GB, 8 vCPU, 32 GB RAM) on the Deep Learning Base GPU AMI (Ubuntu 24.04) — Ubuntu 24.04 or newer is required (glibc ≥ 2.36), 200 GB gp3 disk, ports 22/8081 open, sudo apt-get install -y libomp5 first. Then one line:

curl -fsSL https://inference-server.searchblox.com/install | \
  sudo API_KEY='your-key' BACKEND=cuda MODELS='4b image' bash

The image group adds the editor add-on (~430 MB adapter) and its three weight files (~27 GB): the Q4 diffusion transformer, a Qwen 2.5-VL text/vision encoder, and the VAE. GPU is strongly recommended for this one — a CPU-only edit takes tens of minutes.

Measured on the A10G

OperationMeasured
First edit after restart (loads + converts the full ~27 GB pipeline)~11 min, one-time
Each edit after that (1200×630 input)2 min 51 s
Each edit, larger photo (1280×899 input)4 min 23 s
VRAM during editing~3 GB resident

That VRAM row is the interesting one: the editor streams weights from host RAM per step, so the 20B pipeline fits comfortably next to your chat models on a 24 GB card — the same box was serving text at 119 tok/s between edits. Editing time scales with input resolution; budget three to four-and-a-half minutes per edit at typical sizes and plan batch jobs accordingly. This is a "queue it" capability, not a real-time one — the value is that the queue is yours, on your hardware, at flat cost.

Using it

The endpoint follows the OpenAI images API — multipart form, image + instruction in, base64 image out:

curl -s http://127.0.0.1:8081/v1/images/edits \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F image=@product-shot.png \
  -F prompt="Make the background a bright sunny sky" \
  | python3 -c "import json,base64,sys; \
      d=json.load(sys.stdin)['data'][0]['b64_json']; \
      open('edited.png','wb').write(base64.b64decode(d))"

Any OpenAI SDK's images.edit(...) call works the same way with base_url pointed at your box. Edits queue behind the same admission control as every other endpoint, so a long edit doesn't block your chat traffic.

What it looks like — real edits from this box

Both examples below are unretouched outputs from the curl call above, run on the g5.2xlarge during this benchmark. Left is the input, right is what came back.

Prompt: "Make the background a bright sunny sky" — input 1200×630, warm edit time 171 s:

Input: dark-themed product card Output: same card re-rendered on a bright sunny sky

Note what it preserved: the wordmark, every pill label, and the layout — while re-lighting the entire design for a sky background. This is instruction editing, not style transfer.

Prompt: "Turn the scene into a warm golden sunset" — input 1280×899 (photo), warm edit time 263 s (4 min 23 s):

Input: wind farm under a blue sky Output: the same wind farm at golden sunset

Honest scope

  • GPU required in practice. x86_64 Linux with an NVIDIA card (A100 / A10G / L4 / RTX 40-series / H100). CPU fallback exists but is tens of minutes per edit.
  • Minutes, not seconds. ~3 minutes per edit on an A10G at 1200×630. Smaller inputs are faster; this is for pipelines and workbenches, not live preview.
  • One-time warmup. The first edit after a restart pays the ~11-minute pipeline load; keep the service warm for batch runs.
  • Host RAM matters — more than VRAM. The weight-streaming design maps ~27–29 GB through host memory during the pipeline load. Measured across our boxes: 16 GB hosts cannot load the editor at all; 32 GB is workable but borderline at the load spike — add a swapfile (fallocate -l 16G /swapfile && mkswap /swapfile && swapon /swapfile) as cheap headroom insurance, or use a 64 GB host for comfort, especially with other models resident.

The bottom line

Instruction image editing is usually a per-image API charge on someone else's infrastructure, with your most sensitive pixels in the request body. Here it's a flag on the install line: a 20B editor behind the same private endpoint as your chat, vision, and speech models, ~3 minutes per edit on a $1.21/hour box, and nothing — not the source image, not the instruction, not the result — ever leaves your VPC.

curl -fsSL https://inference-server.searchblox.com/install | \
  sudo API_KEY='your-key' BACKEND=cuda MODELS='4b image' bash

Timings measured 2026-08-31 on g5.2xlarge as described above. GPU text/vision numbers for the same box: One flag, 4× the tokens.