← Blog · Download & Install · Models

Point your OpenAI code at your own box: a drop-in migration

August 2026 · SearchAI Inference Server team

OpenAI-compatibletwo-line change SDK · streaming · tools · visionreal captured output

The reason to move off a per-token API is easy to state — fixed cost, full privacy, no vendor quota. The reason people don't is fear of a rewrite. Here's the good news: there's no rewrite. The server speaks the OpenAI API, so migrating is a base URL and key change. Your existing OpenAI SDK, your streaming code, your tool definitions, your vision calls, your JSON-mode prompts — all keep working. Everything below is captured from the unmodified OpenAI Python SDK talking to a local box.

The entire change (Python)

Two arguments to the client constructor. Nothing else in your code moves:

from openai import OpenAI

# before — hosted, metered
# client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

# after — your own box, fixed cost
client = OpenAI(
    base_url="http://<host>:8081/v1",     # your server
    api_key="<your-server-key>",           # the key the installer generated
)

resp = client.chat.completions.create(          # unchanged
    model="q35-4b",                             # a local model id
    messages=[{"role":"user","content":"..."}],
)

Run against a real server, the standard calls return exactly what your code already expects — this is the actual output:

>>> [m.id for m in client.models.list().data]
['q35-4b', 'qwen-talker-1.7b', 'gemma-4-12B-it', 'q35-0.8b',
 'gemma-4-E2B-it', 'gemma-4-E4B-it', ...]

>>> r = client.chat.completions.create(model="q35-4b",
...     messages=[{"role":"user","content":"In one sentence, what is RAG?"}])
>>> r.choices[0].message.content
'Retrieval-augmented generation (RAG) is a hybrid approach that combines
 the ability of large language models to generate coherent text with an
 external retrieval mechanism ... thereby enhancing accuracy and reducing
 hallucinations.'
>>> r.usage.prompt_tokens, r.usage.completion_tokens
(24, 55)

Same object shape, same choices[0].message.content, same usage counts. Your parsing code doesn't change.

Zero-code path: environment variables

If your app already reads the standard OpenAI env vars, you don't touch the code at all — just set two variables in the environment and restart:

export OPENAI_BASE_URL=http://<host>:8081/v1
export OPENAI_API_KEY=<your-server-key>

The OpenAI SDKs pick these up automatically. Point a staging deployment at your box first, run your existing test suite, and flip production when it's green.

Streaming keeps working

Server-sent events are identical — same stream=True, same delta shape. Captured live:

stream = client.chat.completions.create(
    model="q35-4b", stream=True,
    messages=[{"role":"user","content":"Name three benefits of local inference."}])
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

# → streamed 65 chunks:
# 1. **Enhanced Privacy**: Data stays on the device...
# 2. **Reduced Latency**: Responses are generated locally without network delays...
# 3. **Improved Reliability**: The system functions independently...

Your token-by-token UI code works unchanged.

Node, curl, and frameworks

Node / TypeScript — same two fields:

import OpenAI from "openai";
const client = new OpenAI({
  baseURL: "http://<host>:8081/v1",
  apiKey: "<your-server-key>",
});
const r = await client.chat.completions.create({
  model: "q35-4b",
  messages: [{ role: "user", content: "..." }],
});

curl — just the host and key:

curl http://<host>:8081/v1/chat/completions \
  -H "Authorization: Bearer <your-server-key>" -H "Content-Type: application/json" \
  -d '{"model":"q35-4b","messages":[{"role":"user","content":"hello"}]}'

Frameworks — anything that accepts an OpenAI-compatible base URL (agent frameworks, LangChain, LlamaIndex, chat UIs, MCP-based clients) points at the server the same way: set the base URL and key, pick a local model id. If it talks to OpenAI today, it talks to your box tomorrow.

What carries over — and the two things to change

The surfaces you actually use come across as-is:

  • Chat completions — messages, system prompts, temperature, max_tokens, streaming, usage.
  • Tools / function calling — same tools array and tool_calls response (see the agent post).
  • Vision — image_url content parts (see document extraction).
  • JSON mode — response_format: {"type":"json_object"}.
  • Audio — /v1/audio/transcriptions and /v1/audio/speech (see the speech post).

Two things you do change, both trivial:

  1. Model names. Swap gpt-… for a local id like q35-4b, q35-2b, or qwen38-27b — usually one constant in your config. See the catalog for the full list.
  2. Your prompts may want a quick pass. A prompt tuned for one model is worth a re-check on another; it takes minutes and it's free to iterate locally — the prompt-optimization post covers the workflow.

The takeaway: migrating off a per-token API is a config change, not a project. Point the base URL at your own box, swap the model name, keep the rest of your code — and the meter stops running while your data stays on your network.

Install a server (Linux, one line) and grab the generated key:

curl -fsSL https://inference-server.searchblox.com/install | sudo bash

Then set base_url to http://<host>:8081/v1 and run your existing OpenAI code against it.

All SDK output above was captured from the unmodified openai Python SDK (v1.47) pointed at a running SearchAI Inference Server, August 2026. See the model catalog for model ids and the Getting Started guide for install and sizing.