Point your OpenAI code at your own box: a drop-in migration
August 2026 · SearchAI Inference Server team
OpenAI-compatibletwo-line change SDK · streaming · tools · visionreal captured output
The reason to move off a per-token API is easy to state — fixed cost, full privacy, no vendor quota. The reason people don't is fear of a rewrite. Here's the good news: there's no rewrite. The server speaks the OpenAI API, so migrating is a base URL and key change. Your existing OpenAI SDK, your streaming code, your tool definitions, your vision calls, your JSON-mode prompts — all keep working. Everything below is captured from the unmodified OpenAI Python SDK talking to a local box.
The entire change (Python)
Two arguments to the client constructor. Nothing else in your code moves:
from openai import OpenAI
# before — hosted, metered
# client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
# after — your own box, fixed cost
client = OpenAI(
base_url="http://<host>:8081/v1", # your server
api_key="<your-server-key>", # the key the installer generated
)
resp = client.chat.completions.create( # unchanged
model="q35-4b", # a local model id
messages=[{"role":"user","content":"..."}],
)
Run against a real server, the standard calls return exactly what your code already expects — this is the actual output:
>>> [m.id for m in client.models.list().data]
['q35-4b', 'qwen-talker-1.7b', 'gemma-4-12B-it', 'q35-0.8b',
'gemma-4-E2B-it', 'gemma-4-E4B-it', ...]
>>> r = client.chat.completions.create(model="q35-4b",
... messages=[{"role":"user","content":"In one sentence, what is RAG?"}])
>>> r.choices[0].message.content
'Retrieval-augmented generation (RAG) is a hybrid approach that combines
the ability of large language models to generate coherent text with an
external retrieval mechanism ... thereby enhancing accuracy and reducing
hallucinations.'
>>> r.usage.prompt_tokens, r.usage.completion_tokens
(24, 55)
Same object shape, same choices[0].message.content, same
usage counts. Your parsing code doesn't change.
Zero-code path: environment variables
If your app already reads the standard OpenAI env vars, you don't touch the code at all — just set two variables in the environment and restart:
export OPENAI_BASE_URL=http://<host>:8081/v1
export OPENAI_API_KEY=<your-server-key>
The OpenAI SDKs pick these up automatically. Point a staging deployment at your box first, run your existing test suite, and flip production when it's green.
Streaming keeps working
Server-sent events are identical — same stream=True, same
delta shape. Captured live:
stream = client.chat.completions.create(
model="q35-4b", stream=True,
messages=[{"role":"user","content":"Name three benefits of local inference."}])
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
# → streamed 65 chunks:
# 1. **Enhanced Privacy**: Data stays on the device...
# 2. **Reduced Latency**: Responses are generated locally without network delays...
# 3. **Improved Reliability**: The system functions independently...
Your token-by-token UI code works unchanged.
Node, curl, and frameworks
Node / TypeScript — same two fields:
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "http://<host>:8081/v1",
apiKey: "<your-server-key>",
});
const r = await client.chat.completions.create({
model: "q35-4b",
messages: [{ role: "user", content: "..." }],
});
curl — just the host and key:
curl http://<host>:8081/v1/chat/completions \
-H "Authorization: Bearer <your-server-key>" -H "Content-Type: application/json" \
-d '{"model":"q35-4b","messages":[{"role":"user","content":"hello"}]}'
Frameworks — anything that accepts an OpenAI-compatible base URL (agent frameworks, LangChain, LlamaIndex, chat UIs, MCP-based clients) points at the server the same way: set the base URL and key, pick a local model id. If it talks to OpenAI today, it talks to your box tomorrow.
What carries over — and the two things to change
The surfaces you actually use come across as-is:
- Chat completions — messages, system prompts,
temperature, max_tokens, streaming,
usage. - Tools / function calling — same
toolsarray andtool_callsresponse (see the agent post). - Vision —
image_urlcontent parts (see document extraction). - JSON mode —
response_format: {"type":"json_object"}. - Audio —
/v1/audio/transcriptionsand/v1/audio/speech(see the speech post).
Two things you do change, both trivial:
- Model names. Swap
gpt-…for a local id likeq35-4b,q35-2b, orqwen38-27b— usually one constant in your config. See the catalog for the full list. - Your prompts may want a quick pass. A prompt tuned for one model is worth a re-check on another; it takes minutes and it's free to iterate locally — the prompt-optimization post covers the workflow.
The takeaway: migrating off a per-token API is a config change, not a project. Point the base URL at your own box, swap the model name, keep the rest of your code — and the meter stops running while your data stays on your network.
Install a server (Linux, one line) and grab the generated key:
curl -fsSL https://inference-server.searchblox.com/install | sudo bash
Then set base_url to http://<host>:8081/v1
and run your existing OpenAI code against it.
All SDK output above was captured from the unmodified
openai Python SDK (v1.47) pointed at a running SearchAI
Inference Server, August 2026. See the
model catalog for model ids and the
Getting Started guide for
install and sizing.