Search that sees: multimodal embeddings and reranking on your own box
August 2026 · SearchAI Inference Server team
/v1/embeddings/v1/rerank text + images, one vector space OpenAI & Jina compatibleCPU-only
Until now the server generated answers; finding the right context to
answer from was your problem. This release closes the loop: two
new retrieval models — qwen3vl-embed-2b and
qwen3vl-rerank-2b — serve
/v1/embeddings (OpenAI-compatible) and
/v1/rerank (Jina/Cohere-compatible) from the
same binary, the same port, and the same API key as chat, vision, and
speech. Embed your documents and your images into one vector
space, recall candidates from your vector store, rerank them for
precision, then generate — a complete private RAG pipeline with nothing
leaving your network and no per-token meter running.
Embeddings: /v1/embeddings
The request shape is the OpenAI one, so every existing SDK and vector
store integration works unchanged. Strings embed as text; objects with an
image embed pictures — into the same 2048-dimension
space, so a text query can find an image and vice versa:
curl http://<host>:8081/v1/embeddings \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{"model": "qwen3vl-embed-2b",
"input": ["How do I rotate an API key?",
"Rotate keys under Settings > Security.",
{"image": "https://example.com/settings-screenshot.png"}],
"dimensions": 512}'
Three things to notice:
- Matryoshka dimensions. The model was trained so its
vectors truncate gracefully: ask for
dimensions: 512(or 256, or 64) and you get smaller, re-normalized vectors with most of the retrieval quality — a 4× smaller index without re-embedding anything. - Images by URL or inline. Images can be base64 data
URLs or public http(s) URLs the server fetches itself (10 MB cap,
private and internal addresses refused, and outbound fetch can be
disabled entirely with
SAI_FETCH_MEDIA=0). PNG, JPEG, WebP, GIF, BMP, and TIFF all decode. - Task instructions. An optional
instructionfield prefixes the task description the model embeds under — tuning it per use case is worth 1–5% retrieval quality.
Captured from a fresh install: three passages embedded at 128 dimensions — the two related sentences land at cosine 0.80 while the unrelated one sits at 0.34; a photo and its matching caption land at 0.62 against 0.07 for an off-topic caption. Text-side accuracy is validated against the reference implementation to cosine 0.9994.
Reranking: /v1/rerank
Vector recall is fast but fuzzy — the reranker is the precision stage. It reads the query against each candidate together and scores relevance 0–1, in the same request shape Jina and Cohere clients already use:
curl http://<host>:8081/v1/rerank \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{"model": "qwen3vl-rerank-2b",
"query": "What is the capital of France?",
"documents": ["Paris is the capital and most populous city of France.",
"The mitochondria is the powerhouse of the cell.",
{"image": "data:image/jpeg;base64,..."}],
"top_n": 5}'
Results come back sorted. From the same fresh install: the relevant passage scores 0.72, the irrelevant one 0.14 — a wide, decision-ready margin. And because the reranker shares the embedder's vision tower, documents can be images: a query like "diagram of the payment flow" scores screenshots and figures directly.
The classic two-stage pattern, all on one box:
- Recall — embed the query, pull the top 50–100 candidates from your vector store by cosine similarity.
- Rerank — score those candidates with
/v1/rerank, keep the top 3–5. - Generate — hand the winners to
/v1/chat/completionsas context.
From Python, it's the client you already have
from openai import OpenAI
client = OpenAI(base_url="http://<host>:8081/v1", api_key=KEY)
resp = client.embeddings.create(
model="qwen3vl-embed-2b",
input=["passage one", "passage two"],
dimensions=512,
)
vectors = [d.embedding for d in resp.data] # L2-normalized — dot = cosine
Rerank is one requests.post to /v1/rerank —
or point any Jina-compatible client at the box. The
console now has a
Retrieval tab to try both interactively (live similarity
matrix, ranked score bars) and an API reference tab that
generates ready-to-run cURL, Python, and JavaScript for every endpoint.
Why this matters for private RAG
Most self-hosted stacks bolt together three vendors: an embedding API (per-token), a reranker API (per-call), and a generation API (per-token) — each one a data-egress decision someone has to sign off on. Here the whole pipeline is one 22 MB binary on a CPU host you already run:
Embedding is a prefill-only workload — exactly the shape CPUs batch well — so index builds run wide across cores while the same box keeps serving chat. And for image-heavy corpora (scanned invoices, product photos, slides, screenshots) the shared vector space means you index the pictures themselves, not just captions someone remembered to write.
Get it: one line installs the server with chat plus both retrieval models:
curl -fsSL https://inference-server.searchblox.com/install | MODELS='4b embed rerank' sudo bash
Numbers above were captured August 2026 from a fresh
one-line install on an AWS Graviton (c9g.8xlarge) node:
qwen3vl-embed-2b and qwen3vl-rerank-2b
(Q8_0 — chosen after Q4 measurably hurt embedding fidelity), served by the
public binaries. Both models also run on the
Docker image and macOS.
Reference-implementation parity and endpoint details are in the
console's API tab on any running server.