← Blog · Models · Performance

Search that sees: multimodal embeddings and reranking on your own box

August 2026 · SearchAI Inference Server team

/v1/embeddings/v1/rerank text + images, one vector space OpenAI & Jina compatibleCPU-only

Until now the server generated answers; finding the right context to answer from was your problem. This release closes the loop: two new retrieval models — qwen3vl-embed-2b and qwen3vl-rerank-2b — serve /v1/embeddings (OpenAI-compatible) and /v1/rerank (Jina/Cohere-compatible) from the same binary, the same port, and the same API key as chat, vision, and speech. Embed your documents and your images into one vector space, recall candidates from your vector store, rerank them for precision, then generate — a complete private RAG pipeline with nothing leaving your network and no per-token meter running.

Embeddings: /v1/embeddings

The request shape is the OpenAI one, so every existing SDK and vector store integration works unchanged. Strings embed as text; objects with an image embed pictures — into the same 2048-dimension space, so a text query can find an image and vice versa:

curl http://<host>:8081/v1/embeddings \
  -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  -d '{"model": "qwen3vl-embed-2b",
       "input": ["How do I rotate an API key?",
                 "Rotate keys under Settings > Security.",
                 {"image": "https://example.com/settings-screenshot.png"}],
       "dimensions": 512}'

Three things to notice:

  1. Matryoshka dimensions. The model was trained so its vectors truncate gracefully: ask for dimensions: 512 (or 256, or 64) and you get smaller, re-normalized vectors with most of the retrieval quality — a 4× smaller index without re-embedding anything.
  2. Images by URL or inline. Images can be base64 data URLs or public http(s) URLs the server fetches itself (10 MB cap, private and internal addresses refused, and outbound fetch can be disabled entirely with SAI_FETCH_MEDIA=0). PNG, JPEG, WebP, GIF, BMP, and TIFF all decode.
  3. Task instructions. An optional instruction field prefixes the task description the model embeds under — tuning it per use case is worth 1–5% retrieval quality.

Captured from a fresh install: three passages embedded at 128 dimensions — the two related sentences land at cosine 0.80 while the unrelated one sits at 0.34; a photo and its matching caption land at 0.62 against 0.07 for an off-topic caption. Text-side accuracy is validated against the reference implementation to cosine 0.9994.

Reranking: /v1/rerank

Vector recall is fast but fuzzy — the reranker is the precision stage. It reads the query against each candidate together and scores relevance 0–1, in the same request shape Jina and Cohere clients already use:

curl http://<host>:8081/v1/rerank \
  -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  -d '{"model": "qwen3vl-rerank-2b",
       "query": "What is the capital of France?",
       "documents": ["Paris is the capital and most populous city of France.",
                     "The mitochondria is the powerhouse of the cell.",
                     {"image": "data:image/jpeg;base64,..."}],
       "top_n": 5}'

Results come back sorted. From the same fresh install: the relevant passage scores 0.72, the irrelevant one 0.14 — a wide, decision-ready margin. And because the reranker shares the embedder's vision tower, documents can be images: a query like "diagram of the payment flow" scores screenshots and figures directly.

The classic two-stage pattern, all on one box:

  1. Recall — embed the query, pull the top 50–100 candidates from your vector store by cosine similarity.
  2. Rerank — score those candidates with /v1/rerank, keep the top 3–5.
  3. Generate — hand the winners to /v1/chat/completions as context.

From Python, it's the client you already have

from openai import OpenAI
client = OpenAI(base_url="http://<host>:8081/v1", api_key=KEY)

resp = client.embeddings.create(
    model="qwen3vl-embed-2b",
    input=["passage one", "passage two"],
    dimensions=512,
)
vectors = [d.embedding for d in resp.data]   # L2-normalized — dot = cosine

Rerank is one requests.post to /v1/rerank — or point any Jina-compatible client at the box. The console now has a Retrieval tab to try both interactively (live similarity matrix, ranked score bars) and an API reference tab that generates ready-to-run cURL, Python, and JavaScript for every endpoint.

Why this matters for private RAG

Most self-hosted stacks bolt together three vendors: an embedding API (per-token), a reranker API (per-call), and a generation API (per-token) — each one a data-egress decision someone has to sign off on. Here the whole pipeline is one 22 MB binary on a CPU host you already run:

Stage
On this server
Embed
qwen3vl-embed-2b — text + images, 2048-dim, Matryoshka 64–2048, ~2.5 GB download, runs in 8 GB RAM
Recall
your vector store (any — the vectors are standard L2-normalized floats)
Rerank
qwen3vl-rerank-2b — text or image documents, sorted 0–1 scores
Generate
any chat model in the catalog — with vision, audio, and tools

Embedding is a prefill-only workload — exactly the shape CPUs batch well — so index builds run wide across cores while the same box keeps serving chat. And for image-heavy corpora (scanned invoices, product photos, slides, screenshots) the shared vector space means you index the pictures themselves, not just captions someone remembered to write.

Get it: one line installs the server with chat plus both retrieval models:

curl -fsSL https://inference-server.searchblox.com/install | MODELS='4b embed rerank' sudo bash

Numbers above were captured August 2026 from a fresh one-line install on an AWS Graviton (c9g.8xlarge) node: qwen3vl-embed-2b and qwen3vl-rerank-2b (Q8_0 — chosen after Q4 measurably hurt embedding fidelity), served by the public binaries. Both models also run on the Docker image and macOS. Reference-implementation parity and endpoint details are in the console's API tab on any running server.