Scale from one box to a fleet: model-aware, load-aware, cache-aware routing — no load balancer
August 2026 · SearchAI Inference Server team
clusteringno external LB smart routingmeasured scaling exact commands
Single-stream decode on a CPU is memory-bandwidth bound: past a point, more cores don't make one answer come out faster. So when you outgrow one box, you don't buy a bigger one — you add nodes. The usual cost of that is real infrastructure: a load balancer, a service registry, a routing tier, health checks. This server has none of it. Every node is identical, discovers its peers by gossip, and routes requests itself. You add a box with the same one-line install and one shared secret — that's the whole scaling story.
Adding a node is the install command, twice
Start a seed node, then point each new node at it. Same binary, same model, one shared cluster secret:
# seed node (10.0.0.10)
sudo API_KEY='public-api-key' MODELS='4b' \
CLUSTER_ADVERTISE='10.0.0.10:8081' \
CLUSTER_SECRET='shared-cluster-secret' \
./install-searchai.sh
# every joining node — unique advertise address, point at the seed
sudo API_KEY='public-api-key' MODELS='4b' \
CLUSTER_ADVERTISE='10.0.0.11:8081' \
CLUSTER_SEED='10.0.0.10:8081' \
CLUSTER_SECRET='shared-cluster-secret' \
./install-searchai.sh
That's it. Nodes gossip every 2 seconds, exchanging the full peer list and each node's loaded models, and probe each other directly. There's no coordinator to run and nothing to register — the seed is just the first address a node happens to learn the fleet from.
No external load balancer required. Any node accepts any request and routes it internally, so clients can hit the fleet through a plain DNS round-robin or a single VIP — or just send everything to one node and let it distribute. There is no dedicated routing tier to provision, scale, or pay for.
The routing brain: three criteria, in order
When a request lands on any node, that node decides where it should run by applying three rules in sequence:
- Model-aware. Gossip carries which models each node
has resident, so a request for
q35-4broutes to a node that actually has the 4B loaded — local first, then round-robin among peers that have it, then any healthy peer. This is what lets nodes specialize: put the 4B on two nodes and the 27B on a third, and every client still sends every request to one address. - Load-aware. Among the nodes that have the right model, the least-busy one wins. Bursts spread across the fleet instead of piling onto whoever the client happened to contact.
- Cache-aware. Multi-turn conversations stick to the node holding their prefix cache, so follow-up turns skip re-processing the conversation history — the repeat-prefill savings survive across the cluster, not just on a single box. (This applies to the Gemma models as well as Qwen.)
The order matters: correctness first (a node that can serve the model), then balance (spread the load), then efficiency (reuse the cache). It's a routing policy that would normally live in a separate service — here it's compiled into every node.
Specialize the fleet — different models, one endpoint
Because routing is model-aware, a cluster doesn't have to be uniform. A common shape:
- Two nodes on the 4B for the high-volume chat and RAG traffic,
- one node on the 27B for the occasional deep-reasoning request,
- clients send everything to the same address.
A q35-4b request lands on one of the 4B nodes; a
qwen38-27b request routes to the 27B node — automatically,
because the fleet knows who has what. Mixed CPU types are fine too: route
by capacity and make the fastest box the seed. You grow and reshape the
fleet by changing what each node loads, not by touching a routing config.
What it measures
Two things scale as you add nodes, and it helps to be precise about which. Per-user speed stays full — each node serves its requests at the same decode rate a single box would, because decode is memory-bound per node and nodes are independent. Aggregate capacity grows — more nodes means more simultaneous users served at that full speed, with each node's memory footprint bounded on its own.
Measured on a 3-node 4B cluster (distinct prompts, balanced load):
Scaling improves with concurrency as the balancer spreads requests more evenly, and each node stays memory-bounded independently (RSS held 16 → 31 GB from concurrency 16 → 64, no OOM) — so capacity grows with node count while every node stays within its own envelope.
See it and verify it
The console's Admin tab, on any node, shows the whole fleet — every node, its CPU load and RAM, and its resident models. From the shell, two checks confirm membership and routing:
# membership + which models each node has (uses the cluster secret, not the Bearer key)
curl -sS -H "X-SAI-Cluster: shared-cluster-secret" \
http://10.0.0.10:8081/cluster/state
# load-balanced chat — send to ANY node; it routes for you
curl -sS http://10.0.0.10:8081/v1/chat/completions \
-H "Authorization: Bearer public-api-key" -H 'Content-Type: application/json' \
-d '{"model":"q35-4b","messages":[{"role":"user","content":"hi"}]}'
/cluster/state lists every peer and shows the model loaded
after one gossip interval. Set SAI_ROUTE_LOG=1 and each node's
journal prints where it sent each request, so you can watch distinct
prompts fan out across the fleet under load.
Operational notes
- Network: keep the fleet on a private network — inter-node traffic is plain HTTP. Terminate TLS at your entry load balancer or VIP if clients are external.
- Firewall: allow TCP 8081 between every pair of
nodes (gossip distributes the full peer list and nodes probe each
other directly), plus from your entry point to the nodes. Block
/cluster/*from the public internet — it's the internal control plane, gated by the cluster secret. - Two secrets, two jobs: the
API_KEYis the Bearer key clients use; theCLUSTER_SECRETauthenticates node-to-node traffic. Keep them different. - Seed choice: any node can be the seed; make the fastest node the entry/seed when the fleet is heterogeneous.
The takeaway: scaling here is horizontal and boring in the best way — add a node with the same install and one secret, and the built-in model-aware → load-aware → cache-aware routing handles the rest. No load balancer to run, no coordinator to babysit, no routing tier to pay for. Start on one box; grow to a fleet when you need to, without changing how anything is deployed.
Install the first node (Linux, one line):
curl -fsSL https://inference-server.searchblox.com/install | sudo bash
Cluster mechanics and the measured 3-node numbers are from the SearchAI Inference Server deployment guide, August 2026. See the full Deployment Guide for the cluster section, Performance for the vertical-vs-horizontal scaling detail, and the model catalog for what each node can load.