← Blog · Performance · Deployment Guide

Scale from one box to a fleet: model-aware, load-aware, cache-aware routing — no load balancer

August 2026 · SearchAI Inference Server team

clusteringno external LB smart routingmeasured scaling exact commands

Single-stream decode on a CPU is memory-bandwidth bound: past a point, more cores don't make one answer come out faster. So when you outgrow one box, you don't buy a bigger one — you add nodes. The usual cost of that is real infrastructure: a load balancer, a service registry, a routing tier, health checks. This server has none of it. Every node is identical, discovers its peers by gossip, and routes requests itself. You add a box with the same one-line install and one shared secret — that's the whole scaling story.

Adding a node is the install command, twice

Start a seed node, then point each new node at it. Same binary, same model, one shared cluster secret:

# seed node (10.0.0.10)
sudo API_KEY='public-api-key' MODELS='4b' \
     CLUSTER_ADVERTISE='10.0.0.10:8081' \
     CLUSTER_SECRET='shared-cluster-secret' \
     ./install-searchai.sh

# every joining node — unique advertise address, point at the seed
sudo API_KEY='public-api-key' MODELS='4b' \
     CLUSTER_ADVERTISE='10.0.0.11:8081' \
     CLUSTER_SEED='10.0.0.10:8081' \
     CLUSTER_SECRET='shared-cluster-secret' \
     ./install-searchai.sh

That's it. Nodes gossip every 2 seconds, exchanging the full peer list and each node's loaded models, and probe each other directly. There's no coordinator to run and nothing to register — the seed is just the first address a node happens to learn the fleet from.

No external load balancer required. Any node accepts any request and routes it internally, so clients can hit the fleet through a plain DNS round-robin or a single VIP — or just send everything to one node and let it distribute. There is no dedicated routing tier to provision, scale, or pay for.

The routing brain: three criteria, in order

When a request lands on any node, that node decides where it should run by applying three rules in sequence:

  1. Model-aware. Gossip carries which models each node has resident, so a request for q35-4b routes to a node that actually has the 4B loaded — local first, then round-robin among peers that have it, then any healthy peer. This is what lets nodes specialize: put the 4B on two nodes and the 27B on a third, and every client still sends every request to one address.
  2. Load-aware. Among the nodes that have the right model, the least-busy one wins. Bursts spread across the fleet instead of piling onto whoever the client happened to contact.
  3. Cache-aware. Multi-turn conversations stick to the node holding their prefix cache, so follow-up turns skip re-processing the conversation history — the repeat-prefill savings survive across the cluster, not just on a single box. (This applies to the Gemma models as well as Qwen.)

The order matters: correctness first (a node that can serve the model), then balance (spread the load), then efficiency (reuse the cache). It's a routing policy that would normally live in a separate service — here it's compiled into every node.

Specialize the fleet — different models, one endpoint

Because routing is model-aware, a cluster doesn't have to be uniform. A common shape:

  • Two nodes on the 4B for the high-volume chat and RAG traffic,
  • one node on the 27B for the occasional deep-reasoning request,
  • clients send everything to the same address.

A q35-4b request lands on one of the 4B nodes; a qwen38-27b request routes to the 27B node — automatically, because the fleet knows who has what. Mixed CPU types are fine too: route by capacity and make the fastest box the seed. You grow and reshape the fleet by changing what each node loads, not by touching a routing config.

What it measures

Two things scale as you add nodes, and it helps to be precise about which. Per-user speed stays full — each node serves its requests at the same decode rate a single box would, because decode is memory-bound per node and nodes are independent. Aggregate capacity grows — more nodes means more simultaneous users served at that full speed, with each node's memory footprint bounded on its own.

Measured on a 3-node 4B cluster (distinct prompts, balanced load):

Concurrency
Cluster tok/s
vs 1 node
Per-node RSS
16
43.8
1.49×
16 GB
32
43.2
1.52×
23 GB
64
47.7
1.65×
31 GB

Scaling improves with concurrency as the balancer spreads requests more evenly, and each node stays memory-bounded independently (RSS held 16 → 31 GB from concurrency 16 → 64, no OOM) — so capacity grows with node count while every node stays within its own envelope.

See it and verify it

The console's Admin tab, on any node, shows the whole fleet — every node, its CPU load and RAM, and its resident models. From the shell, two checks confirm membership and routing:

# membership + which models each node has (uses the cluster secret, not the Bearer key)
curl -sS -H "X-SAI-Cluster: shared-cluster-secret" \
  http://10.0.0.10:8081/cluster/state

# load-balanced chat — send to ANY node; it routes for you
curl -sS http://10.0.0.10:8081/v1/chat/completions \
  -H "Authorization: Bearer public-api-key" -H 'Content-Type: application/json' \
  -d '{"model":"q35-4b","messages":[{"role":"user","content":"hi"}]}'

/cluster/state lists every peer and shows the model loaded after one gossip interval. Set SAI_ROUTE_LOG=1 and each node's journal prints where it sent each request, so you can watch distinct prompts fan out across the fleet under load.

Operational notes

  • Network: keep the fleet on a private network — inter-node traffic is plain HTTP. Terminate TLS at your entry load balancer or VIP if clients are external.
  • Firewall: allow TCP 8081 between every pair of nodes (gossip distributes the full peer list and nodes probe each other directly), plus from your entry point to the nodes. Block /cluster/* from the public internet — it's the internal control plane, gated by the cluster secret.
  • Two secrets, two jobs: the API_KEY is the Bearer key clients use; the CLUSTER_SECRET authenticates node-to-node traffic. Keep them different.
  • Seed choice: any node can be the seed; make the fastest node the entry/seed when the fleet is heterogeneous.

The takeaway: scaling here is horizontal and boring in the best way — add a node with the same install and one secret, and the built-in model-aware → load-aware → cache-aware routing handles the rest. No load balancer to run, no coordinator to babysit, no routing tier to pay for. Start on one box; grow to a fleet when you need to, without changing how anything is deployed.

Install the first node (Linux, one line):

curl -fsSL https://inference-server.searchblox.com/install | sudo bash

Cluster mechanics and the measured 3-node numbers are from the SearchAI Inference Server deployment guide, August 2026. See the full Deployment Guide for the cluster section, Performance for the vertical-vs-horizontal scaling detail, and the model catalog for what each node can load.