← Blog · Use Cases · Docs

Decide, then generate: pairing /v1/score with a 4B model for fast agent workflows

September 2026 · SearchAI Inference Server team

new in v1.4.0POST /v1/score ~355 ms decision · q35-4b~84 ms exact-repeat local · OpenAI-compatible

A lot of agent work is really two jobs wearing one coat. First a decision — which category is this, which tool do I call, is this safe to answer — and then a generation — the drafted reply, the extracted fields, the summary. The usual reflex is to make one model do both in a single generation: ask it to "think about the category and then write the answer." That's slower than it needs to be, and it's not deterministic where you most want it to be.

As of v1.4.0, there's a cleaner split. Use POST /v1/score for the fast, deterministic decision in one prefill with no decode. Then hand the path you actually chose to a 4B chat model — Spark-X2.5 4B or Qwen 3.5 4B — over POST /v1/chat/completions for the generation. Both endpoints run on the same local server, behind the same OpenAI-compatible API, and no data leaves the box.

Why split the step

Scoring is cheap and deterministic. It's a single prefill pass over the prompt and the candidate choices — the server reads the input and reports the relative probability of each choice. There is no token-by-token decode, so there's no generation to wait for and nothing to sample: the same input produces the same decision, every time.

That gives you two wins at once:

  • Reliable routing without a full generation. You get a decision you can branch on directly — no parsing free-text, no coaxing JSON out of a chat model, no risk of the model narrating instead of answering.
  • You only pay for generation on the path you chose. The expensive part — decoding a real reply — runs once, on the one branch that matters, instead of every request carrying the cost of the model reasoning its way to a category first.

It's a good fit anywhere a decision gates work: support triage, RAG answer-vs-escalate, guardrails before you generate, and agent tool selection.

The score endpoint

/v1/score is supported on the Qwen 3.5 and Spark-X2.5 families. The friendly request form is small:

{
  "model": "q35-4b",
  "query": "Customer message text here",
  "choices": ["billing", "shipping", "technical support", "none of the above"],
  "instruction": "Route this support message to the correct queue."
}

and the response gives you the pick plus the full distribution:

{
  "decision": "billing",
  "probabilities": [0.71, 0.09, 0.14, 0.06],
  "scores": [ ... ]
}

It's deterministic, and it's fast: about 355 ms on q35-4b and 394 ms on the Spark family for a fresh call, dropping to roughly 84 ms on exact-repeats as the prefix cache reuses the shared prompt. The instruction field is optional; query and choices are the core.

A worked example: support triage

Here's the two-step pattern end to end. Step 1 scores the incoming message against a fixed set of queues and returns a decision:

curl -s http://127.0.0.1:8081/v1/score \
  -H "Authorization: Bearer $SEARCHAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "q35-4b",
    "query": "I was charged twice for my subscription this month.",
    "choices": ["billing", "shipping", "technical support", "none of the above"],
    "instruction": "Route this support message to the correct queue."
  }'
# => {"decision":"billing","probabilities":[0.71,0.09,0.14,0.06], ...}

Step 2 takes that decision and calls /v1/chat/completions on a 4B chat model to draft the actual reply — here grounded in the billing policy, because the decision was "billing":

curl -s http://127.0.0.1:8081/v1/chat/completions \
  -H "Authorization: Bearer $SEARCHAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spark-x25-4b",
    "messages": [
      {"role":"system","content":"You are a billing support agent. Ground every reply in the billing policy below.\n"},
      {"role":"user","content":"I was charged twice for my subscription this month."}
    ]
  }'

The glue between the two is tiny — read the decision, branch to the right generator prompt:

DECISION=$(curl -s http://127.0.0.1:8081/v1/score \
  -H "Authorization: Bearer $SEARCHAI_API_KEY" -H "Content-Type: application/json" \
  -d "{\"model\":\"q35-4b\",\"query\":\"$MSG\",
       \"choices\":[\"billing\",\"shipping\",\"technical support\",\"none of the above\"],
       \"instruction\":\"Route this support message to the correct queue.\"}" \
  | jq -r .decision)

case "$DECISION" in
  billing)            SYS="$BILLING_POLICY" ;;
  shipping)           SYS="$SHIPPING_POLICY" ;;
  "technical support") SYS="$TECH_POLICY" ;;
  *)                  SYS="$GENERAL_POLICY" ;;
esac

curl -s http://127.0.0.1:8081/v1/chat/completions \
  -H "Authorization: Bearer $SEARCHAI_API_KEY" -H "Content-Type: application/json" \
  -d "{\"model\":\"spark-x25-4b\",
       \"messages\":[{\"role\":\"system\",\"content\":$(jq -Rs . <<<"$SYS")},
                     {\"role\":\"user\",\"content\":$(jq -Rs . <<<"$MSG")}]}"

Swap spark-x25-4b for q35-4b on the generation step if you prefer Qwen 3.5 — same request shape, same server. Spark models also carry a reasoning_content (thinking) channel you can enable for the harder replies.

Where the prefix cache earns its keep

Triage prompts are highly repetitive: the same instruction and the same list of choices, message after message. The score endpoint's prefix cache reuses that shared prefix, which is why exact-repeats land near 84 ms instead of the ~355 ms of a cold call. In a routing loop where only the customer message changes, most of the decision cost is already paid — you're mostly scoring the delta.

More places the pattern fits

  • RAG answer-vs-escalate. Score whether the retrieved context is sufficient to answer before you spend a generation on it — answer on one branch, escalate to a human or a bigger model on the other.
  • Guardrails before generate. A deterministic gate — safe / needs-review / refuse — in front of the chat model, with no decode cost for the requests that get blocked.
  • Agent tool selection. Score the next action against the set of available tools, then generate only the arguments for the tool you picked.

In every case the shape is the same: a cheap deterministic decision, then generation only on the path you committed to.

Getting it

Install the server, fetch a 4B model, and both endpoints are live on the same process:

curl -fsSL https://inference-server.searchblox.com/install | sudo bash
sudo ./fetch-models.sh 4b          # fetches q35-4b
# Spark-X2.5 4B loads from the console Models tab
# then call POST /v1/score for the decision
# and POST /v1/chat/completions for the generation

Timings in this post: /v1/score latency measured with a single 4B model loaded — ~355 ms on q35-4b, ~394 ms on the Spark family, ~84 ms on exact-repeats via prefix-cache reuse. The score endpoint is supported on the Qwen 3.5 and Spark-X2.5 families; both are served locally over the OpenAI-compatible API, and no request data leaves the box.