Ten things to decide on-box: fast local classification with /v1/score
September 2026 · SearchAI Inference Server team
new in v1.4.0one prefill · no decode deterministic~355 ms · q35-4b on-box · nothing leaves the host
A lot of what an LLM does in production isn't writing — it's deciding. Which queue does this ticket go to? Is this message spam? Does this text contain PII? For those jobs you don't want a paragraph back; you want one of a fixed set of answers, fast, and the same answer every time.
New in v1.4.0, /v1/score — "Jev-style" fixed-answer scoring — does exactly that. The server prefills the prompt once and reads a restricted softmax over just the allowed answers. There is no decode loop: no tokens are generated one at a time. The result is a decision that is far faster than a chat call and fully deterministic — and it happens entirely on-box, so no data leaves the host. It's supported on the Qwen 3.5 and Spark X2.5 model families.
One call, one decision
The friendly form is the one you'll reach for most: give it a
query and a list of choices, get back a
decision plus the probability mass over each choice.
curl -s http://127.0.0.1:8081/v1/score \
-H "Authorization: Bearer $SEARCHAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "q35-4b",
"instruction": "Route this support message to the right queue.",
"query": "My card was charged twice for the same order.",
"choices": ["billing", "shipping", "technical", "none of the above"]
}'
# response
{ "decision": "billing",
"probabilities": { "billing": 0.94, "shipping": 0.01,
"technical": 0.02, "none of the above": 0.03 },
"scores": [ ... ] }
You supply 2–26 choices. Because the probability mass is
forced onto the listed choices, always add an explicit escape
hatch like "none of the above" — otherwise a query that
matches nothing will still be pushed onto whichever choice is least
wrong.
There's also a raw form at parity with SGLang's
/v1/score: pass label_token_ids (and an
optional apply_softmax) and get nested
scores [[...]] back. Use it when you're
scoring specific token ids directly rather than natural-language
choices.
Ten things to decide on-box
1. Intent routing. Drop an inbound message into the right
support queue before a human ever sees it. One score call replaces a
brittle keyword rule set and returns the queue plus a confidence you
can threshold on. choices:
["billing","shipping","technical","account","none of the above"]
2. Ticket triage and priority. Grade an incoming ticket by
urgency so the queue orders itself. "Production is down for all users"
and "typo on the pricing page" land in different buckets in ~355 ms,
every time. choices:
["p1 urgent","p2 high","p3 normal","p4 low"]
3. Guardrail yes/no. A binary allow/block gate in front of a
more expensive pipeline. Because it's deterministic, the same input
always gets the same verdict — which is what you want from a
guardrail. choices: ["allow","block"]
4. PII / sensitivity flagging. Flag whether a chunk of text
carries personal or otherwise sensitive data before it's logged,
indexed, or forwarded. Since the check runs on-box, the sensitive text
never has to leave the machine to find out it's sensitive.
choices:
["contains pii","no pii","unsure"]
5. Sentiment. Classify a review, a survey reply, or a chat
turn into a fixed sentiment scale. At a few hundred milliseconds a
decision, you can score a firehose of feedback in near real time.
choices:
["positive","neutral","negative","mixed"]
6. Language / locale routing. Decide which language a message
is in so it goes to the right localized model, template, or agent.
choices:
["english","spanish","french","german","other"]
7. Spam / abuse detection. Score whether a submission is spam
or abusive at the edge of your intake, deterministically, so the same
payload can't slip through by luck on a retry. choices:
["spam","abuse","legitimate"]
8. Approval gates. Decide whether an action can auto-approve
or must escalate to a human. A refund under a threshold auto-approves;
anything ambiguous escalates — and the decision is repeatable, which
matters when you have to explain it later. choices:
["auto-approve","escalate to human","reject"]
9. Tool / skill selection for an agent. Instead of asking a
model to write which tool to call, ask it to pick one from
the tools it actually has. One prefill, no decode, and the answer is
constrained to real tool names. choices:
["search_docs","run_sql","send_email","none of the above"]
10. Document / type classification. Sort an incoming document
into a type so the right extractor or workflow picks it up. Feed the
first page or a summary as the query and read the label.
choices:
["invoice","contract","resume","receipt","other"]
Keep the stable part first
The decision is fast on its own — q35-4b ~355 ms (about 2.6× faster than generating the answer) and spark ~394 ms (about 1.6×). But an exact repeat over the same choice set lands in about 84 ms, because the prefix cache reuses the already-prefilled prompt.
To get those hits, order your prompt stable part first: put
the instruction and the label set up front, and vary only
the query. When the leading tokens don't change from call
to call, the cache does the heavy lifting and you pay only for the new
query.
Why it's fast
A chat completion runs a full decode loop: prefill the prompt, then generate the answer one token at a time, each step a forward pass. /v1/score skips all of that. It does a single prefill and reads a restricted softmax over the allowed answers — no tokens are generated, so there's no per-token cost and no loop to wait through. Add the prefix cache for repeated choice sets and the marginal cost of a decision drops to the new query alone. The same mechanics make it deterministic: no sampling, same input in, same decision out.
Getting started
Install the server, fetch a Qwen 3.5 or Spark X2.5 model, and point
/v1/score at it with the "model" field:
curl -fsSL https://inference-server.searchblox.com/install | sudo bash
# fetch q35-4b (or a spark-x25 model) from the console catalog, then:
curl -s http://127.0.0.1:8081/v1/score \
-H "Authorization: Bearer $SEARCHAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "q35-4b",
"query": "My card was charged twice.",
"choices": ["billing","shipping","technical","none of the above"] }'
Numbers in this post: decision latencies measured on
CPU — q35-4b ~355 ms (2.6× vs. generating the answer), spark ~394 ms
(1.6×), ~84 ms on an exact-repeat via the prefix cache. /v1/score is
new in v1.4.0 and supported on the Qwen 3.5 and Spark X2.5 families;
2–26 choices in the friendly form, with a raw label_token_ids
form at parity with SGLang's /v1/score.