← Blog · Download & Install · Docs

Ten things to decide on-box: fast local classification with /v1/score

September 2026 · SearchAI Inference Server team

new in v1.4.0one prefill · no decode deterministic~355 ms · q35-4b on-box · nothing leaves the host

A lot of what an LLM does in production isn't writing — it's deciding. Which queue does this ticket go to? Is this message spam? Does this text contain PII? For those jobs you don't want a paragraph back; you want one of a fixed set of answers, fast, and the same answer every time.

New in v1.4.0, /v1/score — "Jev-style" fixed-answer scoring — does exactly that. The server prefills the prompt once and reads a restricted softmax over just the allowed answers. There is no decode loop: no tokens are generated one at a time. The result is a decision that is far faster than a chat call and fully deterministic — and it happens entirely on-box, so no data leaves the host. It's supported on the Qwen 3.5 and Spark X2.5 model families.

One call, one decision

The friendly form is the one you'll reach for most: give it a query and a list of choices, get back a decision plus the probability mass over each choice.

curl -s http://127.0.0.1:8081/v1/score \
  -H "Authorization: Bearer $SEARCHAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "q35-4b",
    "instruction": "Route this support message to the right queue.",
    "query": "My card was charged twice for the same order.",
    "choices": ["billing", "shipping", "technical", "none of the above"]
  }'

# response
{ "decision": "billing",
  "probabilities": { "billing": 0.94, "shipping": 0.01,
                     "technical": 0.02, "none of the above": 0.03 },
  "scores": [ ... ] }

You supply 2–26 choices. Because the probability mass is forced onto the listed choices, always add an explicit escape hatch like "none of the above" — otherwise a query that matches nothing will still be pushed onto whichever choice is least wrong.

There's also a raw form at parity with SGLang's /v1/score: pass label_token_ids (and an optional apply_softmax) and get nested scores [[...]] back. Use it when you're scoring specific token ids directly rather than natural-language choices.

Ten things to decide on-box

1. Intent routing. Drop an inbound message into the right support queue before a human ever sees it. One score call replaces a brittle keyword rule set and returns the queue plus a confidence you can threshold on. choices: ["billing","shipping","technical","account","none of the above"]

2. Ticket triage and priority. Grade an incoming ticket by urgency so the queue orders itself. "Production is down for all users" and "typo on the pricing page" land in different buckets in ~355 ms, every time. choices: ["p1 urgent","p2 high","p3 normal","p4 low"]

3. Guardrail yes/no. A binary allow/block gate in front of a more expensive pipeline. Because it's deterministic, the same input always gets the same verdict — which is what you want from a guardrail. choices: ["allow","block"]

4. PII / sensitivity flagging. Flag whether a chunk of text carries personal or otherwise sensitive data before it's logged, indexed, or forwarded. Since the check runs on-box, the sensitive text never has to leave the machine to find out it's sensitive. choices: ["contains pii","no pii","unsure"]

5. Sentiment. Classify a review, a survey reply, or a chat turn into a fixed sentiment scale. At a few hundred milliseconds a decision, you can score a firehose of feedback in near real time. choices: ["positive","neutral","negative","mixed"]

6. Language / locale routing. Decide which language a message is in so it goes to the right localized model, template, or agent. choices: ["english","spanish","french","german","other"]

7. Spam / abuse detection. Score whether a submission is spam or abusive at the edge of your intake, deterministically, so the same payload can't slip through by luck on a retry. choices: ["spam","abuse","legitimate"]

8. Approval gates. Decide whether an action can auto-approve or must escalate to a human. A refund under a threshold auto-approves; anything ambiguous escalates — and the decision is repeatable, which matters when you have to explain it later. choices: ["auto-approve","escalate to human","reject"]

9. Tool / skill selection for an agent. Instead of asking a model to write which tool to call, ask it to pick one from the tools it actually has. One prefill, no decode, and the answer is constrained to real tool names. choices: ["search_docs","run_sql","send_email","none of the above"]

10. Document / type classification. Sort an incoming document into a type so the right extractor or workflow picks it up. Feed the first page or a summary as the query and read the label. choices: ["invoice","contract","resume","receipt","other"]

Keep the stable part first

The decision is fast on its own — q35-4b ~355 ms (about 2.6× faster than generating the answer) and spark ~394 ms (about 1.6×). But an exact repeat over the same choice set lands in about 84 ms, because the prefix cache reuses the already-prefilled prompt.

To get those hits, order your prompt stable part first: put the instruction and the label set up front, and vary only the query. When the leading tokens don't change from call to call, the cache does the heavy lifting and you pay only for the new query.

Why it's fast

A chat completion runs a full decode loop: prefill the prompt, then generate the answer one token at a time, each step a forward pass. /v1/score skips all of that. It does a single prefill and reads a restricted softmax over the allowed answers — no tokens are generated, so there's no per-token cost and no loop to wait through. Add the prefix cache for repeated choice sets and the marginal cost of a decision drops to the new query alone. The same mechanics make it deterministic: no sampling, same input in, same decision out.

Getting started

Install the server, fetch a Qwen 3.5 or Spark X2.5 model, and point /v1/score at it with the "model" field:

curl -fsSL https://inference-server.searchblox.com/install | sudo bash
# fetch q35-4b (or a spark-x25 model) from the console catalog, then:

curl -s http://127.0.0.1:8081/v1/score \
  -H "Authorization: Bearer $SEARCHAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "q35-4b",
        "query": "My card was charged twice.",
        "choices": ["billing","shipping","technical","none of the above"] }'

Numbers in this post: decision latencies measured on CPU — q35-4b ~355 ms (2.6× vs. generating the answer), spark ~394 ms (1.6×), ~84 ms on an exact-repeat via the prefix cache. /v1/score is new in v1.4.0 and supported on the Qwen 3.5 and Spark X2.5 families; 2–26 choices in the friendly form, with a raw label_token_ids form at parity with SGLang's /v1/score.