← Blog · Download & Install · Use Cases

v1.3.0: double your AI throughput — no new hardware, no new risks

September 2026 · SearchAI Inference Server team

deep analysis in half the time 2× team throughput on one box instant follow-ups in chat answers unchanged — nothing to re-test

Most software upgrades ask you for something — new hardware, a migration project, or a round of re-testing because the outputs changed. v1.3.0 asks for none of that. It's a pure engineering release: the models you already run now do more than twice the work on the same machine, and every answer they produce is guaranteed identical to the previous version. If you validated prompts, reports, or extraction pipelines on v1.2, they behave exactly the same on v1.3 — just faster.

Here's what that speed actually buys your team, and which jobs the two newest models are best at.

What the speed means in practice

If you use it for…BeforeNow
Deep document analysis (Flash-Next)a contract review that ran ~an hourdone in under 30 minutes — 2.3× faster
A shared team assistant (Spark)the server slowed down when several people used it2.2× the throughput when requests overlap — one box serves the team
Back-and-forth conversationsevery follow-up re-read the whole chat historyfollow-ups start near-instantly — a 9× faster response start in our tests

Measured on standard CPU hardware. The same relative gains apply on AWS Graviton, Intel/AMD servers, Apple silicon, and Windows — no GPU required.

Spark-X2.5: the everyday workhorse — now built for teams

Spark-X2.5 (4B and 1.7B) is the model tier for the AI work that happens constantly: classifying incoming requests, pulling fields out of forms, routing tickets, summarizing threads, answering from your knowledge base. It was already fast for one user. v1.3.0 makes it fast for a department: when several requests arrive together, the server now processes them through the model in one combined pass instead of one at a time.

Where it's a perfect fit:

  • Support & ticket triage — every incoming email or ticket classified, prioritized, and routed in seconds, even at 9 AM when everything arrives at once. See the measured triage walkthrough.
  • Document & form processing at volume — contracts, invoices, surveys, resumes turned into clean structured data. One server now works through a stack more than twice as fast when jobs run in parallel. See document extraction (~33 s per contract, measured).
  • A private assistant for the whole team — because follow-up questions in a conversation now start instantly, and concurrent users no longer queue behind each other. Pairs with built-in sessions & memory.

And the privacy case is the whole point: HR files, contracts, and customer data never leave your network, because the model runs on your machine.

Flash-Next: frontier-class reasoning, privately — now twice as fast

Qwen3.8-Flash-Next is the opposite tier: a 125-billion-parameter reasoning model that thinks before it answers. You don't use it for instant replies — you use it for work you'd otherwise send to an expensive frontier API, on material you can't send anywhere: analyzing a contract clause by clause, drafting a compliance report against your policies, digesting a long technical document into an executive brief.

In v1.2 that was an overnight-and-review workflow. At 2.3× the speed, it becomes a same-meeting workflow: queue the analysis, get the draft back while the topic is still open. A detailed 500-word analysis that took ~10 minutes of generation now lands in about 4.

Where it's a perfect fit:

  • Compliance & incident reporting — policy + incident record in, review-ready report with clause citations out. See the measured compliance walkthrough.
  • Contract and policy review — clause extraction, risk flags, and plain-English summaries of documents that legally cannot transit a third-party API.
  • Research & due-diligence digests — long reports in, structured briefs out, running in the background while the team works.

A practical pattern our customers use: Flash-Next plans, Spark executes. The big model breaks the job down once; the fast model does the hundred small steps. Both live on the same server and you pick per request with one field.

Why "identical answers" is a business feature

Speed upgrades in AI usually come with a quiet cost: the model's answers shift a little, and every validated workflow needs re-checking. We engineered v1.3.0 to a stricter standard — every optimization had to produce exactly the output the previous version produced, verified token by token across every model family before release.

For a business that means the upgrade is genuinely free: no re-validation of extraction templates, no drifted report formats, no surprise changes in a compliance pipeline. Faster, and provably the same.

Get it

v1.3.0 is live on Linux, macOS, Windows, and Docker. Upgrading is one line — models, conversations, and settings carry over untouched:

curl -fsSL https://inference-server.searchblox.com/install | sudo bash

New here? Start with the use-case walkthroughs — each one is a real workload with measured timings you can reproduce on your own hardware, free under the Elastic License 2.0.

Speed figures measured on standard CPU hardware (Apple M4-class and AWS instances); relative gains carry across platforms. Output-identity verified by token-exact comparison suites across all served model families at release.