← Blog · Download & Install · Performance

Two prompts, one small model: a worked example of testing locally to cut tokens and latency

August 2026 · SearchAI Inference Server team

worked exampleprompt testing local inferencetoken savings real captured runs

A companion piece to Optimize the prompt, shrink the model made the economic case. This one is hands-on: two prompts, the same small model, and the actual outputs and token counts we captured running them through a local server. No leaderboards — just the before and after a developer sees in an afternoon of prompt testing.

Every number and every model response below is copied verbatim from q35-2b (a 2-billion-parameter model, ~1.9 GB, runs on a 16 GB box) served locally at greedy decoding. Nothing is idealized; you can reproduce all of it.

Example 1 — pulling structured data out of messy text

The task every integration needs: turn a human sentence into JSON your code can consume. Here's the input, an order note:

Order #4471 came in today from Acme Corp — they want 3 units of the
widget-X (part WX-9), plus 1 replacement cable, ship to their Denver
office by March 5th. Rush.

The prompt most people write first:

Extract the order info from this: {text}

What the 2B model returns:

Based on the text provided, here is the extracted order information:

*   **Order Number:** 4471
*   **Ship To:** Denver office (Acme Corp)
*   **Due Date:** March 5th
*   **Urgency:** Rush

67 prompt tokens · 59 completion tokens

It reads fine to a human and is useless to your code: it's prose, not JSON, and — the quiet failure — it silently dropped both line items. No widget-X, no quantity, no cable. Ship this and your order pipeline loses what the customer actually ordered.

The prompt after five minutes of testing — a system message that names the job, pins the schema, and forbids prose:

You are an order-extraction service. Return ONLY valid JSON, no
prose, matching exactly this schema:
{"order_id": string, "customer": string,
 "line_items": [{"part": string, "qty": number}],
 "ship_to": string, "due_date": "YYYY-MM-DD", "rush": boolean}
If a field is missing use null.

Same model, same input, new answer:

{"order_id": "4471", "customer": "Acme Corp",
 "line_items": [{"part": "widget-X (WX-9)", "qty": 3},
                {"part": "replacement cable", "qty": 1}],
 "ship_to": "Denver office", "due_date": "2025-03-05", "rush": true}

139 prompt tokens · 81 completion tokens

Both line items recovered. Quantities correct. "March 5th" normalized to 2025-03-05. rush is a real boolean. This parses on the first try — from a 2B model that fits on a laptop-class box. The model didn't get smarter; the prompt got specific.

Example 2 — where the token and latency savings show up

Classification is the case that makes the cost of a lazy prompt impossible to miss. The task: route a support ticket into one bucket. The input:

Hi, I was charged twice for my subscription this month and I want
the duplicate refunded ASAP. Order ARR-8823.

The obvious prompt:

What is this support ticket about? {ticket}

The 2B model, being helpful, writes an essay:

This support ticket is a request for a **duplicate charge refund**.

Here are the key details from your message:
*   **Issue**: The user was charged twice for their subscription...
*   **Action Required**: A refund for the duplicate amount...
*   **Order Reference**: `ARR-8823`...

**Next Steps:** You should copy and paste this entire message...

120 completion tokens · 2,698 ms

Your router can't use any of that without more parsing, and you paid for 120 tokens and 2.7 seconds to get it.

The tested prompt — constrain the label set and the output shape:

Classify the support ticket into EXACTLY ONE category from this
list: [billing, technical, account, shipping, other]. Reply with
ONLY the category word, lowercase, nothing else.

The entire response:

billing

1 completion token · 298 ms

Same model, same ticket. The difference is only the prompt, and the difference is not small:

Metric
Obvious prompt
Tested prompt
Change
Output tokens
120
1
120× fewer
Latency
2,698 ms
298 ms
9× faster
Usable by code?
No (prose)
Yes (label)
correctness

Multiply that by every ticket, every day. The tokens you don't generate and the seconds you don't wait are recovered on every single request, forever — and it came from a prompt edit, not a bigger machine.

Why do this on a local model?

Prompt testing is inherently iterative — you'll run a prompt dozens of times before it's right. Doing that loop locally changes its economics and its honesty:

  • Every iteration is free. The runs above cost nothing but a few seconds of CPU. Burn a thousand comparisons finding the prompt that works and the bill is the same flat instance price — there's no per-token meter running while you experiment.
  • You test the model you ship. Optimize against q35-2b locally and that's exactly what serves the request in production — no "works on the frontier API, degrades on the cheap tier" surprise, because there's only one tier.
  • The savings compound where it's cheapest. The smaller the model, the more a sharp prompt matters — and the smaller the model, the more requests per box you serve. A well-prompted 2B on one node quietly handles work people assume needs a big model.
  • Your data stays on your infrastructure. The tickets, orders, and documents you're prompting over never leave the box you own.

The loop, in the built-in console

You don't need a script to do any of this. The server ships a console at http://<host>:8081/console built for exactly this loop:

  1. Compare — put your prompt on two models side by side (or your local model against an external OpenAI-compatible endpoint through the proxy). Both stream with live tokens/sec, so the speed half of the trade-off is measured, not guessed.
  2. Optimize — one click sends the prompt and both outputs to a reviewer model that critiques them, flags what the prompt left ambiguous, and proposes improved variants with a Use button.
  3. Re-compare — apply a variant, run again, watch the output tokens and latency drop like they did above. Each loop is seconds and zero dollars.

When the prompt holds up on your own inputs, every prompt card has Copy cURL / JSON / Python — the API is OpenAI-compatible, so shipping it is a two-line change in your app.

The takeaway is small and repeatable: before you reach for a bigger model, spend five minutes testing the prompt on the small one. The examples here weren't cherry-picked wins — they're the ordinary result of naming the job, pinning the output shape, and cutting the open-ended ask. Do it locally, where iterating is free and you're testing the exact model you'll deploy.

Install (Linux, one line): curl -fsSL https://inference-server.searchblox.com/install | sudo bash

Then open http://<host>:8081/console and run your first comparison on a prompt you already have.

Every model response and token/latency figure above was captured from q35-2b served on a SearchAI Inference Server node, greedy decoding, August 2026. See Getting Started for sizing and the console walkthrough, and the companion post Optimize the prompt, shrink the model for the 380-prompt evaluation behind the argument.