← Blog · Models · Performance

Speech in, speech out: ASR and TTS with the local models

August 2026 · SearchAI Inference Server team

speech-to-texttext-to-speech local & privatereal captured output voice cloning

Chat and vision get the attention, but the server also does speech both directions — transcription (ASR) and synthesis (TTS) — from the same OpenAI-compatible endpoints, entirely on your own box. No cloud speech API, no per-minute meter, and for the workloads where audio is sensitive (support calls, clinical dictation, voicemail, meetings) nothing ever leaves the network. Install the asr and tts model groups and the endpoints light up. Everything below is captured from a running server.

Speech to text: /v1/audio/transcriptions

POST a WAV, get text back — the same shape as the OpenAI transcription API, so existing clients work unchanged:

curl http://<host>:8081/v1/audio/transcriptions \
  -H "Authorization: Bearer $KEY" \
  -F model=qwen3-asr-1.7b \
  -F file=@call.wav

Real transcripts from three sample clips — accurate, punctuated, with numbers and names intact:

Spoken audio
Transcript
account balance
"Your account balance is four thousand two hundred dollars and fifteen cents."
appointment
"Please schedule a follow-up appointment with Doctor Patel for next Tuesday at ten thirty."
insurance claim
"I would like to file a claim for water damage that occurred on March third."

The short clips transcribe in a couple of seconds on CPU. Because it's just text out, the transcript flows straight into the rest of the server: summarize the call with a chat request, extract fields to JSON with the read-then-structure pattern, or route it — all on the same box, same key.

Text to speech: /v1/audio/speech

The other direction — text in, a WAV out:

curl http://<host>:8081/v1/audio/speech \
  -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
  -d '{"model":"qwen-talker-1.7b","voice":"samantha",
       "input":"Your appointment with Doctor Patel is confirmed for next Tuesday at ten thirty in the morning."}' \
  --output reply.wav

Here's the actual audio the server produced from that request — 24 kHz, about five seconds, generated locally:

Server-generated with the built-in samantha voice from the fictional sentence above.

The round trip: a private voice assistant

Chain the three capabilities and you have a voice loop that never touches a third party:

  1. Listen — /v1/audio/transcriptions turns the caller's audio into text.
  2. Think — /v1/chat/completions answers, grounds against your documents, or calls a tool with that text.
  3. Speak — /v1/audio/speech voices the answer back as audio.

Every hop is a local model on the same node. For a support line, a kiosk, an accessibility feature, or an IVR replacement, the audio and the transcript stay inside your network end to end.

Your own voice: file-based cloning

The samantha preset is the default, but you can add a voice by dropping a short reference recording into the models directory — no training run, no upload:

# 10-30 s of clean mono speech, named for the voice
cp narrator.wav /var/lib/searchai/models/tts-voices/narrator.wav

# then request it
... -d '{"model":"qwen-talker-1.7b","voice":"narrator","input":"..."}'

The first request with a new voice computes its speaker embedding from that file; after that it's instant. Only clone voices you have permission to use.

The point: speech in and speech out are already in the box you installed — same API, same key, same privacy guarantee as the rest. No separate transcription vendor, no per-minute billing, no audio crossing the network boundary.

Install (Linux) with the audio groups:

curl -fsSL https://inference-server.searchblox.com/install | MODELS='4b asr tts' sudo bash

ASR/TTS transcripts and the audio sample above were captured from qwen3-asr-1.7b and qwen-talker-1.7b on a running SearchAI Inference Server, August 2026, from synthetic clips. Audio work wants a 32 GB host alongside a chat model — see Getting Started for sizing and the model catalog for the full line-up.