Speech in, speech out: ASR and TTS with the local models
August 2026 · SearchAI Inference Server team
speech-to-texttext-to-speech local & privatereal captured output voice cloning
Chat and vision get the attention, but the server also does
speech both directions — transcription (ASR) and
synthesis (TTS) — from the same OpenAI-compatible endpoints, entirely on
your own box. No cloud speech API, no per-minute meter, and for the
workloads where audio is sensitive (support calls, clinical dictation,
voicemail, meetings) nothing ever leaves the network. Install the
asr and tts model groups and the endpoints light
up. Everything below is captured from a running server.
Speech to text: /v1/audio/transcriptions
POST a WAV, get text back — the same shape as the OpenAI transcription API, so existing clients work unchanged:
curl http://<host>:8081/v1/audio/transcriptions \
-H "Authorization: Bearer $KEY" \
-F model=qwen3-asr-1.7b \
-F file=@call.wav
Real transcripts from three sample clips — accurate, punctuated, with numbers and names intact:
The short clips transcribe in a couple of seconds on CPU. Because it's just text out, the transcript flows straight into the rest of the server: summarize the call with a chat request, extract fields to JSON with the read-then-structure pattern, or route it — all on the same box, same key.
Text to speech: /v1/audio/speech
The other direction — text in, a WAV out:
curl http://<host>:8081/v1/audio/speech \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{"model":"qwen-talker-1.7b","voice":"samantha",
"input":"Your appointment with Doctor Patel is confirmed for next Tuesday at ten thirty in the morning."}' \
--output reply.wav
Here's the actual audio the server produced from that request — 24 kHz, about five seconds, generated locally:
Server-generated with the built-in samantha
voice from the fictional sentence above.
The round trip: a private voice assistant
Chain the three capabilities and you have a voice loop that never touches a third party:
- Listen —
/v1/audio/transcriptionsturns the caller's audio into text. - Think —
/v1/chat/completionsanswers, grounds against your documents, or calls a tool with that text. - Speak —
/v1/audio/speechvoices the answer back as audio.
Every hop is a local model on the same node. For a support line, a kiosk, an accessibility feature, or an IVR replacement, the audio and the transcript stay inside your network end to end.
Your own voice: file-based cloning
The samantha preset is the default, but you can add a
voice by dropping a short reference recording into the models directory —
no training run, no upload:
# 10-30 s of clean mono speech, named for the voice
cp narrator.wav /var/lib/searchai/models/tts-voices/narrator.wav
# then request it
... -d '{"model":"qwen-talker-1.7b","voice":"narrator","input":"..."}'
The first request with a new voice computes its speaker embedding from that file; after that it's instant. Only clone voices you have permission to use.
The point: speech in and speech out are already in the box you installed — same API, same key, same privacy guarantee as the rest. No separate transcription vendor, no per-minute billing, no audio crossing the network boundary.
Install (Linux) with the audio groups:
curl -fsSL https://inference-server.searchblox.com/install | MODELS='4b asr tts' sudo bash
ASR/TTS transcripts and the audio sample above were
captured from qwen3-asr-1.7b and qwen-talker-1.7b
on a running SearchAI Inference Server, August 2026, from synthetic
clips. Audio work wants a 32 GB host alongside a chat model — see
Getting Started for sizing
and the model catalog for the full line-up.