How to Build a Voice Assistant That Responds in Under 1 Second
October 6, 2026 · technical guide · 9 minutes read
Most voice assistants feel slow. You say something, wait, wait more, then finally get a response. Google Assistant averages 1.5-2 seconds. Alexa is similar. Siri can hit 3 seconds on a bad day. Users tolerate it because there's no better option.
But 2 seconds is forever in conversation. Humans respond in 200-300ms. At 2 seconds, the assistant feels like it's thinking too hard or didn't hear you. The interaction stops feeling like dialogue.
Amalia hits ~0.87 seconds from end of speech to first sound of response. This guide explains how.
The Four-Stage Pipeline
Every voice assistant has the same basic flow:
- Voice Activity Detection (VAD): Is the user speaking? Have they finished?
- Speech-to-Text (STT): Convert audio to text
- Language Model (LLM): Generate a response (and optionally trigger actions)
- Text-to-Speech (TTS): Convert response text back to audio
The naive approach runs these sequentially: wait for VAD, then send to STT, wait for full transcription, send to LLM, wait for complete response, send to TTS, wait for full audio, then play. That's 3-4 seconds minimum.
The fast approach pipelines everything and streams where possible.
Stage 1: Voice Activity Detection (200ms)
You can't start transcription until the user stops talking. But you also can't wait too long — a half-second pause kills the conversation flow.
Amalia uses Silero VAD running offline on the device:
- Tiny model (~5 MB), runs in real-time on any phone
- Outputs speech probability for each audio frame
- When probability drops below threshold for 600ms → user is done
Why 600ms? Less than that and you cut off natural pauses ("set a timer for… ten minutes"). More than that and the assistant feels sluggish.
Cost: ~200ms (the silence detection window after speech ends)
More on VAD: how voice activity detection works
Stage 2: Speech Recognition (220ms)
Once you have the audio, it needs to become text. Traditional STT models (on-device or cloud) take 500-800ms. That's the biggest bottleneck.
Amalia uses Groq Whisper v3 Turbo:
- Groq LPU (Language Processing Unit) is custom silicon for inference
- Whisper v3 Turbo runs at 170-220ms for typical utterances (under 10 seconds of audio)
- API is simple: send WAV, get JSON with transcription
Why Groq and not on-device Whisper? On-device Whisper (whisper.cpp) takes 1-2 seconds even on flagship phones. You can't build a sub-second assistant with 2-second STT.
Cost: ~220ms
Implementation details: Groq Whisper API integration
Stage 3: Language Model Inference (150ms to first token)
Now you have text. The LLM needs to:
- Understand the request
- Decide what actions to take (if any)
- Generate a natural language response
Traditional approach: wait for the full response, then parse it. That's another 2-3 seconds.
The Groq LPU Advantage
Amalia uses Qwen3 27B on Groq LPU infrastructure:
- Time to first token: 100-150ms (vs 500-800ms on standard GPU inference)
- Streaming: Tokens arrive as they're generated, not after the full response is done
- JSON mode: The model outputs structured data
{"reply":"...","tools":[...]}reliably
Why does first token matter? Because you can start TTS synthesis as soon as the first few tokens arrive. You don't wait for the full reply.
Cost: ~150ms to first token, but TTS starts immediately so user perceives less
What is LPU and why it's faster: Groq LPU explained
Stage 4: Text-to-Speech Streaming (Parallel)
This is where most assistants fail. They wait for the LLM to finish, then send the entire text to TTS, wait for the full audio to generate, then start playback. That adds 1-2 seconds.
Streaming TTS Changes Everything
Amalia uses Fish Audio drama-3 with streaming:
- TTS starts as soon as the first sentence is ready (you don't need the full LLM response)
- Audio is streamed as raw PCM chunks, not MP3 (no decode latency)
- Chunks are fed directly to Android AudioTrack — playback starts immediately
This overlaps LLM generation and TTS synthesis. While the model is generating token 50, the TTS is already speaking token 10.
Cost: ~0ms perceived (runs in parallel with LLM)
Why PCM over MP3: streaming TTS architecture
The Full Timeline (Real Numbers)
Here's the actual measured latency from Amalia in production:
0ms User stops speaking ↓ ~200ms VAD detects silence, audio buffer ready ↓ ~420ms Groq Whisper completes transcription ↓ ~570ms Qwen3 first token arrives, TTS starts ↓ ~870ms First PCM chunk hits AudioTrack, user hears voice
Total: 0.87 seconds.
In good conditions (short utterance, fast network), this drops to 0.7s. In poor conditions (long speech, slow connection), it rises to 1.2s. But it averages 0.87s, which is 2x faster than Google Assistant.
The Architecture Principles
Four rules for sub-second latency:
1. Never Wait for Full Data
Don't wait for the full LLM response before starting TTS. Don't wait for the full TTS audio before starting playback. Stream everything.
2. Run Stages in Parallel
As soon as VAD fires, start preparing the STT request. As soon as first tokens arrive, start TTS. Overlap everything possible.
3. Choose Speed-Optimized Infrastructure
Groq LPU is 3-5x faster than standard GPU inference for the same model size. That's the difference between 1 second and 3 seconds.
4. Keep Hot Paths Offline
VAD runs on-device so there's zero network latency before you know speech has ended. If VAD required a network call, you'd lose another 100-200ms.
The Trade-Offs
This architecture isn't free:
- Requires internet: STT and LLM are cloud-based. No offline mode (yet).
- API costs: Groq and Fish Audio are paid services. Free tiers cover normal use, but heavy users pay.
- Complexity: Streaming pipelines are harder to debug than sequential ones.
But the result is worth it: a voice assistant that feels responsive, not sluggish.
Can You Go Faster?
Probably. The current bottleneck is STT (220ms). If Groq releases Whisper Turbo v4 or if someone builds faster ASR, you could hit 0.6-0.7s consistently.
The theoretical minimum is ~100ms (network round-trip + model inference + audio buffering). We're at 870ms, so there's still room.
How to Implement This Yourself
If you're building a voice assistant:
- Use Silero VAD for on-device speech detection
- Use Groq for STT and LLM (or another low-latency provider)
- Stream your TTS — don't wait for full audio generation
- Feed raw PCM to AudioTrack — skip MP3 encoding/decoding
- Measure everything — instrument each stage and optimize the slowest one
Full implementation walkthrough: how to build your own voice assistant
Try It Yourself
Want to experience sub-second voice latency? Download Amalia and see how fast an open source assistant can be. The code is MIT-licensed — fork it, modify it, ship it.
Related: open source voice assistant overview · latency analysis (Russian) · Groq Whisper integration