Offline Voice Assistant for Android: Options and Trade-offs
October 6, 2026 · guide · 9 minutes read
Google Assistant requires internet. Alexa requires internet. Siri requires internet. Every mainstream voice assistant sends your voice to the cloud, processes it there, and sends back a response. That's fast, accurate, and convenient — until you lose signal, travel abroad, or care about privacy.
Offline voice assistants run entirely on your phone. No cloud, no tracking, no data leaving your device. But there are trade-offs. This guide covers what's possible in 2026, what works well, and what still sucks.
Why Offline Matters
Three reasons people want offline voice assistants:
- Privacy: Your voice never leaves the device. No one can intercept, store, or analyze it.
- Reliability: Works in airplane mode, underground, or in poor network conditions.
- Cost: No API fees, no data usage, no monthly subscriptions.
The downside: offline models are slower, less accurate, and harder to set up than cloud alternatives.
The Three Components of an Offline Assistant
A voice assistant has three stages:
- Speech-to-Text (STT): Audio → text
- Language Model (LLM): Text → response + actions
- Text-to-Speech (TTS): Text → audio
For a fully offline assistant, all three must run on-device. Let's look at each.
Offline Speech-to-Text (STT)
Option 1: Vosk
Vosk is the most popular offline STT for Android:
- Model size: 40 MB (small model) to 1 GB (large model)
- Speed: Real-time on flagship phones, 1.5-2x real-time on mid-range
- Accuracy: Good for English, acceptable for other languages
- License: Apache 2.0 (open source)
Vosk works by loading a model into memory and feeding audio chunks. It returns partial results as you speak. Implementation guide: Vosk on Android (Russian).
Option 2: Whisper.cpp
OpenAI's Whisper model, ported to C++ for on-device use:
- Model size: 70 MB (tiny) to 1.5 GB (medium)
- Speed: 1-2 seconds on flagship phones (Snapdragon 8 Gen 3), 3-5 seconds on mid-range
- Accuracy: Excellent — better than Vosk, especially for accents and multilingual
- License: MIT (open source)
Whisper.cpp is slower than Vosk but significantly more accurate. Good for transcription, less ideal for real-time conversation. Details: Whisper on Android (Russian).
Option 3: Android Speech Recognizer (Limited Offline)
Android has built-in offline STT, but:
- Only works if you've downloaded language packs via Google
- Still sends data to Google for "quality improvement" (even in offline mode)
- Accuracy is worse than cloud mode
Not recommended for privacy-focused use cases.
Speed Comparison (Real Numbers)
| STT | Speed (3s audio) | Accuracy | Model Size |
|---|---|---|---|
| Groq Whisper (cloud) | 220ms | Excellent | N/A (cloud) |
| Vosk (small model) | 3.5s | Good | 40 MB |
| Whisper.cpp (tiny) | 1.5s | Excellent | 70 MB |
| Whisper.cpp (medium) | 4s | Excellent+ | 1.5 GB |
Recommendation: Whisper.cpp tiny model. Best balance of speed, accuracy, and size.
Offline Language Models (LLM)
Once you have text, the LLM decides what to say and what actions to take. Running LLMs on-device in 2026:
Option 1: Gemini Nano
Google's on-device LLM, available on Pixel 8+ and select flagships:
- Model size: ~3 GB
- Speed: 500ms to first token on Pixel 8 Pro
- Quality: Good for summarization, acceptable for conversation
- Limitation: Proprietary, only on supported devices, no control over system prompt
Option 2: Llama.cpp (Llama 3.2 3B)
Open source quantized models run via llama.cpp on Android:
- Model size: 1.5 GB (Q4 quantized)
- Speed: 1-2 seconds to first token on flagship, 3-5 seconds on mid-range
- Quality: Decent for simple tasks, struggles with complex reasoning
- Limitation: Slow compared to cloud, requires 6+ GB RAM
Option 3: Phi-3 Mini
Microsoft's small LLM, optimized for mobile:
- Model size: 2 GB (4-bit quantized)
- Speed: Similar to Llama 3.2 3B
- Quality: Better reasoning than Llama 3B, worse than Llama 7B
The Honest Truth About On-Device LLMs
In 2026, on-device LLMs are okay but not great:
- Responses take 2-5 seconds (vs 150ms for Groq cloud LLMs)
- Models hallucinate more often (smaller = less knowledge)
- Battery drain is significant (5-10% per 10 minutes of heavy use)
- Phone gets hot under sustained load
For command-only assistants ("turn on Wi-Fi", "set a timer"), you don't need an LLM at all — just pattern matching. For conversational AI, cloud LLMs are 10x better.
Offline Text-to-Speech (TTS)
This is the easiest part. Good offline TTS has existed for years:
Option 1: Android TTS (Built-in)
- Works offline if you've downloaded voice data
- Sounds robotic but intelligible
- Free, no setup
Option 2: Piper TTS
- Neural TTS running on-device
- Much better quality than Android TTS
- Model size: 10-50 MB per voice
- Real-time synthesis on any modern phone
Option 3: eSpeak NG
- Tiny (under 5 MB)
- Extremely robotic voice
- Only use if storage is critical
Recommendation: Piper TTS. Sounds human enough, runs fast, open source.
Fully Offline Voice Assistants in 2026
Dicio
Open source offline assistant for Android:
- Uses Vosk for STT
- No LLM — pattern-based command matching
- Android TTS for speech
- Works fully offline
- Limitation: Command-only, no conversation
Good for: "open YouTube", "what's the weather", "set a timer"
Bad for: "why is the sky blue", "tell me about quantum physics"
DIY: Whisper.cpp + Llama.cpp + Piper
If you're a developer, you can build a fully offline assistant:
- Whisper.cpp tiny for STT (~1.5s)
- Llama 3.2 3B for LLM (~2s to first token)
- Piper TTS for speech (~100ms)
Total latency: 3.5-4 seconds from speech to response. Slow, but private.
Guide: how to build a voice assistant (adapt for on-device models).
Hybrid Approach: Best of Both Worlds
Instead of fully offline or fully cloud, consider a hybrid:
- VAD (voice detection): Offline (Silero VAD)
- STT: Cloud (Groq Whisper) with offline fallback (Vosk)
- LLM: Cloud (Qwen3 on Groq) with offline fallback (Llama 3.2 3B)
- TTS: Cloud (Fish Audio) with offline fallback (Piper)
This gives you:
- Fast responses when online (~0.87s)
- Degraded but functional experience when offline (~3.5s)
- Privacy when needed (airplane mode = fully offline)
This is the roadmap for Amalia — fast cloud by default, offline fallback for privacy-conscious users.
The Privacy Question
Is offline really more private?
- Yes: If you use fully open source models (Whisper.cpp, Llama.cpp, Piper), your data never leaves the device.
- No: If you use proprietary on-device models (Gemini Nano, Android STT), Google still collects usage data and model queries.
For true privacy, use fully open source offline models. For convenience, use cloud with your own API keys (you control the data, not Google/Amazon).
When Offline Makes Sense
Use an offline voice assistant if:
- You travel frequently to areas with poor connectivity
- You need guaranteed privacy (journalist, activist, security professional)
- You have unlimited storage and don't mind 2-4 GB of models
- You're okay with 3-5 second latency
When Cloud Makes More Sense
Use a cloud voice assistant if:
- You value speed (sub-second responses)
- You need high accuracy (multilingual, accents, complex queries)
- You have reliable internet
- You're okay with API costs ($0-10/month for typical use)
Try Both
Want to experience the speed of cloud inference? Download Amalia — open source, sub-second responses, privacy-focused architecture. Offline mode coming soon.
Want fully offline? Try Dicio from F-Droid — command-based, but works without internet.
Related: open source voice assistants · Whisper on Android (Russian) · voice assistant privacy