Groq Whisper API Integration for Android Apps
October 6, 2026 · developer guide · 8 minutes read
Groq's Whisper v3 Turbo is the fastest production speech-to-text API available in 2026. It runs on custom LPU silicon and transcribes audio in 170-220ms — 3x faster than standard cloud STT services. This guide shows how to integrate it into an Android app.
Why Groq Whisper for Android
Three reasons:
- Speed: 170-220ms latency for typical utterances. Fast enough for real-time voice assistants.
- Accuracy: OpenAI Whisper v3 model — best-in-class multilingual recognition.
- Free tier: Groq offers free API access with generous limits for development and personal use.
On-device Whisper (whisper.cpp) takes 1-2 seconds on flagship phones. Cloud Whisper on standard infrastructure takes 500-800ms. Groq LPU does it in 220ms.
Prerequisites
You'll need:
- Android project with Kotlin
- OkHttp for networking (
com.squareup.okhttp3:okhttp:4.12.0) - Groq API key (get it free at console.groq.com)
- RECORD_AUDIO permission in your manifest
Step 1: Record Audio
Groq Whisper expects audio in specific formats. The safest bet is WAV, 16kHz, mono, 16-bit PCM.
Using AudioRecord
import android.media.AudioFormat
import android.media.AudioRecord
import android.media.MediaRecorder
import java.io.ByteArrayOutputStream
class AudioRecorder {
private val sampleRate = 16000
private val channelConfig = AudioFormat.CHANNEL_IN_MONO
private val audioFormat = AudioFormat.ENCODING_PCM_16BIT
private val bufferSize = AudioRecord.getMinBufferSize(
sampleRate, channelConfig, audioFormat
)
private var audioRecord: AudioRecord? = null
private val audioBuffer = ByteArrayOutputStream()
fun startRecording() {
audioRecord = AudioRecord(
MediaRecorder.AudioSource.VOICE_RECOGNITION,
sampleRate,
channelConfig,
audioFormat,
bufferSize
)
audioRecord?.startRecording()
// Read audio in a background thread
val buffer = ShortArray(bufferSize / 2)
while (isRecording) {
val read = audioRecord?.read(buffer, 0, buffer.size) ?: 0
if (read > 0) {
// Convert shorts to bytes and write
buffer.take(read).forEach { sample ->
audioBuffer.write(sample.toInt() and 0xFF)
audioBuffer.write((sample.toInt() shr 8) and 0xFF)
}
}
}
}
fun stopRecording(): ByteArray {
audioRecord?.stop()
audioRecord?.release()
return audioBuffer.toByteArray()
}
}
Convert to WAV Format
Raw PCM needs a WAV header. Here's a minimal WAV writer:
fun pcmToWav(pcmData: ByteArray, sampleRate: Int = 16000): ByteArray {
val output = ByteArrayOutputStream()
val channels = 1
val bitsPerSample = 16
// RIFF header
output.write("RIFF".toByteArray())
output.write(intToBytes(36 + pcmData.size))
output.write("WAVE".toByteArray())
// fmt chunk
output.write("fmt ".toByteArray())
output.write(intToBytes(16)) // chunk size
output.write(shortToBytes(1)) // audio format (PCM)
output.write(shortToBytes(channels))
output.write(intToBytes(sampleRate))
output.write(intToBytes(sampleRate * channels * bitsPerSample / 8))
output.write(shortToBytes(channels * bitsPerSample / 8))
output.write(shortToBytes(bitsPerSample))
// data chunk
output.write("data".toByteArray())
output.write(intToBytes(pcmData.size))
output.write(pcmData)
return output.toByteArray()
}
fun intToBytes(value: Int) = byteArrayOf(
(value and 0xFF).toByte(),
((value shr 8) and 0xFF).toByte(),
((value shr 16) and 0xFF).toByte(),
((value shr 24) and 0xFF).toByte()
)
fun shortToBytes(value: Int) = byteArrayOf(
(value and 0xFF).toByte(),
((value shr 8) and 0xFF).toByte()
)
Step 2: Call Groq Whisper API
Groq uses OpenAI-compatible API. Send a multipart form with the audio file.
API Client
import okhttp3.*
import okhttp3.MediaType.Companion.toMediaType
import org.json.JSONObject
import java.io.IOException
class GroqWhisperClient(private val apiKey: String) {
private val client = OkHttpClient.Builder()
.connectTimeout(10, TimeUnit.SECONDS)
.readTimeout(30, TimeUnit.SECONDS)
.build()
fun transcribe(wavData: ByteArray): String? {
val requestBody = MultipartBody.Builder()
.setType(MultipartBody.FORM)
.addFormDataPart(
"file",
"audio.wav",
wavData.toRequestBody("audio/wav".toMediaType())
)
.addFormDataPart("model", "whisper-large-v3-turbo")
.addFormDataPart("language", "en") // optional, auto-detect if omitted
.addFormDataPart("response_format", "json")
.build()
val request = Request.Builder()
.url("https://api.groq.com/openai/v1/audio/transcriptions")
.header("Authorization", "Bearer $apiKey")
.post(requestBody)
.build()
return try {
val response = client.newCall(request).execute()
if (response.isSuccessful) {
val json = JSONObject(response.body?.string() ?: "{}")
json.optString("text", null)
} else {
println("Groq error: ${response.code} ${response.body?.string()}")
null
}
} catch (e: IOException) {
println("Network error: ${e.message}")
null
}
}
}
Usage
// In a coroutine or background thread
val recorder = AudioRecorder()
recorder.startRecording()
// ... user speaks ...
val pcmData = recorder.stopRecording()
val wavData = pcmToWav(pcmData)
val client = GroqWhisperClient("your-api-key-here")
val transcription = client.transcribe(wavData)
println("User said: $transcription")
Step 3: Error Handling
Production code needs to handle:
Rate Limits
Groq free tier has limits. Handle 429 responses:
if (response.code == 429) {
val retryAfter = response.header("Retry-After")?.toIntOrNull() ?: 60
println("Rate limited. Retry after $retryAfter seconds")
// Show user-friendly error or queue for retry
}
Network Failures
Mobile networks are unreliable. Implement retry with exponential backoff:
suspend fun transcribeWithRetry(wavData: ByteArray, maxRetries: Int = 3): String? {
repeat(maxRetries) { attempt ->
val result = client.transcribe(wavData)
if (result != null) return result
if (attempt < maxRetries - 1) {
delay((2.0.pow(attempt) * 1000).toLong()) // 1s, 2s, 4s
}
}
return null
}
Audio Format Validation
Groq rejects invalid audio. Validate before sending:
fun validateAudio(wavData: ByteArray): Boolean {
if (wavData.size < 1000) {
println("Audio too short (less than 0.06 seconds)")
return false
}
if (wavData.size > 25 * 1024 * 1024) {
println("Audio too large (max 25 MB)")
return false
}
return true
}
Optimization Tips
1. Use Connection Pooling
Reuse OkHttpClient instances. Creating a new client for each request wastes time on TLS handshakes.
2. Compress Audio
Groq also accepts MP3, Opus, and other formats. For long recordings, MP3 reduces upload time:
// Use LAME or Android MediaCodec to encode MP3 val mp3Data = encodeMp3(pcmData, sampleRate = 16000, bitrate = 32) // 32kbps is enough for voice
But beware: encoding adds 50-100ms. Only worth it for >10 second recordings.
3. Send Early, Stream Late
Don't wait for the user to finish speaking. Start the API call as soon as VAD detects silence. Overlap network latency with the user's perception delay.
4. Cache Common Phrases
If your app has command phrases ("set a timer", "turn on Wi-Fi"), cache their transcriptions locally. Check the cache before calling the API.
Groq vs Alternatives
| Provider | Latency | Cost | Notes |
|---|---|---|---|
| Groq Whisper | 170-220ms | Free tier + paid | Fastest, best for real-time |
| OpenAI Whisper API | 500-800ms | $0.006/min | Reliable, good accuracy |
| Google Speech-to-Text | 400-600ms | Free tier limited | Good for GCP users |
| On-device whisper.cpp | 1000-2000ms | Free (offline) | Slow, but works offline |
Real-World Performance
In Amalia (open source voice assistant), Groq Whisper handles:
- Short commands ("set a timer") — 170ms average
- Medium phrases (1-2 sentences) — 220ms average
- Long speech (10+ seconds) — 400ms average
Network latency adds 50-150ms depending on location and connection quality. Total STT time: 220-370ms in production.
Common Pitfalls
1. Audio Format Mismatches
Groq is strict about formats. If you get a 400 error, check your WAV header. Use a tool like ffprobe to validate locally before debugging the API.
2. Blocking the Main Thread
Network calls block. Always run API calls in a coroutine or background thread:
lifecycleScope.launch(Dispatchers.IO) {
val transcription = client.transcribe(wavData)
withContext(Dispatchers.Main) {
// Update UI
}
}
3. Not Handling Silence
If the user doesn't speak, Whisper returns an empty string. Check for that before processing:
if (transcription.isNullOrBlank()) {
println("No speech detected")
return
}
Full Implementation Example
See Amalia's source code for a complete production implementation. The relevant files:
AudioRecorder.kt— Recording with VADGroqClient.kt— API wrapper with retry logicVoicePipeline.kt— Full STT → LLM → TTS flow
Next Steps
Once you have Groq Whisper working, the next bottleneck is the LLM. To keep latency low:
- Use Groq LPU for LLM inference too (Qwen3, Llama, etc.)
- Stream the LLM response and start TTS early
- Use structured output (JSON mode) to avoid parsing delays
Full voice assistant guide: how to build a voice assistant
Related: building a sub-second voice assistant · open source voice assistants · Android audio APIs (Russian)