Groq Whisper API Integration for Android Apps

October 6, 2026 · developer guide · 8 minutes read

Groq's Whisper v3 Turbo is the fastest production speech-to-text API available in 2026. It runs on custom LPU silicon and transcribes audio in 170-220ms — 3x faster than standard cloud STT services. This guide shows how to integrate it into an Android app.

Why Groq Whisper for Android

Three reasons:

On-device Whisper (whisper.cpp) takes 1-2 seconds on flagship phones. Cloud Whisper on standard infrastructure takes 500-800ms. Groq LPU does it in 220ms.

Prerequisites

You'll need:

Step 1: Record Audio

Groq Whisper expects audio in specific formats. The safest bet is WAV, 16kHz, mono, 16-bit PCM.

Using AudioRecord

import android.media.AudioFormat
import android.media.AudioRecord
import android.media.MediaRecorder
import java.io.ByteArrayOutputStream

class AudioRecorder {
    private val sampleRate = 16000
    private val channelConfig = AudioFormat.CHANNEL_IN_MONO
    private val audioFormat = AudioFormat.ENCODING_PCM_16BIT
    private val bufferSize = AudioRecord.getMinBufferSize(
        sampleRate, channelConfig, audioFormat
    )
    
    private var audioRecord: AudioRecord? = null
    private val audioBuffer = ByteArrayOutputStream()
    
    fun startRecording() {
        audioRecord = AudioRecord(
            MediaRecorder.AudioSource.VOICE_RECOGNITION,
            sampleRate,
            channelConfig,
            audioFormat,
            bufferSize
        )
        audioRecord?.startRecording()
        
        // Read audio in a background thread
        val buffer = ShortArray(bufferSize / 2)
        while (isRecording) {
            val read = audioRecord?.read(buffer, 0, buffer.size) ?: 0
            if (read > 0) {
                // Convert shorts to bytes and write
                buffer.take(read).forEach { sample ->
                    audioBuffer.write(sample.toInt() and 0xFF)
                    audioBuffer.write((sample.toInt() shr 8) and 0xFF)
                }
            }
        }
    }
    
    fun stopRecording(): ByteArray {
        audioRecord?.stop()
        audioRecord?.release()
        return audioBuffer.toByteArray()
    }
}

Convert to WAV Format

Raw PCM needs a WAV header. Here's a minimal WAV writer:

fun pcmToWav(pcmData: ByteArray, sampleRate: Int = 16000): ByteArray {
    val output = ByteArrayOutputStream()
    val channels = 1
    val bitsPerSample = 16
    
    // RIFF header
    output.write("RIFF".toByteArray())
    output.write(intToBytes(36 + pcmData.size))
    output.write("WAVE".toByteArray())
    
    // fmt chunk
    output.write("fmt ".toByteArray())
    output.write(intToBytes(16)) // chunk size
    output.write(shortToBytes(1)) // audio format (PCM)
    output.write(shortToBytes(channels))
    output.write(intToBytes(sampleRate))
    output.write(intToBytes(sampleRate * channels * bitsPerSample / 8))
    output.write(shortToBytes(channels * bitsPerSample / 8))
    output.write(shortToBytes(bitsPerSample))
    
    // data chunk
    output.write("data".toByteArray())
    output.write(intToBytes(pcmData.size))
    output.write(pcmData)
    
    return output.toByteArray()
}

fun intToBytes(value: Int) = byteArrayOf(
    (value and 0xFF).toByte(),
    ((value shr 8) and 0xFF).toByte(),
    ((value shr 16) and 0xFF).toByte(),
    ((value shr 24) and 0xFF).toByte()
)

fun shortToBytes(value: Int) = byteArrayOf(
    (value and 0xFF).toByte(),
    ((value shr 8) and 0xFF).toByte()
)

Step 2: Call Groq Whisper API

Groq uses OpenAI-compatible API. Send a multipart form with the audio file.

API Client

import okhttp3.*
import okhttp3.MediaType.Companion.toMediaType
import org.json.JSONObject
import java.io.IOException

class GroqWhisperClient(private val apiKey: String) {
    private val client = OkHttpClient.Builder()
        .connectTimeout(10, TimeUnit.SECONDS)
        .readTimeout(30, TimeUnit.SECONDS)
        .build()
    
    fun transcribe(wavData: ByteArray): String? {
        val requestBody = MultipartBody.Builder()
            .setType(MultipartBody.FORM)
            .addFormDataPart(
                "file",
                "audio.wav",
                wavData.toRequestBody("audio/wav".toMediaType())
            )
            .addFormDataPart("model", "whisper-large-v3-turbo")
            .addFormDataPart("language", "en") // optional, auto-detect if omitted
            .addFormDataPart("response_format", "json")
            .build()
        
        val request = Request.Builder()
            .url("https://api.groq.com/openai/v1/audio/transcriptions")
            .header("Authorization", "Bearer $apiKey")
            .post(requestBody)
            .build()
        
        return try {
            val response = client.newCall(request).execute()
            if (response.isSuccessful) {
                val json = JSONObject(response.body?.string() ?: "{}")
                json.optString("text", null)
            } else {
                println("Groq error: ${response.code} ${response.body?.string()}")
                null
            }
        } catch (e: IOException) {
            println("Network error: ${e.message}")
            null
        }
    }
}

Usage

// In a coroutine or background thread
val recorder = AudioRecorder()
recorder.startRecording()

// ... user speaks ...

val pcmData = recorder.stopRecording()
val wavData = pcmToWav(pcmData)

val client = GroqWhisperClient("your-api-key-here")
val transcription = client.transcribe(wavData)

println("User said: $transcription")

Step 3: Error Handling

Production code needs to handle:

Rate Limits

Groq free tier has limits. Handle 429 responses:

if (response.code == 429) {
    val retryAfter = response.header("Retry-After")?.toIntOrNull() ?: 60
    println("Rate limited. Retry after $retryAfter seconds")
    // Show user-friendly error or queue for retry
}

Network Failures

Mobile networks are unreliable. Implement retry with exponential backoff:

suspend fun transcribeWithRetry(wavData: ByteArray, maxRetries: Int = 3): String? {
    repeat(maxRetries) { attempt ->
        val result = client.transcribe(wavData)
        if (result != null) return result
        
        if (attempt < maxRetries - 1) {
            delay((2.0.pow(attempt) * 1000).toLong()) // 1s, 2s, 4s
        }
    }
    return null
}

Audio Format Validation

Groq rejects invalid audio. Validate before sending:

fun validateAudio(wavData: ByteArray): Boolean {
    if (wavData.size < 1000) {
        println("Audio too short (less than 0.06 seconds)")
        return false
    }
    if (wavData.size > 25 * 1024 * 1024) {
        println("Audio too large (max 25 MB)")
        return false
    }
    return true
}

Optimization Tips

1. Use Connection Pooling

Reuse OkHttpClient instances. Creating a new client for each request wastes time on TLS handshakes.

2. Compress Audio

Groq also accepts MP3, Opus, and other formats. For long recordings, MP3 reduces upload time:

// Use LAME or Android MediaCodec to encode MP3
val mp3Data = encodeMp3(pcmData, sampleRate = 16000, bitrate = 32) // 32kbps is enough for voice

But beware: encoding adds 50-100ms. Only worth it for >10 second recordings.

3. Send Early, Stream Late

Don't wait for the user to finish speaking. Start the API call as soon as VAD detects silence. Overlap network latency with the user's perception delay.

4. Cache Common Phrases

If your app has command phrases ("set a timer", "turn on Wi-Fi"), cache their transcriptions locally. Check the cache before calling the API.

Groq vs Alternatives

Provider Latency Cost Notes
Groq Whisper 170-220ms Free tier + paid Fastest, best for real-time
OpenAI Whisper API 500-800ms $0.006/min Reliable, good accuracy
Google Speech-to-Text 400-600ms Free tier limited Good for GCP users
On-device whisper.cpp 1000-2000ms Free (offline) Slow, but works offline

Real-World Performance

In Amalia (open source voice assistant), Groq Whisper handles:

Network latency adds 50-150ms depending on location and connection quality. Total STT time: 220-370ms in production.

Common Pitfalls

1. Audio Format Mismatches

Groq is strict about formats. If you get a 400 error, check your WAV header. Use a tool like ffprobe to validate locally before debugging the API.

2. Blocking the Main Thread

Network calls block. Always run API calls in a coroutine or background thread:

lifecycleScope.launch(Dispatchers.IO) {
    val transcription = client.transcribe(wavData)
    withContext(Dispatchers.Main) {
        // Update UI
    }
}

3. Not Handling Silence

If the user doesn't speak, Whisper returns an empty string. Check for that before processing:

if (transcription.isNullOrBlank()) {
    println("No speech detected")
    return
}

Full Implementation Example

See Amalia's source code for a complete production implementation. The relevant files:

Next Steps

Once you have Groq Whisper working, the next bottleneck is the LLM. To keep latency low:

Full voice assistant guide: how to build a voice assistant

Related: building a sub-second voice assistant · open source voice assistants · Android audio APIs (Russian)