Speech recognition localization fails when the screen switches language but the recognizer does not. A user can choose Japanese in the app, tap the microphone, and receive an English transcript because the app reused a stale recognizer or accepted the device default. On another phone, the requested language may not be available at all. Fixing this requires one locale policy, capability checks, a restart rule for active recognition, and a typed input fallback.
The implementation below covers that lifecycle on iOS and Android: locale resolution, recognizer setup, language changes, offline use, failure recovery, and tests on installed devices.
Start with an explicit voice input contract
A microphone button is not a language policy. Before calling either platform API, decide what language the feature should recognize and what happens when the platform cannot provide it.
Use the app's effective content locale as the default recognition language. If the product has a separate voice language setting, make that choice visible and give it higher precedence. The device locale can fill an unset preference, but it should not silently override an app language the user selected.
Use this precedence order unless the product has a documented reason to differ:
- A voice language chosen for the current feature.
- The saved app language.
- The device language when the app follows the system.
- A documented source locale as the final fallback.
Resolve that value to a canonical BCP 47 tag such as en-US, ja-JP, or pt-BR. Keep the requested tag separate from the tag the recognizer actually accepts. This distinction matters when an app supports a regional locale but the installed recognition service exposes only a broader language.
The contract also needs a failure outcome. If recognition is unavailable, the text field must still work. A user should never lose the ability to complete a form because a speech service, language model, permission, or network connection is missing.
Model recognition as a locale-bound session
Treat each recognition session as immutable configuration. It should record the requested language, the resolved platform language, whether offline processing is required, and the text revision to which results may apply.
VoiceSession = {
session_id,
input_id,
input_revision,
requested_language,
resolved_language,
mode,
started_at
}
RecognitionResult = {
session_id,
transcript,
is_final,
detected_language,
confidence_if_available
}
The session identity blocks a common race. A user starts dictating in English, changes the app to French, then starts another request before the first callback arrives. The old result must not overwrite the new field. Apply a callback only when its session ID, input ID, revision, and active language still match the current input state.
Do not mutate an active session to another locale. Cancel it, preserve text that the product has already accepted, then create a new session. Platform recognizers often capture language and service configuration when they are constructed or started. Replacing the session makes the lifecycle visible and testable.
Configure iOS from the effective locale
Apple's SFSpeechRecognizer represents a recognizer for a locale. It also reports service availability and whether on-device recognition is supported. Build a small adapter that creates the recognizer only after the app has resolved the current voice language.
struct VoiceConfiguration: Equatable {
let languageTag: String
let requiresOnDeviceRecognition: Bool
}
final class IOSVoiceRecognizer {
private var configuration: VoiceConfiguration?
private var recognizer: SFSpeechRecognizer?
func prepare(_ next: VoiceConfiguration) -> VoiceAvailability {
guard let locale = Locale(identifier: next.languageTag) as Locale? else {
return .unsupportedLanguage
}
let candidate = SFSpeechRecognizer(locale: locale)
guard let candidate else {
return .unsupportedLanguage
}
guard candidate.isAvailable else {
return .temporarilyUnavailable
}
if next.requiresOnDeviceRecognition && !candidate.supportsOnDeviceRecognition {
return .offlineUnavailable
}
configuration = next
recognizer = candidate
return .available
}
}
This application adapter does not promise that the platform will complete a request. Keep permission handling and audio session ownership outside prepare. The adapter answers whether the resolved locale and processing mode can start a request.
Rebuild this adapter when the effective voice language changes. Also invalidate it when an interruption, route change, or service availability change ends the current session. Do not retain a recognizer merely because the microphone view remains mounted.
On-device support needs a product rule. If privacy or offline operation requires local recognition, fail closed when the selected locale cannot run locally and keep typed input available. If server recognition is acceptable, explain the network requirement in the feature's privacy and error copy instead of silently changing modes.
Configure Android without trusting defaults
Android's RecognizerIntent accepts an explicit BCP 47 language through EXTRA_LANGUAGE. Current Android documentation also defines extras for allowed language detection tags, automatic language switching, and an offline preference.
Pass the resolved language on every request. A practitioner question about setting the Android speech recognition language has accumulated because relying on implicit recognizer settings is unreliable across apps and devices.
fun recognitionIntent(config: VoiceConfiguration): Intent {
return Intent(RecognizerIntent.ACTION_RECOGNIZE_SPEECH).apply {
putExtra(
RecognizerIntent.EXTRA_LANGUAGE_MODEL,
RecognizerIntent.LANGUAGE_MODEL_FREE_FORM
)
putExtra(RecognizerIntent.EXTRA_LANGUAGE, config.languageTag)
putExtra(RecognizerIntent.EXTRA_PARTIAL_RESULTS, true)
putExtra(RecognizerIntent.EXTRA_PREFER_OFFLINE, config.preferOffline)
}
}
EXTRA_PREFER_OFFLINE expresses a preference. The installed recognition service determines actual support and behavior. Your adapter must report what happened rather than assuming that one intent flag guarantees offline processing.
Use automatic language switching only when the product expects multilingual speech. Restrict it to languages the app supports and can present correctly. Otherwise, the detector may choose a language that the app cannot format, label, or verify. Keep the detected tag with the result so analytics and debugging can record a category without capturing the transcript.
Installed services differ. A report of Japanese working in an emulator but defaulting to English on a Pixel shows why emulator success is not release evidence. Query available capabilities where the platform permits it, handle missing recognition activities, and test the same signed build on real devices.
Decide when multilingual switching belongs
Single-language input should stay single-language. A French search box in a French app does not need automatic switching merely because the API offers it. Explicit language configuration gives the user predictable recognition and gives QA a finite matrix.
Enable switching when users commonly mix supported languages in one task, such as messaging, travel search, or bilingual notes. Even then, set an allowed language list. Define whether a transcript may contain several detected languages or whether the recognizer may select one language for the whole session.
Show the active voice language before recording. A compact label beside the microphone is enough. If automatic switching is active, describe that mode without promising perfect detection. Offer a manual language choice when detection repeatedly selects the wrong language.
Do not use detected speech language to change the whole app locale. Recognition is evidence about one utterance, not consent to rewrite navigation, stored preferences, or notification settings.
Preserve partial results through cancellation
Partial transcripts improve perceived speed, but they create state problems. A partial result is provisional. Do not submit it, trigger a destructive action, or save it as the final field value until the recognizer marks the result final or the user explicitly accepts it.
Keep three values in the input component:
- The text committed before recognition started.
- The current provisional transcript.
- The final text accepted from the active session.
If the user cancels, return to the committed text unless the product provides an explicit Keep partial action. If permission is denied or recognition fails before a final result, preserve manually typed text. If a locale change cancels the session, do not carry a provisional transcript into the new language session.
This separation also prevents late callbacks from replacing newer typing. Increment the input revision whenever the user edits the field, changes the language, or starts another session. Reject results tied to an older revision.
Give each failure a useful response
Convert platform errors into product states. The user does not need an iOS error code or an Android recognizer constant. The app still needs enough categories to choose a recovery action.
Use categories such as permission denied, unsupported language, service unavailable, offline model unavailable, network unavailable, no speech, cancelled, interrupted, and internal failure. Keep cancellation separate from failure. Treat no speech as a retryable empty result rather than an exception.
For unsupported language, keep typing active and offer another voice language only if the product supports one. For temporary service or network failure, keep the resolved language and allow a retry. For permanent permission denial, show the platform settings route only after the user asks to use voice again.
Prompts, permission rationale, error copy, accessibility labels, and the active language label belong in the ordinary app localization workflow. The microphone and speech permission text follows the same rules as localized app permission prompts and privacy rationale. Speech recognition returns user content. It should never translate or generate the surrounding interface copy.
Verify the whole lifecycle on installed builds
Unit tests can cover locale precedence, tag normalization, adapter configuration, session identity, input revisions, and error mapping. Fake recognizers should emit partial, final, cancelled, delayed, and wrong-session callbacks in deterministic order.
Platform tests must cover behavior that a fake cannot prove:
- Clean install with microphone and speech permissions unset.
- A saved app language that differs from the device language.
- Language change before recording, during recording, and after a partial result.
- Supported and unsupported locale tags.
- Online recognition, airplane mode, and an offline preference.
- Service interruption, cancellation, silence, and rapid restart.
- Automatic switching restricted to the supported language list.
- Typed edits made while an old callback is pending.
- iOS devices with and without on-device support for the selected locale.
- Android phones with different installed recognition services.
- Screen reader focus on the language label, microphone state, transcript, and retry action.
Use synthetic phrases with names, numbers, punctuation, mixed scripts, and short ambiguous words. Do not put transcripts in analytics or crash reports. Operational telemetry can record requested language, resolved language, platform category, completion state, and duration without recording what the person said.
Simulators and emulators are useful for state transitions, but they cannot prove language availability. Release only after the signed build passes on representative physical devices with the app language, recognizer language, connectivity, and processing mode that users will encounter.
Next action
Write the locale precedence function and VoiceSession identity before adding another recognizer call. Then build one vertical slice: choose an app language different from the device language, start voice input, apply partial and final results only to the active session, switch languages, and verify that typing remains available through every failure. Run that slice on one iPhone and two Android devices with different recognition services.
References
- Apple: SFSpeechRecognizer, for locale-bound recognizer construction, service availability, and on-device capability.
- Android: RecognizerIntent, for recognition language tags, automatic switching controls, and offline preference.
- Stack Overflow: Set the Android speech recognition language, for the recurring practitioner need to override recognizer defaults.
- Stack Overflow: Locale differs between emulator and Pixel, for a device-specific wrong-language report that motivates physical device verification.
