Before you start

  1. Complete the Ascuta setup wizard and allow microphone access.
  2. Choose an engine: a downloaded on-device model, the system speech engine, or your own cloud provider.
  3. If you want the floating bubble, enable Ascuta under Android Accessibility settings.
  4. Test in a normal text field, such as a messaging or notes app, before using it in your main apps.

Ascuta only works with a keyboard’s microphone when that keyboard asks Android for standard voice input. Most keyboards keep their own microphone button, so Ascuta also offers two keyboard-independent ways to dictate: the floating bubble and the Ascuta Keyboard.

Gboard

Use the floating bubble, or switch to the Ascuta Keyboard

Gboard’s microphone button opens Google voice typing and does not hand it to another recognizer. Choose one of these:

  1. Floating bubble (keeps Gboard): leave Gboard selected, focus a text field, tap the Ascuta bubble above the app, speak, then tap stop. Ascuta inserts the text into the field.
  2. Ascuta Keyboard: enable it in Android keyboard settings, switch to it with the globe key, and use its microphone. It commits the transcript and normally returns to Gboard.

You never have to replace Gboard permanently: Ascuta is an auxiliary voice keyboard that you switch to only when you want to dictate.

Microsoft SwiftKey

SwiftKey’s microphone can open Ascuta

SwiftKey invokes Android’s standard speech-recognition interface, which Ascuta can serve:

  1. Focus a text field and tap the microphone button on the SwiftKey toolbar.
  2. If Android asks which voice-input service to use, choose Ascuta. If a panel opens directly, it is the Ascuta voice-input popup.
  3. Speak, then stop. Ascuta returns one result to SwiftKey and the focused field.

This is the most direct keyboard-microphone integration. If SwiftKey’s microphone is bound to another service, use the bubble or the Ascuta Keyboard instead.

Samsung Keyboard

Use the floating bubble, or switch to the Ascuta Keyboard

Samsung Keyboard keeps its microphone connected to Samsung voice input. Use the same two paths as with Gboard:

  1. Keep Samsung Keyboard selected and tap the Ascuta floating bubble, or
  2. Switch to the Ascuta Keyboard with the globe key and tap its microphone.

The bubble is convenient when you want to keep Samsung Keyboard visible; the Ascuta Keyboard is useful when you want the text to appear in the field as you dictate.

Ascuta Keyboard

Enable it once, switch to it when needed

  1. Open Ascuta and tap Voice typing keyboard, or open Android keyboard settings manually.
  2. Enable Ascuta voice typing.
  3. Focus a text field in any app and open the keyboard switcher (globe or keyboard button), then select Ascuta.
  4. Tap the microphone to start and tap again to stop.
  5. Ascuta commits the final transcript into the focused field and normally switches back to your previous keyboard.

This is an auxiliary keyboard, meant to be selected temporarily. With the system speech engine, partial words appear while you speak; other engines insert the finished result when recording stops. If the switcher does not show Ascuta, confirm it is enabled and restart the target app.

Floating bubble

Works independently of your keyboard

  1. Enable the Ascuta Accessibility Service when Android asks.
  2. Leave Gboard, SwiftKey, Samsung Keyboard, or any other keyboard selected.
  3. Focus a normal text field and tap the floating Ascuta button. Press and hold it to dictate, and release to finish.
  4. Speak, then tap stop. Ascuta inserts the result into the focused field.

Some apps use secure or custom editors that block accessibility text insertion. In that case Ascuta copies the transcript to the clipboard so you can paste it manually.

Use Ascuta when an app asks Android for voice input

Some apps and websites open Android’s standard speech-recognition interface instead of using a keyboard-specific voice feature. When that happens, Android can show the Ascuta voice-input popup.

  1. Tap the app or website’s voice-search or microphone control.
  2. Choose Ascuta if Android asks which voice-input service to use.
  3. Record and stop in the Ascuta popup.
  4. Ascuta returns one final result to the requesting app.

This flow is controlled by the calling app. If it does not use Android’s standard request, use the bubble or the Ascuta Keyboard instead.

Other keyboards

Any keyboard: use the bubble or the Ascuta Keyboard

HeliBoard, OpenBoard, Fleksy, Grammarly, Yandex Keyboard, Typewise and others all work the same way as Gboard and Samsung Keyboard:

  1. Keep your keyboard and dictate with the floating bubble, or
  2. Switch to the Ascuta Keyboard when you want the words to land in the field as you speak.

If your keyboard exposes Android’s standard voice-input action (often a microphone key or a long-press option), it may also open the Ascuta popup directly.

Voice models

Pick an engine, then download one model for it

Everything below except cloud runs fully on your phone after a one-time download. Sizes are approximate; use Wi-Fi for the large ones.

  1. System (ML Kit): nothing to download. Streams with live words while you speak, using the phone’s built-in speech model where available.
  2. Whisper — Tiny (~111 MB): fastest, for older phones. English-only or multilingual variant.
  3. Whisper — Base (~198 MB, recommended): the balanced default. English-only or multilingual variant.
  4. Whisper — Small (~610 MB): most accurate Whisper, for difficult audio. English-only or multilingual variant.
  5. Parakeet TDT 0.6B v3 (~465 MB): best quality for 25 European languages with auto-detect.
  6. Parakeet TDT 0.6B v2 (~450 MB): great quality, English only.
  7. Parakeet TDT-CTC 110M (~100 MB): fastest, English only.
  8. Nemotron 3.5 streaming (~475 MB): live words while you speak, 32 languages with auto-detect.
  9. Gemma transcription (~2.6–4.9 GB): accurate on-device audio models (Gemma 4 E2B/E4B open; Gemma 3n gated behind a Hugging Face token). Best for short clips of about 30 seconds; needs a powerful phone (see below).
  10. Cloud (OpenAI-compatible): no download. Uses your own API key; audio goes straight from your phone to that provider — Ascuta runs no relay server.

Short on space? Start with Whisper Base or Parakeet CTC. You can switch engines anytime without losing the others.

Post-processing models

Polish the transcript after dictation

Post-processing is a Lifetime Pro feature. Two paths:

  1. On-device (Gemma 4 E2B ~2.6 GB recommended, or E4B ~3.7 GB): open downloads, work offline, use your editable prompt. The gated Gemma 3n variants additionally need a Hugging Face token.
  2. Cloud (OpenAI-compatible): uses your own API key and your editable prompt.

Post-processing actions

Combine them; a custom prompt overrides all

  1. Cleanup: fixes punctuation, capitalization, spelling, and clear grammar slips — staying in the transcript’s language, never paraphrasing.
  2. Summarize: condenses the transcript to its essentials, in the same language.
  3. Voice commands: applies spoken formatting and editing instructions (“delete the last sentence”, “make it formal”) during cleanup.
  4. Custom prompt: your own instruction replaces the three actions above entirely.

Cleanup and summarize can run together; voice commands apply on top. If the result looks invented rather than edited, Ascuta falls back to the raw transcript.

Agent mode

Dictate instructions instead of words

With Agent mode on, your speech is treated as an instruction to carry out, not text to write down. Say “reply to this message saying I’ll be ten minutes late” and Ascuta inserts the finished reply — no greeting, no explanation, ready to send.

  1. Toggle it by double-tapping the microphone on the bubble, the Ascuta Keyboard, or the voice-input popup. The mic is tinted while Agent mode is on. Each surface remembers its own toggle.
  2. Starting behavior comes from settings: start every dictation in Agent mode, or pick up where you left off with “keep last mode”.
  3. Screen context (bubble only): while you record, Ascuta reads the visible screen and summarizes it — its language and what it is about. Your reply then fits the conversation: same language, same names and details, continuing the current topic. Turn it off in settings to dictate without screen context.

Agent mode needs Lifetime Pro and a selected post-processing model, since the same on-device or cloud model executes your instruction. Without a readable screen it simply runs the instruction on its own.

Hardware & storage

What your phone needs

  1. Android 11 or newer.
  2. Small models (Whisper Tiny/Base, Parakeet CTC, system engine) run on most phones.
  3. Large models (Whisper Small, Parakeet 0.6B, Nemotron) want a mid-range phone and ~500–600 MB of free space each.
  4. Gemma models need 8 GB or more of RAM and 2.6–4.9 GB of free space. Loading takes about 10 seconds and the model stays cached for fast repeats.
  5. Idle footprint: models unload automatically when idle (configurable: after each dictation, after 5 minutes, or kept loaded), so they cost no RAM when you are not dictating.

Supported languages

Coverage depends on the engine you pick

  1. Whisper multilingual (Tiny / Base / Small): up to 99 languages with auto-detect — the full list is below.
  2. Whisper English-only variants: English. Smaller download when you only dictate in English.
  3. Parakeet TDT 0.6B v3: 25 European languages with auto-detect: Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hungarian, Italian, Latvian, Lithuanian, Maltese, Polish, Portuguese, Romanian, Slovak, Slovenian, Spanish, Swedish, Russian, Ukrainian.
  4. Parakeet TDT 0.6B v2 and TDT-CTC 110M: English only.
  5. Nemotron streaming: 32 locales with auto-detect: English, Spanish, French, Italian, Portuguese, Dutch, German, Turkish, Russian, Arabic, Hindi, Japanese, Korean, Vietnamese, Ukrainian, Polish, Swedish, Czech, Norwegian, Danish, Bulgarian, Finnish, Croatian, Slovak, Mandarin, Hungarian, Romanian, Estonian.
  6. System (ML Kit): follows the app Language setting — about two dozen locales including English, French, Italian, German, Spanish, Hindi, Japanese, Portuguese, Turkish, Polish, Chinese, Korean, Russian, Vietnamese, Dutch, Danish, Swedish, Thai, Indonesian, and Arabic. Anything else falls back to US English.
  7. Gemma transcription and cloud mode: multilingual; coverage follows the underlying model or provider (OpenAI Whisper covers up to 99 languages).

Tip: set Language explicitly instead of Auto-detect for the most consistent results, especially with live transcription.

Whisper multilingual covers: English, Chinese, German, Spanish, Russian, Korean, French, Japanese, Portuguese, Turkish, Polish, Catalan, Dutch, Arabic, Swedish, Italian, Indonesian, Hindi, Finnish, Vietnamese, Hebrew, Ukrainian, Greek, Malay, Czech, Romanian, Danish, Hungarian, Tamil, Norwegian, Thai, Urdu, Croatian, Bulgarian, Lithuanian, Latin, Maori, Malayalam, Welsh, Slovak, Telugu, Persian, Latvian, Bengali, Serbian, Azerbaijani, Slovenian, Kannada, Estonian, Macedonian, Breton, Basque, Icelandic, Armenian, Nepali, Mongolian, Bosnian, Kazakh, Albanian, Swahili, Galician, Marathi, Punjabi, Sinhala, Khmer, Shona, Yoruba, Somali, Afrikaans, Occitan, Georgian, Belarusian, Tajik, Sindhi, Gujarati, Amharic, Yiddish, Lao, Uzbek, Faroese, Haitian Creole, Pashto, Turkmen, Nynorsk, Maltese, Sanskrit, Luxembourgish, Myanmar, Tibetan, Tagalog, Malagasy, Assamese, Tatar, Hawaiian, Lingala, Hausa, Bashkir, Javanese, Sundanese.

Troubleshooting

No microphone or no words

Check Android microphone permission, close any other app using the microphone, and retry from Ascuta’s test screen.

The bubble does nothing

Confirm the Accessibility Service is enabled and that a standard text field is focused. Some secure or custom fields reject insertion.

No local transcription

Open the Models tab and verify a model for the selected engine is installed. Large Gemma models take longer to load.

Cloud transcription fails

Check the configured API key, provider endpoint, network connection, account limits, and the provider’s request format.

For privacy details, see the Privacy Policy.