Skip to content
WM KeyboardWM Keyboard
Accessibility

Voice typing

Dictation through your device's speech engine: a full panel, a compact bar over the keys, or a draggable bar with the keyboard out of the way.

Voice typing turns speech into text using your device’s own speech recognizer: no setup, works in any language your device supports. Tap the mic, talk, and words appear as you speak.

Screenshot pending
The voice panel: a pulsing mic button on the left, the handwriting-style action rail on the right.

Open Voice typing from the toolbar (or the toolbox if you haven’t pinned it). What opens depends on Voice typing view, covered below. There are three of them:

  • Full panel (the default) replaces the whole keyboard with a large mic button and a level-driven pulse ring, plus a four-key action rail down the right side: Delete, Space, a context-aware Enter (it shows a search icon in a search box instead of promising a newline it won’t insert), and a key that closes the panel and returns to normal typing.
  • Compact bar over the keys replaces the suggestion strip with a small mic and status text, while the keys stay visible underneath: dictate a sentence, then tap a letter key to fix a misheard word without leaving the panel-free view.
  • Collapsed bar, keyboard hidden hands the keyboard’s whole window back to the app you’re typing in and leaves a small draggable pill of dictation controls floating over it. See The collapsed bar.

You aren’t locked into the one you picked. The panel and the compact bar each carry a collapse button down to the pill, and the pill carries a button back up, so switching is one tap from wherever you are.

Tap the mic to start listening; tap again to finish. In the full panel only, you can also long-press the mic for walkie-talkie dictation: holding it past about 600 ms starts listening only while your finger stays down, and releasing stops it immediately. A quick tap-and-release instead toggles listening on, same as before. Neither bar has this. Both mics are plain tap-to-toggle buttons with no hold gesture.

Hold the mic in the full panel to dictate walkie-talkie style.

Recognized words stream into the text field live as you talk (this is the system recognizer only; see Offline voice for how Whisper’s panel differs). WM Keyboard also works out the spacing around what you dictate from the characters already next to the cursor, so starting mid-sentence or dictating over a selection doesn’t leave a double space or a missing one.

Every word you dictate is learned the same way a typed word is, so frequent dictated vocabulary starts showing up in suggestions and autocorrect too.

  • Password and other secure fields never get a mic. Tapping the voice tool in one shows a brief “You cannot use voice typing in a password field” toast and opens nothing. If you land in a password field with a view already open, the panel reads “Voice typing does not work in password fields.” and the compact bar reads “Not available in password fields”.
  • No microphone permission shows an explanation and an “Allow microphone” button. Tapping it opens a system permission prompt through a small trampoline screen (IMEs can’t show permission dialogs directly), and the keyboard rechecks automatically once you’re back. The bars have room for one line, so they say “The app needs microphone permission” with an Allow chip.
  • No speech recognizer on the device (rare, typically no Google app installed) shows “Speech recognition is not available on your device.” in the panel, or “Speech recognition is not available” on either bar.

The third surface isn’t a keyboard at all. Pick Collapsed bar, keyboard hidden, or tap the collapse button in the panel or the compact bar, and the keyboard gives its whole window back to the app you’re typing in. What’s left is a small pill of dictation controls floating on top. Only the pill itself is touchable: taps anywhere around it go through to the app underneath, so you can scroll the thing you’re dictating into while you talk.

Screenshot pending
The collapsed bar floating over an app, with the keyboard out of the way.

Drag it wherever you want it. Where it settles is remembered, so it comes back in the same place next time:

  • Lying flat, it snaps to the nearest of three rests across the screen (left, centre or right) and keeps whatever height you dragged it to. It starts centred and docked at the bottom.
  • Standing upright, it docks to whichever side edge you let go nearest, and keeps its position along that edge. It starts on the right, halfway down.

The hamburger swaps the row of controls for a second page: back, undo, the language chip, a button that stands the bar upright (or lays it flat again), and a shortcut into the voice settings. The first page has the status line, undo when there’s something to undo, an exit button, backspace, and the mic.

That exit button is two different buttons depending on how you got here, which is worth knowing because they don’t do the same thing:

  • Arrived by collapsing from the panel or the compact bar, you get an expand button that takes you back to the surface you came from. It undoes the switch.
  • Chose the collapsed bar in settings, and you get a keyboard button instead. That brings the keys back for now, and the bar stays your default for the next time you open voice typing.

The bar isn’t just for one field. It stands in for the keyboard on every new field until you bring the keyboard back, and it survives the keyboard’s process being killed in between. Two things put the keys back on their own without disarming it: a password field, which never gets the bar, and a panel opened over the top (a hardware shortcut can do that). Both end the current dictation session, and closing the panel brings the bar back.

Screenshot pending
The Undo control, shown after a phrase commits.

An Undo control appears (bottom-left in the panel, as a small icon on either bar) right after a phrase commits, as long as it’s still sitting immediately before your cursor. Tapping it removes exactly that dictated text in one step. It disappears again once you dictate something else, edit the text some other way, or the text it would remove is no longer where dictation left it.

WM KeyboardToolsVoice typingKeep listening

With Keep listening on (the default), finishing one sentence doesn’t stop dictation: the next listening session starts automatically so you can keep talking sentence after sentence without re-tapping the mic. Turn it off and each session ends after one utterance; you tap the mic again for the next one.

Continuous mode is also patient with silence: a pause that would normally time out the recognizer just restarts listening quietly instead of surfacing an error, but only for two silent retries in a row. After that it gives up and goes idle, so an abandoned open mic doesn’t sit there listening forever.

Dictation stops for good (not just pausing) the moment you type a key, swipe, or tap a suggestion. Partial results build up as one continuous chunk, so typing in the middle of an utterance would corrupt it rather than just interrupting it; whichever view you’re in stays open afterward so you can tap the mic to resume. It also stops when you close the panel, dismiss either bar, or switch to a different tool panel.

WM KeyboardToolsVoice typingSpoken punctuation

With this on (the default), saying certain words types the punctuation mark instead of the word itself. It only applies to finished phrases, never to the live partial text, and only recognizes two language groups: English and Bengali.

Say (English)TypesSay (Bengali)Types
“new paragraph”two line breaks“নতুন প্যারা”two line breaks
“new line”line break“নতুন লাইন”line break
“question mark”?“প্রশ্নবোধক চিহ্ন” / “প্রশ্নবোধক”?
“exclamation mark” / “exclamation point”!“বিস্ময়সূচক চিহ্ন” / “বিস্ময়বোধক”!
“full stop” / “period”.“দাঁড়ি”
“comma”,“কমা”,
“colon”:“কোলন”:
“semicolon”;“সেমিকোলন”;

Punctuation attaches straight to the word before it with no space. The trade-off is deliberate: turn this off if you actually want to say the word “period” or “comma” out loud rather than typing the mark. A sentence that genuinely uses one of these words as a word will type the symbol instead while this is on.

Voice typing doesn’t have its own language picker: it dictates in whatever language your active keyboard layout is set to. If you have both English and at least one non-English language enabled, a small EN / বাং chip appears in the corner of the panel; tapping it switches your active layout to the other language’s first enabled layout, and dictation follows along. With only one language enabled, or with several non-English languages and no English one, no chip appears.

The system recognizer prefers running fully on-device (Android 12 and up) when your phone’s model for that language is installed: it’s faster, needs no network, and skips the little beep before it starts listening. If a language’s on-device model can’t handle recognition, WM Keyboard falls back to the network recognizer transparently.

On Android 13 and up, if an on-device model for your language is available but not installed, a chip offers to download it, named for the language you are actually dictating in: “Get Spanish for offline dictation”, “Get Bengali for offline dictation”, and so on. A dictation language the catalogue doesn’t recognise reads “Get this language for offline dictation” rather than guessing. Tapping it starts the system’s own download and shows progress until it’s installed.

Screenshot pending
Offline model download chip in the voice panel.
WM KeyboardToolsVoice typing
SettingDefaultWhat it does
Voice typing viewFull panelA choice of three, not a switch: which surface the mic button opens. Full panel, Compact bar over the keys, or Collapsed bar, keyboard hidden.
Keep listeningOnStarts the next sentence automatically after each one commits.
Spoken punctuationOnSaying “comma”, “question mark” or “দাঁড়ি” types the mark.

All three sit in a group headed Voice typing, on the Voice typing tool’s own screen. An edition with the offline engine built in adds an Engine group above it holding the Recognition engine choice (see Offline voice (Whisper) for what that adds). With the system recognizer in use, the screen ends with this note about what it sends:

Recognition uses the speech recognizer of your device. On Android 12 and later it runs on your device when the language model is installed. If it is not, the audio goes to the recognizer service while you speak. Press and hold the microphone to speak in walkie-talkie style. It stops when you let go.

On-device recognition (Android 12+, once the language’s model is installed) never sends audio anywhere. Without that model installed, dictating with the system recognizer sends your audio to the OS’s speech recognition service while you talk (the same behavior any app using Android’s standard speech APIs has). See Network policy for the fuller picture of what WM Keyboard can and can’t send.

If you want a hard guarantee that audio never leaves the device regardless of which model is installed, Offline voice typing with Whisper is a second dictation engine that transcribes entirely on-device; it’s a bigger download and doesn’t stream words live the way the system recognizer does, but it makes no network calls at all beyond the one-time model download.

Errors show specific messages. A network problem says “No connection. Check your network and try again.”; the recognizer being busy (or another app holding the mic) says “The microphone is busy. Close the other apps that use it and try again.”; a language the device can’t recognise at all says “Speech recognition does not support this language on your device.” A network hiccup mid-utterance keeps whatever was already heard instead of throwing it away.

Switching tools closes dictation. Opening any other panel (Handwriting, Clipboard, Emoji, and so on) ends the current dictation session first.

All three rows carry the usual restore button once you move one off the default listed above; see Putting one setting back. Where the collapsed bar last rested isn’t one of those rows: it has no control on the settings screen, and you reset it by dragging the bar somewhere else.

The three views aren’t full equivalents. Continuous mode, spoken punctuation, and undo work the same way whichever one you’re in: those live in the dictation session itself, not the view. Two things are panel-only, though: hold-to-talk, and the offline-model download chip. The EN / বাং language-switch chip is on the panel and, behind the hamburger, on the collapsed bar, but not on the compact bar. Dictation still follows your active layout’s language there, you just switch layouts some other way (spacebar swipe, the language picker) instead of tapping a chip. Neither bar prompts you to download an offline model.

Related: Offline voice (Whisper) covers the second dictation engine, including its model catalog, per-language routing, downloads, and the full offline privacy guarantee.