Skip to content
WM KeyboardWM Keyboard
Accessibility

Scanner (OCR, QR, documents)

Point the camera at text, codes, or paper and get the result in your field.

Three camera-driven tools live under one Scanners group. Text scan (OCR) reads printed words into your field. QR and barcode scanner decodes a code and inserts or opens its value. Document scanner captures a paper document as a cropped, cleaned-up image. Full edition only: all three only ship in the full build. The QR and document scanners run on Google ML Kit; text scan reads with ML Kit or with Tesseract, depending on the language.

Screenshot pending
Text scan (OCR): a full-bleed camera view with recognized words ready to tap.

Open Text scan (OCR) from the toolbar or toolbox to replace the keyboard with a full-bleed camera view. It covers the toolbar row too, not only the key area. Point it at printed text and it recognizes words on-device. A torch toggle appears if your device has a flash.

Two engines read the text, both on your device:

  • ML Kit reads Latin-script text only: English, French, Spanish and the like. It needs no download.
  • Tesseract reads more than 100 languages in most scripts, for example Bangla, Hindi, Arabic, Russian, Greek, Thai, Chinese, Japanese and Korean. Each language needs a one-time download of its data, from 0.4 MB to 12 MB.

With Text recognition on Automatic (the default), Latin-script languages read with ML Kit and every other script reads with Tesseract. The scanner reads in the language you’re typing in when you open it. If you’ve turned on more than one language, a chip at the bottom left of the viewfinder names the text’s language, and tapping it steps to the next one without switching your keyboard. On Automatic, all your Latin-script languages share one stop on the chip, since ML Kit reads them the same way.

A language without Tesseract data of its own reads with the main one for its script: Maithili and other Devanagari languages read with Hindi’s, and a Latin-script language with no data of its own reads with English’s.

The first time you scan in a language that needs Tesseract data, the viewfinder is replaced by a Download button that says how big the data is. The download shows its progress in the panel and can be cancelled there. With data saver set to ask, the first press explains that it uses mobile data and a second press starts it. Set to block, the button says why it can’t. The data comes from the tessdata_fast 4.1.0 set of the Tesseract project on GitHub and stays in the app’s private storage; you can also download or delete each language on the tool’s settings page, and delete it under Storage.

After a capture, the recognized words appear as tappable chips grouped by line:

Screenshot pending
OCR result: recognized words as chips, grouped by line, ready to copy or insert.
  • With Start with everything selected on (the default), every word starts selected and tapping a chip deselects it, so you can trim the capture down to what you need. Turn it off and chips start empty, so you tap to pick words one at a time instead.
  • Copy and Insert act on whatever’s currently selected.
  • Scan again discards the capture and returns to the viewfinder. A select-all/deselect-all toggle flips every chip at once.

Tesseract reads each photo twice, once with a threshold set from each small area of the photo and once with one threshold for the whole photo, and keeps whichever read found more text. The first copes with uneven light: a shadow across the page, the dark corners of a phone photo, or the torch’s bright spot. The second copes with light text on a dark screen. “No text found” means both reads came back empty. If recognition itself fails, the panel says Could not read instead. If a downloaded language won’t load, the scanner deletes that language’s data and offers the download again.

If your device has no camera, the panel shows a message instead of a blank viewfinder.

QR and barcode scanner is also full-bleed and decodes continuously as you point the camera, with no shutter tap needed. It reads QR codes plus the common barcode formats: Aztec, Codabar, Code 39/93/128, Data Matrix, EAN-8/13, ITF, PDF417, and UPC-A/E, labeled by format in the result.

Screenshot pending
QR and barcode scanner: live camera view with pinch-to-zoom, freezing on the first code it decodes.

The first code it decodes freezes into a result card with Scan again, Copy and Insert. If the decoded value is a plain URL (https://…, http://…, or www.…), an Open button appears above that row. Anything else is text only, so a scanned value like a Wi-Fi network payload inserts as raw text rather than turning into a connect action. Pinch to zoom while scanning, and a small pill reads out the zoom (“1.2x” style) above roughly 1.05x. One particular link gets a little extra recognition on the result card, over on Easter eggs.

All three scanners have their own settings screen, and all three start with the same header rows before anything tool-specific:

  • Turn on: Text scan is in the “recommended” and “show me everything” starter sets that setup lands, so it arrives on unless you picked the simplest answer. Neither QR and barcode scanner nor Document scanner is in any starter set, so setup leaves both switched off unless you pressed Everything on its tools page. Turn them on here.
  • Icon colour: only shown while colorful tool icons is on globally, which it is by default. Text scan is green, QR and barcode teal, Document scanner a light blue.
  • Trigger words: under a Keyword shortcut heading. It’s ocr for Text scan, scan for QR and barcode, and docscan for Document scanner. Type one of them on its own and the suggestion strip offers to open that tool. Commas separate several words, and an empty field means the tool never offers itself.
  • Match capitals: off by default on all three, so a trigger word matches however you capitalise it. Turn it on and only the exact spelling counts.
  • Reset to defaults: joins the group once you’ve edited that tool’s trigger words, and puts them back.
WM KeyboardToolsQR and barcode scannerOptions
  • Insert automatically: off by default. With it on, the scanned code’s text types into your field the instant it’s spotted, with no confirm tap.
  • Vibrate on detection: on by default. A short buzz when a code is spotted.
  • Load link details: off by default. When a scanned value is a web link, it fetches the page’s title and description to show above the result. That needs a network connection. Power Saving mode force-disables this row while its background-network toggle is on.

Recognition for both OCR and the QR/barcode scanner runs entirely on-device, so no photo is uploaded and both work offline once any Tesseract language data they need is downloaded. Link previews are the one exception, since fetching a page’s title needs the network. See network policy for what else does and doesn’t touch the network.

WM KeyboardToolsText scan (OCR)Options
  • Start with everything selected: on by default, described above.
  • Text recognition: Automatic (the default), ML Kit or Tesseract. See which languages it reads. ML Kit ignores the text’s language, so the viewfinder shows no language chip with it. Tesseract reads Latin script too, usually slower than ML Kit.
WM KeyboardToolsText scan (OCR)Tesseract languagesNext release

One row for each Tesseract language the languages you type in need, with its size, a Download button, its progress, or a delete button once it’s on your device. Two languages that read with the same data share one row, for example Norwegian Bokmål and Nynorsk. The list is hidden while Text recognition is on ML Kit.

Document scanner hands the whole flow to Google’s own scanner UI: edge detection, cropping, and cleanup happen there, outside the keyboard, since an IME can’t host that kind of full-screen flow itself. Gallery import is allowed alongside the camera, and pages come back as JPEGs.

Screenshot pending
Document scanner: Google's scan UI for capturing and cropping a page.

Once you’re back in your app, each captured page is inserted exactly like a Camera-tool photo, sent to the field one page at a time. If Google Play services or the scanner module isn’t available on your device, you get a toast saying so instead of a crash.

WM KeyboardToolsDocument scannerOptions
  • Save to the gallery: off by default. Also keeps scanned pages in the WM Keyboard album under Pictures.
  • Camera permission works differently across the three tools. OCR and QR/barcode check for it themselves, and when it isn’t granted they offer an Allow the camera button. That opens the same disclosure-then-system-prompt trampoline the Camera tool uses, and returns you to the panel. The document scanner skips that prompt entirely. Google’s scanner activity manages its own camera permission internally.
  • None of the three scanners work before you’ve unlocked your device once after a restart (no direct-boot support), unlike some other tools that do work at the lock screen.
  • The trigger words in Options are what make a standalone ocr, scan or docscan surface a chip into the matching tool. Smart chips covers that mechanism in full.
  • Open any of the three from Settings → Tools → Scanners, or by typing its keyword as described above.

OCR depends on ML Kit’s text-recognition library and on a Tesseract library built into the app, about 1.7 MB on a 64-bit phone. QR/barcode scanning depends on ML Kit’s barcode-scanning library. Document scanning depends on Google Play services’ document scanner module. All three are compiled into the full build and left out of lite entirely, rather than hidden behind a toggle, so the Scanners group doesn’t appear in the lite toolbox or toolbar at all. See Full vs Lite if you’re deciding between builds.