All tools

OCR — make a scan searchable

Turn a scanned or photographed PDF into real text you can select, search and edit.

Adds an invisible text layer over the page image, so the document looks exactly the same but the words become real text. Pages that already have text are left alone.

Options

Processed and handed straight back. Nothing is written to disk.

About ocr — make a scan searchable

A scanned page is a photograph of a document, not a document. That is why you cannot select the words, search for a name, or copy a number out of it — as far as any program is concerned the page is just a picture. Text recognition fixes that by reading the picture and writing what it finds back into the file as an invisible text layer sitting exactly over the words.

The page looks completely unchanged afterwards. Nothing is redrawn and no quality is lost; the image you scanned is still the image you see. What changes is that the words are now real text, so you can select them, search the document, copy a figure into a spreadsheet, and edit the page in the editor like any other PDF.

Pages that already contain text are passed through untouched. Recognising a page that was created digitally would be a step backwards, because it would replace crisp, exact text with a program’s best guess at a picture of text.

Questions about ocr — make a scan searchable

Will this change how my document looks?

No. The scanned image is left exactly as it is. The recognised text is added as an invisible layer on top, so the page prints and displays identically — it simply becomes selectable and searchable.

How accurate is it?

On a clean, straight scan of ordinary printed text it is very good. Accuracy drops with low resolution, skewed or curled pages, shadows from phone photos, handwriting, unusual fonts and heavy background patterns. Always check important numbers, and expect the occasional word boundary in the wrong place.

Can I edit a scan after running this?

Yes, and that is the main reason to do it. Once the text layer exists you can open the document in the editor, click a word and retype it. The editor will offer to run recognition for you when it notices you have opened a scan.

Which languages are supported?

Whichever language packs are installed on the server. English is always included; others can be added by installing more Tesseract language data.

Why does it take a few seconds?

Every page has to be rendered at high resolution and then read character by character. A few seconds per page is normal, and larger or higher-quality settings take longer.