Skip to content
ALL PDF TOOLS37
TESSERACT · PDF-LIBTOOL 181 OF 190

OCR a PDF — Make a Scanned Document Searchable

Adds an invisible text layer to a scan so you can search, select and copy it. The pages look exactly as they did.

ENGINETESSERACT · PDF-LIB
ACCEPTSPDF
MAX SIZEMEMORY-BOUND
UPLOADNEVER

Nothing is enforced about file size, but 50 pages is: each page is rendered at about 300 dpi and held in wasm memory beside the language model, and beyond that a tab is likely to run out before it finishes.

01Drop a scanned PDF onto the page, or click the drop zone to pick one.
02Wait while each page is read — a long scan takes a while, and the strip says which page it is on.
03Press Download to save the searchable PDF. It looks identical; the text is invisible.

About OCR PDF

A scanned PDF is a stack of photographs. It may look like a document, but there is no text in it at all — searching finds nothing, selecting gets you nothing, and a screen reader has nothing to read. OCR is what puts the words back, and this page does it in your tab: Tesseract, compiled to WebAssembly, renders each page, reads it, and writes what it found back into the file as text nobody can see. What comes out is the same PDF plus a searchable layer. That last point is worth being precise about, because it is where tools of this kind usually mislead. The pages are not rebuilt. The scan's own images, their compression and any vector artwork are exactly the bytes that arrived; the recognised words are appended as invisible text at the coordinates they were found. So the document cannot silently reflow, lose a signature or change how it prints — and equally, this is not a conversion into something editable. The text is a layer over the picture, not a retyping of it. The accuracy is OCR accuracy, and here the errors are harder to notice than usual. On a page of readable text you can see when a word has been misread. Here you cannot: what you see is the original scan, which is right, while what you copy is the guess, which may not be. A misread digit in an invoice total does not announce itself. Check anything that matters against the page. Layout is the other limit. Reading runs down the page, so a two-column article comes back as the whole left column followed by the whole right one, and a table loses its rows and columns entirely. Handwriting is not supported at all — that is a different class of model, not a setting. Only English is served. And if your PDF was exported rather than scanned, it already has real text: PDF to Text will give you the characters the file stores, exactly, with no guessing involved, and this tool would only add a second and less accurate copy on top.

Questions

Does this change how my PDF looks?

No, and that is the point of doing it this way. The recognised words are written into the file as text in rendering mode 3 — the PDF instruction for 'lay this text out and draw none of it' — so a viewer runs the whole text machinery and rasterises nothing. Your pages are not rebuilt or re-encoded: the scan's own images, their compression and any vector artwork are exactly the bytes you gave us, and the text is appended alongside them. Rendered side by side, a page before and after is pixel-for-pixel identical. That is measured, not assumed — in two independent renderers, at every page rotation.

How accurate is the searchable text?

It is OCR, so it is a guess, and here the mistakes are harder to catch than usual. On a page of ordinary text you can see when a word has been misread. Here you cannot: what you see is the original scan, which is right, and what you search or copy is the recognition, which may not be. A misread digit in an invoice total does not announce itself. Clean, high-contrast, straight scans of printed text do very well; photographs of pages, low-resolution scans, heavy skew and text over busy backgrounds all do noticeably worse. The page shows Tesseract's own mean confidence, but treat it as a hint — a confident wrong answer is entirely possible.

Can I edit the text now?

No. This makes a document searchable, not editable, and the difference is the whole design. The words are an invisible layer positioned over a picture of the page — they are not a retyping of it, and the picture underneath is still a picture. Nothing on this site can turn a scan into an editable document, and anything that claims to is really rebuilding the page from its own guess, which is how a signature or a figure quietly changes. If you want to add text, white something out or sign a page, PDF Editor does that on top of the page as it stands.

My PDF already has selectable text — should I use this?

No, and the page will warn you if it spots that. A PDF that was exported rather than scanned already carries the real characters the file stores, which are exact rather than recognised. Adding OCR on top means search and copy return everything twice, and the less accurate copy is the one you just added. PDF to Text will give you that existing text as a .txt file with no guessing involved. The one case where this tool still makes sense is a mixed document — a scan bound into a digital report — where some pages have text and some do not; the warning names which pages already have it.

Why did my file get bigger?

Because a text layer is new content. It is usually small — the words themselves plus a font dictionary of about 1.3 KB for the whole document, since the font carries no glyph outlines at all (nothing is ever drawn, so none are needed). Growth of a few kilobytes per page is normal. Nothing is removed to make room, and the original page content is untouched, so the increase is close to the size of the text you gained.

Is my document uploaded?

No. Tesseract is compiled to WebAssembly and served from this site as a file — about 4.9 MB, fetched the first time you run OCR and not before — and it reads pages your browser has already rendered in the tab. The PDF never leaves your device at any point, which matters more here than almost anywhere else on this site: the documents people run OCR over are contracts, invoices, medical records and identity papers. It is the same engine Image to Text uses, so if you have used that already, this costs no extra download.

Is my file uploaded to a server?

No. Transmute processes everything locally in your browser using JavaScript and WebAssembly. Your files never leave your device — there is no server, no upload, no cloud processing.

Related