Make a scanned PDF searchable, without uploading the scan

A scanned document is a stack of photographs as far as your computer is concerned — you can't search it, select from it, or have a screen reader read it aloud. OCR fixes that by recognising the text and storing it alongside the image.

Online OCR often means sending the scan to someone else's server, or hitting a free-tier page cap. Here it runs in your browser, so your scans — which are frequently identity documents and medical records — stay on your device.

The recognised text is added as an invisible layer on top of the original scan, so the document looks exactly the same and the image is never re-encoded. You also get the recognised text as a .txt file, pages that already have real text can be skipped, and you can pick the languages, the pages and how sharply each page is drawn before it is read. A progress line and a Cancel button stay on screen while it works.

What this does — and what it doesn't

  • The recognition engine is downloaded on demand and is a few megabytes, so the first page takes noticeably longer than the rest. It's never loaded unless you ask for OCR.
  • Accuracy is good on clean, straight, 300-DPI printed text and drops sharply on handwriting, low-resolution scans and skewed pages. Pages are read upright, so a scan that is sideways or upside down reads poorly; rotate it first.
  • The searchable PDF is made for Latin-alphabet languages only (English, German, French, Spanish, Portuguese, Italian), because the built-in font can't draw other scripts. For Chinese, Japanese, Arabic and Hindi you get the recognised text as a .txt file. A run covers at most 100 pages; pick a range for longer files.
  • Expect a few seconds per page on a laptop and considerably longer on a phone.

OCR PDF — frequently asked questions

Does OCR change how my document looks?
No. The recognised text is added as an invisible layer over the existing scan — the image itself isn't re-encoded or altered, so the document looks identical but is now searchable.
How accurate is it?
Very good on clean printed text scanned straight at around 300 DPI. Noticeably worse on skewed pages, low-resolution scans, unusual fonts, and poor on handwriting.
Why is the first page slow?
The recognition engine and its language data are downloaded the first time you use it, then cached by your browser. Subsequent pages and later visits are much faster.
Is my PDF uploaded anywhere?
No. Every operation runs in your browser using JavaScript and WebAssembly — the file never leaves your device. You can verify it: load the page, disconnect from the internet, and the tool still works. There is no upload endpoint to send it to.

Spotted a bug or have a suggestion?

Found something broken, have an idea, or just want to say thanks about any of our tools? Every message reaches a real person.