F FlyerUtility

OCR PDF

Turn a scanned PDF into one you can search, select and copy from

Advertisement · 728×90

What this tool does

  • Recognises text in scans — using Tesseract, the open-source OCR engine, compiled to WebAssembly
  • Searchable PDF output — an invisible text layer is placed over the original image, so the page still looks like the scan but is fully searchable and selectable
  • Or plain text — export just the recognised words as a .txt file
  • Over 20 languages — including French, German, Spanish, Italian, Portuguese, Dutch, Russian, Arabic, Hindi, Bengali, Chinese, Japanese and Korean
  • Per-page confidence — see how sure the engine was, so you know which pages to proofread
  • Skips pages that already have text — so a mixed document is not needlessly rasterised

Why in-browser OCR is a big deal

OCR is the one PDF operation that almost every free site does on a server, because it needs a real recognition engine. That means your scanned contracts, medical letters and bank statements — exactly the documents people need OCR for — get uploaded to someone else's machine.

Here the engine itself runs on your device as WebAssembly. The only thing downloaded is the language model, fetched from a public CDN the first time you use a language and then cached by your browser. Your document is never transmitted.

Getting good results

  • Resolution matters most — 300 DPI is the sweet spot; the quality setting here controls how finely each page is rendered before recognition
  • Straighten first — a skewed page costs far more accuracy than a slightly blurry one, so rotate pages the right way up before running OCR
  • Pick the right language — the wrong model will confidently return nonsense
  • Clean scans beat photos — even lighting and no shadow across the page

Frequently asked questions

Are my PDFs uploaded anywhere?

No — recognition runs on your device via WebAssembly. Only the language training data is downloaded from a CDN; your document is never sent anywhere.

Why is the first run slow?

The language model is a few megabytes and has to be downloaded once, after which your browser caches it. Recognition itself takes a few seconds per page because it is real work being done locally rather than on a server farm.

How accurate is it?

On a clean 300 DPI scan of printed text, typically 95-99% of characters. Accuracy drops on low-resolution scans, skewed pages, unusual fonts and handwriting — handwriting is not supported in any practical sense.

Does the page still look like the original?

Yes. The scanned image is kept exactly as it was and the recognised text is written behind it as invisible glyphs, which is how searchable PDFs are normally built.

Can I OCR a huge document?

You can, but it is processed page by page in your browser, so a 500-page scan will take a long while and a lot of memory. Split it first if you only need part of it.