About this tool
OCR — Optical Character Recognition — adds an invisible text layer to a scanned PDF so the document becomes searchable and copy-paste-able. The page images are left untouched, so the file looks identical to the original; only the underlying text data changes. We support English, French, German, and Spanish recognition, picked at upload time so the engine can use the right language dictionary.
OCR is most useful on scanned documents — camera scans of contracts, copier scans of paperwork, photographs of pages of text. It works on regular PDFs too, but adds little: a PDF exported from Word or a browser already contains real text the OS can index, so OCR finds nothing new to do. The free tier processes up to 10 pages per upload; Pro raises that to 50. The page caps exist because OCR is meaningfully more CPU-intensive than other operations on this site.
How it works
Drop your scanned PDF
Drag a scanned PDF onto the upload area or click to pick one. It is uploaded over HTTPS for processing and is not intentionally retained after the request.
Pick the language
Choose the language that matches the document — English, French, German, or Spanish. Picking the wrong one drops accuracy sharply.
Click Extract text
We run the recognition pass over every page and add the recognised text as an invisible layer beneath each image.
Download the searchable PDF
Save the result locally. The file looks the same as before, but you can now search, select, and copy text out of it.
When to use it
- Making a stack of scanned contracts searchable so you can find clauses by keyword instead of reading each page.
- Indexing scanned receipts and invoices so accounting tools and search bars can pull text from them.
- Converting photographed documents into a searchable archive — old letters, handwritten notes (when print-clear), legacy paperwork.
- Preparing scanned PDFs for keyword search inside document management systems that don't OCR on their own.
- Re-using passages of text from scanned reports without retyping them by hand — copy from the searchable PDF instead.
- Improving accessibility: screen readers can announce text from an OCR'd PDF; raw scans give them nothing to read.
Frequently asked questions
- What does OCR mean?
- OCR stands for Optical Character Recognition. It's the process of looking at an image of text — like a scanned page or a photo of a document — and turning it into actual text characters a computer can read. Once a PDF has been OCR'd, you can search it, copy text out of it, and screen readers can announce it.
- Will the PDF look different after OCR?
- No. The visual rendering is preserved exactly — the page images are untouched. We add an invisible text layer underneath the image so the document is searchable and selectable, but every pixel you see stays the same as the original scan.
- Which languages are supported?
- English, French, German, and Spanish. Pick the language that matches the dominant text on the page when you upload — accuracy drops sharply if the language is wrong, since the engine uses language-specific dictionaries to decide between similar-looking characters.
- Why is the page cap 10 free / 50 Pro?
- OCR is CPU-intensive — recognising every glyph on every page takes meaningfully more compute than a simple PDF rewrite. The page caps stop a single upload from monopolising the server and slowing things down for everyone else. If you have a longer document, split it first and OCR the relevant section.
- Will OCR work on a phone screenshot of a page of text?
- Usually yes, if the photo is clear, in focus, and well-lit. Blurry shots, heavy shadows, or extreme angles produce noticeably worse results — letters get misread, words get joined or split, and accuracy can drop below useful. For best results, scan or photograph straight-on with even lighting.
- Why doesn't OCR seem to do much on my regular PDF?
- Because a regular PDF — one exported from Word, a browser, or a design tool — already contains the text as text. You can already search and copy from it. OCR is for documents whose text exists only as images, like a scan of a paper contract or a photo of a receipt. On a normal PDF the operation is essentially a no-op.