OCR PDF
Run OCR on a scanned PDF to add a searchable text layer, so you can find and copy text. Supports multiple languages, entirely in your browser.
Enviar arquivo PDF
Sobre OCR
OCR (Reconhecimento Óptico de Caracteres) extrai texto de documentos digitalizados e imagens. Para melhores resultados, use digitalizações de alta qualidade e selecione o(s) idioma(s) correto(s).
Sobre esta ferramenta
A scanned PDF is a photograph of a document. It looks like text to you and is completely opaque to your computer - Ctrl+F finds nothing, you cannot copy a sentence, and screen readers have nothing to read. Which is why a folder of scanned contracts is effectively unsearchable.
OCR fixes that by recognising the characters in the image and adding an invisible text layer positioned behind them. The page looks identical; search, copy and text selection start working.
Recognition runs in your browser, which is unusual for OCR and the reason this is safe to use on documents you cannot send to a third-party service. Multiple languages are supported, and picking the right one measurably improves accuracy - the engine uses language models, not just letter shapes.
Como usar
Load the scanned PDF
Drop in the file. Higher-resolution scans give better results - 300 DPI is the usual target.
Choose the language
Select the language of the document. This matters more than people expect.
Run recognition
Pages are processed in turn. Expect a few seconds per page; the language data downloads once on first use.
Download the searchable PDF
The result looks the same but the text is now selectable and searchable.
Casos de uso
Making an archive searchable
Years of scanned invoices become a set you can search by supplier name.
Quoting from a scanned document
Copy a clause out of a scanned contract instead of retyping it.
Accessibility
Screen readers cannot read an image. A text layer makes the document accessible.
Perguntas frequentes
How accurate is it?
On a clean 300 DPI scan of printed text, typically above 95 percent. Accuracy falls with low resolution, skew, faint print and handwriting.
Does the page look different afterwards?
No. The text layer is invisible and sits behind the image, so the appearance is unchanged.
Is my document uploaded for processing?
No. Recognition runs in your browser, which is what makes it usable for confidential material.
How can I improve the results?
Deskew the scan first, make sure the resolution is at least 300 DPI, and select the correct language before running.