7 min read
How to extract text from a PDF
Sometimes you need the words out of a PDF, not the PDF itself. Maybe you want to quote a paragraph, reuse a table, or search a scanned contract you cannot select. Here is how to extract text from a PDF for free, including scanned pages that need OCR, all inside your browser so the document never leaves your device.
Two kinds of PDF text
Before you extract anything, it helps to know which kind of PDF you have, because it decides which method works.
- Digital PDFs were created from a document, like a Word file or a web page. The text is real, so you can already select and copy it. Extraction here just pulls all of it out cleanly at once.
- Scanned PDFs are pictures of pages. The text looks readable to you, but to the computer it is an image, so you cannot select a single word. Getting text out of these needs OCR, short for optical character recognition, which reads the shapes of the letters and turns them back into words.
The good news is one tool handles both. It pulls the real text from digital PDFs and runs OCR on scanned ones.
Extract text step by step
The process is short whether your file is digital or scanned. Open the tool, let it read the pages, then copy or download the result.
- Open the extract text from PDF tool and drop your file in. It loads in your browser and is not uploaded.
- The tool reads each page. For a scanned document, it runs OCR to recognize the letters, which takes a few seconds per page.
- Review the extracted text on screen. Scanned pages may need a quick check for the odd misread character.
- Copy the text you need, or download it all as a plain text file.
How OCR reads a scanned page
OCR is what makes it possible to get text out of a scan or a photographed page. Instead of treating the page as one flat image, it looks at the picture, finds the shapes that look like letters, and matches them to characters. The result is real, selectable text built from what was only a picture a moment before.
A few things help OCR do its best work:
- Clear scans. Sharp, well lit pages read far better than dark or blurry ones. If your scan is rough, redoing the capture beats fixing errors by hand.
- Straight pages. Text that sits level is easier to recognize than text on a slant.
- Standard fonts. Ordinary printed type reads more accurately than decorative or handwritten text, which OCR still finds hard.
If a page came out crooked, straighten it first with the rotate PDF tool, then extract the text. A level page gives OCR a much better chance.
What to do with the extracted text
Once the words are out, you can use them however you need. Common uses include the following.
- Quote a passage in an email or report without retyping it.
- Copy a table or list into a spreadsheet.
- Make an old scanned document searchable by pulling its text.
- Feed the words into a translator or a summary.
If your real goal is a fully editable document rather than plain text, a different route may suit you better. The PDF to Word tool rebuilds the file as an editable Word document, keeping much of the layout, which is handy when you want to reshape the content rather than just copy it.
Handling tables and columns
Plain paragraphs come out of a PDF cleanly, but structured layouts need a little more attention. Text extraction reads the words, not the grid they sit in, so a table or a two column page can arrive in an order that surprises you.
- Tables. The cell contents come through, but the rows and columns may run together. Paste the result into a spreadsheet and tidy the alignment there, where it is easy to split values into columns.
- Two column pages. Newsletters and academic papers often use columns. The reading order can jump from one column to the other, so check that sentences follow on before you reuse them.
- Headers and footers. Page numbers and running titles get pulled in with the body. Cropping the page first, or simply deleting those lines after extraction, keeps the text clean.
For a simple report or letter, none of this applies and the text comes out ready to use. It is only dense, gridded layouts that reward a second look. When you do hit a stubborn table, extracting the text and cleaning it in a spreadsheet is still far faster than retyping every value by hand from the original page.
Extract text or edit the PDF
Extraction is the right move when you want the words somewhere else, like an email, a spreadsheet, or a search box. But sometimes you do not want to move the text at all. You just want to change a word or two on the page itself. In that case, extraction is the long way around.
When your goal is to fix a typo, update a figure, or reword a sentence in place, open the file in the edit PDF tool instead. It lets you change the text where it sits, so you keep the original layout and skip the copy and paste entirely. Reach for extraction when the words need to leave the document, and for editing when they need to stay put.
Your document stays private
The PDFs people extract text from are often the sensitive ones, like signed contracts, medical letters, or financial records that arrived as scans. Many OCR websites upload your file to their servers to read it, which puts a copy of your private pages on hardware you do not control.
This tool works on your device instead. The reading and the OCR both happen inside your own browser, so the PDF and the text pulled from it never leave your machine. Nothing is stored elsewhere. That means you can extract the words from a confidential agreement or a personal record without exposing it to anyone during the process.
Tips for cleaner results
A little preparation makes the extracted text more accurate and easier to reuse.
- Start with the highest quality scan you have. Better input means fewer misread characters.
- Crop away edges and margins first with the crop PDF tool so OCR focuses on the text, not on shadows or borders around the page.
- Proofread numbers carefully, since OCR sometimes confuses characters like the digit zero and the letter O.
- For a form you actually need to complete rather than copy, skip extraction and open it in the fill a PDF tool instead.
With these steps, even an old scanned page becomes text you can search, copy, and reuse, and it all stays on your own device from start to finish. What was once locked inside a picture of a page is now words you can work with, and getting there cost you nothing but a few seconds of the tool reading each page.
Frequently asked questions
How do I extract text from a scanned PDF?
Open the extract text from PDF tool and drop in your file. It runs OCR on scanned pages, recognizing the letters and turning them into text you can copy or download.
What is OCR?
OCR, or optical character recognition, reads the shapes of letters in a scanned image and converts them into real, selectable text, so you can copy words from a page that was only a picture.
Is my PDF uploaded when I extract text?
No. The tool reads the pages and runs OCR in your browser on your own device, so the file and the extracted text never leave your machine.
Why is some extracted text wrong?
OCR accuracy depends on the scan. Blurry, dark, or crooked pages produce more errors. Start with a clear, straight scan and proofread numbers, which are easiest to misread.
Should I extract text or convert the PDF to Word?
Extract text when you just need the words to copy. If you want an editable document that keeps the layout, use the PDF to Word tool instead.