PDFMod

PDF to Text

All tools

Drop a PDF here

Or click to choose one. It stays on this machine — we never see your file.


The words, out of the PDF.

Drop a PDF and get its text as a plain file — in reading order for prose, or with the columns and spacing roughly kept for anything with a table in it. If a page is a scan with no text to copy, you’re told, and pointed at OCR, rather than handed an empty file. It all happens on this machine; your document goes nowhere at all.

$99 once and the whole toolkit is yours — Acrobat is about $155 a year and stops working the day you stop paying. And the contract, the statement or the letter you’re copying the text out of stays on your laptop: we never see it, because nothing you open here is stored on a server.

Questions people actually ask

Where does my file go?
Nowhere. The text is read inside this page, on your machine. The app is served with a Content-Security-Policy that only permits connections back to its own origin, so the browser refuses to send your file anywhere — watch the network tab while it works.
What’s the difference between the two modes?
Reading order joins the text the way the document lays it out, with the spaces put back — it reads like prose and is right for an article or a letter. Layout preserved drops each line onto a character grid at the position it sits on the page, so a receipt’s totals stay under their headings and two columns stay two columns. That second mode is an approximation — a monospace projection of a proportional page — not an exact reproduction.
It gave me nothing / said the PDF is a scan. Why?
Because a scanned page is a picture of words with no text behind it — there is genuinely nothing to copy. Rather than hand you an empty file, PDFMod detects that and points you at OCR, which is in this same toolkit and also runs on your machine: it reads the words off the image and writes them back as a real text layer. Run that first, then come back here.
Will it get tables and columns right?
Roughly, in the layout-preserved mode — it keeps things lined up using each line’s position, which is enough to read a table or a two-column page. It is not a perfect reconstruction, and it says so. Right-to-left scripts like Arabic and Hebrew are handed to us in visual order, so those lines can come out reversed; you’re warned when the document contains them.

A scanned PDF has no text to pull out — run OCR first, it’s in this toolkit. If you want a picture of each page instead of its words, use PDF to JPG.