PDF to Word

How to convert a scanned PDF to Word

A scan contains no text whatsoever. Getting words out of one means recognising shapes, and that fails in its own particular ways.

There are two documents that both call themselves PDFs and have almost nothing in common. One was exported from a word processor and contains real text. The other is a photograph of a piece of paper. They open in the same viewer, print the same way, and behave completely differently the moment you try to get anything out of them.

If your conversion produced a Word file containing one enormous image and no editable words, you have the second kind.

Why a scan is a different problem

Converting an ordinary PDF is a matter of reading text that is already in the file and working out how it was grouped. Difficult, but the words themselves are right there.

A scanned PDF has no words in it. Someone put paper on a glass plate and the scanner recorded which parts were dark. What ended up in the file is a grid of pixels. The letters exist in the same sense that they exist in a photograph of a street sign: perfectly legible to you, and completely opaque to software.

Getting text out means looking at the picture and deducing which characters those shapes represent. That is optical character recognition, and it is a fundamentally different operation with a fundamentally different failure mode. Conversion of a real PDF makes structural mistakes: text in the wrong order, a table that is not a table. Recognition makes textual mistakes. It gets the words themselves wrong.

A born-digital PDF compared with a scanned PDF at the same magnification On the left, a born-digital page where dragging the cursor highlights individual words and the letterforms stay sharp when magnified. On the right, a scanned page where the whole page selects as a single image and magnifying reveals a coarse grid of pixels instead of letters. Born-digital Scanned magnified edges stay sharp Words highlight one at a time. magnified pixels, not letters The page selects as one image.
The five-second testDrag your cursor across a sentence. If individual words highlight, there is real text in the file. If the whole page highlights as one block, or nothing highlights at all, you are looking at a picture and no amount of ordinary conversion will get words out of it.

Recognition does not fail loudly. It does not leave gaps or question marks where it was unsure. It returns its best guess as ordinary text, formatted exactly like everything around it.

A recognised document always looks finished. Whether it is correct is a separate question, and one only you can answer.

Recovering the text with Google Docs

Google Drive runs optical character recognition automatically when you open a PDF with Google Docs, at no cost, from any browser. For most people this is the shortest route from a scan to editable text.

Running the scan through Google Docs

  1. Drag the scanned PDF into a Google Drive window and let it upload.
  2. Right-click the file, choose Open with, then Google Docs.
  3. Wait. Recognition takes noticeably longer than an ordinary conversion, and longer again on a multi-page scan.
  4. Choose File, then Download, then Microsoft Word (.docx).

The original scan is placed above the recovered text. Google Docs typically inserts the page image and puts the recognised text underneath it, which is convenient for checking and confusing if you were not expecting it. Delete the images once you have proofread the text against them.

Two things to weigh. Your document is uploaded to Google Drive to be processed, which matters a great deal for the kind of material that tends to arrive as a scan. And layout is simplified heavily, so a scanned form or a multi-column page will come back as text in roughly the right order rather than as a reproduction of the page.

Other routes, and one Word cannot do

It is worth stating plainly, because people spend real time on it: Microsoft Word does not perform optical character recognition on a scanned PDF. Word's PDF conversion reads text that already exists in the file. Point it at a scan and it will generally place the page into the document as an image. Nothing is wrong with your copy of Word. It is not a feature it has.

On a Mac, Preview can select text inside a scanned page using the system's built-in text recognition, which is genuinely useful for pulling a paragraph or an address out of a scan without converting anything. Select the text on the page, copy it, and paste it into Word. It works on a page or two and becomes tedious across a long document.

Beyond that, the category to look for is dedicated OCR software, which is where recognition accuracy and layout reconstruction get the most attention. That is also where the privacy question is easiest to settle, since software running on your own machine does not upload anything.

Scans are disproportionately the documents you should not upload. People scan things because the only copy was on paper, and that tends to mean signed contracts, medical records, identity documents, bank statements, court filings and land registry papers. Before uploading a scan to any online service, check what it does with the file and how long it keeps it.

What recognition reliably gets wrong

Recognition works on shape. It looks at a cluster of dark pixels and decides which character has that outline. Most of the time this is unambiguous. Sometimes two different characters have very nearly the same outline, and at that point the software is genuinely guessing.

Character pairs that optical character recognition confuses, and the effect on a line of figures Five pairs of characters that share almost the same shape: r n read as m, c l read as d, the digit zero read as capital O, the digit one read as lowercase L, and the digit five read as capital S. Below, an invoice line showing how these substitutions turn a correct reference and amount into a wrong one that still looks entirely plausible. Same shape, different character rn m cl d 0 O 1 l 5 S on the page recognised as In prose you would notice. In a figure you would not. INV-1058 Balance due 1,505.00 INV-l0S8 Balance due 1,SOS.00
Why numbers are the dangerSurrounding words usually resolve an ambiguous letter, which is why prose survives recognition reasonably well. A reference code or a column of figures carries no context at all, so nothing corrects the guess and the result reads as a perfectly ordinary number.

This is the part worth internalising. In running text, context does the correcting for you: if the software offers modem where the page said rnodern, the sentence around it makes the error obvious, and a spell checker catches most of what is left.

Numbers have no context. A serial number, an invoice reference, a policy number, an account number, a dosage, a measurement, a column of figures: each character stands alone, nothing constrains what it could be, and a spell checker has no opinion about it. An incorrectly recognised figure is not flagged, does not look odd, and reads exactly like a correct one.

Handwriting is a separate matter again. Recognition tuned for printed characters is generally not reliable on handwriting, and a signature or a handwritten annotation on an otherwise printed form should be treated as an image rather than as text.

What makes a scan recover well

Recognition accuracy is largely decided before the software ever runs, by the quality of the image it is given. If you still have access to the paper, rescanning well is by far the highest-value thing you can do.

Resolution

  • 300 dpi is the long-standing baseline for document recognition and is what most scanning guidance recommends
  • Below roughly 200 dpi, the fine detail that separates similar characters is simply not captured
  • Going well above 300 dpi mostly produces larger files, unless the source has very small print or footnotes

The page itself

  • Straight matters more than people expect: a page scanned at a slight angle degrades results, and most scanning software can deskew automatically
  • Strong, even contrast between ink and paper beats a darker or more dramatic image
  • Photocopies of photocopies compound their own artefacts, so use the earliest generation you have
  • Coloured or patterned backgrounds behind text interfere with separating one from the other

Why a phone photo underperforms

  • Uneven lighting leaves part of the page brighter than the rest, so a single threshold cannot suit all of it
  • Shadows, frequently from the phone itself, fall across the text
  • Paper curvature stretches letterforms, particularly near a bound spine
  • Perspective distortion makes the far edge of the page smaller than the near edge
  • A scanning app that flattens and corrects the image first removes most of this

Layout is a second problem

Suppose recognition performs perfectly and every character is correct. You still do not have your document back.

What recognition produces is words, and where on the page each one sat. Deciding that a particular group of words forms a table, or that this page has two columns which should be read in sequence rather than straight across, or that this line is a heading, is an entirely separate inference layered on top.

Which means a scanned document has to clear two independent hurdles: the characters have to be right, and then the structure has to be reconstructed from their positions. That second hurdle is the same one every PDF conversion faces, described in detail in what a PDF actually stores. A scan simply has to clear it after already clearing the first.

Scanned tables are the usual casualty. The ruled lines are pixels rather than drawn objects, so they have to be detected as lines before anything can be inferred from them, and a faint or broken rule frequently is not detected at all.

Proofreading the result

Treat a recognised document the way you would treat a translation: fluent, confident, and requiring verification against the source. A spell checker will not help, because recognition errors are almost always real words.

Check first, always

  • Every number that carries meaning: amounts, dates, quantities, percentages
  • Reference codes, account numbers, policy numbers and anything alphanumeric
  • Proper nouns, particularly names of people and places
  • Anything you would have to be able to defend if it were challenged

Check next

  • Reading order, if the original had columns, sidebars or callouts
  • Table contents against the scan, cell by cell where the figures matter
  • Anything that was handwritten, which should be assumed unreliable
  • Page boundaries, where a paragraph split across two scanned pages often loses or gains a line

Worth keeping

  • Keep the original scan alongside the recovered document, permanently
  • The scan is the record. The recovered text is a convenience built on top of it
  • If the document has any legal or financial weight, the scan is what you would need to produce

Common questions

Upload the scanned PDF to Google Drive, right-click it and choose Open with Google Docs. Google runs optical character recognition automatically and opens the recovered text as an editable document. Download it through File, Download, Microsoft Word (.docx), then proofread it against the original scan.

No. Word's PDF conversion reads text that already exists inside the file. A scanned PDF holds a picture of a page and no text at all, so Word generally places the image into the document instead of producing editable words. You need a tool that performs optical character recognition first.

Try to drag your cursor across a sentence. If individual words highlight, there is real text in the file. If nothing highlights, or the whole page highlights as one block, the page is an image. Zooming in is the second clue: real text stays crisp at any magnification, a scan becomes visibly pixelated.

Recognition matches shapes, and several characters have almost identical shapes: zero and capital O, one and lowercase L, five and S. In ordinary prose the surrounding words resolve the ambiguity. In a column of figures or a reference code there is no context to resolve it, so nothing catches the error and the wrong value looks completely normal.

300 dpi is the long-standing baseline for document recognition and what most scanning guidance recommends. Below roughly 200 dpi, letterforms lose the detail that distinguishes similar characters and accuracy drops off sharply. Scanning far above 300 dpi mainly produces larger files, unless the source has very small print.

Recognising characters and reconstructing layout are two separate problems, and the second is harder. Recognition returns words and their positions. Deciding that a group of words forms a table, or that a page has two columns to be read in sequence, is a further inference on top. Accurate text arranged in the wrong structure is a common outcome.

You can, but expect materially worse results than a flatbed scan. Photos bring uneven lighting, shadows, paper curvature and perspective distortion, all of which deform letterforms before recognition starts. If the content matters, especially if it contains figures, use a proper scan or a scanning app that flattens and corrects the image first.

Scanned documents are disproportionately the material you should not upload: signed contracts, medical records, identity documents, bank statements, court filings. An online service processes and stores the file on a server you do not control. Check the retention policy before uploading, or use software that runs locally on your own machine.