Translating Scanned DocumentsWhat actually happens between your scan and the finished file

You can drop a scanned page into a translation tool today and get a translation back. It genuinely works. But something happens in between that decides whether the result is reliable or not - and this is the step most people never see.

First, what a scan actually is

A scanned page is a photograph. To a computer it is a picture of a document, not a document - the same way a photo of a road sign is a picture and not text. There are no words in the file, only light and dark areas arranged in a pattern that happens to look like writing to a human eye.

So before anything can be translated, something has to look at that picture and work out which letters it shows. That reading step is called “character recognition”, and this is where most errors in scanned document translation originate.

The reading step has not disappeared from modern tools. It moved inside them. DeepL states in its own documentation “translating a PDF relies on character recognition and therefore carries a higher error rate than translating a normal text file”.

The convenience is real - but so is the step, and so is its error rate.

Three ways to do it

Which one we choose depends on the quality of your scan, the language, what the document is for, and how confidential it is. Each has genuine advantages.

You drag the scan straight into the tool and a translated file comes back. The reading step still happens (inside the software) - it just happens out of sight.

Layout rebuilt approximatelyDeepLGoogle Cloud Translation (Advanced)Sarvam Studio

How it works

  1. 1The tool detects that the file is an image rather than selectable text
  2. 2It runs its own character recognition to turn the picture of the text into actual text
  3. 3It translates the recognised text
  4. 4It rebuilds a document that approximates the original layout

Strengths

Fastest route by a wide margin - no separate step to manage
Good results on clean, high-resolution scans in widely supported languages
One supplier, one workflow, one bill

Limitations

You cannot see or correct the recognition before it is translated
Recognition errors become fluent, confident mistranslations further down the line
IMPORTANT:Quality drops sharply on poor scans, unusual fonts and image-heavy pages
What happens to the formatting

The tool rebuilds the page rather than editing it. Simple one-column pages come back looking close to the original. Tables, columns and anything with text wrapped around images will need manual work.

Confidentiality

The file is uploaded to the vendor. For NDA and regulated work we use enterprise tiers with no-training terms, never free consumer tiers.

We use this route for: Clean scans in common scripts where the client needs the meaning quickly and the layout is straightforward.

“Does it keep the formatting?”

ANSWER: Yes but partly, and never automatically for anything complex.

When a tool says it preserves formatting, it means it attempts to rebuild a page that resembles the original. It is not editing your document - it cannot, because your document is a picture. It is constructing a new one and trying to make it look similar.

There is also a problem, which no software can solve. Translated text is almost never the same length as the original: German runs roughly 30% longer than English, Chinese noticeably shorter. Every line length on the page changes, so a layout built for the original text cannot fit the translation without adjustment. That adjustment is design work, and design work is done by people, manually.

What reliably needs manual work

  • Tables - cell boundaries are frequently lost or misplaced
  • Multi-column pages - software often reads straight across, interleaving the columns into a single unreadable stream
  • Text wrapped around images or set over coloured backgrounds
  • Headers, footers, page numbers and figure captions
  • Anything that must match the original page for certification purposes

We quote this as production time rather than leaving it as a surprise. It is a real part of the job on any scanned document with formatting.

What makes a scan hard to read

Worth knowing before you send files, because in several of these cases a better source file is cheaper than the workaround.

Low scan resolution

Major

Characters blur into each other and the software guesses. Below roughly 300 dpi, accuracy falls away quickly.

What we do: We ask for a rescan at 300+ dpi (if that is practically feasible for a client) before quoting, wherever that is possible.

Skewed or rotated pages

Moderate

Text read at an angle produces scrambled lines, and the reading order of the page can come out wrong.

What we do: Pages are de-skewed and reoriented before recognition begins.

Handwriting and signatures

Major

Handwriting recognition remains far less reliable than print. Signatures are usually unreadable by design.

What we do: Handwritten sections are transcribed by a person. Signatures are marked as such rather than invented.

Stamps, seals and watermarks

Moderate

These overlap the text underneath, and the software may merge the two into nonsense or drop the covered words entirely.

What we do: Affected areas are flagged and transcribed manually - significant for immigration and legal documents, where the stamp is often the point.

Tables and multi-column layouts

Major

Software reads straight across the page, so a two-column page can come back with the two columns interleaved into a single unreadable stream.

What we do: Reading order is verified before translation, and the layout is rebuilt by our DTP team afterwards.

Decorative or custom fonts

Moderate

Recognition is trained on common typefaces. Unusual ones produce systematic, repeated errors throughout the document.

What we do: A sample page is checked first so the error pattern is caught before the full run.

Text over images or coloured backgrounds

Moderate

Low contrast between text and background causes characters to be missed altogether rather than misread - which is harder to notice.

What we do: Contrast is adjusted before recognition, and the page is checked against the original.

Photographs of documents rather than scans

Major

Phone photos bring uneven lighting, shadow, curvature and perspective distortion all at once.

What we do: We ask for a flatbed scan. If none exists, we quote for the additional cleanup time.

Why the language changes the difficulty

Character recognition works far better in some writing systems than others. The hardest ones happen to be where our human network is strongest.

Latin

Well supported

English, French, German, Spanish, Italian and most European languages

The most heavily researched scripts, with decades of recognition work behind them. Letters sit separately and in a predictable line.

Our approach: Standard recognition, then verification against the source.

Indic

Hardest for software

Hindi, Bengali, Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Odia, Punjabi and more

Indic scripts join consonants into combined shapes, and vowel marks attach above, below and around the letters. One combined character can be several sounds, and software regularly splits or merges them wrongly. Recognition also has far less training data than Latin.

Our approach: Indic-specialist recognition tools, then verification by a native speaker of that language before translation begins. This is where we differ most from a general-purpose provider.

Arabic

Hardest for software

Arabic and its dialects, Urdu, Persian, Pashto

Letters change shape depending on their position in a word and join into flowing ligatures, so there is no clean gap between characters. Text runs right to left while embedded numbers run left to right, which confuses reading order.

Our approach: Arabic-capable recognition with native-speaker verification, and manual attention to reading order and mixed-direction numbers.

CJK

Moderate

Chinese, Japanese, Korean

Thousands of distinct characters, some differing by a single stroke. Well supported by the major vendors, but small errors are easy to miss.

Our approach: Standard recognition with native-speaker verification.

Cyrillic

Well supported

Russian, Ukrainian, Bulgarian, Serbian and others

Well supported. The main risk is characters that look identical to Latin ones being tagged as the wrong script.

Our approach: Standard recognition with a script-consistency check.

Rare African scripts and minority languages

Hardest for software

Ge'ez/Amharic, Tifinagh, N'Ko and many minority languages

Little or no commercial recognition support exists, because there is not enough training data for vendors to build it.

Our approach: Manual transcription by a native speaker. No automated route is reliable here, and we do not pretend otherwise.

Our procedure, step by step

The same sequence runs on every scanned job, and each step has a named owner.

01

Assess the source file

We check whether the file is a real scan or a digital PDF, and grade the scan for resolution, skew, contrast, handwriting, stamps and layout complexity.

Why: A digital PDF needs no recognition at all and costs less. Grading the scan first is what makes the quote accurate rather than optimistic.

Carried out by: Project manager

02

Choose the route and record why

Based on the assessment, the script involved and the confidentiality requirement, we select bundled recognition, a separate reviewed step, or a vision model - and note the reason on the job file.

Why: Engine and tool selection is a documented decision, which is what a quality audit looks for.

Carried out by: Project in charge, against our engine-selection procedure

03

Recognise the text

The scan is converted into editable text using the tool selected at step 2.

Why: Everything downstream depends on this being right.

Carried out by: Production team

04

Verify recognition against the original - before any translation

A person who reads the source language compares the recognised text against the original page and corrects every misreading, including reading order, numbers, dates and names.

Why: This is the step most workflows skip, and it is the most important one. A recognition error does not look like an error after translation - it reads as a fluent, confident sentence that happens to be wrong. Catching it here costs minutes; catching it after delivery may not happen at all.

Carried out by: Native speaker of the source language

05

Translate

The verified text enters our normal workflow at the service tier the client has chosen, with translation memory and approved glossaries (if available) applied.

Why: From this point the job is an ordinary translation, with the same quality controls as any other.

Carried out by: Qualified translator, plus reviewer as per the agreed tier

06

Rebuild the layout

Our DTP team reconstructs tables, columns, headers, figure captions and pagination to match the source document.

Why: No tool reproduces a scanned layout reliably. Treating this as a real production step rather than an afterthought is why the delivered file matches the original.

Carried out by: DTP team

07

Final check against the source page

The finished file is compared side by side with the original scan - content, layout and completeness.

Why: Confirms nothing was dropped along the way. Sections lost at the recognition stage are invisible in the translated file unless someone checks against the source.

Carried out by: Reviewer independent of the translator

Four things people often assume

Modern AI translates a scan directly, so no recognition step is needed.

The recognition step moved inside the tool rather than disappearing. DeepL states plainly in its own documentation that PDF translation relies on character recognition, and therefore carries a higher error rate. Vision models genuinely skip conventional recognition software, but they are still interpreting an image - with their own distinct failure modes.

Formatting is preserved automatically.

"Preserved" in vendor language means the layout is rebuilt approximately. Translated text is rarely the same length as the original - German runs about 30% longer than English - so the rebuilt page shifts regardless. Tables, columns and captions need manual effort.

If the translation reads well, the recognition must have been correct.

The opposite is the real risk. A misread character produces a translation that is grammatical, fluent BUT wrong. Fluency is not evidence of accuracy, which is exactly why verification happens before translation rather than after the translation.

A photo of a document is as good as a scan.

Photos add shadow, curvature and perspective distortion on top of the usual problems. They are workable, but they cost more and carry more risk than a flatbed scan of the same page.

Have scanned documents to translate?

Send us a sample page. We will tell you which route suits it, what the layout will need, and what it will cost.