[{"data":1,"prerenderedAt":456},["ShallowReactive",2],{"blog-post-fr-translate-scanned-pdf-with-ocr":3},{"id":4,"title":5,"body":6,"category":444,"coverImage":445,"date":446,"description":447,"extension":448,"meta":449,"navigation":450,"path":451,"seo":452,"slug":453,"stem":454,"__hash__":455},"blog/en/blog/translate-scanned-pdf-with-ocr.md","How to Translate a Scanned PDF with OCR: A Step-by-Step Guide",{"type":7,"value":8,"toc":418},"minimark",[9,13,24,63,67,75,78,92,103,107,112,115,118,129,133,140,147,154,163,167,178,181,184,193,197,200,203,233,236,245,249,252,255,258,262,273,287,290,293,302,306,309,342,345,349,352,355,358,362,366,369,373,376,380,383,387,390,394,397,401,404,408,411],[10,11,12],"p",{},"You open a PDF and try to select a sentence, but the whole page behaves like a single image. Copy and paste does nothing, and a normal PDF translator cannot find any text to translate. That usually means you have a scanned or image-based PDF.",[10,14,15,16,23],{},"The solution is to run optical character recognition (OCR) before translation. OCR identifies the words visible on each page, after which a PDF translation app can translate the recognized text and rebuild it inside the document layout. This guide shows how to complete that workflow with ",[17,18,22],"a",{"href":19,"rel":20},"https://apps.apple.com/app/id6758205667",[21],"nofollow","Doco Translate"," on Mac.",[25,26,29,34],"div",{"className":27},[28],"tldr-block",[30,31,33],"h2",{"id":32},"tldr","TL;DR",[35,36,37,45,51,57],"ul",{},[38,39,40,44],"li",{},[41,42,43],"strong",{},"What:"," Use OCR to recognize text in a scanned or image-based PDF, then translate the recognized content.",[38,46,47,50],{},[41,48,49],{},"Why:"," A scan stores pages as images, so a translator cannot extract the words as it would from a normal text PDF.",[38,52,53,56],{},[41,54,55],{},"How:"," Choose the source and target languages in Doco Translate, open the scan, let the built-in OCR process it, review the bilingual result, and export a translated or bilingual PDF.",[38,58,59,62],{},[41,60,61],{},"Result:"," You get a readable translation that stays connected to the original page structure, with selectable text in the exported PDF, though OCR and layout still need checking.",[30,64,66],{"id":65},"does-your-pdf-need-ocr","Does your PDF need OCR?",[10,68,69,70,74],{},"A digital PDF usually contains text objects. You can select a sentence, search for a word, or copy a paragraph into another app. A scanned PDF is different: each page may be only a photograph of a paper document, even though the file still ends in ",[71,72,73],"code",{},".pdf",".",[10,76,77],{},"Try these quick checks:",[35,79,80,83,86,89],{},[38,81,82],{},"Drag across a sentence. If the whole page is selected as one image—or nothing is selected—the page probably needs OCR.",[38,84,85],{},"Search for a word that is clearly visible. No result is another sign that the text has not been encoded.",[38,87,88],{},"Zoom in closely. If the letters become visibly pixelated, you may be looking at a page image rather than rendered text.",[38,90,91],{},"Check several pages. Some PDFs are mixed: a digital cover may be followed by scanned pages.",[10,93,94,95,98,99,102],{},"OCR is the bridge between the page image and the translation engine. It detects text regions, recognizes characters, and creates text that can be translated. The translation can then be placed back into the page structure. These are separate stages, so a recognition error such as ",[71,96,97],{},"1"," instead of ",[71,100,101],{},"I"," will also affect the translation unless you correct it.",[30,104,106],{"id":105},"how-to-translate-a-scanned-pdf-with-ocr","How to translate a scanned PDF with OCR",[108,109,111],"h3",{"id":110},"_1-start-with-the-clearest-scan-available","1. Start with the clearest scan available",[10,113,114],{},"OCR works best when the page is sharp, straight, evenly lit, and high contrast. If you still have the paper document, rescanning a blurry or tilted page can save more time than correcting dozens of recognition errors later.",[10,116,117],{},"Before importing the file, rotate sideways pages and check that text near the margins has not been cut off. Heavy shadows, bleed-through from the reverse side, watermarks, handwritten notes, and decorative typefaces can all reduce recognition quality.",[25,119,122],{"className":120},[121],"feature-image",[10,123,124],{},[125,126],"img",{"alt":127,"src":128},"Compare a wrinkled scanned PDF with the cleaned page produced after OCR","/images/blog/translate-scanned-pdf-with-ocr/scanned-pdf-before-after.jpg",[108,130,132],{"id":131},"_2-choose-the-source-language-manually","2. Choose the source language manually",[10,134,135,136,139],{},"Download ",[17,137,22],{"href":19,"rel":138},[21]," from the Mac App Store. Once it is installed, open the app and find the translation options above the file picker. Select the language used in the scanned document, choose the target language, and select a translation service.",[10,141,142,143,146],{},"Do not leave the source language on ",[41,144,145],{},"Auto Detect"," for a difficult scan. A scanned page does not contain the reliable text metadata available in a digital PDF. Giving OCR the correct language helps it distinguish characters and extract the text more accurately. If the document is mainly French with a few English terms, choose French as the source language and review the English terms afterward.",[10,148,149,150,153],{},"For a quick translation, ",[41,151,152],{},"Auto Select"," chooses an available basic translation service. You can also use Apple Translate or a configured AI service. The service affects how the recognized text is translated; it does not fix characters that OCR misread.",[25,155,157],{"className":156},[121],[10,158,159],{},[125,160],{"alt":161,"src":162},"Select the source language manually in Doco Translate before running OCR","/images/blog/translate-scanned-pdf-with-ocr/select-source-language-for-ocr.jpg",[108,164,166],{"id":165},"_3-open-the-scanned-pdf","3. Open the scanned PDF",[10,168,169,170,173,174,177],{},"Keep the import type set to ",[41,171,172],{},"Local",", click ",[41,175,176],{},"Open File",", and choose the PDF. Doco Translate opens the translator and begins processing the document.",[10,179,180],{},"When Doco detects a scanned document, it may prompt you to confirm the source language. Treat that prompt as part of OCR setup, not as an optional translation preference. Confirm the language before recognition continues.",[10,182,183],{},"The app processes the document in stages: it parses the pages, recognizes text where OCR is needed, sends the recognized text to the selected translation service, and displays the result. Translation appears progressively, so you can begin checking early pages while later pages continue processing.",[25,185,187],{"className":186},[121],[10,188,189],{},[125,190],{"alt":191,"src":192},"Choose a scanned PDF from the Mac file picker in Doco Translate","/images/blog/translate-scanned-pdf-with-ocr/open-scanned-pdf.jpg",[108,194,196],{"id":195},"_4-compare-the-ocr-translation-with-the-original","4. Compare the OCR translation with the original",[10,198,199],{},"Doco Translate opens in a bilingual view, with the original page on the left and the translated page on the right. Scrolling and zooming stay synchronized. Pointing to an object highlights its corresponding object on the other side, which helps you trace a suspicious translation back to the scanned source.",[10,201,202],{},"Check the places where OCR errors have the largest consequences:",[35,204,205,208,211,214,217],{},[38,206,207],{},"People, company, and place names",[38,209,210],{},"Dates, decimal points, percentages, currencies, and reference numbers",[38,212,213],{},"Table headers and values that must stay in the correct row and column",[38,215,216],{},"Abbreviations, product codes, citations, and specialist terminology",[38,218,219,220,223,224,227,228,223,230,232],{},"Characters that look similar, such as ",[71,221,222],{},"0"," and ",[71,225,226],{},"O",", ",[71,229,97],{},[71,231,101],{},", or punctuation marks",[10,234,235],{},"For multi-column pages, confirm that paragraphs were recognized in the intended reading order. For tables, compare every important value with the source rather than judging the translated wording alone. Mixed-language pages may need extra attention because one OCR language setting cannot describe every phrase equally well.",[25,237,239],{"className":238},[121],[10,240,241],{},[125,242],{"alt":243,"src":244},"Review the scanned source and OCR translation side by side in Doco Translate","/images/blog/translate-scanned-pdf-with-ocr/review-ocr-translation-side-by-side.jpg",[108,246,248],{"id":247},"_5-correct-text-and-layout-problems","5. Correct text and layout problems",[10,250,251],{},"If a translated paragraph needs work, select it in the translated area. You can edit the text directly or translate that paragraph again with a different service. You can also move or resize its container when longer translated text no longer fits comfortably.",[10,253,254],{},"This review matters because OCR, translation, and layout reconstruction solve different problems. A perfectly translated sentence may still overflow its original box, while a well-positioned paragraph may contain a recognition mistake. Review both the language and the page.",[10,256,257],{},"Doco Translate is designed to retain the document structure, including images, tables, columns, and page relationships, but no OCR PDF translator can guarantee an exact result for every scan. Translation length, dense layouts, damaged pages, and complex graphics can require manual adjustment.",[108,259,261],{"id":260},"_6-export-the-translated-or-bilingual-pdf","6. Export the translated or bilingual PDF",[10,263,264,265,268,269,272],{},"When the pages are ready, click ",[41,266,267],{},"Export"," in the upper-right corner. Choose ",[41,270,271],{},"PDF"," as the format, then select one of these modes:",[35,274,275,281],{},[38,276,277,280],{},[41,278,279],{},"Translation Only"," exports the translated pages.",[38,282,283,286],{},[41,284,285],{},"Bilingual"," exports the original and translated content together, alternating each original page with its translated page.",[10,288,289],{},"Choose a file name and location, then wait for export to finish. Exported PDFs contain selectable, editable text rather than image-only pages. PDF export is a Pro feature.",[10,291,292],{},"If the translation is for internal reading, keeping the document in the bilingual reader may be enough. A bilingual export is useful when another person needs to verify the translation against the scan.",[25,294,296],{"className":295},[121],[10,297,298],{},[125,299],{"alt":300,"src":301},"Export the OCR result as a translated or bilingual PDF","/images/blog/translate-scanned-pdf-with-ocr/export-translated-or-bilingual-pdf.jpg",[30,303,305],{"id":304},"how-to-improve-ocr-translation-quality","How to improve OCR translation quality",[10,307,308],{},"If the first result contains too many errors, work backward through the pipeline instead of changing only the translation service.",[310,311,312,318,324,330,336],"ol",{},[38,313,314,317],{},[41,315,316],{},"Improve the source image."," Rescan or replace pages that are blurred, tilted, cropped, faint, or heavily compressed.",[38,319,320,323],{},[41,321,322],{},"Confirm the source language."," A wrong or overly broad language choice can cause errors before translation begins.",[38,325,326,329],{},[41,327,328],{},"Check recognition before style."," Correct names, numbers, dates, and technical terms before polishing sentence flow.",[38,331,332,335],{},[41,333,334],{},"Review complex regions separately."," Tables, footnotes, stamps, handwriting, and text embedded in figures need closer comparison.",[38,337,338,341],{},[41,339,340],{},"Keep the original file."," Treat the translated PDF as a working derivative, especially when the document has legal, medical, financial, safety, or publication consequences.",[10,343,344],{},"For a high-stakes document, OCR and machine translation should speed up the first pass, not replace qualified human review.",[30,346,348],{"id":347},"privacy-local-pdf-parsing-is-not-always-offline-translation","Privacy: local PDF parsing is not always offline translation",[10,350,351],{},"Doco Translate parses the PDF on your device. What happens during translation depends on the service you choose.",[10,353,354],{},"If you use Google Translate, Microsoft Translator, or a cloud AI provider, the text required for translation is sent to that provider under its own terms and privacy policy. For a fully local workflow, use Apple Translate with the required language packs downloaded, or connect Doco Translate to a local model through Ollama, LM Studio, or oMLX.",[10,356,357],{},"If the scan contains confidential or regulated information, follow your organization's policy before selecting a cloud translation service.",[30,359,361],{"id":360},"frequently-asked-questions","Frequently asked questions",[108,363,365],{"id":364},"can-i-translate-a-scanned-pdf-if-i-cannot-select-its-text","Can I translate a scanned PDF if I cannot select its text?",[10,367,368],{},"Yes. That is the main purpose of OCR in this workflow. OCR recognizes the visible characters first, then the recognized text is translated and placed into a readable document view.",[108,370,372],{"id":371},"why-should-i-select-the-source-language-instead-of-using-auto-detect","Why should I select the source language instead of using Auto Detect?",[10,374,375],{},"Scanned pages lack reliable text metadata. Selecting the source language gives OCR a stronger signal for character recognition and usually reduces extraction errors.",[108,377,379],{"id":378},"can-ocr-translate-a-blurry-or-tilted-scan","Can OCR translate a blurry or tilted scan?",[10,381,382],{},"It may recognize some text, but results vary with sharpness, angle, contrast, resolution, typography, and page damage. Straightening or rescanning the page is often the best first fix.",[108,384,386],{"id":385},"will-the-translated-pdf-keep-its-tables-images-and-columns","Will the translated PDF keep its tables, images, and columns?",[10,388,389],{},"Doco Translate rebuilds translated content in the original page structure and is designed to retain those relationships. Complex layouts and longer translated text can still need manual review or adjustment.",[108,391,393],{"id":392},"can-i-translate-a-scanned-pdf-completely-offline","Can I translate a scanned PDF completely offline?",[10,395,396],{},"Yes, if both document processing and translation stay local. In Doco Translate, use Apple Translate with downloaded language packs or a local model workflow such as Ollama, LM Studio, or oMLX. Cloud translation services require sending the text needed for translation to the provider.",[108,398,400],{"id":399},"can-i-export-searchable-text-from-an-image-based-pdf","Can I export searchable text from an image-based PDF?",[10,402,403],{},"Yes. Doco Translate's PDF export contains selectable, editable translated text. You can export the translation alone or create a bilingual PDF for source comparison.",[30,405,407],{"id":406},"turn-an-image-only-pdf-into-a-translation-you-can-review","Turn an image-only PDF into a translation you can review",[10,409,410],{},"A scanned PDF is not untranslatable; it simply needs a recognition step before translation. Start with a clear scan, specify the source language, compare the result with the original, and correct important details before exporting.",[10,412,413,417],{},[17,414,416],{"href":19,"rel":415},[21],"Download Doco Translate"," on your Mac to run OCR, translate the scanned PDF, and review the pages side by side in one workflow.",{"title":419,"searchDepth":420,"depth":420,"links":421},"",2,[422,423,424,433,434,435,443],{"id":32,"depth":420,"text":33},{"id":65,"depth":420,"text":66},{"id":105,"depth":420,"text":106,"children":425},[426,428,429,430,431,432],{"id":110,"depth":427,"text":111},3,{"id":131,"depth":427,"text":132},{"id":165,"depth":427,"text":166},{"id":195,"depth":427,"text":196},{"id":247,"depth":427,"text":248},{"id":260,"depth":427,"text":261},{"id":304,"depth":420,"text":305},{"id":347,"depth":420,"text":348},{"id":360,"depth":420,"text":361,"children":436},[437,438,439,440,441,442],{"id":364,"depth":427,"text":365},{"id":371,"depth":427,"text":372},{"id":378,"depth":427,"text":379},{"id":385,"depth":427,"text":386},{"id":392,"depth":427,"text":393},{"id":399,"depth":427,"text":400},{"id":406,"depth":420,"text":407},"Tutorial","/images/blog/translate-scanned-pdf-with-ocr/cover.jpg","2026-08-19","Learn how to translate scanned and image-based PDFs with OCR, improve text recognition, and preserve tables, images, columns, and the original page layout.","md",{},true,"/en/blog/translate-scanned-pdf-with-ocr",{"title":5,"description":447},"translate-scanned-pdf-with-ocr","en/blog/translate-scanned-pdf-with-ocr","uCOm12bipy2F2i0mOJmC-cELFFPBYaUokWvFLxWyu9E",1787207635331]