Which language options are available for PDF OCR?+
The interface offers automatic language detection plus English, French, German, Spanish, Arabic, Chinese, and Japanese hints. A selected hint tells the AI route which language to expect.
Can PDF OCR make a scanned document searchable?+
Yes. The downloaded searchable PDF uses a rendered copy of each source page as its background and adds recognized text as a near-invisible text layer using an embedded Unicode Latin font. For non-Latin writing systems (such as Arabic, Bengali, Devanagari, or CJK scripts), export to TXT or use the copy-text feature to retain 100% of recognized characters without font-encoding limitations.
Are non-Latin scripts supported in searchable-PDF exports?+
The AI OCR engine recognizes non-Latin languages (including Arabic, Chinese, Japanese, and others), and all recognized characters appear in the in-browser text pane and TXT downloads. However, searchable-PDF exports embed a Latin/Cyrillic/Greek font layer; non-Latin script characters that cannot be encoded in that PDF font layer are omitted from the invisible PDF overlay. Use TXT export or copy the text to preserve full Unicode text for these scripts.
How accurate is AI PDF OCR?+
Accuracy varies with resolution, contrast, skew, fonts, handwriting, tables, and page complexity. AI output can be inaccurate or incomplete. Review results against the source document before relying on them, especially for legal, medical, financial, or other high-stakes use.
What happens to native text pages during OCR?+
When native-text detection is enabled and a page already contains enough extractable text, the tool uses that text without sending the page image for an OCR AI action.
Is my scanned PDF sent to an AI server?+
The original PDF is loaded locally. Pages that need AI OCR are rendered in the browser, and each rendered JPEG page image is sent through the OCR server route to the configured model provider.