How to extract Arabic text from a scanned PDF
Scanned Arabic documents are images, so the text cannot be copied or searched. Optical character recognition fixes that, and it can run entirely on your own device.
Why this happens
A scan is a photograph of a page. The Arabic you can see is pixels, not characters, which is why you cannot select it, search it, or copy it into another document. Optical character recognition reads the shapes and turns them back into text.
Step by step
- Open the OCR PDF tool.
- Add the scanned file or photograph.
- Pick Arabic as the language so the recognizer uses the right character set.
- Run it and download the extracted text.
Worth knowing
Accuracy depends heavily on the scan. A straight, well-lit 300 dpi scan gives good results; a tilted phone photo of a faint fax will not. Always proofread names and numbers, which are where recognition errors hurt most.