Turn scans and photos into structured text.
document parsing
Extracts text, positions, and confidence data from images and scanned documents using PaddleOCR.
When to use it
Use it when extracting text from an image or scanned document.
Give it an image or scanned document; it returns extracted text with position and confidence data.
What you provide
No additional actions listed in the analysis.
The OCR examples require a Python runtime.
PaddlePaddle must be installed for the CPU OCR setup.
The PaddleOCR package must be installed.
pdf2image is needed for the documented scanned-PDF processing path.