What it does
pdfmux extracts text from PDF files and checks its own output page by page. Each page goes to the extraction backend that suits it, such as PyMuPDF, Docling, OCR or an LLM fallback, and pages with low confidence are extracted again. Pages it cannot read are flagged instead of being dropped silently.
Use cases
- 01Extract text from mixed PDFs, including scanned pages
- 02Flag pages that could not be read, instead of losing them
- 03Prepare documents for a search or retrieval pipeline