AcademyMCP serversData and reporting

pdfmux

Data and operations teams that need clean text from PDF reports, contracts or scanned files before analysis or search.

  • Recommendedour rating
  • 83gitHub stars
  • MITlicence
  • 3 days agolast update
claude mcp add pdfmux -- npx -y pdfmux-mcp
README.md

Loading the file...

What it does

pdfmux extracts text from PDF files and checks its own output page by page. Each page goes to the extraction backend that suits it, such as PyMuPDF, Docling, OCR or an LLM fallback, and pages with low confidence are extracted again. Pages it cannot read are flagged instead of being dropped silently.

Use cases

  1. 01Extract text from mixed PDFs, including scanned pages
  2. 02Flag pages that could not be read, instead of losing them
  3. 03Prepare documents for a search or retrieval pipeline

Questions about
pdfmux.

What is pdfmux used for?

Data and operations teams that need clean text from PDF reports, contracts or scanned files before analysis or search. pdfmux extracts text from PDF files and checks its own output page by page. Each page goes to the extraction backend that suits it, such as PyMuPDF, Docling, OCR or an LLM fallback, and pages with low confidence are extracted again. Pages it cannot read are flagged instead of being dropped silently.

How do I install pdfmux?

Run this in your terminal: claude mcp add pdfmux -- npx -y pdfmux-mcp

Is pdfmux open source?

Yes. The code is on GitHub (NameetP/pdfmux) under the MIT licence.

Is pdfmux safe to use?

It passed our automatic scan for credential theft, hidden instructions and risky install commands. Third party open source software. Sabemos AI does not maintain it. Check the code and permissions before you connect it to company data.