PyMuPDF
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other)
About this project
PyMuPDF The PDF engine behind over 50 million monthly downloads, powering AI pipelines worldwide. PyMuPDF is a high-performance Python library for data extraction, analysis, conversion, rendering and manipulation of PDF (and other) documents. Built on top of MuPDF — a lightweight, fast C engine — PyMuPDF gives you precise, low-level control over documents alongside high-level convenience APIs. No mandatory external dependencies. Why PyMuPDF? — Fast — powered by MuPDF, a best-in-class C rendering engine — Accurate — pixel-perfect text extraction with font, color, and position metadata — Versatile — read, write, annotate, redact, merge, split, and…
Technologies
Project health
GitHub
Reviews
Built by
Maintain pymupdf/PyMuPDF? Claiming verifies admin access through your GitHub account and gives you control of this listing.