pdf-inspector
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based P
About this project
pdf-inspector Fast Rust library for PDF classification and text extraction. By default it detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown without OCR. Native Rust and CLI consumers can opt into selective OCR. Includes bindings for Python, Node.js, and browser WebAssembly. Built by Firecrawl to handle text-based PDFs locally in under 200ms, skipping expensive OCR services for the 54% of PDFs that don't need them. Features — Smart classification — Detect TextBased, Scanned, ImageBased, or Mixed PDFs in 10-50ms by sampling content streams. Returns a confidence score (0.0-1.0) and per-page OCR routing. — Text…
Technologies
Project health
GitHub
Reviews
Built by
Maintain firecrawl/pdf-inspector? Claiming verifies admin access through your GitHub account and gives you control of this listing.