pdf-inspector

by firecrawl · Developer Tools

Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based P

New0 ratings18,890 starsActive
Developer ToolsDeveloper ToolRustPython
Open project ↗↓ Download v1.15.0GitHub⚑ Report
pdf-inspector preview

About this project

pdf-inspector Fast Rust library for PDF classification and text extraction. By default it detects whether a PDF is text-based or scanned, extracts text with position awareness, and converts to clean Markdown without OCR. Native Rust and CLI consumers can opt into selective OCR. Includes bindings for Python, Node.js, and browser WebAssembly. Built by Firecrawl to handle text-based PDFs locally in under 200ms, skipping expensive OCR services for the 54% of PDFs that don't need them. Features — Smart classification — Detect TextBased, Scanned, ImageBased, or Mixed PDFs in 10-50ms by sampling content streams. Returns a confidence score (0.0-1.0) and per-page OCR routing. — Text…

Technologies

JavaScriptPythonHTMLRustMarkdownnodejsocr-routing

Project health

Actively maintained
Last update5 days ago
Contributors12
Latest releasev1.15.0
Open issues & PRs213
LicenseMIT
On GitHubsince 2026

GitHub

18,890
stars
1,265
forks
12
contributors
213
open issues & PRs
Rust
language
5 days ago
last commit
View on GitHub ↗All releases ↗

Reviews

out of 5 · 0 ratings
★★★★★
0%
★★★★
0%
★★★
0%
★★
0%
0%
Sign in to write a review
No reviews yet
Be the first to review pdf-inspector.

Built by

firecrawl
Imported from GitHub · not yet claimed on GitPalace
View developer pageSign in with GitHub to claim

Maintain firecrawl/pdf-inspector? Claiming verifies admin access through your GitHub account and gives you control of this listing.

You might also like