Dolphin
The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
About this project
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting Dolphin-v2 is an enhanced universal document parsing model that substantially improves upon the original Dolphin. It seamlessly handles any document type—whether digital-born or photographed—through a document-type-aware two-stage architecture with scalable anchor prompting. 📑 Overview Document image parsing is challenging due to diverse document types and complexly intertwined elements such as text paragraphs, figures, formulas, tables, and code blocks. Dolphin-v2 addresses these challenges through a document-type-aware two-stage approach: 1. 🔍 Stage 1: Document type classification (digital vs. photographed) + layout…
Technologies
Project health
GitHub
Reviews
Built by
Maintain bytedance/Dolphin? Claiming verifies admin access through your GitHub account and gives you control of this listing.


