OmniParser
A simple screen parsing tool towards pure vision based GUI agent
About this project
OmniParser: Screen Parsing tool for Pure Vision Based GUI Agent 📢 [Project Page] [V2 Blog Post] [Models V2] [Models V1.5] [HuggingFace Space Demo] OmniParser is a comprehensive method for parsing user interface screenshots into structured and easy-to-understand elements, which significantly enhances the ability of GPT-4V to generate actions that can be accurately grounded in the corresponding regions of the interface. News — [2026/7] We add a YOLOv9-E interactive region detector. Its inference-only weight is available in Hugging Face PR #37. — [2025/3] We support local logging of trajecotry so that you can use OmniParser+OmniTool to build training data pipeline for your…
Technologies
Project health
GitHub
Reviews
Built by
Maintain microsoft/OmniParser? Claiming verifies admin access through your GitHub account and gives you control of this listing.
