vllm-omni
A framework for efficient model inference with omni-modality models
About this project
Easy, fast, and cheap omni-modality model serving for everyone Documentation DeepWiki User Forum Developer Slack WeChat Paper Slides Latest News 🔥 — [2026/08] We released 0.28.0, featuring production-ready MiniMax H3 serving on GPU and NPU, a unified AR/DiT paged KV cache runtime, and enhanced realtime full-duplex serving for the MiniCPM-o series. — [2026/08] VeRL-Omni v0.2.0 is released: faster diffusion RL powered by vLLM-Omni (request-level/step-wise batching with FA3), rebuilt Qwen3-Omni multimodal training (DPO & GSPO), plus LTX-2.3, Qwen-Image-Edit support and more. See the release notes. — [2026/08] We released 0.26.0 - aligned with the vLLM 0.26 release line, featuring…
Technologies
Project health
GitHub
Reviews
Built by
Maintain vllm-project/vllm-omni? Claiming verifies admin access through your GitHub account and gives you control of this listing.
