vllm-omni

by vllm-project · Machine Learning

A framework for efficient model inference with omni-modality models

New0 ratings6,705 starsActive
Machine LearningDeveloper ToolPythonShell
Open project ↗↓ Download v0.28.0GitHub⚑ Report
vllm-omni preview

About this project

Easy, fast, and cheap omni-modality model serving for everyone Documentation DeepWiki User Forum Developer Slack WeChat Paper Slides Latest News 🔥 — [2026/08] We released 0.28.0, featuring production-ready MiniMax H3 serving on GPU and NPU, a unified AR/DiT paged KV cache runtime, and enhanced realtime full-duplex serving for the MiniCPM-o series. — [2026/08] VeRL-Omni v0.2.0 is released: faster diffusion RL powered by vLLM-Omni (request-level/step-wise batching with FA3), rebuilt Qwen3-Omni multimodal training (DPO & GSPO), plus LTX-2.3, Qwen-Image-Edit support and more. See the release notes. — [2026/08] We released 0.26.0 - aligned with the vLLM 0.26 release line, featuring…

Technologies

ShellPythonC++CudaJinjadiffusionimage-generationaudio-generation

Project health

Actively maintained
Last updateyesterday
Contributors384
Latest releasev0.28.0
Open issues & PRs1,900
LicenseApache-2.0
On GitHubsince 2025

GitHub

6,705
stars
1,662
forks
384
contributors
1,900
open issues & PRs
Python
language
yesterday
last commit
View on GitHub ↗All releases ↗

Reviews

out of 5 · 0 ratings
★★★★★
0%
★★★★
0%
★★★
0%
★★
0%
0%
Sign in to write a review
No reviews yet
Be the first to review vllm-omni.

Built by

vllm-project
Imported from GitHub · not yet claimed on GitPalace
View developer pageSign in with GitHub to claim

Maintain vllm-project/vllm-omni? Claiming verifies admin access through your GitHub account and gives you control of this listing.

You might also like