vllm

by vllm-project · Machine Learning

A high-throughput and memory-efficient inference and serving engine for LLMs

New0 ratings91,098 starsActive
Machine LearningAI ApplicationPythonRust
Open project ↗↓ Download v0.28.0GitHub⚑ Report
vllm preview

About this project

Easy, fast, and cheap LLM serving for everyone Documentation Blog Paper Twitter/X User Forum Developer Slack 🔥 We have built a vLLM website to help you get started with vLLM. Please visit vllm.ai to learn more. For events, please visit vllm.ai/events to join us. About vLLM is a fast and easy-to-use library for LLM inference and serving. Originally developed in the Sky Computing Lab at UC Berkeley, vLLM has grown into one of the most active open-source AI projects built and maintained by a diverse community of many dozens of academic institutions and companies from over 2000 contributors. vLLM is fast with: — State-of-the-art serving throughput — Efficient management of…

Technologies

ShellPythonC++RustCudaamdblackwell

Project health

Actively maintained
Last update2 days ago
Contributors453
Latest releasev0.28.0
Open issues & PRs7,674
LicenseApache-2.0
On GitHubsince 2023

GitHub

91,098
stars
21,798
forks
453
contributors
7,674
open issues & PRs
Python
language
2 days ago
last commit
View on GitHub ↗All releases ↗

Reviews

out of 5 · 0 ratings
★★★★★
0%
★★★★
0%
★★★
0%
★★
0%
0%
Sign in to write a review
No reviews yet
Be the first to review vllm.

Built by

vllm-project
Imported from GitHub · not yet claimed on GitPalace
View developer pageSign in with GitHub to claim

Maintain vllm-project/vllm? Claiming verifies admin access through your GitHub account and gives you control of this listing.

You might also like