vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
About this project
Easy, fast, and cheap LLM serving for everyone Documentation Blog Paper Twitter/X User Forum Developer Slack 🔥 We have built a vLLM website to help you get started with vLLM. Please visit vllm.ai to learn more. For events, please visit vllm.ai/events to join us. About vLLM is a fast and easy-to-use library for LLM inference and serving. Originally developed in the Sky Computing Lab at UC Berkeley, vLLM has grown into one of the most active open-source AI projects built and maintained by a diverse community of many dozens of academic institutions and companies from over 2000 contributors. vLLM is fast with: — State-of-the-art serving throughput — Efficient management of…
Technologies
Project health
GitHub
Reviews
Built by
Maintain vllm-project/vllm? Claiming verifies admin access through your GitHub account and gives you control of this listing.
