lmdeploy
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
About this project
📘Documentation 🛠️Quick Start 🤔Reporting Issues English 简体中文 日本語 👋 join us on Latest News 🎉 2026 — \[2026/08\] Our paper, “LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind,” has been accepted to EuroSys 2027. — \[2026/04\] PyPI has expanded the storage quota for LMDeploy and wheel uploads have resumed. v0.12.3 is now available on PyPI, so you can install it directly via pip install lmdeploy. — \[2026/02\] Support Qwen3.5 — \[2026/02\] Support vllm-project/llm-compressor 4bit symmetric/asymmetric quantization. Refer here for a detailed guide 2025 — \[2025/09\] TurboMind supports MXFP4 on NVIDIA GPUs starting from V100, achieving 1.5x the…
Technologies
Project health
GitHub
Reviews
Built by
Maintain InternLM/lmdeploy? Claiming verifies admin access through your GitHub account and gives you control of this listing.