lmdeploy

by InternLM · AI

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

New0 ratings8,047 starsActive
AIAI ApplicationPythonC++
Open project ↗↓ Download v0.17.0GitHub⚑ Report
lmdeploy preview

About this project

📘Documentation 🛠️Quick Start 🤔Reporting Issues English 简体中文 日本語 👋 join us on Latest News 🎉 2026 — \[2026/08\] Our paper, “LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind,” has been accepted to EuroSys 2027. — \[2026/04\] PyPI has expanded the storage quota for LMDeploy and wheel uploads have resumed. v0.12.3 is now available on PyPI, so you can install it directly via pip install lmdeploy. — \[2026/02\] Support Qwen3.5 — \[2026/02\] Support vllm-project/llm-compressor 4bit symmetric/asymmetric quantization. Refer here for a detailed guide 2025 — \[2025/09\] TurboMind supports MXFP4 on NVIDIA GPUs starting from V100, achieving 1.5x the…

Technologies

ShellPythonC++CudaCMakecuda-kernelscodellamadeepspeed

Project health

Actively maintained
Last updateyesterday
Contributors153
Latest releasev0.17.0
Open issues & PRs605
LicenseApache-2.0
On GitHubsince 2023

GitHub

8,047
stars
739
forks
153
contributors
605
open issues & PRs
Python
language
yesterday
last commit
View on GitHub ↗All releases ↗

Reviews

out of 5 · 0 ratings
★★★★★
0%
★★★★
0%
★★★
0%
★★
0%
0%
Sign in to write a review
No reviews yet
Be the first to review lmdeploy.

Built by

InternLM
Imported from GitHub · not yet claimed on GitPalace
View developer pageSign in with GitHub to claim

Maintain InternLM/lmdeploy? Claiming verifies admin access through your GitHub account and gives you control of this listing.

You might also like