gpustack

by gpustack · AI

A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

New0 ratings5,618 starsActive
AIAI ApplicationPythonJinja
Open project ↗↓ Download v2.2.3GitHub⚑ Report
gpustack preview

About this project

English 简体中文 日本語 Overview GPUStack is an open-source GPU cluster manager for AI model serving and GPU instance provisioning. It configures and orchestrates inference engines — vLLM, SGLang, TensorRT-LLM, or your own — and lets you launch SSH-accessible GPU instances on demand. Its core features include: — Multi-Cluster GPU Management. Manages GPU clusters across multiple environments. This includes on-premises servers, Kubernetes clusters, and cloud providers. — Pluggable Inference Engines. Automatically configures high-performance inference engines such as vLLM, SGLang, and TensorRT-LLM. You can also add custom inference engines as needed. — Day 0 Model Support. GPUStack's…

Technologies

ShellPythonPowerShellDockerfiledeepseekCudaJinjaascend

Project health

Actively maintained
Last updateyesterday
Contributors52
Latest releasev2.2.3
Open issues & PRs684
LicenseApache-2.0
On GitHubsince 2024

GitHub

5,618
stars
636
forks
52
contributors
684
open issues & PRs
Python
language
yesterday
last commit
View on GitHub ↗All releases ↗

Reviews

out of 5 · 0 ratings
★★★★★
0%
★★★★
0%
★★★
0%
★★
0%
0%
Sign in to write a review
No reviews yet
Be the first to review gpustack.

Built by

gpustack
Imported from GitHub · not yet claimed on GitPalace
View developer pageSign in with GitHub to claim

Maintain gpustack/gpustack? Claiming verifies admin access through your GitHub account and gives you control of this listing.

You might also like