PowerInfer

by Tiiny-AI · AI

High-speed Large Language Model Serving for Local Deployment

New0 ratings9,772 starsActive
AIAI ApplicationC++C
GitHub⚑ Report
PowerInfer — image 1
PowerInfer — image 2
1 / 2

About this project

PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU TL;DR PowerInfer is a CPU/GPU LLM inference engine leveraging activation locality for your device. Project Kanban Latest News 🔥 — [2026/1/5] We released Tiiny AI Pocket Lab, the world's first pocket-size supercomputer. It runs GPT-OSS-120B (int4) locally at 20 tokens/s. Featured at CES 2026. — [2025/7/27] We released SmallThinker-21BA3B-Instruct and SmallThinker-4BA0.6B-Instruct. We also released a corresponding framework for efficient on-device inference. — [2024/6/11] We are thrilled to introduce PowerInfer-2, our highly optimized inference framework designed specifically for smartphones. With…

Technologies

PythonC++Objective-CCCudallmlarge-language-modelsllama

Project health

Low recent activity
Last update4 months ago
Contributors391
Latest release
Open issues & PRs129
LicenseMIT
On GitHubsince 2023

GitHub

9,772
stars
598
forks
391
contributors
129
open issues & PRs
C++
language
4 months ago
last commit
View on GitHub ↗

Reviews

out of 5 · 0 ratings
★★★★★
0%
★★★★
0%
★★★
0%
★★
0%
0%
Sign in to write a review
No reviews yet
Be the first to review PowerInfer.

Built by

Tiiny-AI
Imported from GitHub · not yet claimed on GitPalace
View developer pageSign in with GitHub to claim

Maintain Tiiny-AI/PowerInfer? Claiming verifies admin access through your GitHub account and gives you control of this listing.

You might also like