PowerInfer
High-speed Large Language Model Serving for Local Deployment
About this project
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU TL;DR PowerInfer is a CPU/GPU LLM inference engine leveraging activation locality for your device. Project Kanban Latest News 🔥 — [2026/1/5] We released Tiiny AI Pocket Lab, the world's first pocket-size supercomputer. It runs GPT-OSS-120B (int4) locally at 20 tokens/s. Featured at CES 2026. — [2025/7/27] We released SmallThinker-21BA3B-Instruct and SmallThinker-4BA0.6B-Instruct. We also released a corresponding framework for efficient on-device inference. — [2024/6/11] We are thrilled to introduce PowerInfer-2, our highly optimized inference framework designed specifically for smartphones. With…
Technologies
Project health
GitHub
Reviews
Built by
Maintain Tiiny-AI/PowerInfer? Claiming verifies admin access through your GitHub account and gives you control of this listing.

