llama.cpp
by ggml-org · Developer Tools
LLM inference in C/C++
★ New0 ratings127,277 stars● Active
Developer ToolsDeveloper ToolC++C
About this project
llama.cpp is a C/C++ implementation of large language model and vision language model inference, built on the ggml library with the goal of running LLMs with minimal setup and state-of-the-art performance. It ships as a command-line tool for interactive sessions, a server with a REST API and built-in web UI, and a library exposing the llama API for embedding inference in other programs. It can be installed via pre-built binaries, Docker, or built from source, making it suited to developers and enthusiasts who want to run models locally on their own machines.
Technologies
TypeScriptPythonC++CCudaggml
Project health
● Actively maintained
Last update2 days ago
Contributors446
Latest releasev0.4.0
Open issues & PRs2,449
LicenseMIT
On GitHubsince 2023
GitHub
127,277
stars
22,838
forks
446
contributors
2,449
open issues & PRs
C++
language
2 days ago
last commit
Reviews
No reviews yet
Be the first to review llama.cpp.
Built by
ggml-org
Imported from GitHub · not yet claimed on GitPalace
Maintain ggml-org/llama.cpp? Claiming verifies admin access through your GitHub account and gives you control of this listing.