ktransformers
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
About this project
A Flexible Framework for Experiencing Cutting-edge LLM Inference/Fine-tune Optimizations 🎯 Overview 🚀 Inference 🎓 SFT 🔥 Citation 🚀 Roadmap(2026Q2) 🎯 Overview KTransformers is a research project focused on efficient inference and fine-tuning of large language models through CPU-GPU heterogeneous computing. The project now exposes two user-facing capabilities from the kt-kernel source tree: Inference and SFT. 🔥 Updates Aug 26, 2026: Added native support for GLM-5.3-flash, bringing 1M-token context and multimodal input to consumer GPUs. (Tutorial) Aug 25, 2026: Uploaded a new easy-to-use KTransformers × LlamaFactory MoE Fine-Tuning Cookbook, covering hardware checks,…
Technologies
Project health
GitHub
Reviews
Built by
Maintain kvcache-ai/ktransformers? Claiming verifies admin access through your GitHub account and gives you control of this listing.