A high-throughput and memory-efficient inference and serving engine for LLMs
A framework for efficient model inference with omni-modality models
A programmable Mixture-of-Models router for heterogeneous LLM inference
Cost-efficient and pluggable Infrastructure components for GenAI inference