
Description
Serving LLMs usually drags in PyTorch and Python, bloating images to many GB. PegaInfer is a pure Rust and CUDA LLM inference engine with no PyTorch.
It exposes an OpenAI-compatible API, serves Qwen3 to Kimi-K2 and optimizes KV caching.
No PyTorch:Pure Rust.
CUDA:Fast kernels.
OpenAI API:Drop-in.
Models:Qwen3 to Kimi-K2.
It exposes an OpenAI-compatible API, serves Qwen3 to Kimi-K2 and optimizes KV caching.
Features
No PyTorch:Pure Rust.
CUDA:Fast kernels.
OpenAI API:Drop-in.
Models:Qwen3 to Kimi-K2.
