PegaInfer

PegaInfer

Pure Rust and CUDA LLM inference engine

Description

Serving LLMs usually drags in PyTorch and Python, bloating images to many GB. PegaInfer is a pure Rust and CUDA LLM inference engine with no PyTorch.

It exposes an OpenAI-compatible API, serves Qwen3 to Kimi-K2 and optimizes KV caching.

Features



No PyTorch:Pure Rust.

CUDA:Fast kernels.

OpenAI API:Drop-in.

Models:Qwen3 to Kimi-K2.