NVIDIA Model Optimizer

NVIDIA Model Optimizer

Unified model optimization library

Description

Deploying large models is slow and memory-hungry, and quantization, distillation and pruning each have separate tools. NVIDIA Model Optimizer is a unified library of quantization, distillation, pruning, architecture search, speculative decoding and more.

Compressed models go straight to TensorRT-LLM, TensorRT and vLLM for much faster inference.

Features



Quantization:FP8, INT4 and more.

Distill and prune:Smaller models.

Speculative decoding:Faster generation.

Deployment:TensorRT-LLM and vLLM.
Tags:llmgpu