
Description
Deploying large models is slow and memory-hungry, and quantization, distillation and pruning each have separate tools. NVIDIA Model Optimizer is a unified library of quantization, distillation, pruning, architecture search, speculative decoding and more.
Compressed models go straight to TensorRT-LLM, TensorRT and vLLM for much faster inference.
Quantization:FP8, INT4 and more.
Distill and prune:Smaller models.
Speculative decoding:Faster generation.
Deployment:TensorRT-LLM and vLLM.
Compressed models go straight to TensorRT-LLM, TensorRT and vLLM for much faster inference.
Features
Quantization:FP8, INT4 and more.
Distill and prune:Smaller models.
Speculative decoding:Faster generation.
Deployment:TensorRT-LLM and vLLM.
