
Description
Feeding silence and noise to speech models wastes compute and slows responses. Silero VAD is a pre-trained, enterprise-grade voice activity detector that finds where someone is speaking.
About 1 MB, it processes a chunk in under a millisecond on PyTorch or ONNX across languages.
Accurate:Speech vs noise.
Tiny:~1 MB.
Fast:Sub-millisecond.
Runtimes:PyTorch and ONNX.
About 1 MB, it processes a chunk in under a millisecond on PyTorch or ONNX across languages.
Features
Accurate:Speech vs noise.
Tiny:~1 MB.
Fast:Sub-millisecond.
Runtimes:PyTorch and ONNX.
