
Description
Offline TTS and voice cloning on an ordinary PC usually needs a GPU, and small models sound stiff. MOSS-TTS-Nano from Fudan's OpenMOSS is a 100M-parameter multilingual TTS model that runs in real time on CPU.
It supports voice cloning, 48 kHz stereo and streaming, strong in Chinese and English, with ONNX builds and a demo.
Real-time CPU:No GPU.
Cloning:From seconds of audio.
Quality:48 kHz stereo.
Streaming:Plays as it generates.
It supports voice cloning, 48 kHz stereo and streaming, strong in Chinese and English, with ONNX builds and a demo.
Features
Real-time CPU:No GPU.
Cloning:From seconds of audio.
Quality:48 kHz stereo.
Streaming:Plays as it generates.

