Xiaomi launched OmniVoice—a public AI model capable of voicing texts in virtually all languages and mimicking voices.

Xiaomi launched OmniVoice—a public AI model capable of voicing texts in virtually all languages and mimicking voices.

76 software

Xiaomi unveiled the open AI module OmniVoice

The company announced the launch of the AI model OmniVoice, which converts text into speech. In addition to voice synthesis in hundreds of languages, the model can clone voices and generate pronunciations with user‑defined settings.

What’s new in OmniVoice?
Metric Description Languages
Support for more than 200 languages, including rare and low‑resource ones. Even with less than 10 hours of data, the model produces speech “almost on any language.” | In tests across 102 languages, speech intelligibility was comparable to human performance, and in some cases exceeded it. On 24 languages, the model outperformed several commercial systems in similarity and intelligibility. | Simplified bidirectional transformer network without intermediate token‑prediction modules. This reduces training time to one day on 100 000 hours of data and increases inference speed (up to 40× real time in PyTorch). | • “Random acoustic code masking” – speeds up training. • Connecting a large language model during pretraining improves pronunciation and intelligibility.

Practical capabilities
1. Voice cloning

- Generate speech based on described characteristics (age, gender, tone, accent, dialect).
- Ability to create whisper and other stylistic variations without reference audio.

2. Processing source recordings

- Noise removal and extraction of clean voice features even from imperfect files.

3. Intonation control

- Add sighs, laughter, and other nuances for a more natural sound.

4. Precise pronunciation correction

- Manual tuning of complex sounds, such as polyphonic Chinese characters or English proper names.

Conclusion
OmniVoice offers a competitive alternative to existing commercial speech‑synthesis systems. With its simple design, fast training and inference speeds, and extensive customization options, the model is ready for integration into consumer applications and services.

Comments (0)

Share your thoughts — please be polite and stay on topic.

No comments yet. Leave a comment — share your opinion!

To leave a comment, please log in.

Log in to comment