Benchmark result: OmniVoice across EN/DE/AR/ES/ZH
I included OmniVoice int8 in a local voice-cloning benchmark across English, German, Modern Standard Arabic, Spanish, and Mandarin Chinese:
https://www.soniqo.audio/blog/voice-cloning-benchmarks
In this run, OmniVoice was the strongest all-around row set: 0.707 mean speaker cosine across all five languages, 0.0% ASR error, and mean RTF 0.45.
The benchmark uses Google FLEURS references and includes reference audio, generated audio, speaker similarity, WER/CER, generated audio length, and RTF for each row.
This is an engineering benchmark rather than a MOS study, but I thought the OmniVoice results were worth sharing here.
Publishing the reference audio and the generated audio next to the numbers is what makes this checkable. Most benchmark posts give you a table and no way to hear whether you agree with it. Including RTF for each row is useful too, since that is usually the figure people leave out.