phoonnx-omnivoice
ONNX build of OmniVoice by k2-fsa (Xiaomi Corp., authors Han Zhu et al.), packaged for phoonnx.
OmniVoice is a zero-shot text-to-speech model for 600+ languages. It is a masked diffusion language model: a Qwen3-0.6B backbone writes eight streams of Higgs Audio V2 codec tokens, starting from an all-MASK grid and unmasking the most confident slots over 32 steps. Attention is bidirectional, and every step is a full-sequence forward.
Why this build exists
An earlier community export, onnx-community/OmniVoice-Onnx, runs the backbone through
the ONNX Runtime contrib operator com.microsoft::GroupQueryAttention. That operator is
unconditionally causal. OmniVoice is not a causal model, so the export computes the
wrong hidden states. We measured this against upstream PyTorch on a fixed input:
| Graph | Reference | Result |
|---|---|---|
audio_embeddings_encoder (community, fp32) |
upstream torch | cos 1.0000000, rel 6e-4 β correct |
audio_heads_decoder (community, fp16) |
upstream torch | cos 1.0000000, rel 2e-4 β correct |
llm_decoder (community, int4) |
upstream torch, bidirectional | cos 0.954 β wrong |
llm_decoder (community, int4) |
upstream torch, causal | cos 0.99945 β matches a causal model |
| community chain, end to end | upstream torch | 18.11 % greedy-token agreement |
The community llm_decoder reproduces a causal OmniVoice, which is a different model.
omnivoice_backbone.onnx here is a fresh export from the PyTorch checkpoint with the
bidirectional mask kept intact:
| Graph | Reference | Result |
|---|---|---|
omnivoice_backbone.onnx (fp32) |
upstream torch | cos 1.0000000, rel 2.9e-7, 100.00 % greedy-token agreement |
| full sampler, 32 steps, greedy | upstream _generate_iterative |
100.00 % codec-token agreement |
The community Higgs codec graphs are exact and are mirrored here unchanged:
| Graph | Reference | Result |
|---|---|---|
acoustic_encoder + semantic_encoder + quantizer_encoder |
upstream torch encode | 100.00 % codec-code agreement (all 8 codebooks) |
higgs_decoder |
upstream torch decode | max abs diff 0.0 (bit-identical) |
No quantized variant of the backbone is published: the int4 community build was rejected on the agreement test above, and we have not yet produced a quantized export that passes.
Files
| File | What it is |
|---|---|
omnivoice_backbone.onnx (+ .onnx_data) |
Qwen3 backbone + audio embeddings + audio heads, one graph. (input_ids[B,8,S] int64, audio_mask[B,S] bool) -> logits[B,8,S,1025]. Bidirectional; no KV cache. |
acoustic_encoder.onnx |
reference wav @24 kHz (1,1,T) -> acoustic features (1,256,T') |
semantic_encoder.onnx |
reference wav @16 kHz (1,T) -> semantic features (1,768,T') |
quantizer_encoder.onnx |
acoustic + semantic -> reference codes (8,1,T') |
higgs_decoder.onnx |
codes (8,1,T') -> waveform @24 kHz |
tokenizer.json, tokenizer_config.json |
the model's own Qwen3 subword BPE |
config.json |
names the phoonnx engine and the graph roles |
Classifier-free guidance runs the backbone twice per step β once over the full prompt and
once over the target span alone. Upstream batches both rows behind a [2B,1,S,S] block
mask; for a single item that is the same as two forwards of different lengths.
Usage
from phoonnx.model_manager import TTSModelManager
voice = TTSModelManager().load_voice("omnivoice/en")
audio = voice.synthesize(
"Machine learning models can now speak in hundreds of languages.",
speaker_reference="reference.wav",
speaker_reference_text="Transcription of the reference clip.",
)
speaker_reference_text is not optional in practice: OmniVoice joins the reference
transcription to the target text into a single prompt string, and cloning is noticeably
worse without it.
Verified languages
Word/character error rate on FLEURS, scored with OpenVoiceOS ONNX ASR models and always reported against a floor β the same ASR run over the real human FLEURS recording. See the phoonnx pull request for the full table. Only a sample of the model's 600+ claimed languages has been measured; the rest are untested, not verified.
Licensing and attribution
OmniVoice's code is Apache-2.0. Its released weights are CC-BY-NC because of their training data (Emilia and others), and this repository inherits that: the ONNX graphs are a format conversion of those weights. Non-commercial use only.
Model and method: Zhu Han, Ye Lingxuan, Kang Wei, Yao Zengwei, Guo Liyong, Kuang Fangjun,
Han Zhifeng, Zhuang Weiji, Lin Long and Povey Daniel β OmniVoice: Towards Omnilingual
Zero-Shot Text-to-Speech with Diffusion Language Models, arXiv 2604.00688, 2026.
The Higgs codec graphs are mirrored from onnx-community/OmniVoice-Onnx.
Do not use this model for voice cloning without the speaker's consent, or for impersonation or fraud.
@article{zhu2026omnivoice,
title={OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models},
author={Zhu, Han and Ye, Lingxuan and Kang, Wei and Yao, Zengwei and Guo, Liyong and
Kuang, Fangjun and Han, Zhifeng and Zhuang, Weiji and Lin, Long and Povey, Daniel},
journal={arXiv preprint arXiv:2604.00688},
year={2026}
}
- Downloads last month
- -