phoonnx-omnivoice

ONNX build of OmniVoice by k2-fsa (Xiaomi Corp., authors Han Zhu et al.), packaged for phoonnx.

OmniVoice is a zero-shot text-to-speech model for 600+ languages. It is a masked diffusion language model: a Qwen3-0.6B backbone writes eight streams of Higgs Audio V2 codec tokens, starting from an all-MASK grid and unmasking the most confident slots over 32 steps. Attention is bidirectional, and every step is a full-sequence forward.

Why this build exists

An earlier community export, onnx-community/OmniVoice-Onnx, runs the backbone through the ONNX Runtime contrib operator com.microsoft::GroupQueryAttention. That operator is unconditionally causal. OmniVoice is not a causal model, so the export computes the wrong hidden states. We measured this against upstream PyTorch on a fixed input:

Graph Reference Result
audio_embeddings_encoder (community, fp32) upstream torch cos 1.0000000, rel 6e-4 β€” correct
audio_heads_decoder (community, fp16) upstream torch cos 1.0000000, rel 2e-4 β€” correct
llm_decoder (community, int4) upstream torch, bidirectional cos 0.954 β€” wrong
llm_decoder (community, int4) upstream torch, causal cos 0.99945 β€” matches a causal model
community chain, end to end upstream torch 18.11 % greedy-token agreement

The community llm_decoder reproduces a causal OmniVoice, which is a different model. omnivoice_backbone.onnx here is a fresh export from the PyTorch checkpoint with the bidirectional mask kept intact:

Graph Reference Result
omnivoice_backbone.onnx (fp32) upstream torch cos 1.0000000, rel 2.9e-7, 100.00 % greedy-token agreement
full sampler, 32 steps, greedy upstream _generate_iterative 100.00 % codec-token agreement

The community Higgs codec graphs are exact and are mirrored here unchanged:

Graph Reference Result
acoustic_encoder + semantic_encoder + quantizer_encoder upstream torch encode 100.00 % codec-code agreement (all 8 codebooks)
higgs_decoder upstream torch decode max abs diff 0.0 (bit-identical)

No quantized variant of the backbone is published: the int4 community build was rejected on the agreement test above, and we have not yet produced a quantized export that passes.

Files

File What it is
omnivoice_backbone.onnx (+ .onnx_data) Qwen3 backbone + audio embeddings + audio heads, one graph. (input_ids[B,8,S] int64, audio_mask[B,S] bool) -> logits[B,8,S,1025]. Bidirectional; no KV cache.
acoustic_encoder.onnx reference wav @24 kHz (1,1,T) -> acoustic features (1,256,T')
semantic_encoder.onnx reference wav @16 kHz (1,T) -> semantic features (1,768,T')
quantizer_encoder.onnx acoustic + semantic -> reference codes (8,1,T')
higgs_decoder.onnx codes (8,1,T') -> waveform @24 kHz
tokenizer.json, tokenizer_config.json the model's own Qwen3 subword BPE
config.json names the phoonnx engine and the graph roles

Classifier-free guidance runs the backbone twice per step β€” once over the full prompt and once over the target span alone. Upstream batches both rows behind a [2B,1,S,S] block mask; for a single item that is the same as two forwards of different lengths.

Usage

from phoonnx.model_manager import TTSModelManager

voice = TTSModelManager().load_voice("omnivoice/en")
audio = voice.synthesize(
    "Machine learning models can now speak in hundreds of languages.",
    speaker_reference="reference.wav",
    speaker_reference_text="Transcription of the reference clip.",
)

speaker_reference_text is not optional in practice: OmniVoice joins the reference transcription to the target text into a single prompt string, and cloning is noticeably worse without it.

Verified languages

Word/character error rate on FLEURS, scored with OpenVoiceOS ONNX ASR models and always reported against a floor β€” the same ASR run over the real human FLEURS recording. See the phoonnx pull request for the full table. Only a sample of the model's 600+ claimed languages has been measured; the rest are untested, not verified.

Licensing and attribution

OmniVoice's code is Apache-2.0. Its released weights are CC-BY-NC because of their training data (Emilia and others), and this repository inherits that: the ONNX graphs are a format conversion of those weights. Non-commercial use only.

Model and method: Zhu Han, Ye Lingxuan, Kang Wei, Yao Zengwei, Guo Liyong, Kuang Fangjun, Han Zhifeng, Zhuang Weiji, Lin Long and Povey Daniel β€” OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models, arXiv 2604.00688, 2026. The Higgs codec graphs are mirrored from onnx-community/OmniVoice-Onnx.

Do not use this model for voice cloning without the speaker's consent, or for impersonation or fraud.

@article{zhu2026omnivoice,
  title={OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models},
  author={Zhu, Han and Ye, Lingxuan and Kang, Wei and Yao, Zengwei and Guo, Liyong and
          Kuang, Fangjun and Han, Zhifeng and Zhuang, Weiji and Lin, Long and Povey, Daniel},
  journal={arXiv preprint arXiv:2604.00688},
  year={2026}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for OpenVoiceOS/phoonnx-omnivoice

Finetuned
Qwen/Qwen3-0.6B
Finetuned
k2-fsa/OmniVoice
Quantized
(27)
this model

Collection including OpenVoiceOS/phoonnx-omnivoice

Paper for OpenVoiceOS/phoonnx-omnivoice