Baharani v13 - VoxCPM2 LoRA (Bahraini Arabic)

Correction (2026-08-28): do not use this adapter for accent transfer

A later, much larger evaluation found that the base model with a reference wav outperforms this LoRA, and that the LoRA's contribution to Bahraini accent is statistically indistinguishable from zero while it roughly triples the error rate. The recommendation further down this card -- "reference conditioned (recommended)" with the adapter loaded -- is superseded. Load openbmb/VoxCPM2 on its own and pass a reference wav instead.

What was measured

Accent distance to a real Bahraini speaker (lower is better; a distance of 0.91 is what two genuine Bahraini speakers score against each other, so that is the noise floor). CER is normalized character error rate.

configuration accent vs. real speaker CER n
base, no reference wav 1.50 4.98% 60
base + reference wav 0.34 5.78% 60
base + reference + prompt pair 0.28 7.20% 60
v14 LoRA (EMA) + reference 0.38 17.23% 60
v13 step 9000 (this repo) + reference 0.99 13.25% 6

Paired bootstrap over the same 60 texts, B=4000. Adding the reference wav to plain base moves accent by -1.16 (95% CI [-1.30, -0.98], p~0.0000) -- 1.27x the between-speaker noise floor, and the largest effect measured anywhere in this project. Adding a LoRA on top of that reference changes accent by an amount whose confidence interval spans zero (p = 0.28 to 0.37) while tripling CER; the v14 adapter loses to plain base on 45 of 60 texts. Three successive fine-tuning generations (v11, v12, v13/v14) each failed to beat in-context conditioning that costs nothing.

Note the asymmetry in n: the 0.99 figure for this checkpoint comes from a 6-text pilot, while the 60-text run was done on v14. This repo's checkpoint was not itself re-scored at n=60, so treat 0.99 as indicative rather than tight. It is reported here because it is the only measurement of this checkpoint that exists, and it is worse than base + reference by a wide margin.

What to use instead

No adapter, reference_wav_path alone, 8 diffusion timesteps:

from voxcpm import VoxCPM
m = VoxCPM.from_pretrained("openbmb/VoxCPM2")   # no LoRA
wav = m.generate(text=text, reference_wav_path="your_bahraini_ref.wav",
                 inference_timesteps=8)

A reference clip needs no transcript, so adding a new voice needs only a wav file. Measured on one A40 at these settings: 0.39 s to first byte and a real-time factor of 0.83 while streaming, i.e. faster than real time with about 15% headroom.

Honest limits on the above

  • The reference speakers were present in the v14 training data, so the comparison is seen-speaker and flatters the LoRA rather than the base model.
  • CER/WER come from Whisper large-v3, which is biased toward Modern Standard Arabic and penalizes Bahraini dialect in absolute terms. Read the deltas between rows, not the absolute percentages.
  • The accent metric is spectral-statistical (formants F1/F2/F3, F2-F1, f0 median and range, voiced-segment rate), not perceptual.
  • Single training seed, and no human listening test has been run. If careful listening contradicts this table, the table is wrong.

Everything below this box is the original v13 card, left unedited. Its training-run documentation and its correction of the earlier "cosine schedule" claim are still accurate; only the recommendation to use the adapter is not.

LoRA adapter for openbmb/VoxCPM2, fine-tuned on Bahraini / Gulf Arabic speech from wldaldakheel/baharani-mix.

Published checkpoint: step 9000 of 15000, selected by validation loss. This is not the final checkpoint โ€” the run overfit after step 9000. See Checkpoint selection.

Known limitation: hard ~3.04 s minimum output duration

Every v13 checkpoint refuses to emit audio shorter than about 3.04 s, no matter how short the input text is. Short prompts are stretched, not padded with silence.

The cause is in the training data, not the hyperparameters: the shortest segment in the dataset is exactly 3.00 s, and 37.8% of the 97,087 training rows sit exactly on that floor, so the model never saw a sub-3 s example. Latent frames are 0.16 s and 3.04 s is 19 frames โ€” every observed duration is an exact multiple of 0.16 s.

Measured on six short (5โ€“7 word) Arabic prompts at a fixed seed, reference-conditioned:

sample 00 01 02 03 04 05
untuned base VoxCPM2 2.08 2.24 2.40 3.04 4.16 2.24
v13 step 9000 3.04 3.20 3.04 3.04 5.12 3.04

The untuned base tracks text length; v13 floors. If your application needs short utterances (single words, short confirmations, IVR fragments), this adapter is not ready โ€” that requires a dataset rebuilt with sub-3 s segments. Everything else below is otherwise sound.

Usage

from voxcpm import VoxCPM
from voxcpm.model.voxcpm import LoRAConfig
import soundfile as sf

cfg = LoRAConfig(
    enable_lm=True, enable_dit=True, enable_proj=False,
    r=32, alpha=64, dropout=0.0,
)

model = VoxCPM.from_pretrained(
    "openbmb/VoxCPM2",
    lora_weights_path="best",     # this repo's best/ directory
    lora_config=cfg,
    load_denoiser=False,          # avoids a hard modelscope dependency
)

# Reference-conditioned (recommended โ€” voice cloning from a reference clip)
wav = model.generate(
    text="ุดุฎุจุงุฑูƒ ุดู„ูˆู†ูƒ ุงู„ูŠูˆู… ูŠุง ุฎูˆูŠ",
    reference_wav_path="ref.wav",
    seed=1234,
)
sf.write("out.wav", wav, 48000)

# Zero-shot (no reference) โ€” works, but see the note below
wav = model.generate(text="ุดุฎุจุงุฑูƒ ุดู„ูˆู†ูƒ ุงู„ูŠูˆู… ูŠุง ุฎูˆูŠ", seed=1234)

load_denoiser=False matters: with the denoiser enabled, from_pretrained imports zipenhancer, which requires modelscope. Without that flag and without modelscope installed, loading fails with ModuleNotFoundError: No module named 'modelscope'.

The adapter is r=32 / alpha=64 specifically so it can be hot-swapped by servers that hard-code r=32 (upstream VoxCPM issue #283). A rank-64 adapter is not servable there.

Prefer reference-conditioned generation

Zero-shot output from this adapter should not be used to judge quality. The untuned base model is also erratic in zero-shot mode on these prompts (3 of 6 samples show a low spectral rolloff), so zero-shot differences are not attributable to the fine-tune. Reference-conditioned output is level-consistent (โˆ’16 to โˆ’21 dBFS, no level collapse) and shows no degenerate behaviour. Both sets are in samples/ so you can compare directly.

Checkpoint selection

Validation ran every 3000 steps. Total validation loss bottomed at step 9000 and rose afterwards โ€” the run overfit over its last 6000 steps:

step loss/total loss/diff loss/stop
0 1.156592 0.981617 0.174975
3000 1.023304 0.968512 0.054792
6000 1.004614 0.943909 0.060706
9000 0.984550 0.942001 0.042549
12000 1.007939 0.933669 0.074270
14999 1.004583 0.958660 0.045923

A second, independent signal agrees. The trainer generates reference-conditioned audio at each validation step; comparing generated duration against the reference duration, step 9000 is the only step where both references are tracked:

step ref 4.0 s ref 6.0 s ratios
0 3.84 7.20 0.96 / 1.20
3000 4.00 5.12 1.00 / 0.85
6000 4.00 4.00 1.00 / 0.67
9000 4.00 6.08 1.00 / 1.01
12000 4.00 5.12 1.00 / 0.85
14999 4.00 4.00 1.00 / 0.67

Steps 6000 and 14999 both collapse a 6.0 s reference to 4.0 s. Step 9000 also retains the most length variation in its own output (5.12 s on sample 04, tracking the base model's 4.16 s, where every other checkpoint flattens to ~3.04 s).

Training recipe (as actually run)

Parameter Value
Base model openbmb/VoxCPM2
LoRA rank / alpha 32 / 64 (ratio 2.0)
LoRA dropout 0.0
Adapted modules q/k/v/o projections in both LM and DiT; projection layers not adapted
Learning rate 1e-4
LR schedule 300-step warmup, then effectively constant (see note)
Weight decay 0.01
Max grad norm 1.0
Steps 15000 (2.47 epochs); step 9000 published
Batch 2 ร— 8 grad-accum = 16 effective, max_batch_tokens 4096
Precision bfloat16
Checkpoint interval 1000 steps
Validation interval 3000 steps
Dataset 97,087 train / 1,252 validation segments
Audio 16 kHz in, 48 kHz out
Hardware 1ร— A40 46 GB, ~10.5 h

Note on the LR schedule. The config requests lr_scheduler: cosine, but the cosine horizon was not tied to the 15000 steps actually run, so the learning rate fell only from 1.00e-4 to 9.5e-5 across the entire run โ€” a 5% decay. Training was effectively constant-LR. Earlier Baharani model cards claimed a cosine schedule; that claim was wrong and is not repeated here. training/train.log records the per-step LR, so this is checkable.

Note on 48 kHz. The VoxCPM2 VAE is 16 kHz in / 48 kHz out. Output is written at 48 kHz, but no genuine content above the 16 kHz input band is recovered, and rebuilding the dataset at a higher sample rate would not change that.

Contents

Path What
best/ step 9000 weights (lora_weights.safetensors, 72,397,184 B) + LoRA config
best/best_step.txt, best/best_val_loss.txt 9000, 0.984550
lora_config.json provenance: selected step, val loss, md5, LoRA geometry
lora_config.yaml the exact training config used
samples/mode_b_step9000/ 6 reference-conditioned samples, step 9000
samples/mode_a_step9000/ 6 zero-shot samples, step 9000
samples/mode_b_base_control/ the same 6 prompts on the untuned base โ€” the control
samples/durations.json measured durations for all three sets
training/train.log full training log: per-step loss, LR, every validation line
scripts/runner_v13.sh the run script as executed (see caveat)

Deliberately not published

  • The final checkpoint (step 15000). Higher validation loss than step 9000, and it collapses a 6.0 s reference to 4.0 s.
  • The EMA average. Its per-checkpoint decay placed 86% of the weight on the two worst checkpoints (step 14999 = 0.631, step 14000 = 0.232) โ€” the opposite of what the validation curve supports. A uniform average over steps 8000โ€“12000 would be the sound version; it has not been built yet.
  • Optimizer state. Not needed to serve or to re-average, and 145 MB per checkpoint.

Caveat on scripts/runner_v13.sh

Published for reproducibility, but it is the version as run, and it contains two bugs that killed the post-training stage: the sample-generation call omits load_denoiser=False (so it crashes on the missing modelscope), and the sample-count line reads an empty glob under set -euo pipefail, which aborts the script. Because of those, the original run never reached its own upload stage โ€” this repository was assembled and verified separately. The script also does not apply an LR schedule over the real step count, as noted above.

Verification

best/lora_weights.safetensors
  md5    c611b2f14669713f5ae6ff942f5da0fb
  bytes  72397184
  source checkpoints/step_0009000, md5-verified against the training host

License

Apache 2.0, matching the base model.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for 9FA/baharani_v13_lora

Base model

openbmb/VoxCPM2
Adapter
(31)
this model