Instructions to use 9FA/baharani_v13_lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- VoxCPM
How to use 9FA/baharani_v13_lora with VoxCPM:
import soundfile as sf from voxcpm import VoxCPM model = VoxCPM.from_pretrained("9FA/baharani_v13_lora") wav = model.generate( text="VoxCPM is an innovative end-to-end TTS model from ModelBest, designed to generate highly expressive speech.", prompt_wav_path=None, # optional: path to a prompt speech for voice cloning prompt_text=None, # optional: reference text cfg_value=2.0, # LM guidance on LocDiT, higher for better adherence to the prompt, but maybe worse inference_timesteps=10, # LocDiT inference timesteps, higher for better result, lower for fast speed normalize=True, # enable external TN tool denoise=True, # enable external Denoise tool retry_badcase=True, # enable retrying mode for some bad cases (unstoppable) retry_badcase_max_times=3, # maximum retrying times retry_badcase_ratio_threshold=6.0, # maximum length restriction for bad case detection (simple but effective), it could be adjusted for slow pace speech ) sf.write("output.wav", wav, 16000) print("saved: output.wav") - Notebooks
- Google Colab
- Kaggle
Baharani v13 - VoxCPM2 LoRA (Bahraini Arabic)
Correction (2026-08-28): do not use this adapter for accent transfer
A later, much larger evaluation found that the base model with a reference wav outperforms this LoRA, and that the LoRA's contribution to Bahraini accent is statistically indistinguishable from zero while it roughly triples the error rate. The recommendation further down this card -- "reference conditioned (recommended)" with the adapter loaded -- is superseded. Load
openbmb/VoxCPM2on its own and pass a reference wav instead.What was measured
Accent distance to a real Bahraini speaker (lower is better; a distance of 0.91 is what two genuine Bahraini speakers score against each other, so that is the noise floor). CER is normalized character error rate.
configuration accent vs. real speaker CER n base, no reference wav 1.50 4.98% 60 base + reference wav 0.34 5.78% 60 base + reference + prompt pair 0.28 7.20% 60 v14 LoRA (EMA) + reference 0.38 17.23% 60 v13 step 9000 (this repo) + reference 0.99 13.25% 6 Paired bootstrap over the same 60 texts, B=4000. Adding the reference wav to plain base moves accent by -1.16 (95% CI [-1.30, -0.98], p~0.0000) -- 1.27x the between-speaker noise floor, and the largest effect measured anywhere in this project. Adding a LoRA on top of that reference changes accent by an amount whose confidence interval spans zero (p = 0.28 to 0.37) while tripling CER; the v14 adapter loses to plain base on 45 of 60 texts. Three successive fine-tuning generations (v11, v12, v13/v14) each failed to beat in-context conditioning that costs nothing.
Note the asymmetry in
n: the 0.99 figure for this checkpoint comes from a 6-text pilot, while the 60-text run was done on v14. This repo's checkpoint was not itself re-scored at n=60, so treat 0.99 as indicative rather than tight. It is reported here because it is the only measurement of this checkpoint that exists, and it is worse than base + reference by a wide margin.What to use instead
No adapter,
reference_wav_pathalone, 8 diffusion timesteps:from voxcpm import VoxCPM m = VoxCPM.from_pretrained("openbmb/VoxCPM2") # no LoRA wav = m.generate(text=text, reference_wav_path="your_bahraini_ref.wav", inference_timesteps=8)A reference clip needs no transcript, so adding a new voice needs only a wav file. Measured on one A40 at these settings: 0.39 s to first byte and a real-time factor of 0.83 while streaming, i.e. faster than real time with about 15% headroom.
Honest limits on the above
- The reference speakers were present in the v14 training data, so the comparison is seen-speaker and flatters the LoRA rather than the base model.
- CER/WER come from Whisper
large-v3, which is biased toward Modern Standard Arabic and penalizes Bahraini dialect in absolute terms. Read the deltas between rows, not the absolute percentages.- The accent metric is spectral-statistical (formants F1/F2/F3, F2-F1, f0 median and range, voiced-segment rate), not perceptual.
- Single training seed, and no human listening test has been run. If careful listening contradicts this table, the table is wrong.
Everything below this box is the original v13 card, left unedited. Its training-run documentation and its correction of the earlier "cosine schedule" claim are still accurate; only the recommendation to use the adapter is not.
LoRA adapter for openbmb/VoxCPM2, fine-tuned on Bahraini / Gulf Arabic speech from wldaldakheel/baharani-mix.
Published checkpoint: step 9000 of 15000, selected by validation loss. This is not the final checkpoint โ the run overfit after step 9000. See Checkpoint selection.
Known limitation: hard ~3.04 s minimum output duration
Every v13 checkpoint refuses to emit audio shorter than about 3.04 s, no matter how short the input text is. Short prompts are stretched, not padded with silence.
The cause is in the training data, not the hyperparameters: the shortest segment in the dataset is exactly 3.00 s, and 37.8% of the 97,087 training rows sit exactly on that floor, so the model never saw a sub-3 s example. Latent frames are 0.16 s and 3.04 s is 19 frames โ every observed duration is an exact multiple of 0.16 s.
Measured on six short (5โ7 word) Arabic prompts at a fixed seed, reference-conditioned:
| sample | 00 | 01 | 02 | 03 | 04 | 05 |
|---|---|---|---|---|---|---|
| untuned base VoxCPM2 | 2.08 | 2.24 | 2.40 | 3.04 | 4.16 | 2.24 |
| v13 step 9000 | 3.04 | 3.20 | 3.04 | 3.04 | 5.12 | 3.04 |
The untuned base tracks text length; v13 floors. If your application needs short utterances (single words, short confirmations, IVR fragments), this adapter is not ready โ that requires a dataset rebuilt with sub-3 s segments. Everything else below is otherwise sound.
Usage
from voxcpm import VoxCPM
from voxcpm.model.voxcpm import LoRAConfig
import soundfile as sf
cfg = LoRAConfig(
enable_lm=True, enable_dit=True, enable_proj=False,
r=32, alpha=64, dropout=0.0,
)
model = VoxCPM.from_pretrained(
"openbmb/VoxCPM2",
lora_weights_path="best", # this repo's best/ directory
lora_config=cfg,
load_denoiser=False, # avoids a hard modelscope dependency
)
# Reference-conditioned (recommended โ voice cloning from a reference clip)
wav = model.generate(
text="ุดุฎุจุงุฑู ุดูููู ุงูููู
ูุง ุฎูู",
reference_wav_path="ref.wav",
seed=1234,
)
sf.write("out.wav", wav, 48000)
# Zero-shot (no reference) โ works, but see the note below
wav = model.generate(text="ุดุฎุจุงุฑู ุดูููู ุงูููู
ูุง ุฎูู", seed=1234)
load_denoiser=False matters: with the denoiser enabled, from_pretrained imports
zipenhancer, which requires modelscope. Without that flag and without modelscope installed,
loading fails with ModuleNotFoundError: No module named 'modelscope'.
The adapter is r=32 / alpha=64 specifically so it can be hot-swapped by servers that hard-code r=32 (upstream VoxCPM issue #283). A rank-64 adapter is not servable there.
Prefer reference-conditioned generation
Zero-shot output from this adapter should not be used to judge quality. The untuned base model
is also erratic in zero-shot mode on these prompts (3 of 6 samples show a low spectral rolloff),
so zero-shot differences are not attributable to the fine-tune. Reference-conditioned output is
level-consistent (โ16 to โ21 dBFS, no level collapse) and shows no degenerate behaviour. Both
sets are in samples/ so you can compare directly.
Checkpoint selection
Validation ran every 3000 steps. Total validation loss bottomed at step 9000 and rose afterwards โ the run overfit over its last 6000 steps:
| step | loss/total | loss/diff | loss/stop |
|---|---|---|---|
| 0 | 1.156592 | 0.981617 | 0.174975 |
| 3000 | 1.023304 | 0.968512 | 0.054792 |
| 6000 | 1.004614 | 0.943909 | 0.060706 |
| 9000 | 0.984550 | 0.942001 | 0.042549 |
| 12000 | 1.007939 | 0.933669 | 0.074270 |
| 14999 | 1.004583 | 0.958660 | 0.045923 |
A second, independent signal agrees. The trainer generates reference-conditioned audio at each validation step; comparing generated duration against the reference duration, step 9000 is the only step where both references are tracked:
| step | ref 4.0 s | ref 6.0 s | ratios |
|---|---|---|---|
| 0 | 3.84 | 7.20 | 0.96 / 1.20 |
| 3000 | 4.00 | 5.12 | 1.00 / 0.85 |
| 6000 | 4.00 | 4.00 | 1.00 / 0.67 |
| 9000 | 4.00 | 6.08 | 1.00 / 1.01 |
| 12000 | 4.00 | 5.12 | 1.00 / 0.85 |
| 14999 | 4.00 | 4.00 | 1.00 / 0.67 |
Steps 6000 and 14999 both collapse a 6.0 s reference to 4.0 s. Step 9000 also retains the most length variation in its own output (5.12 s on sample 04, tracking the base model's 4.16 s, where every other checkpoint flattens to ~3.04 s).
Training recipe (as actually run)
| Parameter | Value |
|---|---|
| Base model | openbmb/VoxCPM2 |
| LoRA rank / alpha | 32 / 64 (ratio 2.0) |
| LoRA dropout | 0.0 |
| Adapted modules | q/k/v/o projections in both LM and DiT; projection layers not adapted |
| Learning rate | 1e-4 |
| LR schedule | 300-step warmup, then effectively constant (see note) |
| Weight decay | 0.01 |
| Max grad norm | 1.0 |
| Steps | 15000 (2.47 epochs); step 9000 published |
| Batch | 2 ร 8 grad-accum = 16 effective, max_batch_tokens 4096 |
| Precision | bfloat16 |
| Checkpoint interval | 1000 steps |
| Validation interval | 3000 steps |
| Dataset | 97,087 train / 1,252 validation segments |
| Audio | 16 kHz in, 48 kHz out |
| Hardware | 1ร A40 46 GB, ~10.5 h |
Note on the LR schedule. The config requests lr_scheduler: cosine, but the cosine horizon
was not tied to the 15000 steps actually run, so the learning rate fell only from 1.00e-4 to
9.5e-5 across the entire run โ a 5% decay. Training was effectively constant-LR. Earlier
Baharani model cards claimed a cosine schedule; that claim was wrong and is not repeated here.
training/train.log records the per-step LR, so this is checkable.
Note on 48 kHz. The VoxCPM2 VAE is 16 kHz in / 48 kHz out. Output is written at 48 kHz, but no genuine content above the 16 kHz input band is recovered, and rebuilding the dataset at a higher sample rate would not change that.
Contents
| Path | What |
|---|---|
best/ |
step 9000 weights (lora_weights.safetensors, 72,397,184 B) + LoRA config |
best/best_step.txt, best/best_val_loss.txt |
9000, 0.984550 |
lora_config.json |
provenance: selected step, val loss, md5, LoRA geometry |
lora_config.yaml |
the exact training config used |
samples/mode_b_step9000/ |
6 reference-conditioned samples, step 9000 |
samples/mode_a_step9000/ |
6 zero-shot samples, step 9000 |
samples/mode_b_base_control/ |
the same 6 prompts on the untuned base โ the control |
samples/durations.json |
measured durations for all three sets |
training/train.log |
full training log: per-step loss, LR, every validation line |
scripts/runner_v13.sh |
the run script as executed (see caveat) |
Deliberately not published
- The final checkpoint (step 15000). Higher validation loss than step 9000, and it collapses a 6.0 s reference to 4.0 s.
- The EMA average. Its per-checkpoint decay placed 86% of the weight on the two worst checkpoints (step 14999 = 0.631, step 14000 = 0.232) โ the opposite of what the validation curve supports. A uniform average over steps 8000โ12000 would be the sound version; it has not been built yet.
- Optimizer state. Not needed to serve or to re-average, and 145 MB per checkpoint.
Caveat on scripts/runner_v13.sh
Published for reproducibility, but it is the version as run, and it contains two bugs that killed
the post-training stage: the sample-generation call omits load_denoiser=False (so it crashes on
the missing modelscope), and the sample-count line reads an empty glob under
set -euo pipefail, which aborts the script. Because of those, the original run never reached its
own upload stage โ this repository was assembled and verified separately. The script also does not
apply an LR schedule over the real step count, as noted above.
Verification
best/lora_weights.safetensors
md5 c611b2f14669713f5ae6ff942f5da0fb
bytes 72397184
source checkpoints/step_0009000, md5-verified against the training host
License
Apache 2.0, matching the base model.
Model tree for 9FA/baharani_v13_lora
Base model
openbmb/VoxCPM2