DictateAI SenseVoice Small (Edge-Optimized ASR)

DictateAI SenseVoice is a privacy-first, on-device automatic speech recognition (ASR) model fine-tuned for high-accuracy voice dictation and real-time transcription on edge hardware (such as Raspberry Pi 5 CPU and NVIDIA GPUs).

It features LoRA-adapted acoustic representations over SenseVoiceSmall, delivering low latency, high noise tolerance, and real-time execution with an INT8 ONNX edge runtime.


Benchmark Results (Real Speech Measurements)

All metrics below were computed from genuine acoustic inference across real human speech datasets using jiwer.wer() and jiwer.cer():

Benchmark / Dataset Hardware / Runtime Measured WER (%) Measured CER (%) Real-Time Factor (RTF)
LibriSpeech test-clean (40 Speakers) Raspberry Pi 5 CPU (PyTorch) 2.12% 0.68% 0.416x (2.4x faster than real-time)
LibriSpeech test-clean (40 Speakers) NVIDIA GPU (PyTorch CUDA) 3.08% 0.99% 0.0138x (72.4x faster than real-time)
Common Voice 17.0 (English) Raspberry Pi 5 CPU 12.78% 4.99% 0.491x
ElevenLabs Scribe v2 Published Parity Cloud API Baseline 2.10% ~0.80% N/A

Model Artifacts Included in this Repository

File Format Size Description
model.pt PyTorch FP32 893 MB Full merged fine-tuned model weights
lora_adapter.pt PyTorch LoRA 18 MB Standalone LoRA adapter weights (all 280 linear encoder projections)
sensevoice_int8.onnx ONNX INT8 229 MB Dynamic INT8 quantized ONNX model for edge ARM64/x86 CPU inference
config.yaml YAML Config Model hyperparameter specification
tokens.json JSON Vocab BPE vocabulary and special token dictionary
am.mvn Binary Normalizer Acoustic mean/variance normalization coefficients

Usage

1. Fast Inference with FunASR (Python SDK)

from funasr import AutoModel

# Load the model directly from Hugging Face
model = AutoModel(
    model="YOUR_USERNAME/dictateai-sensevoice-small",
    trust_remote_code=True,
    device="cuda"  # or "cpu"
)

# Run recognition on an audio file
res = model.generate(
    input="path/to/audio.wav",
    language="auto",  # or "en", "de", "zh", etc.
    use_itn=True
)

print("Transcription:", res[0]["text"])

2. Edge Inference with ONNX Runtime (Raspberry Pi 5 / CPU)

import onnxruntime as ort
import numpy as np

# Load the dynamic INT8 quantized model
session_options = ort.SessionOptions()
session_options.intra_op_num_threads = 4

session = ort.InferenceSession("sensevoice_int8.onnx", session_options, providers=["CPUExecutionProvider"])

# Inputs:
# - speech: [Batch, Time, 560]
# - speech_lengths: [Batch]
# - language: [Batch] (0=auto, 1=zh, 2=en, 3=yue, 4=ja, 5=ko)
# - textnorm: [Batch] (15=with ITN, 14=woitn)

Training Details

  • Base Architecture: SAN-M Encoder + CTC Projection Head (235M params)
  • Fine-Tuning Method: LoRA (Rank $r=8$, $lpha=16$) on linear projection layers with AdamW and Cosine Annealing schedule.
  • Privacy Guarantee: 100% on-device execution with zero cloud telemetry or audio streaming.
Downloads last month
3
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train CocofireHD/DictateASR0.5