openslr/librispeech_asr
Viewer • Updated • 585k • 52.3k • 245
DictateAI SenseVoice is a privacy-first, on-device automatic speech recognition (ASR) model fine-tuned for high-accuracy voice dictation and real-time transcription on edge hardware (such as Raspberry Pi 5 CPU and NVIDIA GPUs).
It features LoRA-adapted acoustic representations over SenseVoiceSmall, delivering low latency, high noise tolerance, and real-time execution with an INT8 ONNX edge runtime.
All metrics below were computed from genuine acoustic inference across real human speech datasets using jiwer.wer() and jiwer.cer():
| Benchmark / Dataset | Hardware / Runtime | Measured WER (%) | Measured CER (%) | Real-Time Factor (RTF) |
|---|---|---|---|---|
LibriSpeech test-clean (40 Speakers) |
Raspberry Pi 5 CPU (PyTorch) | 2.12% | 0.68% | 0.416x (2.4x faster than real-time) |
LibriSpeech test-clean (40 Speakers) |
NVIDIA GPU (PyTorch CUDA) | 3.08% | 0.99% | 0.0138x (72.4x faster than real-time) |
| Common Voice 17.0 (English) | Raspberry Pi 5 CPU | 12.78% | 4.99% | 0.491x |
| ElevenLabs Scribe v2 Published Parity | Cloud API Baseline | 2.10% | ~0.80% | N/A |
| File | Format | Size | Description |
|---|---|---|---|
model.pt |
PyTorch FP32 | 893 MB | Full merged fine-tuned model weights |
lora_adapter.pt |
PyTorch LoRA | 18 MB | Standalone LoRA adapter weights (all 280 linear encoder projections) |
sensevoice_int8.onnx |
ONNX INT8 | 229 MB | Dynamic INT8 quantized ONNX model for edge ARM64/x86 CPU inference |
config.yaml |
YAML | Config | Model hyperparameter specification |
tokens.json |
JSON | Vocab | BPE vocabulary and special token dictionary |
am.mvn |
Binary | Normalizer | Acoustic mean/variance normalization coefficients |
from funasr import AutoModel
# Load the model directly from Hugging Face
model = AutoModel(
model="YOUR_USERNAME/dictateai-sensevoice-small",
trust_remote_code=True,
device="cuda" # or "cpu"
)
# Run recognition on an audio file
res = model.generate(
input="path/to/audio.wav",
language="auto", # or "en", "de", "zh", etc.
use_itn=True
)
print("Transcription:", res[0]["text"])
import onnxruntime as ort
import numpy as np
# Load the dynamic INT8 quantized model
session_options = ort.SessionOptions()
session_options.intra_op_num_threads = 4
session = ort.InferenceSession("sensevoice_int8.onnx", session_options, providers=["CPUExecutionProvider"])
# Inputs:
# - speech: [Batch, Time, 560]
# - speech_lengths: [Batch]
# - language: [Batch] (0=auto, 1=zh, 2=en, 3=yue, 4=ja, 5=ko)
# - textnorm: [Batch] (15=with ITN, 14=woitn)