Nemotron 3.5 Core ML encoder export experiment
Experimental 320 ms Nemotron bundle for testing the ANE specialization failure on A16 described in speech-swift issue 503. Only the encoder conversion target changes from iOS 18 to iOS 17. The export preserves all 296 published encoder palettes and copies the decoder, joint network, and runtime metadata from the published bundle without changes.
| Property | Value |
|---|---|
| Architecture | Cache-aware FastConformer encoder and RNN-T decoder |
| Parameters | 600 million |
| Format | Compiled Core ML .mlmodelc |
| Encoder weights | Published 8-bit palette indices and FP16 lookup tables |
| Decoder and joint weights | FP16, unchanged |
| Runtime file size | 613 MiB |
| Audio | 16 kHz mono |
| Streaming chunk | 320 ms |
| Encoder minimum deployment target | iOS 17 |
| Decoder and joint minimum deployment target | iOS 18, unchanged |
| Baseline revision | 447095fe87b480b5e6a15367f135303d479de8ac |
| Status | Experimental; affected-iPhone validation pending |
This bundle does not lower the Swift SDK's iOS 18 requirement. A successful macOS load does not establish compatibility with the affected iPhone.
Files
| File | Size | Purpose |
|---|---|---|
encoder.mlmodelc/ |
565.4 MiB | Encoder compiled for the iOS 17 operation set |
decoder.mlmodelc/ |
28.5 MiB | Unchanged published RNN-T decoder |
joint.mlmodelc/ |
18.0 MiB | Unchanged published joint network |
config.json |
589 B | Unchanged streaming geometry |
vocab.json |
230.6 KiB | Unchanged vocabulary |
languages.json |
2.0 KiB | Unchanged language map |
tokenizer.model |
397.0 KiB | Unchanged SentencePiece tokenizer |
experiment.json |
Small JSON; exact inventory included | Source revisions, export versions, and artifact hashes |
TESTING.md |
6.5 KiB | Published versus candidate A/B instructions |
testing/LoadProbe.swift |
7.8 KiB | Standalone iOS/macOS Core ML load probe |
testing/test_audio.wav |
625.0 KiB | Real-speech smoke-test fixture, 20 s, mono 16 kHz |
testing/mac-validation.json |
30.9 KiB | Local load, encoder-output, and SDK test results |
testing/fixture-provenance.json |
<1 KiB | Original fixture and resampling hashes |
Validation
M5 Pro, macOS 26.6.2 (25G83):
- All 296 encoder palette lookup tables and index arrays match the published values byte for byte. The decoder, joint, and runtime metadata are unchanged.
- All six encoder outputs match the published CPU encoder exactly across 67 calls with real speech, silence, partial chunks, and full streaming caches.
- Four SDK tests passed in each of three separate processes: published CPU+ANE, candidate CPU+ANE, and candidate CPU-only. Batch, streaming, and word-boosted transcripts matched the published result. The boosting engagement test also passed in all three configurations.
- The supplied probe compiles with Swift 6 on macOS and typechecks for arm64 iOS 18. The export-helper suite passed all 14 unit tests.
The SDK tests used speech-swift revision
7fc8f6c2b7847cad17641cf294b2854e20d936a8. Validation covers one English
fixture; it does not establish multilingual accuracy or A16 compatibility.
Local load and encoder timing
The standalone Core ML probe ran one bundle and compute configuration per process. Prediction measurements use synthetic input, not full ASR.
| Bundle / compute units | First observed encoder load | Same-process reload | Median repeated encoder prediction |
|---|---|---|---|
| Published / CPU+ANE | 10.53 s | 76.5 ms | 8.09 ms |
| Candidate / CPU+ANE, first process | 6.63 s | 62.7 ms | 9.10 ms |
| Candidate / CPU+ANE, second process | 130.0 ms | 65.4 ms | 8.45 ms |
| Candidate / CPU-only | 2.77 s | 53.1 ms | 15.50 ms |
Core ML system-cache state was uncontrolled, so first-load times are observations rather than a speedup claim. CPU+ANE compute plans for both encoders prefer the ANE for 1,592 operations and CPU for 86. Planned placement does not prove execution on the ANE. Preserve device compiler logs alongside the reports.
See testing/mac-validation.json for the measurements. Repeated loads and
real-speech tests on the affected iPhone remain the acceptance check.
Usage
Download the snapshot into a separate directory and follow TESTING.md.
On macOS, the Python Core ML loader can check the candidate encoder directly:
import coremltools as ct
from huggingface_hub import snapshot_download
bundle = snapshot_download(
"aufklarer/Nemotron-3.5-ASR-Streaming-0.6B-CoreML-INT8-iOS17-Encoder-Test",
local_dir="nemotron-ios17-encoder-test",
)
encoder = ct.models.CompiledMLModel(
f"{bundle}/encoder.mlmodelc",
compute_units=ct.ComputeUnit.CPU_AND_NE,
)
To obtain a JSON load report on macOS:
swiftc -O -parse-as-library -D LOAD_PROBE_CLI nemotron-ios17-encoder-test/testing/LoadProbe.swift -o /tmp/nemotron-load-probe
/tmp/nemotron-load-probe nemotron-ios17-encoder-test ane candidate-ane.json
With speech-swift v0.0.27 or later, use the unchanged local-bundle API:
import CoreML
import NemotronStreamingASR
let model = try await NemotronStreamingASRModel.fromLocal(
bundleDir: candidateBundleURL,
computeUnits: .cpuAndNeuralEngine)
Preserve the stock bundle for the baseline comparison. The model remains an experiment until the affected iPhone passes repeated-load and real-speech batch/streaming checks.
Source and license
Upstream: NVIDIA Nemotron 3.5 ASR Streaming 0.6B,
revision f3d333391852ba876df169dcc9ba902d25b6ab0b.
Baseline: published Core ML INT8 bundle. The model uses the upstream OpenMDW 1.1 license.
Links
- Downloads last month
- -
Model tree for aufklarer/Nemotron-3.5-ASR-Streaming-0.6B-CoreML-INT8-iOS17-Encoder-Test
Base model
nvidia/nemotron-3.5-asr-streaming-0.6b