CAM++ zh-CN common โ€” LiteRT

3D-Speaker CAM++ (campplus_cn_common, 192-d) converted PyTorch โ†’ LiteRT via litert-torch, static input (1, 500, 80), fp32.

Front end (matches the validated ggml/rapidspeech pipeline): kaldi log-mel fbank โ€” 25 ms / 10 ms, povey window, 80 bins (20 Hzโ€“8 kHz, kaldi mel scale), dither 0, waveform ร—2ยนโต, natural log โ€” then global CMN over the utterance, sliced into 500-frame windows; per-window embeddings are L2-normalized, mean-pooled and L2-normalized again.

Parity gates (2026-07-24):

  • tflite vs PyTorch source, 14-input battery (9 real-speech fbank windows + 5 random): worst element-wise max|ฮ”| 1.7e-05, worst cosine 0.99999994.
  • Windowed tflite pipeline vs the rapidspeech ggml CAM++ reference embedding (10 s real speech): cosine 0.9924 โ€” cross-window speaker-linking thresholds tuned on the ggml pipeline carry over.

Converted for VoxSumDroid's all-LiteRT migration.

Downloads last month
24
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support