You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Lmaana 2.2

Lmaana 2.2 is a Moroccan Darija automatic speech recognition model built by fully fine-tuning OmniASR CTC 3B v2 with a CTC objective. It was trained on Dataset13 clean v4 with two-node FSDP and evaluated with greedy CTC decoding.

The released checkpoint is step_10000, the final checkpoint in the planned training schedule. All validation and test results below were computed with this reproducible saved checkpoint.

Model summary

Field Value
Release Lmaana 2.2
Base model OmniASR CTC 3B v2
Adaptation Full fine-tuning with CTC
Target language Moroccan Darija (ary-Arab)
Training data Dataset13 clean v4
Training setup Two-node FSDP, one L40S GPU per node
Planned training budget 10,000 steps
Released checkpoint Step 10,000
Tokenizer OmniASR written tokenizer v2
Decoder Greedy CTC

Dataset

Dataset13 clean v4 contains approximately 330.74 hours of Moroccan Darija speech. Audio duration is derived from audio_size at 16 kHz.

Split Parquet files Rows Hours
Train 53 52,907 263.84
Validation 7 6,894 33.70
Test 7 6,576 33.20

The validation reader reported 6,893 evaluated examples, one fewer than the 6,894 physical validation rows. Release metrics always use the number of examples actually processed by the evaluator. The analysis notebooks audit durations, missing values, duplicates, text length, and Latin-character usage as a simple proxy for code-switching.

Results

The validation and test sets are disjoint Dataset13 clean v4 splits. UER is the tokenizer-unit error rate reported by fairseq2 and is distinct from CER. Test CER was computed at corpus level from the normalized references and greedy hypotheses after removing whitespace.

Split Examples CTC loss UER (%) CER (%) WER (%)
Validation 6,893 155.0191 15.5020 Not measured 38.0618
Test 6,576 160.8470 15.7283 16.3063 38.4762

The validation-to-test WER gap is 0.4144 percentage points, which indicates close agreement between the two splits under the same normalization and greedy decoding protocol.

Compared with the previous step_8000 candidate, test WER improved by 0.6869 percentage points, from 39.1631% to 38.4762%.

Validation and test comparison

Cross-corpus evaluation

The unchanged Dataset13-trained step_10000 checkpoint was also evaluated on the frozen Lmaana clean v1 (MoulSot) test split. No MoulSot example was used to train this release. The split contains 1,957 examples from six held-out sources and approximately 3.46 hours of audio; its audit reports no source overlap with the Lmaana clean training split.

Test corpus Examples CTC loss UER (%) CER (%) WER (%)
Dataset13 clean v4 6,576 160.8470 15.7283 16.3063 38.4762
Lmaana clean v1 (MoulSot) 1,957 75.2370 13.5374 13.9285 40.7776

MoulSot has lower UER and CER but higher WER. This pattern motivates a separate analysis of word boundaries, orthographic conventions, and word-level error operations before any mixed-corpus fine-tuning.

Checkpoint evolution

Validation improved throughout the run. WER is the primary selection metric; UER and CTC loss are supporting diagnostics.

Step Validation CTC loss Validation UER (%) Validation WER (%)
1,000 205.0993 19.5169 48.2411
3,000 175.3598 17.4674 43.2970
5,000 168.1747 16.5931 41.1742
7,000 158.0662 15.7977 38.8891
8,000 156.4570 15.6952 38.7391
10,000 155.0191 15.5020 38.0618

Checkpoint trade-off

Training dynamics

Validation WER decreased throughout the run, from 48.24% at step 1,000 to 38.06% at the final step 10,000 checkpoint. The learning rate began decaying after step 5,000, and no final-stage validation regression was observed.

Training dynamics

Validation progress

Learning-rate schedule

Training configuration

Setting Value
GPUs 2 x NVIDIA L40S
Distributed strategy Two-node FSDP, one process per GPU
Precision bfloat16 automatic mixed precision
Optimizer learning rate 1e-5
Gradient accumulation 8 batches
Maximum audio length 320,000 samples, or 20 seconds at 16 kHz
Validation interval 500 steps
Checkpoint interval 1,000 steps
Training schedule 10,000 steps

Evaluation protocol

  • Dataset13 clean v4 validation and test partitions are disjoint.
  • Audio normalization is enabled.
  • Evaluation uses bfloat16 automatic mixed precision on one GPU.
  • Decoding is greedy CTC without a language model or beam search.
  • WER and CER are corpus-level metrics.
  • CER removes whitespace before computing character edit distance.
  • UER is computed on tokenizer units and is reported separately from CER.
  • The test set contains 6,576 examples and approximately 33.20 hours.

Intended use

This model is intended for:

  • research on Moroccan Darija speech recognition;
  • evaluation of CTC-based ASR systems;
  • transcription experiments on audio similar to Dataset13;
  • reproducible comparison with other Darija ASR systems.

It should not be treated as a certified transcription service, a general Arabic speech recognizer, or a system suitable for high-stakes decisions without task-specific evaluation.

Limitations

  • Dataset13 may not represent all Moroccan accents, speakers, recording conditions, or code-switching patterns.
  • WER is sensitive to text normalization, spelling conventions, punctuation, and segmentation.
  • Performance may degrade on noisy audio, reverberant recordings, children's speech, rare names, and domains absent from training data.
  • The model should be evaluated separately on code-switched and non-code-switched subsets before making deployment claims.

Access

The Model Card, evaluation results, plots, and research code are public. Model weights are gated on Hugging Face. Users must request access and remain responsible for complying with the base-model and dataset terms.

Reproduce the evaluation

The release was evaluated through the OmniASR fairseq2 workflow. Clone this repository alongside a compatible omnilingual-asr checkout, install the analysis and Hub dependencies, and adapt the absolute model, tokenizer, and dataset paths in the configuration:

git clone https://github.com/BADR-JOULAlI/darija-asr-v2.git
cd darija-asr-v2
python -m pip install -e ".[analysis,hub]"

python -u -m workflows.recipes.wav2vec2.asr.eval \
  outputs/evaluations/omniasr_ctc_3b_dataset13_step10000_test \
  --config-file configs/omniasr_ctc_3b/test_step10000.yaml

The provided PBS launcher performs FSDP consolidation when needed and runs the same evaluation recipe:

qsub scripts/omniasr_ctc_3b_test_step10000.pbs

The released artifact can be checked after download with:

echo "b22dd8ece46eb24316a96d342f53f8864881b6827ecdec9923006b03ad806471  consolidated.pt" \
  | sha256sum --check

Analysis notebooks

Reproducibility

The public research repository contains versioned configurations, the FSDP launcher, the consolidation utility, evaluation recipe, notebooks, and figure generator:

https://github.com/BADR-JOULAlI/darija-asr-v2

The two FSDP model shards were consolidated before single-GPU evaluation. The released consolidated.pt checksum is:

SHA-256  b22dd8ece46eb24316a96d342f53f8864881b6827ecdec9923006b03ad806471

The original audio data remains outside Git. Reproduction requires authorized access to Dataset13 clean v4 and OmniASR written tokenizer v2.

Base model and data attribution

This work is an adaptation of OmniASR CTC 3B v2. Users must review and comply with the base model license, the Dataset13 terms, and the terms of any other source corpora before downloading or redistributing the weights.

Citation

@misc{lmaana_2_2,
  title  = {Lmaana 2.2: CTC Speech Recognition for Moroccan Darija},
  author = {Skiredj, Abderrahman},
  year   = {2026},
  note   = {Dataset13 clean v4; released checkpoint at step 10000}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results