Lmaana 2.2
Lmaana 2.2 is a Moroccan Darija automatic speech recognition model built by fully fine-tuning OmniASR CTC 3B v2 with a CTC objective. It was trained on Dataset13 clean v4 with two-node FSDP and evaluated with greedy CTC decoding.
The released checkpoint is step_10000, the final checkpoint in the planned
training schedule. All validation and test results below were computed with
this reproducible saved checkpoint.
Model summary
| Field | Value |
|---|---|
| Release | Lmaana 2.2 |
| Base model | OmniASR CTC 3B v2 |
| Adaptation | Full fine-tuning with CTC |
| Target language | Moroccan Darija (ary-Arab) |
| Training data | Dataset13 clean v4 |
| Training setup | Two-node FSDP, one L40S GPU per node |
| Planned training budget | 10,000 steps |
| Released checkpoint | Step 10,000 |
| Tokenizer | OmniASR written tokenizer v2 |
| Decoder | Greedy CTC |
Dataset
Dataset13 clean v4 contains approximately 330.74 hours of Moroccan Darija
speech. Audio duration is derived from audio_size at 16 kHz.
| Split | Parquet files | Rows | Hours |
|---|---|---|---|
| Train | 53 | 52,907 | 263.84 |
| Validation | 7 | 6,894 | 33.70 |
| Test | 7 | 6,576 | 33.20 |
The validation reader reported 6,893 evaluated examples, one fewer than the 6,894 physical validation rows. Release metrics always use the number of examples actually processed by the evaluator. The analysis notebooks audit durations, missing values, duplicates, text length, and Latin-character usage as a simple proxy for code-switching.
Results
The validation and test sets are disjoint Dataset13 clean v4 splits. UER is the tokenizer-unit error rate reported by fairseq2 and is distinct from CER. Test CER was computed at corpus level from the normalized references and greedy hypotheses after removing whitespace.
| Split | Examples | CTC loss | UER (%) | CER (%) | WER (%) |
|---|---|---|---|---|---|
| Validation | 6,893 | 155.0191 | 15.5020 | Not measured | 38.0618 |
| Test | 6,576 | 160.8470 | 15.7283 | 16.3063 | 38.4762 |
The validation-to-test WER gap is 0.4144 percentage points, which indicates close agreement between the two splits under the same normalization and greedy decoding protocol.
Compared with the previous step_8000 candidate, test WER improved by
0.6869 percentage points, from 39.1631% to 38.4762%.
Cross-corpus evaluation
The unchanged Dataset13-trained step_10000 checkpoint was also evaluated on
the frozen Lmaana clean v1 (MoulSot) test split. No MoulSot example was used to
train this release. The split contains 1,957 examples from six held-out sources
and approximately 3.46 hours of audio; its audit reports no source overlap with
the Lmaana clean training split.
| Test corpus | Examples | CTC loss | UER (%) | CER (%) | WER (%) |
|---|---|---|---|---|---|
| Dataset13 clean v4 | 6,576 | 160.8470 | 15.7283 | 16.3063 | 38.4762 |
| Lmaana clean v1 (MoulSot) | 1,957 | 75.2370 | 13.5374 | 13.9285 | 40.7776 |
MoulSot has lower UER and CER but higher WER. This pattern motivates a separate analysis of word boundaries, orthographic conventions, and word-level error operations before any mixed-corpus fine-tuning.
Checkpoint evolution
Validation improved throughout the run. WER is the primary selection metric; UER and CTC loss are supporting diagnostics.
| Step | Validation CTC loss | Validation UER (%) | Validation WER (%) |
|---|---|---|---|
| 1,000 | 205.0993 | 19.5169 | 48.2411 |
| 3,000 | 175.3598 | 17.4674 | 43.2970 |
| 5,000 | 168.1747 | 16.5931 | 41.1742 |
| 7,000 | 158.0662 | 15.7977 | 38.8891 |
| 8,000 | 156.4570 | 15.6952 | 38.7391 |
| 10,000 | 155.0191 | 15.5020 | 38.0618 |
Training dynamics
Validation WER decreased throughout the run, from 48.24% at step 1,000 to 38.06% at the final step 10,000 checkpoint. The learning rate began decaying after step 5,000, and no final-stage validation regression was observed.
Training configuration
| Setting | Value |
|---|---|
| GPUs | 2 x NVIDIA L40S |
| Distributed strategy | Two-node FSDP, one process per GPU |
| Precision | bfloat16 automatic mixed precision |
| Optimizer learning rate | 1e-5 |
| Gradient accumulation | 8 batches |
| Maximum audio length | 320,000 samples, or 20 seconds at 16 kHz |
| Validation interval | 500 steps |
| Checkpoint interval | 1,000 steps |
| Training schedule | 10,000 steps |
Evaluation protocol
- Dataset13 clean v4 validation and test partitions are disjoint.
- Audio normalization is enabled.
- Evaluation uses bfloat16 automatic mixed precision on one GPU.
- Decoding is greedy CTC without a language model or beam search.
- WER and CER are corpus-level metrics.
- CER removes whitespace before computing character edit distance.
- UER is computed on tokenizer units and is reported separately from CER.
- The test set contains 6,576 examples and approximately 33.20 hours.
Intended use
This model is intended for:
- research on Moroccan Darija speech recognition;
- evaluation of CTC-based ASR systems;
- transcription experiments on audio similar to Dataset13;
- reproducible comparison with other Darija ASR systems.
It should not be treated as a certified transcription service, a general Arabic speech recognizer, or a system suitable for high-stakes decisions without task-specific evaluation.
Limitations
- Dataset13 may not represent all Moroccan accents, speakers, recording conditions, or code-switching patterns.
- WER is sensitive to text normalization, spelling conventions, punctuation, and segmentation.
- Performance may degrade on noisy audio, reverberant recordings, children's speech, rare names, and domains absent from training data.
- The model should be evaluated separately on code-switched and non-code-switched subsets before making deployment claims.
Access
The Model Card, evaluation results, plots, and research code are public. Model weights are gated on Hugging Face. Users must request access and remain responsible for complying with the base-model and dataset terms.
Reproduce the evaluation
The release was evaluated through the OmniASR fairseq2 workflow. Clone this
repository alongside a compatible omnilingual-asr checkout, install the
analysis and Hub dependencies, and adapt the absolute model, tokenizer, and
dataset paths in the configuration:
git clone https://github.com/BADR-JOULAlI/darija-asr-v2.git
cd darija-asr-v2
python -m pip install -e ".[analysis,hub]"
python -u -m workflows.recipes.wav2vec2.asr.eval \
outputs/evaluations/omniasr_ctc_3b_dataset13_step10000_test \
--config-file configs/omniasr_ctc_3b/test_step10000.yaml
The provided PBS launcher performs FSDP consolidation when needed and runs the same evaluation recipe:
qsub scripts/omniasr_ctc_3b_test_step10000.pbs
The released artifact can be checked after download with:
echo "b22dd8ece46eb24316a96d342f53f8864881b6827ecdec9923006b03ad806471 consolidated.pt" \
| sha256sum --check
Analysis notebooks
- Dataset audit
- Training dynamics
- Final results report
- All-checkpoint comparison
- Cross-corpus error analysis
Reproducibility
The public research repository contains versioned configurations, the FSDP launcher, the consolidation utility, evaluation recipe, notebooks, and figure generator:
https://github.com/BADR-JOULAlI/darija-asr-v2
The two FSDP model shards were consolidated before single-GPU evaluation. The
released consolidated.pt checksum is:
SHA-256 b22dd8ece46eb24316a96d342f53f8864881b6827ecdec9923006b03ad806471
The original audio data remains outside Git. Reproduction requires authorized access to Dataset13 clean v4 and OmniASR written tokenizer v2.
Base model and data attribution
This work is an adaptation of OmniASR CTC 3B v2. Users must review and comply with the base model license, the Dataset13 terms, and the terms of any other source corpora before downloading or redistributing the weights.
Citation
@misc{lmaana_2_2,
title = {Lmaana 2.2: CTC Speech Recognition for Moroccan Darija},
author = {Skiredj, Abderrahman},
year = {2026},
note = {Dataset13 clean v4; released checkpoint at step 10000}
}
Evaluation results
- Test WER on Dataset13 clean v4test set self-reported38.476
- Test CER on Dataset13 clean v4test set self-reported16.306
- Cross-corpus Test WER on Lmaana clean v1 (MoulSot)test set self-reported40.778
- Cross-corpus Test CER on Lmaana clean v1 (MoulSot)test set self-reported13.928




