Lmaana 2.1
Lmaana 2.1 is a research automatic speech recognition checkpoint for Moroccan Darija. It is a conservative continuation of the retained Lmaana V5 model, adapted for 5,000 steps on Dataset13 and Lmaana clean. The latter is derived from MoulSot and contains revised transcripts for Arabic/Latin code switching, including corrected French expressions.
Release format: native fairseq2/OmniASR model shard. This repository is not a Transformers
from_pretrained()export and does not currently provide a hosted Hugging Face inference widget.
At a Glance
| Property | Value |
|---|---|
| Task | Automatic speech recognition (speech to text) |
| Primary language | Moroccan Darija (ary), including Arabic/Latin code switching |
| Architecture | OmniASR CTC 1b_v2 (wav2vec2_asr) |
| Parameters | 975,065,300 in the upstream 1B v2 architecture |
| Framework | fairseq2 / OmniASR |
| Starting checkpoint | Retained Lmaana V5 |
| Released checkpoint | Step 5,000 |
| Tokenizer | omniASR_tokenizer_written_v2 character tokenizer |
| Evaluation decoding | Greedy CTC, no external language model |
The released checkpoint improves Dataset13 test WER from 44.2243% to
44.0393%. On Lmaana clean, character-level error is essentially preserved,
while WER changes from 45.1385% to 45.1690%. This is a measured incremental
release, not a claim of a large universal quality gain.
Key Test Results
The test partitions were evaluated once, after checkpoint selection using only the validation partitions.
| Model | Test corpus | CTC loss | CER/UER | WER | Examples |
|---|---|---|---|---|---|
| Retained V5 | Dataset13 | 183.8770 | 17.7674% | 44.2243% | 6,576 |
| Lmaana 2.1 | Dataset13 | 183.6260 | 17.7029% | 44.0393% | 6,576 |
| Retained V5 | Lmaana clean | 75.2100 | 14.6099% | 45.1385% | 1,957 |
| Lmaana 2.1 | Lmaana clean | 75.0515 | 14.6048% | 45.1690% | 1,957 |
Relative to V5, Lmaana 2.1 improves Dataset13 test CER/UER by 0.0645
percentage point and WER by 0.1850 point. Lmaana clean test CER/UER improves
by 0.0051 point, while WER regresses by 0.0305 point. It is therefore the
best balanced continuation selected in this experiment. It is not an improvement on
every individual metric.
Model Lineage
Lmaana 2.1 uses the architecture and tokenizer of Meta's OmniASR CTC 1B v2. It does not start directly from the official base checkpoint: it continues the retained Lmaana V5 checkpoint, which already contains earlier Darija adaptation. The published weights are therefore a derived research artifact and should not be presented as an official Meta release.
The historical figure contains two evaluation phases. Lmaana v0-v2 were evaluated on Dataset13 and the original MoulSot partitions. Later V5 and Lmaana 2.1 results use Dataset13 and Lmaana clean. Corrected labels and changed test membership make absolute MoulSot and Lmaana-clean scores non-comparable. Trends should be interpreted within each evaluation phase.
Training Data
| Corpus | Training examples | Duration | Effective sampling weight |
|---|---|---|---|
| Dataset13 | 52,892 | 263.76 h | 74.21% |
| Lmaana clean | 74,243 | 91.68 h | 25.79% |
Dataset13 contributes most of the acoustic duration and has segments averaging
about 18.0 s. Lmaana clean contributes denser supervision from shorter
segments averaging about 4.4 s. Validation and test partitions were excluded
from training.
Lmaana clean is based on MoulSot audio with revised transcripts, especially for French and Latin-script spans in code-switched speech. Dataset terms, consent, and redistribution rights must be reviewed separately from the model weights.
Adaptation Procedure
A previous curriculum initialized from the official OmniASR base underperformed the retained V5 model. Lmaana 2.1 therefore uses a guarded continuation from V5:
| Parameter | Value |
|---|---|
| Training steps | 5,000 |
| Peak learning rate | 1e-7 |
| Encoder freeze | First 500 steps |
| Gradient accumulation | 8 batches |
beta_corpus / beta_language |
1.0 / 1.0 |
| Maximum training audio length | 320,000 samples (20 s at 16 kHz) |
| Maximum batch elements | 1,280,000 |
| Audio normalization | Enabled |
| Validation/checkpoint interval | 500 steps |
The retained V5 model was included as step 0 in checkpoint selection. A
candidate could regress by at most 0.05 CER/UER point and 0.10 WER point on
either validation set. Step 5,000 ranked first on the four validation metrics.
Evaluation Protocol
- Checkpoint selection used Dataset13 and Lmaana-clean validation splits.
- The held-out test splits were evaluated only after selection.
- Decoding was greedy CTC with no language model, rescoring, or beam search.
- All comparisons in a table use the same tokenizer, evaluator, and split.
UERis the metric name emitted by fairseq2. With this character tokenizer, it is reported here asCER/UER; it should not be assumed identical to every external CER implementation without matching normalization rules.- No confidence intervals or multi-seed variance estimates are available.
Dataset13 and Lmaana-clean numbers are not repeated measurements of one corpus. They differ in source material, segment length, transcript conventions, and code-switching distribution. The strongest evidence is the within-dataset comparison between retained V5 and Lmaana 2.1.
Validation Results
| Checkpoint | Dataset13 CER/UER | Dataset13 WER | Lmaana CER/UER | Lmaana WER |
|---|---|---|---|---|
| Retained V5 baseline | 17.3908% | 43.2622% | 13.1994% | 40.4155% |
| Step 3,500 | 17.3347% | 43.1245% | 13.2053% | 40.4015% |
| Lmaana 2.1, step 5,000 | 17.3252% | 43.1238% | 13.1925% | 40.3535% |
Both retained adaptation checkpoints satisfy the validation guard. Step 5,000 is furthest toward the lower-left optimum in the trade-off figure, although the absolute changes remain small.
Training Dynamics
The figures are generated from the fairseq2 logs. Training loss varies with sampled batches while validation loss remains nearly flat. Together with the small WER changes, this indicates late-stage adaptation near a plateau. The process comparison includes runs with different initialization points; it compares trajectories, not equal-cost training experiments.
Executed Report Snapshots
The following snapshots come from the executed final-analysis report. They are included as visual evidence of the rendered notebook outputs; the generated figures above remain the canonical, publication-quality plots.
Training and validation CTC loss
Comparison of adaptation trajectories
Qualitative transcription examples
The qualitative table groups examples whose token-level error count improved, regressed, or remained unchanged relative to V5. These examples are illustrative and were selected from automatic alignment outputs. They should not be treated as a human listening study or as a representative estimate of all error types.
Download
Install the download client and OmniASR runtime:
python -m pip install --upgrade huggingface_hub omnilingual-asr
Download the complete repository:
from huggingface_hub import snapshot_download
local_path = snapshot_download(repo_id="sailu4/lmaana-2.1")
print(local_path)
The selected weights are stored at
checkpoint/model/pp_00/tp_00/sdp_00.pt, and the matching tokenizer is under
tokenizer/. metadata/ contains the training configuration, checkpoint
selection, validation summary, and test summary. manifest.json records file
sizes and SHA-256 hashes.
Inference and Integration
The checkpoint uses the upstream wav2vec2_asr 1b_v2 architecture, but it is
published as a native fairseq2 model shard rather than a Transformers model.
Consumers must register a local fairseq2 model asset that points to the
downloaded checkpoint and tokenizer, then load it through the OmniASR
ASRInferencePipeline or the matching fairseq2 ASR recipe. See the
official OmniASR inference documentation
for the pipeline API and the
project evaluation script
for the exact evaluation integration used in this release.
Recommended input is single-channel speech resampled to 16 kHz. Training examples were capped at 20 seconds; longer recordings should be segmented before transcription. Upstream OmniASR currently documents a sub-40-second limit for CTC inference.
The included metadata/lmaana_clean_v5_long_model_asset.yaml documents the
training initialization asset. Its original machine path is not a portable
inference path; replace it with the absolute path of the downloaded Lmaana 2.1
checkpoint when creating a local asset card.
Intended Use
- Research on Moroccan Darija speech recognition.
- Transcription experiments with Arabic/Latin code switching.
- Evaluation of transcript cleaning and mixed-replay adaptation.
- Continued fairseq2/OmniASR adaptation with a newly initialized optimizer.
Out-of-Scope Use
This checkpoint has not been validated for medical, legal, financial, safety, or other high-stakes decisions. It is not a speaker-identification, diarization, translation, or calibrated confidence-scoring system. Do not use it for covert surveillance or to process speech without an appropriate legal basis and the required consent.
Limitations and Risks
- Improvements over V5 are small and come from one training run.
- Greedy CTC decoding can omit, repeat, or substitute words and can produce plausible but incorrect text.
- Performance may vary across regions, speakers, ages, recording devices, background noise, speaking styles, and code-switching patterns.
- The evaluation sets do not establish fairness across demographic groups.
- French and Latin-script corrections improve label consistency but do not guarantee correct normalization for every code-switched expression.
- The model does not provide calibrated uncertainty. Human review remains necessary for consequential transcripts.
License and Attribution
The repository is marked license: other because the release combines a
derived OmniASR checkpoint, tokenizer, project code, and datasets with separate
terms. Users must review the
upstream OmniASR repository and license
as well as the applicable Dataset13, MoulSot, and Lmaana data terms. This card
does not grant rights to redistribute source audio or third-party data.
Reproducibility
- Source code and protocol: BADR-JOULAlI/darija-ctc-model
- Training configuration:
metadata/lmaana_clean_v5_long.yaml - Validation-based selection:
metadata/selection.json - Final metrics:
metadata/validation_summary.jsonandmetadata/test_summary.json - Integrity information:
manifest.json
Citation
If this checkpoint is useful in your work, cite the model release and the upstream OmniASR project:
@misc{joulali2026lmaana21,
author = {Badr Joulali},
title = {Lmaana 2.1: Clean Darija Adaptation with OmniASR CTC 1B},
year = {2026},
howpublished = {Hugging Face model release},
url = {https://huggingface.co/sailu4/lmaana-2.1}
}
Questions and reproducibility issues can be reported through the project issue tracker.











