Instructions to use MuseMesh/mume-tokenizer-32k with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MuseMesh/mume-tokenizer-32k with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("MuseMesh/mume-tokenizer-32k", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Mume tokenizer 32k
The 32,000-piece SentencePiece unigram tokenizer shared by the Muse Mesh English and Math models (mume-english-125m, mume-math-125m). One vocabulary for English, mathematical text with LaTeX, and Sanskrit in SLP1 transliteration.
From Muse Mesh (Hugging Face). Part of the Muse Mesh English and Math models (collections). Every training run, including the ones not released, is logged at mume.ai/sansar/runs (moving to mume.ai/lab).
Quick start
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("MuseMesh/mume-tokenizer-32k", revision="v0.1.0", trust_remote_code=True)
ids = tok("Let $f(x) = x^2 + 1$. Then f(3) = 10.")["input_ids"]
print(tok.convert_ids_to_tokens(ids))
print(tok.decode(ids)) # the same string
or with SentencePiece directly: sentencepiece.SentencePieceProcessor(model_file="tokenizer.model"). trust_remote_code=True loads tokenization_mume.py, a thin wrapper around SentencePiece, so the ids are exactly those the models were trained on. Without it, AutoTokenizer loads tokenizer.json (the tokenizers fast version of the same vocabulary): identical on short texts, but SentencePiece scores a whole document as one lattice in float32 and so settles a few near-ties differently on long documents (24 of 14,983 FineWeb and 34 of 2,704 OpenWebMath validation documents differ, by a token or two). Use the exact path with the models.
What it is
| Model | SentencePiece unigram, 32,000 pieces |
| Normalisation | none (identity): text in is text out; a ▁ marks each space and one is prepended |
| Digits | ASCII digits always split one per token (2026 -> ▁ 2 0 2 6) |
| Unknown characters | byte fallback (256 <0x..> pieces): nothing maps to <unk> |
| Special tokens | <unk> 0, <s> 1, </s> 2 (end of document); no BOS/EOS added on encode; a literal </s> in text is read as characters |
| User symbols | । and ॥ (Sanskrit daṇḍas) |
| Character coverage | 0.9999 |
Training sample
1.2 GB, 400 MB from each of three sources, sampled by hash of the record id (2195 s to train):
| source | bytes | lines | what |
|---|---|---|---|
| English | 400,000,339 | 2,048,890 | FineWeb sample-10BT files 001-005 and 007 (25% of records) |
| Math | 400,000,494 | 2,517,260 | OpenWebMath shards 0-3 (25% of records) |
| Sanskrit | 400,016,178 | 2,121,518 | the Sansar corpus training slice without OCR (15% of records), transliterated to SLP1 |
The Sanskrit part is the reason for the licence (see Licence). The English and Math models never saw Sanskrit in training; the Sanskrit pieces never occur in their training data.
Measured
Tokens per whitespace word on the frozen evaluation sets (MuseMesh/mume-eval-suites), against GPT-2's 50,257-piece tokenizer, and the release check (ids equal to SentencePiece's, per record):
| set | group | tokens / word | GPT-2 tokens / word | bytes / token | exact class = SentencePiece | fast tokenizer.json = SentencePiece |
|---|---|---|---|---|---|---|
fineweb_val |
english | 1.523 | 1.347 | 3.931 | 14,983 / 14,983 | 14,959 / 14,983 |
fineweb_val_clean |
english | 1.501 | 1.337 | 3.963 | 11,048 / 11,048 | 11,041 / 11,048 |
wikitext103_test |
english | 1.509 | 1.174 | 3.538 | 62 / 62 | 62 / 62 |
enwik8_test |
english | 2.650 | 2.283 | 2.910 | 1 / 1 | 0 / 1 |
text8_test |
english | 1.339 | 1.174 | 4.365 | 1 / 1 | 0 / 1 |
owm_val |
math | 1.940 | 1.755 | 3.266 | 2,704 / 2,704 | 2,670 / 2,704 |
owm_val_clean |
math | 1.869 | 1.701 | 3.328 | 1,126 / 1,126 | 1,119 / 1,126 |
gsm8k_test |
math | 1.715 | 1.332 | 2.911 | 1,319 / 1,319 | 1,319 / 1,319 |
math_test |
math | 2.652 | 2.524 | 2.329 | 5,000 / 5,000 | 4,997 / 5,000 |
math_test_clean |
math | 2.578 | 2.433 | 2.383 | 3,754 / 3,754 | 3,751 / 3,754 |
With 32k pieces shared three ways, FineWeb text takes 13% more tokens than with GPT-2's 50k vocabulary, WikiText 28% more.
Sanskrit (the Sansar E0 held-out sets): tokens per word in SLP1, the form the tokenizer was trained on, and on raw Devanagari, which it was not (Devanagari letters fall back to bytes):
| set | SLP1 | Devanagari |
|---|---|---|
dcs_gold |
2.359 | 9.918 |
prose |
2.152 | 8.399 |
ood |
2.619 | 9.698 |
vedic |
3.269 | 10.175 |
gita |
2.362 | 9.248 |
All 58,338 texts of the check (the ten English and Math sets, the five Sanskrit sets in both scripts) give SentencePiece's ids through the exact class (58338/58338) and decode back to the input wherever SentencePiece's own decode does (1 text with a literal ▁ excepted); the fast tokenizer matches on 58257/58338. Edge cases (empty strings, runs of spaces, tabs and newlines, digits, LaTeX, emoji, CJK, literal special-token text): 35/35 exact, 35/35 fast. The one string that does not round-trip is a literal ▁ (U+2581), which SentencePiece itself reads as a space.
Files
| file | what |
|---|---|
tokenizer.model, tokenizer.vocab |
the SentencePiece model and its piece list with scores |
tokenizer.json |
the same vocabulary for the tokenizers library |
tokenization_mume.py, tokenizer_config.json, special_tokens_map.json |
transformers glue |
training_config.json |
trainer settings, the sample's sources and sizes |
eval/verification.json |
the measurements and checks above |
LICENSE, LICENSE-CODE, CHANGELOG.md |
licences and version history |
Attribution
| data | Hugging Face | reference | licence |
|---|---|---|---|
| FineWeb | HuggingFaceFW/fineweb | Penedo et al. 2024, The FineWeb Datasets, arXiv:2406.17557 | ODC-By 1.0; use is also subject to the Common Crawl Terms of Use |
| OpenWebMath | open-web-math/open-web-math | Paster et al. 2023, OpenWebMath, arXiv:2310.06786 | ODC-By 1.0; use is also subject to the Common Crawl Terms of Use |
Sanskrit sample: the Sansar corpus training slice (sources and licences listed in MuseMesh/sansar-sanskrit-corpus; this sample also includes sources that the open corpus leaves out).
Licence
This release is for research and non-commercial use.
tokenizer.model,tokenizer.vocab,tokenizer.json: CC BY-NC 4.0 (LICENSE), attribution to "Muse Mesh Private Limited".- Code (
tokenization_mume.py): Apache-2.0 (LICENSE-CODE). - Commercial use: contact kushal@muse-mesh.com.
Why non-commercial: the training sample's Sanskrit includes text licensed for non-commercial use only (GRETIL, CC BY-NC-SA 4.0; Muktabodha, CC BY-NC 4.0) and text without a licence statement. An Apache-2.0 tokenizer trained without those sources is planned as a separate repository, with models retrained on it.
Citation
@misc{mume_tokenizer_32k_2026,
title = {Mume tokenizer 32k},
author = {Muse Mesh},
year = {2026},
note = {v0.1.0},
url = {https://huggingface.co/MuseMesh/mume-tokenizer-32k}
}
Contact: kushal@muse-mesh.com