Transformers
English
Sanskrit
tokenizer
sentencepiece
unigram
english
math
custom_code

Mume tokenizer 32k

The 32,000-piece SentencePiece unigram tokenizer shared by the Muse Mesh English and Math models (mume-english-125m, mume-math-125m). One vocabulary for English, mathematical text with LaTeX, and Sanskrit in SLP1 transliteration.

From Muse Mesh (Hugging Face). Part of the Muse Mesh English and Math models (collections). Every training run, including the ones not released, is logged at mume.ai/sansar/runs (moving to mume.ai/lab).

Quick start

from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("MuseMesh/mume-tokenizer-32k", revision="v0.1.0", trust_remote_code=True)
ids = tok("Let $f(x) = x^2 + 1$. Then f(3) = 10.")["input_ids"]
print(tok.convert_ids_to_tokens(ids))
print(tok.decode(ids))      # the same string

or with SentencePiece directly: sentencepiece.SentencePieceProcessor(model_file="tokenizer.model"). trust_remote_code=True loads tokenization_mume.py, a thin wrapper around SentencePiece, so the ids are exactly those the models were trained on. Without it, AutoTokenizer loads tokenizer.json (the tokenizers fast version of the same vocabulary): identical on short texts, but SentencePiece scores a whole document as one lattice in float32 and so settles a few near-ties differently on long documents (24 of 14,983 FineWeb and 34 of 2,704 OpenWebMath validation documents differ, by a token or two). Use the exact path with the models.

What it is

Model SentencePiece unigram, 32,000 pieces
Normalisation none (identity): text in is text out; a ▁ marks each space and one is prepended
Digits ASCII digits always split one per token (2026 -> ▁ 2 0 2 6)
Unknown characters byte fallback (256 <0x..> pieces): nothing maps to <unk>
Special tokens <unk> 0, <s> 1, </s> 2 (end of document); no BOS/EOS added on encode; a literal </s> in text is read as characters
User symbols । and ॥ (Sanskrit daṇḍas)
Character coverage 0.9999

Training sample

1.2 GB, 400 MB from each of three sources, sampled by hash of the record id (2195 s to train):

source bytes lines what
English 400,000,339 2,048,890 FineWeb sample-10BT files 001-005 and 007 (25% of records)
Math 400,000,494 2,517,260 OpenWebMath shards 0-3 (25% of records)
Sanskrit 400,016,178 2,121,518 the Sansar corpus training slice without OCR (15% of records), transliterated to SLP1

The Sanskrit part is the reason for the licence (see Licence). The English and Math models never saw Sanskrit in training; the Sanskrit pieces never occur in their training data.

Measured

Tokens per whitespace word on the frozen evaluation sets (MuseMesh/mume-eval-suites), against GPT-2's 50,257-piece tokenizer, and the release check (ids equal to SentencePiece's, per record):

set group tokens / word GPT-2 tokens / word bytes / token exact class = SentencePiece fast tokenizer.json = SentencePiece
fineweb_val english 1.523 1.347 3.931 14,983 / 14,983 14,959 / 14,983
fineweb_val_clean english 1.501 1.337 3.963 11,048 / 11,048 11,041 / 11,048
wikitext103_test english 1.509 1.174 3.538 62 / 62 62 / 62
enwik8_test english 2.650 2.283 2.910 1 / 1 0 / 1
text8_test english 1.339 1.174 4.365 1 / 1 0 / 1
owm_val math 1.940 1.755 3.266 2,704 / 2,704 2,670 / 2,704
owm_val_clean math 1.869 1.701 3.328 1,126 / 1,126 1,119 / 1,126
gsm8k_test math 1.715 1.332 2.911 1,319 / 1,319 1,319 / 1,319
math_test math 2.652 2.524 2.329 5,000 / 5,000 4,997 / 5,000
math_test_clean math 2.578 2.433 2.383 3,754 / 3,754 3,751 / 3,754

With 32k pieces shared three ways, FineWeb text takes 13% more tokens than with GPT-2's 50k vocabulary, WikiText 28% more.

Sanskrit (the Sansar E0 held-out sets): tokens per word in SLP1, the form the tokenizer was trained on, and on raw Devanagari, which it was not (Devanagari letters fall back to bytes):

set SLP1 Devanagari
dcs_gold 2.359 9.918
prose 2.152 8.399
ood 2.619 9.698
vedic 3.269 10.175
gita 2.362 9.248

All 58,338 texts of the check (the ten English and Math sets, the five Sanskrit sets in both scripts) give SentencePiece's ids through the exact class (58338/58338) and decode back to the input wherever SentencePiece's own decode does (1 text with a literal ▁ excepted); the fast tokenizer matches on 58257/58338. Edge cases (empty strings, runs of spaces, tabs and newlines, digits, LaTeX, emoji, CJK, literal special-token text): 35/35 exact, 35/35 fast. The one string that does not round-trip is a literal ▁ (U+2581), which SentencePiece itself reads as a space.

Files

file what
tokenizer.model, tokenizer.vocab the SentencePiece model and its piece list with scores
tokenizer.json the same vocabulary for the tokenizers library
tokenization_mume.py, tokenizer_config.json, special_tokens_map.json transformers glue
training_config.json trainer settings, the sample's sources and sizes
eval/verification.json the measurements and checks above
LICENSE, LICENSE-CODE, CHANGELOG.md licences and version history

Attribution

data Hugging Face reference licence
FineWeb HuggingFaceFW/fineweb Penedo et al. 2024, The FineWeb Datasets, arXiv:2406.17557 ODC-By 1.0; use is also subject to the Common Crawl Terms of Use
OpenWebMath open-web-math/open-web-math Paster et al. 2023, OpenWebMath, arXiv:2310.06786 ODC-By 1.0; use is also subject to the Common Crawl Terms of Use

Sanskrit sample: the Sansar corpus training slice (sources and licences listed in MuseMesh/sansar-sanskrit-corpus; this sample also includes sources that the open corpus leaves out).

Licence

This release is for research and non-commercial use.

  • tokenizer.model, tokenizer.vocab, tokenizer.json: CC BY-NC 4.0 (LICENSE), attribution to "Muse Mesh Private Limited".
  • Code (tokenization_mume.py): Apache-2.0 (LICENSE-CODE).
  • Commercial use: contact kushal@muse-mesh.com.

Why non-commercial: the training sample's Sanskrit includes text licensed for non-commercial use only (GRETIL, CC BY-NC-SA 4.0; Muktabodha, CC BY-NC 4.0) and text without a licence statement. An Apache-2.0 tokenizer trained without those sources is planned as a separate repository, with models retrained on it.

Citation

@misc{mume_tokenizer_32k_2026,
  title  = {Mume tokenizer 32k},
  author = {Muse Mesh},
  year   = {2026},
  note   = {v0.1.0},
  url    = {https://huggingface.co/MuseMesh/mume-tokenizer-32k}
}

Contact: kushal@muse-mesh.com

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train MuseMesh/mume-tokenizer-32k

Collection including MuseMesh/mume-tokenizer-32k

Papers for MuseMesh/mume-tokenizer-32k