Kev-0.8B for LiteRT

jaredpalmer/kev-0.8b is a decision model by Jared Palmer. It reads a text, called the state, and typed questions about it: yes/no (noul), multiple choice (choice) and rating (score). For each question it returns an answer with probabilities. It never generates text. A pointer head of two linear layers reads the hidden states of Qwen/Qwen3.5-0.8B-Base, adapted with a rank-16 LoRA, and turns them into those probabilities. Requests and responses follow the /v1/systemone shape of the author's server (github.com/jaredpalmer/kev).

This repository holds the backbone as eight LiteRT files, with the LoRA folded into the weights. Six are row graphs, for rows of up to 64, 128, 256, 512, 1,024 and 2,048 tokens. Two are shared-state pairs, which run the state once (up to 128 or 256 tokens) and then each question from it (up to 64 tokens each). The repository also holds the pointer head, the tokenizer files, a Python host, a Kotlin snippet, test fixtures and the conversion scripts. The files return hidden states. The host builds the input rows and applies the head.

The reference for every check is the author's fp32 PyTorch code. Every file ran on the Apple M4 Max CPU and Metal GPU (float32 precision) with ai-edge-litert 2.2.0, and on the Galaxy S26 GPU (FP16_WITH_FP32_ACCUM precision) with LiteRT 2.2.0. The 64-, 128- and 256-token files also ran on the Galaxy S26 NPU. A near tie is a question whose two most likely options in the reference are 0.02 or less apart. On every other question, every file gave the reference's most likely option in these runs. Their probabilities, near ties included, differed from the reference by at most 0.0104 on the desktop, 0.0110 on the S26 GPU and 0.0134 on the S26 NPU.

An invented support ticket with three questions, the state abbreviated:

{
  "state": "Ticket #48213, opened by Mara Quellen.\n\nHi, I ordered the Thistlebeam desk lamp (order TB-20931) on September 14 and was charged twice on my card … Could you refund the duplicate charge? I need it sorted before my card statement closes on Friday. …",
  "questions": {
    "team": {"type": "choice", "instructions": "Which team should handle this ticket?",
             "criteria": {"billing": "Charges, refunds and invoices", "shipping": "Deliveries, tracking and lost parcels",
                          "returns": "Exchanges and sending a product back", "technical": "Product faults and setup help"}},
    "deadline": {"type": "noul", "instructions": "Does the customer ask for action by a specific deadline?",
                 "criteria": {"true": "The ticket names a day or date by which something must happen",
                              "false": "No deadline is stated"}},
    "mood": {"type": "score", "instructions": "How upset is the customer?", "criteria": ["Calm", "Annoyed", "Angry"]}
  }
}

On desktop CPU, the host returned this response in each of its three modes (auto, row and pair). It is examples/run_example.expected.json, which leaves out latency_ms.

{
  "model": "kev-latest",
  "answers": {
    "team": {"type": "choice", "choice": "billing", "confidence": 0.7618,
             "probabilities": {"billing": 0.8213, "shipping": 0.014, "returns": 0.0883, "technical": 0.0764}},
    "deadline": {"type": "noul", "noul": 0.4461},
    "mood": {"type": "score", "score": 0.8542, "legend": {"0": "Calm", "1": "Annoyed", "2": "Angry"},
             "probabilities": {"0": 0.3672, "1": 0.4114, "2": 0.2214}, "confidence": 0.1172}
  },
  "usage": {"input_tokens": 216, "output_tokens": 189}
}

Other formats:

We have not checked their contents.

What changed on 2026-10-05

  • Files: six row graphs instead of three, with new graphs for rows of up to 64, 128 and 256 tokens, and two shared-state pairs. The L512, L1024 and L2048 files keep their names but have new contents and new SHA-256 values. The head weights (head/kev_0.8b_pointer_head.safetensors), the tokenizer files, fixtures/, LICENSE, NOTICE and the example's expected response are unchanged. The JSON beside the head weights has a new graph entry.
  • Kernel: the Gated DeltaNet chunk kernel reaches the same result by different steps. In PyTorch fp32 on all 402 test questions, the read-out hidden states moved by at most 2.3e-5 and the probabilities by at most 4.7e-6, and no most likely option changed. The 512-token fp32 graph went from 20,873 to 8,327 operators.
  • GPU precision: at the GPU's default precision (float16 activations), the 2026-10-04 files gave NaN on 18 of 393 test rows. The new files stay finite there, but their probabilities move outside the tolerance. Python's GpuOptions offers only float32 and that default, so the desktop GPU stays at float32.
  • Android GPU: with FP16_WITH_FP32_ACCUM (float16 storage, float32 accumulation) in the Kotlin and C APIs, every file stays within the tolerance on the Galaxy S26.
  • NPU: the L64, L128 and L256 files also run within the tolerance on the Galaxy S26's Qualcomm NPU (HTP). They carry 24 extra single-element SUM operators, which change no value.
  • Host: the 64-, 128- and 256-token graphs, the pair route, and the options mode (auto, row, pair), pair_ratio, handover and constant_tensor_sharing.

Files

File Bytes Role
kev-0.8b_rowprefill_L64_fp16fc_i8emb.tflite 1,258,031,552 Row graph for rows of up to 64 tokens. GPU, CPU and NPU.
kev-0.8b_rowprefill_L128_fp16fc_i8emb.tflite 1,258,444,912 Row graph for rows of up to 128 tokens. GPU, CPU and NPU.
kev-0.8b_rowprefill_L256_fp16fc_i8emb.tflite 1,259,246,704 Row graph for rows of up to 256 tokens. GPU, CPU and NPU.
kev-0.8b_rowprefill_L512_fp16fc_i8emb.tflite 1,261,233,328 Row graph for rows of up to 512 tokens. GPU and CPU.
kev-0.8b_rowprefill_L1024_fp16fc_i8emb.tflite 1,266,799,568 Row graph for rows of up to 1,024 tokens. GPU and CPU.
kev-0.8b_rowprefill_L2048_fp16fc_i8emb.tflite 1,284,223,520 Row graph for rows of up to 2,048 tokens. GPU and CPU.
kev-0.8b_sharedstate_Ls128_Lq64_fp16fc_i8emb.tflite 1,261,368,160 Shared-state pair: a state of up to 128 tokens, then questions of up to 64 tokens each.
kev-0.8b_sharedstate_Ls256_Lq64_fp16fc_i8emb.tflite 1,261,918,016 Shared-state pair: a state of up to 256 tokens, then questions of up to 64 tokens each.
head/kev_0.8b_pointer_head.safetensors 2,099,632 Pointer head: q.weight [256, 1024], q.bias, k.weight, k.bias, float32.
head/kev_0.8b_pointer_head.json 3,817 Temperature, delimiter and pad token ids, the head formula, and the contract of the row graphs, the pairs and the NPU files.
tokenizer/tokenizer.json 19,989,325 The source repository's tokenizer.json, unchanged.
tokenizer/tokenizer_config.json 1,128 The source repository's tokenizer_config.json, unchanged.
host/kev_litert.py 54,212 Python host: request to rows, row graph or pair, pointer head and response. It also runs from the command line. The same file serves Kev-4B.
host/requirements-host.txt 393 Pinned packages for the Python host.
examples/ run_example.py, the example above, and run_example.expected.json, its expected response.
android/CardSnippet.kt 7,041 The Kotlin block below: one class for a row graph, one for a pair.
android/measure/ Sources of the measurement activity and the NPU runner that produced the phone numbers below. They are not a sample app. A sample app in a separate repository (github.com/john-rocky/LiteRT-Models/tree/main/kev) runs the model files on the GPU and, for rows of up to 256 tokens, on the NPU.
fixtures/ Test requests (156 with their text, 221 by reference) and the script that rebuilds the full set; the reference's results for all 402 questions; 12 tokenizer probes; the SemIf license.
conversion/ Conversion, check and timing scripts, their environment files, and a README with the commands in order.
REPRODUCE.md 30,558 How to reproduce the files and the checks.
LICENSE 11,358 Apache License 2.0 text.
NOTICE 2,013 Attribution.
SHA256SUMS Lists every file with its SHA-256.

All eight files hold the same weights. A row is the state plus one question. A graph computes all of its positions whatever the row length, so use the smallest row graph that holds the row. Any one file works alone: the Python host loads the files that are present, picks a row graph or a pair for each question, and refuses a row that no file holds instead of cutting it.

Each file holds the token embedding as an int8 table, the 24 layers of the backbone and the final RMSNorm. It has no LM head and returns hidden states. The FULLY_CONNECTED operators (222 in the L128 file) read float16 weights through DEQUANTIZE operators (ai-edge-quantizer float casting). The embedding table, [248,320 × 1,024], is int8 with one scale per row. Activations stay in float32, and so does one FULLY_CONNECTED in each pair, the one that looks up the positions with a one-hot vector.

Minimal usage

Python: desktop CPU or GPU

The host needs tokenizers, numpy, safetensors and ai-edge-litert (2.2.0), pinned with their dependencies in host/requirements-host.txt. It does not need PyTorch or the kev package.

hf download litert-community/Kev-0.8B-LiteRT --local-dir Kev-0.8B-LiteRT
cd Kev-0.8B-LiteRT
pip install -r host/requirements-host.txt
python examples/run_example.py --check

examples/run_example.py sends the ticket request above on the CPU and prints the response. --check also compares it with examples/run_example.expected.json, and --mode row or --mode pair sends it through one route only. The example needs one model file, the Ls128 pair or the L256 graph: add --include "*_L256_*" "head/*" "tokenizer/*" "host/*" "examples/*" to the download to fetch the L256 graph and the files around it only.

In your own code:

import sys
sys.path.insert(0, "host")  # run from the repository root
from kev_litert import KevLiteRT

request = {"state": "Order #1182 arrived with a cracked screen.",
           "questions": {"refund": {"type": "noul", "instructions": "Should we offer a refund?"}}}

with KevLiteRT.from_dir(".") as kev:
    print(kev.decide(request))

from_dir finds the row graphs and pairs that are present, the head and the tokenizer in this repository's layout. It compiles a file only when a request needs it. decide() returns the response: model, answers, usage and latency_ms. The default is the CPU with 4 threads; threads= changes the count. accelerator="gpu" runs on the GPU with float32 precision (GpuOptions(enforce_f32=True)). A row that fits no file raises RowTooLong, and a non-finite hidden state raises NonFiniteOutput.

A request with several questions on one state can take a pair. In the default mode, auto, the ticket above goes through the Ls128 pair when that file is present: the state runs once, then one question step per question.

sys.path.insert(0, "examples")
from run_example import REQUEST as ticket  # the ticket above: one state, three questions

with KevLiteRT.from_dir(".", accelerator="gpu") as kev:  # mode="auto" is the default
    route = kev.route(kev.encode_rows(ticket))
    print(route["pair"], route["pair_questions"])  # (128, 64) 3
    print(kev.decide(ticket))

mode="row" or mode="pair" forces one route; the rule of auto is under Host contract. On the GPU, the host keeps one copy of a pair's weights for its two signatures (constant_tensor_sharing=True). With constant_tensor_sharing=False, a 5-question request on the Ls128 pair took 163.6 to 164.0 ms instead of 199.4 to 200.1 ms on Metal at float32, and the process footprint after one request was 6.3 GB instead of 3.0 GB.

From the command line:

python host/kev_litert.py --graph kev-0.8b_*.tflite \
    --head head/kev_0.8b_pointer_head.safetensors --tokenizer tokenizer/tokenizer.json --request request.json

--graph takes one or more row graphs and pairs; here the shell passes all eight. --accel gpu selects the GPU at float32 precision, and --request - reads the request from stdin. --mode, --pair-ratio, --handover and --no-constant-tensor-sharing set the route and the pair options.

Kotlin: Android GPU

// SPDX-License-Identifier: Apache-2.0
package com.kev.snippet

import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import com.google.ai.edge.litert.Environment
import com.google.ai.edge.litert.TensorBuffer
import java.io.File

private const val PAD = 248044 // <|endoftext|>, right padding
private const val STATE = 248060
private const val QUESTION = 248061
private const val OPTION_END = 248050
private const val DECIDE = 248062

private fun gpu(precision: CompiledModel.GpuOptions.Precision, constantTensorSharing: Boolean? = null) =
  CompiledModel.Options(Accelerator.GPU).apply {
    gpuOptions = CompiledModel.GpuOptions(constantTensorSharing = constantTensorSharing, precision = precision)
  }

/** The rows the pointer head reads: [decide (the last token), each option's 248050], d floats each. */
private fun readoutRows(hidden: FloatArray, length: Int, last: Int, optionEnds: IntArray): List<FloatArray> {
  val d = hidden.size / length
  return (listOf(last) + optionEnds.toList()).map { hidden.copyOfRange(it * d, (it + 1) * d) }
}

/**
 * One Kev row-prefill graph (`*_rowprefill_L{64,128,256,512,1024,2048}_fp16fc_i8emb.tflite`) on the GPU:
 * ids int32 [1, L] + valid float32 [1, L] -> hidden float32 [1, L, d] (d = 1024 for Kev-0.8B, 2560 for Kev-4B).
 * precision: FP16_WITH_FP32_ACCUM (float16 storage, float32 accumulation) or FP32. The default precision (plain float16
 * activations) moves the probabilities outside the parity tolerance. Kev-4B has not run on a phone: on a 12 GB Galaxy
 * S26 both tries, at FP32 and at FP16_WITH_FP32_ACCUM, used up the memory while the graph compiled.
 */
class KevRowGraph(
  file: File,
  private val length: Int,
  env: Environment,
  precision: CompiledModel.GpuOptions.Precision = CompiledModel.GpuOptions.Precision.FP16_WITH_FP32_ACCUM,
) : AutoCloseable {
  private val model = CompiledModel.create(file.absolutePath, gpu(precision), env)
  private val inputs = listOf("ids", "valid").associateWith { model.createInputBuffer(it, "serving_default") }
  private val outputs = mapOf("hidden" to model.createOutputBuffer("hidden", "serving_default"))

  /**
   * row = [248060] + state + [248061] + instructions + for each option ([248049] + option + [248050]) + [248062], with
   * the Kev repository's tokenizer.json and no special tokens added; optionEnds = the index of each option's 248050.
   */
  fun readout(row: IntArray, optionEnds: IntArray): List<FloatArray> {
    require(row.size <= length && row.last() == DECIDE) { "a row ends with 248062 and fits $length tokens" }
    require(optionEnds.all { row[it] == OPTION_END }) { "optionEnds must point at 248050 tokens" }
    inputs.getValue("ids").writeInt(IntArray(length) { if (it < row.size) row[it] else PAD })
    inputs.getValue("valid").writeFloat(FloatArray(length) { if (it < row.size) 1f else 0f })
    model.run(inputs, outputs, "serving_default")
    return readoutRows(outputs.getValue("hidden").readFloat(), length, row.size - 1, optionEnds)
  }

  override fun close() {
    (inputs.values + outputs.values).forEach { it.close() }
    model.close()
  }
}

/**
 * One Kev shared-state pair (`*_sharedstate_Ls{Ls}_Lq{Lq}_fp16fc_i8emb.tflite`) on the GPU: `state_prefill_<Ls>` runs
 * the request's state once, then `question_step_<Ls>_<Lq>` runs each question from it. state_prefill's output buffers go
 * to question_step as its inputs (no copy through the app). constantTensorSharing (default true) keeps one copy of the
 * weights on the GPU for the two signatures. Without it the pair is faster, but the GPU holds the weights twice: on a
 * Galaxy S26 (12 GB) at FP16_WITH_FP32_ACCUM, Kev-0.8B's Ls 128 pair answered a 5-question request in 430.9 ms without
 * sharing and 624.8 ms with it (median of the requests timed with the GPU clock not capped), and the phone's smallest
 * MemAvailable in the Ls 128 and Ls 256 runs (gate and timing) was 2.7 to 3.1 GB without sharing and 5.6 to 6.1 GB
 * with it.
 * layers = 24 for Kev-0.8B, 32 for Kev-4B: every fourth layer (3, 7, ...) is an attention layer (state k_<l>, v_<l>),
 * the others are Gated DeltaNet layers (gdn_state_<l>, conv_tail_<l>).
 */
class KevPairGraph(
  file: File,
  private val ls: Int,
  private val lq: Int,
  env: Environment,
  layers: Int = 24,
  precision: CompiledModel.GpuOptions.Precision = CompiledModel.GpuOptions.Precision.FP16_WITH_FP32_ACCUM,
  constantTensorSharing: Boolean = true,
) : AutoCloseable {
  private val model =
    CompiledModel.create(file.absolutePath, gpu(precision, if (constantTensorSharing) true else null), env)
  private val sigState = "state_prefill_$ls"
  private val sigQuestion = "question_step_${ls}_$lq"
  private val stateNames =
    (0 until layers).flatMap { if (it % 4 == 3) listOf("k_$it", "v_$it") else listOf("gdn_state_$it", "conv_tail_$it") }
  private val stateIn = listOf("ids", "valid").associateWith { model.createInputBuffer(it, sigState) }
  private val stateOut = stateNames.associateWith { model.createOutputBuffer(it, sigState) }
  private val questionIn = listOf("ids", "valid", "state_valid").associateWith { model.createInputBuffer(it, sigQuestion) }
  private val questionOut = mapOf("hidden" to model.createOutputBuffer("hidden", sigQuestion))
  private val questionInputs: Map<String, TensorBuffer> = questionIn + stateOut

  /** state = [248060] + the state's tokens, at most Ls. Run once per request, before its questions. */
  fun runState(state: IntArray) {
    require(state.size <= ls && state.first() == STATE) { "a state starts with 248060 and fits $ls tokens" }
    val valid = FloatArray(ls) { if (it < state.size) 1f else 0f }
    stateIn.getValue("ids").writeInt(IntArray(ls) { if (it < state.size) state[it] else PAD })
    stateIn.getValue("valid").writeFloat(valid)
    model.run(stateIn, stateOut, sigState)
    questionIn.getValue("state_valid").writeFloat(valid)
  }

  /**
   * branch = [248061] + instructions + for each option ([248049] + option + [248050]) + [248062], at most Lq tokens: the
   * row without the state. optionEnds index the branch. Returns the same rows as KevRowGraph.readout on the whole row,
   * up to float rounding.
   */
  fun readout(branch: IntArray, optionEnds: IntArray): List<FloatArray> {
    require(branch.size <= lq && branch.first() == QUESTION && branch.last() == DECIDE) {
      "a question starts with 248061, ends with 248062 and fits $lq tokens"
    }
    require(optionEnds.all { branch[it] == OPTION_END }) { "optionEnds must point at 248050 tokens" }
    questionIn.getValue("ids").writeInt(IntArray(lq) { if (it < branch.size) branch[it] else PAD })
    questionIn.getValue("valid").writeFloat(FloatArray(lq) { if (it < branch.size) 1f else 0f })
    model.run(questionInputs, questionOut, sigQuestion)
    return readoutRows(questionOut.getValue("hidden").readFloat(), lq, branch.size - 1, optionEnds)
  }

  override fun close() {
    (stateIn.values + stateOut.values + questionIn.values + questionOut.values).forEach { it.close() }
    model.close()
  }
}

Before calling readout, the app builds row and optionEnds: it tokenizes the state, the instructions and each option with this repository's tokenizer.json (no special tokens added, <|name|> rewritten to <¦name¦>) and joins them with the delimiter ids of the host contract below. With KevPairGraph, the app passes [248060] and the state's tokens to runState once per request, then each question's branch to readout, with optionEnds counted within the branch. Afterwards, the app applies the pointer head in float32 to the returned vectors, with the weights in head/kev_0.8b_pointer_head.safetensors and the temperature in the JSON beside it, and takes the softmax.

Both classes default to FP16_WITH_FP32_ACCUM (float16 storage, float32 accumulation), which keeps every file within the tolerance on the Galaxy S26; FP32 stays closer to the reference on rows longer than 1,024 tokens (see GPU precision). constantTensorSharing, on by default in KevPairGraph, keeps one copy of the weights for the pair's two signatures. Without it, a 5-question request on the Ls128 pair took 430.9 ms instead of 624.8 ms on the S26, but the pair runs peaked at 6.6 to 7.0 GB of process memory (VmHWM) instead of 3.1 to 3.2 GB.

Host contract

host/kev_litert.py implements this contract. It ports the request handling of the kev package at tag kev-1.0.

  1. Render the request. A JSON state becomes key: value lines. The options become text: no[: description] and yes[: description] for noul, name[: description] for choice, and the level text for score.
  2. Tokenize each text with tokenizer/tokenizer.json and add_special_tokens=False. Rewrite <|name|> in caller text to <¦name¦> before tokenizing, so that caller text cannot produce a delimiter.
  3. Build one row per question: [248060] + state + [248061] + instructions, then [248049] + option + [248050] for each option, then [248062]. The five delimiters are state (248060), question (248061), option start (248049), option end (248050) and decide (248062). The part of a row from 248061 on is the question's branch.
  4. Send each question to a row graph whose length L holds its row, or to a pair (see Routing). Pad ids on the right with 248044. valid is 1.0 on real tokens and 0.0 on padding.
  5. Run the row graph's serving_default signature. The inputs are ids int32 [1, L] and valid float32 [1, L]. The output is hidden float32 [1, L, 1024], the hidden states after the final RMSNorm at every position. Positions are the constants 0 to L−1. The graph has no input or output for model state, so each row starts from zero.
  6. Read out in float32. The decide vector is the hidden state of the row's last real token (248062). Option k's vector is the hidden state of its closing 248050. Then z_k = ((h_opt_k · Wkᵀ + bk) · (h_decide · Wqᵀ + bq)) / 16 / T, with T = 2.3510958125672174, and p = softmax(z).
  7. Answer. Noul returns noul = p(yes). Choice returns choice (the most likely option), probabilities and confidence = (p_max − 1/K) / (1 − 1/K) for K options. Score returns score = Σ i·p_i over the levels counted from 0, legend, probabilities and confidence = max(0, 1 − E|level − mode| / D), where D is the mean absolute deviation of a uniform distribution over the levels. Numbers are rounded to 4 decimals, as the author's round_prob does.

Shared-state pairs

A pair is one file with two signatures that share the weights:

  • state_prefill_<Ls> takes ids and valid [1, Ls], holding [248060] and the state's tokens. It returns 48 state tensors: gdn_state_<l> [1, 16, 128, 128] and conv_tail_<l> [1, 3, 6144] for the 18 Gated DeltaNet layers, and k_<l> and v_<l> [1, 2, Ls, 256] for the 6 attention layers (l = 3, 7, 11, 15, 19, 23). At Ls 128 they hold 23,347,200 bytes per request.
  • question_step_<Ls>_64 takes one question's branch as ids and valid [1, 64], the valid given to state_prefill as state_valid [1, Ls], and the 48 state tensors. It returns hidden float32 [1, 64, 1024]. The graph places the branch's positions after the state. Readout indices count within the branch: decide is its last real token, and each option is read at its 248050.

A state belongs to the file that made it: pass a file's state_prefill outputs unchanged to the same file's question_step, and never mix states between files. The host gives the output buffers to question_step as its inputs (handover="direct", the default) or reads them back and writes them (handover="host"). Both give the same bits on the Mac and the S26.

On the GPU, constant_tensor_sharing=True (the host's default; in Kotlin constantTensorSharing = true) keeps one copy of the weights for the two signatures. Without it, each signature holds its own copy, which is faster but uses more memory (see the speed tables). Both give the same probabilities bit for bit on the Mac and the S26.

Routing

In mode="auto", the default, the host takes the smallest pair whose Ls holds the state. For the n questions whose branches fit its Lq, it compares R, the sum of the lengths of the smallest row graphs that hold their rows, with P = Ls + n × Lq. Those questions take the pair only when R > 1.5 × P (pair_ratio); every other question takes its row graph. mode="row" uses only the row graphs, and mode="pair" sends every question through a pair or raises PairDoesNotFit. Of the 377 test requests, auto sends 9 through a pair (3 on Ls128, 6 on Ls256) and 368 through the row graphs.

Use this repository's tokenizer.json, which is the source repository's file, and read it with tokenizers.Tokenizer.from_file. It stores the tokenizer pipeline that transformers builds for the author's code (Qwen2Tokenizer). The base repository's own tokenizer.json, read the same way, gives different ids for text with combining marks, such as Devanagari, and for a few special-looking strings such as <think>. It differs on 4 of the 12 probes in fixtures/tokenizer_probes.json. The host refuses it.

Measured agreement and speed

Agreement with the author's fp32 code

The reference is the author's code at tag kev-1.0 on the CPU in float32, with the author's lock file (torch 2.8.0, transformers 5.17.0, peft 0.21.0): Checkpoint.load(cpu, dtype=float32), then forward in the row form. Probabilities are compared after the temperature.

The test set has 377 requests with 402 questions. The author's evals/v4/transfer-v4/development.jsonl gives 220 of them: records 1 to 60 (MMLU, 4 options), 20 from each of its 7 other sources and 20 score questions. SemIf authored144 gives 144 items, mapped to 3-option choice questions. The 12 invented records, written for this conversion, hold 37 questions of all three types. One more request is a control, described below. Without the control there are 401 questions: 280 choice, 93 noul and 28 score. The rows of 392 questions have at most 366 tokens. The other 9 questions sit on 3 long invented states, with rows of 1,369 to 1,805 tokens. Only the L2048 file holds them.

"Same most likely option" counts the questions that are not near ties, and "Near ties" the near-tie questions that kept the reference's most likely option. Max |Δp| is the largest absolute difference of any option's probability, and mean |Δp| is the mean over all options. The tolerance set before the runs was: the same most likely option on every question that is not a near tie, max |Δp| of at most 0.02 and mean |Δp| of at most 0.002.

"Control" is the max |Δp| of a control request against the reference of the question it copies, with one word of the instructions changed ("correctly" to "incorrectly"). It must exceed 0.02, which shows that the tolerance catches a one-word change. Its row has 94 tokens, so the L64 file does not hold it. Each file runs the questions it holds.

File Runtime Questions Same most likely option Near ties Max |Δp| Mean |Δp| Control
L64 Mac CPU (8 threads) and Metal (float32) 72 71/71 1/1 0.0078 8.8e-4 —
L128 Mac CPU (8 threads) and Metal (float32) 321 311/311 9/10 0.0104 9.7e-4 0.0477
L256 Mac CPU (8 threads) and Metal (float32) 385 371/371 12/14 0.0104 9.1e-4 0.0477
L512 Mac CPU (8 threads) and Metal (float32) 392 377/377 13/15 0.0104 9.1e-4 0.0477
L1024 Mac CPU (8 threads) and Metal (float32) 392 377/377 13/15 0.0104 9.1e-4 0.0477
L2048 Mac CPU (8 threads) and Metal (float32) 401 386/386 (long rows 9/9, max 0.0015) 13/15 0.0104 8.9e-4 0.0477
Pair Ls128 Mac CPU; Metal (float32) with and without sharing 313 306/306 6/7 0.0078 9.5e-4 0.0477
Pair Ls256 Mac CPU; Metal (float32) with and without sharing 340 329/329 9/11 0.0078 9.1e-4 0.0477
L128 Metal, default precision (float16 activations) 321 311/311 9/10 0.0332 3.6e-3 0.0469
L2048 Metal, default precision (float16 activations) 401 385/386 11/15 0.0352 4.1e-3 0.0539

The CPU and Metal give the same statistics on each line. On the same file they agree within 6.9e-6 (row graphs) and 6.4e-6 (pairs). A pair and the row route, with each question on its smallest row graph, differ by at most 3.7e-6 on the CPU and 4.4e-6 on Metal.

At float32, the near ties that changed are the same two questions on every file that holds their rows: tv4_023 (reference top-two gap 0.0020) and own_sensor_08 (8.1e-5). The 2026-10-04 files gave the same numbers, because the int8 embedding table sets the difference from the reference.

Run on the published files, the Python host reproduced these runs bit for bit: 452 of 452 question runs through the row graphs (CPU with 4 threads, and Metal at float32), 314 through the Ls128 pair and 341 through the Ls256 pair. In auto mode on the CPU with 4 threads, over all 377 requests, it gave the reference's most likely option on 386 of the 386 questions that are not near ties, with max |Δp| 0.0104 and mean |Δp| 8.9e-4. Rebuilt from the scripts in conversion/ alone, the L128 file and the Ls128 pair have the same SHA-256 as the published files.

Galaxy S26 agreement

These runs used one Galaxy S26 (SM-S942Q, Snapdragon SM8850, Android 16, 12 GB) with LiteRT 2.2.0. The phone saved the read-out hidden states, and the Mac scored them against the reference with the same tolerance. On the GPU (Kotlin CompiledModel API, OpenCL), every row graph ran whole on the delegate in one partition. Logcat shows Replacing 3912 out of 3912 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for L64, and 4,759, 6,019, 8,515, 13,555 and 23,635 of as many nodes for L128 to L2048.

File GPU precision Questions Same most likely option Near ties Max |Δp| Mean |Δp| Control
L64 FP16_WITH_FP32_ACCUM 72 71/71 1/1 0.0076 1.01e-3 —
L128 FP16_WITH_FP32_ACCUM 321 311/311 9/10 0.0077 1.08e-3 0.0462
L256 FP16_WITH_FP32_ACCUM 385 371/371 13/14 0.0092 1.04e-3 0.0472
L512 FP16_WITH_FP32_ACCUM 392 377/377 13/15 0.0091 1.06e-3 0.0473
L1024 FP16_WITH_FP32_ACCUM 392 377/377 13/15 0.0091 1.06e-3 0.0473
L2048 FP16_WITH_FP32_ACCUM 401 386/386 (long rows 9/9) 13/15 0.0093 1.08e-3 0.0473
Pair Ls128, with and without sharing (same bits) FP16_WITH_FP32_ACCUM 313 306/306 6/7 0.0101 1.10e-3 0.0467
Pair Ls256, with and without sharing (same bits) FP16_WITH_FP32_ACCUM 340 329/329 9/11 0.0110 1.08e-3 0.0467
L128 FP32 321 311/311 9/10 0.0104 9.7e-4 0.0477
L2048, the 9 long rows only FP32 9 9/9 — 0.0015 4.3e-4 0.0477 (the control row of the same run)
L2048, the 9 long rows only FP16_WITH_FP32_ACCUM 9 9/9 — 0.0093 1.84e-3 — (part of the L2048 line above)

On the NPU (the Qualcomm HTP), a small runner on the LiteRT 2.2.0 C API ran the files with the accelerators NPU and CPU. The int8 embedding lookup runs on the CPU and every other operator on the NPU: the dispatch delegate takes 2 of the 3 nodes. The phone compiles each graph when it loads it (JIT).

File Questions Same most likely option Near ties Max |Δp| Mean |Δp| Control
L64 72 71/71 1/1 0.0105 1.18e-3 —
L128 321 311/311 10/10 0.0134 1.29e-3 0.0476
L256 385 371/371 13/14 0.0134 1.24e-3 0.0476

With the NPU alone as the accelerator, without the CPU, LiteRT 2.2.0 fails to compile the graphs. The L512 and longer files and the pairs were not run on the NPU.

The L64, L128 and L256 files carry 24 extra single-element SUM operators. They sit between the FULLY_CONNECTED operators that form two kinds of gate, the gate of the Gated DeltaNet output norm and the attention output gate, and the sigmoids that read them. On the HTP, a sigmoid that reads a FULLY_CONNECTED output directly was off by about 2.3e-3 (absolute, root mean square), and graphs without the SUMs moved the probabilities by up to 0.085. The SUMs change no value: on the Mac CPU and GPU and on the S26 GPU, the files give the same bits as the graphs without them.

Desktop speed (Apple M4 Max)

These times come from an Apple M4 Max (macOS 27.0) with ai-edge-litert 2.2.0 through the Python CompiledModel API, measured while no other GPU job ran. A call runs one question's row. A time is the wall clock of writing the inputs, running and reading the output back: the median of 20 calls after 5 warm-up calls. A range spans two passes.

Graph (row tokens) GPU (Metal), float32 precision CPU, 8 threads
L64 (64) 24.8 to 24.9 ms 69.0 to 69.1 ms
L128 (80) 34.0 to 34.1 ms 109.4 to 109.5 ms
L256 (128 to 142; one call of a 5-question request) 54.4 ms 187.9 to 188.2 ms
L512 (300) 96.7 to 97.0 ms 358.8 to 388.3 ms
L1024 (1,000) 187.9 to 188.2 ms 789.9 to 798.5 ms
L2048 (1,805) 398.7 to 398.8 ms 1,634 to 1,660 ms

The next table times whole requests on Metal at float32. Row graphs: the sum of the calls, each row on the smallest graph that holds it. Pair: the state once, then one step per question, with the state handed over directly.

Request Pair Row graphs Pair, sharing Pair, no sharing
1 question, state 30 tokens (one row of 94 tokens: L128) Ls128 33.9 ms 75.3 to 75.7 ms 62.3 to 62.5 ms
2 questions, state 109 tokens (L256 × 2) Ls128 108.5 ms 106.3 to 106.8 ms 87.5 to 87.8 ms
3 questions, state 109 tokens (L256 × 3) Ls128 162.7 to 162.9 ms 137.2 to 137.9 ms 113.0 to 113.3 ms
5 questions, state 99 tokens (L256 × 3, L128 × 2) Ls128 230.0 to 231.8 ms 199.4 to 200.1 ms 163.6 to 164.0 ms
2 questions, state 150 tokens (L256 × 2) Ls256 108.0 to 108.8 ms 130.8 ms 110.7 to 110.9 ms
3 questions, state 167 tokens (L256 × 3) Ls256 162.7 to 162.8 ms 162.1 to 162.3 ms 136.3 to 136.4 ms

On the Ls128 pair without sharing, state_prefill took 36.8 to 37.0 ms and question_step 25.3 to 25.4 ms per question; with sharing, 44.2 to 44.5 ms and 30.9 to 31.1 ms. On the CPU with 8 threads, the Ls128 pair took 194.7 to 195.4, 340.5 to 341.8 and 484.4 to 487.4 ms for the requests of 1, 3 and 5 questions. On Metal, with a new process per run, the process footprint with the Ls128 pair after one request was 6.3 GB without sharing and 3.0 GB with it, and its peak during the compile 9.3 to 9.5 GB and 3.3 GB. With one row graph, again one new process per run, the footprint was 3.5 GB after compile, 3.7 GB after one call and 6.7 GB at its peak during the compile for L128, and 3.6, 3.8 and 6.6 GB for L512. Creating the CompiledModel on Metal took 4.2 to 4.4 s for L64, 5.0 to 6.5 s for L128, 5.4 to 5.6 s for L256, 5.4 to 7.0 s for L512, 6.8 to 8.2 s for L1024 and 10.9 to 12.2 s for L2048, and 2.7 to 4.1 s for a pair with sharing and 9.0 to 10.0 s without; on the CPU with 8 threads it took 0.2 to 2.5 s.

Galaxy S26 GPU speed

These times come from the same Galaxy S26 with LiteRT 2.2.0 through the Kotlin CompiledModel API on the GPU (OpenCL). A time is write + run + readFloat. On the GPU, run() returns at once and the work finishes when the output is read, so the times include the read.

The phone lowers its clock caps under load, and a compile alone can start it. In a run on a 128-token graph at FP32, the 7.6 s compile took the GPU from 41 to 60 °C. The calls began right away: the 17 made while the GPU cap was lowered took 143.4 to 161.5 ms, and the 8 made after it was lifted took 134.6 to 137.3 ms. So each timing run waits after the compile until the GPU is back to its temperature before the compile. Every call is logged with its time and matched with the phone's state, read every 2 seconds.

Cool is the median of the calls made while neither the GPU's clock cap (1,300 MHz) nor any CPU cap was lowered. Sustained is the median of the calls made after a cap was lowered. At FP16_WITH_FP32_ACCUM, the L64 and L128 runs (1.4 and 2.5 s of calls) ran with neither cap lowered; in the L256, L512, L1024 and L2048 runs a cap came 5.4 to 7.4 s into the back-to-back calls (warm-up included).

Graph (row tokens) FP16_WITH_FP32_ACCUM, cool (calls) FP16_WITH_FP32_ACCUM, sustained (calls) FP32, cool (calls)
L64 (64) 57.4 ms (20) — (cap not lowered) 85.3 ms (20)
L128 (80) 102.4 ms (20) — (cap not lowered) 137.4 ms (18)
L256 (128 to 142; one call of a 5-question request) 199.4 ms (32); the request: 992.9 ms (7 requests) 220.5 ms (68); the request: 1,088.3 ms (13 requests) 270.2 ms (30); the request: 1,355.6 ms (6 requests)
L512 (300) 387.0 ms (9) 393.4 ms (11) 565.1 ms (16)
L1024 (1,000) 819.8 ms (7) 847.7 ms (13) 1,219.0 ms (3)
L2048 (1,805) 1,808.2 ms (2) 2,147.6 ms (18) 3,119.6 ms (2)

Requests at FP16_WITH_FP32_ACCUM are timed as on the desktop, with 2 warm-up and 6 timed requests per line. A cell is the median of the requests timed while the GPU cap was not lowered, with their number in parentheses. The app loads one graph per run, so the 5-question row value is the sum of two runs: 205.6 ms for its 2 rows on L128 and 594.4 ms for its 3 rows on L256. The CPU clock was capped during the pair requests without sharing of 3 and 5 questions and during the row requests on L256: all of them for both 2-question requests and for the 3-question request with the 167-token state, 2 of 6 for the other 3-question request and 4 of 6 for the 5-question request's rows. The other cells ran with neither clock capped.

Request Pair Row graphs Pair, sharing Pair, no sharing
1 question, state 30 tokens (L128) Ls128 102.7 ms (6) 249.1 ms (6) 179.3 ms (6)
2 questions, state 109 tokens (L256 × 2) Ls128 394.7 ms (6) 341.8 ms (2) 242.1 ms (6)
3 questions, state 109 tokens (L256 × 3) Ls128 594.2 ms (6) 441.5 ms (5) 304.4 ms (4)
5 questions, state 99 tokens (L256 × 3, L128 × 2) Ls128 800.0 ms (6 + 6) 624.8 ms (2) 430.9 ms (6)
2 questions, state 150 tokens (L256 × 2) Ls256 399.0 ms (6) 445.2 ms (5) 345.3 ms (6)
3 questions, state 167 tokens (L256 × 3) Ls256 593.5 ms (5) 538.4 ms (2) 407.0 ms (6)

Memory during these runs, on the 12 GB phone (the largest or smallest value of a run; GB = the /proc kB × 1,024 / 10⁹):

Run Process VmHWM GPU (kgsl) Smallest MemAvailable
Pair with sharing (4 runs) 3.1 to 3.2 GB 1.7 to 1.8 GB 5.6 to 6.1 GB
Pair without sharing (4 runs, no low-memory kill) 6.6 to 7.0 GB 3.2 GB 2.7 to 3.1 GB
One row graph, FP16_WITH_FP32_ACCUM 4.6 to 5.7 GB (L64 to L1024); 5.4 to 6.1 GB (L2048) 1.6 to 1.8 GB (L64 to L1024); 2.1 to 2.2 GB (L2048) 3.7 to 4.3 GB (L64 to L1024); 2.4 to 2.9 GB (L2048)
One row graph, FP32 5.5 to 6.7 GB 2.6 to 2.9 GB (L64 to L1024); 3.4 GB (L2048) 2.6 to 3.4 GB

The row-graph runs did not use constant tensor sharing (the measurement app's default). Not measured: the memory and the times of a row graph with sharing.

Compiling on the GPU, once per process, took 5.8 to 11.0 s for L64 to L1024 and 12.2 and 14.3 s for L2048 at FP16_WITH_FP32_ACCUM; at FP32 it took 7.0 to 8.5 s for L64 to L256, 8.4 s for L512, 9.2 s for L1024 and 15.3 s for L2048 in the timing runs, and 22.0 s for L2048 in an 11-row run. A pair took 9.9 to 10.7 s with sharing and 11.9 to 16.3 s without.

Many requests in a row lower the GPU and CPU clocks. In a run where the GPU cap fell to 578 MHz, the 5-question request on the pair with sharing took 1,081.0 ms instead of 624.8 ms. If the screen locks, the app moves to the background CPU set and the computation stops; keep the app in the foreground with the screen on.

Galaxy S26 NPU speed

These times come from the runner of the NPU agreement runs, with each graph loaded from its cache. A time is write + run + read.

Graph Same row, 20 calls (median) 20 different rows, once each (median) GPU at FP16_WITH_FP32_ACCUM, same runner and row, 20 calls (median)
L64 44.0 ms 44.1 ms 56.9 ms
L128 67.8 ms; 71.1 ms (40 calls, another process) 69.9 ms; 72.4 ms (the other process) 101.9 ms; 102.3 ms (40 calls)
L256 126.7 ms 126.4 ms 194.9 ms

NPU times move by about 5% from process to process (L128: 67.8 and 71.1 ms). The GPU times for L64 and L256 come from graphs with the same weights but without the 24 SUMs; on L128 the published file took 101.9 ms and the graph without the SUMs 101.1 ms. In the phone's state read during these runs, the GPU clock cap was never lowered; the CPU cores ran at their maximum in the L64 run, while in the L128 and L256 runs the cores of cpufreq/policy6 were capped at 4.26 to 4.65 GHz of 4.74 GHz (and in the L256 run those of policy0 at 3.51 of 3.63 GHz). The NPU's own clock is not among the values read.

The initial load of a file compiles it on the phone: 55.7 s for L64, 191.2 s for L128 and 263.3 s for L256, with a process VmHWM of 5.5 to 6.2 GB and a smallest MemAvailable of 3.3 to 4.2 GB meanwhile. Each file leaves a cache of 1.27 to 1.29 GB. Later loads read it in 1.5 to 1.7 s, with a VmHWM of 2.3 to 2.4 GB.

The NPU needs LiteRT 2.2.0's NPU runtime libraries (the Qualcomm compiler plugin and dispatch library) and the HTP libraries of QAIRT 2.47; this repository does not include them.

A sample app (see Files) also ran the published L64, L128 and L256 files through the Kotlin CompiledModel API, with the accelerators NPU and CPU. It used the 181 questions whose text the repository carries (the 144 SemIf questions and the 37 invented ones). Each file stayed within the tolerance on the rows it holds: max |Δp| 0.0105 on L64 (34 questions), 0.0134 on L128 (147) and 0.0134 on L256 (172).

In the app, one call took 45.1 ms on L64, 65.8 ms on L128 (68.5 ms in another process) and 121.9 ms on L256. These times are medians of 60 calls after 5 warm-up calls, with each graph loaded from its cache. A time is write + run + read.

Inside the app, the initial compile took 81.6 s for L64, 179.0 s for L128 and 298.2 s for L256, one file per process. Later loads read the cache in 0.9 to 1.5 s. Not measured: other phones or SoCs, and files compiled ahead of time (on small graphs, AOT gave the same bits as the JIT).

In two debug launches, the app also tried the Ls128 pair, with the accelerators NPU and CPU. Both times, Android's low-memory killer (lmkd) stopped the app about 5 minutes into compiling the pair's two signatures on the 12 GB phone, before any cache file was written. So the app compiles the pairs for the GPU only.

This release and the 2026-10-04 release

The 2026-10-04 values are that card's, on the same devices and rows: the median of 20 calls after 5 warm-up calls, with sustained load included on the S26. The NPU time is the L256 graph's, which computes all 256 positions at any row length.

Row 2026-10-04 This release
S26, about 130 tokens (one call of the 5-question request) L512 at FP32: 630 ms L256 at FP16_WITH_FP32_ACCUM: 199.4 ms cool, 220.5 ms sustained; at FP32: 270.2 ms cool; on the NPU: 126.7 ms
S26, the 5-question request L512 × 5 at FP32: 3,150 ms Pair at FP16_WITH_FP32_ACCUM: 624.8 ms with sharing, 430.9 ms without; row graphs: 800.0 ms
S26, 300 tokens L512 at FP32: 663 ms L512 at FP16_WITH_FP32_ACCUM: 387.0 ms cool, 393.4 ms sustained; at FP32: 565.1 ms cool, 563.1 ms sustained
S26, 1,000 tokens L1024 at FP32: 1,345 ms L1024 at FP16_WITH_FP32_ACCUM: 819.8 ms cool, 847.7 ms sustained; at FP32: 1,219.0 ms cool, 1,316.8 ms sustained
S26, 1,805 tokens L2048 at FP32: 3,331 ms (one call at any row length, 25 calls) L2048 at FP16_WITH_FP32_ACCUM: 1,808.2 ms cool, 2,147.6 ms sustained; at FP32: 3,119.6 ms cool, 3,121.2 ms sustained
Mac (Metal, float32), about 130 tokens L512: 139.9 ms L256: 54.4 ms
Mac, the 5-question request L512 × 5: 699.9 ms Pair: 163.6 to 164.0 ms without sharing, 199.4 to 200.1 ms with; row graphs: 230.0 to 231.8 ms
Mac, 1,805 tokens L2048: 474.2 ms L2048: 398.7 to 398.8 ms

GPU precision

Run the desktop GPU at float32 precision. On Android, use FP16_WITH_FP32_ACCUM or FP32.

At the GPU's default precision (float16 activations), the 2026-10-04 files gave non-finite (NaN) hidden states on 18 of 393 test rows. On Metal, the L128 and L2048 files of this release stay finite at that precision, but their probabilities move outside the tolerance: max |Δp| 0.0332 and mean 3.6e-3 on L128, 0.0352 and 4.1e-3 on L2048. Python's GpuOptions offers only float32 (enforce_f32=True) and the default, so the Python host runs the GPU at float32 by default. It raises NonFiniteOutput instead of answering from a non-finite row.

The Kotlin and C APIs add FP16_WITH_FP32_ACCUM: float16 storage with float32 accumulation. With it, every file stays within the tolerance on the Galaxy S26 GPU. One call took 57.4 ms against 85.3 ms at FP32 on L64, 102.4 against 137.4 ms on L128 and 199.4 against 270.2 ms on L256 (cool).

On long rows, FP32 stays closer to the reference. On the 9 rows longer than 1,024 tokens, FP16_WITH_FP32_ACCUM gives max |Δp| 0.0093 and mean 1.84e-3: inside the tolerance of 0.02 and 0.002, with little room on the mean. FP32 gives 0.0015 and 4.3e-4 on the same rows. On the other 392 questions, FP16_WITH_FP32_ACCUM gives a mean of 1.06e-3. FP32 costs time on the long graphs: one call took 387.0 ms against 565.1 ms at FP32 on L512, 819.8 against 1,219.0 ms on L1024 and 1,808.2 against 3,119.6 ms on L2048 (cool; both L2048 values are medians of 2 calls).

Limits

  • The model never generates text. A row holds at most 2,048 tokens: the state plus one question. A pair holds a state of up to 256 tokens and question branches of up to 64 tokens; a question that does not fit takes a row graph. Longer states are not handled. The author's server accepts states of up to 65,536 tokens.
  • A choice question takes at most 255 options.
  • The GPU's default precision (float16 activations) is outside the tolerance. Use float32 on the desktop and FP16_WITH_FP32_ACCUM or FP32 on Android. Rows longer than 1,024 tokens stay closer to the reference at FP32: 0.0015 and 4.3e-4 on the 9 long rows, against 0.0093 and 1.84e-3 at FP16_WITH_FP32_ACCUM.
  • Probabilities differ from the reference by up to 0.0104 on the desktop at float32, 0.0110 on the S26 GPU and 0.0134 on the S26 NPU. On the desktop, the int8 embedding table sets this difference. A question whose two most likely options are 0.002 or less apart can change its answer: on the desktop, 2 of 401 did.
  • The NPU runs only the L64, L128 and L256 files. Its initial compile on the phone takes 55.7 to 263.3 s, and its cache takes 1.27 to 1.29 GB per file. It was measured on one Galaxy S26 only; inside a sample app, the initial compile took 81.6 to 298.2 s.
  • A pair without constant tensor sharing uses about twice the memory: a process footprint of 6.3 GB against 3.0 GB on the Mac after one request, and a VmHWM of 6.6 to 7.0 GB against 3.1 to 3.2 GB on the S26.
  • Continuous load slows the phone; compare the cool and sustained columns above.
  • Kev-4B is in a separate repository, litert-community/Kev-4B-LiteRT. It does not fit the 12 GB Galaxy S26: in both tries, memory ran out during compile.
  • The agreement numbers measure how closely the conversion follows the author's fp32 code, not task accuracy. This conversion did not measure accuracy or calibration again; see the source model card. The temperature is the author's fitted value, unchanged.
  • Input languages and intended uses follow the source model card.
  • Measured on one Mac (Apple M4 Max) and one Galaxy S26.

Provenance, conversion and license

  • Source: jaredpalmer/kev-0.8b at tag v1.0 (commit bf75a6a8848ea6960ff2ed108d9ed44c2941174f) by Jared Palmer, Apache-2.0. Base: Qwen/Qwen3.5-0.8B-Base at revision dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68, Apache-2.0.
  • Model: the base with a rank-16 LoRA on 12 projection types and a pointer head of two linear layers with 256 dimensions. The base has 24 layers (18 Gated DeltaNet layers, a form of linear attention, and 6 full-attention layers), hidden size 1024 and a vocabulary of 248,320.
  • Weights: the author's scripts/merge_lora_checkpoint.py at tag kev-1.0 folded the LoRA into 186 weights in float32. Read with the author's loader, the folded checkpoint gives bit-identical probabilities to the adapter checkpoint on all 402 questions.
  • Graph: the Qwen3.5 text model of transformers 5.14.1, re-authored for export. The Gated DeltaNet chunk kernel is rewritten at rank 4 or less: the tail padding becomes a concat, and the diagonal and triangular masks become constants. Attention copies the GQA keys and values by concat. The hidden states before and after this rewrite differ by 3.2e-5. On top of it, this release's final kernel changes how the chunk kernel computes, not what it computes; conversion/README.md describes each change. Exported with litert-torch 0.9.4 and quantized with ai-edge-quantizer 0.9.0.
  • Operators: 3,912 in L64, 4,759 in L128, 6,019 in L256, 8,515 in L512, 13,555 in L1024 and 23,635 in L2048. Each pair has 4,975 (Ls128) or 6,235 (Ls256) in state_prefill and 3,965 in question_step. These are also the node counts the S26 GPU delegate took. None is CUSTOM, no tensor is int64 and no tensor has a rank above 4.
  • Scripts: conversion/ holds every conversion, check and timing script, with its environment files and a README of the commands, the final kernel's changes and the float16 headroom of the norms. android/measure/ holds the sources behind the phone numbers. REPRODUCE.md lists the sources, the environments and the steps. Kev-4B-LiteRT ships the same conversion/ folder.
  • Test data: fixtures/ holds the SemIf records (MIT, from github.com/TheoLeeCJ/SemIf at commit ca3ba65f) and the 12 invented records with their text. The 221 transfer-v4 records, the control included, are listed by reference only: file, tag, line, line SHA-256 and _meta.id. Their source datasets carry different licenses: on the Hub cards, tweet_eval is unknown and SciQ is CC BY-NC 3.0. fixtures/rebuild_requests.py restores them from the author's GitHub repository at tag kev-1.0.
  • Training data and evaluation: see the source model card.

License: Kev-0.8B and Qwen3.5-0.8B-Base are licensed under Apache 2.0. This repository is released under the same license (LICENSE), except the SemIf fixture records, which are MIT (fixtures/LICENSE-SemIf-MIT.txt). Attribution, including the code adapted from the kev package and from litert-torch, is in NOTICE.

Downloads last month
80
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for litert-community/Kev-0.8B-LiteRT

Finetuned
(2)
this model