OLMo-2-1B-Instruct β€” LiteRT-LM (blockwise int4)

allenai/OLMo-2-0425-1B-Instruct converted to the LiteRT-LM (.litertlm) format for on-device inference with Google's LiteRT-LM runtime (the engine behind the official litert-community/* models).

OLMo-2 is AllenAI's fully-open model family (Apache-2.0; open weights, data, and training code). This 1B variant is small enough to run on a phone β€” verified on iPhone 17 Pro. Converted with the official upstream litert-torch β€” no fork.

File model.litertlm (~0.93 GB)
Quantization int4 weights β€” blockwise (block 32) + OCTAV optimal-clipping, symmetric; embedding INT8
Compute integer
Context (KV cache) 4096
Base model allenai/OLMo-2-0425-1B-Instruct
Decode speed ~24 tok/s (iPhone 17 Pro; loads 5.2 s, ~1.2 GB footprint) Β· ~138 tok/s (Mac M-series, Metal GPU)

Usage

Run with the LiteRT-LM runtime:

litert_lm_main \
  --model_path model.litertlm \
  --backend gpu \
  --input_prompt "Explain on-device AI in one sentence."

The .litertlm bundle carries the tokenizer and the prompt template (OLMo-2's native TΓΌlu format β€” <|user|> / <|assistant|>, stop token <|endoftext|>), so no separate tokenizer files are needed.

Run on Android

Update (July 2026): Google AI Edge Gallery v1.0.16+ can import litert-lm models directly from Hugging Face inside the app (tap +) β€” no computer or adb needed. The manual steps below are only required on older builds or for sideloading a local file.

The easiest way to try this model on a phone is the official Google AI Edge Gallery app:

  1. Install a recent Gallery (package com.google.ai.edge.gallery, APK from the repo's releases β€” 1.0.15+ supports .litertlm).
  2. Download model.litertlm and push it to the device:
    adb push model.litertlm /sdcard/Download/
    
  3. In the app, tap + (bottom-right), pick the file, and choose CPU or GPU. At ~0.93 GB this 1B fits comfortably on an 8 GB phone.
  4. Chat β€” the bundle already carries the tokenizer and OLMo-2 prompt template.

See the Gallery Importing Local Models guide for details. To embed it in your own Android app, use the LiteRT-LM Kotlin API (com.google.ai.edge.litertlm:litertlm-android).

Run on desktop (LiteRT-LM CLI)

The same .litertlm bundle runs on macOS / Linux / Windows with the official LiteRT-LM CLI β€” including as a local OpenAI-compatible API server:

pip install litert-lm
litert-lm import --from-huggingface-repo mlboydaisuke/OLMo-2-1B-Instruct-LiteRT model.litertlm olmo-2-1b-instruct-litert
litert-lm run olmo-2-1b-instruct-litert     # interactive chat in the terminal
litert-lm serve           # local OpenAI-compatible API server

Quality β€” GSM8K

Measured on GSM8K (n=100, greedy, 0-shot chain-of-thought, identical prompt and answer-extraction for every row).

Configuration GSM8K
bf16 (reference) 72.0%
This model β€” LiteRT int4 (BOCTAV4) 63.0%

63 % is a strong, coherent, non-degenerate score for a 1B (the \boxed{}-style answers terminate cleanly at <|endoftext|>). At 1B, 4-bit quantization costs ~9 pt vs bf16 β€” a small model has less redundancy to absorb int4 rounding than a 3B+ (where the same recipe is at parity). An int8 build recovers only ~2 pt (65 %) for +60 % size, so int4 is shipped as the best size/quality trade-off for on-device.

Conversion

Converted with the official upstream litert-torch export_hf (clean git worktree at upstream/main, dev-fork patches excluded). Olmo2ForCausalLM rides the stock converter with no custom code: QK-norm and OLMo-2's reordered post-norm lower to generic ops. The int4 recipe is blockwise (block 32) + OCTAV with the embedding at INT8.

2026-08-30 β€” tokenizer section replaced (weights unchanged)

The tokenizer.json embedded in model.litertlm carried the GPT-2 default pre-tokenizer instead of the model's own regex, so digit groups and punctuation followed by a newline were split differently from the upstream tokenizer on every turn (the role markers <|user|>\n / <|assistant|>\n alone differed by 3 tokens per turn). model.litertlm now embeds the upstream tokenizer.json byte for byte.

Tokenizer-only change: every section of the bundle except the tokenizer is byte-identical to the previous file (verified by sha256 per section), so the weights, the graph and the chat template are unchanged and the speed and memory numbers on this card still describe this file per token β€” only the file's own sha256 differs. Verified on the LiteRT-LM runtime: the default turn, 7 probe strings, the 223 standalone characters U+00A1–U+017F and every special token now tokenize identically to the upstream tokenizer, and the four ASCII-only test questions answer byte-identically to the previous file (same ids in, same tokens out). Prompts containing accented letters, symbols or emoji reach the model differently from before, so individual answers to such prompts can change. Unless a row says otherwise, the accuracy figures on this card were measured on the previous file, and any on-device rows were measured on the previous file too β€” the on-device gate has not been re-run on this one (the runtime's tokenizer code is the same on macOS and on device; the weights and graph are byte-identical). If you downloaded before 2026-08-30, re-download.

License

Apache-2.0, inherited from the base model allenai/OLMo-2-0425-1B-Instruct.


Want a different model on-device? Open a request β€” free, open weights only; the export and its measured numbers get published publicly.

Downloads last month
50
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for mlboydaisuke/OLMo-2-1B-Instruct-LiteRT