feat(tokenizer): OLMo 3 Think chat template #118

Merged
rcsheets merged 1 commit from feat/olmo3-chat-template into main 2026-10-04 06:28:35 +00:00
Owner

Renders allenai/Olmo-3-7B-Think's chat template; the same file ships with its NVFP4 quantizations. The template is selected by its fingerprint (aad5e783...). Before this change, chat completions against those checkpoints returned 501.

The template

It is ChatML, but not Granite's. These are the differences, pinned by goldens from HF apply_chat_template:

  • System turns: with no system message anywhere in the conversation, a default turn comes first ("You are OLMo, a helpful function-calling AI assistant built by Ai2. ... <functions></functions>"). With one, every system message stays in place and gets "You do not currently have access to any functions. <functions></functions>" appended.
  • Turns are rendered verbatim: no empty think block, no trimming.
  • Generation prompt: always <|im_start|>assistant\n<think>. The template has no enable_thinking.
  • Control tokens: only <|im_start|> and <|im_end|> are control tokens. <think>, </think> and <functions> are plain text in OLMo's vocab, so the reasoning split uses a TextMarker (#113).

What changes

  • encodeOlmo3Think plus its chatTemplates entry: reasoningEnd: "</think>", thinkingDefault: true.
  • tokenizer.ErrChatOption: returned when a template cannot render what the request asks. Here that is enable_thinking: false against a template that always reasons. The server maps it to 400, alongside ErrRequestTooLarge, rather than rendering <think></think>, a prompt the model was never trained on.
  • make olmo3-testdata caches the tokenizer, tokenizer_config.json and chat_template.jinja for the tests, like the Granite and Nemotron targets.

Verification

  • go build ./..., go vet ./..., go test ./... pass; gofmt is clean.
  • TestHFEncodeChatOlmo3Goldens: 11 conversations from testdata/generate_hf_chat_olmo3.py (transformers 5.14.1) match token for token, with thinking unset and with thinking asked for. They cover generate_hf_chat.py's set (no system, system first, multi-turn, whitespace, empty, multi-line, non-ASCII, a mid-conversation system message) plus two with several system messages or a padded assistant turn.
  • TestOlmo3ThinkingOffRefused: enable_thinking: false yields ErrChatOption.
  • TestOlmo3ReasoningEnd: the marker is a TextMarker, and its forced close equals the tokenizer's encoding of "\n</think>\n", ending in the merged >\n token.
  • TestErrorMapping: ErrChatOption is a 400.
  • End to end: run on the RTX PRO 6000 with this branch merged locally onto #117 (the OLMo 3 architecture), serving Olmo-3-7B-Think-nvfp4, greedy, "What is 17*23? Answer briefly.":
request finish reasoning_tokens log reasoning_cut answer
default stop (`< im_end >`) 1335
max_thinking_tokens_hard: 64 stop 64 hard 391
soft 32, hard 128 stop 37 (cut at a sentence boundary) soft 391 (the answer continues working the problem)
enable_thinking: false HTTP 400

With the soft cut, OLMo's answer goes on working the problem in prose rather than stating the result. Granite showed the same pattern in the thinking-budget measurements. That is a property of the model under a cut, not of rendering.

🤖 Generated with Claude Code

Renders `allenai/Olmo-3-7B-Think`'s chat template; the same file ships with its NVFP4 quantizations. The template is selected by its fingerprint (`aad5e783...`). Before this change, chat completions against those checkpoints returned 501. ## The template It is ChatML, but not Granite's. These are the differences, pinned by goldens from HF `apply_chat_template`: - **System turns:** with no system message anywhere in the conversation, a default turn comes first ("You are OLMo, a helpful function-calling AI assistant built by Ai2. ... `<functions></functions>`"). With one, every system message stays in place and gets "You do not currently have access to any functions. `<functions></functions>`" appended. - **Turns** are rendered verbatim: no empty think block, no trimming. - **Generation prompt:** always `<|im_start|>assistant\n<think>`. The template has no `enable_thinking`. - **Control tokens:** only `<|im_start|>` and `<|im_end|>` are control tokens. `<think>`, `</think>` and `<functions>` are plain text in OLMo's vocab, so the reasoning split uses a `TextMarker` (#113). ## What changes - **`encodeOlmo3Think`** plus its `chatTemplates` entry: `reasoningEnd: "</think>"`, `thinkingDefault: true`. - **`tokenizer.ErrChatOption`:** returned when a template cannot render what the request asks. Here that is `enable_thinking: false` against a template that always reasons. The server maps it to **400**, alongside `ErrRequestTooLarge`, rather than rendering `<think></think>`, a prompt the model was never trained on. - **`make olmo3-testdata`** caches the tokenizer, `tokenizer_config.json` and `chat_template.jinja` for the tests, like the Granite and Nemotron targets. ## Verification - `go build ./...`, `go vet ./...`, `go test ./...` pass; gofmt is clean. - `TestHFEncodeChatOlmo3Goldens`: 11 conversations from `testdata/generate_hf_chat_olmo3.py` (transformers 5.14.1) match token for token, with thinking unset and with thinking asked for. They cover generate_hf_chat.py's set (no system, system first, multi-turn, whitespace, empty, multi-line, non-ASCII, a mid-conversation system message) plus two with several system messages or a padded assistant turn. - `TestOlmo3ThinkingOffRefused`: `enable_thinking: false` yields `ErrChatOption`. - `TestOlmo3ReasoningEnd`: the marker is a `TextMarker`, and its forced close equals the tokenizer's encoding of `"\n</think>\n"`, ending in the merged `>\n` token. - `TestErrorMapping`: `ErrChatOption` is a 400. - **End to end:** run on the RTX PRO 6000 with this branch merged locally onto #117 (the OLMo 3 architecture), serving Olmo-3-7B-Think-nvfp4, greedy, "What is 17*23? Answer briefly.": | request | finish | reasoning_tokens | log `reasoning_cut` | answer | | --- | --- | --- | --- | --- | | default | stop (`<|im_end|>`) | 1335 | -- | 391 | | `max_thinking_tokens_hard: 64` | stop | 64 | hard | 391 | | soft 32, hard 128 | stop | 37 (cut at a sentence boundary) | soft | 391 (the answer continues working the problem) | | `enable_thinking: false` | HTTP 400 | | | | With the soft cut, OLMo's answer goes on working the problem in prose rather than stating the result. Granite showed the same pattern in the thinking-budget measurements. That is a property of the model under a cut, not of rendering. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
feat(tokenizer): OLMo 3 Think chat template
All checks were successful
ci / test_and_build (pull_request) Successful in 48s
8329db28e7
Renders allenai/Olmo-3-7B-Think's ChatML template, selected by its
fingerprint: a default "You are OLMo" system turn when the conversation
has none, the "no functions" sentence appended to each system message,
turns verbatim, and a generation prompt that always ends
"assistant\n<think>". The template has no enable_thinking, so a request
that explicitly turns thinking off gets ErrChatOption, which the server
reports as a 400, rather than a prompt the model was never trained on.
Its "</think>" is plain text in the vocab, so the reasoning split uses a
TextMarker.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator

Automated review by pr-reviewer v0.54.0 | Safety Check | Nemotron 3 Nano | tracking id r-c1eb9f-778ec4
This is an AI-generated review and may contain mistakes.

Status: ❌ Failed


Review failed. Tracking id r-c1eb9f-778ec4 — see logs for details.

Comment @pr-reviewer-bot retry to try again.

<!-- pr-reviewer:review --> *Automated review by [pr-reviewer](https://git.brooktrails.org/brooktrails/pr-reviewer) v0.54.0 | Safety Check | Nemotron 3 Nano | tracking id `r-c1eb9f-778ec4`* *This is an AI-generated review and may contain mistakes.* **Status:** ❌ Failed --- Review failed. Tracking id `r-c1eb9f-778ec4` — see logs for details. Comment `@pr-reviewer-bot retry` to try again.
rcsheets deleted branch feat/olmo3-chat-template 2026-10-04 06:28:36 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brooktrails/gllm!118
No description provided.