feat(tokenizer): OLMo 3 Think chat template #118
Loading…
Reference in a new issue
No description provided.
Delete branch "feat/olmo3-chat-template"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Renders
allenai/Olmo-3-7B-Think's chat template; the same file ships with its NVFP4 quantizations. The template is selected by its fingerprint (aad5e783...). Before this change, chat completions against those checkpoints returned 501.The template
It is ChatML, but not Granite's. These are the differences, pinned by goldens from HF
apply_chat_template:<functions></functions>"). With one, every system message stays in place and gets "You do not currently have access to any functions.<functions></functions>" appended.<|im_start|>assistant\n<think>. The template has noenable_thinking.<|im_start|>and<|im_end|>are control tokens.<think>,</think>and<functions>are plain text in OLMo's vocab, so the reasoning split uses aTextMarker(#113).What changes
encodeOlmo3Thinkplus itschatTemplatesentry:reasoningEnd: "</think>",thinkingDefault: true.tokenizer.ErrChatOption: returned when a template cannot render what the request asks. Here that isenable_thinking: falseagainst a template that always reasons. The server maps it to 400, alongsideErrRequestTooLarge, rather than rendering<think></think>, a prompt the model was never trained on.make olmo3-testdatacaches the tokenizer,tokenizer_config.jsonandchat_template.jinjafor the tests, like the Granite and Nemotron targets.Verification
go build ./...,go vet ./...,go test ./...pass; gofmt is clean.TestHFEncodeChatOlmo3Goldens: 11 conversations fromtestdata/generate_hf_chat_olmo3.py(transformers 5.14.1) match token for token, with thinking unset and with thinking asked for. They cover generate_hf_chat.py's set (no system, system first, multi-turn, whitespace, empty, multi-line, non-ASCII, a mid-conversation system message) plus two with several system messages or a padded assistant turn.TestOlmo3ThinkingOffRefused:enable_thinking: falseyieldsErrChatOption.TestOlmo3ReasoningEnd: the marker is aTextMarker, and its forced close equals the tokenizer's encoding of"\n</think>\n", ending in the merged>\ntoken.TestErrorMapping:ErrChatOptionis a 400.reasoning_cutmax_thinking_tokens_hard: 64enable_thinking: falseWith the soft cut, OLMo's answer goes on working the problem in prose rather than stating the result. Granite showed the same pattern in the thinking-budget measurements. That is a property of the model under a cut, not of rendering.
🤖 Generated with Claude Code
Automated review by pr-reviewer v0.54.0 | Safety Check | Nemotron 3 Nano | tracking id
r-c1eb9f-778ec4This is an AI-generated review and may contain mistakes.
Status: ❌ Failed
Review failed. Tracking id
r-c1eb9f-778ec4— see logs for details.Comment
@pr-reviewer-bot retryto try again.