feat: stop on generation_config's eos_token_id as well as the tokenizer's EOS #114

Merged
rcsheets merged 1 commit from feat/generation-config-eos into main 2026-10-04 06:11:19 +00:00
Owner

gllm ended a turn only on the tokenizer's single EOS id (tokenizer_config.json's eos_token). A checkpoint's generation_config.json can name other ids in eos_token_id, and HF generate and vLLM stop on all of them. For some checkpoints the difference decides whether a reply ends at all:

checkpoint tokenizer EOS generation_config eos_token_id
allenai/Olmo-3-7B-Think `< endoftext
NVIDIA-Nemotron-3-Nano-30B-A3B `< im_end
GLM-5.3-Flash -- 154820, 154827, 154829
granite-4.2-8b `< im_end

OLMo 3's chat turns end in <|im_end|>, which its tokenizer's EOS does not name. Without this change gllm would generate past the end of every reply, writing further turns until max_tokens.

What changes

  • config.Load reads generation_config.json's eos_token_id, written as one id or as a list, into config.Model.StopTokenIDs.
    • A missing file, a missing field, or null gives no extra ids.
    • A file present in any other shape is an error, because dropping it silently would bring the runaway back.
  • The engine builds its stop set once (stopIDs): the tokenizer's EOS plus the generation-config ids. It uses every id in that set wherever it used the EOS:
    • finishing a turn with finish_reason: stop;
    • masking stops while a pending grammar is incomplete;
    • leaving the final stop token out of reasoning_tokens for a block that never closed.
  • The grammar: grammar.NewVocab(tok, extraStops...) treats each extra id as EOS.
    • Stops are masked while the JSON value is incomplete.
    • Once the value is complete, all stops are allowed and nothing else is. The model then ends with whichever it prefers, which for a chat model is its end-of-turn token.
    • Accepting a stop does not change the matcher's state.
  • AGENTS.md and the README no longer list multi-id EOS as missing for GLM.

Effect on checkpoints already served: Granite's ids are unchanged. Nemotron 3 now also stops on </s>, as transformers does.

Verification

  • go build ./..., go vet ./..., go test ./... pass; gofmt is clean.
  • TestLoadGenerationConfig covers one id, a list (OLMo 3's shape), no field, null, no file, and two malformed files.
  • TestGenerateStopTokenIDs runs the tiny CPU model greedily, names its third generated token as a generation-config stop, and checks that the turn ends there with finish_reason: stop.
  • TestExtraStops checks that an extra stop is masked before the value starts, leaves the state unchanged when accepted, and is allowed beside EOS, and only beside EOS, once {} is complete.

No GPU run: the change is in config loading and step-loop bookkeeping, covered by the CPU tests above.

🤖 Generated with Claude Code

gllm ended a turn only on the tokenizer's single EOS id (`tokenizer_config.json`'s `eos_token`). A checkpoint's `generation_config.json` can name other ids in `eos_token_id`, and HF `generate` and vLLM stop on all of them. For some checkpoints the difference decides whether a reply ends at all: | checkpoint | tokenizer EOS | generation_config `eos_token_id` | | --- | --- | --- | | allenai/Olmo-3-7B-Think | `<|endoftext|>` (100257) | 100265 `<|im_end|>`, 100257 | | NVIDIA-Nemotron-3-Nano-30B-A3B | `<|im_end|>` (11) | 2 `</s>`, 11 | | GLM-5.3-Flash | -- | 154820, 154827, 154829 | | granite-4.2-8b | `<|im_end|>` (100257) | 100257 | OLMo 3's chat turns end in `<|im_end|>`, which its tokenizer's EOS does not name. Without this change gllm would generate past the end of every reply, writing further turns until `max_tokens`. ## What changes - **`config.Load`** reads `generation_config.json`'s `eos_token_id`, written as one id or as a list, into `config.Model.StopTokenIDs`. - A missing file, a missing field, or `null` gives no extra ids. - A file present in any other shape is an error, because dropping it silently would bring the runaway back. - **The engine** builds its stop set once (`stopIDs`): the tokenizer's EOS plus the generation-config ids. It uses every id in that set wherever it used the EOS: - finishing a turn with `finish_reason: stop`; - masking stops while a pending grammar is incomplete; - leaving the final stop token out of `reasoning_tokens` for a block that never closed. - **The grammar:** `grammar.NewVocab(tok, extraStops...)` treats each extra id as EOS. - Stops are masked while the JSON value is incomplete. - Once the value is complete, all stops are allowed and nothing else is. The model then ends with whichever it prefers, which for a chat model is its end-of-turn token. - Accepting a stop does not change the matcher's state. - AGENTS.md and the README no longer list multi-id EOS as missing for GLM. Effect on checkpoints already served: Granite's ids are unchanged. Nemotron 3 now also stops on `</s>`, as transformers does. ## Verification - `go build ./...`, `go vet ./...`, `go test ./...` pass; gofmt is clean. - `TestLoadGenerationConfig` covers one id, a list (OLMo 3's shape), no field, `null`, no file, and two malformed files. - `TestGenerateStopTokenIDs` runs the tiny CPU model greedily, names its third generated token as a generation-config stop, and checks that the turn ends there with `finish_reason: stop`. - `TestExtraStops` checks that an extra stop is masked before the value starts, leaves the state unchanged when accepted, and is allowed beside EOS, and only beside EOS, once `{}` is complete. No GPU run: the change is in config loading and step-loop bookkeeping, covered by the CPU tests above. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
feat: stop on generation_config's eos_token_id as well as the tokenizer's EOS
All checks were successful
ci / test_and_build (pull_request) Successful in 1m1s
efe739d64d
The engine and the grammar ended a turn only on the tokenizer's single
EOS id. A checkpoint's generation_config.json can name others, and HF
generate and vLLM stop on those: OLMo 3's tokenizer names <|endoftext|>
while its chat turns end in <|im_end|>, so gllm would run on past every
reply; GLM ships three ids; Nemotron 3 adds </s> beside <|im_end|>.

config.Load now reads eos_token_id (one id or a list) into
StopTokenIDs. The engine stops on any of them, masks all of them while
a grammar is incomplete, and grammar.NewVocab takes them as extra stops
that are allowed exactly when the value is complete.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator

Automated review by pr-reviewer v0.54.0 | Safety Check | Nemotron 3 Nano | tracking id r-c1e247-2af2fe
This is an AI-generated review and may contain mistakes.

Status: ❌ Failed


Review failed. Tracking id r-c1e247-2af2fe — see logs for details.

Comment @pr-reviewer-bot retry to try again.

<!-- pr-reviewer:review --> *Automated review by [pr-reviewer](https://git.brooktrails.org/brooktrails/pr-reviewer) v0.54.0 | Safety Check | Nemotron 3 Nano | tracking id `r-c1e247-2af2fe`* *This is an AI-generated review and may contain mistakes.* **Status:** ❌ Failed --- Review failed. Tracking id `r-c1e247-2af2fe` — see logs for details. Comment `@pr-reviewer-bot retry` to try again.
rcsheets deleted branch feat/generation-config-eos 2026-10-04 06:11:19 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brooktrails/gllm!114
No description provided.