fix(engine): apply a response_format grammar to the answer, not the reasoning #109

Merged
rcsheets merged 1 commit from fix/grammar-after-reasoning into main 2026-10-02 19:57:52 +00:00
Owner

Problem

A chat completion with response_format against a model whose chat template thinks by default (Granite 4, Nemotron 3) returned empty content with a full completion_tokens count and finish_reason: stop.

The grammar masked the logits from the first generated token. The prompt for such a turn ends in an open <think>, and the token that closes the block (</think>) is not JSON, so it could never be sampled. The block stayed open, the model wrote its entire schema-valid answer inside it, and the reasoning split filed all of it as reasoning. A client that reads only content saw nothing.

It started when requests that do not say began to follow the template's enable_thinking default, which is on for both models.

Change

  • The grammar constrains the answer only. sampleBatch masks and advances the acceptor once the reasoning block has closed; the marker itself is not fed to it. Reasoning is free text.
  • A pending grammar masks EOS while the block is open. A turn that ends mid-thought returns no answer, which is what the constraint exists to rule out. max_tokens still bounds the turn, so the worst case is a length finish with no content rather than a stop finish with no content.
  • reasoning_tokens is reported for a block that never closed. It was 0, which read as a turn that had not reasoned; it now counts every generated token but a final EOS.

Instrumentation

  • The engine: request complete line gains content_bytes, plus reasoning_tokens, reasoning_bytes and reasoning_closed on a reasoning turn, and constrained=true under a grammar.
  • A request that generated tokens but returned no content logs a WARN, engine: request generated tokens but no content, with a cause: reasoning block open at max_tokens, block never closed, a stop string matching at the start of the answer, or output made of control tokens only.

Behavior change for callers

A response_format request to a thinking-by-default model now actually reasons before it answers, and the reasoning is spent from the same max_tokens. A caller that wants the whole budget for the answer sends chat_template_kwargs: {"enable_thinking": false}.

Verification

  • TestGenerateChatReasoningConstrained: the constrained run reproduces the unconstrained reasoning up to the marker, then emits a valid JSON object as the answer. Fails with the fix reverted.
  • TestGenerateChatReasoningConstrainedNoEOS: with a grammar pending the turn does not end inside the reasoning block, and an unclosed block reports its tokens as reasoning.
  • go build ./..., go vet ./..., go test ./... pass on the CPU backend. Not run against a real checkpoint or on the GPU.

🤖 Generated with Claude Code

## Problem A chat completion with `response_format` against a model whose chat template thinks by default (Granite 4, Nemotron 3) returned empty `content` with a full `completion_tokens` count and `finish_reason: stop`. The grammar masked the logits from the first generated token. The prompt for such a turn ends in an open `<think>`, and the token that closes the block (`</think>`) is not JSON, so it could never be sampled. The block stayed open, the model wrote its entire schema-valid answer inside it, and the reasoning split filed all of it as reasoning. A client that reads only `content` saw nothing. It started when requests that do not say began to follow the template's `enable_thinking` default, which is on for both models. ## Change - **The grammar constrains the answer only.** `sampleBatch` masks and advances the acceptor once the reasoning block has closed; the marker itself is not fed to it. Reasoning is free text. - **A pending grammar masks EOS while the block is open.** A turn that ends mid-thought returns no answer, which is what the constraint exists to rule out. `max_tokens` still bounds the turn, so the worst case is a `length` finish with no content rather than a `stop` finish with no content. - **`reasoning_tokens` is reported for a block that never closed.** It was 0, which read as a turn that had not reasoned; it now counts every generated token but a final EOS. ## Instrumentation - The `engine: request complete` line gains `content_bytes`, plus `reasoning_tokens`, `reasoning_bytes` and `reasoning_closed` on a reasoning turn, and `constrained=true` under a grammar. - A request that generated tokens but returned no content logs a WARN, `engine: request generated tokens but no content`, with a `cause`: reasoning block open at `max_tokens`, block never closed, a stop string matching at the start of the answer, or output made of control tokens only. ## Behavior change for callers A `response_format` request to a thinking-by-default model now actually reasons before it answers, and the reasoning is spent from the same `max_tokens`. A caller that wants the whole budget for the answer sends `chat_template_kwargs: {"enable_thinking": false}`. ## Verification - `TestGenerateChatReasoningConstrained`: the constrained run reproduces the unconstrained reasoning up to the marker, then emits a valid JSON object as the answer. Fails with the fix reverted. - `TestGenerateChatReasoningConstrainedNoEOS`: with a grammar pending the turn does not end inside the reasoning block, and an unclosed block reports its tokens as reasoning. - `go build ./...`, `go vet ./...`, `go test ./...` pass on the CPU backend. Not run against a real checkpoint or on the GPU. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
fix(engine): apply a response_format grammar to the answer, not the reasoning
All checks were successful
ci / test_and_build (pull_request) Successful in 34s
ddcd6ecb0d
A request with response_format against a model whose chat template thinks
by default (Granite 4, Nemotron 3) came back with empty content and a full
completion token count. The grammar masked the logits from the first
generated token, and the reasoning-end marker is not JSON, so it could
never be sampled: the reasoning block stayed open, the model wrote its
whole constrained answer inside it, and the split filed all of it as
reasoning.

The grammar now masks and advances only once the reasoning block has
closed, and the marker itself is not fed to the acceptor. While the block
is open a pending grammar still masks EOS, since a turn that ends
mid-thought returns no answer at all, which is what the constraint exists
to rule out; max_tokens still bounds the turn.

Two things made this hard to see from a response, and both change here:

- A reasoning block that never closed reported reasoning_tokens as 0, as
  if the turn had not reasoned. It now counts every generated token but a
  final EOS.
- The "request complete" log line now says where the tokens went
  (content_bytes, and reasoning_tokens / reasoning_bytes /
  reasoning_closed on a reasoning turn, constrained=true under a grammar),
  and a request that generated tokens but no content logs a WARN naming
  the cause.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator

Automated review by pr-reviewer v0.52.3 | Safety Check | Granite | tracking id r-c00a57-7636b2
This is an AI-generated review and may contain mistakes.

Status: ❌ Failed


Review failed. Tracking id r-c00a57-7636b2 — see logs for details.

Comment @pr-reviewer-bot retry to try again.

<!-- pr-reviewer:review --> *Automated review by [pr-reviewer](https://git.brooktrails.org/brooktrails/pr-reviewer) v0.52.3 | Safety Check | Granite | tracking id `r-c00a57-7636b2`* *This is an AI-generated review and may contain mistakes.* **Status:** ❌ Failed --- Review failed. Tracking id `r-c00a57-7636b2` — see logs for details. Comment `@pr-reviewer-bot retry` to try again.
rcsheets deleted branch fix/grammar-after-reasoning 2026-10-02 19:57:52 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brooktrails/gllm!109
No description provided.