fix(engine): apply a response_format grammar to the answer, not the reasoning #109
Loading…
Reference in a new issue
No description provided.
Delete branch "fix/grammar-after-reasoning"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem
A chat completion with
response_formatagainst a model whose chat template thinks by default (Granite 4, Nemotron 3) returned emptycontentwith a fullcompletion_tokenscount andfinish_reason: stop.The grammar masked the logits from the first generated token. The prompt for such a turn ends in an open
<think>, and the token that closes the block (</think>) is not JSON, so it could never be sampled. The block stayed open, the model wrote its entire schema-valid answer inside it, and the reasoning split filed all of it as reasoning. A client that reads onlycontentsaw nothing.It started when requests that do not say began to follow the template's
enable_thinkingdefault, which is on for both models.Change
sampleBatchmasks and advances the acceptor once the reasoning block has closed; the marker itself is not fed to it. Reasoning is free text.max_tokensstill bounds the turn, so the worst case is alengthfinish with no content rather than astopfinish with no content.reasoning_tokensis reported for a block that never closed. It was 0, which read as a turn that had not reasoned; it now counts every generated token but a final EOS.Instrumentation
engine: request completeline gainscontent_bytes, plusreasoning_tokens,reasoning_bytesandreasoning_closedon a reasoning turn, andconstrained=trueunder a grammar.engine: request generated tokens but no content, with acause: reasoning block open atmax_tokens, block never closed, a stop string matching at the start of the answer, or output made of control tokens only.Behavior change for callers
A
response_formatrequest to a thinking-by-default model now actually reasons before it answers, and the reasoning is spent from the samemax_tokens. A caller that wants the whole budget for the answer sendschat_template_kwargs: {"enable_thinking": false}.Verification
TestGenerateChatReasoningConstrained: the constrained run reproduces the unconstrained reasoning up to the marker, then emits a valid JSON object as the answer. Fails with the fix reverted.TestGenerateChatReasoningConstrainedNoEOS: with a grammar pending the turn does not end inside the reasoning block, and an unclosed block reports its tokens as reasoning.go build ./...,go vet ./...,go test ./...pass on the CPU backend. Not run against a real checkpoint or on the GPU.🤖 Generated with Claude Code
Automated review by pr-reviewer v0.52.3 | Safety Check | Granite | tracking id
r-c00a57-7636b2This is an AI-generated review and may contain mistakes.
Status: ❌ Failed
Review failed. Tracking id
r-c00a57-7636b2— see logs for details.Comment
@pr-reviewer-bot retryto try again.