feat: per-config thinking switch and stored reasoning for gllm #118
Loading…
Reference in a new issue
No description provided.
Delete branch "feat/thinking-option"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem
Granite 4 and Nemotron 3 reason before they answer unless told not to, and gllm follows the model's default for a request that does not say. pr-reviewer had no way to say either way, and when a model did reason, the reasoning was discarded: a call that spent most of its output budget thinking showed a large token count beside a short response, and one cut off mid-thought showed an empty response and no evidence of why.
Change
A per-config Thinking setting.
model_configsgains a nullablethinkingcolumn, edited on/admin/configsbeside temperature: on, off, or "model default".chat_template_kwargs: {"enable_thinking": ...}on all four calls (preflight, discovery, extract, review). Unset sends nothing, so the model's chat template decides; the field must be absent rather than defaulted, since gllm rejects an explicittruefor a model with no reasoning mode.Stored and displayed reasoning. The gllm backend reads
reasoning_contentandusage.completion_tokens_details.reasoning_tokens, and each call'sreview_passesrow stores them (reasoning,reasoning_tokens). The comparison page shows the reasoning inside each shared pass's "Input and output", and as a collapsible "Model reasoning" in a config's own column for its review call. A call that did not reason adds nothing to the page. Nothing decides on the reasoning; it is there to be read.Things to know before merging or enabling
000021_model_config_thinkingand000022_review_pass_reasoning.000020is taken by #116. golang-migrate only moves forward from the database's current version, so if this deploys before #116, #116's000020would be skipped on that database. Merge #116 first, or renumber whichever lands second.reasoning_content; without that fix the page shows a reasoning token count with no text.max_tokens. The review caps are per config, but preflight and discovery are fixed at 1024 in code and may be tight with thinking on. Not changed here.review_passes.docs/gllm-schema-notes.mdgains a "Thinking" section covering the above.Verification
TestGLLMThinkingOnTheWire: the kwarg is present with the right value, or absent, on each of the four calls;TestVLLMIgnoresThinking.TestGLLMReasoningCaptured,TestGLLMReasoningWithoutAnswer: reasoning and its token count reach each result, including a turn with no answer.TestConfigThinkingRoundTrips,TestPassReasoningRoundTrips(Postgres): all three thinking states survive a round trip, clearing writes NULL back, and reasoning survives the two-phase pass upsert.TestConfigsListRendersThinking,TestComparePageShowsReasoning: the edit form preselects the stored choice; reasoning renders where its call is shown.gofmt,go vet,go build, andgo test -race ./...pass, the tracker tests against a local Postgres 16. Not exercised against a live gllm.🤖 Generated with Claude Code
Automated review by pr-reviewer v0.52.3 | Safety Check | Granite | tracking id
r-c00aef-6a9dc0This is an AI-generated review and may contain mistakes.
Status: ❌ Failed
Review failed. Tracking id
r-c00aef-6a9dc0— see logs for details.Comment
@pr-reviewer-bot retryto try again.