docs(readme): thinking-budget guidance measured on four models #119

Open
rcsheets wants to merge 1 commit from docs/thinking-budgets-guidance into main
Owner

Rewrites the README's Thinking budgets section from the full set of measurements now available. The previous text had three problems:

  • It came from one model, Nemotron 3 Nano, on 12 problems.
  • It said Granite was unmeasured.
  • Its unbudgeted baseline was a 3000-token cap, which understated every model that reasons longer than that.

What the measurements were

  • Tasks: 54 problems with checkable numeric answers: 12 word problems, 28 computations and 14 Python program traces whose answers come from running them.
  • Models and sampling, each at its model card's recommendation:
model temperature top_p
Nemotron 3 Nano NVFP4 1.0 1.0
Granite 4.2 8b 1.0 0.95
Granite 4.2 30b 1.0 0.95
OLMo 3 7B Think, BF16 and NVFP4 0.6 0.95
  • Settings:
    • hard budgets of 128, 256 and 512, each alone and with soft budgets 32 and 64 below;
    • for Granite 30b and OLMo, also 1024 and 2048, with soft budgets 64 and 128 below;
    • an unbudgeted baseline at max_tokens 32768.
  • Seeds and pairing: 3 to 8 seeds per setting. Every setting runs on the same problems and seeds, so differences are sign-tested.
  • Grading: the final "Answer:" line, or the last \boxed{} when the model drops that line.
  • Leaks: three signals checked against hand labels: a stray </think> or <tool_call>, a run to max_tokens, and three or more self-talk phrases.

The raw data and analysis scripts are outside the repository (the thinking-budget experiment directory).

What the section now says

  • Cap at the model's own reasoning length. A budget costs accuracy until it reaches the model's natural reasoning length: Nemotron keeps 97% at 512, while Granite 30b and OLMo need 1024 for 95% or more. A 128 cap costs 17-57 points.
  • Soft budgets only for tight caps. A soft budget below the hard one helps only when the hard one is tight: 6-17 points at 128 on every model, each p <= 0.01. From 256 up it gained nothing significant, and in two cases it hurt.
  • Leak rates by family: at most about 6% for Nemotron and Granite. OLMo leaks 8-15% at 512 and below and 4-8% above, writing its own </think> even after the template close.
  • Thinking off is an option: Granite 4.2 with thinking off scores 98-99% for fewer tokens than the budgets that match it. Nemotron drops to 62%.
  • NVFP4 matches BF16 for OLMo 3, budget for budget.

Docs only; no code changes.

🤖 Generated with Claude Code

Rewrites the README's *Thinking budgets* section from the full set of measurements now available. The previous text had three problems: - It came from one model, Nemotron 3 Nano, on 12 problems. - It said Granite was unmeasured. - Its unbudgeted baseline was a 3000-token cap, which understated every model that reasons longer than that. ## What the measurements were - **Tasks:** 54 problems with checkable numeric answers: 12 word problems, 28 computations and 14 Python program traces whose answers come from running them. - **Models and sampling,** each at its model card's recommendation: | model | temperature | top_p | | --- | --- | --- | | Nemotron 3 Nano NVFP4 | 1.0 | 1.0 | | Granite 4.2 8b | 1.0 | 0.95 | | Granite 4.2 30b | 1.0 | 0.95 | | OLMo 3 7B Think, BF16 and NVFP4 | 0.6 | 0.95 | - **Settings:** - hard budgets of 128, 256 and 512, each alone and with soft budgets 32 and 64 below; - for Granite 30b and OLMo, also 1024 and 2048, with soft budgets 64 and 128 below; - an unbudgeted baseline at `max_tokens` 32768. - **Seeds and pairing:** 3 to 8 seeds per setting. Every setting runs on the same problems and seeds, so differences are sign-tested. - **Grading:** the final "Answer:" line, or the last `\boxed{}` when the model drops that line. - **Leaks:** three signals checked against hand labels: a stray `</think>` or `<tool_call>`, a run to `max_tokens`, and three or more self-talk phrases. The raw data and analysis scripts are outside the repository (the thinking-budget experiment directory). ## What the section now says - **Cap at the model's own reasoning length.** A budget costs accuracy until it reaches the model's natural reasoning length: Nemotron keeps 97% at 512, while Granite 30b and OLMo need 1024 for 95% or more. A 128 cap costs 17-57 points. - **Soft budgets only for tight caps.** A soft budget below the hard one helps only when the hard one is tight: 6-17 points at 128 on every model, each p <= 0.01. From 256 up it gained nothing significant, and in two cases it hurt. - **Leak rates by family:** at most about 6% for Nemotron and Granite. OLMo leaks 8-15% at 512 and below and 4-8% above, writing its own `</think>` even after the template close. - **Thinking off is an option:** Granite 4.2 with thinking off scores 98-99% for fewer tokens than the budgets that match it. Nemotron drops to 62%. - **NVFP4 matches BF16** for OLMo 3, budget for budget. Docs only; no code changes. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
docs(readme): thinking-budget guidance measured on four models
All checks were successful
ci / test_and_build (pull_request) Successful in 48s
83575145a8
The Thinking budgets section rested on one model and 12 problems, and
its unbudgeted baseline was a 3000-token cap. Replace it with results
from 54 problems with checkable answers across Nemotron 3 Nano, Granite
4.2 8b and 30b, and OLMo 3 7B Think: each model at its recommended
sampling, settings paired and sign-tested, and the unbudgeted baseline
at full context. Covers accuracy by hard budget with and without a soft
budget, when a soft budget helps (only at tight budgets), leak rates by
family, thinking off as an alternative, and the NVFP4 comparison.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator

Automated review by pr-reviewer v0.54.0 | Safety Check | Nemotron 3 Nano | tracking id r-c4892a-2b2265
This is an AI-generated review and may contain mistakes.

Status: ❌ Failed


Review failed. Tracking id r-c4892a-2b2265 — see logs for details.

Comment @pr-reviewer-bot retry to try again.

<!-- pr-reviewer:review --> *Automated review by [pr-reviewer](https://git.brooktrails.org/brooktrails/pr-reviewer) v0.54.0 | Safety Check | Nemotron 3 Nano | tracking id `r-c4892a-2b2265`* *This is an AI-generated review and may contain mistakes.* **Status:** ❌ Failed --- Review failed. Tracking id `r-c4892a-2b2265` — see logs for details. Comment `@pr-reviewer-bot retry` to try again.
All checks were successful
ci / test_and_build (pull_request) Successful in 48s
This pull request doesn't have enough approvals yet. 0 of 1 approvals granted.
You are not authorized to merge this pull request.
View command line instructions

Checkout

From your project repository, check out a new branch and test the changes.
git fetch -u origin docs/thinking-budgets-guidance:docs/thinking-budgets-guidance
git switch docs/thinking-budgets-guidance
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brooktrails/gllm!119
No description provided.