Greedy outputs depend on batch composition (NVFP4 GEMV/GEMM switch at 32 rows is a candidate) #98

Open
opened 2026-09-28 18:44:10 +00:00 by rcsheets · 0 comments
Owner

Observation

At temperature 0, gllm returns different outputs for the same prompt depending on what else is in the batch. Requests served one at a time are reproducible.

Measured on NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 (gllm 51ef3b7, RTX PRO 6000 Blackwell) with cmd/glbench/configs/nemotron3-vs-vllm.yaml. Each workload cycles 8 distinct prompts, so every prompt is sent several times per run.

cell requests prompts with >1 distinct greedy output (of 8) most variants of one prompt
throughput, concurrency 8, run A 32 5 3
throughput, concurrency 8, run B 32 4 3
throughput, concurrency 32, run A 64 6 6
throughput, concurrency 32, run B 64 6 7
long_context, concurrency 1 (A vs B) 4 0 of 4 differ between runs -
parity, concurrency 1 (A vs B) 4 0 of 4 differ between runs -

Comparing the two runs request by request, 8 of 32 outputs differ at concurrency 8 and 37 of 64 at concurrency 32. The outputs stay coherent and diverge only partway through, when a token near a tie flips. One example, "Give three tips for debugging a CUDA out-of-memory error.", diverged at character 248:

A: ... OOM often stems from a single allocation that's larger than expected (e.g., a tensor that accidentally
B: ... OOM often stems from a hidden allocation (e.g., large tensors, temporary buffers, or unexpected copies)

For reference, vLLM 0.30.0 (default settings) in the same run also varies across batches: all 8 prompts produced more than one output, with up to 8 variants of a single prompt. Some batch dependence is normal for batched inference, so this is probably not a correctness bug. It is worth deciding on explicitly, though, because it:

  • makes batched glbench parity cells noisy (exact-match rates between two identical gllm runs were 75% and 42%), and
  • means a client cannot rely on greedy decoding being reproducible under load.

Candidate cause (unverified)

matMulNVFP4 (internal/backend/cuda/matmul.go) picks its kernel by row count:

  • n <= nvfp4GEMVMaxRows (32): fused-dequant GEMV with f32 activations
  • above that: unpack to bf16 and run the tensor-core GEMM with bf16-rounded activations

In the gathered MoE path, n for an expert's matmul is the number of (token, expert) pairs routed to that expert in the step. That depends on which other sequences share the step, so the same token's expert FFN can run at f32 or bf16 activation precision depending on its neighbours. Decode steps at concurrency 8 stay below 32 rows per expert. A prefill that shares a step with other prompts can cross the threshold, so the divergence would enter through prefill.

Other things that may contribute: reduction order in any kernel whose work split depends on batch shape (split-K, the attention dispatch threshold), and the order in which gathered expert outputs are summed.

Next steps

  1. Confirm the mechanism. Force the GEMV path for every row count (or the GEMM path for every row count), then check whether concurrency-8 outputs become batch-invariant.
  2. Decide the policy:
    • accept batch variance and document it (like vLLM's default), or
    • make the NVFP4 paths agree on activation precision, e.g. round the GEMV's activations to bf16 too, so the kernel switch changes speed and not results. The second option is cheap if the precision loss is acceptable.
  3. Make glbench's batched cells report run-to-run variance so batch-induced divergence is not mistaken for a cross-engine parity gap.

🤖 Generated with Claude Code

## Observation At temperature 0, gllm returns different outputs for the same prompt depending on what else is in the batch. Requests served one at a time are reproducible. Measured on NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 (gllm 51ef3b7, RTX PRO 6000 Blackwell) with `cmd/glbench/configs/nemotron3-vs-vllm.yaml`. Each workload cycles 8 distinct prompts, so every prompt is sent several times per run. | cell | requests | prompts with >1 distinct greedy output (of 8) | most variants of one prompt | | --- | ---: | ---: | ---: | | throughput, concurrency 8, run A | 32 | 5 | 3 | | throughput, concurrency 8, run B | 32 | 4 | 3 | | throughput, concurrency 32, run A | 64 | 6 | 6 | | throughput, concurrency 32, run B | 64 | 6 | 7 | | long_context, concurrency 1 (A vs B) | 4 | 0 of 4 differ between runs | - | | parity, concurrency 1 (A vs B) | 4 | 0 of 4 differ between runs | - | Comparing the two runs request by request, 8 of 32 outputs differ at concurrency 8 and 37 of 64 at concurrency 32. The outputs stay coherent and diverge only partway through, when a token near a tie flips. One example, "Give three tips for debugging a CUDA out-of-memory error.", diverged at character 248: ``` A: ... OOM often stems from a single allocation that's larger than expected (e.g., a tensor that accidentally B: ... OOM often stems from a hidden allocation (e.g., large tensors, temporary buffers, or unexpected copies) ``` For reference, vLLM 0.30.0 (default settings) in the same run also varies across batches: all 8 prompts produced more than one output, with up to 8 variants of a single prompt. Some batch dependence is normal for batched inference, so this is probably not a correctness bug. It is worth deciding on explicitly, though, because it: - makes batched glbench parity cells noisy (exact-match rates between two identical gllm runs were 75% and 42%), and - means a client cannot rely on greedy decoding being reproducible under load. ## Candidate cause (unverified) `matMulNVFP4` (`internal/backend/cuda/matmul.go`) picks its kernel by row count: - `n <= nvfp4GEMVMaxRows` (32): fused-dequant GEMV with **f32 activations** - above that: unpack to bf16 and run the tensor-core GEMM with **bf16-rounded activations** In the gathered MoE path, `n` for an expert's matmul is the number of (token, expert) pairs routed to that expert in the step. That depends on which other sequences share the step, so the same token's expert FFN can run at f32 or bf16 activation precision depending on its neighbours. Decode steps at concurrency 8 stay below 32 rows per expert. A prefill that shares a step with other prompts can cross the threshold, so the divergence would enter through prefill. Other things that may contribute: reduction order in any kernel whose work split depends on batch shape (split-K, the attention dispatch threshold), and the order in which gathered expert outputs are summed. ## Next steps 1. Confirm the mechanism. Force the GEMV path for every row count (or the GEMM path for every row count), then check whether concurrency-8 outputs become batch-invariant. 2. Decide the policy: - accept batch variance and document it (like vLLM's default), or - make the NVFP4 paths agree on activation precision, e.g. round the GEMV's activations to bf16 too, so the kernel switch changes speed and not results. The second option is cheap if the precision loss is acceptable. 3. Make glbench's batched cells report run-to-run variance so batch-induced divergence is not mistaken for a cross-engine parity gap. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brooktrails/gllm#98
No description provided.