- Go 94.2%
- Cuda 4.4%
- C 0.7%
- Makefile 0.5%
- Python 0.2%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
|
|
||
| .forgejo/workflows | ||
| cmd | ||
| docs | ||
| internal | ||
| safetensors | ||
| .actrc | ||
| .changelog.env | ||
| .dockerignore | ||
| .gitignore | ||
| AGENTS.md | ||
| CHANGELOG.es.md | ||
| CHANGELOG.md | ||
| CHANGELOG.zh.md | ||
| CLAUDE.md | ||
| Dockerfile | ||
| Dockerfile.cuda | ||
| Dockerfile.cuda-prerelease | ||
| go.mod | ||
| go.sum | ||
| LICENSE | ||
| Makefile | ||
| README.md | ||
gllm
Kind of like vLLM, but in Go.
A minimal (but expandable) LLM inference server: OpenAI-compatible HTTP API, continuous batching, paged KV cache, CUDA backend via cgo.
Initial target: Mistral Small on an NVIDIA Blackwell GPU.
Status
Works end to end on the CPU reference backend (validated against HF transformers and mistral-common); CUDA is next. Mistral NeMo Instruct 2407 (12B) is the smallest known-compatible real checkpoint. Roughly in dependency order:
- safetensors loading (single-file and sharded checkpoints)
- config.json parsing (incl.
text_confighoisting for Mistral 3/4 multimodal checkpoints) - paged KV cache block allocator
- prefix KV caching (
--prefix-cache, off by default): content-addressed, refcounted block reuse across requests sharing a prompt prefix -- suffix-only prefill, lazy LRU eviction of released blocks, preempted sequences reuse their own blocks on recompute; hit rate and block states at/metrics(gllm_prefix_cache_*). Design notes in docs/prefix-kv-caching.md - continuous-batching scheduler (whole-prompt prefill, preemption with exact recompute when the cache fills, request cancellation; chunked prefill TODO)
- Tekken tokenizer (
tekken.json, v11 pattern; hand-written linear-time splitter, goldens from mistral-common; legacy v3 files, e.g. Mistral NeMo, with synthesized specials and inline-system chat layout) - Hugging Face tokenizer (
tokenizer.jsonbyte-level BPE, GPT-2/GPT-4 lineage; hand-written splitters for all three pre-tokenizer shapes --Sequence[Split(GPT-4 pattern), ByteLevel],Sequence[Split(tekken v11 pattern), ByteLevel](sharing the tekken splitter), and a bareByteLevelthat applies the GPT-2 pattern itself --ignore_mergesfast path, added/special tokens, BOS/EOS fromtokenizer_config.json; goldens from the tokenizers crate. GLM family, Llama 3, Granite 4, Nemotron 3, ...) - CPU reference ops (correctness baseline for kernels; float32, float64 accumulation)
- CUDA kernels: embedding, add, rmsnorm, rope, silu_and_mul, append_kv, paged_attention, and the Mamba2 ops (causal_conv1d, ssm_scan, gated_rmsnorm, relu2) (sm_120; validated op-by-op and end-to-end against the CPU reference and HF goldens)
- cuBLASLt GEMMs (row-major HF weights via a column-major op derivation)
- Mistral weight loading + forward pass (BF16 checkpoints widened to F32; validated against HF transformers on a tiny checked-in checkpoint). Multimodal nesting (
model.language_model./language_model.model., vision tower skipped), modelopt MIXED_PRECISION quant (attention FP8, MLP NVFP4), and YaRN rope scaling all load and are validated on tiny fixtures.Mistral-Medium-3.5-128B-NVFP4runs end to end on one 96 GB GPU (generates coherent text): NVFP4 block scales stay FP8 + a global scalar (dequant in the matmul), and the auto KV-cache sizer reserves the forward's peak scratch and caps the batch to fit. - BF16 weights resident on device (
--weights-dtype, defaultauto): a BF16 checkpoint's 2-D weights stay unwidened, halving resident weight bytes at no loss (widening bf16 is exact) and running a bf16 tensor-core GEMM on CUDA.--weights-dtype f32widens instead, for f32 activations when chasing a numerical difference;gllm planplans the same choice. Honored by themistralfamily and Nemotron-H; for an architecture whose loader always widens (mistral4, GLM),autoresolves to f32 andbf16is refused, rather than the engine claiming a residency it does not get - Granite (
GraniteForCausalLM), served by the same package: identical tensor layout plus four scalar multipliers, of whichattention_multiplierreplaces1/sqrt(head_dim)rather than compounding with it. Validated prefill+decode against HF transformers on a tiny checkpoint, CPU and CUDA.granite-4.2-30bruns end to end on one 96 GB GPU (54.5 GiB of BF16 weights + a 35.6 GiB / 72k-token KV cache) - [~] GLM-5.2 (
GlmMoeDsaForCausalLM): MLA + DSA sparse-attention indexer + 256-expert MoE. Host float32 CPU reference forward, validated prefill+decode against HF transformers on a tiny checkpoint. NVFP4 mixed loading (modelopt-packed experts dequantized on load, attention/shared/dense kept BF16) validated end to end on a tiny NVFP4 checkpoint. Remaining: device kernels + packed-on-device NVFP4 compute (host f32 dequant does not scale), chat template (the real ~470 GB checkpoint does not fit one GPU regardless) - [~] OLMo 3 (
Olmo3ForCausalLM), served by the same package: post-sublayer norms, q/k norms, three sliding-window layers (plain rope) to each full-attention one (YaRN). Validated prefill+decode against HF transformers on a tiny checkpoint, CPU and CUDA.allenai/Olmo-3-7B-Think(BF16) and its compressed-tensors NVFP4 quantization load in seconds on one RTX PRO 6000 and complete text coherently, including a fact recalled from 6,400 tokens back, past the 4096-token window. Remaining: its chat template - [~] Nemotron-H (
NemotronHForCausalLM, Nemotron 3 Nano): a Mamba2 / attention / MoE hybrid. Each running sequence holds a fixed-size recurrent state slot beside its KV blocks (design in docs/hybrid-state-cache.md);--prefix-cacheis refused, since a reused KV prefix has no state behind it. CPU reference forward validated prefill, decode, chunked prefill and batched sequences against HF transformers on tiny checkpoints, full precision and modelopt NVFP4; the real 30B-A3B checkpoint loads (gllm plan: 18.0 GiB of weights, its BF16 tensors kept unwidened). Chat renders its template (the same ChatML-with-think layout as Granite, goldens from HFapply_chat_template), with reasoning split on</think>. The Mamba2 ops have CUDA kernels, and the tiny fixtures match the same goldens on the GPU. The real 30B-A3B NVFP4 checkpoint serves on one RTX PRO 6000: it loads in 10 s and answers coherently, with thinking on by default (reasoning split intoreasoning_content) and off on request, at about 42 ms per decoded token under a 300 W cap. Remaining: throughput (a chunked prefill scan, gathered routed experts) - sampling (greedy, temperature/top-k/top-p, per-request seeds, EOS/length/stop-string finish)
- streaming (SSE) and the real Mistral chat template (token-level [INST]/[SYSTEM_PROMPT], goldens from mistral-common)
- Granite and Nemotron 3 chat template (token-level ChatML-with-think, goldens from HF
apply_chat_templatefor each), selected by fingerprinting the checkpoint'schat_template.jinjaso an unrecognized template reports not-implemented rather than rendering another model's layout - reasoning models:
chat_template_kwargs: {"enable_thinking": true}asks a capable model to think first (falseasks it not to; left out, the model's chat template decides, which for Granite 4 and Nemotron 3 means thinking), and its chain of thought comes back inreasoning_content(message and streamed delta) plususage.completion_tokens_details.reasoning_tokens-- vLLM's spelling throughout.max_thinking_tokens_soft/max_thinking_tokens_hardcap the reasoning block: past the soft budget the engine ends it at the next sentence boundary, past the hard one immediately, by emitting the template's own closing sequence (\n</think>\n) itself; the model then goes on to answer. Split by token id where the marker is an added token (Granite 4, Nemotron 3), so nothing a user or the model writes in prose can forge the boundary; a template whose marker is ordinary text is split by finding it in the decoded output. glchat's/thinkingsets thinking and its budgets mid-conversation (--think/--think=falsestill works at launch but is deprecated), and glchat shows reasoning dimmed above each reply and never replays it into the context - energy accounting: each request's
gllmmetadata carries the energy its batches used, in joules (batch_gpu_energy_j,proportional_gpu_energy_j), measured from the GPU board's cumulative NVML energy counter, sampled every 25 ms and interpolated onto each step's wall-clock window, so the billed energy adds up to what the device used./metricshas the device total and the attributed part; perfstats logs average power and J per token. CPU energy has anenergy.Meterslot for a privileged RAPL reader (RAPL is root-only) and is omitted until one is configured - logprobs: OpenAI
logprobs/top_logprobsfor generated tokens and vLLM'sprompt_logprobsextension for prompt positions (raw log-softmax of the final logits, before temperature/top-k/top-p; per-token rawbytes; prompt path reduced host-side in bounded position chunks; non-streaming) - [~] grammar-constrained decoding:
response_format: {"type":"json_object"}guarantees valid JSON via a byte-level JSON acceptor that masks each step's logits to tokens keeping the output on a valid path (internal/grammar; forces a top-level object and EOS once complete). On a turn that reasons first the constraint applies to the answer only: the reasoning block is free text, and the grammar starts at the first token after it closes.{"type":"json_schema"}constrains to a specific schema -- the same acceptor plus a schema cursor, so a required property cannot be dropped, an invented key never starts, and anenumvalue cannot be freeform text. Supported keywords aretype,properties,required,enum,items,additionalProperties: false,minItems/maxItems,minimum/maximumand theirexclusiveforms, and thex-gllm-orderedextension (opt-in fixed emission order, taken from therequiredarray); anything else is a 400 at request time rather than a silently unenforced constraint. Numeric bounds are decided on the digits as they arrive, not on the finished number: underminimum: 100the digit5is refused where it is emitted, since no suffix could rescue it and the model would otherwise generate digits it is never allowed to stop. Still to do:$ref,oneOf/anyOf,pattern - Prometheus metrics at
/metrics(token throughput counters, weights/KV/scratch/device-memory gauges, scheduler load), plus an optional periodic perf-stats log (--log-perfstats-interval=30s) over the same registry - per-request metadata: engine request id (
X-Request-Id+ agllmresponse object) and per-request CPU/GPU time accounting, plus measured GPU energy (see energy accounting above) POST /tokenize(vLLM-compatible extension): count apromptormessages(chat template applied) without a forward pass, so a client can size a request against the batch budget before sending; the count matches generation's prefill exactly- KV cache sizing from free VRAM (
--kv-cache auto; a unit-aware budget -- blocks/tokens/bytes -- and--max-model-lencap allocation to min(useful, affordable); host RAM sized the same way, cgroup-limit aware) - FP8 (E4M3) weight quantization: loads pre-quantized compressed-tensors / llm-compressor checkpoints (F8_E4M3 weights + per-channel or per-tensor
weight_scale), keeps them compact, and dequantizes in MatMul;--fp8-computepicks weight-only dequant (W8A16, default) or dynamic activation quant (W8A8). CPU reference validated end to end; the CUDA path dequantizes to a scratch buffer and reuses the cuBLASLt GEMM (weight-only) - NVFP4 quantization (Blackwell 4-bit microscaling): loads pre-quantized compressed-tensors
nvfp4-pack-quantizedcheckpoints (packed E2M1 weights two-per-byte + per-16-group E4M3weight_scale+ per-tensor FP32weight_global_scale), folds the two-level scale into one per-group scale at load, and unpacks/dequantizes in MatMul (weight-only W4A16). CPU reference validated end to end; the CUDA path fuses the dequant into a GEMV-style kernel for batches up tonvfp4GEMVMaxRowsrows (the common decode case), and falls back to unpacking+dequantizing to a scratch buffer + cuBLASLt GEMM above that (GPU-validated against the CPU reference). Still well behind vLLM's native FP4 tensor-core kernels -- see docs/perf-optimization-plan.md - settings characterization: gllm (or glbench beside it) sweeps a grid of inference settings -- sampling, reasoning budgets, quantized compute paths -- against a task set with checkable answers, and suggests settings for a goal the operator states (fewest tokens at a given accuracy, no reasoning leaks, ...). Today this is done by hand; the Thinking budgets section below is the first result, and AGENTS.md (Knobs for the operator's goals) sets out how new knobs should support it
Layout
cmd/gllm/ server binary
cmd/glchat/ reference terminal chat client (bubbletea)
internal/glchat/ instance discovery, streaming client, TUI
safetensors/ safetensors reading library (lazy per-tensor
reads, sharded checkpoints, spec validation;
importable outside gllm)
internal/engine/ step loop; ties everything together
internal/server/ OpenAI-compatible HTTP API
internal/scheduler/ request queue, batch assembly
internal/kvcache/ paged KV cache block allocator
internal/model/ model interface + architecture registry
internal/model/mistral/ Llama-shaped dense family: Mistral, Granite (GQA + RoPE + SwiGLU)
internal/model/glm/ GLM-5.2 (MLA + DSA indexer + MoE; CPU reference)
internal/model/nemotronh/ Nemotron-H (Mamba2 + attention + MoE hybrid)
internal/model/moe/ shared MoE routing (sigmoid / noaux_tc router)
internal/backend/ compute backend interface (op set)
internal/backend/cpu/ reference backend, always available
internal/backend/cuda/ cgo CUDA backend (build tag `cuda`)
internal/backend/cuda/kernels/ .cu kernels, built by nvcc
internal/tokenizer/ Tekken / HF tokenizer loading
internal/config/ HF config.json parsing
internal/quant/ FP8 (e4m3) + NVFP4 (e2m1) codecs + weight quantizers
internal/quantformat/ how a checkpoint spells a quantized weight (NVFP4 dialects, FP8 scales)
internal/tensor/ dtype/shape descriptor shared by the above
Building
Pure Go (CPU backend only -- no CUDA toolkit needed):
make build # or: go build ./...
With CUDA (requires CUDA 12.8+ for Blackwell):
make cuda # sm_120: GeForce RTX 50-series / RTX PRO
make cuda CUDA_ARCH=100 # sm_100: B100 / B200
Container image (pure-Go CPU build; the Kubernetes deployment lives in the brooktrails/infra repo under apps/gllm):
make image # tags harbor.brooktrails.org/library/gllm:<git sha>
make push
make image-cuda # CUDA build (Dockerfile.cuda), tagged :<git sha>-cuda
make push-cuda
Running
gllm serve --model /path/to/mistral-small --addr :8000
curl localhost:8000/v1/models
From the container image
The image built by make image (and pushed by CI to
harbor.brooktrails.org/brooktrails/gllm, tagged latest, the release
version, and the git short sha) is the pure-Go CPU build -- no CUDA. It is
built FROM scratch with the static gllm binary as the entrypoint, so
everything after the image name is a gllm argument. Mount the model
directory as a volume and publish the port:
docker run --rm -p 8000:8000 \
-v /path/to/mistral-small:/model:ro \
harbor.brooktrails.org/brooktrails/gllm:latest \
serve --model /model --served-model-name mistral-small --addr :8000
curl localhost:8000/v1/models
Pass --served-model-name: the API model id defaults to the checkpoint
directory's basename, which for a container mounted at /model is the
meaningless id model -- and requests naming the model anything else are
rejected with a 404.
podman run takes the same arguments. The container runs as the
unprivileged user 65534 (nobody), so the model files must be readable by
that uid -- world-readable files are enough; with rootless podman, add
--userns=keep-id:uid=65534,gid=65534 to map your own uid onto the
container user instead. On SELinux hosts use :ro,z on the volume so the
mount gets a container-accessible label.
There is no shell or anything else in the CPU image; to poke around inside
it, docker exec will not help -- inspect it with docker create +
docker cp, or run other subcommands directly
(e.g. ... gllm:latest plan --model /model --device-memory 96GiB).
That emptiness is also why the binary carries its own probe. A container
healthcheck runs inside the container, where there is no curl to call
/readyz with -- not in the CPU image, which has nothing at all, and not in
the CUDA image either, whose base has a shell but no HTTP client. So use
gllm health, which probes a running server and turns the answer into an
exit status:
gllm health # readiness (/readyz), localhost:8000
gllm health --endpoint host:8000 # somewhere else; :8000 alone also works
gllm health --live # liveness (/livez) instead
As a healthcheck, in a podman quadlet or compose file:
HealthCmd=/gllm health
Readiness is the default because it answers the question a load balancer asks: it reports 503 while the model loads and once a drain begins. Liveness stays 200 for as long as the process answers HTTP at all, so restarting on it would kill instances that are merely loading or draining. Give the healthcheck a start period longer than a weight load, and prefer not to restart on failure -- a slow load must not be mistaken for a dead server.
CI also pushes a CUDA-enabled variant, tagged with a -cuda suffix
(latest-cuda, <version>-cuda, <sha>-cuda) and built by
Dockerfile.cuda for sm_86 and sm_120 (override
CUDA_ARCHS to change that). Its runtime base is NVIDIA's CUDA runtime
image, so unlike the CPU image it does have a shell. The host needs the
NVIDIA driver and nvidia-container-toolkit; hand the GPU to the container
with --gpus all (docker) or --device nvidia.com/gpu=all (podman, CDI):
docker run --rm --gpus all -p 8000:8000 \
-v /path/to/mistral-small:/model:ro \
harbor.brooktrails.org/brooktrails/gllm:latest-cuda \
serve --model /model --served-model-name mistral-small --addr :8000
A quantized checkpoint is detected from its tensor dtypes and config.json
(quantization_config); no flag is needed to serve one. FP8 and NVFP4 are both
recognized automatically. --fp8-compute chooses how FP8 weights are used:
weight-only (dequantize weights, keep activations full precision; the default)
or w8a8 (also quantize activations, FP8 x FP8 -- reference on CPU; the CUDA
backend falls back to weight-only). NVFP4 checkpoints serve weight-only (W4A16):
the packed 4-bit weights are unpacked and dequantized in MatMul, with no flag to
set. For which FP8 checkpoints to try and how, see
docs/fp8-checkpoints.md.
To decide where to run a model before committing a GPU, gllm plan estimates
whether it fits on a device of a given size and the KV headroom it would leave,
without a GPU and without holding the weights resident:
gllm plan --model /path/to/model --device-memory 96GiB
gllm plan --model /path/to/model --device-memory 96GiB --json # for an orchestrator
It measures the weight footprint by running the real loader against a sizing
backend that tallies allocations instead of reserving them (so quantized
compaction, checkpoint nesting, and multimodal skips are exact, not guessed) and
splits the remaining memory with the same auto:fill math serve uses. The
weights must be readable, but only one tensor is held at a time. --max-seqs,
--max-model-len, --max-batch-tokens, and --kv-block-size mirror serve so
the plan reflects how the model would actually run, including
--max-batch-tokens auto: the budget sized to the longest prompt the device can
serve (the model's context or --max-model-len when the KV cache left over still
holds that many tokens, less when it does not), which matters because prefill is
not chunked and the default 8192-token budget is also the prompt ceiling. Architectures whose weights
are not device-resident (GLM's host-float32 reference forward) are reported as an
error rather than a bogus footprint.
Every response carries an X-Request-Id header and a gllm metadata object
(the request id, the batches it ran in, and its CPU/GPU time). Disable the body
object with serve --request-metadata=false or a per-request "gllm_metadata": false; the header is always sent.
On the CUDA backend the object also reports the energy the request used, in joules, measured from the GPU board's own energy counter rather than estimated:
"gllm": {
"request_id": "gdqz...",
"batch_gpu_time_us": 1366444,
"proportional_gpu_time_us": 1366444,
"batch_gpu_energy_j": 160.766,
"proportional_gpu_energy_j": 160.766,
...
}
Like the times, energy comes two ways, because continuous batching runs many
requests in one forward pass. batch_gpu_energy_j is what the GPU used during
every step the request was part of, shared with the other requests in those
steps. proportional_gpu_energy_j is the request's share, split by scheduled
tokens. Both are measured energy and include the card's static draw. The
counter is sampled every 25 ms and interpolated onto each step, so across
requests the billed energy adds up to what the card used while running them.
Where there is no meter (the CPU backend, or a driver without NVML's energy
counter) the fields are left out rather than reported as zero. CPU energy uses
the same fields (batch_cpu_energy_j, proportional_cpu_energy_j) but needs a
CPU energy source: RAPL counters are root-only, so it will come from a
privileged reader and is not reported yet.
Prometheus metrics are served at GET /metrics (token throughput counters,
weights/KV/scratch/device-memory gauges, scheduler load, and energy:
gllm_gpu_energy_joules_total is everything the GPU used while serving, idle
included, and gllm_gpu_energy_attributed_joules_total the part billed to
requests). For a log-only view, serve --log-perfstats-interval=30s logs a
performance summary every 30s from the same metrics -- the served model name
(served_model), prompt tokens ingested
and tokens generated in the interval (with per-second rates), a memory breakdown
(weights, KV cache, transient scratch, device free/total), the scheduler load,
and the GPU's energy, average power and joules per generated token over the
interval; omit the flag (or set 0) to disable it.
Two gllm-internal management endpoints (not part of the OpenAI API) support
orchestration: GET /v1/internal/status returns a JSON snapshot -- the served
model, resolved capacity (max seqs / model len / batch tokens, KV cache size),
device footprint and free/total memory, live running/waiting load, build
version, whether HTTP drain is enabled, and (when set) the request limit and
completed count -- so an orchestrator can route requests and make
model-placement decisions. POST /v1/internal/drain begins the same graceful
drain as a SIGTERM (stop accepting new requests, finish in-flight, then exit),
returning 202 immediately and idempotent, so an instance can be retired over
the wire without shell or signal access to the box. The drain endpoint is gated
on a shared secret: it is disabled unless serve --drain-secret=<secret> is
set, and callers must then present that secret verbatim in the X-Drain-Secret
header (SIGTERM drains regardless).
--power-budget declares the GPU power management limit this node is expected
to be running under, and gllm checks the device against it: 600W expects
exactly that, min=450W sets a throughput floor, max=300W an electrical
ceiling, min=450W,max=600W a range. gllm only ever reads the cap -- setting
one is device-global, outlives the process, and needs privileges a serving
process should not hold -- so the flag is an assertion about how the node was
provisioned, not a request to change it. The two directions are not treated
alike, because they are not the same problem. A cap below the budget costs
throughput on this node and nothing else, so it warns and serves. A cap above
it risks tripping an upstream breaker and taking down every host behind it, and
since a cap bounds draw rather than causing it, admitting traffic is what turns
a too-high cap into amps -- so it refuses to start. Append :warn or :fail to
override both. --power-breach-action decides what a running instance does if
the cap is raised past the ceiling while it serves: drain (the default: stop
admitting, finish in-flight, exit), warn, or exit (stop without waiting).
The power envelope is logged at startup and reported in the status payload
(power_limit_watts, power_usage_watts, power_capped, throttle,
power_budget, power_budget_met) and at /metrics
(gllm_device_power_limit_watts, gllm_device_power_usage_watts,
gllm_device_power_capped), whether or not a budget is declared. It needs the
driver's NVML, which is loaded on demand: without it the power fields are simply
absent.
Both endpoints answer from the moment the port binds, before the
(minutes-long) model load finishes: during the load the status snapshot is
{"state": "loading", "model_dir": ..., "elapsed_seconds": ...} so an
orchestrator can tell "starting" from "dead", the drain endpoint aborts the
load and exits (same secret gating), and every other endpoint answers 503
with a Retry-After hint instead of hanging. Once the load completes the
status payload reports "state": "serving" (then "draining" after a drain
is triggered).
Two probe endpoints follow the k8s liveness/readiness split. GET /livez is
liveness: 200 whenever the process is up and answering HTTP, including
during the load and a graceful drain -- point restart-on-failure supervision
here, since restarting an instance for loading or draining defeats both.
GET /readyz is readiness: 200 only while the instance should receive
traffic -- 503 during the load and 503 again once a drain has been
triggered, so a load balancer stops routing new requests while in-flight ones
finish. GET /healthz is a deprecated, temporary alias of /readyz from
before the split; migrate probes to /livez + /readyz. Each IP still
probing the alias is warned about once in the server log, so the log names
every prober left to migrate.
At --log-level debug each request logs an engine: request received line when
it is admitted (paired with the INFO engine: request complete line by
request_id), so tailing the log shows what is in flight. A request that does
not complete gets an INFO closing line too, with its request_id and an err
field: engine: request cancelled when the client went away, engine: request failed otherwise. Add
serve --debug-with-prompt-preview to include the text of the first few prompt
tokens as a prompt_preview field, with control tokens shown by name (e.g.
<s>[INST]hi there) rather than dropped (off by default -- it logs prompt
content); --debug-prompt-preview-size=N sets how many tokens.
serve --debug-with-response-preview is the response-side counterpart: at
--log-level debug it adds the text of the first few generated tokens to the
engine: request complete line as a response_preview field (also off by
default -- it logs response content); --debug-response-preview-size=N sets
how many tokens. The cancelled and failed lines carry it as well, showing what
was generated before the request stopped.
The engine: request complete line also says where the generated tokens went:
content_bytes is the size of the answer, and a turn that reasons first adds
reasoning_tokens, reasoning_bytes and reasoning_closed (whether the model
ever closed its reasoning block); constrained=true marks a response_format
request. A request that generated tokens and returned no content at all also
gets a WARN, engine: request generated tokens but no content, whose cause
names which case it was: a reasoning block still open at max_tokens or at
the end of the turn, a stop string matching at the start of the answer, or
output made of nothing but control tokens.
Either size can be -1 for no limit. A whole prompt or response is too much
for a log line, so an unlimited preview is written to a file instead --
<request_id>.prompt.txt or <request_id>.response.txt under
--debug-preview-dir -- and the log line carries its path as prompt_file or
response_file in place of the preview field. Without --debug-preview-dir,
gllm makes a fresh private directory for the run, gllm-previews-* in
$TMPDIR (or /tmp), and logs its path when the first file is written. The
files are owner-readable only and gllm never deletes them, so clear the
directory yourself after a debugging session.
Thinking budgets
A chat request to a reasoning model can cap its reasoning block with
max_thinking_tokens_soft (past it, the reasoning ends at the next sentence
boundary) and max_thinking_tokens_hard (it never runs past it). gllm closes
the block itself by writing the chat template's own \n</think>\n. The
failure to watch for is a leak: the model does not register that its
thinking is over and carries on reasoning in the answer, often ending with a
</think> of its own -- a stray </think> or <tool_call> in content is
the tell. Measured on Nemotron 3 Nano (12 problems with checkable answers, 4
seeds at the recommended temperature 1.0, 515 budget-cut turns):
- Set a soft budget whenever you set a hard one. A hard cut always lands
mid-thought: with only
hard: 64, 5 of 48 turns leaked. Any soft budget below the hard one brought leaks under 1% (3 of 390), and to none at all with a hard budget of 128 or more. - Leave 32 to 64 tokens between them. Past the soft budget the model reached a sentence boundary within a median 11 tokens (90% within 27, the longest 62), so a gap of 32 lets the soft budget make most cuts and 64 leaves the hard one as a rare backstop. Gaps of 16, 32 and 64 leaked alike, so a wider gap costs nothing beyond the tokens the model may use.
- Budget for accuracy, not just length. Unbudgeted, the model reasoned ~480 tokens and answered all 48 correctly; capped at 64 or 128 it answered 29-38, at 256 it answered 37-41. Tight budgets also shrink the answer less than the cap suggests: the model moves the work into a worked solution in the visible answer (a 64-token cap still averaged 200-265 tokens per turn, against 560 unbudgeted).
A reasonable starting point is soft = hard - 64 with hard no lower than
about 256 for multi-step problems.
These numbers are for one model; Granite 4.2's template closes reasoning the
same way, but its budgets have not been measured.
To exercise the whole stack on a CPU-only machine -- including building a tiny
servable checkpoint to curl -- see
docs/end-to-end-cpu.md.
Chatting with a server
glchat is a reference terminal client, built by make build alongside the
server (or on its own with make glchat):
glchat # find a local server and chat with it
glchat --endpoint host:9000 # talk to a specific one
With no flags it probes the port range gllm serve binds (:8000 plus the 20
it walks forward to when that is taken), asking each for /v1/internal/status.
That confirms an instance is live and past its model load, and reports what it
is serving -- so the model name never has to be typed. One instance connects
straight through; several open a picker ordered by load. Replies stream token
by token; ctrl+c interrupts a generation and ctrl+d (EOF) quits, as do
/quit and /exit. pgup/pgdn
(and shift+up/shift+down by the line) scroll the transcript; scrolling up
during a turn stops the stream following, and returning to the bottom resumes
it. Every turn closes
with a dim rule saying how it ended -- end of turn, truncated at the token cap,
interrupted, failed -- so a finished answer looks finished instead of merely
having stopped, along with what the client observed of it:
-- end of turn 2.3s 142 tok 70.0 tok/s ttft 310ms gdqz7k2p9v4m ----
Those timings are the client's own, measured around the request and so
including queue wait and the network -- what you actually waited. The trailing
gdqz... is the server's request id for the turn, the string to grep its log
with; it moves to a line of its own rather than being dropped when the rule
runs out of room, and it is shown for a failed turn too. For the server's view
of the same turn (per-batch CPU/GPU time), ask the server: it rides on the
response as the vendor gllm object, and /metrics has the fleet-wide
series.
A meter above the input says how much of the context window the conversation has taken and how much is left:
context [========--- ] 3.0k/8.2k 36% prefill batch cap
= is what the next request will send -- the whole transcript through the
model's chat template -- - the room reserved for the reply (--max-tokens),
and the blanks what neither has claimed. The count comes from the server's own
tokenizer over /tokenize, since a client cannot derive it; between turns it
falls back on the last reply's usage and marks the number ~ until the exact
one lands.
The window it measures against is the one the instance can actually serve,
which is often not the context length the model advertises. A sequence longer
than the whole KV cache can never be scheduled, and with --prefix-cache off
(the default) every request prefills its entire prompt, so the conversation has
to fit in one batch as well. A 131k-context model served with an 8k batch
budget holds an 8k conversation, and the meter says so -- naming the limit that
binds, as above -- rather than filling to 6% and then failing the request.
Input that starts with a / is a command to the client rather than a message
to the model. /read <path> sends a file's contents as your turn -- the text
itself, so the turns around it say what to do with it -- /width toggles the
text column limit (below), /thinking sets reasoning and its budgets (below),
/quit leaves, and /help lists what is available:
/read internal/engine/engine.go
/read ~/notes/today.md
The whole rest of the line is the path, so one with spaces in it needs no
quoting, and a leading ~ is expanded here rather than by a shell that never
saw the line. What to do with the file goes in the message after it.
/thinking changes what every later turn asks of a reasoning model.
on, off or default (the chat template decides, as without --think)
says whether it reasons; soft=N and hard=N set max_thinking_tokens_soft
and max_thinking_tokens_hard, and =off clears one. Words combine, words
left out keep their value, and bare /thinking reports the setting:
/thinking on soft=512 hard=1024
/thinking hard=off
Typing a / opens a list of what matches under the input, narrowing as you go:
/re
/read <path> send a file's contents as your message
start with a space to send it as a message -- the space is stripped
The parse stays out of the way of ordinary typing: only a single line whose
first word is a bare name (/read, /help) is a command, so /usr/lib/libc.so is missing, a pasted // TODO, and anything spanning more than one line are
sent as typed -- and none of them opens the list. A command-shaped word that is
not a command is reported instead of being sent, so a typo does not quietly
spend a turn. A message that really does have to open with /read is sent by
putting a space in front of it, which is stripped on the way out; the list
vanishing as you type that space is the confirmation that the line is now an
ordinary message.
On a terminal wider than 80 columns the transcript's text and the input wrap at
80, since prose stretched across a wide screen is hard to read. --width N
sets the limit and --width 0 lets the text fill the terminal; while chatting,
/width toggles between the limit and the full width, and /width N sets a
new limit.
gllm chat runs the same client. It is a wrapper that execs the glchat
binary rather than a subcommand proper, so the server binary does not link the
TUI's dependencies; it is only offered when that binary is present next to
gllm or on PATH.