Kind of like vLLM, but in Go
  • Go 94.2%
  • Cuda 4.4%
  • C 0.7%
  • Makefile 0.5%
  • Python 0.2%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
Charley Sheets a1f427aff8
All checks were successful
ci / test_and_build (push) Successful in 1m3s
Build and push to Harbor / build-push (push) Successful in 11m47s
Merge pull request 'feat(tokenizer): OLMo 3 Think chat template' (#118) from feat/olmo3-chat-template into main
Reviewed-on: #118
2026-10-04 06:28:30 +00:00
.forgejo/workflows build(cuda): move the CUDA image to 13.4.1, relax its driver requirement 2026-09-27 18:22:39 -07:00
cmd feat(glchat): deprecate --think in favor of /thinking 2026-10-02 22:59:45 -07:00
docs fix(engine): reserve every layer kind's scratch for a hybrid forward 2026-10-01 23:44:30 -07:00
internal Merge pull request 'feat(tokenizer): OLMo 3 Think chat template' (#118) from feat/olmo3-chat-template into main 2026-10-04 06:28:30 +00:00
safetensors fix(safetensors): locate tensors from shard headers, not the index 2026-09-15 03:28:41 -07:00
.actrc ci: point act at the real forgejo-runner-go image for local runs 2026-06-30 19:09:02 -07:00
.changelog.env feat(changelog): adopt changelog-action with EN/ES/ZH changelogs 2026-07-25 09:45:11 +00:00
.dockerignore feat: add CUDA-enabled container image variant 2026-07-21 23:02:31 -07:00
.gitignore feat(glbench): serial gllm/vLLM performance + parity harness (Go) 2026-07-04 00:39:20 -07:00
AGENTS.md feat(model): OLMo 3 (Olmo3ForCausalLM) in the dense-transformer package 2026-10-03 22:56:38 -07:00
CHANGELOG.es.md fix(ci): only conventional-commit markers are breaking; 0.x bumps minor 2026-09-16 00:41:37 -07:00
CHANGELOG.md fix(ci): only conventional-commit markers are breaking; 0.x bumps minor 2026-09-16 00:41:37 -07:00
CHANGELOG.zh.md fix(ci): only conventional-commit markers are breaking; 0.x bumps minor 2026-09-16 00:41:37 -07:00
CLAUDE.md docs: refine agent guidance 2026-06-23 13:19:55 -07:00
Dockerfile chore(go): move the toolchain to Go 1.27 2026-09-16 01:14:58 -07:00
Dockerfile.cuda build(cuda): move the CUDA image to 13.4.1, relax its driver requirement 2026-09-27 18:22:39 -07:00
Dockerfile.cuda-prerelease build(cuda): move the CUDA image to 13.4.1, relax its driver requirement 2026-09-27 18:22:39 -07:00
go.mod chore(go): move the toolchain to Go 1.27 2026-09-16 01:14:58 -07:00
go.sum chore(deps): update all module dependencies 2026-07-21 22:14:48 -07:00
LICENSE Initial commit 2026-06-09 19:39:52 +00:00
Makefile feat(tokenizer): OLMo 3 Think chat template 2026-10-03 22:59:47 -07:00
README.md feat(model): OLMo 3 (Olmo3ForCausalLM) in the dense-transformer package 2026-10-03 22:56:38 -07:00

gllm

Kind of like vLLM, but in Go.

A minimal (but expandable) LLM inference server: OpenAI-compatible HTTP API, continuous batching, paged KV cache, CUDA backend via cgo.

Initial target: Mistral Small on an NVIDIA Blackwell GPU.

Status

Works end to end on the CPU reference backend (validated against HF transformers and mistral-common); CUDA is next. Mistral NeMo Instruct 2407 (12B) is the smallest known-compatible real checkpoint. Roughly in dependency order:

  • safetensors loading (single-file and sharded checkpoints)
  • config.json parsing (incl. text_config hoisting for Mistral 3/4 multimodal checkpoints)
  • paged KV cache block allocator
  • prefix KV caching (--prefix-cache, off by default): content-addressed, refcounted block reuse across requests sharing a prompt prefix -- suffix-only prefill, lazy LRU eviction of released blocks, preempted sequences reuse their own blocks on recompute; hit rate and block states at /metrics (gllm_prefix_cache_*). Design notes in docs/prefix-kv-caching.md
  • continuous-batching scheduler (whole-prompt prefill, preemption with exact recompute when the cache fills, request cancellation; chunked prefill TODO)
  • Tekken tokenizer (tekken.json, v11 pattern; hand-written linear-time splitter, goldens from mistral-common; legacy v3 files, e.g. Mistral NeMo, with synthesized specials and inline-system chat layout)
  • Hugging Face tokenizer (tokenizer.json byte-level BPE, GPT-2/GPT-4 lineage; hand-written splitters for all three pre-tokenizer shapes -- Sequence[Split(GPT-4 pattern), ByteLevel], Sequence[Split(tekken v11 pattern), ByteLevel] (sharing the tekken splitter), and a bare ByteLevel that applies the GPT-2 pattern itself -- ignore_merges fast path, added/special tokens, BOS/EOS from tokenizer_config.json; goldens from the tokenizers crate. GLM family, Llama 3, Granite 4, Nemotron 3, ...)
  • CPU reference ops (correctness baseline for kernels; float32, float64 accumulation)
  • CUDA kernels: embedding, add, rmsnorm, rope, silu_and_mul, append_kv, paged_attention, and the Mamba2 ops (causal_conv1d, ssm_scan, gated_rmsnorm, relu2) (sm_120; validated op-by-op and end-to-end against the CPU reference and HF goldens)
  • cuBLASLt GEMMs (row-major HF weights via a column-major op derivation)
  • Mistral weight loading + forward pass (BF16 checkpoints widened to F32; validated against HF transformers on a tiny checked-in checkpoint). Multimodal nesting (model.language_model. / language_model.model., vision tower skipped), modelopt MIXED_PRECISION quant (attention FP8, MLP NVFP4), and YaRN rope scaling all load and are validated on tiny fixtures. Mistral-Medium-3.5-128B-NVFP4 runs end to end on one 96 GB GPU (generates coherent text): NVFP4 block scales stay FP8 + a global scalar (dequant in the matmul), and the auto KV-cache sizer reserves the forward's peak scratch and caps the batch to fit.
  • BF16 weights resident on device (--weights-dtype, default auto): a BF16 checkpoint's 2-D weights stay unwidened, halving resident weight bytes at no loss (widening bf16 is exact) and running a bf16 tensor-core GEMM on CUDA. --weights-dtype f32 widens instead, for f32 activations when chasing a numerical difference; gllm plan plans the same choice. Honored by the mistral family and Nemotron-H; for an architecture whose loader always widens (mistral4, GLM), auto resolves to f32 and bf16 is refused, rather than the engine claiming a residency it does not get
  • Granite (GraniteForCausalLM), served by the same package: identical tensor layout plus four scalar multipliers, of which attention_multiplier replaces 1/sqrt(head_dim) rather than compounding with it. Validated prefill+decode against HF transformers on a tiny checkpoint, CPU and CUDA. granite-4.2-30b runs end to end on one 96 GB GPU (54.5 GiB of BF16 weights + a 35.6 GiB / 72k-token KV cache)
  • [~] GLM-5.2 (GlmMoeDsaForCausalLM): MLA + DSA sparse-attention indexer + 256-expert MoE. Host float32 CPU reference forward, validated prefill+decode against HF transformers on a tiny checkpoint. NVFP4 mixed loading (modelopt-packed experts dequantized on load, attention/shared/dense kept BF16) validated end to end on a tiny NVFP4 checkpoint. Remaining: device kernels + packed-on-device NVFP4 compute (host f32 dequant does not scale), chat template (the real ~470 GB checkpoint does not fit one GPU regardless)
  • [~] OLMo 3 (Olmo3ForCausalLM), served by the same package: post-sublayer norms, q/k norms, three sliding-window layers (plain rope) to each full-attention one (YaRN). Validated prefill+decode against HF transformers on a tiny checkpoint, CPU and CUDA. allenai/Olmo-3-7B-Think (BF16) and its compressed-tensors NVFP4 quantization load in seconds on one RTX PRO 6000 and complete text coherently, including a fact recalled from 6,400 tokens back, past the 4096-token window. Remaining: its chat template
  • [~] Nemotron-H (NemotronHForCausalLM, Nemotron 3 Nano): a Mamba2 / attention / MoE hybrid. Each running sequence holds a fixed-size recurrent state slot beside its KV blocks (design in docs/hybrid-state-cache.md); --prefix-cache is refused, since a reused KV prefix has no state behind it. CPU reference forward validated prefill, decode, chunked prefill and batched sequences against HF transformers on tiny checkpoints, full precision and modelopt NVFP4; the real 30B-A3B checkpoint loads (gllm plan: 18.0 GiB of weights, its BF16 tensors kept unwidened). Chat renders its template (the same ChatML-with-think layout as Granite, goldens from HF apply_chat_template), with reasoning split on </think>. The Mamba2 ops have CUDA kernels, and the tiny fixtures match the same goldens on the GPU. The real 30B-A3B NVFP4 checkpoint serves on one RTX PRO 6000: it loads in 10 s and answers coherently, with thinking on by default (reasoning split into reasoning_content) and off on request, at about 42 ms per decoded token under a 300 W cap. Remaining: throughput (a chunked prefill scan, gathered routed experts)
  • sampling (greedy, temperature/top-k/top-p, per-request seeds, EOS/length/stop-string finish)
  • streaming (SSE) and the real Mistral chat template (token-level [INST]/[SYSTEM_PROMPT], goldens from mistral-common)
  • Granite and Nemotron 3 chat template (token-level ChatML-with-think, goldens from HF apply_chat_template for each), selected by fingerprinting the checkpoint's chat_template.jinja so an unrecognized template reports not-implemented rather than rendering another model's layout
  • reasoning models: chat_template_kwargs: {"enable_thinking": true} asks a capable model to think first (false asks it not to; left out, the model's chat template decides, which for Granite 4 and Nemotron 3 means thinking), and its chain of thought comes back in reasoning_content (message and streamed delta) plus usage.completion_tokens_details.reasoning_tokens -- vLLM's spelling throughout. max_thinking_tokens_soft / max_thinking_tokens_hard cap the reasoning block: past the soft budget the engine ends it at the next sentence boundary, past the hard one immediately, by emitting the template's own closing sequence (\n</think>\n) itself; the model then goes on to answer. Split by token id where the marker is an added token (Granite 4, Nemotron 3), so nothing a user or the model writes in prose can forge the boundary; a template whose marker is ordinary text is split by finding it in the decoded output. glchat's /thinking sets thinking and its budgets mid-conversation (--think / --think=false still works at launch but is deprecated), and glchat shows reasoning dimmed above each reply and never replays it into the context
  • energy accounting: each request's gllm metadata carries the energy its batches used, in joules (batch_gpu_energy_j, proportional_gpu_energy_j), measured from the GPU board's cumulative NVML energy counter, sampled every 25 ms and interpolated onto each step's wall-clock window, so the billed energy adds up to what the device used. /metrics has the device total and the attributed part; perfstats logs average power and J per token. CPU energy has an energy.Meter slot for a privileged RAPL reader (RAPL is root-only) and is omitted until one is configured
  • logprobs: OpenAI logprobs/top_logprobs for generated tokens and vLLM's prompt_logprobs extension for prompt positions (raw log-softmax of the final logits, before temperature/top-k/top-p; per-token raw bytes; prompt path reduced host-side in bounded position chunks; non-streaming)
  • [~] grammar-constrained decoding: response_format: {"type":"json_object"} guarantees valid JSON via a byte-level JSON acceptor that masks each step's logits to tokens keeping the output on a valid path (internal/grammar; forces a top-level object and EOS once complete). On a turn that reasons first the constraint applies to the answer only: the reasoning block is free text, and the grammar starts at the first token after it closes. {"type":"json_schema"} constrains to a specific schema -- the same acceptor plus a schema cursor, so a required property cannot be dropped, an invented key never starts, and an enum value cannot be freeform text. Supported keywords are type, properties, required, enum, items, additionalProperties: false, minItems/maxItems, minimum/maximum and their exclusive forms, and the x-gllm-ordered extension (opt-in fixed emission order, taken from the required array); anything else is a 400 at request time rather than a silently unenforced constraint. Numeric bounds are decided on the digits as they arrive, not on the finished number: under minimum: 100 the digit 5 is refused where it is emitted, since no suffix could rescue it and the model would otherwise generate digits it is never allowed to stop. Still to do: $ref, oneOf/anyOf, pattern
  • Prometheus metrics at /metrics (token throughput counters, weights/KV/scratch/device-memory gauges, scheduler load), plus an optional periodic perf-stats log (--log-perfstats-interval=30s) over the same registry
  • per-request metadata: engine request id (X-Request-Id + a gllm response object) and per-request CPU/GPU time accounting, plus measured GPU energy (see energy accounting above)
  • POST /tokenize (vLLM-compatible extension): count a prompt or messages (chat template applied) without a forward pass, so a client can size a request against the batch budget before sending; the count matches generation's prefill exactly
  • KV cache sizing from free VRAM (--kv-cache auto; a unit-aware budget -- blocks/tokens/bytes -- and --max-model-len cap allocation to min(useful, affordable); host RAM sized the same way, cgroup-limit aware)
  • FP8 (E4M3) weight quantization: loads pre-quantized compressed-tensors / llm-compressor checkpoints (F8_E4M3 weights + per-channel or per-tensor weight_scale), keeps them compact, and dequantizes in MatMul; --fp8-compute picks weight-only dequant (W8A16, default) or dynamic activation quant (W8A8). CPU reference validated end to end; the CUDA path dequantizes to a scratch buffer and reuses the cuBLASLt GEMM (weight-only)
  • NVFP4 quantization (Blackwell 4-bit microscaling): loads pre-quantized compressed-tensors nvfp4-pack-quantized checkpoints (packed E2M1 weights two-per-byte + per-16-group E4M3 weight_scale + per-tensor FP32 weight_global_scale), folds the two-level scale into one per-group scale at load, and unpacks/dequantizes in MatMul (weight-only W4A16). CPU reference validated end to end; the CUDA path fuses the dequant into a GEMV-style kernel for batches up to nvfp4GEMVMaxRows rows (the common decode case), and falls back to unpacking+dequantizing to a scratch buffer + cuBLASLt GEMM above that (GPU-validated against the CPU reference). Still well behind vLLM's native FP4 tensor-core kernels -- see docs/perf-optimization-plan.md
  • settings characterization: gllm (or glbench beside it) sweeps a grid of inference settings -- sampling, reasoning budgets, quantized compute paths -- against a task set with checkable answers, and suggests settings for a goal the operator states (fewest tokens at a given accuracy, no reasoning leaks, ...). Today this is done by hand; the Thinking budgets section below is the first result, and AGENTS.md (Knobs for the operator's goals) sets out how new knobs should support it

Layout

cmd/gllm/                       server binary
cmd/glchat/                     reference terminal chat client (bubbletea)
internal/glchat/                instance discovery, streaming client, TUI
safetensors/                    safetensors reading library (lazy per-tensor
                                reads, sharded checkpoints, spec validation;
                                importable outside gllm)
internal/engine/                step loop; ties everything together
internal/server/                OpenAI-compatible HTTP API
internal/scheduler/             request queue, batch assembly
internal/kvcache/               paged KV cache block allocator
internal/model/                 model interface + architecture registry
internal/model/mistral/         Llama-shaped dense family: Mistral, Granite (GQA + RoPE + SwiGLU)
internal/model/glm/             GLM-5.2 (MLA + DSA indexer + MoE; CPU reference)
internal/model/nemotronh/       Nemotron-H (Mamba2 + attention + MoE hybrid)
internal/model/moe/             shared MoE routing (sigmoid / noaux_tc router)
internal/backend/               compute backend interface (op set)
internal/backend/cpu/           reference backend, always available
internal/backend/cuda/          cgo CUDA backend (build tag `cuda`)
internal/backend/cuda/kernels/  .cu kernels, built by nvcc
internal/tokenizer/             Tekken / HF tokenizer loading
internal/config/                HF config.json parsing
internal/quant/                 FP8 (e4m3) + NVFP4 (e2m1) codecs + weight quantizers
internal/quantformat/           how a checkpoint spells a quantized weight (NVFP4 dialects, FP8 scales)
internal/tensor/                dtype/shape descriptor shared by the above

Building

Pure Go (CPU backend only -- no CUDA toolkit needed):

make build          # or: go build ./...

With CUDA (requires CUDA 12.8+ for Blackwell):

make cuda                   # sm_120: GeForce RTX 50-series / RTX PRO
make cuda CUDA_ARCH=100     # sm_100: B100 / B200

Container image (pure-Go CPU build; the Kubernetes deployment lives in the brooktrails/infra repo under apps/gllm):

make image                  # tags harbor.brooktrails.org/library/gllm:<git sha>
make push
make image-cuda             # CUDA build (Dockerfile.cuda), tagged :<git sha>-cuda
make push-cuda

Running

gllm serve --model /path/to/mistral-small --addr :8000
curl localhost:8000/v1/models

From the container image

The image built by make image (and pushed by CI to harbor.brooktrails.org/brooktrails/gllm, tagged latest, the release version, and the git short sha) is the pure-Go CPU build -- no CUDA. It is built FROM scratch with the static gllm binary as the entrypoint, so everything after the image name is a gllm argument. Mount the model directory as a volume and publish the port:

docker run --rm -p 8000:8000 \
    -v /path/to/mistral-small:/model:ro \
    harbor.brooktrails.org/brooktrails/gllm:latest \
    serve --model /model --served-model-name mistral-small --addr :8000
curl localhost:8000/v1/models

Pass --served-model-name: the API model id defaults to the checkpoint directory's basename, which for a container mounted at /model is the meaningless id model -- and requests naming the model anything else are rejected with a 404.

podman run takes the same arguments. The container runs as the unprivileged user 65534 (nobody), so the model files must be readable by that uid -- world-readable files are enough; with rootless podman, add --userns=keep-id:uid=65534,gid=65534 to map your own uid onto the container user instead. On SELinux hosts use :ro,z on the volume so the mount gets a container-accessible label.

There is no shell or anything else in the CPU image; to poke around inside it, docker exec will not help -- inspect it with docker create + docker cp, or run other subcommands directly (e.g. ... gllm:latest plan --model /model --device-memory 96GiB).

That emptiness is also why the binary carries its own probe. A container healthcheck runs inside the container, where there is no curl to call /readyz with -- not in the CPU image, which has nothing at all, and not in the CUDA image either, whose base has a shell but no HTTP client. So use gllm health, which probes a running server and turns the answer into an exit status:

gllm health                          # readiness (/readyz), localhost:8000
gllm health --endpoint host:8000     # somewhere else; :8000 alone also works
gllm health --live                   # liveness (/livez) instead

As a healthcheck, in a podman quadlet or compose file:

HealthCmd=/gllm health

Readiness is the default because it answers the question a load balancer asks: it reports 503 while the model loads and once a drain begins. Liveness stays 200 for as long as the process answers HTTP at all, so restarting on it would kill instances that are merely loading or draining. Give the healthcheck a start period longer than a weight load, and prefer not to restart on failure -- a slow load must not be mistaken for a dead server.

CI also pushes a CUDA-enabled variant, tagged with a -cuda suffix (latest-cuda, <version>-cuda, <sha>-cuda) and built by Dockerfile.cuda for sm_86 and sm_120 (override CUDA_ARCHS to change that). Its runtime base is NVIDIA's CUDA runtime image, so unlike the CPU image it does have a shell. The host needs the NVIDIA driver and nvidia-container-toolkit; hand the GPU to the container with --gpus all (docker) or --device nvidia.com/gpu=all (podman, CDI):

docker run --rm --gpus all -p 8000:8000 \
    -v /path/to/mistral-small:/model:ro \
    harbor.brooktrails.org/brooktrails/gllm:latest-cuda \
    serve --model /model --served-model-name mistral-small --addr :8000

A quantized checkpoint is detected from its tensor dtypes and config.json (quantization_config); no flag is needed to serve one. FP8 and NVFP4 are both recognized automatically. --fp8-compute chooses how FP8 weights are used: weight-only (dequantize weights, keep activations full precision; the default) or w8a8 (also quantize activations, FP8 x FP8 -- reference on CPU; the CUDA backend falls back to weight-only). NVFP4 checkpoints serve weight-only (W4A16): the packed 4-bit weights are unpacked and dequantized in MatMul, with no flag to set. For which FP8 checkpoints to try and how, see docs/fp8-checkpoints.md.

To decide where to run a model before committing a GPU, gllm plan estimates whether it fits on a device of a given size and the KV headroom it would leave, without a GPU and without holding the weights resident:

gllm plan --model /path/to/model --device-memory 96GiB
gllm plan --model /path/to/model --device-memory 96GiB --json   # for an orchestrator

It measures the weight footprint by running the real loader against a sizing backend that tallies allocations instead of reserving them (so quantized compaction, checkpoint nesting, and multimodal skips are exact, not guessed) and splits the remaining memory with the same auto:fill math serve uses. The weights must be readable, but only one tensor is held at a time. --max-seqs, --max-model-len, --max-batch-tokens, and --kv-block-size mirror serve so the plan reflects how the model would actually run, including --max-batch-tokens auto: the budget sized to the longest prompt the device can serve (the model's context or --max-model-len when the KV cache left over still holds that many tokens, less when it does not), which matters because prefill is not chunked and the default 8192-token budget is also the prompt ceiling. Architectures whose weights are not device-resident (GLM's host-float32 reference forward) are reported as an error rather than a bogus footprint.

Every response carries an X-Request-Id header and a gllm metadata object (the request id, the batches it ran in, and its CPU/GPU time). Disable the body object with serve --request-metadata=false or a per-request "gllm_metadata": false; the header is always sent.

On the CUDA backend the object also reports the energy the request used, in joules, measured from the GPU board's own energy counter rather than estimated:

"gllm": {
  "request_id": "gdqz...",
  "batch_gpu_time_us": 1366444,
  "proportional_gpu_time_us": 1366444,
  "batch_gpu_energy_j": 160.766,
  "proportional_gpu_energy_j": 160.766,
  ...
}

Like the times, energy comes two ways, because continuous batching runs many requests in one forward pass. batch_gpu_energy_j is what the GPU used during every step the request was part of, shared with the other requests in those steps. proportional_gpu_energy_j is the request's share, split by scheduled tokens. Both are measured energy and include the card's static draw. The counter is sampled every 25 ms and interpolated onto each step, so across requests the billed energy adds up to what the card used while running them. Where there is no meter (the CPU backend, or a driver without NVML's energy counter) the fields are left out rather than reported as zero. CPU energy uses the same fields (batch_cpu_energy_j, proportional_cpu_energy_j) but needs a CPU energy source: RAPL counters are root-only, so it will come from a privileged reader and is not reported yet.

Prometheus metrics are served at GET /metrics (token throughput counters, weights/KV/scratch/device-memory gauges, scheduler load, and energy: gllm_gpu_energy_joules_total is everything the GPU used while serving, idle included, and gllm_gpu_energy_attributed_joules_total the part billed to requests). For a log-only view, serve --log-perfstats-interval=30s logs a performance summary every 30s from the same metrics -- the served model name (served_model), prompt tokens ingested and tokens generated in the interval (with per-second rates), a memory breakdown (weights, KV cache, transient scratch, device free/total), the scheduler load, and the GPU's energy, average power and joules per generated token over the interval; omit the flag (or set 0) to disable it.

Two gllm-internal management endpoints (not part of the OpenAI API) support orchestration: GET /v1/internal/status returns a JSON snapshot -- the served model, resolved capacity (max seqs / model len / batch tokens, KV cache size), device footprint and free/total memory, live running/waiting load, build version, whether HTTP drain is enabled, and (when set) the request limit and completed count -- so an orchestrator can route requests and make model-placement decisions. POST /v1/internal/drain begins the same graceful drain as a SIGTERM (stop accepting new requests, finish in-flight, then exit), returning 202 immediately and idempotent, so an instance can be retired over the wire without shell or signal access to the box. The drain endpoint is gated on a shared secret: it is disabled unless serve --drain-secret=<secret> is set, and callers must then present that secret verbatim in the X-Drain-Secret header (SIGTERM drains regardless).

--power-budget declares the GPU power management limit this node is expected to be running under, and gllm checks the device against it: 600W expects exactly that, min=450W sets a throughput floor, max=300W an electrical ceiling, min=450W,max=600W a range. gllm only ever reads the cap -- setting one is device-global, outlives the process, and needs privileges a serving process should not hold -- so the flag is an assertion about how the node was provisioned, not a request to change it. The two directions are not treated alike, because they are not the same problem. A cap below the budget costs throughput on this node and nothing else, so it warns and serves. A cap above it risks tripping an upstream breaker and taking down every host behind it, and since a cap bounds draw rather than causing it, admitting traffic is what turns a too-high cap into amps -- so it refuses to start. Append :warn or :fail to override both. --power-breach-action decides what a running instance does if the cap is raised past the ceiling while it serves: drain (the default: stop admitting, finish in-flight, exit), warn, or exit (stop without waiting). The power envelope is logged at startup and reported in the status payload (power_limit_watts, power_usage_watts, power_capped, throttle, power_budget, power_budget_met) and at /metrics (gllm_device_power_limit_watts, gllm_device_power_usage_watts, gllm_device_power_capped), whether or not a budget is declared. It needs the driver's NVML, which is loaded on demand: without it the power fields are simply absent.

Both endpoints answer from the moment the port binds, before the (minutes-long) model load finishes: during the load the status snapshot is {"state": "loading", "model_dir": ..., "elapsed_seconds": ...} so an orchestrator can tell "starting" from "dead", the drain endpoint aborts the load and exits (same secret gating), and every other endpoint answers 503 with a Retry-After hint instead of hanging. Once the load completes the status payload reports "state": "serving" (then "draining" after a drain is triggered).

Two probe endpoints follow the k8s liveness/readiness split. GET /livez is liveness: 200 whenever the process is up and answering HTTP, including during the load and a graceful drain -- point restart-on-failure supervision here, since restarting an instance for loading or draining defeats both. GET /readyz is readiness: 200 only while the instance should receive traffic -- 503 during the load and 503 again once a drain has been triggered, so a load balancer stops routing new requests while in-flight ones finish. GET /healthz is a deprecated, temporary alias of /readyz from before the split; migrate probes to /livez + /readyz. Each IP still probing the alias is warned about once in the server log, so the log names every prober left to migrate.

At --log-level debug each request logs an engine: request received line when it is admitted (paired with the INFO engine: request complete line by request_id), so tailing the log shows what is in flight. A request that does not complete gets an INFO closing line too, with its request_id and an err field: engine: request cancelled when the client went away, engine: request failed otherwise. Add serve --debug-with-prompt-preview to include the text of the first few prompt tokens as a prompt_preview field, with control tokens shown by name (e.g. <s>[INST]hi there) rather than dropped (off by default -- it logs prompt content); --debug-prompt-preview-size=N sets how many tokens. serve --debug-with-response-preview is the response-side counterpart: at --log-level debug it adds the text of the first few generated tokens to the engine: request complete line as a response_preview field (also off by default -- it logs response content); --debug-response-preview-size=N sets how many tokens. The cancelled and failed lines carry it as well, showing what was generated before the request stopped.

The engine: request complete line also says where the generated tokens went: content_bytes is the size of the answer, and a turn that reasons first adds reasoning_tokens, reasoning_bytes and reasoning_closed (whether the model ever closed its reasoning block); constrained=true marks a response_format request. A request that generated tokens and returned no content at all also gets a WARN, engine: request generated tokens but no content, whose cause names which case it was: a reasoning block still open at max_tokens or at the end of the turn, a stop string matching at the start of the answer, or output made of nothing but control tokens.

Either size can be -1 for no limit. A whole prompt or response is too much for a log line, so an unlimited preview is written to a file instead -- <request_id>.prompt.txt or <request_id>.response.txt under --debug-preview-dir -- and the log line carries its path as prompt_file or response_file in place of the preview field. Without --debug-preview-dir, gllm makes a fresh private directory for the run, gllm-previews-* in $TMPDIR (or /tmp), and logs its path when the first file is written. The files are owner-readable only and gllm never deletes them, so clear the directory yourself after a debugging session.

Thinking budgets

A chat request to a reasoning model can cap its reasoning block with max_thinking_tokens_soft (past it, the reasoning ends at the next sentence boundary) and max_thinking_tokens_hard (it never runs past it). gllm closes the block itself by writing the chat template's own \n</think>\n. The failure to watch for is a leak: the model does not register that its thinking is over and carries on reasoning in the answer, often ending with a </think> of its own -- a stray </think> or <tool_call> in content is the tell. Measured on Nemotron 3 Nano (12 problems with checkable answers, 4 seeds at the recommended temperature 1.0, 515 budget-cut turns):

  • Set a soft budget whenever you set a hard one. A hard cut always lands mid-thought: with only hard: 64, 5 of 48 turns leaked. Any soft budget below the hard one brought leaks under 1% (3 of 390), and to none at all with a hard budget of 128 or more.
  • Leave 32 to 64 tokens between them. Past the soft budget the model reached a sentence boundary within a median 11 tokens (90% within 27, the longest 62), so a gap of 32 lets the soft budget make most cuts and 64 leaves the hard one as a rare backstop. Gaps of 16, 32 and 64 leaked alike, so a wider gap costs nothing beyond the tokens the model may use.
  • Budget for accuracy, not just length. Unbudgeted, the model reasoned ~480 tokens and answered all 48 correctly; capped at 64 or 128 it answered 29-38, at 256 it answered 37-41. Tight budgets also shrink the answer less than the cap suggests: the model moves the work into a worked solution in the visible answer (a 64-token cap still averaged 200-265 tokens per turn, against 560 unbudgeted).

A reasonable starting point is soft = hard - 64 with hard no lower than about 256 for multi-step problems. These numbers are for one model; Granite 4.2's template closes reasoning the same way, but its budgets have not been measured.

To exercise the whole stack on a CPU-only machine -- including building a tiny servable checkpoint to curl -- see docs/end-to-end-cpu.md.

Chatting with a server

glchat is a reference terminal client, built by make build alongside the server (or on its own with make glchat):

glchat                          # find a local server and chat with it
glchat --endpoint host:9000     # talk to a specific one

With no flags it probes the port range gllm serve binds (:8000 plus the 20 it walks forward to when that is taken), asking each for /v1/internal/status. That confirms an instance is live and past its model load, and reports what it is serving -- so the model name never has to be typed. One instance connects straight through; several open a picker ordered by load. Replies stream token by token; ctrl+c interrupts a generation and ctrl+d (EOF) quits, as do /quit and /exit. pgup/pgdn (and shift+up/shift+down by the line) scroll the transcript; scrolling up during a turn stops the stream following, and returning to the bottom resumes it. Every turn closes with a dim rule saying how it ended -- end of turn, truncated at the token cap, interrupted, failed -- so a finished answer looks finished instead of merely having stopped, along with what the client observed of it:

-- end of turn  2.3s  142 tok  70.0 tok/s  ttft 310ms  gdqz7k2p9v4m ----

Those timings are the client's own, measured around the request and so including queue wait and the network -- what you actually waited. The trailing gdqz... is the server's request id for the turn, the string to grep its log with; it moves to a line of its own rather than being dropped when the rule runs out of room, and it is shown for a failed turn too. For the server's view of the same turn (per-batch CPU/GPU time), ask the server: it rides on the response as the vendor gllm object, and /metrics has the fleet-wide series.

A meter above the input says how much of the context window the conversation has taken and how much is left:

context [========---             ]  3.0k/8.2k  36%  prefill batch cap

= is what the next request will send -- the whole transcript through the model's chat template -- - the room reserved for the reply (--max-tokens), and the blanks what neither has claimed. The count comes from the server's own tokenizer over /tokenize, since a client cannot derive it; between turns it falls back on the last reply's usage and marks the number ~ until the exact one lands.

The window it measures against is the one the instance can actually serve, which is often not the context length the model advertises. A sequence longer than the whole KV cache can never be scheduled, and with --prefix-cache off (the default) every request prefills its entire prompt, so the conversation has to fit in one batch as well. A 131k-context model served with an 8k batch budget holds an 8k conversation, and the meter says so -- naming the limit that binds, as above -- rather than filling to 6% and then failing the request.

Input that starts with a / is a command to the client rather than a message to the model. /read <path> sends a file's contents as your turn -- the text itself, so the turns around it say what to do with it -- /width toggles the text column limit (below), /thinking sets reasoning and its budgets (below), /quit leaves, and /help lists what is available:

/read internal/engine/engine.go
/read ~/notes/today.md

The whole rest of the line is the path, so one with spaces in it needs no quoting, and a leading ~ is expanded here rather than by a shell that never saw the line. What to do with the file goes in the message after it.

/thinking changes what every later turn asks of a reasoning model. on, off or default (the chat template decides, as without --think) says whether it reasons; soft=N and hard=N set max_thinking_tokens_soft and max_thinking_tokens_hard, and =off clears one. Words combine, words left out keep their value, and bare /thinking reports the setting:

/thinking on soft=512 hard=1024
/thinking hard=off

Typing a / opens a list of what matches under the input, narrowing as you go:

/re
  /read <path>  send a file's contents as your message
  start with a space to send it as a message -- the space is stripped

The parse stays out of the way of ordinary typing: only a single line whose first word is a bare name (/read, /help) is a command, so /usr/lib/libc.so is missing, a pasted // TODO, and anything spanning more than one line are sent as typed -- and none of them opens the list. A command-shaped word that is not a command is reported instead of being sent, so a typo does not quietly spend a turn. A message that really does have to open with /read is sent by putting a space in front of it, which is stripped on the way out; the list vanishing as you type that space is the confirmation that the line is now an ordinary message.

On a terminal wider than 80 columns the transcript's text and the input wrap at 80, since prose stretched across a wide screen is hard to read. --width N sets the limit and --width 0 lets the text fill the terminal; while chatting, /width toggles between the limit and the full width, and /width N sets a new limit.

gllm chat runs the same client. It is a wrapper that execs the glchat binary rather than a subcommand proper, so the server binary does not link the TUI's dependencies; it is only offered when that binary is present next to gllm or on PATH.