feat: discover upstream model names and rewrite the model field #4

Merged
rcsheets merged 1 commit from feat/upstream-model-discovery into main 2026-07-20 09:51:14 +00:00
Owner

The exposed model ID is a product name; the ID a backend answers to is
whatever it was loaded with, frequently a build artifact. A client that
reads GET /v1/models and echoes back what it saw gets rejected:

model "mistral-small-4" does not exist;
this server serves "Mistral-Small-4-119B-2603-NVFP4"

Reconcile the two by asking the backend what it calls itself. Each raw
backend gets a supervisor goroutine that probes its /v1/models, retrying
indefinitely (1s backing off to 30s) until it sees exactly one model, then
backing off to a 5m check that only watches for the name to change. A
weights rebuild renames the model under a running SLP without a restart.

Discovery never gates startup: supervisors run alongside the server, so a
down or still-loading backend delays only its own rewrite rule. A route
with no name yet forwards unchanged, exactly as before this change, which
keeps a backend outage from becoming an SLP outage.

Ambiguity never guesses. Zero or several models means keeping the last
known good name (or plain pass-through), a rate-limited WARNING naming the
models seen, and slp_upstream_model_ambiguous at 1 -- the state where SLP
serves under a name it can no longer confirm must be alertable. Also
exports slp_discovery_probes_total, slp_upstream_model_changes_total, and
slp_upstream_model_info carrying the live mapping.

An explicit backend model: pins the name and skips probing, for multi-model
backends discovery cannot disambiguate or upstreams without a catalog.
Editions are never probed: a filter is launched --serve-as the edition ID,
so probing would only rediscover the ID SLP already exposes.

The rewrite touches only the model field. Every other field is carried
across as json.RawMessage, so numeric precision and string escaping are
preserved; only key order changes, which JSON does not define.

Amends README claims this makes untrue: "no rewriting", "never fans out to
ask the backends", and the unqualified "Stateless".

Adds the repo's first tests, covering the rewrite, discovery convergence,
drift, ambiguity, pinning, and shutdown.

Co-Authored-By: Claude Opus 4.8 noreply@anthropic.com

The exposed model ID is a product name; the ID a backend answers to is whatever it was loaded with, frequently a build artifact. A client that reads GET /v1/models and echoes back what it saw gets rejected: model "mistral-small-4" does not exist; this server serves "Mistral-Small-4-119B-2603-NVFP4" Reconcile the two by asking the backend what it calls itself. Each raw backend gets a supervisor goroutine that probes its /v1/models, retrying indefinitely (1s backing off to 30s) until it sees exactly one model, then backing off to a 5m check that only watches for the name to change. A weights rebuild renames the model under a running SLP without a restart. Discovery never gates startup: supervisors run alongside the server, so a down or still-loading backend delays only its own rewrite rule. A route with no name yet forwards unchanged, exactly as before this change, which keeps a backend outage from becoming an SLP outage. Ambiguity never guesses. Zero or several models means keeping the last known good name (or plain pass-through), a rate-limited WARNING naming the models seen, and slp_upstream_model_ambiguous at 1 -- the state where SLP serves under a name it can no longer confirm must be alertable. Also exports slp_discovery_probes_total, slp_upstream_model_changes_total, and slp_upstream_model_info carrying the live mapping. An explicit backend model: pins the name and skips probing, for multi-model backends discovery cannot disambiguate or upstreams without a catalog. Editions are never probed: a filter is launched --serve-as the edition ID, so probing would only rediscover the ID SLP already exposes. The rewrite touches only the model field. Every other field is carried across as json.RawMessage, so numeric precision and string escaping are preserved; only key order changes, which JSON does not define. Amends README claims this makes untrue: "no rewriting", "never fans out to ask the backends", and the unqualified "Stateless". Adds the repo's first tests, covering the rewrite, discovery convergence, drift, ambiguity, pinning, and shutdown. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
feat: discover upstream model names and rewrite the model field
All checks were successful
ci / check (pull_request) Successful in 1m25s
d0b5cf9183
The exposed model ID is a product name; the ID a backend answers to is
whatever it was loaded with, frequently a build artifact. A client that
reads GET /v1/models and echoes back what it saw gets rejected:

  model "mistral-small-4" does not exist;
  this server serves "Mistral-Small-4-119B-2603-NVFP4"

Reconcile the two by asking the backend what it calls itself. Each raw
backend gets a supervisor goroutine that probes its /v1/models, retrying
indefinitely (1s backing off to 30s) until it sees exactly one model, then
backing off to a 5m check that only watches for the name to change. A
weights rebuild renames the model under a running SLP without a restart.

Discovery never gates startup: supervisors run alongside the server, so a
down or still-loading backend delays only its own rewrite rule. A route
with no name yet forwards unchanged, exactly as before this change, which
keeps a backend outage from becoming an SLP outage.

Ambiguity never guesses. Zero or several models means keeping the last
known good name (or plain pass-through), a rate-limited WARNING naming the
models seen, and slp_upstream_model_ambiguous at 1 -- the state where SLP
serves under a name it can no longer confirm must be alertable. Also
exports slp_discovery_probes_total, slp_upstream_model_changes_total, and
slp_upstream_model_info carrying the live mapping.

An explicit backend model: pins the name and skips probing, for multi-model
backends discovery cannot disambiguate or upstreams without a catalog.
Editions are never probed: a filter is launched --serve-as the edition ID,
so probing would only rediscover the ID SLP already exposes.

The rewrite touches only the model field. Every other field is carried
across as json.RawMessage, so numeric precision and string escaping are
preserved; only key order changes, which JSON does not define.

Amends README claims this makes untrue: "no rewriting", "never fans out to
ask the backends", and the unqualified "Stateless".

Adds the repo's first tests, covering the rewrite, discovery convergence,
drift, ambiguity, pinning, and shutdown.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
rcsheets force-pushed feat/upstream-model-discovery from d0b5cf9183
All checks were successful
ci / check (pull_request) Successful in 1m25s
to a957f92a30
All checks were successful
ci / check (pull_request) Successful in 45s
2026-07-20 09:44:43 +00:00
Compare
rcsheets force-pushed feat/upstream-model-discovery from a957f92a30
All checks were successful
ci / check (pull_request) Successful in 45s
to 96a4fdf78f
All checks were successful
ci / check (pull_request) Successful in 45s
2026-07-20 09:50:19 +00:00
Compare
rcsheets deleted branch feat/upstream-model-discovery 2026-07-20 09:51:14 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brooktrails/slp!4
No description provided.