feat(proxy): gate routes on the loaded model with expect: #15

Merged
rcsheets merged 1 commit from feat/expect-gating into main 2026-09-29 21:05:17 +00:00
Owner

Trying models on one GPU means one backend URL serving whichever model was
loaded last. Declared as one backend per model, discovery gets this wrong
silently: every backend on the URL probes the same /v1/models, adopts the one
loaded name, and a request for qwen-32b is rewritten to the loaded Mistral and
comes back 200.

expect: is a case-sensitive glob (path.Match) naming the upstream model a
backend is for. A gated route is probed like any discovered one but serves
only while exactly one reported model matches, rewriting to that name;
otherwise it answers 503 "not currently loaded" with Retry-After, before
anything reaches the backend, and without naming what is loaded. Editions of
a gated backend are gated with it.

Probes decide whether a route serves, never what it is: the table is still
exactly what config compiled, nothing is added, removed, repointed, or
substituted. A probe error keeps the last known state, which is safe because
the rewrite uses the exact last-seen name -- a backend swapped meanwhile
rejects it rather than answering as the wrong model.

Swaps are noticed quickly: gated routes monitor every 15s, a refusal or an
upstream 404 nudges the supervisor to probe early (never blocking the
request, at most one probe a second per route), and a route that sees a swap
nudges the other gated routes on its URL, since one of them was just loaded.

Without expect nothing changes. Ungated backends sharing a URL are allowed --
that is how you spell "whatever is loaded" -- with a startup warning;
expect together with model: is a startup error.

GET /v1/models still lists every id. Clients that send X-SLP-Catalog: v1 get
a namespaced "slp": {"loaded": ...} per entry and the header echoed; an
unknown version gets the plain catalog and a fixed X-SLP-Catalog-Warning that
never reflects the request. Every catalog response carries Vary. ?debug gains
expect, loaded, and an unloaded state; metrics gain slp_upstream_model_loaded,
probe result "unloaded", and request status "not_loaded".

Co-Authored-By: Claude Opus 5.5 noreply@anthropic.com

Trying models on one GPU means one backend URL serving whichever model was loaded last. Declared as one backend per model, discovery gets this wrong silently: every backend on the URL probes the same /v1/models, adopts the one loaded name, and a request for qwen-32b is rewritten to the loaded Mistral and comes back 200. expect: is a case-sensitive glob (path.Match) naming the upstream model a backend is for. A gated route is probed like any discovered one but serves only while exactly one reported model matches, rewriting to that name; otherwise it answers 503 "not currently loaded" with Retry-After, before anything reaches the backend, and without naming what is loaded. Editions of a gated backend are gated with it. Probes decide whether a route serves, never what it is: the table is still exactly what config compiled, nothing is added, removed, repointed, or substituted. A probe error keeps the last known state, which is safe because the rewrite uses the exact last-seen name -- a backend swapped meanwhile rejects it rather than answering as the wrong model. Swaps are noticed quickly: gated routes monitor every 15s, a refusal or an upstream 404 nudges the supervisor to probe early (never blocking the request, at most one probe a second per route), and a route that sees a swap nudges the other gated routes on its URL, since one of them was just loaded. Without expect nothing changes. Ungated backends sharing a URL are allowed -- that is how you spell "whatever is loaded" -- with a startup warning; expect together with model: is a startup error. GET /v1/models still lists every id. Clients that send X-SLP-Catalog: v1 get a namespaced "slp": {"loaded": ...} per entry and the header echoed; an unknown version gets the plain catalog and a fixed X-SLP-Catalog-Warning that never reflects the request. Every catalog response carries Vary. ?debug gains expect, loaded, and an unloaded state; metrics gain slp_upstream_model_loaded, probe result "unloaded", and request status "not_loaded". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
feat(proxy): gate routes on the loaded model with expect:
All checks were successful
ci / check (pull_request) Successful in 45s
6da06d86c5
Trying models on one GPU means one backend URL serving whichever model was
loaded last. Declared as one backend per model, discovery gets this wrong
silently: every backend on the URL probes the same /v1/models, adopts the one
loaded name, and a request for qwen-32b is rewritten to the loaded Mistral and
comes back 200.

expect: is a case-sensitive glob (path.Match) naming the upstream model a
backend is for. A gated route is probed like any discovered one but serves
only while exactly one reported model matches, rewriting to that name;
otherwise it answers 503 "not currently loaded" with Retry-After, before
anything reaches the backend, and without naming what is loaded. Editions of
a gated backend are gated with it.

Probes decide whether a route serves, never what it is: the table is still
exactly what config compiled, nothing is added, removed, repointed, or
substituted. A probe error keeps the last known state, which is safe because
the rewrite uses the exact last-seen name -- a backend swapped meanwhile
rejects it rather than answering as the wrong model.

Swaps are noticed quickly: gated routes monitor every 15s, a refusal or an
upstream 404 nudges the supervisor to probe early (never blocking the
request, at most one probe a second per route), and a route that sees a swap
nudges the other gated routes on its URL, since one of them was just loaded.

Without expect nothing changes. Ungated backends sharing a URL are allowed --
that is how you spell "whatever is loaded" -- with a startup warning;
expect together with model: is a startup error.

GET /v1/models still lists every id. Clients that send X-SLP-Catalog: v1 get
a namespaced "slp": {"loaded": ...} per entry and the header echoed; an
unknown version gets the plain catalog and a fixed X-SLP-Catalog-Warning that
never reflects the request. Every catalog response carries Vary. ?debug gains
expect, loaded, and an unloaded state; metrics gain slp_upstream_model_loaded,
probe result "unloaded", and request status "not_loaded".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator

Automated review by pr-reviewer v0.52.3 | Safety Check | Ministral 3 Instruct | tracking id r-bc23f2-08a865
This is an AI-generated review and may contain mistakes.

Status: ❌ Failed


This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id r-bc23f2-08a865.

Comment @pr-reviewer-bot retry to try again.

<!-- pr-reviewer:review --> *Automated review by [pr-reviewer](https://git.brooktrails.org/brooktrails/pr-reviewer) v0.52.3 | Safety Check | Ministral 3 Instruct | tracking id `r-bc23f2-08a865`* *This is an AI-generated review and may contain mistakes.* **Status:** ❌ Failed --- This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id `r-bc23f2-08a865`. Comment `@pr-reviewer-bot retry` to try again.
rcsheets deleted branch feat/expect-gating 2026-09-29 21:05:18 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brooktrails/slp!15
No description provided.