feat(proxy): gate routes on the loaded model with expect: #15
Loading…
Reference in a new issue
No description provided.
Delete branch "feat/expect-gating"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Trying models on one GPU means one backend URL serving whichever model was
loaded last. Declared as one backend per model, discovery gets this wrong
silently: every backend on the URL probes the same /v1/models, adopts the one
loaded name, and a request for qwen-32b is rewritten to the loaded Mistral and
comes back 200.
expect: is a case-sensitive glob (path.Match) naming the upstream model a
backend is for. A gated route is probed like any discovered one but serves
only while exactly one reported model matches, rewriting to that name;
otherwise it answers 503 "not currently loaded" with Retry-After, before
anything reaches the backend, and without naming what is loaded. Editions of
a gated backend are gated with it.
Probes decide whether a route serves, never what it is: the table is still
exactly what config compiled, nothing is added, removed, repointed, or
substituted. A probe error keeps the last known state, which is safe because
the rewrite uses the exact last-seen name -- a backend swapped meanwhile
rejects it rather than answering as the wrong model.
Swaps are noticed quickly: gated routes monitor every 15s, a refusal or an
upstream 404 nudges the supervisor to probe early (never blocking the
request, at most one probe a second per route), and a route that sees a swap
nudges the other gated routes on its URL, since one of them was just loaded.
Without expect nothing changes. Ungated backends sharing a URL are allowed --
that is how you spell "whatever is loaded" -- with a startup warning;
expect together with model: is a startup error.
GET /v1/models still lists every id. Clients that send X-SLP-Catalog: v1 get
a namespaced "slp": {"loaded": ...} per entry and the header echoed; an
unknown version gets the plain catalog and a fixed X-SLP-Catalog-Warning that
never reflects the request. Every catalog response carries Vary. ?debug gains
expect, loaded, and an unloaded state; metrics gain slp_upstream_model_loaded,
probe result "unloaded", and request status "not_loaded".
Co-Authored-By: Claude Opus 5.5 noreply@anthropic.com
Trying models on one GPU means one backend URL serving whichever model was loaded last. Declared as one backend per model, discovery gets this wrong silently: every backend on the URL probes the same /v1/models, adopts the one loaded name, and a request for qwen-32b is rewritten to the loaded Mistral and comes back 200. expect: is a case-sensitive glob (path.Match) naming the upstream model a backend is for. A gated route is probed like any discovered one but serves only while exactly one reported model matches, rewriting to that name; otherwise it answers 503 "not currently loaded" with Retry-After, before anything reaches the backend, and without naming what is loaded. Editions of a gated backend are gated with it. Probes decide whether a route serves, never what it is: the table is still exactly what config compiled, nothing is added, removed, repointed, or substituted. A probe error keeps the last known state, which is safe because the rewrite uses the exact last-seen name -- a backend swapped meanwhile rejects it rather than answering as the wrong model. Swaps are noticed quickly: gated routes monitor every 15s, a refusal or an upstream 404 nudges the supervisor to probe early (never blocking the request, at most one probe a second per route), and a route that sees a swap nudges the other gated routes on its URL, since one of them was just loaded. Without expect nothing changes. Ungated backends sharing a URL are allowed -- that is how you spell "whatever is loaded" -- with a startup warning; expect together with model: is a startup error. GET /v1/models still lists every id. Clients that send X-SLP-Catalog: v1 get a namespaced "slp": {"loaded": ...} per entry and the header echoed; an unknown version gets the plain catalog and a fixed X-SLP-Catalog-Warning that never reflects the request. Every catalog response carries Vary. ?debug gains expect, loaded, and an unloaded state; metrics gain slp_upstream_model_loaded, probe result "unloaded", and request status "not_loaded". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>Automated review by pr-reviewer v0.52.3 | Safety Check | Ministral 3 Instruct | tracking id
r-bc23f2-08a865This is an AI-generated review and may contain mistakes.
Status: ❌ Failed
This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id
r-bc23f2-08a865.Comment
@pr-reviewer-bot retryto try again.