feat: register runners through the Forgejo runner API, one Pod per runner #55

Merged
rcsheets merged 2 commits from feat/forgejo-runner-registration-api into main 2026-09-27 19:50:46 +00:00
Owner

Part of #30 — the "plumbing" half: migrate to Forgejo's runner registration API and bring the operator up to current Forgejo (16.x) and forgejo-runner (13.x). Ephemeral pools and OIDC are follow-ups.

Why

  • Deregistration was silently broken. The client listed/deleted runners under /api/v1/admin/runners, which no longer exists (moved to /admin/actions/runners), and used admin paths even for org/repo-scoped pools — so deleting a pool left its runners registered.
  • The registration-token endpoints are deprecated; Forgejo now registers a runner directly (POST …/actions/runners → {id, uuid, token}).
  • Since Forgejo 15, 404 also means "not visible to this token" (no more 403).

What changes

Kubernetes backend: one Pod per runner. The operator now owns each runner's identity, so a Deployment (whose replicas share one Secret) can't express it. Slot n gets:

  • <pool>-n-credentials — Secret with uuid/token; the durable record of the registration
  • <pool>-n — Pod running forgejo-runner daemon --url --uuid --token-url file:… (token never in the Pod spec)
  • <pool>-n-data — PVC, when storage is configured

The register init container, .runner file and its chmod workaround are gone, and so is the leak of one stale Forgejo runner per restart of an emptyDir pod.

The controller replaces a Pod when its spec/registration changes or it terminates for good; re-registers a runner deleted on the Forgejo side (confirmed by a direct 404 first); on scale-down deletes the Pod, then deregisters once it's gone; deregisters everything on pool deletion. A Forgejo outage still doesn't block spec changes (#54 behaviour kept). Stalled now names pods not ready after 10 min and why (e.g. pool-0 (runner: ImagePullBackOff)).

Remote backend. Proto carries runner uuid/token/Forgejo ID instead of a registration token (field 2 reserved). Ephemeral VMs register a Forgejo ephemeral runner and run one-job --wait (the runner refuses ephemeral runners in daemon mode). Token written to a root-only file; provisioner validates credential formats since they land in a shell script. Scale-down deregisters the VM's runner.

Also: runner images/samples → forgejo-runner 13 (minimum v12.8.0 for --token-url, so existing 12-latest pools keep working); e2e runs Forgejo 16 and asserts the runner is really registered via the API; status.runners[].registeredAt optional; RBAC gains pods, drops Deployment writes; manager caches only its own Pods. Separate commit: Go toolchain → 1.27.

⚠️ Rollout / migration

On first reconcile, a pool still on the legacy Deployment (once Forgejo is reachable):

  1. deregisters every runner in its scope named exactly <pool> (the old self-registered ones, including leaked stale ones),
  2. deletes the <pool> Deployment, <pool>-registration Secret, and <pool>-data PVC (held only .runner + scratch),
  3. registers <pool>-0… and starts per-runner Pods.

Caveat: two same-named pools in different namespaces sharing a Forgejo scope would deregister each other's legacy runners. Running jobs on a pool's old runner are interrupted when its Deployment is deleted (same as a Recreate rollout).

Testing

  • make test (generate, fmt, vet, tests) passes; new tests drive the reconciler against an in-memory fake Forgejo (per-replica registration, image change, Forgejo failing, scale-down ordering, lost registration, legacy migration, pool deletion, stall detection) plus client, pod-builder and cloud-init tests.
  • Not run locally: e2e (needs kind) and golangci-lint — CI is the first run of both.
  • The remote backend remains scaffolding (insecure gRPC) and hasn't been exercised end to end.

Not touched, noticed in passing: host-mode jobs can read the runner token file (same exposure as .runner before); CI workflows install protoc plugins @latest, contrary to CLAUDE.md.

🤖 Generated with Claude Code

Part of #30 — the "plumbing" half: migrate to Forgejo's runner registration API and bring the operator up to current Forgejo (16.x) and forgejo-runner (13.x). Ephemeral pools and OIDC are follow-ups. ## Why - **Deregistration was silently broken.** The client listed/deleted runners under `/api/v1/admin/runners`, which no longer exists (moved to `/admin/actions/runners`), and used admin paths even for org/repo-scoped pools — so deleting a pool left its runners registered. - The registration-token endpoints are **deprecated**; Forgejo now registers a runner directly (`POST …/actions/runners` → `{id, uuid, token}`). - Since Forgejo 15, 404 also means "not visible to this token" (no more 403). ## What changes **Kubernetes backend: one Pod per runner.** The operator now owns each runner's identity, so a Deployment (whose replicas share one Secret) can't express it. Slot *n* gets: - `<pool>-n-credentials` — Secret with uuid/token; the durable record of the registration - `<pool>-n` — Pod running `forgejo-runner daemon --url --uuid --token-url file:…` (token never in the Pod spec) - `<pool>-n-data` — PVC, when storage is configured The `register` init container, `.runner` file and its chmod workaround are gone, and so is the leak of one stale Forgejo runner per restart of an emptyDir pod. The controller replaces a Pod when its spec/registration changes or it terminates for good; re-registers a runner deleted on the Forgejo side (confirmed by a direct 404 first); on scale-down deletes the Pod, then deregisters once it's gone; deregisters everything on pool deletion. A Forgejo outage still doesn't block spec changes (#54 behaviour kept). `Stalled` now names pods not ready after 10 min and why (e.g. `pool-0 (runner: ImagePullBackOff)`). **Remote backend.** Proto carries runner uuid/token/Forgejo ID instead of a registration token (field 2 reserved). Ephemeral VMs register a Forgejo ephemeral runner and run `one-job --wait` (the runner refuses ephemeral runners in daemon mode). Token written to a root-only file; provisioner validates credential formats since they land in a shell script. Scale-down deregisters the VM's runner. **Also:** runner images/samples → forgejo-runner 13 (minimum v12.8.0 for `--token-url`, so existing `12-latest` pools keep working); e2e runs Forgejo 16 and asserts the runner is really registered via the API; `status.runners[].registeredAt` optional; RBAC gains `pods`, drops Deployment writes; manager caches only its own Pods. Separate commit: Go toolchain → 1.27. ## ⚠️ Rollout / migration On first reconcile, a pool still on the legacy Deployment (once Forgejo is reachable): 1. deregisters every runner **in its scope named exactly `<pool>`** (the old self-registered ones, including leaked stale ones), 2. deletes the `<pool>` Deployment, `<pool>-registration` Secret, and `<pool>-data` PVC (held only `.runner` + scratch), 3. registers `<pool>-0…` and starts per-runner Pods. Caveat: two same-named pools in different namespaces sharing a Forgejo scope would deregister each other's legacy runners. Running jobs on a pool's old runner are interrupted when its Deployment is deleted (same as a Recreate rollout). ## Testing - `make test` (generate, fmt, vet, tests) passes; new tests drive the reconciler against an in-memory fake Forgejo (per-replica registration, image change, Forgejo failing, scale-down ordering, lost registration, legacy migration, pool deletion, stall detection) plus client, pod-builder and cloud-init tests. - **Not run locally:** e2e (needs kind) and `golangci-lint` — CI is the first run of both. - The remote backend remains scaffolding (insecure gRPC) and hasn't been exercised end to end. Not touched, noticed in passing: host-mode jobs can read the runner token file (same exposure as `.runner` before); CI workflows install protoc plugins `@latest`, contrary to CLAUDE.md. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Raise the go directive to 1.27 and move every place that builds with Go
onto 1.27 images: both Containerfiles (golang:1.27-alpine) and all four
CI workflows (golang:1.27-bookworm). The documented minimum in README
and CLAUDE.md said "Go 1.23+", long out of date; it now says 1.27+.

go mod tidy also merged the two direct-dependency require blocks into
one; no dependency versions changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
feat: register runners through the Forgejo runner API, one Pod per runner
All checks were successful
CI (next Go) / next-go (tip) (pull_request) Successful in 3m46s
CI / ci (pull_request) Successful in 1m52s
E2E smoke test / e2e (pull_request) Successful in 3m53s
8b68d53446
Forgejo deprecated the registration-token endpoints in favour of an HTTP
API that registers a runner directly (POST .../actions/runners) and returns
its id, uuid and token. The operator now uses it: it owns every runner's
identity from the start instead of letting pods register themselves. Part
of #30, targeting the current Forgejo (16.x) and forgejo-runner (13.x).

Fixes deregistration, which had silently stopped working: the client
listed and deleted runners under /api/v1/admin/runners, which no longer
exists (runners moved to /admin/actions/runners), and always used admin
paths even for org- and repo-scoped pools. The client is now scope-aware
for register, list (paginated), get and delete, and maps 404 to
ErrNotFound -- since Forgejo 15 that also means "not visible to this
token", as the API stopped answering 403.

Kubernetes backend: each replica is its own registered runner, so the
single Deployment (whose replicas can only share one Secret, i.e. one
identity) is replaced by operator-managed Pods. Slot n gets a credentials
Secret <pool>-n-credentials (the durable record of the registration), a
Pod <pool>-n running `forgejo-runner daemon --url --uuid --token-url`,
and with storage a PVC <pool>-n-data. The register init container, the
.runner file and its group-write workaround are gone, as is the leak of
one stale Forgejo runner per restart of an emptyDir-backed pod. The
controller replaces a Pod when its spec or registration changes or it
terminates for good, re-registers a runner deleted on the Forgejo side
(after confirming the 404 directly), and on scale-down deletes the Pod
first and deregisters once it is gone. Pool deletion deregisters every
runner. A Forgejo outage still does not block spec changes. The Stalled
condition now names runner pods not ready after 10 minutes and why.

Migration: a pool still on the legacy Deployment has every runner in its
scope named exactly <pool> deregistered, then its Deployment,
<pool>-registration Secret and <pool>-data PVC deleted. It waits for
Forgejo to be reachable, so the old runners serve jobs until then.

Remote backend: the provisioner proto carries the runner's uuid, token
and Forgejo ID instead of a registration token (field 2 reserved). An
ephemeral VM registers a Forgejo ephemeral runner and runs `one-job
--wait` (the runner refuses ephemeral runners in daemon mode); the token
goes to a root-only file, and the provisioner rejects credentials that
are not in the format Forgejo issues. Scale-down deregisters the VM's
runner.

Also: runner images and samples move to forgejo-runner 13 (v12.8.0 is
the minimum, for --token-url); the e2e smoke test runs Forgejo 16 and
checks that the runner is really registered; status.runners[].registeredAt
is optional, so not-yet-registered runners pass validation; RBAC gains
pods and drops deployment writes; and the manager caches only its own
Pods rather than every Pod in the cluster.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator

Automated review by pr-reviewer v0.52.3 | Safety Check | Mistral Small | tracking id r-b970fd-4d63d5
This is an AI-generated review and may contain mistakes.

Status: ❌ Failed


This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id r-b970fd-4d63d5.

Comment @pr-reviewer-bot retry to try again.

<!-- pr-reviewer:review --> *Automated review by [pr-reviewer](https://git.brooktrails.org/brooktrails/pr-reviewer) v0.52.3 | Safety Check | Mistral Small | tracking id `r-b970fd-4d63d5`* *This is an AI-generated review and may contain mistakes.* **Status:** ❌ Failed --- This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id `r-b970fd-4d63d5`. Comment `@pr-reviewer-bot retry` to try again.
rcsheets deleted branch feat/forgejo-runner-registration-api 2026-09-27 19:50:46 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brooktrails/forgejo-runner-operator!55
No description provided.