feat: register runners through the Forgejo runner API, one Pod per runner #55
Loading…
Reference in a new issue
No description provided.
Delete branch "feat/forgejo-runner-registration-api"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Part of #30 — the "plumbing" half: migrate to Forgejo's runner registration API and bring the operator up to current Forgejo (16.x) and forgejo-runner (13.x). Ephemeral pools and OIDC are follow-ups.
Why
/api/v1/admin/runners, which no longer exists (moved to/admin/actions/runners), and used admin paths even for org/repo-scoped pools — so deleting a pool left its runners registered.POST …/actions/runners→{id, uuid, token}).What changes
Kubernetes backend: one Pod per runner. The operator now owns each runner's identity, so a Deployment (whose replicas share one Secret) can't express it. Slot n gets:
<pool>-n-credentials— Secret with uuid/token; the durable record of the registration<pool>-n— Pod runningforgejo-runner daemon --url --uuid --token-url file:…(token never in the Pod spec)<pool>-n-data— PVC, when storage is configuredThe
registerinit container,.runnerfile and its chmod workaround are gone, and so is the leak of one stale Forgejo runner per restart of an emptyDir pod.The controller replaces a Pod when its spec/registration changes or it terminates for good; re-registers a runner deleted on the Forgejo side (confirmed by a direct 404 first); on scale-down deletes the Pod, then deregisters once it's gone; deregisters everything on pool deletion. A Forgejo outage still doesn't block spec changes (#54 behaviour kept).
Stallednow names pods not ready after 10 min and why (e.g.pool-0 (runner: ImagePullBackOff)).Remote backend. Proto carries runner uuid/token/Forgejo ID instead of a registration token (field 2 reserved). Ephemeral VMs register a Forgejo ephemeral runner and run
one-job --wait(the runner refuses ephemeral runners in daemon mode). Token written to a root-only file; provisioner validates credential formats since they land in a shell script. Scale-down deregisters the VM's runner.Also: runner images/samples → forgejo-runner 13 (minimum v12.8.0 for
--token-url, so existing12-latestpools keep working); e2e runs Forgejo 16 and asserts the runner is really registered via the API;status.runners[].registeredAtoptional; RBAC gainspods, drops Deployment writes; manager caches only its own Pods. Separate commit: Go toolchain → 1.27.⚠️ Rollout / migration
On first reconcile, a pool still on the legacy Deployment (once Forgejo is reachable):
<pool>(the old self-registered ones, including leaked stale ones),<pool>Deployment,<pool>-registrationSecret, and<pool>-dataPVC (held only.runner+ scratch),<pool>-0…and starts per-runner Pods.Caveat: two same-named pools in different namespaces sharing a Forgejo scope would deregister each other's legacy runners. Running jobs on a pool's old runner are interrupted when its Deployment is deleted (same as a Recreate rollout).
Testing
make test(generate, fmt, vet, tests) passes; new tests drive the reconciler against an in-memory fake Forgejo (per-replica registration, image change, Forgejo failing, scale-down ordering, lost registration, legacy migration, pool deletion, stall detection) plus client, pod-builder and cloud-init tests.golangci-lint— CI is the first run of both.Not touched, noticed in passing: host-mode jobs can read the runner token file (same exposure as
.runnerbefore); CI workflows install protoc plugins@latest, contrary to CLAUDE.md.🤖 Generated with Claude Code
Automated review by pr-reviewer v0.52.3 | Safety Check | Mistral Small | tracking id
r-b970fd-4d63d5This is an AI-generated review and may contain mistakes.
Status: ❌ Failed
This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id
r-b970fd-4d63d5.Comment
@pr-reviewer-bot retryto try again.