feat(kubernetes): ephemeral pools, one runner per job #57

Merged
rcsheets merged 1 commit from feat/ephemeral-kubernetes-pools into main 2026-09-28 01:22:16 +00:00
Owner

Part of #30: ephemeral mode for Kubernetes pools — one Forgejo ephemeral runner, in its own Pod, per job.

spec:
  labels: [ephemeral, dind]
  maxReplicas: 4          # required: caps concurrent jobs
  backend:
    type: kubernetes
    kubernetes:
      image: code.forgejo.org/forgejo/runner:13
      ephemeral: true
      privileged: true

How it works

  • Polling, not webhooks. Forgejo sends no event when a job is queued (only action_run_* when runs finish — see #56), so the operator polls GET …/actions/runners/jobs?labels=… in the pool's scope every 10s for waiting jobs whose runs-on labels are all among the pool's.
  • One Pod per job, pinned by handle. For each waiting job without a Pod: register an ephemeral runner, store its credentials, start <pool>-job-<id>-<suffix> running forgejo-runner one-job --handle <handle>. The handle names one attempt of one job, so the Pod runs exactly that job — or, if another runner took it first, nothing (exit 2, recorded as unclaimed). Overlapping pools or a stale cache cost an idle Pod, never a wrong job.
  • Capacity. maxReplicas caps concurrent job Pods; waiting jobs start oldest first; the Ready condition reports jobs waiting for capacity. Ready means "able to take jobs", so an idle pool with 0 Pods is Ready.
  • Pods complete. restartPolicy: Never; the DinD sidecar is a native sidecar (init container, restartPolicy: Always, TCP startup probe on 2375) so the Pod finishes when the runner exits and the runner waits for dockerd. Needs Kubernetes ≥ 1.29 (cluster is 1.34).
  • Afterwards. Every finished job → status.history entry (job, duration, outcome), runner deregistered, credentials deleted. Completed Pods deleted at once; unclaimed kept 10 min, failed 1 h (for logs). A job attempt is retried at most 3 times with growing backoff, then a JobRunnerGaveUp event — so a systematically broken image can't start a Pod every poll.
  • Robustness. Forgejo unreachable → nothing starts, nothing pending deregistration is discarded, Ready=False with the reason. Switching a pool between ephemeral and long-lived retires the other mode's runners; legacy-Deployment migration runs in both modes; pool deletion deregisters job runners.
  • Validation. Admission webhook rejects replicas/storage on ephemeral pools and requires maxReplicas ≥ 1.

Safety for existing pools

The Pod builder is now shared between long-lived and job Pods. Long-lived Pods are byte-identical to v0.9.0's (verified by snapshot before/after), and a new test pins their pod-spec-hash values — a change there would restart every runner on upgrade.

Testing

  • make test passes. New reconciler tests against the fake Forgejo: one Pod per waiting job (no duplicates), maxReplicas cap, completed cleanup + history, unclaimed retention + backoff, retry cap, Forgejo outage and recovery, switching modes both ways. Mutation checks: removing the duplicate guard or the backoff fails the tests.
  • e2e now commits a workflow (runs-on: e2e-ephemeral) to its Forgejo 16 instance and asserts a job Pod starts and an ephemeral runner is registered for the real queued job — the only check of the jobs endpoint's real label semantics. Unverified locally: that a contents-API commit triggers the workflow; CI is the first run.
  • golangci-lint not run locally.

Tunables (constants for now)

Poll interval 10s; retention 10 min (unclaimed) / 1 h (failed); max 3 tries per job attempt. Easy to promote to CRD fields if we want per-pool control.

🤖 Generated with Claude Code

Part of #30: **ephemeral mode for Kubernetes pools** — one Forgejo ephemeral runner, in its own Pod, per job. ```yaml spec: labels: [ephemeral, dind] maxReplicas: 4 # required: caps concurrent jobs backend: type: kubernetes kubernetes: image: code.forgejo.org/forgejo/runner:13 ephemeral: true privileged: true ``` ## How it works - **Polling, not webhooks.** Forgejo sends no event when a job is queued (only `action_run_*` when runs finish — see #56), so the operator polls `GET …/actions/runners/jobs?labels=…` in the pool's scope every 10s for waiting jobs whose `runs-on` labels are all among the pool's. - **One Pod per job, pinned by handle.** For each waiting job without a Pod: register an ephemeral runner, store its credentials, start `<pool>-job-<id>-<suffix>` running `forgejo-runner one-job --handle <handle>`. The handle names one attempt of one job, so the Pod runs exactly that job — or, if another runner took it first, nothing (exit 2, recorded as `unclaimed`). Overlapping pools or a stale cache cost an idle Pod, never a wrong job. - **Capacity.** `maxReplicas` caps concurrent job Pods; waiting jobs start oldest first; the Ready condition reports jobs waiting for capacity. Ready means "able to take jobs", so an idle pool with 0 Pods is Ready. - **Pods complete.** `restartPolicy: Never`; the DinD sidecar is a **native sidecar** (init container, `restartPolicy: Always`, TCP startup probe on 2375) so the Pod finishes when the runner exits and the runner waits for dockerd. Needs Kubernetes ≥ 1.29 (cluster is 1.34). - **Afterwards.** Every finished job → `status.history` entry (job, duration, outcome), runner deregistered, credentials deleted. Completed Pods deleted at once; `unclaimed` kept 10 min, failed 1 h (for logs). A job attempt is retried at most **3** times with growing backoff, then a `JobRunnerGaveUp` event — so a systematically broken image can't start a Pod every poll. - **Robustness.** Forgejo unreachable → nothing starts, nothing pending deregistration is discarded, Ready=False with the reason. Switching a pool between ephemeral and long-lived retires the other mode's runners; legacy-Deployment migration runs in both modes; pool deletion deregisters job runners. - **Validation.** Admission webhook rejects `replicas`/`storage` on ephemeral pools and requires `maxReplicas ≥ 1`. ## Safety for existing pools The Pod builder is now shared between long-lived and job Pods. Long-lived Pods are **byte-identical** to v0.9.0's (verified by snapshot before/after), and a new test pins their `pod-spec-hash` values — a change there would restart every runner on upgrade. ## Testing - `make test` passes. New reconciler tests against the fake Forgejo: one Pod per waiting job (no duplicates), `maxReplicas` cap, completed cleanup + history, unclaimed retention + backoff, retry cap, Forgejo outage and recovery, switching modes both ways. Mutation checks: removing the duplicate guard or the backoff fails the tests. - **e2e** now commits a workflow (`runs-on: e2e-ephemeral`) to its Forgejo 16 instance and asserts a job Pod starts and an ephemeral runner is registered for the real queued job — the only check of the jobs endpoint's real label semantics. Unverified locally: that a contents-API commit triggers the workflow; CI is the first run. - `golangci-lint` not run locally. ## Tunables (constants for now) Poll interval 10s; retention 10 min (unclaimed) / 1 h (failed); max 3 tries per job attempt. Easy to promote to CRD fields if we want per-pool control. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
feat(kubernetes): ephemeral pools, one runner per job
All checks were successful
CI (next Go) / next-go (tip) (pull_request) Successful in 3m1s
CI / ci (pull_request) Successful in 1m52s
E2E smoke test / e2e (pull_request) Successful in 3m28s
131ad9309f
Adds an ephemeral mode for Kubernetes-backend pools
(spec.backend.kubernetes.ephemeral: true), the next part of #30. Instead
of a fixed set of long-lived runners, the pool runs one Forgejo ephemeral
runner per job: nothing is reused between jobs, and an idle pool runs no
Pods.

Forgejo sends no webhook when a job is queued (only when a run finishes,
see #56), so the operator polls the runner jobs endpoint in the pool's
scope every 10s for waiting jobs whose runs-on labels are all among the
pool's. For each one without a Pod it registers an ephemeral runner and
starts a Pod running `forgejo-runner one-job --handle <handle>`. The
handle names one attempt of one job, so a Pod runs exactly the job it was
started for; if another runner takes the job first, the Pod runs nothing,
exits 2 and is recorded as unclaimed. Overlapping pools or a stale cache
therefore cost an idle Pod, never a wrong job. spec.maxReplicas caps
concurrent job Pods (required); waiting jobs are started oldest first and
the Ready condition reports any left waiting for capacity. Ready now means
"able to take jobs", so an idle pool with no Pods is Ready.

Job Pods never restart, and a DinD sidecar is a native sidecar (init
container with restartPolicy Always, Kubernetes 1.29+) with a TCP startup
probe, so the Pod completes when the runner exits and the runner waits
for dockerd.

Every finished job is recorded in status.history (job, duration,
outcome), its runner deregistered and its credentials deleted. Completed
Pods are deleted at once; unclaimed Pods are kept 10 minutes and failed
ones an hour, for their logs. A job attempt is retried at most 3 times,
with a growing wait, before a JobRunnerGaveUp event -- otherwise a
systematic failure (an image whose runner can't start) would start a Pod
every poll. While Forgejo is unreachable nothing starts and nothing still
to be deregistered is discarded.

Switching a pool between ephemeral and long-lived retires the other
mode's runners, a legacy Deployment is migrated in either mode, and pool
deletion deregisters job runners too. The admission webhook rejects
replicas or storage on an ephemeral pool and requires maxReplicas >= 1.

The Pod builder is shared between long-lived and job Pods; long-lived
Pods are byte-identical to v0.9.0's, and a new test pins their spec
hashes, since a change there restarts every runner on upgrade.

The e2e smoke test now commits a workflow to its Forgejo 16 instance and
checks that the ephemeral pool starts a job Pod and registers an
ephemeral runner for the real queued job. Adds a sample
(config/samples/ephemeral_dind.yaml) and a README section.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator

Automated review by pr-reviewer v0.52.3 | Safety Check | Mistral Small | tracking id r-b9bf1d-218e49
This is an AI-generated review and may contain mistakes.

Status: ❌ Failed


Review failed. Tracking id r-b9bf1d-218e49 — see logs for details.

Comment @pr-reviewer-bot retry to try again.

<!-- pr-reviewer:review --> *Automated review by [pr-reviewer](https://git.brooktrails.org/brooktrails/pr-reviewer) v0.52.3 | Safety Check | Mistral Small | tracking id `r-b9bf1d-218e49`* *This is an AI-generated review and may contain mistakes.* **Status:** ❌ Failed --- Review failed. Tracking id `r-b9bf1d-218e49` — see logs for details. Comment `@pr-reviewer-bot retry` to try again.
rcsheets deleted branch feat/ephemeral-kubernetes-pools 2026-09-28 01:22:16 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brooktrails/forgejo-runner-operator!57
No description provided.