feat(kubernetes): ephemeral pools, one runner per job #57
Loading…
Reference in a new issue
No description provided.
Delete branch "feat/ephemeral-kubernetes-pools"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Part of #30: ephemeral mode for Kubernetes pools — one Forgejo ephemeral runner, in its own Pod, per job.
How it works
action_run_*when runs finish — see #56), so the operator pollsGET …/actions/runners/jobs?labels=…in the pool's scope every 10s for waiting jobs whoseruns-onlabels are all among the pool's.<pool>-job-<id>-<suffix>runningforgejo-runner one-job --handle <handle>. The handle names one attempt of one job, so the Pod runs exactly that job — or, if another runner took it first, nothing (exit 2, recorded asunclaimed). Overlapping pools or a stale cache cost an idle Pod, never a wrong job.maxReplicascaps concurrent job Pods; waiting jobs start oldest first; the Ready condition reports jobs waiting for capacity. Ready means "able to take jobs", so an idle pool with 0 Pods is Ready.restartPolicy: Never; the DinD sidecar is a native sidecar (init container,restartPolicy: Always, TCP startup probe on 2375) so the Pod finishes when the runner exits and the runner waits for dockerd. Needs Kubernetes ≥ 1.29 (cluster is 1.34).status.historyentry (job, duration, outcome), runner deregistered, credentials deleted. Completed Pods deleted at once;unclaimedkept 10 min, failed 1 h (for logs). A job attempt is retried at most 3 times with growing backoff, then aJobRunnerGaveUpevent — so a systematically broken image can't start a Pod every poll.replicas/storageon ephemeral pools and requiresmaxReplicas ≥ 1.Safety for existing pools
The Pod builder is now shared between long-lived and job Pods. Long-lived Pods are byte-identical to v0.9.0's (verified by snapshot before/after), and a new test pins their
pod-spec-hashvalues — a change there would restart every runner on upgrade.Testing
make testpasses. New reconciler tests against the fake Forgejo: one Pod per waiting job (no duplicates),maxReplicascap, completed cleanup + history, unclaimed retention + backoff, retry cap, Forgejo outage and recovery, switching modes both ways. Mutation checks: removing the duplicate guard or the backoff fails the tests.runs-on: e2e-ephemeral) to its Forgejo 16 instance and asserts a job Pod starts and an ephemeral runner is registered for the real queued job — the only check of the jobs endpoint's real label semantics. Unverified locally: that a contents-API commit triggers the workflow; CI is the first run.golangci-lintnot run locally.Tunables (constants for now)
Poll interval 10s; retention 10 min (unclaimed) / 1 h (failed); max 3 tries per job attempt. Easy to promote to CRD fields if we want per-pool control.
🤖 Generated with Claude Code
Automated review by pr-reviewer v0.52.3 | Safety Check | Mistral Small | tracking id
r-b9bf1d-218e49This is an AI-generated review and may contain mistakes.
Status: ❌ Failed
Review failed. Tracking id
r-b9bf1d-218e49— see logs for details.Comment
@pr-reviewer-bot retryto try again.