No description
  • Go 89.6%
  • Shell 6.4%
  • Makefile 2.9%
  • HCL 0.9%
  • Dockerfile 0.2%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-09-28 01:29:59 +00:00
.forgejo/workflows build: move the toolchain to Go 1.27 2026-09-27 12:25:16 -07:00
api/v1alpha1 feat(kubernetes): ephemeral pools, one runner per job 2026-09-27 18:12:36 -07:00
cmd feat: register runners through the Forgejo runner API, one Pod per runner 2026-09-27 12:38:57 -07:00
config feat(kubernetes): ephemeral pools, one runner per job 2026-09-27 18:12:36 -07:00
hack feat(kubernetes): ephemeral pools, one runner per job 2026-09-27 18:12:36 -07:00
internal feat(kubernetes): ephemeral pools, one runner per job 2026-09-27 18:12:36 -07:00
proto/provisioner/v1 feat: register runners through the Forgejo runner API, one Pod per runner 2026-09-27 12:38:57 -07:00
runner-images feat: register runners through the Forgejo runner API, one Pod per runner 2026-09-27 12:38:57 -07:00
.gitignore feat: scaffold Forgejo Runner Operator 2026-02-14 03:51:22 -08:00
CHANGELOG.md docs(changelog): update for v0.10.0 [skip ci] 2026-09-28 01:29:59 +00:00
CLAUDE.md build: move the toolchain to Go 1.27 2026-09-27 12:25:16 -07:00
Containerfile build: move the toolchain to Go 1.27 2026-09-27 12:25:16 -07:00
Containerfile.provisioner build: move the toolchain to Go 1.27 2026-09-27 12:25:16 -07:00
go.mod build: move the toolchain to Go 1.27 2026-09-27 12:25:16 -07:00
go.sum chore: update x/net to 0.57.0 2026-07-15 19:42:02 -07:00
Makefile feat: register runners through the Forgejo runner API, one Pod per runner 2026-09-27 12:38:57 -07:00
README.md feat(kubernetes): ephemeral pools, one runner per job 2026-09-27 18:12:36 -07:00
renovate.json feat: add CI/CD pipelines, e2e smoke test, Go 1.26 toolchain, and Renovate 2026-03-20 13:33:27 +00:00

Forgejo Runner Operator

A Kubernetes operator that manages pools of Forgejo Actions runners, supporting both in-cluster Pods (with Docker-in-Docker) and off-cluster ephemeral VMs.

Core Principle: Ephemerality Without Invisibility

Every disposable resource — a runner pod, a VM, a registration, a job execution — is also an instrumentation point. The shorter something lives, the more important it is to capture what it did while it existed.

  • Structured event log: every lifecycle transition emits a structured event to a durable store.
  • Metrics at every boundary: registration latency, job queue wait time, execution duration, VM boot time.
  • Status on the CRD: RunnerPool.status carries a full accounting including recently-destroyed runners.
  • Log forwarding from ephemeral VMs: logs ship to a collector before the runner accepts work.

Architecture

The operator consists of three components:

  1. Runner Controller (in-cluster) — reconciles RunnerPool CRDs, manages runner Pods, each registered as its own Forgejo runner (kubernetes backend) or calls the provisioner (remote backend).
  2. Webhook Receiver (in-cluster, same binary) — receives Forgejo workflow_job webhooks for remote-backend scale-up. Forgejo does not currently send workflow_job events, so this path is inert; see #56. Kubernetes ephemeral pools poll Forgejo for waiting jobs instead.
  3. Provisioner Agent (on VM host, separate binary) — manages ephemeral VM lifecycle via a pluggable hypervisor driver (Firecracker recommended).

Quick Start

Prerequisites

  • Go 1.27+
  • A Kubernetes cluster with kubectl configured
  • A Forgejo instance with Actions enabled

Build

# Resolve dependencies
go mod tidy

# Generate CRD manifests and deepcopy methods
make generate

# Build both binaries
make build

# Build container images
make podman-build

Install CRDs

make install

Deploy the operator

# Edit config/manager/manager.yaml to set your image
make deploy

Create a RunnerPool

# Create the namespace and token secret
kubectl create namespace forgejo-runners-trusted
kubectl create secret generic forgejo-admin-token \
  --namespace=forgejo-runners-trusted \
  --from-literal=token=<YOUR_FORGEJO_API_TOKEN>

# Apply the sample RunnerPool
kubectl apply -f config/samples/trusted_dind.yaml
# or a repo-enrolled trusted pool:
kubectl apply -f config/samples/trusted_repo_forgejo_runner_operator.yaml

For trusted runners, enroll specific projects by setting spec.forgejo.scope:

  • repository + spec.forgejo.repository: owner/repo for a single project
  • organization + spec.forgejo.organization: org for a whole org

This lets you run multiple trusted pools (for example, one per enrolled project) without exposing those runners to unrelated repositories.

Ephemeral pools

A Kubernetes pool can run one runner per job instead of a fixed set of long-lived runners: set spec.backend.kubernetes.ephemeral: true and cap concurrency with spec.maxReplicas (spec.replicas and storage are not allowed).

kubectl apply -f config/samples/ephemeral_dind.yaml

The operator polls Forgejo every 10s for waiting jobs whose runs-on labels are all among the pool's labels. For each, it registers a Forgejo ephemeral runner and starts a Pod running forgejo-runner one-job --handle <job>, which takes exactly that job and exits. An idle pool runs no Pods. Every job gets a fresh runner, workspace and (for privileged pools) DinD daemon.

Each finished job is recorded in status.history (job, duration, outcome) and its Pod, runner and credentials are removed. A Pod whose job was taken by another runner first (unclaimed) is kept for 10 minutes, and one that failed for an hour, so their logs can be read; a job whose runners keep failing is not retried more than three times until those expire.

Observe

# Watch the pool status
kubectl get runnerpools -A

# Detailed status including runner history
kubectl describe runnerpool trusted-dind -n forgejo-runners-trusted

# Prometheus metrics
curl http://localhost:8443/metrics | grep forgejo_runner_operator

CRD: RunnerPool

A single CRD covers both runner topologies. The backend field selects the execution strategy. Use spec.forgejo.scope (global, organization, or repository) to control which Forgejo namespace can request runners from the pool.

See config/samples/ for example manifests.

Project Layout

cmd/
  controller/       Operator entrypoint
  provisioner/      Provisioner agent entrypoint
api/v1alpha1/       CRD type definitions
internal/
  controller/       Reconciliation loop
  backend/          Execution strategy implementations
  forgejo/          Forgejo API client
  webhook/          Webhook receiver
  metrics/          Prometheus metrics
  audit/            Append-only audit logger
  provisioner/
    server/         gRPC server
    driver/         Hypervisor drivers (Firecracker, etc.)
    events/         Event ring buffer
proto/              Protobuf definitions
config/
  crd/              CRD manifests
  rbac/             RBAC resources
  manager/          Controller deployment
  samples/          Example RunnerPool manifests

Development Phases

  • Phase 0: Self-registering init container (in infra repo, no operator needed)
  • Phase 1: Operator scaffolding + in-cluster Kubernetes backend
  • Phase 2: Webhook receiver + scale-to-zero autoscaling
  • Phase 3: Provisioner agent + remote ephemeral VM backend
  • Phase 4: mTLS hardening, audit logging, Grafana dashboards

Changelog

CHANGELOG.md is intentionally machine-written and human-read: each release entry is generated from that release's commit range and committed automatically by CI — it is not hand-edited, and there is no "Unreleased" section to keep up to date. The entries are plain-language summaries for people running the operator; if you want commit-level detail, use git log or a release's compare link instead.

Because entries are generated from commits, the way to influence what shows up in the changelog is to write a good commit message — the body is summarized along with the subject, so a detailed message yields a richer entry.

License

Apache License 2.0