A write-up, not a feature. State what CI is good for, what it is not, and which gaps are worth closing versus documenting. Facts to start from:
- Every runner change so far cost an outage. The drain in #179 shipped broken (KillMode) and was caught by reading a journal by hand. No test exercises a real systemd stop.
- Dedupe (#177), path filters (#169, #171, #172) and the reaper are three mechanisms with their own edge cases; #176 was a case where they combined into a green build that never ran.
- Nothing measures it: no queue latency, no time-to-claim, no count of builds ended by the reaper rather than by report.
- One runner, one host, shared with the forge.
End with a decision: CICD is a feature of the forge. That sets the bar for everything else on the list.
Risks of widening the runner's scope, as of 2026-09-07
- The build home is one directory shared by every build on the runner,
mounted read-write into every container as HOME. A stranger's build
can poison the module cache or plant a
.gitconfigthat the forge's own build honours. #144 mounted the same shared home. Fixed under A below. - No network flag: a build gets podman's default network, so bay1 is free compute with a clean IP, and 127.0.0.1:22 is reachable from inside.
- No memory cap, on the forge's own host. Fixed under A below.
- Queue starvation: one runner, oldest-first, 45 minutes per build, no per-account cap on builds queued, schedules down to a minute.
- Rootless podman needed NoNewPrivileges, ProtectKernelTunables and RestrictSUIDSGID relaxed; an escape is on the forge's host.
Handled already: the runner's key and other workspaces are unreachable (the nightly canary proves it), fork heads build without secrets, images are provisioned never pulled, logs are capped at 2 MiB.
A, done regardless: CI stays scoped, and two bugs go
- A build home per repository, not one for the runner.
- A memory cap that leaves the forge alive.
- The Admin page and FAQ say that CI on gitbay.org builds only the repositories the operator names; self-hosters get the runner for their own.
B, if CI is a feature: a second machine and a tier per trust level
The runner already polls over SSH from anywhere, admin runners lists
several, and -repos scopes each, which is most of a tiered design:
- The runner on bay1 keeps building krz/gitbay and the canary.
- A second small VPS runs a runner with no scope,
--network noneor an egress allowlist, a memory cap, a lower CPU share, and nothing else on it. Disposable: a compromise costs a reinstall, not the forge. - Per-account limits server-side: builds queued at once, build minutes
per day, a floor on schedule intervals. Same shape and place as
admin user limits. - Claim order stops being oldest-first across the instance: each runner takes the oldest build in its scope, and an unscoped runner rotates across accounts so one cannot hold the queue.
Under any answer
- One table on the wiki that says, for every push shape, which builds dedupe, path filters and the reaper produce and what status the commit gets; a property test over push shapes asserting it. #176 was found by accident.
- A real stop test: the runner under a throwaway systemd unit, claim, SIGTERM, assert the build reports. The drain shipped broken because only the process was tested, never the unit.
- Three numbers on
admin runnersor the dashboard: queue-to-claim time, builds ended by the reaper rather than by report, queue depth.
referenced in commit 0caaaedf8f by cmc: runner, deploy, wiki: a build home per repository, and caps on the runner unit
2026-09-07 19:47 UTC