make deploy-runner restarts gitbay-runner unconditionally. A build the old
process had claimed is killed with it and stays running on the server: the new
process knows nothing about it, and the scheduler's reaper does not fail it for
a long time — build 1037 sat running for over 25 minutes after the restart at
03:10:41, with nothing executing on the host (podman ps empty, load 0.01),
until cancelled by hand. Every deploy today that overlapped a build did this.
Under podman there is a second half: the container is under the user slice, not
the service cgroup, so a service stop does not end it either. The drop-in's
ExecStopPost ends the pause process, which takes the container down, but that
is a side effect rather than a drain.
Three fixes, in increasing size:
- The reaper should key on the runner's presence, not a 45-minute timeout. A
runner reports claims and logs over one ssh session; when that session ends
without
runner done, the build is orphaned and can be failed within a minute. make deploy-runnershould wait for the runner to be idle before restarting it, or the runner should drain on SIGTERM: finish the current build, then exit. systemd'sTimeoutStopSecwould bound that.- Failing that, the runner should on start-up report any build it finds recorded as claimed by its account and still running, so the server can resolve it.
referenced in commit ec2d05573b by cmc: runner: retry reporting a build's outcome when the server is unreachable
2026-09-07 04:54 UTC