Restarting the runner orphans its in-flight build as running until the reaper #179

closed cmc opened this on 2026-09-07 03:37 UTC · milestone v1.15.0

Discussion

cmc 2026-09-07 03:37 UTC

make deploy-runner restarts gitbay-runner unconditionally. A build the old process had claimed is killed with it and stays running on the server: the new process knows nothing about it, and the scheduler's reaper does not fail it for a long time — build 1037 sat running for over 25 minutes after the restart at 03:10:41, with nothing executing on the host (podman ps empty, load 0.01), until cancelled by hand. Every deploy today that overlapped a build did this.

Under podman there is a second half: the container is under the user slice, not the service cgroup, so a service stop does not end it either. The drop-in's ExecStopPost ends the pause process, which takes the container down, but that is a side effect rather than a drain.

Three fixes, in increasing size:

  • The reaper should key on the runner's presence, not a 45-minute timeout. A runner reports claims and logs over one ssh session; when that session ends without runner done, the build is orphaned and can be failed within a minute.
  • make deploy-runner should wait for the runner to be idle before restarting it, or the runner should drain on SIGTERM: finish the current build, then exit. systemd's TimeoutStopSec would bound that.
  • Failing that, the runner should on start-up report any build it finds recorded as claimed by its account and still running, so the server can resolve it.

Ref #144, #152.

cmc 2026-09-07 03:39 UTC

A second variant, from the server side: make deploy restarted gitbayd at 03:38:09 while the runner was reporting build 1155 (ci/build on c4a6cdaa):

runner: reporting build 1155: exit status 255 (ssh: connect to host 127.0.0.1 port 22: Connection refused

The build had finished; only the report failed, and the runner did not retry — it moved on to the next claim and the server kept the build as running. Cancelled by hand as build 1043.

So the runner should retry runner done (and runner log) on a transient ssh failure, with a bound, rather than dropping a finished build's result on the floor. That is independent of the restart-drain half above and smaller.

referenced in commit ec2d05573b by cmc: runner: retry reporting a build's outcome when the server is unreachable

2026-09-07 04:54 UTC

referenced in commit 631ccdc137 by cmc: store, control, wiki: fail a build whose log stream ended with no outcome

2026-09-07 06:36 UTC

closed by commit 65f453dc9b by cmc: runner, deploy, wiki: drain on SIGTERM

2026-09-07 06:47 UTC
cmc 2026-09-07 06:52 UTC

The drain shipped in 65f453d does not reach production builds as deployed. On the first restart of the new runner (bay1, 06:50:04 UTC) the journal shows, in the same instant as draining: finishing builds in flight:

build 1174: log session ended: exit status 255
gitbay-build-1174[4175962]: Terminated

The unit runs with the default KillMode=control-group, so systemd signals every process in the cgroup — the ssh log session and the container included — not only the runner. Build 1174 still reported success because its steps had finished three seconds earlier. A step in flight would have been killed.

Fix is KillMode=mixed in the drop-in: SIGTERM to the main process only, the rest killed after it exits.

referenced in commit 5ffa892991 by cmc: deploy, wiki: KillMode=mixed so a stop reaches only the runner

2026-09-07 07:02 UTC

referenced in commit 7e7a919b3b by cmc: e2e, wiki: stop tests for the runner's drain

2026-09-08 06:55 UTC