gitbay-runner marks a build failed when its log stream breaks, even though the step itself succeeded. The log is truncated at the same instant, so the evidence of what happened is gone too.
The mechanism
run() in cmd/gitbay-runner/main.go opens one long-lived ssh … runner log <id> session and hands its stdin to every step as both stdout and stderr:
cmd.Stdout, cmd.Stderr = sink, sink
sink is not an *os.File, so os/exec runs a copy goroutine per step. cmd.Wait() returns the error from that goroutine when the process itself exited cleanly. If the log session has gone away, the copy hits EPIPE, Wait() returns non-nil, and the step is reported failed:
case err := <-done:
if err != nil {
fmt.Fprintf(sink, "step failed: %v\n", err) // also lost — same dead pipe
return false
}
So a green go test becomes a red build, and the diagnostic line explaining it is written to the pipe that just died.
Evidence
Build 148 (test @ fe8361c849) ran 05:19:08–05:23:44 UTC. gitbayd restarted at 05:20:44, inside that window, which drops every SSH session including the log stream. The build was recorded as a failure.
The same commit is fine. Running the same suite on bay1 as the ci-runner user:
ok gitbay.org/gitbay/e2e 272.656s
EXIT=0
The stored logs match the theory — the failure stops mid-stream and never reaches step failed::
build 148 (failure): 460 bytes, 6 lines, last line "ok gitbay.org/gitbay/cmd/gitbayd"
build 138 (success): 1569 bytes, 29 lines, every package present
Builds 127, 131, 144 and 151 show the same signature: test failure followed by a pass on the identical commit.
Why it looks like a flaky suite
Any interruption of the log session fails the build: a deploy restart, a dropped connection, or the server side returning early. runRunnerLog in internal/control/build.go exits the read loop the moment AppendBuildLog errors, so a transient SQLITE_BUSY against the live database ends the session and takes the build down with it — the appends contend with ordinary forge traffic and with VACUUM INTO from the backup timers.
This trains you to hit rerun, which is the failure mode #62 was about, from a different direction.
Suggested shape
- A broken log sink must not fail a build. Track the step's own exit status separately from the copy error, and treat a dead sink as "logging degraded" rather than "step failed".
- Write step output to a local file and stream from it, so the record survives a dropped session.
- Do not end the server's read loop on a single append error; retry, or drop the chunk and keep the session.
Separately, a silent cap
AppendBuildLog drops everything past MaxBuildLog (2 MiB) with no error and no marker:
UPDATE builds SET log = log || ?
WHERE id = ? AND length(log) < ?
Not the cause here — both logs above are far under 2 MiB — but a timed-out e2e package dumps every goroutine, which clears 2 MiB easily, and the tail is exactly where the failure is. It should record that it truncated.
referenced in commit 62742f1870 by cmc: e2e: reserve every port before releasing any of them
2026-09-01 07:08 UTC