os/exec surfaces a write error on a step's stdout through cmd.Wait(), and the runner handed the log session's pipe to every step as both stdout and stderr. When the session went away the copy goroutine hit EPIPE, Wait() returned non-nil for a process that had exited 0, and the step was reported failed — with the step failed: line explaining it written to the same dead pipe.
The result was a green suite recorded as a red build with a log that just stopped, which reads as a build that died mid-step. That is what sent me hunting for an e2e flake that does not exist.
Confirmed, not inferred
Build 148 ran 05:19:08–05:23:44 UTC. gitbayd restarted at 05:20:44, inside the window, dropping every SSH session including the log stream. The same commit passes the same suite in the same environment — ok gitbay.org/gitbay/e2e 272.656s, exit 0, run on bay1 as ci-runner. Stored logs: 460 bytes and 6 lines on the failure against 1569 and 29 on the success, with the failure never reaching step failed:.
Builds 127, 131, 144 and 151 share the signature.
The fix
logSink swallows write errors so Wait() reflects only the step's own exit status. It latches after the first failure rather than retrying a broken pipe once per write for the rest of a long build, and reports that the log is incomplete instead of leaving a gap that looks like a crash. Its mutex also closes the pre-existing race between the timeout path's Fprintf and the copy goroutine still draining.
runner log was the other half: it ended its read loop on any append error, so a single transient SQLITE_BUSY against the live database broke the runner's pipe and took the build down with it. It now drops the chunk, warns, and keeps draining.
AppendBuildLog dropped everything past the 2 MiB cap silently. It appends a notice once; the SQL bounds match exactly once, so later chunks fall through without repeating it.
Tests
TestBrokenLogSinkDoesNotFailTheStep reproduces the production failure — with the old error-returning sink it fails with step reported failed because its log sink broke: broken pipe. TestGenuineStepFailureStillFails pins the converse, so this cannot become a fix that swallows real failures. Plus latching, forwarding, and the truncation notice.
Full suite green.
Not done
The issue also suggested spooling step output to a local file and streaming from that, so a log survives a dropped session entirely. That is a larger redesign of the runner's output path and the build-correctness bug is fixed without it. Worth doing separately if incomplete logs prove annoying in practice.