Closes #246.
The test job took about 11m30s and grew with every test added. Nearly
all of it was e2e, run one test at a time.
Two changes:
Build each binary once per process. startInstance ran go build
for gitbayd on every test. Go's build cache covers the compile but not
the final link, so a run linked a byte-identical 35MB binary 208 times
at about 0.8s each. The gitbay and gitbay-runner helpers did the same.
One sync.Once per binary, under a directory TestMain owns, built
lazily so -run of one test does not link binaries it has no use for.
Run the tests in parallel. t.Parallel on 202 of 207. The five that
call t.Setenv stay serial, since it panics in a parallel test.
freePorts had to be replaced first, and that is the part worth
reviewing. It picked a port by binding :0, reading the number back and
closing, and between that close and the bind inside gitbayd the kernel
could hand the same port to another instance picking at that moment.
Serially the window is never contended. In parallel it is, and the loser
did not fail cleanly — waitForPort only asks whether something is
listening, so a test whose port was taken would talk to another test's
server. An atomic counter cannot collide within a process whatever the
interleaving; each candidate is still bind-tested, so ports other
programs hold are skipped.
The minimum-timeout guard (#143) goes with it: it refused a whole-package run under fifteen minutes because the suite took six to fifteen, and now refuses only reasonable commands.
Measured locally, 12 cores, go test ./... -count=1:
| wall | |
|---|---|
| before | 637s |
| build hoist | 401s |
| and parallel | 120s |
Identical results at each step: 658 pass, 3 skip, 0 fail.
That 120s was the fastest of six runs after the change; the others were 136s, 147s, 149s, 166s and 176s. The spread is this laptop under a long series of heavy runs, not the suite — all six were green with the same counts. Two or three minutes is the honest figure.
bay1 has 4 cores against this laptop's 12, so CI will not see 5.3x. The
build hoist does not depend on core count and should land harder there,
since a slower box spends longer on each of the 208 links. Worth reading
the first CI run before deciding whether to set -parallel above
GOMAXPROCS — the tests wait on sockets and subprocesses more than they
compute, so 4 may not be the right number for a 4-core box.
Bazel was considered and rejected; the reasoning is in #246.