Make host monitoring speak, scan for vulnerabilities in CI, snapshot the database hourly !145

merged merged by cmc on 2026-09-01 04:45 UTC · krz/gitbay:sec-28-remaining into main

Discussion

cmc

The three items #28 was narrowed to.

Monitoring reported nowhere. gitbay-monitor.sh exited 0 at its second line whenever /etc/gitbay/monitor.url was absent, which it has been on bay1 since the script landed — the hourly timer computed disk, service and cert status and dropped all of it, and an unset webhook was indistinguishable from a healthy host. It now writes the reading to journald every run, reports a failed webhook post instead of swallowing it with || true, and exits non-zero on an alert so the unit shows up in systemctl --failed. Verified against stubbed df/systemctl: healthy exits 0, disk at 91% and a dead service each exit 1.

govulncheck was not in the loop. It lived only in deploy/audit.sh, run by hand. Now a CI job, separate from test so an advisory published against unchanged code cannot be mistaken for a test failure. It exits 0 against the tree today — the one module-graph finding, GO-2026-5932, is uncalled — so the job is not red on arrival.

Database mode. store.Open chmods to 0640. The directory above is 0750 so 0644 was not a live exposure, but the file carries token hashes and addresses. SQLite gives -wal/-shm the main file's mode, so all three follow. The test fails at 0644 and passes at 0640.

Recovery point. admin backup --db-only writes the SQLite snapshot without the repositories; gitbay-db-backup.timer runs it hourly keeping 48. The nightly full archive and its 7-day retention are untouched.

Plain hourly full backups do not work here: each is 550MB on bay1, and the retention is "keep the last 7", so hourly would have silently cut the history window from 7 days to 7 hours. Holding 7 days at hourly needs 168 archives, about 92GB. The split costs ~250MB and puts the tighter recovery point on the data that warrants it — repositories are git and have other copies, while issues, merge requests and comments have none.

Continuous replication was considered and rejected: it would take the database to seconds but leave repositories on the nightly archive, so a restore could produce a database referencing commits the repository backup does not have. Recorded on the wiki's Admin page, pushed separately.

Full suite green locally.

Two things found and not touched: gitbay-offsite.timer and its restic script exist on bay1 but are nowhere in deploy/cloud-init.yaml, so a rebuild from cloud-init comes up with no offsite backup; and the -wal/-shm files already on bay1 keep 0644 after deploy, since SQLite only sets the mode on files it creates.

Ref #28