.gitbay/wiki/Admin.org

1ed9fb9399b21da8e6cf45be8389792aac82d5bc
gitbay/.gitbay/wiki/Admin.org rendered · source · history · blame · raw

806 lines · 40689 bytes

  1#+title: gitbay admin guide
  2
  3One static binary (=gitbayd=), one SQLite file, bare repositories on
  4disk, and the system =git=. Schema migrations run automatically on
  5startup and on every admin command.
  6
  7* Install
  8
  9Build from source (=go build ./cmd/gitbayd=), install via the vanity
 10module path (=go install gitbay.org/gitbay/cmd/gitbayd@latest=), or use
 11a release build: =deploy/release.sh <tag>= cross-compiles reproducible
 12linux/amd64, linux/arm64, and darwin/arm64 binaries with a SHA256SUMS
 13manifest (CGO off, trimpath, stripped — byte-identical per commit and
 14toolchain).
 15
 16#+begin_src sh
 17install -m 755 gitbayd /usr/local/bin/
 18adduser --system --group --home /var/lib/gitbay --shell /usr/sbin/nologin gitbay
 19install -d -o gitbay -g gitbay -m 750 /var/lib/gitbay
 20gitbayd --config /etc/gitbay/config.toml check-config
 21#+end_src
 22
 23=deploy/= in the source tree has a cloud-init file, a hardened systemd
 24unit, and a nightly backup timer. Run as the unprivileged =gitbay= user;
 25the unit's =AmbientCapabilities=CAP_NET_BIND_SERVICE= covers ports
 2622/80/443 without root.
 27
 28** The SSH port decision
 29
 30- =ssh.mode = "embedded"= (default): gitbayd itself listens, normally on
 31  22 — move the host's admin sshd to another port. Remotes read
 32  =git@host:owner/repo= with no port gymnastics.
 33- =ssh.mode = "system"=: the host sshd owns 22 and invokes gitbayd via
 34  =AuthorizedKeysCommand=:
 35  #+begin_example
 36  AuthorizedKeysCommand /usr/local/bin/gitbayd --config /etc/gitbay/config.toml authorized-keys %t %k
 37  AuthorizedKeysCommandUser gitbay
 38  #+end_example
 39  sshd requires that binary to be root-owned and not group/world
 40  writable. Unknown keys fail authentication inside sshd, so system mode
 41  requires =registration.mode = "closed"= (check-config enforces this).
 42  Under the forced command the CLI's leading =--term=<cols>[,color]=
 43  argument works as is; a client that sets =GITBAY_TERM= instead needs
 44  =AcceptEnv GITBAY_TERM= in =sshd_config=, since that env request is
 45  handled by the host's sshd, not gitbayd.
 46
 47* Configuration reference
 48
 49=/etc/gitbay/config.toml=. =check-config= validates and names every
 50contradiction; =--no-host-checks= skips port/path probes.
 51=gitbayd admin config show= prints the configuration in effect as TOML,
 52every default filled in and =smtp_pass= redacted. A file that fails
 53validation still prints, followed by the contradiction.
 54
 55** [server]
 56- =root= (default =/var/lib/gitbay=) — repositories, database, host
 57  keys, ACME cache all live here.
 58- =site_url= (required) — canonical =https://host=; drives ACME, clone
 59  URLs, mail links.
 60- =source_repo= (optional, =owner/name=) — the repository this instance
 61  develops itself in. Startup warns when the running build's commit is
 62  not on that repository's default branch, which is how a binary built
 63  from an unmerged branch stops being invisible. Leave it unset unless
 64  the instance hosts its own source.
 65
 66** [ssh]
 67- =mode= — =embedded= | =system= (above).
 68- =port= (22) — embedded listener port.
 69- =host_keys= — list of private key paths; empty generates an ed25519
 70  key at =<root>/ssh/host_ed25519=.
 71
 72** [http]
 73- =addr= (=:443=), =tls= — =acme= | =files= | =off=.
 74- =acme=: certificates via TLS-ALPN-01 on the HTTPS port, cached at
 75  =<root>/acme=; =acme_email= for the CA account; =acme_http_addr=
 76  (=:80=, ="off"= to disable) adds HTTP-01 and an https redirect —
 77  failing to bind it is a warning, not fatal. Requires an =https://=
 78  site_url with a public DNS name.
 79- =files=: =cert_file= + =key_file=.
 80- =off=: plain HTTP — development, or behind a TLS-terminating proxy.
 81
 82=trusted_proxies= lists the addresses or CIDRs of reverse proxies in
 83front of the daemon. A request from one of them is attributed, for API
 84rate limiting, to the last =X-Forwarded-For= hop that is not itself a
 85trusted proxy; from anyone else the header is ignored. Empty, the
 86default, is right when gitbayd terminates TLS itself.
 87
 88A reverse proxy in front of gitbay must not buffer responses, or the
 89build page's live log arrives only when the build ends;
 90=X-Accel-Buffering: no= covers nginx.
 91
 92** [web]
 93- =mode= — =view_only= (default) | =accounts=. In view_only the mutating
 94  web routes are never registered; in accounts, browser sessions are
 95  minted over SSH (=web login=), and users with write access can create
 96  repos, comment, and make simple file edits (which commit unsigned,
 97  honestly). =password_auth= is reserved and currently rejected.
 98- =title= — the instance's display name in the rail and page titles;
 99  empty falls back to the site host. Lower case is the convention for
100  gitbay itself.
101- =privacy_notice= — operator text shown on =/privacy= under the fixed
102  statement.
103
104** [registration]
105- =mode= — =closed= (default) | =invite= | =open=. invite/open require
106  [mail]. See the user guide for the flows.
107
108- =pending_expiry= (empty, never) — a duration such as ="168h"=; a
109  self-registered account still unverified after that long is removed,
110  hourly and at start, audited as =pending.expired=.
111
112- =notify_admin= (false) — mail every instance admin when an account
113  becomes active: an invite redeemed, or an open-mode signup that
114  verified its address. The unverified row an open signup creates is
115  not reported, because anyone can post the form and mailing on that
116  would point a flood at the admins. Recipients are the verified
117  primary addresses of active admins who have activity mail on, the
118  same rule any other notice follows, so an admin with no verified
119  address hears nothing. The notice is queued, so a dead SMTP host
120  shows up in the admin page's Mail table instead of failing the
121  registration. Requires [mail].
122
123** [mail]
124- =smtp_host= (host:port, 587 assumed), =from=, optional =smtp_user= /
125  =smtp_pass=. STARTTLS when offered. Required for invite/open
126  registration and self-service =email add=; in closed mode you may omit
127  it entirely and assert addresses by hand (below).
128
129** [push]
130Push notifications to Apple devices, delivered by gitbayd talking to
131APNs directly over HTTP/2, authenticated by an ES256 JWT signed with an
132operator-supplied provider key. Off unless configured.
133
134- =enabled= (false).
135- =key_file= — path to the =.p8= provider key from Apple's developer
136  portal (Certificates, Identifiers & Profiles → Keys). It belongs at
137  =/etc/gitbay/apns.p8=, mode 0600, owned by the account gitbayd runs
138  as. Read and validated at startup: it must parse as a PEM-wrapped
139  PKCS#8 EC (P-256) private key, or the daemon refuses to start rather
140  than fill a queue nobody is watching.
141- =key_id=, =team_id= — the key's id and your Apple developer team id,
142  both from the same portal page.
143- =topic= — the app's bundle identifier. *An APNs key belongs to a
144  bundle ID.* gitbay.org pushes to the App Store build under its own
145  bundle id; a self-hoster who wants push ships their own iOS build
146  under their own bundle id, with its own =.p8= key from their own
147  developer account, and points =topic= at that id. There is no way to
148  push to someone else's build, by design — this is Apple's model, not
149  gitbay's.
150- =environment= — =production= or =sandbox=, naming the APNs host
151  rather than taking a URL, so a typo cannot aim the key at a host that
152  is not Apple's.
153
154All five of =key_file=, =key_id=, =team_id=, =topic= and =environment=
155are required when =enabled= is true; validation runs at config load,
156so a misconfigured =[push]= is caught before the daemon serves
157anything. The delivery queue (a device's undelivered and attempted
158pushes) is capped the same way the mail queue is, by =[retention]
159push=.
160
161** [api]
162- =enabled= (false) — the JSON API surface; see [[API]]. Off
163  means no credential-bearing HTTP endpoint exists at all.
164
165** [webhooks]
166- =allow_local= (false) — permit webhook targets on loopback/private
167  addresses. Leave off unless you know why you need it (SSRF).
168
169** [limits]
170- =clone_timeout= (3600s) — cap on =repo import= fetches.
171- =max_blob_bytes= (100MB) — cap on raw file serving over the web.
172- =max_asset_bytes= (512MB) — cap per uploaded release asset.
173- =max_snippet_bytes= (1MB) — cap per snippet file.
174- =max_snippets_per_user= (0, unlimited) — snippets an account may own.
175- =max_repos_per_user= (0, unlimited) — repositories an account may own
176  directly; =repo create=, =fork= and =import= refuse past it.
177  Organizations are not capped.
178- =max_bytes_per_user= (0, unlimited) — disk the account's own
179  repositories may take; a push may be no larger than what is left.
180- =max_pack_bytes=, =ssh_auth_rate= — reserved, not yet enforced.
181
182** [git_daemon]
183- =enabled= (false), =port= (9418) — the anonymous =git://= listener.
184  Serves only public repositories that additionally ran
185  =repo settings git-daemon <repo> on=.
186
187** [mirrors]
188- =pull_interval_minutes= (15) — how often pull mirrors fetch their
189  upstream. Push mirrors sync shortly after each local ref update.
190  Mirror URLs pass the same SSRF rules as webhook targets.
191
192** [go_import]
193Vanity Go module paths, one per line: ="host/module" = "owner/repo"=.
194Requests with =?go-get=1= at or under the module path answer with the
195go-import meta tag pointing at the repository's HTTPS clone URL, so
196=go install host/module/cmd/...@latest= resolves. The repository should
197be public (the module path itself confirms it exists).
198
199* Users, email, invites
200
201Every =gitbayd admin= subcommand except =backup=, =gc= and the one-shot
202backfills is a wrapper that dispatches the registry command of the same
203name as the host: an admin context with no account behind it, so its
204audit rows carry no actor and =source: host=. The same commands run in an
205instance admin's SSH session (=ssh git@<host> admin ...=) and write the
206same rows with the key fingerprint as source. One implementation, two
207credentials.
208
209#+begin_src sh
210gitbayd admin user create alice --key alice.pub --email a@example.org --verified [--admin]
211gitbayd admin email verify alice a@example.org   # admin assertion, no SMTP needed
212gitbayd admin invite --email b@example.org       # mails a code; prints it if no SMTP
213#+end_src
214
215"Verified" means SMTP-confirmed or host-admin-asserted; the database
216records which. Verified emails are what make commit signatures
217meaningful — an unverified address never produces a =verified= badge.
218
219* Audit and account control
220
221The audit log is the security feed (events are the product feed): every
222successful mutating command with its argv and source credential (SSH key
223fingerprint or API), registrations, admin actions, force-pushes, and
224auth failures/throttling. Secrets never appear — they travel on stdin,
225never in argv.
226
227#+begin_src sh
228gitbayd admin audit [--actor u|-] [--action prefix] [--since 24h|7d|date] [--limit n] [--json]
229ssh git@<host> audit ...             # the same, from an admin session
230ssh git@<host> admin user list [--state active|pending|disabled|admin]
231ssh git@<host> admin user show <name>   # keys, emails, orgs, tokens, sessions
232ssh git@<host> admin user limits <name> [--repos n|default] [--bytes n|default]   # per-account caps
233ssh git@<host> admin user promote <name>   # grant instance admin
234ssh git@<host> admin user demote <name>    # remove it; the last admin is refused
235gitbayd admin user promote <name>    # host-local: recovery when no admin key is reachable
236gitbayd admin user disable <name>    # suspend: SSH, web, API all refused;
237gitbayd admin user enable <name>     #   sessions dropped, nothing deleted
238gitbayd admin user delete <name> --yes  # only for accounts anchoring nothing:
239                                     #   refused (with each blocker named) while
240                                     #   the account owns repos, authored
241                                     #   issues/MRs/comments/reviews, or is an
242                                     #   org's only admin
243#+end_src
244
245=--actor= takes a username, or =-= for rows with no actor: host commands
246and auth failures. =--action= is a prefix, so =cmd repo= catches every
247repository command and =admin= every host or admin-session action.
248=--since= is a duration back from now (=30m=, =24h=, =7d=) or a date.
249
250=admin user list= pages by username (=--limit=, =--cursor=) and carries
251each account's state and =last_seen=, the newest use of any of its SSH
252keys or API tokens. =admin user show= adds the keys with their last use,
253each address with how it was verified, PGP keys, org roles, the owned
254repository count, API token names, and live browser sessions. Both are
255Both are refused to non-admins, like =audit=, on every surface.
256
257=/admin/users= is the same list in a browser, linked from the admin
258page: the state filter the command takes, keyset paging on its cursor,
259and a row per account with promote, demote, disable and enable, each
260dispatching the command. Demote and disable ask for the username to be
261typed, since both take someone's access away. Creating and deleting an
262account, issuing an invite and asserting an address stay on the command
263line: each takes a key, mints a credential, or cannot be undone. A
264non-admin gets the 404 a missing page would, so the URL confirms
265nothing.
266
267Promotion needs an active account: a pending or disabled one is refused.
268Demotion is refused when it would leave no admin, over SSH, in the
269browser and on the host alike, so the host-local =promote= is the way
270back in when the only admin key is lost.
271
272Instance admin carries no right on anyone's repository: policy does not
273consult it, and a private repository still answers not-found to an
274admin. Moderation goes through explicit overrides that skip the access
275check and write their own audit row:
276
277#+begin_src sh
278ssh git@<host> admin repo list [--owner o] [--visibility public|private]  # size, last push
279ssh git@<host> admin repo archive|unarchive <owner/name>
280ssh git@<host> admin repo visibility <owner/name> public|private
281ssh git@<host> admin repo delete <owner/name> --yes
282#+end_src
283
284Each lands in the audit log as =admin repo.<action>= naming the
285repository, on top of the =cmd= row every mutating command gets.
286
287=limits.ssh_auth_rate= (10) throttles per-IP authentication *failures*
288per minute — successful auths never count and clear the slate.
289=limits.max_pack_bytes= is enforced as =receive.maxInputSize= on every
290push.
291
292=limits.write_rate= (60) bounds *mutating commands per account per
293minute*. It is counted in the dispatcher, so SSH, the JSON API and the
294web spend one budget and a caller cannot refresh it by changing surface;
295=limits.api_rate= stays in front of it, bounding a network source rather
296than an account. A command is one token whatever it writes, so a bundle
297import costs one and only a loop of separate commands spends the budget.
298Read-only commands, the runner protocol (a build streams its log in many
299small writes) and the host CLI are exempt. Refusals exit 4 and say when
300to retry. A negative value turns the limit off; it matters most with
301=registration = "open"=, where every write also queues notification mail
302and webhook deliveries.
303
304* Queues
305
306Every background worker keeps a backlog and a failure state. An instance
307admin reads them all in one place:
308
309#+begin_src sh
310gitbay dashboard --json | jq .queues   # webhooks, mail, push, mirrors, builds, deps
311#+end_src
312
313Per worker: pending, retrying (pending with a failed attempt) and
314dead-lettered counts with the oldest pending age, and the retrying or
315failed rows themselves, capped at twenty each. Builds list what is
316running and then what is pending, each since when, so a build no runner
317is scoped to claim is visible here rather than only in its repository;
318mirrors list the ones whose last sync failed;
319dependency checks list the ones whose last check errored. Non-admins get
320no =queues= key at all.
321
322Push rows name the device id, never the token. Watch this one after
323configuring =[push]=: a =key_id= or =team_id= Apple did not issue passes
324config validation, which can only check that the =.p8= parses, and then
325every send comes back =403 InvalidProviderToken= and dead-letters on its
326first attempt.
327
328In accounts mode the same read renders at =/admin=, linked from the rail
329for admins. Anyone else gets a 404 there.
330
331A dead-lettered mail is logged as =notification dead-lettered mail=<id>=,
332with the address redacted out of the relay's error. The id is the queue
333row: find it in the Mail table on =/admin=, or in =dashboard --json=,
334where the recipient and the unredacted error are. That is deliberate —
335see the Threat-Model page.
336
337* Maintenance
338
339#+begin_src sh
340gitbayd admin stats [--json]         # counts, database size, per-repo disk
341ssh git@<host> admin stats [--json]  # the same, from an admin session
342gitbayd admin gc [--repo owner/name] # git gc: repack and prune; per-repo sizes
343gitbayd admin gc --aggressive        # thorough repack; slow, rarely needed
344gitbayd admin gc --lfs               # also drop LFS objects no pointer names (older than a day)
345#+end_src
346
347A history rewrite leaves the commits it removed reachable through
348=refs/merge-requests/N/head= of the merge requests that landed them, so
349they stay fetchable by anyone who can read the repository. Nothing drops
350a head ref on its own — an open or source-gone MR is merged through it,
351and a merged or closed one keeps its diff readable through it — so the
352cleanup is a command an instance admin runs, naming the MRs:
353
354#+begin_src sh
355ssh git@<host> admin mr prune owner/name 1 2 3 --yes
356#+end_src
357
358It refuses an open or source-gone MR, deletes the named refs, runs
359=git gc --prune=now= on that one repository so the objects go at once
360rather than after git's two-week grace, leaves a system comment on each
361MR, and audits as =admin mr.prune=. The MR keeps its title, comments,
362reviews and head sha; =mr diff= and the MR page say the head is gone.
363Run it when nothing is pushing to that repository: without the grace, a
364push caught between leaving quarantine and writing its ref loses its
365objects. Objects also survive in offsite backups until those are
366rewritten; see "Removing a repository's history from every snapshot".
367
368=deploy/cloud-init.yaml= ships a =gitbay-gc.timer= that runs =admin gc=
369weekly (Sunday 07:00 UTC). Imported repositories keep whatever pack
370layout the source sent, so a first manual =admin gc= after a bulk
371import is worthwhile.
372
373* Backup and restore
374
375#+begin_src sh
376gitbayd admin backup --out /var/backups/gitbay/backup.tar.gz
377gitbayd admin backup --verify /var/backups/gitbay/backup.tar.gz   # read it back
378#+end_src
379
380One archive: a consistent SQLite snapshot (taken *before* the
381repositories are read, so the database never references objects the
382archive missed), every repository, and the SSH host keys. Excluded:
383hook socket, regenerated hook scripts, WAL files. Safe to run against a
384live daemon.
385
386=--verify= reads an archive back: the snapshot must pass SQLite's
387integrity check, and every repository the snapshot names must be in the
388archive. A database-only archive is checked for integrity and says so.
389Exit is non-zero on damage or a missing repository.
390
391Restore: extract into an empty directory, point =server.root= at it,
392start gitbayd. Host keys are preserved, so clients keep their
393known_hosts entries; hooks regenerate at startup.
394
395** Schedule and recovery point
396
397Two timers, because the two halves of the data have different exposure.
398
399- =gitbay-backup.timer=, nightly. The full archive above, last 7 kept.
400- =gitbay-db-backup.timer=, hourly. =admin backup --db-only=, which
401  writes the SQLite snapshot alone, last 48 kept. A few MB against the
402  full archive's hundreds, which is what makes the frequency affordable.
403
404The split follows what a loss would actually cost. Repositories are git,
405so a mirror or any clone is a second copy; the database is the only copy
406of issues, merge requests, comments and review state. So the recovery
407point is about an hour for the data that exists nowhere else, and a day
408for the data that does.
409
410Continuous replication (litestream and similar) was considered and not
411adopted. It would take the database's recovery point to seconds, but the
412repositories would still be on the nightly archive, so a restore could
413produce a database referencing commits the repository backup does not
414have. Consistency between the two halves is worth more here than latency
415on one of them. Revisit if repository replication becomes continuous
416too.
417
418** Offsite copies
419
420bay1 also takes a nightly restic snapshot of =/var/lib/gitbay= and
421=/var/lib/gitbay-stage= (the staged database copy) to an S3 bucket at
422Scaleway, with a key that can only add snapshots. The key that can
423remove them lives on the operator's machine, in
424=~/.config/gitbay/offsite.env=, and never on bay1: a compromised host
425cannot destroy its own history. Forgetting, pruning and rewriting all
426run from there.
427
428*** Removing a repository's history from every snapshot
429
430A history rewrite plus =admin mr prune= takes commits off the server,
431but every snapshot taken before it still holds them, and the retention
432window is the only thing that ages them out. To remove them now,
433rewrite the snapshots without that repository rather than forgetting
434the snapshots: everything else in them stays restorable. The next
435nightly run adds the repository back in its current state.
436
437Repositories are stored under the name they had on disk when each
438snapshot was taken, so a renamed repository needs every name it has
439carried. Check what an older snapshot holds before choosing the paths:
440
441#+begin_src sh
442set -a; . ~/.config/gitbay/offsite.env; set +a
443restic $RESTIC_OPTS snapshots
444restic $RESTIC_OPTS ls <old-snapshot> /var/lib/gitbay/repos/<owner>
445#+end_src
446
447Then dry-run, apply, prune, and confirm nothing matches:
448
449#+begin_src sh
450EXCL="--exclude /var/lib/gitbay/repos/<owner>/<name>.git --exclude /var/lib/gitbay/repos/<owner>/<old-name>.git"
451restic $RESTIC_OPTS rewrite --dry-run $EXCL     # "would modify N snapshots"
452restic $RESTIC_OPTS rewrite --forget $EXCL      # new snapshots replace the originals
453restic $RESTIC_OPTS prune                       # drops the data nothing references
454restic $RESTIC_OPTS find <name>.git <old-name>.git   # expect no output
455#+end_src
456
457=--forget= is what makes the originals go; without it the rewritten
458snapshots sit beside them and the data stays referenced. Snapshot IDs
459change; their times do not. Done for krz/keycask (formerly rust-pass)
460on 2026-09-18, across 22 snapshots.
461
462* Upgrades
463
464Replace the binary, restart the unit. Migrations apply automatically and
465are transactional; hook scripts under =<root>/hooks= are rewritten at
466startup to point at the current binary path.
467
468* CI runner
469
470=gitbay-runner= executes builds queued by pushes and merge requests. It
471polls over SSH with a key of scope =runner=, which reaches only the
472runner protocol and read-only git (a runner executes arbitrary
473repository code, so the key it holds must not do more). A runner key
474claims builds only for the repositories it is attached to, by =repo
475runner add= from a repository admin or an instance admin; an admin key
476claims any. Users attach their own runners: see the Users page. For an
477instance runner, run it as a dedicated unprivileged user on a non-admin
478account. =admin user create --key= registers a full-scope key, so the
479runner key is added afterwards through a bootstrap key that is then
480removed, and attached to each repository it should build:
481
482#+begin_src sh
483useradd --system --create-home --home-dir /var/lib/gitbay-runner ci-runner
484sudo -u ci-runner ssh-keygen -t ed25519 -N "" -f /var/lib/gitbay-runner/.ssh/id_ed25519
485ssh-keygen -t ed25519 -N "" -f /tmp/ci-bootstrap
486gitbayd --config /etc/gitbay/config.toml admin user create ci --key /tmp/ci-bootstrap.pub
487ssh -i /tmp/ci-bootstrap git@127.0.0.1 keys add --scope runner < /var/lib/gitbay-runner/.ssh/id_ed25519.pub
488ssh -i /tmp/ci-bootstrap git@127.0.0.1 keys remove "$(ssh-keygen -lf /tmp/ci-bootstrap.pub | awk '{print $2}')"
489rm /tmp/ci-bootstrap /tmp/ci-bootstrap.pub
490gitbay-runner -remote git@127.0.0.1 -workdir /var/lib/gitbay-runner/work
491#+end_src
492
493#+begin_src sh
494gitbay repo runner add krz/site < /var/lib/gitbay-runner/.ssh/id_ed25519.pub
495#+end_src
496
497=-jobs N= runs N builds at once. Claiming is one transaction that
498selects and updates, and each build works in its own =build-<id>=
499directory, so workers do not collide; idle polls are staggered across
500the interval so N of them do not wake together. The drop-in's weights
501below are per service, not per build, so raising =-jobs= divides them
502rather than multiplying the host's load.
503
504=admin runners= shows which account each runner polls as, and what each
505may claim. A runner key claims builds only for the repositories it is
506attached to: with none attached it claims nothing, and =-repos= may only
507narrow within them. An admin's full-scope key claims any repository —
508that is what =-repos= was for — and still works for the protocol during
509a rotation. A merge request head from a fork is built in the target
510repository as untrusted: the claim carries no secrets, and only a runner
511started with =-untrusted= takes it. Same-repository heads were built by
512their branch push and are not built again.
513
514=make deploy-runner= also installs
515=deploy/gitbay-runner.override.conf= as a systemd drop-in: =Nice=10=,
516=CPUWeight=30=, =IOWeight=30=, so a build never starves the host's sshd,
517the daemon or the backup timers, and =NoNewPrivileges=,
518=ProtectSystem=full=, =ProtectKernelTunables=, =ProtectControlGroups=
519and =RestrictSUIDSGID=, so a step cannot reach outside its workspace
520and the runner's home. The e2e suite alone starts sixty
521daemon instances; without the drop-in a deploy's copy over the admin
522sshd stalled. Both deploy targets copy with =rsync --partial=, which
523resumes a stalled transfer.
524
525Under =-isolation podman=, the default and what bay1 runs, each build
526is confined to a container (see Container isolation below). Under
527=-isolation none= steps run directly on the host as the runner's user,
528so treat that machine as executing whatever your users push, and
529install the toolchains your builds need on it.
530
531A runner claims the oldest pending build among the repositories its key
532is attached to — for an admin key, the oldest in the instance. =-repos=
533narrows within that set, which is what makes a runner outside the server
534practical: one on a machine that should build a single project, or that
535holds credentials for one deployment, stays on it.
536
537Oldest-first is across everything the key may claim, so a repository
538with a deep queue holds every other repository the same runner serves;
539bay1 measured a 15-minute average wait on a day of merge request
540stacks from one repository. A runner attached to one repository cannot
541be starved. That is the rule, decided in krz/gitbay#207: a runner
542serving several repositories takes them oldest-first, and an operator
543who wants one repository never to wait on another runs a second
544runner attached to it alone. Nothing caps what an account queues,
545and nothing needs to (krz/gitbay#206): a schedule tick queues nothing
546while the job's last build is pending or running, so a repository
547with no runner holds one row per scheduled job rather than one per
548tick, and a build runs only on a runner its owner attaches, so a busy
549schedule spends the owner's compute. Pushes are bounded by what an
550account can push.
551
552#+begin_src sh
553gitbay-runner -remote git@gitbay.org -repos krz/site,krz/docs \
554  -workdir /var/lib/gitbay-runner/work
555#+end_src
556
557Add =-untrusted= only with =-isolation podman=.
558
559gitbay.org's runner is attached to the forge's own repositories and
560the isolation canary, nothing else, because it shares the host with
561the forge; its unit names no =-repos=, the attachments are the
562boundary. Any other repository builds on a runner its owner attaches.
563
564=-repos= narrows an admin runner; for a runner key the attachments are
565the boundary, held by the server, and =-repos= may only name
566repositories among them. =-untrusted= makes a runner claim merge
567request heads from forks; the bay1 unit sets it because it isolates in
568podman. A runner without it builds trusted commits only.
569
570=gitbay dashboard= and =ssh git@<host> admin runners= list every key
571that has polled as a runner: the account, the key's fingerprint, when it
572last polled, the repositories it may claim — its attachments for a runner
573key, the =-repos= it asked for or =any= for an admin key — and the build
574it holds; =admin runners remove <fingerprint>= (=forget= until the next release) drops the row for a key
575that polled by mistake, the key itself untouched. =admin runners= also
576heads the list with the queue: builds
577pending now, and over the last day how many were claimed, how long they
578waited to be claimed (average and worst), and
579how many the reaper ended instead of a runner reporting them. A build a runner claimed and never
580reported is failed by the scheduler's minute tick, whether or not any
581runner is still alive: within about two minutes of its log stream ending
582with no outcome reported — the runner reports right after closing the
583stream, retrying for half a minute if gitbayd is unreachable — or, if no
584stream was ever seen, at the deadline.
585
586Instance admin on the runner account only authorizes the claim/report
587protocol; it grants no repo access. A build that pushes back — a pages
588deploy, an archive publish, an automated MR branch — needs an explicit
589grant on that repo: =repo access grant <owner/name> ci write=. Private
590repos likewise need at least read for the clone.
591
592** Container isolation
593
594Builds run in a rootless podman container, one per job, with the
595workspace bind mounted and nothing else: the clone happens outside with
596the runner's key, so a step cannot read it. =-isolation none= keeps the
597old behaviour — steps on the host as the runner's user — for an instance
598where every repository is trusted. There is no automatic fallback: a
599runner started with =-isolation podman= that cannot find a working
600podman exits rather than running a build unsandboxed.
601
602gitbay's own jobs name =localhost/gitbay-ci:2=, built from
603=deploy/Containerfile.ci= on the runner host. A job's image must carry
604what its steps need: the suite drives real git, git-lfs, gpg and sshd and
605asserts they exist before running, so the stock runner default would fail
606it immediately. Build or rebuild it with:
607
608#+begin_src sh
609ssh -p 2222 root@<host> 'cat > /tmp/Containerfile.ci' < deploy/Containerfile.ci
610ssh -p 2222 root@<host> 'su - ci-runner -s /bin/sh -c \
611  "podman build -t localhost/gitbay-ci:2 -f /tmp/Containerfile.ci /tmp"'
612#+end_src
613
614The tag is deliberate rather than =:latest=: changing the file means
615bumping the tag in =.gitbay/ci.yml=, so a running branch's image does not
616change under it.
617
618=-image= names the image a job runs in when it declares none, and is
619required under =-isolation podman=: there is no built-in default,
620because an image this host does not have would fail every build. A job
621overrides it with =image:= in =.gitbay/ci.yml=, validated as a reference
622so a config file cannot turn it into podman arguments.
623
624=-cpus= and =-memory= cap one build (podman's units, e.g. =-cpus 2
625-memory 4g=); unset means uncapped. The runner applies them itself: it
626creates a cgroup per build under its own delegated service cgroup,
627writes the limits there, and starts every podman process for the build
628inside it, with podman's cgroup handling off. Podman's own =--memory=
629and =--cpus= never applied under rootless cgroupfs, which is what a
630system service gets (krz/gitbay#188). The unit therefore needs
631=Delegate=yes=, which the drop-in sets; without it the runner refuses
632to start when a limit is set, and logs that builds run unconfined when
633none is. bay1 runs =-cpus 3 -memory 6g= per build inside =MemoryMax=6G=
634and =CPUQuota=300%= on the unit, on a 7.7GB four-core host with no
635swap: the memory cap is what keeps the forge alive when a build
636allocates without bound, and it sits above the e2e suite's 5GB peak
637rather than at a fair share. =OOMPolicy=continue= keeps systemd from
638stopping the runner when a build is OOM-killed.
639
640Each repository gets its own build home under the runner's workdir,
641mounted into its containers as =HOME=. Caches persist between builds of
642one repository and are never read by another's.
643
644*Images are provisioned, never pulled by a build.* The runner passes
645=--pull=never=. Two reasons, and the second is the better one: the
646service runs with =RestrictSUIDSGID=yes= so podman cannot unpack a layer
647holding a setuid file, which is nearly every distribution image; and on
648an instance where anyone can push a =ci.yml=, =image:= would otherwise
649mean "fetch and run anything from the internet". An operator pulls or
650builds what is allowed and a build picks among those. A job naming an
651image the host does not have fails with a message saying so.
652
653#+begin_src sh
654su - ci-runner -s /bin/sh -c "podman pull docker.io/library/alpine:3.20"
655su - ci-runner -s /bin/sh -c "podman images"
656#+end_src
657
658Prepare a host before pointing an isolating runner at it:
659
660#+begin_src sh
661ssh -p 2222 root@<host> 'sh -s' < deploy/runner-podman-setup.sh
662make deploy-runner
663#+end_src
664
665The script installs podman, delegates a subuid/subgid range to
666=ci-runner=, checks that user namespaces are enabled rather than
667assuming, enables lingering, and verifies rootless podman actually runs
668as that user. It is idempotent.
669
670The drop-in sets =NoNewPrivileges=no=, without which rootless podman
671cannot call =newuidmap= and the runner refuses to start. That is a
672considered trade, explained in the file and in the Threat-Model; if you
673run with =-isolation none=, set it back to =yes=.
674
675*Restarting the runner is safe.* On SIGTERM it stops claiming, finishes
676the build in flight, reports it, and exits; the drop-in's
677=TimeoutStopSec=50min= covers the longest build, and its =KillMode=mixed=
678is what makes the signal reach the runner alone — under systemd's default
679the build's container and the log session are signalled with it, and the
680runner drains a build that is already dead. So =make deploy-runner=
681waits for a running build rather than orphaning it, and a build's result
682is retried for half a minute if gitbayd is restarting at that moment. A
683second SIGTERM ends the runner at once, abandoning the build to the
684reaper. The suite checks all three: =TestRunnerDrainsOnSIGTERM= signals
685the process, =TestRunnerDropInLetsTheDrainHappen= reads the drop-in's
686=KillMode= and =TimeoutStopSec=, and =TestRunnerDrainUnderSystemd= runs
687the runner as a transient user unit under =systemd-run= and stops it
688under both kill modes. That last one needs a systemd user manager, so it
689skips in the container CI runs in; run it on a Linux host with
690=go test ./e2e -run TestRunnerDrainUnderSystemd -v=.
691
692*Validate podman mode on a scratch repository before pointing the runner
693at real ones.* Every deploy that switched the whole instance to
694containers and failed took CI down with it. Instead: create a throwaway
695repository the runner account can read (public, or granted read — a
696private one is "not found" to the runner and the build stays pending),
697give it one job that names the CI image, and deploy the runner with
698=-repos= naming only that repository. The production unit, with its real
699hardening, then claims nothing else; other repositories' builds queue
700until =-repos= is switched back, which is a pause, not an outage.
701
702#+begin_src sh
703gitbay repo create cmc/ci-smoke          # then push a .gitbay/ci.yml naming the image
704sed -i 's#-repos krz/gitbay #-repos cmc/ci-smoke #' /etc/systemd/system/gitbay-runner.service.d/override.conf
705systemctl daemon-reload && systemctl restart gitbay-runner
706gitbay build log cmc/ci-smoke 1         # green: switch -repos back, redeploy
707#+end_src
708
709*Do not deploy an isolating runner to a host that has not been
710prepared.* The runner is specified to refuse to start without a working
711podman rather than fall back to running builds unsandboxed — a fallback
712that silently drops isolation is worse than a stopped runner, because
713nothing surfaces it. On an unprepared host that refusal stops every
714build on the instance.
715
716The service drop-in carries =Delegate=yes= for rootless cgroup
717management and =ReadWritePaths= for podman's store under
718=/var/lib/gitbay-runner=, which =ProtectSystem=full= would otherwise
719make read-only. Those paths are prefixed =-= so they are ignored when
720absent: the drop-in installs on unprepared hosts too, and a unit that
721refused to start would stop every build.
722The nightly canary on =cmc/ci-smoke= only runs if the runner's =-repos=
723names that repository too; a scoped runner claims nothing else.
724=gitbay-runner-prune.timer= prunes unused images weekly, as the runner's
725user: rootless storage belongs to that user, and root's prune would not
726see it. An unpruned image store on a 40GB host is a slow outage.
727
728* LFS storage
729
730Objects live content-addressed under =[lfs] root= (default
731=<server.root>/lfs=); =[lfs] max_object_bytes= caps a single object
732(512MB default). Storage sits behind a small interface — an
733S3-compatible backend is a drop-in with the server proxying, and
734presigned URLs a later optimization. LFS objects do not travel with
735push mirrors (mirrors move git refs only), and gc does not yet collect
736orphaned objects.
737
738* Pages
739
740=[pages] domain = "example.site"= serves public repos' =pages= branches
741on =<owner>.<domain>=. DNS needs a wildcard record =*.<domain>= to the
742server; ACME issues per-subdomain certificates on demand (only for
743owners that exist). The domain must not be the site host or a parent of
744it — pages content runs its own scripts and must stay off the forge's
745origin.
746
747Users with repo admin claim custom domains with =repo domain add=.
748Claims activate only after a DNS TXT challenge proves control of the
749domain (=repo domain verify=, audit-logged); pending claims serve
750nothing, get no certificates, and expire after 7 days. ACME issues
751certificates only for verified hosts, so stray DNS pointed at the
752server gets nothing.
753
754* Security
755
756The [[Threat-Model]] file is the reference for what the forge
757trusts and refuses to do. Operational checklist:
758
759- *Software checks.* =deploy/audit.sh= runs =go vet=, =govulncheck=
760  (the module list is deliberately short — review it on each release),
761  and a short fuzz pass over every attacker-facing parser (pkt-line,
762  commit, SSHSIG armor, OpenPGP key, SSH tokenizer). Run it before
763  tagging a release. CI's own =vuln= job runs =govulncheck= nightly
764  against main rather than per push, because =@latest= scans today's
765  advisory database and an advisory lands without anyone pushing;
766  =build trigger krz/gitbay vuln= runs it on demand.
767- *Web responses* carry a scripts-forbidden CSP, =X-Frame-Options:
768  DENY=, =nosniff=, =no-referrer=, and HSTS when TLS is on — no
769  configuration needed.
770- *Host sandboxing.* The systemd unit in =deploy/cloud-init.yaml= runs
771  gitbayd unprivileged with =ProtectSystem=strict=, =PrivateDevices=,
772  =LockPersonality=, =MemoryDenyWriteExecute=,
773  =SystemCallFilter=@system-service=, and =RestrictAddressFamilies= to
774  INET/INET6/UNIX. It keeps =CAP_NET_BIND_SERVICE= only, to bind 22/80/443.
775- *OS patches* apply via =unattended-upgrades= (security origins,
776  auto-reboot 04:30 if required).
777- *Admin sshd (2222)* is throttled by =MaxStartups=/=MaxAuthTries= and
778  watched by =fail2ban=; gitbayd's own port 22 is throttled by
779  =limits.ssh_auth_rate= (auth failures per IP per minute), and every
780  account's writes by =limits.write_rate=.
781- *Monitoring.* =gitbay-monitor.timer= writes a reading hourly to
782journald and, when =/etc/gitbay/monitor.url= exists, posts it to that
783webhook: disk, service, the daemon's own =/healthz= answer, certificate
784expiry, and the age of the newest full backup and database snapshot.
785It exits non-zero on an alert so the unit shows in =systemctl
786--failed=: a stopped service, =/healthz= not answering =ok=, disk ≥ 85%,
787a certificate under 21 days, a full backup over 25 hours old, or a
788database snapshot over 2 hours old.
789
790=GET /healthz= is unauthenticated and cache-free: whether the database
791answers and which commit serves, 503 when it does not.
792- *Database.* =gitbay.db= and its WAL live under =/var/lib/gitbay= (mode
793  0750, owned by =gitbay=). The nightly archive plus provider snapshots
794  are the recovery path; for tighter RPO, add continuous replication
795  (litestream) against the same file — it coexists with the WAL.
796
797* Odds and ends
798
799- deleting a fork marks MRs sourced from it =source_gone=; their diffs
800  remain viewable and mergeable because the target repo owns the
801  objects.
802- =refs/merge-requests/*= is server-owned and unpushable by clients;
803  only =admin mr prune= removes one (see Maintenance).
804- audit-relevant activity (issue/MR lifecycle, imports, pushes) lands in
805  the =events= table, which also feeds webhooks.
806- the daemon idles under 10MB RSS; the smallest VPS tier is adequate.