.gitbay/wiki/Admin.org

v1.31.0
gitbay/.gitbay/wiki/Admin.org rendered · source · history · blame · raw

760 lines · 38305 bytes

  1#+title: gitbay admin guide
  2
  3One static binary (=gitbayd=), one SQLite file, bare repositories on
  4disk, and the system =git=. Schema migrations run automatically on
  5startup and on every admin command.
  6
  7* Install
  8
  9Build from source (=go build ./cmd/gitbayd=), install via the vanity
 10module path (=go install gitbay.org/gitbay/cmd/gitbayd@latest=), or use
 11a release build: =deploy/release.sh <tag>= cross-compiles reproducible
 12linux/amd64, linux/arm64, and darwin/arm64 binaries with a SHA256SUMS
 13manifest (CGO off, trimpath, stripped — byte-identical per commit and
 14toolchain).
 15
 16#+begin_src sh
 17install -m 755 gitbayd /usr/local/bin/
 18adduser --system --group --home /var/lib/gitbay --shell /usr/sbin/nologin gitbay
 19install -d -o gitbay -g gitbay -m 750 /var/lib/gitbay
 20gitbayd --config /etc/gitbay/config.toml check-config
 21#+end_src
 22
 23=deploy/= in the source tree has a cloud-init file, a hardened systemd
 24unit, and a nightly backup timer. Run as the unprivileged =gitbay= user;
 25the unit's =AmbientCapabilities=CAP_NET_BIND_SERVICE= covers ports
 2622/80/443 without root.
 27
 28** The SSH port decision
 29
 30- =ssh.mode = "embedded"= (default): gitbayd itself listens, normally on
 31  22 — move the host's admin sshd to another port. Remotes read
 32  =git@host:owner/repo= with no port gymnastics.
 33- =ssh.mode = "system"=: the host sshd owns 22 and invokes gitbayd via
 34  =AuthorizedKeysCommand=:
 35  #+begin_example
 36  AuthorizedKeysCommand /usr/local/bin/gitbayd --config /etc/gitbay/config.toml authorized-keys %t %k
 37  AuthorizedKeysCommandUser gitbay
 38  #+end_example
 39  sshd requires that binary to be root-owned and not group/world
 40  writable. Unknown keys fail authentication inside sshd, so system mode
 41  requires =registration.mode = "closed"= (check-config enforces this).
 42
 43* Configuration reference
 44
 45=/etc/gitbay/config.toml=. =check-config= validates and names every
 46contradiction; =--no-host-checks= skips port/path probes.
 47=gitbayd admin config show= prints the configuration in effect as TOML,
 48every default filled in and =smtp_pass= redacted. A file that fails
 49validation still prints, followed by the contradiction.
 50
 51** [server]
 52- =root= (default =/var/lib/gitbay=) — repositories, database, host
 53  keys, ACME cache all live here.
 54- =site_url= (required) — canonical =https://host=; drives ACME, clone
 55  URLs, mail links.
 56- =source_repo= (optional, =owner/name=) — the repository this instance
 57  develops itself in. Startup warns when the running build's commit is
 58  not on that repository's default branch, which is how a binary built
 59  from an unmerged branch stops being invisible. Leave it unset unless
 60  the instance hosts its own source.
 61
 62** [ssh]
 63- =mode= — =embedded= | =system= (above).
 64- =port= (22) — embedded listener port.
 65- =host_keys= — list of private key paths; empty generates an ed25519
 66  key at =<root>/ssh/host_ed25519=.
 67
 68** [http]
 69- =addr= (=:443=), =tls= — =acme= | =files= | =off=.
 70- =acme=: certificates via TLS-ALPN-01 on the HTTPS port, cached at
 71  =<root>/acme=; =acme_email= for the CA account; =acme_http_addr=
 72  (=:80=, ="off"= to disable) adds HTTP-01 and an https redirect —
 73  failing to bind it is a warning, not fatal. Requires an =https://=
 74  site_url with a public DNS name.
 75- =files=: =cert_file= + =key_file=.
 76- =off=: plain HTTP — development, or behind a TLS-terminating proxy.
 77
 78=trusted_proxies= lists the addresses or CIDRs of reverse proxies in
 79front of the daemon. A request from one of them is attributed, for API
 80rate limiting, to the last =X-Forwarded-For= hop that is not itself a
 81trusted proxy; from anyone else the header is ignored. Empty, the
 82default, is right when gitbayd terminates TLS itself.
 83
 84** [web]
 85- =mode= — =view_only= (default) | =accounts=. In view_only the mutating
 86  web routes are never registered; in accounts, browser sessions are
 87  minted over SSH (=web login=), and users with write access can create
 88  repos, comment, and make simple file edits (which commit unsigned,
 89  honestly). =password_auth= is reserved and currently rejected.
 90- =title= — the instance's display name in the rail and page titles;
 91  empty falls back to the site host. Lower case is the convention for
 92  gitbay itself.
 93- =privacy_notice= — operator text shown on =/privacy= under the fixed
 94  statement.
 95
 96** [registration]
 97- =mode= — =closed= (default) | =invite= | =open=. invite/open require
 98  [mail]. See the user guide for the flows.
 99
100- =pending_expiry= (empty, never) — a duration such as ="168h"=; a
101  self-registered account still unverified after that long is removed,
102  hourly and at start, audited as =pending.expired=.
103
104- =notify_admin= (false) — mail every instance admin when an account
105  becomes active: an invite redeemed, or an open-mode signup that
106  verified its address. The unverified row an open signup creates is
107  not reported, because anyone can post the form and mailing on that
108  would point a flood at the admins. Recipients are the verified
109  primary addresses of active admins who have activity mail on, the
110  same rule any other notice follows, so an admin with no verified
111  address hears nothing. The notice is queued, so a dead SMTP host
112  shows up in the admin page's Mail table instead of failing the
113  registration. Requires [mail].
114
115** [mail]
116- =smtp_host= (host:port, 587 assumed), =from=, optional =smtp_user= /
117  =smtp_pass=. STARTTLS when offered. Required for invite/open
118  registration and self-service =email add=; in closed mode you may omit
119  it entirely and assert addresses by hand (below).
120
121** [api]
122- =enabled= (false) — the JSON API surface; see [[API]]. Off
123  means no credential-bearing HTTP endpoint exists at all.
124
125** [webhooks]
126- =allow_local= (false) — permit webhook targets on loopback/private
127  addresses. Leave off unless you know why you need it (SSRF).
128
129** [limits]
130- =clone_timeout= (3600s) — cap on =repo import= fetches.
131- =max_blob_bytes= (100MB) — cap on raw file serving over the web.
132- =max_asset_bytes= (512MB) — cap per uploaded release asset.
133- =max_snippet_bytes= (1MB) — cap per snippet file.
134- =max_snippets_per_user= (0, unlimited) — snippets an account may own.
135- =max_repos_per_user= (0, unlimited) — repositories an account may own
136  directly; =repo create=, =fork= and =import= refuse past it.
137  Organizations are not capped.
138- =max_bytes_per_user= (0, unlimited) — disk the account's own
139  repositories may take; a push may be no larger than what is left.
140- =max_pack_bytes=, =ssh_auth_rate= — reserved, not yet enforced.
141
142** [git_daemon]
143- =enabled= (false), =port= (9418) — the anonymous =git://= listener.
144  Serves only public repositories that additionally ran
145  =repo settings git-daemon <repo> on=.
146
147** [mirrors]
148- =pull_interval_minutes= (15) — how often pull mirrors fetch their
149  upstream. Push mirrors sync shortly after each local ref update.
150  Mirror URLs pass the same SSRF rules as webhook targets.
151
152** [go_import]
153Vanity Go module paths, one per line: ="host/module" = "owner/repo"=.
154Requests with =?go-get=1= at or under the module path answer with the
155go-import meta tag pointing at the repository's HTTPS clone URL, so
156=go install host/module/cmd/...@latest= resolves. The repository should
157be public (the module path itself confirms it exists).
158
159* Users, email, invites
160
161Every =gitbayd admin= subcommand except =backup=, =gc= and the one-shot
162backfills is a wrapper that dispatches the registry command of the same
163name as the host: an admin context with no account behind it, so its
164audit rows carry no actor and =source: host=. The same commands run in an
165instance admin's SSH session (=ssh git@<host> admin ...=) and write the
166same rows with the key fingerprint as source. One implementation, two
167credentials.
168
169#+begin_src sh
170gitbayd admin user create alice --key alice.pub --email a@example.org --verified [--admin]
171gitbayd admin email verify alice a@example.org   # admin assertion, no SMTP needed
172gitbayd admin invite --email b@example.org       # mails a code; prints it if no SMTP
173#+end_src
174
175"Verified" means SMTP-confirmed or host-admin-asserted; the database
176records which. Verified emails are what make commit signatures
177meaningful — an unverified address never produces a =verified= badge.
178
179* Audit and account control
180
181The audit log is the security feed (events are the product feed): every
182successful mutating command with its argv and source credential (SSH key
183fingerprint or API), registrations, admin actions, force-pushes, and
184auth failures/throttling. Secrets never appear — they travel on stdin,
185never in argv.
186
187#+begin_src sh
188gitbayd admin audit [--actor u|-] [--action prefix] [--since 24h|7d|date] [--limit n] [--json]
189ssh git@<host> audit ...             # the same, from an admin session
190ssh git@<host> admin user list [--state active|pending|disabled|admin]
191ssh git@<host> admin user show <name>   # keys, emails, orgs, tokens, sessions
192ssh git@<host> admin user limits <name> [--repos n|default] [--bytes n|default]   # per-account caps
193ssh git@<host> admin user promote <name>   # grant instance admin
194ssh git@<host> admin user demote <name>    # remove it; the last admin is refused
195gitbayd admin user promote <name>    # host-local: recovery when no admin key is reachable
196gitbayd admin user disable <name>    # suspend: SSH, web, API all refused;
197gitbayd admin user enable <name>     #   sessions dropped, nothing deleted
198gitbayd admin user delete <name> --yes  # only for accounts anchoring nothing:
199                                     #   refused (with each blocker named) while
200                                     #   the account owns repos, authored
201                                     #   issues/MRs/comments/reviews, or is an
202                                     #   org's only admin
203#+end_src
204
205=--actor= takes a username, or =-= for rows with no actor: host commands
206and auth failures. =--action= is a prefix, so =cmd repo= catches every
207repository command and =admin= every host or admin-session action.
208=--since= is a duration back from now (=30m=, =24h=, =7d=) or a date.
209
210=admin user list= pages by username (=--limit=, =--cursor=) and carries
211each account's state and =last_seen=, the newest use of any of its SSH
212keys or API tokens. =admin user show= adds the keys with their last use,
213each address with how it was verified, PGP keys, org roles, the owned
214repository count, API token names, and live browser sessions. Both are
215Both are refused to non-admins, like =audit=, on every surface.
216
217=/admin/users= is the same list in a browser, linked from the admin
218page: the state filter the command takes, keyset paging on its cursor,
219and a row per account with promote, demote, disable and enable, each
220dispatching the command. Demote and disable ask for the username to be
221typed, since both take someone's access away. Creating and deleting an
222account, issuing an invite and asserting an address stay on the command
223line: each takes a key, mints a credential, or cannot be undone. A
224non-admin gets the 404 a missing page would, so the URL confirms
225nothing.
226
227Promotion needs an active account: a pending or disabled one is refused.
228Demotion is refused when it would leave no admin, over SSH, in the
229browser and on the host alike, so the host-local =promote= is the way
230back in when the only admin key is lost.
231
232Instance admin carries no right on anyone's repository: policy does not
233consult it, and a private repository still answers not-found to an
234admin. Moderation goes through explicit overrides that skip the access
235check and write their own audit row:
236
237#+begin_src sh
238ssh git@<host> admin repo list [--owner o] [--visibility public|private]  # size, last push
239ssh git@<host> admin repo archive|unarchive <owner/name>
240ssh git@<host> admin repo visibility <owner/name> public|private
241ssh git@<host> admin repo delete <owner/name> --yes
242#+end_src
243
244Each lands in the audit log as =admin repo.<action>= naming the
245repository, on top of the =cmd= row every mutating command gets.
246
247=limits.ssh_auth_rate= (10) throttles per-IP authentication *failures*
248per minute — successful auths never count and clear the slate.
249=limits.max_pack_bytes= is enforced as =receive.maxInputSize= on every
250push.
251
252=limits.write_rate= (60) bounds *mutating commands per account per
253minute*. It is counted in the dispatcher, so SSH, the JSON API and the
254web spend one budget and a caller cannot refresh it by changing surface;
255=limits.api_rate= stays in front of it, bounding a network source rather
256than an account. A command is one token whatever it writes, so a bundle
257import costs one and only a loop of separate commands spends the budget.
258Read-only commands, the runner protocol (a build streams its log in many
259small writes) and the host CLI are exempt. Refusals exit 4 and say when
260to retry. A negative value turns the limit off; it matters most with
261=registration = "open"=, where every write also queues notification mail
262and webhook deliveries.
263
264* Queues
265
266Every background worker keeps a backlog and a failure state. An instance
267admin reads them all in one place:
268
269#+begin_src sh
270gitbay dashboard --json | jq .queues   # webhooks, mail, mirrors, builds, deps
271#+end_src
272
273Per worker: pending, retrying (pending with a failed attempt) and
274dead-lettered counts with the oldest pending age, and the retrying or
275failed rows themselves, capped at twenty each. Builds list what is
276running and then what is pending, each since when, so a build no runner
277is scoped to claim is visible here rather than only in its repository;
278mirrors list the ones whose last sync failed;
279dependency checks list the ones whose last check errored. Non-admins get
280no =queues= key at all.
281
282In accounts mode the same read renders at =/admin=, linked from the rail
283for admins. Anyone else gets a 404 there.
284
285A dead-lettered mail is logged as =notification dead-lettered mail=<id>=,
286with the address redacted out of the relay's error. The id is the queue
287row: find it in the Mail table on =/admin=, or in =dashboard --json=,
288where the recipient and the unredacted error are. That is deliberate —
289see the Threat-Model page.
290
291* Maintenance
292
293#+begin_src sh
294gitbayd admin stats [--json]         # counts, database size, per-repo disk
295ssh git@<host> admin stats [--json]  # the same, from an admin session
296gitbayd admin gc [--repo owner/name] # git gc: repack and prune; per-repo sizes
297gitbayd admin gc --aggressive        # thorough repack; slow, rarely needed
298gitbayd admin gc --lfs               # also drop LFS objects no pointer names (older than a day)
299#+end_src
300
301A history rewrite leaves the commits it removed reachable through
302=refs/merge-requests/N/head= of the merge requests that landed them, so
303they stay fetchable by anyone who can read the repository. Nothing drops
304a head ref on its own — an open or source-gone MR is merged through it,
305and a merged or closed one keeps its diff readable through it — so the
306cleanup is a command an instance admin runs, naming the MRs:
307
308#+begin_src sh
309ssh git@<host> admin mr prune owner/name 1 2 3 --yes
310#+end_src
311
312It refuses an open or source-gone MR, deletes the named refs, runs
313=git gc --prune=now= on that one repository so the objects go at once
314rather than after git's two-week grace, leaves a system comment on each
315MR, and audits as =admin mr.prune=. The MR keeps its title, comments,
316reviews and head sha; =mr diff= and the MR page say the head is gone.
317Run it when nothing is pushing to that repository: without the grace, a
318push caught between leaving quarantine and writing its ref loses its
319objects. Objects also survive in offsite backups until those are
320rewritten; see "Removing a repository's history from every snapshot".
321
322=deploy/cloud-init.yaml= ships a =gitbay-gc.timer= that runs =admin gc=
323weekly (Sunday 07:00 UTC). Imported repositories keep whatever pack
324layout the source sent, so a first manual =admin gc= after a bulk
325import is worthwhile.
326
327* Backup and restore
328
329#+begin_src sh
330gitbayd admin backup --out /var/backups/gitbay/backup.tar.gz
331gitbayd admin backup --verify /var/backups/gitbay/backup.tar.gz   # read it back
332#+end_src
333
334One archive: a consistent SQLite snapshot (taken *before* the
335repositories are read, so the database never references objects the
336archive missed), every repository, and the SSH host keys. Excluded:
337hook socket, regenerated hook scripts, WAL files. Safe to run against a
338live daemon.
339
340=--verify= reads an archive back: the snapshot must pass SQLite's
341integrity check, and every repository the snapshot names must be in the
342archive. A database-only archive is checked for integrity and says so.
343Exit is non-zero on damage or a missing repository.
344
345Restore: extract into an empty directory, point =server.root= at it,
346start gitbayd. Host keys are preserved, so clients keep their
347known_hosts entries; hooks regenerate at startup.
348
349** Schedule and recovery point
350
351Two timers, because the two halves of the data have different exposure.
352
353- =gitbay-backup.timer=, nightly. The full archive above, last 7 kept.
354- =gitbay-db-backup.timer=, hourly. =admin backup --db-only=, which
355  writes the SQLite snapshot alone, last 48 kept. A few MB against the
356  full archive's hundreds, which is what makes the frequency affordable.
357
358The split follows what a loss would actually cost. Repositories are git,
359so a mirror or any clone is a second copy; the database is the only copy
360of issues, merge requests, comments and review state. So the recovery
361point is about an hour for the data that exists nowhere else, and a day
362for the data that does.
363
364Continuous replication (litestream and similar) was considered and not
365adopted. It would take the database's recovery point to seconds, but the
366repositories would still be on the nightly archive, so a restore could
367produce a database referencing commits the repository backup does not
368have. Consistency between the two halves is worth more here than latency
369on one of them. Revisit if repository replication becomes continuous
370too.
371
372** Offsite copies
373
374bay1 also takes a nightly restic snapshot of =/var/lib/gitbay= and
375=/var/lib/gitbay-stage= (the staged database copy) to an S3 bucket at
376Scaleway, with a key that can only add snapshots. The key that can
377remove them lives on the operator's machine, in
378=~/.config/gitbay/offsite.env=, and never on bay1: a compromised host
379cannot destroy its own history. Forgetting, pruning and rewriting all
380run from there.
381
382*** Removing a repository's history from every snapshot
383
384A history rewrite plus =admin mr prune= takes commits off the server,
385but every snapshot taken before it still holds them, and the retention
386window is the only thing that ages them out. To remove them now,
387rewrite the snapshots without that repository rather than forgetting
388the snapshots: everything else in them stays restorable. The next
389nightly run adds the repository back in its current state.
390
391Repositories are stored under the name they had on disk when each
392snapshot was taken, so a renamed repository needs every name it has
393carried. Check what an older snapshot holds before choosing the paths:
394
395#+begin_src sh
396set -a; . ~/.config/gitbay/offsite.env; set +a
397restic $RESTIC_OPTS snapshots
398restic $RESTIC_OPTS ls <old-snapshot> /var/lib/gitbay/repos/<owner>
399#+end_src
400
401Then dry-run, apply, prune, and confirm nothing matches:
402
403#+begin_src sh
404EXCL="--exclude /var/lib/gitbay/repos/<owner>/<name>.git --exclude /var/lib/gitbay/repos/<owner>/<old-name>.git"
405restic $RESTIC_OPTS rewrite --dry-run $EXCL     # "would modify N snapshots"
406restic $RESTIC_OPTS rewrite --forget $EXCL      # new snapshots replace the originals
407restic $RESTIC_OPTS prune                       # drops the data nothing references
408restic $RESTIC_OPTS find <name>.git <old-name>.git   # expect no output
409#+end_src
410
411=--forget= is what makes the originals go; without it the rewritten
412snapshots sit beside them and the data stays referenced. Snapshot IDs
413change; their times do not. Done for krz/keycask (formerly rust-pass)
414on 2026-09-18, across 22 snapshots.
415
416* Upgrades
417
418Replace the binary, restart the unit. Migrations apply automatically and
419are transactional; hook scripts under =<root>/hooks= are rewritten at
420startup to point at the current binary path.
421
422* CI runner
423
424=gitbay-runner= executes builds queued by pushes and merge requests. It
425polls over SSH with a key of scope =runner=, which reaches only the
426runner protocol and read-only git (a runner executes arbitrary
427repository code, so the key it holds must not do more). A runner key
428claims builds only for the repositories it is attached to, by =repo
429runner add= from a repository admin or an instance admin; an admin key
430claims any. Users attach their own runners: see the Users page. For an
431instance runner, run it as a dedicated unprivileged user on a non-admin
432account. =admin user create --key= registers a full-scope key, so the
433runner key is added afterwards through a bootstrap key that is then
434removed, and attached to each repository it should build:
435
436#+begin_src sh
437useradd --system --create-home --home-dir /var/lib/gitbay-runner ci-runner
438sudo -u ci-runner ssh-keygen -t ed25519 -N "" -f /var/lib/gitbay-runner/.ssh/id_ed25519
439ssh-keygen -t ed25519 -N "" -f /tmp/ci-bootstrap
440gitbayd --config /etc/gitbay/config.toml admin user create ci --key /tmp/ci-bootstrap.pub
441ssh -i /tmp/ci-bootstrap git@127.0.0.1 keys add --scope runner < /var/lib/gitbay-runner/.ssh/id_ed25519.pub
442ssh -i /tmp/ci-bootstrap git@127.0.0.1 keys remove "$(ssh-keygen -lf /tmp/ci-bootstrap.pub | awk '{print $2}')"
443rm /tmp/ci-bootstrap /tmp/ci-bootstrap.pub
444gitbay-runner -remote git@127.0.0.1 -workdir /var/lib/gitbay-runner/work
445#+end_src
446
447#+begin_src sh
448gitbay repo runner add krz/site < /var/lib/gitbay-runner/.ssh/id_ed25519.pub
449#+end_src
450
451=-jobs N= runs N builds at once. Claiming is one transaction that
452selects and updates, and each build works in its own =build-<id>=
453directory, so workers do not collide; idle polls are staggered across
454the interval so N of them do not wake together. The drop-in's weights
455below are per service, not per build, so raising =-jobs= divides them
456rather than multiplying the host's load.
457
458=admin runners= shows which account each runner polls as, and what each
459may claim. A runner key claims builds only for the repositories it is
460attached to: with none attached it claims nothing, and =-repos= may only
461narrow within them. An admin's full-scope key claims any repository —
462that is what =-repos= was for — and still works for the protocol during
463a rotation. A merge request head from a fork is built in the target
464repository as untrusted: the claim carries no secrets, and only a runner
465started with =-untrusted= takes it. Same-repository heads were built by
466their branch push and are not built again.
467
468=make deploy-runner= also installs
469=deploy/gitbay-runner.override.conf= as a systemd drop-in: =Nice=10=,
470=CPUWeight=30=, =IOWeight=30=, so a build never starves the host's sshd,
471the daemon or the backup timers, and =NoNewPrivileges=,
472=ProtectSystem=full=, =ProtectKernelTunables=, =ProtectControlGroups=
473and =RestrictSUIDSGID=, so a step cannot reach outside its workspace
474and the runner's home. The e2e suite alone starts sixty
475daemon instances; without the drop-in a deploy's copy over the admin
476sshd stalled. Both deploy targets copy with =rsync --partial=, which
477resumes a stalled transfer.
478
479Under =-isolation podman=, the default and what bay1 runs, each build
480is confined to a container (see Container isolation below). Under
481=-isolation none= steps run directly on the host as the runner's user,
482so treat that machine as executing whatever your users push, and
483install the toolchains your builds need on it.
484
485A runner claims the oldest pending build among the repositories its key
486is attached to — for an admin key, the oldest in the instance. =-repos=
487narrows within that set, which is what makes a runner outside the server
488practical: one on a machine that should build a single project, or that
489holds credentials for one deployment, stays on it.
490
491Oldest-first is across everything the key may claim, so a repository
492with a deep queue holds every other repository the same runner serves;
493bay1 measured a 15-minute average wait on a day of merge request
494stacks from one repository. A runner attached to one repository cannot
495be starved. That is the rule, decided in krz/gitbay#207: a runner
496serving several repositories takes them oldest-first, and an operator
497who wants one repository never to wait on another runs a second
498runner attached to it alone. Nothing caps what an account queues,
499and nothing needs to (krz/gitbay#206): a schedule tick queues nothing
500while the job's last build is pending or running, so a repository
501with no runner holds one row per scheduled job rather than one per
502tick, and a build runs only on a runner its owner attaches, so a busy
503schedule spends the owner's compute. Pushes are bounded by what an
504account can push.
505
506#+begin_src sh
507gitbay-runner -remote git@gitbay.org -repos krz/site,krz/docs \
508  -workdir /var/lib/gitbay-runner/work
509#+end_src
510
511Add =-untrusted= only with =-isolation podman=.
512
513gitbay.org's runner is attached to the forge's own repositories and
514the isolation canary, nothing else, because it shares the host with
515the forge; its unit names no =-repos=, the attachments are the
516boundary. Any other repository builds on a runner its owner attaches.
517
518=-repos= narrows an admin runner; for a runner key the attachments are
519the boundary, held by the server, and =-repos= may only name
520repositories among them. =-untrusted= makes a runner claim merge
521request heads from forks; the bay1 unit sets it because it isolates in
522podman. A runner without it builds trusted commits only.
523
524=gitbay dashboard= and =ssh git@<host> admin runners= list every key
525that has polled as a runner: the account, the key's fingerprint, when it
526last polled, the repositories it may claim — its attachments for a runner
527key, the =-repos= it asked for or =any= for an admin key — and the build
528it holds; =admin runners remove <fingerprint>= (=forget= until the next release) drops the row for a key
529that polled by mistake, the key itself untouched. =admin runners= also
530heads the list with the queue: builds
531pending now, and over the last day how many were claimed, how long they
532waited to be claimed (average and worst), and
533how many the reaper ended instead of a runner reporting them. A build a runner claimed and never
534reported is failed by the scheduler's minute tick, whether or not any
535runner is still alive: within about two minutes of its log stream ending
536with no outcome reported — the runner reports right after closing the
537stream, retrying for half a minute if gitbayd is unreachable — or, if no
538stream was ever seen, at the deadline.
539
540Instance admin on the runner account only authorizes the claim/report
541protocol; it grants no repo access. A build that pushes back — a pages
542deploy, an archive publish, an automated MR branch — needs an explicit
543grant on that repo: =repo access grant <owner/name> ci write=. Private
544repos likewise need at least read for the clone.
545
546** Container isolation
547
548Builds run in a rootless podman container, one per job, with the
549workspace bind mounted and nothing else: the clone happens outside with
550the runner's key, so a step cannot read it. =-isolation none= keeps the
551old behaviour — steps on the host as the runner's user — for an instance
552where every repository is trusted. There is no automatic fallback: a
553runner started with =-isolation podman= that cannot find a working
554podman exits rather than running a build unsandboxed.
555
556gitbay's own jobs name =localhost/gitbay-ci:2=, built from
557=deploy/Containerfile.ci= on the runner host. A job's image must carry
558what its steps need: the suite drives real git, git-lfs, gpg and sshd and
559asserts they exist before running, so the stock runner default would fail
560it immediately. Build or rebuild it with:
561
562#+begin_src sh
563ssh -p 2222 root@<host> 'cat > /tmp/Containerfile.ci' < deploy/Containerfile.ci
564ssh -p 2222 root@<host> 'su - ci-runner -s /bin/sh -c \
565  "podman build -t localhost/gitbay-ci:2 -f /tmp/Containerfile.ci /tmp"'
566#+end_src
567
568The tag is deliberate rather than =:latest=: changing the file means
569bumping the tag in =.gitbay/ci.yml=, so a running branch's image does not
570change under it.
571
572=-image= names the image a job runs in when it declares none, and is
573required under =-isolation podman=: there is no built-in default,
574because an image this host does not have would fail every build. A job
575overrides it with =image:= in =.gitbay/ci.yml=, validated as a reference
576so a config file cannot turn it into podman arguments.
577
578=-cpus= and =-memory= cap one build (podman's units, e.g. =-cpus 2
579-memory 4g=); unset means uncapped. The runner applies them itself: it
580creates a cgroup per build under its own delegated service cgroup,
581writes the limits there, and starts every podman process for the build
582inside it, with podman's cgroup handling off. Podman's own =--memory=
583and =--cpus= never applied under rootless cgroupfs, which is what a
584system service gets (krz/gitbay#188). The unit therefore needs
585=Delegate=yes=, which the drop-in sets; without it the runner refuses
586to start when a limit is set, and logs that builds run unconfined when
587none is. bay1 runs =-cpus 3 -memory 6g= per build inside =MemoryMax=6G=
588and =CPUQuota=300%= on the unit, on a 7.7GB four-core host with no
589swap: the memory cap is what keeps the forge alive when a build
590allocates without bound, and it sits above the e2e suite's 5GB peak
591rather than at a fair share. =OOMPolicy=continue= keeps systemd from
592stopping the runner when a build is OOM-killed.
593
594Each repository gets its own build home under the runner's workdir,
595mounted into its containers as =HOME=. Caches persist between builds of
596one repository and are never read by another's.
597
598*Images are provisioned, never pulled by a build.* The runner passes
599=--pull=never=. Two reasons, and the second is the better one: the
600service runs with =RestrictSUIDSGID=yes= so podman cannot unpack a layer
601holding a setuid file, which is nearly every distribution image; and on
602an instance where anyone can push a =ci.yml=, =image:= would otherwise
603mean "fetch and run anything from the internet". An operator pulls or
604builds what is allowed and a build picks among those. A job naming an
605image the host does not have fails with a message saying so.
606
607#+begin_src sh
608su - ci-runner -s /bin/sh -c "podman pull docker.io/library/alpine:3.20"
609su - ci-runner -s /bin/sh -c "podman images"
610#+end_src
611
612Prepare a host before pointing an isolating runner at it:
613
614#+begin_src sh
615ssh -p 2222 root@<host> 'sh -s' < deploy/runner-podman-setup.sh
616make deploy-runner
617#+end_src
618
619The script installs podman, delegates a subuid/subgid range to
620=ci-runner=, checks that user namespaces are enabled rather than
621assuming, enables lingering, and verifies rootless podman actually runs
622as that user. It is idempotent.
623
624The drop-in sets =NoNewPrivileges=no=, without which rootless podman
625cannot call =newuidmap= and the runner refuses to start. That is a
626considered trade, explained in the file and in the Threat-Model; if you
627run with =-isolation none=, set it back to =yes=.
628
629*Restarting the runner is safe.* On SIGTERM it stops claiming, finishes
630the build in flight, reports it, and exits; the drop-in's
631=TimeoutStopSec=50min= covers the longest build, and its =KillMode=mixed=
632is what makes the signal reach the runner alone — under systemd's default
633the build's container and the log session are signalled with it, and the
634runner drains a build that is already dead. So =make deploy-runner=
635waits for a running build rather than orphaning it, and a build's result
636is retried for half a minute if gitbayd is restarting at that moment. A
637second SIGTERM ends the runner at once, abandoning the build to the
638reaper. The suite checks all three: =TestRunnerDrainsOnSIGTERM= signals
639the process, =TestRunnerDropInLetsTheDrainHappen= reads the drop-in's
640=KillMode= and =TimeoutStopSec=, and =TestRunnerDrainUnderSystemd= runs
641the runner as a transient user unit under =systemd-run= and stops it
642under both kill modes. That last one needs a systemd user manager, so it
643skips in the container CI runs in; run it on a Linux host with
644=go test ./e2e -run TestRunnerDrainUnderSystemd -v=.
645
646*Validate podman mode on a scratch repository before pointing the runner
647at real ones.* Every deploy that switched the whole instance to
648containers and failed took CI down with it. Instead: create a throwaway
649repository the runner account can read (public, or granted read — a
650private one is "not found" to the runner and the build stays pending),
651give it one job that names the CI image, and deploy the runner with
652=-repos= naming only that repository. The production unit, with its real
653hardening, then claims nothing else; other repositories' builds queue
654until =-repos= is switched back, which is a pause, not an outage.
655
656#+begin_src sh
657gitbay repo create cmc/ci-smoke          # then push a .gitbay/ci.yml naming the image
658sed -i 's#-repos krz/gitbay #-repos cmc/ci-smoke #' /etc/systemd/system/gitbay-runner.service.d/override.conf
659systemctl daemon-reload && systemctl restart gitbay-runner
660gitbay build log cmc/ci-smoke 1         # green: switch -repos back, redeploy
661#+end_src
662
663*Do not deploy an isolating runner to a host that has not been
664prepared.* The runner is specified to refuse to start without a working
665podman rather than fall back to running builds unsandboxed — a fallback
666that silently drops isolation is worse than a stopped runner, because
667nothing surfaces it. On an unprepared host that refusal stops every
668build on the instance.
669
670The service drop-in carries =Delegate=yes= for rootless cgroup
671management and =ReadWritePaths= for podman's store under
672=/var/lib/gitbay-runner=, which =ProtectSystem=full= would otherwise
673make read-only. Those paths are prefixed =-= so they are ignored when
674absent: the drop-in installs on unprepared hosts too, and a unit that
675refused to start would stop every build.
676The nightly canary on =cmc/ci-smoke= only runs if the runner's =-repos=
677names that repository too; a scoped runner claims nothing else.
678=gitbay-runner-prune.timer= prunes unused images weekly, as the runner's
679user: rootless storage belongs to that user, and root's prune would not
680see it. An unpruned image store on a 40GB host is a slow outage.
681
682* LFS storage
683
684Objects live content-addressed under =[lfs] root= (default
685=<server.root>/lfs=); =[lfs] max_object_bytes= caps a single object
686(512MB default). Storage sits behind a small interface — an
687S3-compatible backend is a drop-in with the server proxying, and
688presigned URLs a later optimization. LFS objects do not travel with
689push mirrors (mirrors move git refs only), and gc does not yet collect
690orphaned objects.
691
692* Pages
693
694=[pages] domain = "example.site"= serves public repos' =pages= branches
695on =<owner>.<domain>=. DNS needs a wildcard record =*.<domain>= to the
696server; ACME issues per-subdomain certificates on demand (only for
697owners that exist). The domain must not be the site host or a parent of
698it — pages content runs its own scripts and must stay off the forge's
699origin.
700
701Users with repo admin claim custom domains with =repo domain add=.
702Claims activate only after a DNS TXT challenge proves control of the
703domain (=repo domain verify=, audit-logged); pending claims serve
704nothing, get no certificates, and expire after 7 days. ACME issues
705certificates only for verified hosts, so stray DNS pointed at the
706server gets nothing.
707
708* Security
709
710The [[Threat-Model]] file is the reference for what the forge
711trusts and refuses to do. Operational checklist:
712
713- *Software checks.* =deploy/audit.sh= runs =go vet=, =govulncheck=
714  (the module list is deliberately short — review it on each release),
715  and a short fuzz pass over every attacker-facing parser (pkt-line,
716  commit, SSHSIG armor, OpenPGP key, SSH tokenizer). Run it before
717  tagging a release. CI's own =vuln= job runs =govulncheck= nightly
718  against main rather than per push, because =@latest= scans today's
719  advisory database and an advisory lands without anyone pushing;
720  =build trigger krz/gitbay vuln= runs it on demand.
721- *Web responses* carry a scripts-forbidden CSP, =X-Frame-Options:
722  DENY=, =nosniff=, =no-referrer=, and HSTS when TLS is on — no
723  configuration needed.
724- *Host sandboxing.* The systemd unit in =deploy/cloud-init.yaml= runs
725  gitbayd unprivileged with =ProtectSystem=strict=, =PrivateDevices=,
726  =LockPersonality=, =MemoryDenyWriteExecute=,
727  =SystemCallFilter=@system-service=, and =RestrictAddressFamilies= to
728  INET/INET6/UNIX. It keeps =CAP_NET_BIND_SERVICE= only, to bind 22/80/443.
729- *OS patches* apply via =unattended-upgrades= (security origins,
730  auto-reboot 04:30 if required).
731- *Admin sshd (2222)* is throttled by =MaxStartups=/=MaxAuthTries= and
732  watched by =fail2ban=; gitbayd's own port 22 is throttled by
733  =limits.ssh_auth_rate= (auth failures per IP per minute), and every
734  account's writes by =limits.write_rate=.
735- *Monitoring.* =gitbay-monitor.timer= writes a reading hourly to
736journald and, when =/etc/gitbay/monitor.url= exists, posts it to that
737webhook: disk, service, the daemon's own =/healthz= answer, certificate
738expiry, and the age of the newest full backup and database snapshot.
739It exits non-zero on an alert so the unit shows in =systemctl
740--failed=: a stopped service, =/healthz= not answering =ok=, disk ≥ 85%,
741a certificate under 21 days, a full backup over 25 hours old, or a
742database snapshot over 2 hours old.
743
744=GET /healthz= is unauthenticated and cache-free: whether the database
745answers and which commit serves, 503 when it does not.
746- *Database.* =gitbay.db= and its WAL live under =/var/lib/gitbay= (mode
747  0750, owned by =gitbay=). The nightly archive plus provider snapshots
748  are the recovery path; for tighter RPO, add continuous replication
749  (litestream) against the same file — it coexists with the WAL.
750
751* Odds and ends
752
753- deleting a fork marks MRs sourced from it =source_gone=; their diffs
754  remain viewable and mergeable because the target repo owns the
755  objects.
756- =refs/merge-requests/*= is server-owned and unpushable by clients;
757  only =admin mr prune= removes one (see Maintenance).
758- audit-relevant activity (issue/MR lifecycle, imports, pushes) lands in
759  the =events= table, which also feeds webhooks.
760- the daemon idles under 10MB RSS; the smallest VPS tier is adequate.