.gitbay/wiki/Admin.org

v1.23.0
gitbay/.gitbay/wiki/Admin.org rendered · source · history · blame · raw

673 lines · 33877 bytes

  1#+title: gitbay admin guide
  2
  3One static binary (=gitbayd=), one SQLite file, bare repositories on
  4disk, and the system =git=. Schema migrations run automatically on
  5startup and on every admin command.
  6
  7* Install
  8
  9Build from source (=go build ./cmd/gitbayd=), install via the vanity
 10module path (=go install gitbay.org/gitbay/cmd/gitbayd@latest=), or use
 11a release build: =deploy/release.sh <tag>= cross-compiles reproducible
 12linux/amd64, linux/arm64, and darwin/arm64 binaries with a SHA256SUMS
 13manifest (CGO off, trimpath, stripped — byte-identical per commit and
 14toolchain).
 15
 16#+begin_src sh
 17install -m 755 gitbayd /usr/local/bin/
 18adduser --system --group --home /var/lib/gitbay --shell /usr/sbin/nologin gitbay
 19install -d -o gitbay -g gitbay -m 750 /var/lib/gitbay
 20gitbayd --config /etc/gitbay/config.toml check-config
 21#+end_src
 22
 23=deploy/= in the source tree has a cloud-init file, a hardened systemd
 24unit, and a nightly backup timer. Run as the unprivileged =gitbay= user;
 25the unit's =AmbientCapabilities=CAP_NET_BIND_SERVICE= covers ports
 2622/80/443 without root.
 27
 28** The SSH port decision
 29
 30- =ssh.mode = "embedded"= (default): gitbayd itself listens, normally on
 31  22 — move the host's admin sshd to another port. Remotes read
 32  =git@host:owner/repo= with no port gymnastics.
 33- =ssh.mode = "system"=: the host sshd owns 22 and invokes gitbayd via
 34  =AuthorizedKeysCommand=:
 35  #+begin_example
 36  AuthorizedKeysCommand /usr/local/bin/gitbayd --config /etc/gitbay/config.toml authorized-keys %t %k
 37  AuthorizedKeysCommandUser gitbay
 38  #+end_example
 39  sshd requires that binary to be root-owned and not group/world
 40  writable. Unknown keys fail authentication inside sshd, so system mode
 41  requires =registration.mode = "closed"= (check-config enforces this).
 42
 43* Configuration reference
 44
 45=/etc/gitbay/config.toml=. =check-config= validates and names every
 46contradiction; =--no-host-checks= skips port/path probes.
 47=gitbayd admin config show= prints the configuration in effect as TOML,
 48every default filled in and =smtp_pass= redacted. A file that fails
 49validation still prints, followed by the contradiction.
 50
 51** [server]
 52- =root= (default =/var/lib/gitbay=) — repositories, database, host
 53  keys, ACME cache all live here.
 54- =site_url= (required) — canonical =https://host=; drives ACME, clone
 55  URLs, mail links.
 56- =source_repo= (optional, =owner/name=) — the repository this instance
 57  develops itself in. Startup warns when the running build's commit is
 58  not on that repository's default branch, which is how a binary built
 59  from an unmerged branch stops being invisible. Leave it unset unless
 60  the instance hosts its own source.
 61
 62** [ssh]
 63- =mode= — =embedded= | =system= (above).
 64- =port= (22) — embedded listener port.
 65- =host_keys= — list of private key paths; empty generates an ed25519
 66  key at =<root>/ssh/host_ed25519=.
 67
 68** [http]
 69- =addr= (=:443=), =tls= — =acme= | =files= | =off=.
 70- =acme=: certificates via TLS-ALPN-01 on the HTTPS port, cached at
 71  =<root>/acme=; =acme_email= for the CA account; =acme_http_addr=
 72  (=:80=, ="off"= to disable) adds HTTP-01 and an https redirect —
 73  failing to bind it is a warning, not fatal. Requires an =https://=
 74  site_url with a public DNS name.
 75- =files=: =cert_file= + =key_file=.
 76- =off=: plain HTTP — development, or behind a TLS-terminating proxy.
 77
 78=trusted_proxies= lists the addresses or CIDRs of reverse proxies in
 79front of the daemon. A request from one of them is attributed, for API
 80rate limiting, to the last =X-Forwarded-For= hop that is not itself a
 81trusted proxy; from anyone else the header is ignored. Empty, the
 82default, is right when gitbayd terminates TLS itself.
 83
 84** [web]
 85- =mode= — =view_only= (default) | =accounts=. In view_only the mutating
 86  web routes are never registered; in accounts, browser sessions are
 87  minted over SSH (=web login=), and users with write access can create
 88  repos, comment, and make simple file edits (which commit unsigned,
 89  honestly). =password_auth= is reserved and currently rejected.
 90- =title= — the instance's display name in the rail and page titles;
 91  empty falls back to the site host. Lower case is the convention for
 92  gitbay itself.
 93- =privacy_notice= — operator text shown on =/privacy= under the fixed
 94  statement.
 95
 96** [registration]
 97- =mode= — =closed= (default) | =invite= | =open=. invite/open require
 98  [mail]. See the user guide for the flows.
 99
100- =pending_expiry= (empty, never) — a duration such as ="168h"=; a
101  self-registered account still unverified after that long is removed,
102  hourly and at start, audited as =pending.expired=.
103
104** [mail]
105- =smtp_host= (host:port, 587 assumed), =from=, optional =smtp_user= /
106  =smtp_pass=. STARTTLS when offered. Required for invite/open
107  registration and self-service =email add=; in closed mode you may omit
108  it entirely and assert addresses by hand (below).
109
110** [api]
111- =enabled= (false) — the JSON API surface; see [[API]]. Off
112  means no credential-bearing HTTP endpoint exists at all.
113
114** [webhooks]
115- =allow_local= (false) — permit webhook targets on loopback/private
116  addresses. Leave off unless you know why you need it (SSRF).
117
118** [limits]
119- =clone_timeout= (3600s) — cap on =repo import= fetches.
120- =max_blob_bytes= (100MB) — cap on raw file serving over the web.
121- =max_asset_bytes= (512MB) — cap per uploaded release asset.
122- =max_snippet_bytes= (1MB) — cap per snippet file.
123- =max_snippets_per_user= (0, unlimited) — snippets an account may own.
124- =max_repos_per_user= (0, unlimited) — repositories an account may own
125  directly; =repo create=, =fork= and =import= refuse past it.
126  Organizations are not capped.
127- =max_bytes_per_user= (0, unlimited) — disk the account's own
128  repositories may take; a push may be no larger than what is left.
129- =max_pack_bytes=, =ssh_auth_rate= — reserved, not yet enforced.
130
131** [git_daemon]
132- =enabled= (false), =port= (9418) — the anonymous =git://= listener.
133  Serves only public repositories that additionally ran
134  =repo settings git-daemon <repo> on=.
135
136** [mirrors]
137- =pull_interval_minutes= (15) — how often pull mirrors fetch their
138  upstream. Push mirrors sync shortly after each local ref update.
139  Mirror URLs pass the same SSRF rules as webhook targets.
140
141** [go_import]
142Vanity Go module paths, one per line: ="host/module" = "owner/repo"=.
143Requests with =?go-get=1= at or under the module path answer with the
144go-import meta tag pointing at the repository's HTTPS clone URL, so
145=go install host/module/cmd/...@latest= resolves. The repository should
146be public (the module path itself confirms it exists).
147
148* Users, email, invites
149
150Every =gitbayd admin= subcommand except =backup=, =gc= and the one-shot
151backfills is a wrapper that dispatches the registry command of the same
152name as the host: an admin context with no account behind it, so its
153audit rows carry no actor and =source: host=. The same commands run in an
154instance admin's SSH session (=ssh git@<host> admin ...=) and write the
155same rows with the key fingerprint as source. One implementation, two
156credentials.
157
158#+begin_src sh
159gitbayd admin user create alice --key alice.pub --email a@example.org --verified [--admin]
160gitbayd admin email verify alice a@example.org   # admin assertion, no SMTP needed
161gitbayd admin invite --email b@example.org       # mails a code; prints it if no SMTP
162#+end_src
163
164"Verified" means SMTP-confirmed or host-admin-asserted; the database
165records which. Verified emails are what make commit signatures
166meaningful — an unverified address never produces a =verified= badge.
167
168* Audit and account control
169
170The audit log is the security feed (events are the product feed): every
171successful mutating command with its argv and source credential (SSH key
172fingerprint or API), registrations, admin actions, force-pushes, and
173auth failures/throttling. Secrets never appear — they travel on stdin,
174never in argv.
175
176#+begin_src sh
177gitbayd admin audit [--actor u|-] [--action prefix] [--since 24h|7d|date] [--limit n] [--json]
178ssh git@<host> audit ...             # the same, from an admin session (SSH only)
179ssh git@<host> admin user list [--state active|pending|disabled|admin]
180ssh git@<host> admin user show <name>   # keys, emails, orgs, tokens, sessions
181ssh git@<host> admin user limits <name> [--repos n|default] [--bytes n|default]   # per-account caps
182ssh git@<host> admin user promote <name>   # grant instance admin
183ssh git@<host> admin user demote <name>    # remove it; the last admin is refused
184gitbayd admin user promote <name>    # host-local: recovery when no admin key is reachable
185gitbayd admin user disable <name>    # suspend: SSH, web, API all refused;
186gitbayd admin user enable <name>     #   sessions dropped, nothing deleted
187gitbayd admin user delete <name> --yes  # only for accounts anchoring nothing:
188                                     #   refused (with each blocker named) while
189                                     #   the account owns repos, authored
190                                     #   issues/MRs/comments/reviews, or is an
191                                     #   org's only admin
192#+end_src
193
194=--actor= takes a username, or =-= for rows with no actor: host commands
195and auth failures. =--action= is a prefix, so =cmd repo= catches every
196repository command and =admin= every host or admin-session action.
197=--since= is a duration back from now (=30m=, =24h=, =7d=) or a date.
198
199=admin user list= pages by username (=--limit=, =--cursor=) and carries
200each account's state and =last_seen=, the newest use of any of its SSH
201keys or API tokens. =admin user show= adds the keys with their last use,
202each address with how it was verified, PGP keys, org roles, the owned
203repository count, API token names, and live browser sessions. Both are
204SSH-only and refused to non-admins, like =audit=.
205
206Promotion needs an active account: a pending or disabled one is refused.
207Demotion is refused when it would leave no admin, over SSH and on the
208host alike, so the host-local =promote= is the way back in when the only
209admin key is lost.
210
211Instance admin carries no right on anyone's repository: policy does not
212consult it, and a private repository still answers not-found to an
213admin. Moderation goes through explicit overrides that skip the access
214check and write their own audit row:
215
216#+begin_src sh
217ssh git@<host> admin repo list [--owner o] [--visibility public|private]  # size, last push
218ssh git@<host> admin repo archive|unarchive <owner/name>
219ssh git@<host> admin repo visibility <owner/name> public|private
220ssh git@<host> admin repo delete <owner/name> --yes
221#+end_src
222
223Each lands in the audit log as =admin repo.<action>= naming the
224repository, on top of the =cmd= row every mutating command gets.
225
226=limits.ssh_auth_rate= (10) throttles per-IP authentication *failures*
227per minute — successful auths never count and clear the slate.
228=limits.max_pack_bytes= is enforced as =receive.maxInputSize= on every
229push.
230
231=limits.write_rate= (60) bounds *mutating commands per account per
232minute*. It is counted in the dispatcher, so SSH, the JSON API and the
233web spend one budget and a caller cannot refresh it by changing surface;
234=limits.api_rate= stays in front of it, bounding a network source rather
235than an account. A command is one token whatever it writes, so a bundle
236import costs one and only a loop of separate commands spends the budget.
237Read-only commands, the runner protocol (a build streams its log in many
238small writes) and the host CLI are exempt. Refusals exit 4 and say when
239to retry. A negative value turns the limit off; it matters most with
240=registration = "open"=, where every write also queues notification mail
241and webhook deliveries.
242
243* Queues
244
245Every background worker keeps a backlog and a failure state. An instance
246admin reads them all in one place:
247
248#+begin_src sh
249gitbay dashboard --json | jq .queues   # webhooks, mail, mirrors, builds, deps
250#+end_src
251
252Per worker: pending, retrying (pending with a failed attempt) and
253dead-lettered counts with the oldest pending age, and the retrying or
254failed rows themselves, capped at twenty each. Builds list what is
255running and then what is pending, each since when, so a build no runner
256is scoped to claim is visible here rather than only in its repository;
257mirrors list the ones whose last sync failed;
258dependency checks list the ones whose last check errored. Non-admins get
259no =queues= key at all.
260
261In accounts mode the same read renders at =/admin=, linked from the rail
262for admins. Anyone else gets a 404 there.
263
264A dead-lettered mail is logged as =notification dead-lettered mail=<id>=,
265with the address redacted out of the relay's error. The id is the queue
266row: find it in the Mail table on =/admin=, or in =dashboard --json=,
267where the recipient and the unredacted error are. That is deliberate —
268see the Threat-Model page.
269
270* Maintenance
271
272#+begin_src sh
273gitbayd admin stats [--json]         # counts, database size, per-repo disk
274ssh git@<host> admin stats [--json]  # the same, from an admin session
275gitbayd admin gc [--repo owner/name] # git gc: repack and prune; per-repo sizes
276gitbayd admin gc --aggressive        # thorough repack; slow, rarely needed
277gitbayd admin gc --lfs               # also drop LFS objects no pointer names (older than a day)
278#+end_src
279
280=deploy/cloud-init.yaml= ships a =gitbay-gc.timer= that runs =admin gc=
281weekly (Sunday 07:00 UTC). Imported repositories keep whatever pack
282layout the source sent, so a first manual =admin gc= after a bulk
283import is worthwhile.
284
285* Backup and restore
286
287#+begin_src sh
288gitbayd admin backup --out /var/backups/gitbay/backup.tar.gz
289gitbayd admin backup --verify /var/backups/gitbay/backup.tar.gz   # read it back
290#+end_src
291
292One archive: a consistent SQLite snapshot (taken *before* the
293repositories are read, so the database never references objects the
294archive missed), every repository, and the SSH host keys. Excluded:
295hook socket, regenerated hook scripts, WAL files. Safe to run against a
296live daemon.
297
298=--verify= reads an archive back: the snapshot must pass SQLite's
299integrity check, and every repository the snapshot names must be in the
300archive. A database-only archive is checked for integrity and says so.
301Exit is non-zero on damage or a missing repository.
302
303Restore: extract into an empty directory, point =server.root= at it,
304start gitbayd. Host keys are preserved, so clients keep their
305known_hosts entries; hooks regenerate at startup.
306
307** Schedule and recovery point
308
309Two timers, because the two halves of the data have different exposure.
310
311- =gitbay-backup.timer=, nightly. The full archive above, last 7 kept.
312- =gitbay-db-backup.timer=, hourly. =admin backup --db-only=, which
313  writes the SQLite snapshot alone, last 48 kept. A few MB against the
314  full archive's hundreds, which is what makes the frequency affordable.
315
316The split follows what a loss would actually cost. Repositories are git,
317so a mirror or any clone is a second copy; the database is the only copy
318of issues, merge requests, comments and review state. So the recovery
319point is about an hour for the data that exists nowhere else, and a day
320for the data that does.
321
322Continuous replication (litestream and similar) was considered and not
323adopted. It would take the database's recovery point to seconds, but the
324repositories would still be on the nightly archive, so a restore could
325produce a database referencing commits the repository backup does not
326have. Consistency between the two halves is worth more here than latency
327on one of them. Revisit if repository replication becomes continuous
328too.
329
330* Upgrades
331
332Replace the binary, restart the unit. Migrations apply automatically and
333are transactional; hook scripts under =<root>/hooks= are rewritten at
334startup to point at the current binary path.
335
336* CI runner
337
338=gitbay-runner= executes builds queued by pushes and merge requests. It
339polls over SSH with a key of scope =runner=, which reaches only the
340runner protocol and read-only git (a runner executes arbitrary
341repository code, so the key it holds must not do more). A runner key
342claims builds only for the repositories it is attached to, by =repo
343runner add= from a repository admin or an instance admin; an admin key
344claims any. Users attach their own runners: see the Users page. For an
345instance runner, run it as a dedicated unprivileged user on a non-admin
346account. =admin user create --key= registers a full-scope key, so the
347runner key is added afterwards through a bootstrap key that is then
348removed, and attached to each repository it should build:
349
350#+begin_src sh
351useradd --system --create-home --home-dir /var/lib/gitbay-runner ci-runner
352sudo -u ci-runner ssh-keygen -t ed25519 -N "" -f /var/lib/gitbay-runner/.ssh/id_ed25519
353ssh-keygen -t ed25519 -N "" -f /tmp/ci-bootstrap
354gitbayd --config /etc/gitbay/config.toml admin user create ci --key /tmp/ci-bootstrap.pub
355ssh -i /tmp/ci-bootstrap git@127.0.0.1 keys add --scope runner < /var/lib/gitbay-runner/.ssh/id_ed25519.pub
356ssh -i /tmp/ci-bootstrap git@127.0.0.1 keys remove "$(ssh-keygen -lf /tmp/ci-bootstrap.pub | awk '{print $2}')"
357rm /tmp/ci-bootstrap /tmp/ci-bootstrap.pub
358gitbay-runner -remote git@127.0.0.1 -workdir /var/lib/gitbay-runner/work
359#+end_src
360
361#+begin_src sh
362gitbay repo runner add krz/site < /var/lib/gitbay-runner/.ssh/id_ed25519.pub
363#+end_src
364
365=-jobs N= runs N builds at once. Claiming is one transaction that
366selects and updates, and each build works in its own =build-<id>=
367directory, so workers do not collide; idle polls are staggered across
368the interval so N of them do not wake together. The drop-in's weights
369below are per service, not per build, so raising =-jobs= divides them
370rather than multiplying the host's load.
371
372=admin runners= shows which account each runner polls as, and what each
373may claim. A runner key claims builds only for the repositories it is
374attached to: with none attached it claims nothing, and =-repos= may only
375narrow within them. An admin's full-scope key claims any repository —
376that is what =-repos= was for — and still works for the protocol during
377a rotation. A merge request head from a fork is built in the target
378repository as untrusted: the claim carries no secrets, and only a runner
379started with =-untrusted= takes it. Same-repository heads were built by
380their branch push and are not built again.
381
382=make deploy-runner= also installs
383=deploy/gitbay-runner.override.conf= as a systemd drop-in: =Nice=10=,
384=CPUWeight=30=, =IOWeight=30=, so a build never starves the host's sshd,
385the daemon or the backup timers, and =NoNewPrivileges=,
386=ProtectSystem=full=, =ProtectKernelTunables=, =ProtectControlGroups=
387and =RestrictSUIDSGID=, so a step cannot reach outside its workspace
388and the runner's home. The e2e suite alone starts sixty
389daemon instances; without the drop-in a deploy's copy over the admin
390sshd stalled. Both deploy targets copy with =rsync --partial=, which
391resumes a stalled transfer.
392
393Under =-isolation podman=, the default and what bay1 runs, each build
394is confined to a container (see Container isolation below). Under
395=-isolation none= steps run directly on the host as the runner's user,
396so treat that machine as executing whatever your users push, and
397install the toolchains your builds need on it.
398
399A runner claims the oldest pending build among the repositories its key
400is attached to — for an admin key, the oldest in the instance. =-repos=
401narrows within that set, which is what makes a runner outside the server
402practical: one on a machine that should build a single project, or that
403holds credentials for one deployment, stays on it.
404
405Oldest-first is across everything the key may claim, so a repository
406with a deep queue holds every other repository the same runner serves;
407bay1 measured a 15-minute average wait on a day of merge request
408stacks from one repository. A runner attached to one repository cannot
409be starved. That is the rule, decided in krz/gitbay#207: a runner
410serving several repositories takes them oldest-first, and an operator
411who wants one repository never to wait on another runs a second
412runner attached to it alone. Nothing caps what an account queues,
413and nothing needs to (krz/gitbay#206): a schedule tick queues nothing
414while the job's last build is pending or running, so a repository
415with no runner holds one row per scheduled job rather than one per
416tick, and a build runs only on a runner its owner attaches, so a busy
417schedule spends the owner's compute. Pushes are bounded by what an
418account can push.
419
420#+begin_src sh
421gitbay-runner -remote git@gitbay.org -repos krz/site,krz/docs \
422  -workdir /var/lib/gitbay-runner/work
423#+end_src
424
425Add =-untrusted= only with =-isolation podman=.
426
427gitbay.org's runner is attached to the forge's own repositories and
428the isolation canary, nothing else, because it shares the host with
429the forge; its unit names no =-repos=, the attachments are the
430boundary. Any other repository builds on a runner its owner attaches.
431
432=-repos= narrows an admin runner; for a runner key the attachments are
433the boundary, held by the server, and =-repos= may only name
434repositories among them. =-untrusted= makes a runner claim merge
435request heads from forks; the bay1 unit sets it because it isolates in
436podman. A runner without it builds trusted commits only.
437
438=gitbay dashboard= and =ssh git@<host> admin runners= list every key
439that has polled as a runner: the account, the key's fingerprint, when it
440last polled, the repositories it may claim — its attachments for a runner
441key, the =-repos= it asked for or =any= for an admin key — and the build
442it holds; =admin runners remove <fingerprint>= (=forget= until the next release) drops the row for a key
443that polled by mistake, the key itself untouched. =admin runners= also
444heads the list with the queue: builds
445pending now, and over the last day how many were claimed, how long they
446waited to be claimed (average and worst), and
447how many the reaper ended instead of a runner reporting them. A build a runner claimed and never
448reported is failed by the scheduler's minute tick, whether or not any
449runner is still alive: within about two minutes of its log stream ending
450with no outcome reported — the runner reports right after closing the
451stream, retrying for half a minute if gitbayd is unreachable — or, if no
452stream was ever seen, at the deadline.
453
454Instance admin on the runner account only authorizes the claim/report
455protocol; it grants no repo access. A build that pushes back — a pages
456deploy, an archive publish, an automated MR branch — needs an explicit
457grant on that repo: =repo access grant <owner/name> ci write=. Private
458repos likewise need at least read for the clone.
459
460** Container isolation
461
462Builds run in a rootless podman container, one per job, with the
463workspace bind mounted and nothing else: the clone happens outside with
464the runner's key, so a step cannot read it. =-isolation none= keeps the
465old behaviour — steps on the host as the runner's user — for an instance
466where every repository is trusted. There is no automatic fallback: a
467runner started with =-isolation podman= that cannot find a working
468podman exits rather than running a build unsandboxed.
469
470gitbay's own jobs name =localhost/gitbay-ci:2=, built from
471=deploy/Containerfile.ci= on the runner host. A job's image must carry
472what its steps need: the suite drives real git, git-lfs, gpg and sshd and
473asserts they exist before running, so the stock runner default would fail
474it immediately. Build or rebuild it with:
475
476#+begin_src sh
477ssh -p 2222 root@<host> 'cat > /tmp/Containerfile.ci' < deploy/Containerfile.ci
478ssh -p 2222 root@<host> 'su - ci-runner -s /bin/sh -c \
479  "podman build -t localhost/gitbay-ci:2 -f /tmp/Containerfile.ci /tmp"'
480#+end_src
481
482The tag is deliberate rather than =:latest=: changing the file means
483bumping the tag in =.gitbay/ci.yml=, so a running branch's image does not
484change under it.
485
486=-image= names the image a job runs in when it declares none, and is
487required under =-isolation podman=: there is no built-in default,
488because an image this host does not have would fail every build. A job
489overrides it with =image:= in =.gitbay/ci.yml=, validated as a reference
490so a config file cannot turn it into podman arguments.
491
492=-cpus= and =-memory= cap one build (podman's units, e.g. =-cpus 2
493-memory 4g=); unset means uncapped. The runner applies them itself: it
494creates a cgroup per build under its own delegated service cgroup,
495writes the limits there, and starts every podman process for the build
496inside it, with podman's cgroup handling off. Podman's own =--memory=
497and =--cpus= never applied under rootless cgroupfs, which is what a
498system service gets (krz/gitbay#188). The unit therefore needs
499=Delegate=yes=, which the drop-in sets; without it the runner refuses
500to start when a limit is set, and logs that builds run unconfined when
501none is. bay1 runs =-cpus 3 -memory 6g= per build inside =MemoryMax=6G=
502and =CPUQuota=300%= on the unit, on a 7.7GB four-core host with no
503swap: the memory cap is what keeps the forge alive when a build
504allocates without bound, and it sits above the e2e suite's 5GB peak
505rather than at a fair share. =OOMPolicy=continue= keeps systemd from
506stopping the runner when a build is OOM-killed.
507
508Each repository gets its own build home under the runner's workdir,
509mounted into its containers as =HOME=. Caches persist between builds of
510one repository and are never read by another's.
511
512*Images are provisioned, never pulled by a build.* The runner passes
513=--pull=never=. Two reasons, and the second is the better one: the
514service runs with =RestrictSUIDSGID=yes= so podman cannot unpack a layer
515holding a setuid file, which is nearly every distribution image; and on
516an instance where anyone can push a =ci.yml=, =image:= would otherwise
517mean "fetch and run anything from the internet". An operator pulls or
518builds what is allowed and a build picks among those. A job naming an
519image the host does not have fails with a message saying so.
520
521#+begin_src sh
522su - ci-runner -s /bin/sh -c "podman pull docker.io/library/alpine:3.20"
523su - ci-runner -s /bin/sh -c "podman images"
524#+end_src
525
526Prepare a host before pointing an isolating runner at it:
527
528#+begin_src sh
529ssh -p 2222 root@<host> 'sh -s' < deploy/runner-podman-setup.sh
530make deploy-runner
531#+end_src
532
533The script installs podman, delegates a subuid/subgid range to
534=ci-runner=, checks that user namespaces are enabled rather than
535assuming, enables lingering, and verifies rootless podman actually runs
536as that user. It is idempotent.
537
538The drop-in sets =NoNewPrivileges=no=, without which rootless podman
539cannot call =newuidmap= and the runner refuses to start. That is a
540considered trade, explained in the file and in the Threat-Model; if you
541run with =-isolation none=, set it back to =yes=.
542
543*Restarting the runner is safe.* On SIGTERM it stops claiming, finishes
544the build in flight, reports it, and exits; the drop-in's
545=TimeoutStopSec=50min= covers the longest build, and its =KillMode=mixed=
546is what makes the signal reach the runner alone — under systemd's default
547the build's container and the log session are signalled with it, and the
548runner drains a build that is already dead. So =make deploy-runner=
549waits for a running build rather than orphaning it, and a build's result
550is retried for half a minute if gitbayd is restarting at that moment. A
551second SIGTERM ends the runner at once, abandoning the build to the
552reaper. The suite checks all three: =TestRunnerDrainsOnSIGTERM= signals
553the process, =TestRunnerDropInLetsTheDrainHappen= reads the drop-in's
554=KillMode= and =TimeoutStopSec=, and =TestRunnerDrainUnderSystemd= runs
555the runner as a transient user unit under =systemd-run= and stops it
556under both kill modes. That last one needs a systemd user manager, so it
557skips in the container CI runs in; run it on a Linux host with
558=go test ./e2e -run TestRunnerDrainUnderSystemd -v=.
559
560*Validate podman mode on a scratch repository before pointing the runner
561at real ones.* Every deploy that switched the whole instance to
562containers and failed took CI down with it. Instead: create a throwaway
563repository the runner account can read (public, or granted read — a
564private one is "not found" to the runner and the build stays pending),
565give it one job that names the CI image, and deploy the runner with
566=-repos= naming only that repository. The production unit, with its real
567hardening, then claims nothing else; other repositories' builds queue
568until =-repos= is switched back, which is a pause, not an outage.
569
570#+begin_src sh
571gitbay repo create cmc/ci-smoke          # then push a .gitbay/ci.yml naming the image
572sed -i 's#-repos krz/gitbay #-repos cmc/ci-smoke #' /etc/systemd/system/gitbay-runner.service.d/override.conf
573systemctl daemon-reload && systemctl restart gitbay-runner
574gitbay build log cmc/ci-smoke 1         # green: switch -repos back, redeploy
575#+end_src
576
577*Do not deploy an isolating runner to a host that has not been
578prepared.* The runner is specified to refuse to start without a working
579podman rather than fall back to running builds unsandboxed — a fallback
580that silently drops isolation is worse than a stopped runner, because
581nothing surfaces it. On an unprepared host that refusal stops every
582build on the instance.
583
584The service drop-in carries =Delegate=yes= for rootless cgroup
585management and =ReadWritePaths= for podman's store under
586=/var/lib/gitbay-runner=, which =ProtectSystem=full= would otherwise
587make read-only. Those paths are prefixed =-= so they are ignored when
588absent: the drop-in installs on unprepared hosts too, and a unit that
589refused to start would stop every build.
590The nightly canary on =cmc/ci-smoke= only runs if the runner's =-repos=
591names that repository too; a scoped runner claims nothing else.
592=gitbay-runner-prune.timer= prunes unused images weekly, as the runner's
593user: rootless storage belongs to that user, and root's prune would not
594see it. An unpruned image store on a 40GB host is a slow outage.
595
596* LFS storage
597
598Objects live content-addressed under =[lfs] root= (default
599=<server.root>/lfs=); =[lfs] max_object_bytes= caps a single object
600(512MB default). Storage sits behind a small interface — an
601S3-compatible backend is a drop-in with the server proxying, and
602presigned URLs a later optimization. LFS objects do not travel with
603push mirrors (mirrors move git refs only), and gc does not yet collect
604orphaned objects.
605
606* Pages
607
608=[pages] domain = "example.site"= serves public repos' =pages= branches
609on =<owner>.<domain>=. DNS needs a wildcard record =*.<domain>= to the
610server; ACME issues per-subdomain certificates on demand (only for
611owners that exist). The domain must not be the site host or a parent of
612it — pages content runs its own scripts and must stay off the forge's
613origin.
614
615Users with repo admin claim custom domains with =repo domain add=.
616Claims activate only after a DNS TXT challenge proves control of the
617domain (=repo domain verify=, audit-logged); pending claims serve
618nothing, get no certificates, and expire after 7 days. ACME issues
619certificates only for verified hosts, so stray DNS pointed at the
620server gets nothing.
621
622* Security
623
624The [[Threat-Model]] file is the reference for what the forge
625trusts and refuses to do. Operational checklist:
626
627- *Software checks.* =deploy/audit.sh= runs =go vet=, =govulncheck=
628  (the module list is deliberately short — review it on each release),
629  and a short fuzz pass over every attacker-facing parser (pkt-line,
630  commit, SSHSIG armor, OpenPGP key, SSH tokenizer). Run it before
631  tagging a release. CI's own =vuln= job runs =govulncheck= nightly
632  against main rather than per push, because =@latest= scans today's
633  advisory database and an advisory lands without anyone pushing;
634  =build trigger krz/gitbay vuln= runs it on demand.
635- *Web responses* carry a scripts-forbidden CSP, =X-Frame-Options:
636  DENY=, =nosniff=, =no-referrer=, and HSTS when TLS is on — no
637  configuration needed.
638- *Host sandboxing.* The systemd unit in =deploy/cloud-init.yaml= runs
639  gitbayd unprivileged with =ProtectSystem=strict=, =PrivateDevices=,
640  =LockPersonality=, =MemoryDenyWriteExecute=,
641  =SystemCallFilter=@system-service=, and =RestrictAddressFamilies= to
642  INET/INET6/UNIX. It keeps =CAP_NET_BIND_SERVICE= only, to bind 22/80/443.
643- *OS patches* apply via =unattended-upgrades= (security origins,
644  auto-reboot 04:30 if required).
645- *Admin sshd (2222)* is throttled by =MaxStartups=/=MaxAuthTries= and
646  watched by =fail2ban=; gitbayd's own port 22 is throttled by
647  =limits.ssh_auth_rate= (auth failures per IP per minute), and every
648  account's writes by =limits.write_rate=.
649- *Monitoring.* =gitbay-monitor.timer= writes a reading hourly to
650journald and, when =/etc/gitbay/monitor.url= exists, posts it to that
651webhook: disk, service, the daemon's own =/healthz= answer, certificate
652expiry, and the age of the newest full backup and database snapshot.
653It exits non-zero on an alert so the unit shows in =systemctl
654--failed=: a stopped service, =/healthz= not answering =ok=, disk ≥ 85%,
655a certificate under 21 days, a full backup over 25 hours old, or a
656database snapshot over 2 hours old.
657
658=GET /healthz= is unauthenticated and cache-free: whether the database
659answers and which commit serves, 503 when it does not.
660- *Database.* =gitbay.db= and its WAL live under =/var/lib/gitbay= (mode
661  0750, owned by =gitbay=). The nightly archive plus provider snapshots
662  are the recovery path; for tighter RPO, add continuous replication
663  (litestream) against the same file — it coexists with the WAL.
664
665* Odds and ends
666
667- deleting a fork marks MRs sourced from it =source_gone=; their diffs
668  remain viewable and mergeable because the target repo owns the
669  objects.
670- =refs/merge-requests/*= is server-owned and unpushable by clients.
671- audit-relevant activity (issue/MR lifecycle, imports, pushes) lands in
672  the =events= table, which also feeds webhooks.
673- the daemon idles under 10MB RSS; the smallest VPS tier is adequate.