#+title: gitbay admin guide One static binary (=gitbayd=), one SQLite file, bare repositories on disk, and the system =git=. Schema migrations run automatically on startup and on every admin command. * Install Build from source (=go build ./cmd/gitbayd=), install via the vanity module path (=go install gitbay.org/gitbay/cmd/gitbayd@latest=), or use a release build: =deploy/release.sh = cross-compiles reproducible linux/amd64, linux/arm64, and darwin/arm64 binaries with a SHA256SUMS manifest (CGO off, trimpath, stripped — byte-identical per commit and toolchain). #+begin_src sh install -m 755 gitbayd /usr/local/bin/ adduser --system --group --home /var/lib/gitbay --shell /usr/sbin/nologin gitbay install -d -o gitbay -g gitbay -m 750 /var/lib/gitbay gitbayd --config /etc/gitbay/config.toml check-config #+end_src =deploy/= in the source tree has a cloud-init file, a hardened systemd unit, and a nightly backup timer. Run as the unprivileged =gitbay= user; the unit's =AmbientCapabilities=CAP_NET_BIND_SERVICE= covers ports 22/80/443 without root. ** The SSH port decision - =ssh.mode = "embedded"= (default): gitbayd itself listens, normally on 22 — move the host's admin sshd to another port. Remotes read =git@host:owner/repo= with no port gymnastics. - =ssh.mode = "system"=: the host sshd owns 22 and invokes gitbayd via =AuthorizedKeysCommand=: #+begin_example AuthorizedKeysCommand /usr/local/bin/gitbayd --config /etc/gitbay/config.toml authorized-keys %t %k AuthorizedKeysCommandUser gitbay #+end_example sshd requires that binary to be root-owned and not group/world writable. Unknown keys fail authentication inside sshd, so system mode requires =registration.mode = "closed"= (check-config enforces this). * Configuration reference =/etc/gitbay/config.toml=. =check-config= validates and names every contradiction; =--no-host-checks= skips port/path probes. =gitbayd admin config show= prints the configuration in effect as TOML, every default filled in and =smtp_pass= redacted. A file that fails validation still prints, followed by the contradiction. ** [server] - =root= (default =/var/lib/gitbay=) — repositories, database, host keys, ACME cache all live here. - =site_url= (required) — canonical =https://host=; drives ACME, clone URLs, mail links. - =source_repo= (optional, =owner/name=) — the repository this instance develops itself in. Startup warns when the running build's commit is not on that repository's default branch, which is how a binary built from an unmerged branch stops being invisible. Leave it unset unless the instance hosts its own source. ** [ssh] - =mode= — =embedded= | =system= (above). - =port= (22) — embedded listener port. - =host_keys= — list of private key paths; empty generates an ed25519 key at =/ssh/host_ed25519=. ** [http] - =addr= (=:443=), =tls= — =acme= | =files= | =off=. - =acme=: certificates via TLS-ALPN-01 on the HTTPS port, cached at =/acme=; =acme_email= for the CA account; =acme_http_addr= (=:80=, ="off"= to disable) adds HTTP-01 and an https redirect — failing to bind it is a warning, not fatal. Requires an =https://= site_url with a public DNS name. - =files=: =cert_file= + =key_file=. - =off=: plain HTTP — development, or behind a TLS-terminating proxy. =trusted_proxies= lists the addresses or CIDRs of reverse proxies in front of the daemon. A request from one of them is attributed, for API rate limiting, to the last =X-Forwarded-For= hop that is not itself a trusted proxy; from anyone else the header is ignored. Empty, the default, is right when gitbayd terminates TLS itself. A reverse proxy in front of gitbay must not buffer responses, or the build page's live log arrives only when the build ends; =X-Accel-Buffering: no= covers nginx. ** [web] - =mode= — =view_only= (default) | =accounts=. In view_only the mutating web routes are never registered; in accounts, browser sessions are minted over SSH (=web login=), and users with write access can create repos, comment, and make simple file edits (which commit unsigned, honestly). =password_auth= is reserved and currently rejected. - =title= — the instance's display name in the rail and page titles; empty falls back to the site host. Lower case is the convention for gitbay itself. - =privacy_notice= — operator text shown on =/privacy= under the fixed statement. ** [registration] - =mode= — =closed= (default) | =invite= | =open=. invite/open require [mail]. See the user guide for the flows. - =pending_expiry= (empty, never) — a duration such as ="168h"=; a self-registered account still unverified after that long is removed, hourly and at start, audited as =pending.expired=. - =notify_admin= (false) — mail every instance admin when an account becomes active: an invite redeemed, or an open-mode signup that verified its address. The unverified row an open signup creates is not reported, because anyone can post the form and mailing on that would point a flood at the admins. Recipients are the verified primary addresses of active admins who have activity mail on, the same rule any other notice follows, so an admin with no verified address hears nothing. The notice is queued, so a dead SMTP host shows up in the admin page's Mail table instead of failing the registration. Requires [mail]. ** [mail] - =smtp_host= (host:port, 587 assumed), =from=, optional =smtp_user= / =smtp_pass=. STARTTLS when offered. Required for invite/open registration and self-service =email add=; in closed mode you may omit it entirely and assert addresses by hand (below). ** [push] Push notifications to Apple devices, delivered by gitbayd talking to APNs directly over HTTP/2, authenticated by an ES256 JWT signed with an operator-supplied provider key. Off unless configured. - =enabled= (false). - =key_file= — path to the =.p8= provider key from Apple's developer portal (Certificates, Identifiers & Profiles → Keys). It belongs at =/etc/gitbay/apns.p8=, mode 0600, owned by the account gitbayd runs as. Read and validated at startup: it must parse as a PEM-wrapped PKCS#8 EC (P-256) private key, or the daemon refuses to start rather than fill a queue nobody is watching. - =key_id=, =team_id= — the key's id and your Apple developer team id, both from the same portal page. - =topic= — the app's bundle identifier. *An APNs key belongs to a bundle ID.* gitbay.org pushes to the App Store build under its own bundle id; a self-hoster who wants push ships their own iOS build under their own bundle id, with its own =.p8= key from their own developer account, and points =topic= at that id. There is no way to push to someone else's build, by design — this is Apple's model, not gitbay's. - =environment= — =production= or =sandbox=, naming the APNs host rather than taking a URL, so a typo cannot aim the key at a host that is not Apple's. All five of =key_file=, =key_id=, =team_id=, =topic= and =environment= are required when =enabled= is true; validation runs at config load, so a misconfigured =[push]= is caught before the daemon serves anything. The delivery queue (a device's undelivered and attempted pushes) is capped the same way the mail queue is, by =[retention] push=. ** [api] - =enabled= (false) — the JSON API surface; see [[API]]. Off means no credential-bearing HTTP endpoint exists at all. ** [webhooks] - =allow_local= (false) — permit webhook targets on loopback/private addresses. Leave off unless you know why you need it (SSRF). ** [limits] - =clone_timeout= (3600s) — cap on =repo import= fetches. - =max_blob_bytes= (100MB) — cap on raw file serving over the web. - =max_asset_bytes= (512MB) — cap per uploaded release asset. - =max_snippet_bytes= (1MB) — cap per snippet file. - =max_snippets_per_user= (0, unlimited) — snippets an account may own. - =max_repos_per_user= (0, unlimited) — repositories an account may own directly; =repo create=, =fork= and =import= refuse past it. Organizations are not capped. - =max_bytes_per_user= (0, unlimited) — disk the account's own repositories may take; a push may be no larger than what is left. - =max_pack_bytes=, =ssh_auth_rate= — reserved, not yet enforced. ** [git_daemon] - =enabled= (false), =port= (9418) — the anonymous =git://= listener. Serves only public repositories that additionally ran =repo settings git-daemon on=. ** [mirrors] - =pull_interval_minutes= (15) — how often pull mirrors fetch their upstream. Push mirrors sync shortly after each local ref update. Mirror URLs pass the same SSRF rules as webhook targets. ** [go_import] Vanity Go module paths, one per line: ="host/module" = "owner/repo"=. Requests with =?go-get=1= at or under the module path answer with the go-import meta tag pointing at the repository's HTTPS clone URL, so =go install host/module/cmd/...@latest= resolves. The repository should be public (the module path itself confirms it exists). * Users, email, invites Every =gitbayd admin= subcommand except =backup=, =gc= and the one-shot backfills is a wrapper that dispatches the registry command of the same name as the host: an admin context with no account behind it, so its audit rows carry no actor and =source: host=. The same commands run in an instance admin's SSH session (=ssh git@ admin ...=) and write the same rows with the key fingerprint as source. One implementation, two credentials. #+begin_src sh gitbayd admin user create alice --key alice.pub --email a@example.org --verified [--admin] gitbayd admin email verify alice a@example.org # admin assertion, no SMTP needed gitbayd admin invite --email b@example.org # mails a code; prints it if no SMTP #+end_src "Verified" means SMTP-confirmed or host-admin-asserted; the database records which. Verified emails are what make commit signatures meaningful — an unverified address never produces a =verified= badge. * Audit and account control The audit log is the security feed (events are the product feed): every successful mutating command with its argv and source credential (SSH key fingerprint or API), registrations, admin actions, force-pushes, and auth failures/throttling. Secrets never appear — they travel on stdin, never in argv. #+begin_src sh gitbayd admin audit [--actor u|-] [--action prefix] [--since 24h|7d|date] [--limit n] [--json] ssh git@ audit ... # the same, from an admin session ssh git@ admin user list [--state active|pending|disabled|admin] ssh git@ admin user show # keys, emails, orgs, tokens, sessions ssh git@ admin user limits [--repos n|default] [--bytes n|default] # per-account caps ssh git@ admin user promote # grant instance admin ssh git@ admin user demote # remove it; the last admin is refused gitbayd admin user promote # host-local: recovery when no admin key is reachable gitbayd admin user disable # suspend: SSH, web, API all refused; gitbayd admin user enable # sessions dropped, nothing deleted gitbayd admin user delete --yes # only for accounts anchoring nothing: # refused (with each blocker named) while # the account owns repos, authored # issues/MRs/comments/reviews, or is an # org's only admin #+end_src =--actor= takes a username, or =-= for rows with no actor: host commands and auth failures. =--action= is a prefix, so =cmd repo= catches every repository command and =admin= every host or admin-session action. =--since= is a duration back from now (=30m=, =24h=, =7d=) or a date. =admin user list= pages by username (=--limit=, =--cursor=) and carries each account's state and =last_seen=, the newest use of any of its SSH keys or API tokens. =admin user show= adds the keys with their last use, each address with how it was verified, PGP keys, org roles, the owned repository count, API token names, and live browser sessions. Both are Both are refused to non-admins, like =audit=, on every surface. =/admin/users= is the same list in a browser, linked from the admin page: the state filter the command takes, keyset paging on its cursor, and a row per account with promote, demote, disable and enable, each dispatching the command. Demote and disable ask for the username to be typed, since both take someone's access away. Creating and deleting an account, issuing an invite and asserting an address stay on the command line: each takes a key, mints a credential, or cannot be undone. A non-admin gets the 404 a missing page would, so the URL confirms nothing. Promotion needs an active account: a pending or disabled one is refused. Demotion is refused when it would leave no admin, over SSH, in the browser and on the host alike, so the host-local =promote= is the way back in when the only admin key is lost. Instance admin carries no right on anyone's repository: policy does not consult it, and a private repository still answers not-found to an admin. Moderation goes through explicit overrides that skip the access check and write their own audit row: #+begin_src sh ssh git@ admin repo list [--owner o] [--visibility public|private] # size, last push ssh git@ admin repo archive|unarchive ssh git@ admin repo visibility public|private ssh git@ admin repo delete --yes #+end_src Each lands in the audit log as =admin repo.= naming the repository, on top of the =cmd= row every mutating command gets. =limits.ssh_auth_rate= (10) throttles per-IP authentication *failures* per minute — successful auths never count and clear the slate. =limits.max_pack_bytes= is enforced as =receive.maxInputSize= on every push. =limits.write_rate= (60) bounds *mutating commands per account per minute*. It is counted in the dispatcher, so SSH, the JSON API and the web spend one budget and a caller cannot refresh it by changing surface; =limits.api_rate= stays in front of it, bounding a network source rather than an account. A command is one token whatever it writes, so a bundle import costs one and only a loop of separate commands spends the budget. Read-only commands, the runner protocol (a build streams its log in many small writes) and the host CLI are exempt. Refusals exit 4 and say when to retry. A negative value turns the limit off; it matters most with =registration = "open"=, where every write also queues notification mail and webhook deliveries. * Queues Every background worker keeps a backlog and a failure state. An instance admin reads them all in one place: #+begin_src sh gitbay dashboard --json | jq .queues # webhooks, mail, push, mirrors, builds, deps #+end_src Per worker: pending, retrying (pending with a failed attempt) and dead-lettered counts with the oldest pending age, and the retrying or failed rows themselves, capped at twenty each. Builds list what is running and then what is pending, each since when, so a build no runner is scoped to claim is visible here rather than only in its repository; mirrors list the ones whose last sync failed; dependency checks list the ones whose last check errored. Non-admins get no =queues= key at all. Push rows name the device id, never the token. Watch this one after configuring =[push]=: a =key_id= or =team_id= Apple did not issue passes config validation, which can only check that the =.p8= parses, and then every send comes back =403 InvalidProviderToken= and dead-letters on its first attempt. In accounts mode the same read renders at =/admin=, linked from the rail for admins. Anyone else gets a 404 there. A dead-lettered mail is logged as =notification dead-lettered mail==, with the address redacted out of the relay's error. The id is the queue row: find it in the Mail table on =/admin=, or in =dashboard --json=, where the recipient and the unredacted error are. That is deliberate — see the Threat-Model page. * Maintenance #+begin_src sh gitbayd admin stats [--json] # counts, database size, per-repo disk ssh git@ admin stats [--json] # the same, from an admin session gitbayd admin gc [--repo owner/name] # git gc: repack and prune; per-repo sizes gitbayd admin gc --aggressive # thorough repack; slow, rarely needed gitbayd admin gc --lfs # also drop LFS objects no pointer names (older than a day) #+end_src A history rewrite leaves the commits it removed reachable through =refs/merge-requests/N/head= of the merge requests that landed them, so they stay fetchable by anyone who can read the repository. Nothing drops a head ref on its own — an open or source-gone MR is merged through it, and a merged or closed one keeps its diff readable through it — so the cleanup is a command an instance admin runs, naming the MRs: #+begin_src sh ssh git@ admin mr prune owner/name 1 2 3 --yes #+end_src It refuses an open or source-gone MR, deletes the named refs, runs =git gc --prune=now= on that one repository so the objects go at once rather than after git's two-week grace, leaves a system comment on each MR, and audits as =admin mr.prune=. The MR keeps its title, comments, reviews and head sha; =mr diff= and the MR page say the head is gone. Run it when nothing is pushing to that repository: without the grace, a push caught between leaving quarantine and writing its ref loses its objects. Objects also survive in offsite backups until those are rewritten; see "Removing a repository's history from every snapshot". =deploy/cloud-init.yaml= ships a =gitbay-gc.timer= that runs =admin gc= weekly (Sunday 07:00 UTC). Imported repositories keep whatever pack layout the source sent, so a first manual =admin gc= after a bulk import is worthwhile. * Backup and restore #+begin_src sh gitbayd admin backup --out /var/backups/gitbay/backup.tar.gz gitbayd admin backup --verify /var/backups/gitbay/backup.tar.gz # read it back #+end_src One archive: a consistent SQLite snapshot (taken *before* the repositories are read, so the database never references objects the archive missed), every repository, and the SSH host keys. Excluded: hook socket, regenerated hook scripts, WAL files. Safe to run against a live daemon. =--verify= reads an archive back: the snapshot must pass SQLite's integrity check, and every repository the snapshot names must be in the archive. A database-only archive is checked for integrity and says so. Exit is non-zero on damage or a missing repository. Restore: extract into an empty directory, point =server.root= at it, start gitbayd. Host keys are preserved, so clients keep their known_hosts entries; hooks regenerate at startup. ** Schedule and recovery point Two timers, because the two halves of the data have different exposure. - =gitbay-backup.timer=, nightly. The full archive above, last 7 kept. - =gitbay-db-backup.timer=, hourly. =admin backup --db-only=, which writes the SQLite snapshot alone, last 48 kept. A few MB against the full archive's hundreds, which is what makes the frequency affordable. The split follows what a loss would actually cost. Repositories are git, so a mirror or any clone is a second copy; the database is the only copy of issues, merge requests, comments and review state. So the recovery point is about an hour for the data that exists nowhere else, and a day for the data that does. Continuous replication (litestream and similar) was considered and not adopted. It would take the database's recovery point to seconds, but the repositories would still be on the nightly archive, so a restore could produce a database referencing commits the repository backup does not have. Consistency between the two halves is worth more here than latency on one of them. Revisit if repository replication becomes continuous too. ** Offsite copies bay1 also takes a nightly restic snapshot of =/var/lib/gitbay= and =/var/lib/gitbay-stage= (the staged database copy) to an S3 bucket at Scaleway, with a key that can only add snapshots. The key that can remove them lives on the operator's machine, in =~/.config/gitbay/offsite.env=, and never on bay1: a compromised host cannot destroy its own history. Forgetting, pruning and rewriting all run from there. *** Removing a repository's history from every snapshot A history rewrite plus =admin mr prune= takes commits off the server, but every snapshot taken before it still holds them, and the retention window is the only thing that ages them out. To remove them now, rewrite the snapshots without that repository rather than forgetting the snapshots: everything else in them stays restorable. The next nightly run adds the repository back in its current state. Repositories are stored under the name they had on disk when each snapshot was taken, so a renamed repository needs every name it has carried. Check what an older snapshot holds before choosing the paths: #+begin_src sh set -a; . ~/.config/gitbay/offsite.env; set +a restic $RESTIC_OPTS snapshots restic $RESTIC_OPTS ls /var/lib/gitbay/repos/ #+end_src Then dry-run, apply, prune, and confirm nothing matches: #+begin_src sh EXCL="--exclude /var/lib/gitbay/repos//.git --exclude /var/lib/gitbay/repos//.git" restic $RESTIC_OPTS rewrite --dry-run $EXCL # "would modify N snapshots" restic $RESTIC_OPTS rewrite --forget $EXCL # new snapshots replace the originals restic $RESTIC_OPTS prune # drops the data nothing references restic $RESTIC_OPTS find .git .git # expect no output #+end_src =--forget= is what makes the originals go; without it the rewritten snapshots sit beside them and the data stays referenced. Snapshot IDs change; their times do not. Done for krz/keycask (formerly rust-pass) on 2026-09-18, across 22 snapshots. * Upgrades Replace the binary, restart the unit. Migrations apply automatically and are transactional; hook scripts under =/hooks= are rewritten at startup to point at the current binary path. * CI runner =gitbay-runner= executes builds queued by pushes and merge requests. It polls over SSH with a key of scope =runner=, which reaches only the runner protocol and read-only git (a runner executes arbitrary repository code, so the key it holds must not do more). A runner key claims builds only for the repositories it is attached to, by =repo runner add= from a repository admin or an instance admin; an admin key claims any. Users attach their own runners: see the Users page. For an instance runner, run it as a dedicated unprivileged user on a non-admin account. =admin user create --key= registers a full-scope key, so the runner key is added afterwards through a bootstrap key that is then removed, and attached to each repository it should build: #+begin_src sh useradd --system --create-home --home-dir /var/lib/gitbay-runner ci-runner sudo -u ci-runner ssh-keygen -t ed25519 -N "" -f /var/lib/gitbay-runner/.ssh/id_ed25519 ssh-keygen -t ed25519 -N "" -f /tmp/ci-bootstrap gitbayd --config /etc/gitbay/config.toml admin user create ci --key /tmp/ci-bootstrap.pub ssh -i /tmp/ci-bootstrap git@127.0.0.1 keys add --scope runner < /var/lib/gitbay-runner/.ssh/id_ed25519.pub ssh -i /tmp/ci-bootstrap git@127.0.0.1 keys remove "$(ssh-keygen -lf /tmp/ci-bootstrap.pub | awk '{print $2}')" rm /tmp/ci-bootstrap /tmp/ci-bootstrap.pub gitbay-runner -remote git@127.0.0.1 -workdir /var/lib/gitbay-runner/work #+end_src #+begin_src sh gitbay repo runner add krz/site < /var/lib/gitbay-runner/.ssh/id_ed25519.pub #+end_src =-jobs N= runs N builds at once. Claiming is one transaction that selects and updates, and each build works in its own =build-= directory, so workers do not collide; idle polls are staggered across the interval so N of them do not wake together. The drop-in's weights below are per service, not per build, so raising =-jobs= divides them rather than multiplying the host's load. =admin runners= shows which account each runner polls as, and what each may claim. A runner key claims builds only for the repositories it is attached to: with none attached it claims nothing, and =-repos= may only narrow within them. An admin's full-scope key claims any repository — that is what =-repos= was for — and still works for the protocol during a rotation. A merge request head from a fork is built in the target repository as untrusted: the claim carries no secrets, and only a runner started with =-untrusted= takes it. Same-repository heads were built by their branch push and are not built again. =make deploy-runner= also installs =deploy/gitbay-runner.override.conf= as a systemd drop-in: =Nice=10=, =CPUWeight=30=, =IOWeight=30=, so a build never starves the host's sshd, the daemon or the backup timers, and =NoNewPrivileges=, =ProtectSystem=full=, =ProtectKernelTunables=, =ProtectControlGroups= and =RestrictSUIDSGID=, so a step cannot reach outside its workspace and the runner's home. The e2e suite alone starts sixty daemon instances; without the drop-in a deploy's copy over the admin sshd stalled. Both deploy targets copy with =rsync --partial=, which resumes a stalled transfer. Under =-isolation podman=, the default and what bay1 runs, each build is confined to a container (see Container isolation below). Under =-isolation none= steps run directly on the host as the runner's user, so treat that machine as executing whatever your users push, and install the toolchains your builds need on it. A runner claims the oldest pending build among the repositories its key is attached to — for an admin key, the oldest in the instance. =-repos= narrows within that set, which is what makes a runner outside the server practical: one on a machine that should build a single project, or that holds credentials for one deployment, stays on it. Oldest-first is across everything the key may claim, so a repository with a deep queue holds every other repository the same runner serves; bay1 measured a 15-minute average wait on a day of merge request stacks from one repository. A runner attached to one repository cannot be starved. That is the rule, decided in krz/gitbay#207: a runner serving several repositories takes them oldest-first, and an operator who wants one repository never to wait on another runs a second runner attached to it alone. Nothing caps what an account queues, and nothing needs to (krz/gitbay#206): a schedule tick queues nothing while the job's last build is pending or running, so a repository with no runner holds one row per scheduled job rather than one per tick, and a build runs only on a runner its owner attaches, so a busy schedule spends the owner's compute. Pushes are bounded by what an account can push. #+begin_src sh gitbay-runner -remote git@gitbay.org -repos krz/site,krz/docs \ -workdir /var/lib/gitbay-runner/work #+end_src Add =-untrusted= only with =-isolation podman=. gitbay.org's runner is attached to the forge's own repositories and the isolation canary, nothing else, because it shares the host with the forge; its unit names no =-repos=, the attachments are the boundary. Any other repository builds on a runner its owner attaches. =-repos= narrows an admin runner; for a runner key the attachments are the boundary, held by the server, and =-repos= may only name repositories among them. =-untrusted= makes a runner claim merge request heads from forks; the bay1 unit sets it because it isolates in podman. A runner without it builds trusted commits only. =gitbay dashboard= and =ssh git@ admin runners= list every key that has polled as a runner: the account, the key's fingerprint, when it last polled, the repositories it may claim — its attachments for a runner key, the =-repos= it asked for or =any= for an admin key — and the build it holds; =admin runners remove = (=forget= until the next release) drops the row for a key that polled by mistake, the key itself untouched. =admin runners= also heads the list with the queue: builds pending now, and over the last day how many were claimed, how long they waited to be claimed (average and worst), and how many the reaper ended instead of a runner reporting them. A build a runner claimed and never reported is failed by the scheduler's minute tick, whether or not any runner is still alive: within about two minutes of its log stream ending with no outcome reported — the runner reports right after closing the stream, retrying for half a minute if gitbayd is unreachable — or, if no stream was ever seen, at the deadline. Instance admin on the runner account only authorizes the claim/report protocol; it grants no repo access. A build that pushes back — a pages deploy, an archive publish, an automated MR branch — needs an explicit grant on that repo: =repo access grant ci write=. Private repos likewise need at least read for the clone. ** Container isolation Builds run in a rootless podman container, one per job, with the workspace bind mounted and nothing else: the clone happens outside with the runner's key, so a step cannot read it. =-isolation none= keeps the old behaviour — steps on the host as the runner's user — for an instance where every repository is trusted. There is no automatic fallback: a runner started with =-isolation podman= that cannot find a working podman exits rather than running a build unsandboxed. gitbay's own jobs name =localhost/gitbay-ci:2=, built from =deploy/Containerfile.ci= on the runner host. A job's image must carry what its steps need: the suite drives real git, git-lfs, gpg and sshd and asserts they exist before running, so the stock runner default would fail it immediately. Build or rebuild it with: #+begin_src sh ssh -p 2222 root@ 'cat > /tmp/Containerfile.ci' < deploy/Containerfile.ci ssh -p 2222 root@ 'su - ci-runner -s /bin/sh -c \ "podman build -t localhost/gitbay-ci:2 -f /tmp/Containerfile.ci /tmp"' #+end_src The tag is deliberate rather than =:latest=: changing the file means bumping the tag in =.gitbay/ci.yml=, so a running branch's image does not change under it. =-image= names the image a job runs in when it declares none, and is required under =-isolation podman=: there is no built-in default, because an image this host does not have would fail every build. A job overrides it with =image:= in =.gitbay/ci.yml=, validated as a reference so a config file cannot turn it into podman arguments. =-cpus= and =-memory= cap one build (podman's units, e.g. =-cpus 2 -memory 4g=); unset means uncapped. The runner applies them itself: it creates a cgroup per build under its own delegated service cgroup, writes the limits there, and starts every podman process for the build inside it, with podman's cgroup handling off. Podman's own =--memory= and =--cpus= never applied under rootless cgroupfs, which is what a system service gets (krz/gitbay#188). The unit therefore needs =Delegate=yes=, which the drop-in sets; without it the runner refuses to start when a limit is set, and logs that builds run unconfined when none is. bay1 runs =-cpus 3 -memory 6g= per build inside =MemoryMax=6G= and =CPUQuota=300%= on the unit, on a 7.7GB four-core host with no swap: the memory cap is what keeps the forge alive when a build allocates without bound, and it sits above the e2e suite's 5GB peak rather than at a fair share. =OOMPolicy=continue= keeps systemd from stopping the runner when a build is OOM-killed. Each repository gets its own build home under the runner's workdir, mounted into its containers as =HOME=. Caches persist between builds of one repository and are never read by another's. *Images are provisioned, never pulled by a build.* The runner passes =--pull=never=. Two reasons, and the second is the better one: the service runs with =RestrictSUIDSGID=yes= so podman cannot unpack a layer holding a setuid file, which is nearly every distribution image; and on an instance where anyone can push a =ci.yml=, =image:= would otherwise mean "fetch and run anything from the internet". An operator pulls or builds what is allowed and a build picks among those. A job naming an image the host does not have fails with a message saying so. #+begin_src sh su - ci-runner -s /bin/sh -c "podman pull docker.io/library/alpine:3.20" su - ci-runner -s /bin/sh -c "podman images" #+end_src Prepare a host before pointing an isolating runner at it: #+begin_src sh ssh -p 2222 root@ 'sh -s' < deploy/runner-podman-setup.sh make deploy-runner #+end_src The script installs podman, delegates a subuid/subgid range to =ci-runner=, checks that user namespaces are enabled rather than assuming, enables lingering, and verifies rootless podman actually runs as that user. It is idempotent. The drop-in sets =NoNewPrivileges=no=, without which rootless podman cannot call =newuidmap= and the runner refuses to start. That is a considered trade, explained in the file and in the Threat-Model; if you run with =-isolation none=, set it back to =yes=. *Restarting the runner is safe.* On SIGTERM it stops claiming, finishes the build in flight, reports it, and exits; the drop-in's =TimeoutStopSec=50min= covers the longest build, and its =KillMode=mixed= is what makes the signal reach the runner alone — under systemd's default the build's container and the log session are signalled with it, and the runner drains a build that is already dead. So =make deploy-runner= waits for a running build rather than orphaning it, and a build's result is retried for half a minute if gitbayd is restarting at that moment. A second SIGTERM ends the runner at once, abandoning the build to the reaper. The suite checks all three: =TestRunnerDrainsOnSIGTERM= signals the process, =TestRunnerDropInLetsTheDrainHappen= reads the drop-in's =KillMode= and =TimeoutStopSec=, and =TestRunnerDrainUnderSystemd= runs the runner as a transient user unit under =systemd-run= and stops it under both kill modes. That last one needs a systemd user manager, so it skips in the container CI runs in; run it on a Linux host with =go test ./e2e -run TestRunnerDrainUnderSystemd -v=. *Validate podman mode on a scratch repository before pointing the runner at real ones.* Every deploy that switched the whole instance to containers and failed took CI down with it. Instead: create a throwaway repository the runner account can read (public, or granted read — a private one is "not found" to the runner and the build stays pending), give it one job that names the CI image, and deploy the runner with =-repos= naming only that repository. The production unit, with its real hardening, then claims nothing else; other repositories' builds queue until =-repos= is switched back, which is a pause, not an outage. #+begin_src sh gitbay repo create cmc/ci-smoke # then push a .gitbay/ci.yml naming the image sed -i 's#-repos krz/gitbay #-repos cmc/ci-smoke #' /etc/systemd/system/gitbay-runner.service.d/override.conf systemctl daemon-reload && systemctl restart gitbay-runner gitbay build log cmc/ci-smoke 1 # green: switch -repos back, redeploy #+end_src *Do not deploy an isolating runner to a host that has not been prepared.* The runner is specified to refuse to start without a working podman rather than fall back to running builds unsandboxed — a fallback that silently drops isolation is worse than a stopped runner, because nothing surfaces it. On an unprepared host that refusal stops every build on the instance. The service drop-in carries =Delegate=yes= for rootless cgroup management and =ReadWritePaths= for podman's store under =/var/lib/gitbay-runner=, which =ProtectSystem=full= would otherwise make read-only. Those paths are prefixed =-= so they are ignored when absent: the drop-in installs on unprepared hosts too, and a unit that refused to start would stop every build. The nightly canary on =cmc/ci-smoke= only runs if the runner's =-repos= names that repository too; a scoped runner claims nothing else. =gitbay-runner-prune.timer= prunes unused images weekly, as the runner's user: rootless storage belongs to that user, and root's prune would not see it. An unpruned image store on a 40GB host is a slow outage. * LFS storage Objects live content-addressed under =[lfs] root= (default =/lfs=); =[lfs] max_object_bytes= caps a single object (512MB default). Storage sits behind a small interface — an S3-compatible backend is a drop-in with the server proxying, and presigned URLs a later optimization. LFS objects do not travel with push mirrors (mirrors move git refs only), and gc does not yet collect orphaned objects. * Pages =[pages] domain = "example.site"= serves public repos' =pages= branches on =.=. DNS needs a wildcard record =*.= to the server; ACME issues per-subdomain certificates on demand (only for owners that exist). The domain must not be the site host or a parent of it — pages content runs its own scripts and must stay off the forge's origin. Users with repo admin claim custom domains with =repo domain add=. Claims activate only after a DNS TXT challenge proves control of the domain (=repo domain verify=, audit-logged); pending claims serve nothing, get no certificates, and expire after 7 days. ACME issues certificates only for verified hosts, so stray DNS pointed at the server gets nothing. * Security The [[Threat-Model]] file is the reference for what the forge trusts and refuses to do. Operational checklist: - *Software checks.* =deploy/audit.sh= runs =go vet=, =govulncheck= (the module list is deliberately short — review it on each release), and a short fuzz pass over every attacker-facing parser (pkt-line, commit, SSHSIG armor, OpenPGP key, SSH tokenizer). Run it before tagging a release. CI's own =vuln= job runs =govulncheck= nightly against main rather than per push, because =@latest= scans today's advisory database and an advisory lands without anyone pushing; =build trigger krz/gitbay vuln= runs it on demand. - *Web responses* carry a scripts-forbidden CSP, =X-Frame-Options: DENY=, =nosniff=, =no-referrer=, and HSTS when TLS is on — no configuration needed. - *Host sandboxing.* The systemd unit in =deploy/cloud-init.yaml= runs gitbayd unprivileged with =ProtectSystem=strict=, =PrivateDevices=, =LockPersonality=, =MemoryDenyWriteExecute=, =SystemCallFilter=@system-service=, and =RestrictAddressFamilies= to INET/INET6/UNIX. It keeps =CAP_NET_BIND_SERVICE= only, to bind 22/80/443. - *OS patches* apply via =unattended-upgrades= (security origins, auto-reboot 04:30 if required). - *Admin sshd (2222)* is throttled by =MaxStartups=/=MaxAuthTries= and watched by =fail2ban=; gitbayd's own port 22 is throttled by =limits.ssh_auth_rate= (auth failures per IP per minute), and every account's writes by =limits.write_rate=. - *Monitoring.* =gitbay-monitor.timer= writes a reading hourly to journald and, when =/etc/gitbay/monitor.url= exists, posts it to that webhook: disk, service, the daemon's own =/healthz= answer, certificate expiry, and the age of the newest full backup and database snapshot. It exits non-zero on an alert so the unit shows in =systemctl --failed=: a stopped service, =/healthz= not answering =ok=, disk ≥ 85%, a certificate under 21 days, a full backup over 25 hours old, or a database snapshot over 2 hours old. =GET /healthz= is unauthenticated and cache-free: whether the database answers and which commit serves, 503 when it does not. - *Database.* =gitbay.db= and its WAL live under =/var/lib/gitbay= (mode 0750, owned by =gitbay=). The nightly archive plus provider snapshots are the recovery path; for tighter RPO, add continuous replication (litestream) against the same file — it coexists with the WAL. * Odds and ends - deleting a fork marks MRs sourced from it =source_gone=; their diffs remain viewable and mergeable because the target repo owns the objects. - =refs/merge-requests/*= is server-owned and unpushable by clients; only =admin mr prune= removes one (see Maintenance). - audit-relevant activity (issue/MR lifecycle, imports, pushes) lands in the =events= table, which also feeds webhooks. - the daemon idles under 10MB RSS; the smallest VPS tier is adequate.