.gitbay/wiki/Admin.org

v1.40.0
gitbay/.gitbay/wiki/Admin.org rendered · source · history · blame · raw

1313 lines · 70412 bytes

gitbay admin guide

One static binary (gitbayd), one SQLite file, bare repositories on disk, and the system git. Schema migrations run automatically on startup and on every admin command.

Install

Build from source (go build ./cmd/gitbayd), install via the vanity module path (go install gitbay.org/gitbay/cmd/gitbayd@latest), or use a release build: deploy/release.sh <tag> cross-compiles reproducible linux/amd64, linux/arm64, and darwin/arm64 binaries with a SHA256SUMS manifest (CGO off, trimpath, stripped — byte-identical per commit and toolchain).

install -m 755 gitbayd /usr/local/bin/
adduser --system --group --home /var/lib/gitbay --shell /usr/sbin/nologin gitbay
install -d -o gitbay -g gitbay -m 750 /var/lib/gitbay
gitbayd --config /etc/gitbay/config.toml check-config
gitbayd --config /etc/gitbay/config.toml admin secrets init
chown gitbay:gitbay /etc/gitbay/secret.key

admin secrets init refuses when the key file already exists, so a reinstall on the same host should skip it — deploy/install.sh does this with a file check before running it.

deploy/ in the source tree has a cloud-init file, a hardened systemd unit, and a nightly backup timer. Run as the unprivileged gitbay user; the unit's AmbientCapabilities=CAP_NET_BIND_SERVICE covers ports 22/80/443 without root.

The SSH port decision

  • ssh.mode = "embedded" (default): gitbayd itself listens, normally on 22 — move the host's admin sshd to another port. Remotes read git@host:owner/repo with no port gymnastics.
  • ssh.mode = "system": the host sshd owns 22 and invokes gitbayd via AuthorizedKeysCommand:

    AuthorizedKeysCommand /usr/local/bin/gitbayd --config /etc/gitbay/config.toml authorized-keys %t %k
    AuthorizedKeysCommandUser gitbay
    

    sshd requires that binary to be root-owned and not group/world writable. Unknown keys fail authentication inside sshd, so system mode requires registration.mode = "closed" (check-config enforces this). Under the forced command the CLI's leading --term=<cols>[,color] argument works as is; a client that sets GITBAY_TERM instead needs AcceptEnv GITBAY_TERM in sshd_config, since that env request is handled by the host's sshd, not gitbayd.

Configuration reference

/etc/gitbay/config.toml. check-config validates and names every contradiction; --no-host-checks skips port/path probes. gitbayd admin config show prints the configuration in effect as TOML, every default filled in and smtp_pass redacted. A file that fails validation still prints, followed by the contradiction.

[server]

  • root (default /var/lib/gitbay) — repositories, database, host keys, ACME cache all live here.
  • site_url (required) — canonical https://host; drives ACME, clone URLs, mail links.
  • secret_key_file (default /etc/gitbay/secret.key) — the keys that seal CI secrets, webhook secrets, mirror tokens and push device tokens in the database. Must be outside root, mode 0600, readable by the daemon's user. See "Secret key" below.
  • source_repo (optional, owner/name) — the repository this instance develops itself in. Startup warns when the running build's commit is not on that repository's default branch, which is how a binary built from an unmerged branch stops being invisible. Leave it unset unless the instance hosts its own source.

[ssh]

  • mode — embedded | system (above).
  • port (22) — embedded listener port.
  • host_keys — list of private key paths; empty generates an ed25519 key at <root>/ssh/host_ed25519.

[http]

  • addr (:443), tls — acme | files | off.
  • acme: certificates via TLS-ALPN-01 on the HTTPS port, cached at <root>/acme; acme_email for the CA account; acme_http_addr (:80, "off" to disable) adds HTTP-01 and an https redirect — failing to bind it is a warning, not fatal. Requires an https:// site_url with a public DNS name.
  • files: cert_file + key_file.
  • off: plain HTTP — development, or behind a TLS-terminating proxy.
  • The HTTPS listener accepts TLS 1.2 and 1.3 only (serverTLS in cmd/gitbayd/tls.go), with the default cipher suites of the Go release the binary was built with. openssl s_client -connect <host>:443 -tls1_1 fails the handshake.

trusted_proxies lists the addresses or CIDRs of reverse proxies in front of the daemon. A request from one of them is attributed, for API rate limiting, to the last X-Forwarded-For hop that is not itself a trusted proxy; from anyone else the header is ignored. Empty, the default, is right when gitbayd terminates TLS itself.

A reverse proxy in front of gitbay must not buffer responses, or the build page's live log arrives only when the build ends; X-Accel-Buffering: no covers nginx.

[web]

  • mode — view_only (default) | accounts. In view_only the mutating web routes are never registered; in accounts, browser sessions are minted over SSH (web login), and users with write access can create repos, comment, and make simple file edits (which commit unsigned, honestly). password_auth is reserved and currently rejected.
  • title — the instance's display name in the rail and page titles; empty falls back to the site host. Lower case is the convention for gitbay itself.
  • privacy_notice — operator text shown on /privacy under the fixed statement.

[registration]

  • mode — closed (default) | invite | open. invite/open require [mail]. See the user guide for the flows.
  • pending_expiry (empty, never) — a duration such as "168h"; a self-registered account still unverified after that long is removed, hourly and at start, audited as pending.expired.
  • notify_admin (false) — mail every instance admin when an account becomes active: an invite redeemed, or an open-mode signup that verified its address. The unverified row an open signup creates is not reported, because anyone can post the form and mailing on that would point a flood at the admins. Recipients are the verified primary addresses of active admins who have activity mail on, the same rule any other notice follows, so an admin with no verified address hears nothing. The notice is queued, so a dead SMTP host shows up in the admin page's Mail table instead of failing the registration. Requires [mail].

[mail]

  • smtp_host (host:port; 587 assumed, 465 with tls = "implicit"), from, optional smtp_user / smtp_pass. Required for invite/open registration and self-service email add; in closed mode you may omit it entirely and assert addresses by hand (below).
  • tls — starttls (default) or implicit (TLS from the first byte, for relays on 465).
  • require_tls — with starttls, a relay that does not offer STARTTLS gets no mail: delivery fails and retries, and the admin page's Mail table shows the error. Unset, it is on for every relay except localhost and loopback addresses; set false to allow plaintext to a remote relay.

[mail.inbound]

Reply by mail: a reply to issue or merge request mail posts a comment (#295). gitbayd polls a mailbox over IMAP; no port is opened for it. Off unless enabled, and it requires [mail] smtp_host, since the replies answer mail gitbayd sends.

[mail.inbound]
enabled       = true
imap_host     = "imap.example.org"       # host:port; 993, or 143 with tls = "starttls"
tls           = "implicit"               # or "starttls"; there is no plaintext setting
user          = "reply@gitbay.example"
password_file = "/etc/gitbay/imap.pass"  # one line, mode 0600, owned by gitbayd's user
mailbox       = "INBOX"                  # default
poll_interval = "1m"                     # default; at least 10s
reply_address = "reply@gitbay.example"
require_dkim  = true                     # recommended; see below
trusted_authserv_id = "mx.example.org"   # the mail host's Authentication-Results id
  • The password is read from password_file at every connection and never appears in the config, argv or the log. An unreadable file, or one readable by group or others, stops the daemon at start.
  • reply_address is what each Reply-To is built from: reply@gitbay.example becomes reply+<token>@gitbay.example. The mailbox must receive that. Hosts that deliver subaddresses to the mailbox need nothing more; others need a wildcard rule. Migadu, for one, sent threads+x@gitbay.org nowhere until a rewrite from threads+* to threads@gitbay.org was added. A domain catch-all only works if it points at this mailbox, which then also holds everything else sent to the domain. Send a test to <reply_address>+test and expect a refused mail reply audit row with "malformed reply token" within a poll interval (gitbay audit --action 'refused mail'). Point the domain's MX at the mail host as for any other mailbox; nothing about it involves gitbayd.
  • Use a mailbox that holds nothing else. gitbayd reads every unseen message in it and marks each one \Seen when it is handled, posted or refused. A message that fails for a reason that may pass (the database busy, the server refusing that one fetch) stays unseen and is tried again on the next poll, up to five times, then is marked seen and audited. A dropped connection or a timeout ends the poll and counts against no message. A message over 10 MiB is refused by its size without being fetched. A server that sends more than about 11 MiB, or more than a thousand untagged responses, for one command has its connection closed; one poll handles at most ten thousand messages.
  • Whatever the settings, a message is refused when a header field name is not RFC 5322 ftext (From : x, a space or a non-ASCII byte in a name), when it does not have exactly one From, or when it has more than one To, Cc, Message-ID, Content-Type or Content-Transfer-Encoding.
  • require_dkim (default false) should be on for any instance reachable from the internet. With it, a reply is posted only when gitbayd itself verifies one of its DKIM signatures (RFC 6376) with a d= in relaxed alignment with the From domain (the same organizational domain by the public suffix list; d=github.io aligns with nothing) and an h= that covers From, the To or Cc holding the reply address, Content-Type, and Message-ID when the message has one. A Content-Transfer-Encoding outside h= is accepted only when it is 7bit, 8bit or binary, which leave the decoded body as it is; Thunderbird, for one, does not sign it. An unsigned quoted-printable or base64 is refused. The reply address is read only from To or Cc: a reply that reached the mailbox by Bcc, with the address only in Delivered-To, is refused ("reply address not in To or Cc"). It needs nothing from the mail host, so it works where the host adds no Authentication-Results. rsa-sha256 (keys of 1024 bits or more) and ed25519-sha256 are accepted, with simple or relaxed canonicalization; rsa-sha1, a body length tag (l=), an expired x= and a t= more than fifteen minutes ahead are refused. Only the first five signatures are checked. The key is looked up at <s>._domainkey.<d> with a five-second timeout and cached for fifteen minutes (the resolver does not report the record's TTL); a lookup that fails for a reason that may pass (a timeout, SERVFAIL) leaves the message for the next poll, up to the five tries above, while a missing key refuses it. Signatures are checked on the message as fetched. Each passing signature is recorded with the Message-ID, so a copy of the same signed message posts once even with unsigned fields changed.
  • Either require_dkim or trusted_authserv_id passing is enough to authenticate From; when both are set, a reply needs only one of them, and a refusal names both reasons. With only trusted_authserv_id, the reply address may come from any recipient field (Delivered-To, X-Original-To, Envelope-To, To, Cc), since the mail host vouches for the sender and not for the fields. Set both when the mail host adds Authentication-Results for most senders but not all.
  • trusted_authserv_id names the authserv-id the mail host writes at the start of its Authentication-Results header (Gmail's is mx.google.com). With it set, a reply is posted only when the topmost header with that id shows dmarc=pass with header.from equal to the From domain, or dkim=pass with a header.d in relaxed alignment with it (the same organizational domain by the public suffix list; a public suffix such as github.io aligns with nothing). The header is parsed per RFC 8601, so text inside a quoted string or a comment (a quoted MAIL FROM local part, a reason) is never read as a result. Lower headers claiming the same id are the sender's and are not read. This is only safe when the mail host removes incoming Authentication-Results headers that claim its id, as RFC 8601 asks; Gmail, Fastmail and Migadu do. Check yours before relying on it. With neither this nor require_dkim set, the daemon logs a warning at start and admin mail inbound check repeats it: From is then whatever the sender wrote. Mail between two addresses at the same host may carry no Authentication-Results at all: at Migadu, mail from another Migadu-hosted domain is delivered through its outbound path and gets none, so setting the id alone refuses every reply from such users. That mail does carry a DKIM signature aligned with From (for example d=cleberg.net; s=key1; a=rsa-sha256; c=simple/simple; h=from:to:subject:date:message-id:mime-version:content-type), which is why gitbay.org sets require_dkim instead.
  • gitbay admin mail inbound check logs in, opens the mailbox read-only (EXAMINE) and reports the message and unseen counts and how From is authenticated (require_dkim, trusted_authserv_id), so a check never marks a reply seen before the poller reads it. It warns when neither is set. With inbound off it says so and exits 0. Poll failures are logged as mail reply: poll failed with the server and the IMAP error.

A reply is posted when all of these hold, checked when it is read:

  1. It is not an automatic reply (Auto-Submitted, Precedence: bulk and the like).
  2. A recipient header carries a reply address whose token verifies and has not expired (thirty days from the mail it came on).
  3. The token's account exists, is active and not disabled, still has reply by mail on, and was created before the token was: an id freed by a delete and reused is not the account the token named. The same holds for the repository.
  4. From is one of that account's verified addresses, and, with require_dkim or trusted_authserv_id set, a DKIM signature gitbayd verified or the mail host's result authenticated it.
  5. The message has a text/plain part (HTML-only mail is refused, not converted), and what is left after quoted text and the signature are removed is not empty and fits a comment (64 KiB).
  6. No earlier copy of the message (same Message-ID, account and thread) posted.
  7. issue comment or mr comment, dispatched as the account, accepts it: the account can still read the repository, the thread exists, the repository is not archived, the write budget is not spent.

The comment's audit row is cmd issue comment (or mr comment) with source: mail. A refusal writes a refused mail reply row with the reason and the Message-ID, never the message's content, at most sixty a minute, and sends nothing back. See Threat-Model for why the token and the sender address are both required.

[push]

Push notifications to Apple devices, delivered by gitbayd talking to APNs directly over HTTP/2, authenticated by an ES256 JWT signed with an operator-supplied provider key. Off unless configured.

  • enabled (false).
  • key_file — path to the .p8 provider key from Apple's developer portal (Certificates, Identifiers & Profiles → Keys). It belongs at /etc/gitbay/apns.p8, mode 0600, owned by the account gitbayd runs as. Read and validated at startup: it must parse as a PEM-wrapped PKCS#8 EC (P-256) private key, or the daemon refuses to start rather than fill a queue nobody is watching.
  • key_id, team_id — the key's id and your Apple developer team id, both from the same portal page.
  • topic — the app's bundle identifier. An APNs key belongs to a bundle ID. gitbay.org pushes to the App Store build under its own bundle id; a self-hoster who wants push ships their own iOS build under their own bundle id, with its own .p8 key from their own developer account, and points topic at that id. There is no way to push to someone else's build, by design — this is Apple's model, not gitbay's.
  • environment — production or sandbox, naming the APNs host rather than taking a URL, so a typo cannot aim the key at a host that is not Apple's.

All five of key_file, key_id, team_id, topic and environment are required when enabled is true; validation runs at config load, so a misconfigured [push] is caught before the daemon serves anything. The delivery queue (a device's undelivered and attempted pushes) is capped the same way the mail queue is, by [retention] push.

[backup]

  • age_recipients (optional) — age public keys (age1...). When set, admin backup encrypts every archive to them and appends .age to its name. Generate the pair off the host with age-keygen; only the public key goes here, so the host writes archives it cannot read. The restic copy is unaffected: the offsite job stages its own VACUUM INTO of the live database and snapshots /var/lib/gitbay, not the archives.

[api]

  • enabled (false) — the JSON API surface; see API. Off means no credential-bearing HTTP endpoint exists at all.

[webhooks]

  • allow_local (false) — permit webhook, mirror and import targets on loopback, private, shared (100.64.0.0/10), link-local or multicast addresses. Leave off unless you know why you need it (SSRF).

[limits]

  • clone_timeout (3600s) — cap on repo import fetches. An import takes http and https URLs only, passes the same address check as a mirror sync and is pinned the same way, so it needs git 2.37 too.
  • max_blob_bytes (100MB) — cap on raw file serving over the web.
  • max_asset_bytes (512MB) — cap per uploaded release asset.
  • max_snippet_bytes (1MB) — cap per snippet file.
  • max_snippets_per_user (0, unlimited) — snippets an account may own.
  • max_repos_per_user (0, unlimited) — repositories an account may own directly; repo create, fork and import refuse past it. Organizations are not capped.
  • max_bytes_per_user (0, unlimited) — disk the account's own repositories may take; a push may be no larger than what is left.
  • pack_concurrency (3), pack_per_principal (2), pack_queue (32), pack_queue_wait ("60s") — git pack generation (clones, fetches, git archive --remote, web archive downloads) over SSH, smart HTTP and git:// shares one budget: this many at once, this many per account (per client address when anonymous: an IPv4 address, or an IPv6 /64), and this many waiting for at most the wait. Anonymous clients together hold at most pack_concurrency − 1 slots when it is above 1: an account (SSH key, bearer token or web session) can take the last slot when it is free, and anonymous clients cannot hold it. Past that an SSH client gets "the server is busy…" and exit 1, HTTP gets 503 with Retry-After: 30, git:// an ERR line; the daemon logs a pack limit warning naming the transport and whether the client was signed in, at most once a minute per transport. An SSH client that disconnects while queued leaves the queue; an HTTP or git:// one keeps its place until the wait runs out. A running clone is killed when its client disconnects, or when no write to the client completes for two minutes: a client reading below about 550 B/s, or an HTTP request body that takes over two minutes with nothing written back, is cut. Ref listings (info/refs, protocol v2 ls-refs), pushes and repo download are outside the budget. For the three counts 0 means the default and a negative value turns that bound off. The defaults suit a four-core host; see Performance. With ssh.mode = "system" each SSH session is its own process and SSH clones are not counted.
  • max_pack_bytes (2 GiB) — the largest pack one push may send, enforced as receive.maxInputSize and lowered to what an owner's storage quota has left.
  • ssh_auth_rate (10) — SSH authentication failures per client address per minute on the embedded listener; see below.

[git_daemon]

  • enabled (false), port (9418) — the anonymous git:// listener. Serves only public repositories that additionally ran repo settings git-daemon <repo> on.

[mirrors]

  • pull_interval_minutes (15) — how often pull mirrors fetch their upstream. Push mirrors sync shortly after each local ref update. Mirror URLs pass the same SSRF rules as webhook targets, when saved and again before every sync; git then connects only to the addresses that were checked (http.curloptResolve) and does not follow redirects, so a mirror of a renamed repository fails until its URL is updated. Needs git 2.37 or later on the server; with an older git the worker logs an error at start and syncs no mirror, recording the reason on each. Sync ignores the system and global gitconfig.

[go_import]

Vanity Go module paths, one per line: "host/module" = "owner/repo". Requests with ?go-get=1 at or under the module path answer with the go-import meta tag pointing at the repository's HTTPS clone URL, so go install host/module/cmd/...@latest resolves. The repository should be public (the module path itself confirms it exists).

Users, email, invites

Every gitbayd admin subcommand except backup, gc and the one-shot backfills is a wrapper that dispatches the registry command of the same name as the host: an admin context with no account behind it, so its audit rows carry no actor and source: host. The same commands run in an instance admin's SSH session (ssh git@<host> admin ...) and write the same rows with the key fingerprint as source. One implementation, two credentials.

gitbayd admin user create alice --key alice.pub --email a@example.org --verified [--admin]
gitbayd admin email verify alice a@example.org   # admin assertion, no SMTP needed
gitbayd admin invite --email b@example.org       # mails a code; prints it if no SMTP

"Verified" means SMTP-confirmed or host-admin-asserted; the database records which. Verified emails are what make commit signatures meaningful — an unverified address never produces a verified badge.

Audit and account control

The audit log is the security feed (events are the product feed): every successful mutating command with its argv and source credential (SSH key fingerprint or API), every refused one (exit 3 or 4) as refused <command>, refused pushes as refused git-receive-pack (the access check) or refused push (a branch or tag rule, a release anchor or an unsigned commit, with the repository and ref names), hook socket requests failing the peer or push-token check as refused hook (never the token), registrations, admin actions, force-pushes, and auth failures/throttling. A refusal row keeps the flag names and the first positional, not the values. Refusals are recorded up to ten a minute per account and 600 a minute across the instance; past either, one refused.throttled row stands for the rest of that minute. The caps bound the embedded listener, the web and the API; under ssh.mode = "system" each gitbayd shell connection counts separately. Secrets never appear — they travel on stdin, never in argv.

Each row carries the SHA-256 of the row before it. gitbayd admin audit verify opens the store as other admin commands do, applying pending migrations, so run it with the binary that matches the daemon. It recomputes the chain and exits 1 naming the first row that was edited or whose predecessor was removed. The chain is unkeyed: whoever can write the database can recompute every hash after an edit, and verify then finds nothing. It catches an edit only when the later hashes were not recomputed. Retention removing the oldest rows is not a break. Rows written before the chain existed are counted and skipped; when every row is such a row, verify warns and exits 1, since clearing the hash columns looks the same. After an upgrade that clears with the first new audit row.

Removing the newest rows leaves no break, and neither do rows written afterwards under the freed ids. The database cannot show either. The daemon logs every row it writes to its journal, outside the database (journalctl -u gitbayd -g 'INFO audit '), and verify prints the last id and hash: comparing them with the newest journal line is the check for any change, recomputed hashes included. Rows written by host gitbayd admin commands, and by gitbayd shell when ssh.mode = "system", are not copied to the journal.

gitbayd admin audit [--actor u|-] [--action prefix] [--since 24h|7d|date] [--limit n] [--json]
gitbayd admin audit verify           # check the hash chain; exit 1 names the first bad row
ssh git@<host> audit ...             # the same, from an admin session
ssh git@<host> admin user list [--state active|pending|disabled|admin]
ssh git@<host> admin user show <name>   # keys, emails, orgs, tokens, sessions
ssh git@<host> admin user limits <name> [--repos n|default] [--bytes n|default]   # per-account caps
ssh git@<host> admin user promote <name>   # grant instance admin
ssh git@<host> admin user demote <name>    # remove it; the last admin is refused
gitbayd admin user promote <name>    # host-local: recovery when no admin key is reachable
gitbayd admin user disable <name>    # suspend: SSH, web, API all refused;
gitbayd admin user enable <name>     #   sessions dropped, nothing deleted
gitbayd admin user delete <name> --yes  # only for accounts anchoring nothing:
                                     #   refused (with each blocker named) while
                                     #   the account owns repos, authored
                                     #   issues/MRs/comments/reviews, or is an
                                     #   org's only admin

--actor takes a username, or - for rows with no actor: host commands and auth failures. --action is a prefix, so cmd repo catches every repository command and admin every host or admin-session action. --since is a duration back from now (30m, 24h, 7d) or a date.

admin user list pages by username (--limit, --cursor) and carries each account's state and last_seen, the newest use of any of its SSH keys or API tokens. admin user show adds the keys with their last use, each address with how it was verified, PGP keys, org roles, the owned repository count, API token names, and live browser sessions. Both are Both are refused to non-admins, like audit, on every surface.

/admin/users is the same list in a browser, linked from the admin page: the state filter the command takes, keyset paging on its cursor, and a row per account with promote, demote, disable and enable, each dispatching the command. Demote and disable ask for the username to be typed, since both take someone's access away. Creating and deleting an account, issuing an invite and asserting an address stay on the command line: each takes a key, mints a credential, or cannot be undone. A non-admin gets the 404 a missing page would, so the URL confirms nothing.

Promotion needs an active account: a pending or disabled one is refused. Demotion is refused when it would leave no admin, over SSH, in the browser and on the host alike, so the host-local promote is the way back in when the only admin key is lost.

Instance admin carries no right on anyone's repository: policy does not consult it, and a private repository still answers not-found to an admin. Moderation goes through explicit overrides that skip the access check and write their own audit row:

ssh git@<host> admin repo list [--owner o] [--visibility public|private]  # size, last push
ssh git@<host> admin repo archive|unarchive <owner/name>
ssh git@<host> admin repo visibility <owner/name> public|private
ssh git@<host> admin repo delete <owner/name> --yes

Each lands in the audit log as admin repo.<action> naming the repository, on top of the cmd row every mutating command gets.

limits.ssh_auth_rate (10) throttles per-IP authentication failures per minute — successful auths never count and clear the slate. limits.max_pack_bytes is enforced as receive.maxInputSize on every push.

limits.write_rate (60) bounds mutating commands per account per minute. It is counted in the dispatcher, so SSH, the JSON API and the web spend one budget and a caller cannot refresh it by changing surface; limits.api_rate stays in front of it, bounding a network source rather than an account. A command is one token whatever it writes, so a bundle import costs one and only a loop of separate commands spends the budget. Read-only commands, the runner protocol (a build streams its log in many small writes) and the host CLI are exempt. Refusals exit 4 and say when to retry. A negative value turns the limit off; it matters most with registration = "open", where every write also queues notification mail and webhook deliveries.

Queues

Every background worker keeps a backlog and a failure state. An instance admin reads them all in one place:

gitbay dashboard --json | jq .queues   # webhooks, mail, push, mirrors, builds, deps

Per worker: pending, retrying (pending with a failed attempt) and dead-lettered counts with the oldest pending age, and the retrying or failed rows themselves, capped at twenty each. Builds list what is running and then what is pending, each since when, so a build no runner is scoped to claim is visible here rather than only in its repository; mirrors list the ones whose last sync failed; dependency checks list the ones whose last check errored. Non-admins get no queues key at all.

Push rows name the device id, never the token. Watch this one after configuring [push]: a key_id or team_id Apple did not issue passes config validation, which can only check that the .p8 parses, and then every send comes back 403 InvalidProviderToken and dead-letters on its first attempt.

In accounts mode the same read renders at /admin, linked from the rail for admins. Anyone else gets a 404 there.

A dead-lettered mail is logged as notification dead-lettered mail=<id>, with the address redacted out of the relay's error. The id is the queue row: find it in the Mail table on /admin, or in dashboard --json, where the recipient and the unredacted error are. That is deliberate — see the Threat-Model page.

Maintenance

gitbayd admin stats [--json]         # counts, database size, per-repo disk
ssh git@<host> admin stats [--json]  # the same, from an admin session
gitbayd admin gc [--repo owner/name] # git gc: repack and prune; per-repo sizes
gitbayd admin gc --aggressive        # thorough repack; slow, rarely needed
gitbayd admin gc --lfs               # also drop LFS objects no pointer names (older than a day)

A history rewrite leaves the commits it removed reachable through refs/merge-requests/N/head of the merge requests that landed them, so they stay fetchable by anyone who can read the repository. Nothing drops a head ref on its own — an open or source-gone MR is merged through it, and a merged or closed one keeps its diff readable through it — so the cleanup is a command an instance admin runs, naming the MRs:

ssh git@<host> admin mr prune owner/name 1 2 3 --yes

It refuses an open or source-gone MR, deletes the named refs, runs git gc --prune=now on that one repository so the objects go at once rather than after git's two-week grace, leaves a system comment on each MR, and audits as admin mr.prune. The MR keeps its title, comments, reviews and head sha; mr diff and the MR page say the head is gone. Run it when nothing is pushing to that repository: without the grace, a push caught between leaving quarantine and writing its ref loses its objects. Objects also survive in offsite backups until those are rewritten; see "Removing a repository's history from every snapshot".

deploy/cloud-init.yaml ships a gitbay-gc.timer that runs admin gc weekly (Sunday 07:00 UTC). Imported repositories keep whatever pack layout the source sent, so a first manual admin gc after a bulk import is worthwhile.

Backup and restore

gitbayd admin backup --out /var/backups/gitbay/backup.tar.gz
gitbayd admin backup --verify /var/backups/gitbay/backup.tar.gz   # read it back

One archive: a consistent SQLite snapshot (taken before the repositories are read, so the database never references objects the archive missed), every repository, and the SSH host keys. Excluded: hook socket, regenerated hook scripts, WAL files. Safe to run against a live daemon. --out must be outside server.root, or the next full backup would carry the archive. Each run removes the snapshot directories (.gitbay-snap-*) and temporary archives (.*.tmp-*) a killed run left beside its archive once they are a day old, and prints each one it removes.

--verify reads an archive back: the snapshot must pass SQLite's integrity check, every repository the snapshot names must be in the archive, and each must pass git fsck --connectivity-only. Every release asset the snapshot names must be in its repository with its recorded size and sha256, and every LFS object in the archive must hash to its name. LFS objects are named by pointer files in git history rather than by the database, so a pointer whose object was never uploaded is not caught, and with [lfs] root outside server.root the archive holds no LFS objects at all. It extracts the repositories to a temporary directory for the checks, so it needs free space the size of the repositories. A database-only archive is checked for integrity and says so. Exit is non-zero on damage, a missing repository, a missing object, or a missing or altered asset or LFS object.

A full backup holds <root>/backup.lock from its database snapshot to its last repository. While it runs, repo delete, repo rename, repo transfer, admin repo delete, org rename, admin mr prune and gitbayd admin gc refuse with "a backup is running"; retry when it finishes. Database-only backups take no lock. A pack that git's own automatic gc removes during the walk is skipped; each repository's refs are archived before its objects, so the refs still find their objects, and --verify reports it if one does not.

Verifying an archive of unknown origin: run it as an unprivileged user. The connectivity check runs git with --git-dir on each extracted repository, so a directory that is not a repository fails rather than git checking an enclosing one. objects/info/alternates and a commondir directly in a *.git directory are not extracted, so an archive cannot use either to have git read another repository's objects or refs on the host. Git still reads each archived repository's own config.

Archives carry a directory entry for every directory, including an empty one, so a bare repository whose refs are all packed restores as a repository. Archives written before this release do not: extracting one can leave a repository's refs/ directory missing, which stops git from recognizing it as a repository at all. gitbayd admin backup --verify <archive> names the repositories this affects; the fix is mkdir -p <root>/repos/<owner>/<name>.git/refs for each one, after which it opens normally.

With [backup] age_recipients set the archive is <name>.tar.gz.age and --verify needs the private key:

gitbayd admin backup --verify gitbay-20260927-090000.tar.gz.age --identity ~/.config/gitbay/backup-identity.txt
age -d -i ~/.config/gitbay/backup-identity.txt gitbay-20260927-090000.tar.gz.age | tar -xz -C /new/root

The identity lives off the host (with the secret key file and the restic credentials), so verifying an encrypted archive happens there or on a restore host.

Restore: extract into an empty directory, point server.root at it, restore server.secret_key_file from its own copy (mode 0600, owned by the account gitbayd runs as), then start gitbayd. No archive carries the key file, and without it gitbayd refuses to start. Host keys are preserved, so clients keep their known_hosts entries; hooks regenerate at startup.

Schedule and recovery point

Two timers, because the two halves of the data have different exposure.

  • gitbay-backup.timer, nightly. The full archive above, last 7 kept.
  • gitbay-db-backup.timer, hourly. admin backup --db-only, which writes the SQLite snapshot alone, last 48 kept. A few MB against the full archive's hundreds, which is what makes the frequency affordable.

The split follows what a loss would actually cost. Repositories are git, so a mirror or any clone is a second copy; the database is the only copy of issues, merge requests, comments and review state. So the recovery point is about an hour for the data that exists nowhere else, and a day for the data that does.

Continuous replication (litestream and similar) was considered and not adopted. It would take the database's recovery point to seconds, but the repositories would still be on the nightly archive, so a restore could produce a database referencing commits the repository backup does not have. Consistency between the two halves is worth more here than latency on one of them. Revisit if repository replication becomes continuous too.

Offsite copies

bay1 also takes a nightly restic snapshot of /var/lib/gitbay and /var/lib/gitbay-stage (a database copy and config.toml, staged there) to an S3 bucket at Scaleway, with a key that can only add snapshots. Nothing else under /etc/gitbay is in it: not the secret key file, not apns.p8. The key that can remove them lives on the operator's machine, in ~/.config/gitbay/offsite.env, and never on bay1: a compromised host cannot destroy its own history. Forgetting, pruning and rewriting all run from there.

Removing a repository's history from every snapshot

A history rewrite plus admin mr prune takes commits off the server, but every snapshot taken before it still holds them, and the retention window is the only thing that ages them out. To remove them now, rewrite the snapshots without that repository rather than forgetting the snapshots: everything else in them stays restorable. The next nightly run adds the repository back in its current state.

Repositories are stored under the name they had on disk when each snapshot was taken, so a renamed repository needs every name it has carried. Check what an older snapshot holds before choosing the paths:

set -a; . ~/.config/gitbay/offsite.env; set +a
restic $RESTIC_OPTS snapshots
restic $RESTIC_OPTS ls <old-snapshot> /var/lib/gitbay/repos/<owner>

Then dry-run, apply, prune, and confirm nothing matches:

EXCL="--exclude /var/lib/gitbay/repos/<owner>/<name>.git --exclude /var/lib/gitbay/repos/<owner>/<old-name>.git"
restic $RESTIC_OPTS rewrite --dry-run $EXCL     # "would modify N snapshots"
restic $RESTIC_OPTS rewrite --forget $EXCL      # new snapshots replace the originals
restic $RESTIC_OPTS prune                       # drops the data nothing references
restic $RESTIC_OPTS find <name>.git <old-name>.git   # expect no output

--forget is what makes the originals go; without it the rewritten snapshots sit beside them and the data stays referenced. Snapshot IDs change; their times do not. Done for krz/keycask (formerly rust-pass) on 2026-09-18, across 22 snapshots.

Secret key

CI secrets, webhook secrets, mirror tokens and APNs device tokens are stored sealed: AES-256-GCM under a key in server.secret_key_file, each value prefixed with the id of the key that sealed it (gbs1:<id>:). The key file is not in the database, not under server.root, and therefore in neither the local archives nor the main restic repository. It must be copied off the host separately; without it a restored database's secrets cannot be opened, and gitbayd refuses to start against them. A separate restic repository for it and apns.p8 is planned (runbook D of the data-at-rest plan) and not yet in place, so today the only off-host copy is one the operator makes by hand after init and after every rotate.

gitbayd admin secrets init     # once; deploy/install.sh does it on first install
gitbayd admin secrets check    # open every value, count by key
gitbayd admin secrets rotate   # new key, reseal, retire the old one (as root)
  • Missing file: every gitbayd process that opens the database refuses to run and names the path, including serve and, in system mode, authorized-keys. migrate does not need it.
  • Backups: admin backup and admin backup --verify do not need the key file; a restore does.
  • Wrong key: serve stops at startup naming the first row that does not open; secrets check does the same without starting anything.
  • Upgrade: the first start after the upgrade seals every value still in clear and logs sealed secret values.
  • Rotation: rotate adds a key, reseals every value under it in one transaction, then removes the old keys. Run it as root, since it replaces the key file in /etc/gitbay; the file keeps its owner. The daemon re-reads the file when it changes, so it needs no restart. Copy the new file off the host afterwards. init and rotate hold an flock on <key file>.lock while they run, so a second run waits for the first.
  • Push devices are looked up by the SHA-256 of their token (push_devices.token_hash), since two seals of one token differ.

Restore drill

A restore onto a clean host, run quarterly (January, April, July, October) and after any change to the backup code (cmd/gitbayd/backup.go, cmd/gitbayd/restoredrill.go, the offsite job), and recorded below. The disaster it rehearses is losing bay1, so the local archives are gone with it and the sources are the main offsite restic repository (repositories, LFS, the staged database, config.toml), the off-host copy of secret.key and apns.p8 (a keys repository once runbook D creates it; until then the operator's hand-made copy), and the operator's password manager (offsite.env, the keys repository's password and token once it exists, backup-identity.txt). The full steps are runbook C of the data-at-rest plan (docs/plans/2026-09-27-data-at-rest-and-backup.md).

gitbayd admin restore-drill does the archive half. It extracts a full archive into an empty or absent directory, runs every --verify check on the extracted copy, and prints what was restored, the newest issue, issue comment, merge request comment and push in the restored database, and the elapsed time. Exit is non-zero if any check fails.

gitbayd admin restore-drill /var/backups/gitbay/gitbay-20260927-090000.tar.gz --into /srv/drill
gitbayd admin restore-drill <archive>.tar.gz.age --identity backup-identity.txt --into /srv/drill

What to restore, and from where:

Item Source Path on the drill host
Database staged copy in /var/lib/gitbay-stage (restic), or an archive <root>/gitbay.db
Repositories /var/lib/gitbay/repos (restic), or an archive <root>/repos
LFS objects /var/lib/gitbay/lfs (restic), or an archive <root>/lfs (or [lfs] root)
Release assets inside each repository (gitbay-releases/) with the repositories
Host keys /var/lib/gitbay/ssh (restic), or an archive <root>/ssh (or [ssh] host_keys)
config.toml /var/lib/gitbay-stage/config.toml (restic) /etc/gitbay/config.toml
secret.key, apns.p8 the off-host copy; no archive or main snapshot carries them /etc/gitbay/, mode 0600, owned by gitbay

From the offsite path (restic is not run by any gitbay command):

restic restore latest --target / --include /var/lib/gitbay --include /var/lib/gitbay-stage
cp /var/lib/gitbay-stage/gitbay.db /var/lib/gitbay/gitbay.db
gitbayd --config /etc/gitbay/config.toml admin backup --out /tmp/drill.tar.gz
gitbayd admin restore-drill /tmp/drill.tar.gz --into /tmp/drill-root

The second archive is how the restic tree gets the same checks and timestamps; /tmp/drill-root is discarded afterwards.

What to check, each a column below:

  • DB integrity, connectivity, release assets, LFS: restore-drill prints integrity ok, connectivity ok on N repositories, release assets ok: N, LFS objects ok: N. Compare N with gitbayd admin stats --json on the source at the snapshot time.
  • Secrets: gitbayd --config /etc/gitbay/config.toml admin secrets check opens every value.
  • Host key: ssh-keyscan -p 22 <drill-host> matches the source's fingerprint.
  • Config: gitbayd --config /etc/gitbay/config.toml check-config.
  • Service: ssh -p 22 git@<drill-host> whoami and a git clone over SSH succeed.

Time to service runs from the clean host's first root login to the first successful git clone over SSH from it. The recovery point is the time of the newest restic snapshot restored; record beside it the newest issue, comment and push restore-drill printed, which show how much activity the restore carries.

The first drill ran on the operator's laptop rather than a provisioned host, so its time to service has no provisioning in it, and it could not check secrets: no copy of secret.key was on the machine, and gitbayd refuses to start while any sealed value does not open. The sealed values were cleared in the drill copy to reach service. The off-host copy of the key is confirmed to exist (#305); the next drill restores it with the snapshot and runs admin secrets check on the restored copy.

Date Host Snapshot restored (UTC) Newest issue / comment / push Time to service DB integrity Connectivity LFS Release assets Host key Secrets Notes
2026-09-29 laptop (macOS), restic from offsite 2026-09-29 00:19:15 (ccd646cf) 00:05:17 / 00:09:44 / 00:15:31 8m43s (restore 1m49s, checks 33s) ok ok, 69/69 ok, 2 ok, 529 matches not checked (#305) 4.99 GiB restored; 71 sealed values cleared in the drill copy to start

Upgrades

Replace the binary, restart the unit. Migrations apply automatically and are transactional; hook scripts under <root>/hooks are rewritten at startup to point at the current binary path.

CI runner

gitbay-runner executes builds queued by pushes and merge requests. It polls over SSH with a key of scope runner, which reaches only the runner protocol and read-only git (a runner executes arbitrary repository code, so the key it holds must not do more). A runner key claims builds only for the repositories it is attached to, by repo runner add from a repository admin or an instance admin; an admin key claims any. Users attach their own runners: see the Users page. For an instance runner, run it as a dedicated unprivileged user on a non-admin account. admin user create --key registers a full-scope key, so the runner key is added afterwards through a bootstrap key that is then removed, and attached to each repository it should build:

useradd --system --create-home --home-dir /var/lib/gitbay-runner ci-runner
sudo -u ci-runner ssh-keygen -t ed25519 -N "" -f /var/lib/gitbay-runner/.ssh/id_ed25519
ssh-keygen -t ed25519 -N "" -f /tmp/ci-bootstrap
gitbayd --config /etc/gitbay/config.toml admin user create ci --key /tmp/ci-bootstrap.pub
ssh -i /tmp/ci-bootstrap git@127.0.0.1 keys add --scope runner < /var/lib/gitbay-runner/.ssh/id_ed25519.pub
ssh -i /tmp/ci-bootstrap git@127.0.0.1 keys remove "$(ssh-keygen -lf /tmp/ci-bootstrap.pub | awk '{print $2}')"
rm /tmp/ci-bootstrap /tmp/ci-bootstrap.pub
gitbay-runner -remote git@127.0.0.1 -workdir /var/lib/gitbay-runner/work
gitbay repo runner add krz/site < /var/lib/gitbay-runner/.ssh/id_ed25519.pub

-jobs N runs N builds at once. Claiming is one transaction that selects and updates, and each build works in its own build-<id> directory, so workers do not collide; idle polls are staggered across the interval so N of them do not wake together. The drop-in's weights below are per service, not per build, so raising -jobs divides them rather than multiplying the host's load.

admin runners shows which account each runner polls as, and what each may claim. A runner key claims builds only for the repositories it is attached to: with none attached it claims nothing, and -repos may only narrow within them. An admin's full-scope key claims any repository — that is what -repos was for — and still works for the protocol during a rotation. A merge request head from a fork is built in the target repository as untrusted: the claim carries no secrets, and only a runner started with -untrusted takes it. Same-repository heads were built by their branch push and are not built again.

make deploy-runner also installs deploy/gitbay-runner.override.conf as a systemd drop-in: Nice=10, CPUWeight=30, IOWeight=30, so a build never starves the host's sshd, the daemon or the backup timers, and NoNewPrivileges, ProtectSystem=full, ProtectKernelTunables, ProtectControlGroups and RestrictSUIDSGID, so a step cannot reach outside its workspace and the runner's home. The e2e suite alone starts sixty daemon instances; without the drop-in a deploy's copy over the admin sshd stalled. Both deploy targets copy with rsync --partial, which resumes a stalled transfer.

Under -isolation podman, the default and what bay1 runs, each build is confined to a container (see Container isolation below). Under -isolation none steps run directly on the host as the runner's user, so treat that machine as executing whatever your users push, and install the toolchains your builds need on it.

A runner claims the oldest pending build among the repositories its key is attached to — for an admin key, the oldest in the instance. -repos narrows within that set, which is what makes a runner outside the server practical: one on a machine that should build a single project, or that holds credentials for one deployment, stays on it.

Oldest-first is across everything the key may claim, so a repository with a deep queue holds every other repository the same runner serves; bay1 measured a 15-minute average wait on a day of merge request stacks from one repository. A runner attached to one repository cannot be starved. That is the rule, decided in krz/gitbay#207: a runner serving several repositories takes them oldest-first, and an operator who wants one repository never to wait on another runs a second runner attached to it alone. Nothing caps what an account queues, and nothing needs to (krz/gitbay#206): a schedule tick queues nothing while the job's last build is pending or running, so a repository with no runner holds one row per scheduled job rather than one per tick, and a build runs only on a runner its owner attaches, so a busy schedule spends the owner's compute. Pushes are bounded by what an account can push.

gitbay-runner -remote git@gitbay.org -repos krz/site,krz/docs \
  -workdir /var/lib/gitbay-runner/work

Add -untrusted only with -isolation podman.

gitbay.org's runner is attached to the forge's own repositories and the isolation canary, nothing else, because it shares the host with the forge; its unit names no -repos, the attachments are the boundary. Any other repository builds on a runner its owner attaches.

-repos narrows an admin runner; for a runner key the attachments are the boundary, held by the server, and -repos may only name repositories among them. -untrusted makes a runner claim merge request heads from forks; the bay1 unit sets it because it isolates in podman. A runner without it builds trusted commits only.

gitbay dashboard and ssh git@<host> admin runners list every key that has polled as a runner: the account, the key's fingerprint, when it last polled, the repositories it may claim — its attachments for a runner key, the -repos it asked for or any for an admin key — and the build it holds; admin runners remove <fingerprint> (forget until the next release) drops the row for a key that polled by mistake, the key itself untouched. admin runners also heads the list with the queue: builds pending now, and over the last day how many were claimed, how long they waited to be claimed (average and worst), and how many the reaper ended instead of a runner reporting them. A build a runner claimed and never reported is failed by the scheduler's minute tick, whether or not any runner is still alive: within about two minutes of its log stream ending with no outcome reported — the runner reports right after closing the stream, retrying for half a minute if gitbayd is unreachable — or, if no stream was ever seen, at the deadline.

Instance admin on the runner account only authorizes the claim/report protocol; it grants no repo access. A build that pushes back — a pages deploy, an archive publish, an automated MR branch — needs an explicit grant on that repo: repo access grant <owner/name> ci write. Private repos likewise need at least read for the clone.

Container isolation

Builds run in a rootless podman container, one per job, with the workspace bind mounted and nothing else: the clone happens outside with the runner's key, so a step cannot read it. -isolation none keeps the old behaviour — steps on the host as the runner's user — for an instance where every repository is trusted. There is no automatic fallback: a runner started with -isolation podman that cannot find a working podman exits rather than running a build unsandboxed.

gitbay's own jobs name localhost/gitbay-ci:2, built from deploy/Containerfile.ci on the runner host. A job's image must carry what its steps need: the suite drives real git, git-lfs, gpg and sshd and asserts they exist before running, so the stock runner default would fail it immediately. Build or rebuild it with:

ssh -p 2222 root@<host> 'cat > /tmp/Containerfile.ci' < deploy/Containerfile.ci
ssh -p 2222 root@<host> 'su - ci-runner -s /bin/sh -c \
  "podman build -t localhost/gitbay-ci:2 -f /tmp/Containerfile.ci /tmp"'

The tag is deliberate rather than :latest: changing the file means bumping the tag in .gitbay/ci.yml, so a running branch's image does not change under it.

-image names the image a job runs in when it declares none, and is required under -isolation podman: there is no built-in default, because an image this host does not have would fail every build. A job overrides it with image: in .gitbay/ci.yml, validated as a reference so a config file cannot turn it into podman arguments.

-cpus and -memory cap one build (podman's units, e.g. -cpus 2 -memory 4g); unset means uncapped. The runner applies them itself: it creates a cgroup per build under its own delegated service cgroup, writes the limits there, and starts every podman process for the build inside it, with podman's cgroup handling off. Podman's own --memory and --cpus never applied under rootless cgroupfs, which is what a system service gets (krz/gitbay#188). The unit therefore needs Delegate=yes, which the drop-in sets; without it the runner refuses to start when a limit is set, and logs that builds run unconfined when none is. bay1 runs -cpus 3 -memory 6g per build inside MemoryMax=6G and CPUQuota=300% on the unit, on a 7.7GB four-core host with no swap: the memory cap is what keeps the forge alive when a build allocates without bound, and it sits above the e2e suite's 5GB peak rather than at a fair share. OOMPolicy=continue keeps systemd from stopping the runner when a build is OOM-killed.

A trusted build's home is its repository's, under <workdir>/trusted-home/<owner>/<name>, mounted into its containers as HOME: caches persist between trusted builds of one repository and are never read by another's. An untrusted build — a merge request head from a fork — gets <workdir>/build-<id>-home, new and empty, removed when the build ends. Homes under <workdir>/home are from runners before krz/gitbay#255, which shared them with untrusted builds; nothing reads them any more, and they can be deleted.

Images are provisioned, never pulled by a build. The runner passes --pull=never. Two reasons, and the second is the better one: the service runs with RestrictSUIDSGID=yes so podman cannot unpack a layer holding a setuid file, which is nearly every distribution image; and on an instance where anyone can push a ci.yml, image: would otherwise mean "fetch and run anything from the internet". An operator pulls or builds what is allowed and a build picks among those. A job naming an image the host does not have fails with a message saying so.

su - ci-runner -s /bin/sh -c "podman pull docker.io/library/alpine:3.20"
su - ci-runner -s /bin/sh -c "podman images"

Prepare a host before pointing an isolating runner at it:

ssh -p 2222 root@<host> 'sh -s' < deploy/runner-podman-setup.sh
make deploy-runner

The script installs podman, delegates a subuid/subgid range to ci-runner, checks that user namespaces are enabled rather than assuming, enables lingering, and verifies rootless podman actually runs as that user. It is idempotent.

It also installs nftables. make deploy-runner ships deploy/gitbay-runner-egress.nft to /etc/gitbay-runner/egress.nft with gitbay-runner-egress.service, which loads it and which the runner's unit requires; it checks the file with nft -c, reloads the unit, and runs deploy/runner-egress-check.sh as ci-runner before restarting the runner: 127.0.0.1:22 and the public 22 must answer, 2222 must not. The table limits the runner's user to 127.0.0.1:22, DNS on loopback, and 22, 80 and 443 on the host's public address; the Threat-Model page says why.

That table cannot tell a build from the runner, since both run as ci-runner. deploy/gitbay-runner-builds.nft, shipped as /etc/gitbay-runner/builds.nft, matches by cgroup instead: the runner starts each build under builds/trusted or builds/untrusted in its service cgroup, and the table closes the host's loopback to every build, limits an untrusted build to TCP 80 and 443 and DNS, and closes private ranges to both (the CI page has the table). nftables resolves a cgroup path to its id when the table loads, and the service cgroup is new on every start, so the drop-in's ExecStartPre creates the two cgroups, hands them to ci-runner and loads the table before the runner starts; a table that does not load stops the start. make deploy-runner checks the file's syntax before the restart and that the table is loaded after it. If /etc/resolv.conf names a nameserver in a private range, allow it in the file first or builds resolve nothing.

A restart of nftables.service flushes both tables; systemctl reload gitbay-runner-egress restores both.

The egress unit's stop and the rollback below use nft destroy, which needs nftables 1.0.8 or later; check nft --version on a new host.

Checking the tables on a running host, during a build (the flood test below waits a minute before its first probe, for this):

nft list table inet gitbay_builds          # counters on the reject rules
for p in $(pgrep -u ci-runner pasta); do cat /proc/$p/cgroup; done
# 0::/system.slice/gitbay-runner.service/builds/trusted/build-<id>

A pasta process anywhere else — the runner's own runner cgroup, a user slice — means the builds table does not see that build's traffic and only the first table applies. Run deploy/runner-auth-flood-test.sh on the scratch repository (below), once as a push and once with --untrusted. It fails if the runner was locked out or if the build log lacks the lines only a working table produces. The authoritative proof that pasta's sockets are in the build's cgroup is the untrusted run: 169.254.1.2:22 and github.com:22 refused, all twelve logins refused, and the counters on the untrusted chain's rejects rising. The first table lets ci-runner reach both of those addresses.

To take the builds table out, on the host:

sed -i '/^ExecStartPre=+.*builds/d' /etc/systemd/system/gitbay-runner.service.d/override.conf
rm /etc/gitbay-runner/builds.nft
systemctl daemon-reload
nft destroy table inet gitbay_builds

The runner needs no restart. It still places builds under builds/trusted and builds/untrusted, but with the file gone neither a runner start nor systemctl reload gitbay-runner-egress loads the table again, and builds keep the first table and --no-map-gw. A runner that takes -untrusted or polls over loopback refuses to start without build cgroups at all, since the table would match nothing. The next make deploy-runner installs the file and the lines again.

The drop-in sets NoNewPrivileges=no, without which rootless podman cannot call newuidmap and the runner refuses to start. That is a considered trade, explained in the file and in the Threat-Model; if you run with -isolation none, set it back to yes.

Restarting the runner is safe. On SIGTERM it stops claiming, finishes the build in flight, reports it, and exits; the drop-in's TimeoutStopSec=50min covers the longest build, and its KillMode=mixed is what makes the signal reach the runner alone — under systemd's default the build's container and the log session are signalled with it, and the runner drains a build that is already dead. So make deploy-runner waits for a running build rather than orphaning it, and a build's result is retried for half a minute if gitbayd is restarting at that moment. A second SIGTERM ends the runner at once, abandoning the build to the reaper. The suite checks all three: TestRunnerDrainsOnSIGTERM signals the process, TestRunnerDropInLetsTheDrainHappen reads the drop-in's KillMode and TimeoutStopSec, and TestRunnerDrainUnderSystemd runs the runner as a transient user unit under systemd-run and stops it under both kill modes. That last one needs a systemd user manager, so it skips in the container CI runs in; run it on a Linux host with go test ./e2e -run TestRunnerDrainUnderSystemd -v.

Validate podman mode on a scratch repository before pointing the runner at real ones. Every deploy that switched the whole instance to containers and failed took CI down with it. Instead: create a throwaway repository the runner account can read (public, or granted read — a private one is "not found" to the runner and the build stays pending), give it one job that names the CI image, attach the runner's key to it, and deploy the runner with -repos naming only that repository. The production unit, with its real hardening, then claims nothing else; other repositories' builds queue until -repos is removed again, which is a pause, not an outage.

# on a machine with an admin gitbay identity (root on the host has none),
# with the runner's public key copied from the host:
gitbay repo create cmc/runner-scratch   # then push a .gitbay/ci.yml naming the image
scp -P 2222 root@<host>:/var/lib/gitbay-runner/.ssh/id_ed25519.pub runner.pub
gitbay repo runner add cmc/runner-scratch < runner.pub
# on the host:
sed -i 's#^ExecStart=/usr/local/bin/gitbay-runner #&-repos cmc/runner-scratch #' /etc/systemd/system/gitbay-runner.service.d/override.conf
systemctl daemon-reload && systemctl restart gitbay-runner
# back on the admin machine:
gitbay build log cmc/runner-scratch 1   # green: remove -repos, redeploy, delete the scratch repository

Do not deploy an isolating runner to a host that has not been prepared. The runner is specified to refuse to start without a working podman rather than fall back to running builds unsandboxed — a fallback that silently drops isolation is worse than a stopped runner, because nothing surfaces it. On an unprepared host that refusal stops every build on the instance.

The service drop-in carries Delegate=yes for rootless cgroup management and ReadWritePaths for podman's store under /var/lib/gitbay-runner, which ProtectSystem=full would otherwise make read-only. Those paths are prefixed - so they are ignored when absent: the drop-in installs on unprepared hosts too, and a unit that refused to start would stop every build. gitbay-runner-prune.timer prunes unused images weekly, as the runner's user: rootless storage belongs to that user, and root's prune would not see it. An unpruned image store on a 40GB host is a slow outage.

LFS storage

Objects live content-addressed under [lfs] root (default <server.root>/lfs); [lfs] max_object_bytes caps a single object (512MB default). Storage sits behind a small interface — an S3-compatible backend is a drop-in with the server proxying, and presigned URLs a later optimization. LFS objects do not travel with push mirrors (mirrors move git refs only), and gc does not yet collect orphaned objects.

Pages

[pages] domain = "example.site" serves public repos' pages branches on <owner>.<domain>. DNS needs a wildcard record *.<domain> to the server; ACME issues per-subdomain certificates on demand (only for owners that exist). The domain must not be the site host or a parent of it — pages content runs its own scripts and must stay off the forge's origin.

Users with repo admin claim custom domains with repo domain add. Claims activate only after a DNS TXT challenge proves control of the domain (repo domain verify, audit-logged); pending claims serve nothing, get no certificates, and expire after 7 days. ACME issues certificates only for verified hosts, so stray DNS pointed at the server gets nothing.

Security

The Threat-Model file is the reference for what the forge trusts and refuses to do. Operational checklist:

  • Software checks. deploy/audit.sh runs go vet, govulncheck (the module list is deliberately short — review it on each release), and a short fuzz pass over every attacker-facing parser (pkt-line, commit, SSHSIG armor, OpenPGP key, SSH tokenizer). Run it before tagging a release. CI's own vuln job runs govulncheck nightly against main rather than per push, because @latest scans today's advisory database and an advisory lands without anyone pushing; build trigger krz/gitbay vuln runs it on demand.
  • Web responses carry a scripts-forbidden CSP, X-Frame-Options: DENY, nosniff, no-referrer, and HSTS when TLS is on — no configuration needed.
  • Host sandboxing. The systemd unit in deploy/cloud-init.yaml runs gitbayd unprivileged with ProtectSystem=strict, PrivateDevices, LockPersonality, MemoryDenyWriteExecute, SystemCallFilter=@system-service, and RestrictAddressFamilies to INET/INET6/UNIX. It keeps CAP_NET_BIND_SERVICE only, to bind 22/80/443.
  • OS patches apply via unattended-upgrades (security origins, auto-reboot 04:30 if required).
  • Admin sshd (2222) is throttled by MaxStartups=/=MaxAuthTries and watched by fail2ban; gitbayd's own port 22 is throttled by limits.ssh_auth_rate (auth failures per IP per minute), and every account's writes by limits.write_rate.
  • Monitoring. gitbay-monitor.timer writes a reading hourly to

journald and, when /etc/gitbay/monitor.url exists, posts it to that webhook: disk, service, the daemon's own /healthz answer, certificate expiry, and the age of the newest full backup and database snapshot. It exits non-zero on an alert so the unit shows in systemctl --failed: a stopped service, /healthz not answering ok, disk ≥ 85%, a certificate under 21 days, a full backup over 25 hours old, or a database snapshot over 2 hours old.

GET /healthz is unauthenticated and cache-free: whether the database answers and which commit serves, 503 when it does not.

  • Database. gitbay.db and its WAL live under /var/lib/gitbay (mode 0750, owned by gitbay). The nightly archive plus provider snapshots are the recovery path; for tighter RPO, add continuous replication (litestream) against the same file — it coexists with the WAL.

Odds and ends

  • deleting a fork marks MRs sourced from it source_gone; their diffs remain viewable and mergeable because the target repo owns the objects.
  • refs/merge-requests/* is server-owned and unpushable by clients; only admin mr prune removes one (see Maintenance).
  • audit-relevant activity (issue/MR lifecycle, imports, pushes) lands in the events table, which also feeds webhooks.
  • the daemon idles under 10MB RSS; the smallest VPS tier is adequate.