krz/orgo

Lightning fast org-mode static site generator. fast go org-mode static-site-generator

Commit d6911e2a61

d6911e2a61dca2635a9a23f5db41be4239ead50e

parent: f828151d5f

Verified · cmc

cmc <hello@cleberg.net> · 2026-08-11 04:37 UTC

Build the nav from top-level pages only

The nav listed every page, so an n-page site emitted n^2 nav links. At 1,790 pages each
page carried 1,799 links and the output was 284 MB, against 5.5 MB for 179 pages — 52x
the bytes for 10x the input. A nav is a map of a site's top level, not an index of its
contents; section pages reach their siblings through that section's landing page.

On a 1,790-page corpus with 6 top-level pages: full build 0.82s -> 0.39s, output 284 MB
-> 34 MB. Scaling is now linear (179 pages in 0.07s, 1,796 in 0.39s, where the small
case is mostly the fixed cost of loading syntect's syntax definitions).

The larger win is incremental. The site-structure hash — the thing that forces a global
re-render — now covers only the pages that appear in the nav, since those are the only
ones whose title or URL can affect another page. Adding a blog post used to re-render
the entire site; it now renders one page. A top-level page's title still invalidates
everything, correctly, because every page displays it.

Trade-off: on a site whose sections live in subdirectories, only genuinely root-level
pages appear. cleberg.net keeps its landing pages at content/salary/index.org and
friends, so its nav comes out as one entry where the live site shows four. Treating a
directory's index.org as top-level is a one-line change to is_top_level; not done here
because it was not what was asked for.

Layout: unified · split

README.md +32 −12
@@ -281,18 +281,38 @@ therefore returns only what was written, and the report is assembled sequentiall
281281`parallel_builds_are_deterministic_in_output_and_report_order` holds that line, and it was
282282verified by reintroducing the bug and watching it fail.
283283
284### The real scaling limit is not the CPU
285
286Going 10× on corpus size cost 17× in time before parallelism, which is superlinear — and
287parallelism moves that constant without fixing it. The cause is the nav bar: it lists **every**
288page, so an *n*-page site emits *n*² nav links. At 1,790 pages each page carries 1,799 links
289and the output is 284 MB, against 5.5 MB for the 179-page corpus — 52× the bytes for 10× the
290input. Even at the real corpus size this is already visible: 18 KB pages whose nav dwarfs the
291prose, where the live site's nav has about six links.
292
293This is a template and configuration question rather than a bug — *which* pages belong in a
294nav is a decision this project has not made yet — so it is recorded here rather than guessed
295at. Until it is made, a build's cost is dominated by chrome nobody asked for.
284### The real scaling limit was not the CPU
285
286Going 10× on corpus size cost 17× in time, which parallelism improves without fixing: the
287cause was the nav bar listing **every** page, so an *n*-page site emitted *n*² nav links. At
2881,790 pages each page carried 1,799 links and the output was 284 MB, against 5.5 MB for the
289179-page corpus — 52× the bytes for 10× the input.
290
291The nav is now built from **top-level pages only** ([`is_top_level`](src/site.rs)): a nav is a
292map of the site's top level, not an index of its contents, and section pages reach their
293siblings through that section's landing page. Nav size becomes a function of the top level
294rather than of the corpus, and the quadratic disappears.
295
296| 1,790-page corpus (6 top-level pages) | before | after |
297|---|---|---|
298| full build | 0.82s | 0.39s |
299| total output | 284 MB | 34 MB |
300| nav links per page | 1,799 | 6 |
301
302Scaling is now linear: 179 pages in 0.07s and 1,796 in 0.39s, where the small case is mostly
303the fixed cost of loading syntect's syntax definitions.
304
305The same rule sharpened the incremental build, which is the larger win. The site-structure
306hash — the thing that forces a global re-render — now covers only the pages that appear in
307the nav, because those are the only ones whose title or URL affects another page. **Adding a
308blog post used to re-render the entire site; now it renders one page.** A top-level page's
309title still invalidates everything, correctly, since every page displays it.
310
311**Trade-off worth knowing:** on a site whose sections live in subdirectories, only genuinely
312root-level pages appear. cleberg.net keeps its landing pages at `content/salary/index.org`
313and friends, so its nav comes out as a single `index.org` entry where the live site shows
314four. Treating a directory's `index.org` as top-level too is a one-line change to
315`is_top_level` if that is the behaviour you want.
296316
297317**From v0.1 (core subset):** headings with nesting and anchors (every heading is now
298318anchored — `:CUSTOM_ID:`/`:ID:` else a slug of its text) and trailing tags; paragraphs;
src/site.rs +28 −6
@@ -127,18 +127,23 @@ fn prepare_pages(src: &Utf8Path) -> Result<(Vec<PagePrep>, SymbolTable)> {
127127 symbols.index_document(doc);
128128 }
129129
130 // Nav is global; titles come from #+TITLE (falling back to the file stem) and URLs
131 // from each page's output path, which `#+SLUG:` can rename.
132 let entries: Vec<(Utf8PathBuf, String)> = docs
130 // Nav is global chrome; titles come from #+TITLE (falling back to the file stem) and
131 // URLs from each page's output path, which `#+SLUG:` can rename.
132 let all_pages: Vec<(Utf8PathBuf, String)> = docs
133133 .iter()
134134 .map(|d| (output_path(&d.source_path, &d.keywords), page_title(d)))
135135 .collect();
136 let entries: Vec<(Utf8PathBuf, String)> = all_pages
137 .iter()
138 .filter(|(out, _)| is_top_level(out))
139 .cloned()
140 .collect();
136141
137142 // Two sources emitting one page would silently drop a page — and with slugs, a
138143 // collision is a typo away and invisible in the source filenames.
139144 let mut claimed: std::collections::HashMap<&Utf8PathBuf, &Utf8PathBuf> =
140145 std::collections::HashMap::new();
141 for (doc, (out, _)) in docs.iter().zip(&entries) {
146 for (doc, (out, _)) in docs.iter().zip(&all_pages) {
142147 if let Some(other) = claimed.insert(out, &doc.source_path) {
143148 anyhow::bail!(
144149 "output collision: {} and {} both build to {out} (check their #+SLUG:)",
@@ -239,10 +244,16 @@ pub fn build_site(src: &Utf8Path, out: &Utf8Path, opts: &BuildOptions) -> Result
239244 // chrome on every page — is built from every page's (path, title), so a title/path
240245 // change or a page add/remove must re-render every page (else stale nav on disk).
241246 let cfg = BuildConfig::default();
242 // Keyed on the *output* path: a `#+SLUG:` change moves a page's URL, which changes
243 // the nav on every other page even though no source filename moved.
247 // Only the pages that actually appear in the nav belong in the site-structure hash,
248 // because the nav is the only global chrome a page carries. Hashing *every* page
249 // here would mean adding one blog post re-rendered the entire site — correct, but
250 // needlessly: a nested page cannot change any other page's nav.
251 //
252 // Keyed on the *output* path, since a `#+SLUG:` change moves a page's URL — and so
253 // its nav link — even though no source filename moved.
244254 let nav_entries: Vec<(String, String)> = preps
245255 .iter()
256 .filter(|p| is_top_level(&p.output))
246257 .map(|p| (p.output.to_string(), p.title.clone()))
247258 .collect();
248259 let cfg_hash = combine(config_hash(&cfg), site_structure_hash(&nav_entries));
@@ -489,6 +500,17 @@ fn discover(src: &Utf8Path) -> Result<(Vec<Utf8PathBuf>, Vec<Utf8PathBuf>)> {
489500 Ok((org, assets))
490501}
491502
503/// Does this output path sit at the site root?
504///
505/// The nav is the site's global chrome, and listing *every* page in it makes an `n`-page
506/// site emit `n²` nav links — 1,790 pages produced 284 MB of output, most of it nav. A
507/// nav is a map of the site's top level, not an index of its contents, so it is built
508/// from root-level pages only. Section pages reach their siblings through that section's
509/// own landing page.
510fn is_top_level(output: &Utf8Path) -> bool {
511 output.parent().is_none_or(|p| p.as_str().is_empty())
512}
513
492514fn page_title(doc: &Document) -> String {
493515 doc.keywords
494516 .entries
tests/incremental.rs +87
@@ -361,3 +361,90 @@ fn parallel_builds_are_deterministic_in_output_and_report_order() {
361361 );
362362 }
363363}
364
365/// A site with pages in subdirectories.
366fn write_nested_site(src: &Utf8PathBuf) {
367 std::fs::create_dir_all(src.join("blog")).unwrap();
368 write(src, "index.org", "#+TITLE: Home\n\nWelcome.\n");
369 write(src, "about.org", "#+TITLE: About\n\nAbout me.\n");
370 write(&src.join("blog"), "first.org", "#+TITLE: First Post\n\nPost body.\n");
371 write(&src.join("blog"), "second.org", "#+TITLE: Second Post\n\nPost body.\n");
372}
373
374/// The nav is a map of the site's top level, not an index of its contents. Listing every
375/// page made an n-page site emit n² nav links: 1,790 pages produced 284 MB of output,
376/// nearly all of it nav.
377#[test]
378fn nav_lists_only_top_level_pages() {
379 let root = tmpdir("navtop");
380 let src = root.join("src");
381 std::fs::create_dir_all(&src).unwrap();
382 write_nested_site(&src);
383 let out_dir = root.join("out");
384
385 build_site(&src, &out_dir, &BuildOptions::default()).unwrap();
386 let home = std::fs::read_to_string(out_dir.join("index.html")).unwrap();
387 let nav = home
388 .split("<nav>")
389 .nth(1)
390 .and_then(|s| s.split("</nav>").next())
391 .expect("a nav element");
392
393 assert!(nav.contains("About"), "a root-level page belongs in the nav:\n{nav}");
394 assert!(nav.contains("Home"), "the index page belongs in the nav:\n{nav}");
395 assert!(
396 !nav.contains("First Post") && !nav.contains("Second Post"),
397 "pages in subdirectories must not appear in the nav:\n{nav}"
398 );
399
400 // Nested pages still get the nav — they just are not *in* it.
401 let post = std::fs::read_to_string(out_dir.join("blog/first.html")).unwrap();
402 assert!(
403 post.contains("href=\"../about.html\"") && post.contains("href=\"../index.html\""),
404 "a nested page links up to the top-level nav:\n{post}"
405 );
406}
407
408/// The payoff for narrowing the site-structure hash to nav entries. Adding a blog post
409/// cannot change any other page's nav, so it must not re-render the site — which is what
410/// hashing *every* page's (path, title) used to force.
411#[test]
412fn adding_a_nested_page_does_not_rebuild_the_site() {
413 let root = tmpdir("navadd");
414 let src = root.join("src");
415 std::fs::create_dir_all(&src).unwrap();
416 write_nested_site(&src);
417 let out_dir = root.join("out");
418
419 build_site(&src, &out_dir, &BuildOptions::default()).unwrap();
420
421 write(&src.join("blog"), "third.org", "#+TITLE: Third Post\n\nBody.\n");
422 let r = build_site(&src, &out_dir, &BuildOptions::default()).unwrap();
423
424 assert_eq!(
425 r.rendered,
426 vec![out("blog/third.html")],
427 "only the new nested page renders, got: {:?}",
428 r.rendered
429 );
430 assert_eq!(r.skipped.len(), 4, "every pre-existing page is reused");
431}
432
433/// The other half of the same rule: a page that IS in the nav still invalidates
434/// everything when its title changes, because every page renders that title.
435#[test]
436fn retitling_a_top_level_page_still_rebuilds_the_site() {
437 let root = tmpdir("navretitle");
438 let src = root.join("src");
439 std::fs::create_dir_all(&src).unwrap();
440 write_nested_site(&src);
441 let out_dir = root.join("out");
442
443 build_site(&src, &out_dir, &BuildOptions::default()).unwrap();
444 write(&src, "about.org", "#+TITLE: Colophon\n\nAbout me.\n");
445 let r = build_site(&src, &out_dir, &BuildOptions::default()).unwrap();
446
447 assert_eq!(r.rendered.len(), 4, "a nav title change re-renders every page");
448 let post = std::fs::read_to_string(out_dir.join("blog/first.html")).unwrap();
449 assert!(post.contains("Colophon"), "nested pages show the updated nav title");
450}