Commit a52aee8095
Unsigned
Layout: unified · split
.github/workflows/release.yml +1 −1
| @@ -77,7 +77,7 @@ jobs: | |||
| 77 | staging="orgo-${{ github.event.inputs.tag || github.ref_name }}-${{ matrix.target }}" | 77 | staging="orgo-${{ github.event.inputs.tag || github.ref_name }}-${{ matrix.target }}" |
| 78 | mkdir "$staging" | 78 | mkdir "$staging" |
| 79 | cp "target/${{ matrix.target }}/release/orgo" "$staging/" | 79 | cp "target/${{ matrix.target }}/release/orgo" "$staging/" |
| 80 | cp README.md LICENSE CHANGELOG.md "$staging/" | 80 | cp README.org LICENSE CHANGELOG.org "$staging/" |
| 81 | tar czf "$staging.tar.gz" "$staging" | 81 | tar czf "$staging.tar.gz" "$staging" |
| 82 | shasum -a 256 "$staging.tar.gz" > "$staging.tar.gz.sha256" | 82 | shasum -a 256 "$staging.tar.gz" > "$staging.tar.gz.sha256" |
| 83 | 83 | ||
CHANGELOG.md deleted −177
| @@ -1,177 +0,0 @@ | |||
| 1 | # Changelog | ||
| 2 | |||
| 3 | What changed and why, newest first. Entries name the *behaviour* that moved, since that is | ||
| 4 | what a rebuild will show you. | ||
| 5 | |||
| 6 | Two conventions worth knowing before reading: | ||
| 7 | |||
| 8 | - **A cache-format bump is not a change you need to act on.** The incremental cache is | ||
| 9 | versioned and discards itself; a bump means the next build re-renders everything once. | ||
| 10 | - **Output changes are called out.** orgo aims at what Emacs exports from the same | ||
| 11 | file, so an entry that says "now renders X" means your pages will change. That is the | ||
| 12 | product, not a regression — but it belongs in a changelog rather than a diff you find | ||
| 13 | later. | ||
| 14 | |||
| 15 | Versions follow the compatibility promise in the README: config keys, template variables, | ||
| 16 | CLI flags and URLs are the stable surface. | ||
| 17 | |||
| 18 | ## 0.19.1 | ||
| 19 | |||
| 20 | - Footnote back-links carry `aria-label="Back to reference N"`, and the notes section is | ||
| 21 | labelled. A link whose only visible content is `↩` has that glyph as its whole | ||
| 22 | accessible name, so a screen reader announced "left arrow with hook" once per note with | ||
| 23 | no way to tell them apart. | ||
| 24 | |||
| 25 | ## 0.19.0 | ||
| 26 | |||
| 27 | - **Full-content collections.** `include_content = true` gives a listing template each | ||
| 28 | entry's rendered HTML as `entry.content` — a feed that carries whole posts rather than | ||
| 29 | excerpts. Rendered only when the listing is actually rebuilt, so a cached feed costs | ||
| 30 | nothing. | ||
| 31 | - **Fixed: a listing could show a stale excerpt.** Its cache key covered a hand-picked set | ||
| 32 | of fields, and the excerpt was not among them, so rewriting a post's first paragraph | ||
| 33 | left the old text on the index until something unrelated invalidated it. Entries are now | ||
| 34 | hashed through their serialization, which cannot drift from what a template can read. | ||
| 35 | Editing a post's body now rebuilds the listings that show it. | ||
| 36 | - `page.toc` entries carry `number`, so a site with section numbering on can number its | ||
| 37 | contents list to match its headings. | ||
| 38 | |||
| 39 | ## 0.18.0 | ||
| 40 | |||
| 41 | Release engineering, so that a version number is worth reading. | ||
| 42 | |||
| 43 | - **A written compatibility promise.** Config keys, template variables, CLI flags and URLs | ||
| 44 | are the stable surface; the incremental cache, HTML details and the Rust API are not. | ||
| 45 | In the README, and in the guide under *Versioning and upgrades*. | ||
| 46 | - **CI** on Linux and macOS: build, test, clippy as an error, and the documentation site | ||
| 47 | built with `--strict`. Emacs is installed on both, so the oracle suite runs for real | ||
| 48 | instead of skipping. | ||
| 49 | - **A checked MSRV**, 1.88 — which is how it came to be 1.88 rather than the 1.82 | ||
| 50 | orgo's own code needs. The floor comes from dependencies, and nobody finds that out | ||
| 51 | by reasoning about it. | ||
| 52 | - **Release binaries** for macOS (arm64, x86_64) and Linux (gnu, musl), built on tag into | ||
| 53 | a draft release. The tag is checked against `Cargo.toml` before anything is built. | ||
| 54 | - A `LICENSE` file to go with the MIT declaration, crates.io metadata, and a release | ||
| 55 | profile that produces a 5.0 MB binary rather than 6.5 MB. | ||
| 56 | - This changelog, and `RELEASING.md`. | ||
| 57 | |||
| 58 | ## 0.17.0 | ||
| 59 | |||
| 60 | - **Asset directories outside the source.** `[build] assets = ["../theme/static"]` copies | ||
| 61 | a directory's contents to the site root. A site's static files do not always live where | ||
| 62 | its writing does, and copying them next to the writing is how a repository ends up with | ||
| 63 | two of every stylesheet. `watch` and `serve` watch these directories too. Two files | ||
| 64 | claiming one URL is a build error naming both. | ||
| 65 | - **Template hashing is per template.** A page's render key covered every template, so | ||
| 66 | editing a feed template re-rendered the whole site. It now covers the layout the page | ||
| 67 | uses plus what that layout extends, includes or imports. On a 196-page site, editing the | ||
| 68 | feed template renders one page instead of 196. | ||
| 69 | - Cache format 7. | ||
| 70 | |||
| 71 | ## 0.16.0 | ||
| 72 | |||
| 73 | - **Org's entity table.** `\alpha`, `\rarr`, `20\deg` and the other 412 names, generated | ||
| 74 | from Emacs' own `org-entities`. An unknown name stays literal; `#+OPTIONS: e:nil` turns | ||
| 75 | the table off. *Output changes* for any page using entities. | ||
| 76 | - **Table captions.** `#+CAPTION:` above a table becomes a numbered `<caption>`. | ||
| 77 | - **`#+INCLUDE:` reports itself.** It was inert and silent, which publishes a page with | ||
| 78 | content missing and nobody told. Now a diagnostic, and `--strict` makes it a failure. | ||
| 79 | - The Emacs oracle separates deliberate divergence from defects. Every difference from | ||
| 80 | org's exporter is named and justified, and a test asserts there are no others. | ||
| 81 | |||
| 82 | ## 0.15.0 | ||
| 83 | |||
| 84 | Export parity, from a page-by-page diff of a 179-file corpus against the site Emacs | ||
| 85 | publishes from the same sources. **All of these change output.** | ||
| 86 | |||
| 87 | - Heading levels are relative to a document's shallowest heading, as org exports them. | ||
| 88 | - Org's text conversions: `--`, `---`, `...`, and `x^2` / `a_{b}`. Never inside verbatim, | ||
| 89 | code, source blocks or LaTeX. `#+OPTIONS: -:nil`, `^:nil` and `^:{}` all work. | ||
| 90 | - Captioned figures are numbered `Figure N:`. | ||
| 91 | - A caption attaches to the element *directly* below it; a blank line between attaches to | ||
| 92 | nothing. | ||
| 93 | - Checkboxes render as org writes them, which keeps the `[-]` partly-done state a disabled | ||
| 94 | `<input>` could not express. `[@4]` sets a list item's number. | ||
| 95 | - A table's special marker column and its marker rows stay out of the output. | ||
| 96 | - `#+BEGIN_NOTE` and any other unrecognised name is a special block: a div holding parsed | ||
| 97 | org rather than a `<pre>` of literal text. Verse keeps its line breaks. | ||
| 98 | - Emphasis borders forbid whitespace and nothing else, so `="proxied":false=` is verbatim | ||
| 99 | and `~~/.config/doom/config.el~` is a path that starts with a tilde. | ||
| 100 | - Listings sort on the time of day when a timestamp carries one. | ||
| 101 | - Cache format 6. | ||
| 102 | |||
| 103 | ## 0.14.0 | ||
| 104 | |||
| 105 | - **Per-page layouts.** `[[pages]]` rules map a source path to a template, and | ||
| 106 | `#+TEMPLATE:` on a page overrides any rule. A missing template fails the build naming | ||
| 107 | the page, the template, and what does exist. | ||
| 108 | - `page.year`, for grouping a listing by year with minijinja's `groupby`. | ||
| 109 | - An explicit nav can order generated pages among authored ones. `nav.mode = "none"` now | ||
| 110 | really means none. | ||
| 111 | |||
| 112 | ## 0.13.0 | ||
| 113 | |||
| 114 | - Bundled TOML and Org syntax definitions, a `syntaxes_dir` for your own, and org's comma | ||
| 115 | escape (`,* heading` inside a block). | ||
| 116 | |||
| 117 | ## 0.12.0 | ||
| 118 | |||
| 119 | - `serve`: a development server with live reload, bound to loopback. | ||
| 120 | - A documentation site under `docs/`, built by orgo itself. | ||
| 121 | |||
| 122 | ## 0.11.0 | ||
| 123 | |||
| 124 | - Table of contents as `page.toc`, section numbers, and org's `#+OPTIONS:` per-file | ||
| 125 | switches. | ||
| 126 | |||
| 127 | ## 0.10.0 | ||
| 128 | |||
| 129 | - Excerpts, word count, reading time, a `truncate` filter, and `#+DRAFT:` pages. | ||
| 130 | |||
| 131 | ## 0.9.0 | ||
| 132 | |||
| 133 | - `watch`: rebuilds on OS filesystem events, debounced. | ||
| 134 | |||
| 135 | ## 0.8.0 | ||
| 136 | |||
| 137 | - `site.base_url`, the `absolute` and `rfc822` filters, canonical links, and an RSS feed | ||
| 138 | in the scaffold that validates. | ||
| 139 | |||
| 140 | ## 0.7.0 | ||
| 141 | |||
| 142 | - Pagination for large listings, with a `paginator` template context that composes with | ||
| 143 | grouping. | ||
| 144 | |||
| 145 | ## 0.6.0 | ||
| 146 | |||
| 147 | - Grouped collections: one page per tag plus a tag index. | ||
| 148 | - Generated listing pages (`[[collections]]`), sorted indexes, and feeds via XML | ||
| 149 | templates. | ||
| 150 | - A config file, user templates, nav modes, an `init` scaffold, and discovery that will | ||
| 151 | not publish `.git`. | ||
| 152 | |||
| 153 | ## 0.5.0 | ||
| 154 | |||
| 155 | - Parse diagnostics carry `file:line`, and pages render in parallel. | ||
| 156 | - The corpus audit (`orgo audit`) and the `emacs --batch` oracle. | ||
| 157 | - `#+SLUG:` decides a page's output filename — found by auditing a real corpus, where it | ||
| 158 | affected 169 of 182 URLs. | ||
| 159 | |||
| 160 | ## 0.4.0 | ||
| 161 | |||
| 162 | - The full v1 construct scope, with the IN/OUT line under test. | ||
| 163 | |||
| 164 | ## 0.3.0 | ||
| 165 | |||
| 166 | - The incremental build layer: content, config and template hashing, a dependency graph, | ||
| 167 | per-page render keys, and a persisted cache manifest. A full build and an incremental | ||
| 168 | build produce byte-identical output. | ||
| 169 | |||
| 170 | ## 0.2.0 | ||
| 171 | |||
| 172 | - Multi-file site builds: a symbol table, internal link resolution, minijinja templates, | ||
| 173 | tables and footnotes. | ||
| 174 | |||
| 175 | ## 0.1.0 | ||
| 176 | |||
| 177 | - Parse and render a single `.org` file to HTML. | ||
CHANGELOG.org added +156
| @@ -0,0 +1,156 @@ | |||
| 1 | * Changelog | ||
| 2 | What changed and why, newest first. Entries name the /behaviour/ that moved, since that is | ||
| 3 | what a rebuild will show you. | ||
| 4 | |||
| 5 | Two conventions worth knowing before reading: | ||
| 6 | |||
| 7 | - *A cache-format bump is not a change you need to act on.* The incremental cache is | ||
| 8 | versioned and discards itself; a bump means the next build re-renders everything once. | ||
| 9 | - *Output changes are called out.* orgo aims at what Emacs exports from the same | ||
| 10 | file, so an entry that says "now renders X" means your pages will change. That is the | ||
| 11 | product, not a regression — but it belongs in a changelog rather than a diff you find | ||
| 12 | later. | ||
| 13 | |||
| 14 | Versions follow the compatibility promise in the README: config keys, template variables, | ||
| 15 | CLI flags and URLs are the stable surface. | ||
| 16 | |||
| 17 | ** 0.19.1 | ||
| 18 | - Footnote back-links carry =aria-label="Back to reference N"=, and the notes section is | ||
| 19 | labelled. A link whose only visible content is =↩= has that glyph as its whole | ||
| 20 | accessible name, so a screen reader announced "left arrow with hook" once per note with | ||
| 21 | no way to tell them apart. | ||
| 22 | |||
| 23 | ** 0.19.0 | ||
| 24 | - *Full-content collections.* =include_content = true= gives a listing template each | ||
| 25 | entry's rendered HTML as =entry.content= — a feed that carries whole posts rather than | ||
| 26 | excerpts. Rendered only when the listing is actually rebuilt, so a cached feed costs | ||
| 27 | nothing. | ||
| 28 | - *Fixed: a listing could show a stale excerpt.* Its cache key covered a hand-picked set | ||
| 29 | of fields, and the excerpt was not among them, so rewriting a post's first paragraph | ||
| 30 | left the old text on the index until something unrelated invalidated it. Entries are now | ||
| 31 | hashed through their serialization, which cannot drift from what a template can read. | ||
| 32 | Editing a post's body now rebuilds the listings that show it. | ||
| 33 | - =page.toc= entries carry =number=, so a site with section numbering on can number its | ||
| 34 | contents list to match its headings. | ||
| 35 | |||
| 36 | ** 0.18.0 | ||
| 37 | Release engineering, so that a version number is worth reading. | ||
| 38 | |||
| 39 | - *A written compatibility promise.* Config keys, template variables, CLI flags and URLs | ||
| 40 | are the stable surface; the incremental cache, HTML details and the Rust API are not. | ||
| 41 | In the README, and in the guide under /Versioning and upgrades/. | ||
| 42 | - *CI* on Linux and macOS: build, test, clippy as an error, and the documentation site | ||
| 43 | built with =--strict=. Emacs is installed on both, so the oracle suite runs for real | ||
| 44 | instead of skipping. | ||
| 45 | - *A checked MSRV*, 1.88 — which is how it came to be 1.88 rather than the 1.82 | ||
| 46 | orgo's own code needs. The floor comes from dependencies, and nobody finds that out | ||
| 47 | by reasoning about it. | ||
| 48 | - *Release binaries* for macOS (arm64, x86_64) and Linux (gnu, musl), built on tag into | ||
| 49 | a draft release. The tag is checked against =Cargo.toml= before anything is built. | ||
| 50 | - A =LICENSE= file to go with the MIT declaration, crates.io metadata, and a release | ||
| 51 | profile that produces a 5.0 MB binary rather than 6.5 MB. | ||
| 52 | - This changelog, and =RELEASING.org=. | ||
| 53 | |||
| 54 | ** 0.17.0 | ||
| 55 | - *Asset directories outside the source.* =[build] assets = ["../theme/static"]= copies | ||
| 56 | a directory's contents to the site root. A site's static files do not always live where | ||
| 57 | its writing does, and copying them next to the writing is how a repository ends up with | ||
| 58 | two of every stylesheet. =watch= and =serve= watch these directories too. Two files | ||
| 59 | claiming one URL is a build error naming both. | ||
| 60 | - *Template hashing is per template.* A page's render key covered every template, so | ||
| 61 | editing a feed template re-rendered the whole site. It now covers the layout the page | ||
| 62 | uses plus what that layout extends, includes or imports. On a 196-page site, editing the | ||
| 63 | feed template renders one page instead of 196. | ||
| 64 | - Cache format 7. | ||
| 65 | |||
| 66 | ** 0.16.0 | ||
| 67 | - *Org's entity table.* =\alpha=, =\rarr=, =20\deg= and the other 412 names, generated | ||
| 68 | from Emacs' own =org-entities=. An unknown name stays literal; =#+OPTIONS: e:nil= turns | ||
| 69 | the table off. /Output changes/ for any page using entities. | ||
| 70 | - *Table captions.* =#+CAPTION:= above a table becomes a numbered =<caption>=. | ||
| 71 | - *=#+INCLUDE:= reports itself.* It was inert and silent, which publishes a page with | ||
| 72 | content missing and nobody told. Now a diagnostic, and =--strict= makes it a failure. | ||
| 73 | - The Emacs oracle separates deliberate divergence from defects. Every difference from | ||
| 74 | org's exporter is named and justified, and a test asserts there are no others. | ||
| 75 | |||
| 76 | ** 0.15.0 | ||
| 77 | Export parity, from a page-by-page diff of a 179-file corpus against the site Emacs | ||
| 78 | publishes from the same sources. *All of these change output.* | ||
| 79 | |||
| 80 | - Heading levels are relative to a document's shallowest heading, as org exports them. | ||
| 81 | - Org's text conversions: =--=, =---=, =...=, and =x^2= / =a_{b}=. Never inside verbatim, | ||
| 82 | code, source blocks or LaTeX. =#+OPTIONS: -:nil=, =^:nil= and =^:{}= all work. | ||
| 83 | - Captioned figures are numbered =Figure N:=. | ||
| 84 | - A caption attaches to the element /directly/ below it; a blank line between attaches to | ||
| 85 | nothing. | ||
| 86 | - Checkboxes render as org writes them, which keeps the =[-]= partly-done state a disabled | ||
| 87 | =<input>= could not express. =[@4]= sets a list item's number. | ||
| 88 | - A table's special marker column and its marker rows stay out of the output. | ||
| 89 | - =#+BEGIN_NOTE= and any other unrecognised name is a special block: a div holding parsed | ||
| 90 | org rather than a =<pre>= of literal text. Verse keeps its line breaks. | ||
| 91 | - Emphasis borders forbid whitespace and nothing else, so =="proxied":false== is verbatim | ||
| 92 | and =~~/.config/doom/config.el~= is a path that starts with a tilde. | ||
| 93 | - Listings sort on the time of day when a timestamp carries one. | ||
| 94 | - Cache format 6. | ||
| 95 | |||
| 96 | ** 0.14.0 | ||
| 97 | - *Per-page layouts.* =[[pages]]= rules map a source path to a template, and | ||
| 98 | =#+TEMPLATE:= on a page overrides any rule. A missing template fails the build naming | ||
| 99 | the page, the template, and what does exist. | ||
| 100 | - =page.year=, for grouping a listing by year with minijinja's =groupby=. | ||
| 101 | - An explicit nav can order generated pages among authored ones. =nav.mode = "none"= now | ||
| 102 | really means none. | ||
| 103 | |||
| 104 | ** 0.13.0 | ||
| 105 | - Bundled TOML and Org syntax definitions, a =syntaxes_dir= for your own, and org's comma | ||
| 106 | escape (=,* heading= inside a block). | ||
| 107 | |||
| 108 | ** 0.12.0 | ||
| 109 | - =serve=: a development server with live reload, bound to loopback. | ||
| 110 | - A documentation site under =docs/=, built by orgo itself. | ||
| 111 | |||
| 112 | ** 0.11.0 | ||
| 113 | - Table of contents as =page.toc=, section numbers, and org's =#+OPTIONS:= per-file | ||
| 114 | switches. | ||
| 115 | |||
| 116 | ** 0.10.0 | ||
| 117 | - Excerpts, word count, reading time, a =truncate= filter, and =#+DRAFT:= pages. | ||
| 118 | |||
| 119 | ** 0.9.0 | ||
| 120 | - =watch=: rebuilds on OS filesystem events, debounced. | ||
| 121 | |||
| 122 | ** 0.8.0 | ||
| 123 | - =site.base_url=, the =absolute= and =rfc822= filters, canonical links, and an RSS feed | ||
| 124 | in the scaffold that validates. | ||
| 125 | |||
| 126 | ** 0.7.0 | ||
| 127 | - Pagination for large listings, with a =paginator= template context that composes with | ||
| 128 | grouping. | ||
| 129 | |||
| 130 | ** 0.6.0 | ||
| 131 | - Grouped collections: one page per tag plus a tag index. | ||
| 132 | - Generated listing pages (=[[collections]]=), sorted indexes, and feeds via XML | ||
| 133 | templates. | ||
| 134 | - A config file, user templates, nav modes, an =init= scaffold, and discovery that will | ||
| 135 | not publish =.git=. | ||
| 136 | |||
| 137 | ** 0.5.0 | ||
| 138 | - Parse diagnostics carry =file:line=, and pages render in parallel. | ||
| 139 | - The corpus audit (=orgo audit=) and the =emacs --batch= oracle. | ||
| 140 | - =#+SLUG:= decides a page's output filename — found by auditing a real corpus, where it | ||
| 141 | affected 169 of 182 URLs. | ||
| 142 | |||
| 143 | ** 0.4.0 | ||
| 144 | - The full v1 construct scope, with the IN/OUT line under test. | ||
| 145 | |||
| 146 | ** 0.3.0 | ||
| 147 | - The incremental build layer: content, config and template hashing, a dependency graph, | ||
| 148 | per-page render keys, and a persisted cache manifest. A full build and an incremental | ||
| 149 | build produce byte-identical output. | ||
| 150 | |||
| 151 | ** 0.2.0 | ||
| 152 | - Multi-file site builds: a symbol table, internal link resolution, minijinja templates, | ||
| 153 | tables and footnotes. | ||
| 154 | |||
| 155 | ** 0.1.0 | ||
| 156 | - Parse and render a single =.org= file to HTML. | ||
Cargo.toml +1 −1
| @@ -4,7 +4,7 @@ version = "0.19.1" | |||
| 4 | edition = "2021" | 4 | edition = "2021" |
| 5 | description = "Org-mode static site generator that renders the org element tree straight to HTML" | 5 | description = "Org-mode static site generator that renders the org element tree straight to HTML" |
| 6 | license = "0BSD" | 6 | license = "0BSD" |
| 7 | readme = "README.md" | 7 | readme = "README.org" |
| 8 | keywords = ["org-mode", "static-site-generator", "emacs", "html", "blog"] | 8 | keywords = ["org-mode", "static-site-generator", "emacs", "html", "blog"] |
| 9 | categories = ["command-line-utilities", "text-processing"] | 9 | categories = ["command-line-utilities", "text-processing"] |
| 10 | repository = "https://github.com/ccleberg/orgo" | 10 | repository = "https://github.com/ccleberg/orgo" |
README.md deleted −755
| @@ -1,755 +0,0 @@ | |||
| 1 | # orgo | ||
| 2 | |||
| 3 | An org-mode static site generator, in Rust. Org is treated as the *source language*, | ||
| 4 | not an inconvenient input to be normalized into markdown. The org element tree — | ||
| 5 | headings, drawers, blocks, links with their org-specific semantics — **is** the | ||
| 6 | document model, and we render that tree straight to HTML. We never round-trip through | ||
| 7 | a markdown-shaped intermediate representation, because the point is to preserve what | ||
| 8 | markdown cannot express: property drawers, TODO/priority/tag metadata on headings, | ||
| 9 | `#+` directives, ID links, named/captioned blocks, footnote semantics. | ||
| 10 | |||
| 11 | The one non-obvious early commitment is **incremental builds keyed on content | ||
| 12 | hashing**, treated as a first-class architectural concern from day one. The discipline | ||
| 13 | it imposes on the data model — pure, hashable, dependency-tracked units — is the real | ||
| 14 | deliverable, even while the corpus is small enough that a full rebuild is instant. | ||
| 15 | |||
| 16 | **Full documentation is in [`docs/`](docs/)** — a site written in org and built by | ||
| 17 | orgo itself. Build and read it with: | ||
| 18 | |||
| 19 | ```bash | ||
| 20 | cargo run -- serve docs -o docs/_site | ||
| 21 | ``` | ||
| 22 | |||
| 23 | ## Quick start | ||
| 24 | |||
| 25 | ```bash | ||
| 26 | cargo run -- init my-site # config + an editable copy of the layout + a page | ||
| 27 | cargo run -- build my-site -o _site | ||
| 28 | ``` | ||
| 29 | |||
| 30 | Or skip the scaffolding entirely — point it at any directory of `.org` files: | ||
| 31 | |||
| 32 | ```bash | ||
| 33 | cargo run -- build ~/notes -o _site | ||
| 34 | ``` | ||
| 35 | |||
| 36 | **Zero configuration is a supported path, not a demo.** With no `orgo.toml`, no | ||
| 37 | templates and no orgo-specific markup in your files, you get a complete site: pages, | ||
| 38 | navigation, syntax-highlighted code and the stylesheet to colour it. Configuration | ||
| 39 | changes what you get; it is never what makes it work. | ||
| 40 | |||
| 41 | Discovery skips what should not be published — dot-directories such as `.git`, the config | ||
| 42 | file, the templates directory, and the output directory when it sits inside the source, so | ||
| 43 | `orgo build . -o _site` does the obvious thing. | ||
| 44 | |||
| 45 | ## Configuration | ||
| 46 | |||
| 47 | Everything is optional. `orgo init` writes a fully commented `orgo.toml`; every | ||
| 48 | value below is the default. | ||
| 49 | |||
| 50 | ```toml | ||
| 51 | [site] | ||
| 52 | title = "orgo site" | ||
| 53 | base_url = "" # absolute URL, no trailing slash; needed for feeds/canonical links | ||
| 54 | description = "" | ||
| 55 | language = "en" | ||
| 56 | |||
| 57 | [nav] | ||
| 58 | mode = "top-level" # top-level | all | explicit | none | ||
| 59 | # pages = ["index.org", "about.org"] # for mode = "explicit"; order is preserved | ||
| 60 | |||
| 61 | [templates] | ||
| 62 | dir = "templates" # base.html replaces the built-in layout | ||
| 63 | expose_page_list = false | ||
| 64 | |||
| 65 | # [[pages]] # which layout a section renders through; base.html by default | ||
| 66 | # match = "blog" # a source directory or one .org file; most specific rule wins | ||
| 67 | # template = "post.html" | ||
| 68 | |||
| 69 | [highlight] | ||
| 70 | theme = "InspiredGitHub" | ||
| 71 | |||
| 72 | [build] | ||
| 73 | drafts = false | ||
| 74 | assets = [] # extra directories copied to the site root, e.g. ["../theme/static"] | ||
| 75 | |||
| 76 | [html] | ||
| 77 | heading_offset = 1 # a level-1 org heading becomes <h2>, beneath the layout's <h1> | ||
| 78 | ``` | ||
| 79 | |||
| 80 | ### Templates | ||
| 81 | |||
| 82 | Drop a `base.html` into the templates directory and it replaces the built-in layout | ||
| 83 | entirely. Any other `.html` file there is available to `{% include %}` and | ||
| 84 | `{% extends %}`. Templates are [minijinja](https://docs.rs/minijinja) (Jinja2 syntax) and | ||
| 85 | receive: | ||
| 86 | |||
| 87 | | Variable | What it is | | ||
| 88 | |---|---| | ||
| 89 | | `body` | the rendered page HTML — use `{{ body \| safe }}` | | ||
| 90 | | `page` | `.title`, `.url`, `.source`, `.date`, `.date_iso`, `.year`, `.tags`, `.content`, `.excerpt`, `.word_count`, `.reading_time`, `.toc`, `.keywords` | | ||
| 91 | | `site` | `.title`, `.base_url`, `.description`, `.language` | | ||
| 92 | | `nav` | list of `{title, url}`, relative to this page | | ||
| 93 | | `root` | `../`-prefix back to the site root from this page | | ||
| 94 | | `stylesheet` | URL of the generated `syntax.css` | | ||
| 95 | | `pages` | every page's metadata — only when `expose_page_list = true` | | ||
| 96 | |||
| 97 | `page.keywords` carries **every** `#+KEYWORD:` in the file under its lowercased name, so | ||
| 98 | your own metadata works without this crate knowing about it: `#+CUSTOM_THING: x` is | ||
| 99 | `{{ page.keywords.custom_thing }}`. | ||
| 100 | |||
| 101 | `base.html` is the default layout, not the only one. A `[[pages]]` rule gives a section | ||
| 102 | its own — `match = "blog"`, `template = "post.html"` — and `#+TEMPLATE: wide.html` gives | ||
| 103 | one page its own, which wins over any rule. A second layout usually starts with | ||
| 104 | `{% extends "base.html" %}`. | ||
| 105 | |||
| 106 | Editing a template re-renders the pages that use it — template sources are a hash input, | ||
| 107 | so a design change never leaves a site half-updated. | ||
| 108 | |||
| 109 | ### Generated listing pages | ||
| 110 | |||
| 111 | A blog index, an archive, a feed — output files with no source `.org` behind them. | ||
| 112 | Repeat the block for each one: | ||
| 113 | |||
| 114 | ```toml | ||
| 115 | [[collections]] | ||
| 116 | source = "blog" # directory to list; empty means every page | ||
| 117 | output = "blog/index.html" # where to write it | ||
| 118 | template = "list.html" | ||
| 119 | title = "Blog" | ||
| 120 | sort = "date" # date | title | path | ||
| 121 | order = "desc" # desc | asc | ||
| 122 | nav = true # put this listing page in the nav | ||
| 123 | ``` | ||
| 124 | |||
| 125 | The template gets the collection's entries as `pages`, already sorted, plus the usual | ||
| 126 | `site`/`nav`/`root`. It can `{% extends "base.html" %}` to inherit the site chrome: | ||
| 127 | |||
| 128 | ```jinja | ||
| 129 | {% extends "base.html" %} | ||
| 130 | {% block main %} | ||
| 131 | <ul>{% for p in pages %} | ||
| 132 | <li><time datetime="{{ p.date_iso }}">{{ p.date_iso }}</time> | ||
| 133 | <a href="{{ root }}{{ p.url }}">{{ p.title }}</a></li> | ||
| 134 | {% endfor %}</ul> | ||
| 135 | {% endblock %} | ||
| 136 | ``` | ||
| 137 | |||
| 138 | `p.date_iso` is the `YYYY-MM-DD` extracted from `#+DATE:`, whatever org syntax it was | ||
| 139 | written in — `[2025-09-05 Fri 10:21:00]`, `<2024-05-01 Wed>` or bare `2024-05-01`. It is | ||
| 140 | also the sort key; pages without a parseable date sort last, so an undated draft never | ||
| 141 | leads a dated archive. | ||
| 142 | |||
| 143 | #### Pagination | ||
| 144 | |||
| 145 | Set `paginate` to split a long listing across numbered pages: | ||
| 146 | |||
| 147 | ```toml | ||
| 148 | [[collections]] | ||
| 149 | source = "blog" | ||
| 150 | output = "blog/index.html" | ||
| 151 | paginate = 10 | ||
| 152 | paginate_output = "blog/page/{n}.html" # {n} is the 1-based page number | ||
| 153 | ``` | ||
| 154 | |||
| 155 | Page 1 stays at `output`, so a section's canonical URL never moves as its page count | ||
| 156 | changes; only pages 2..N are named by `paginate_output`. The template gets a `paginator`: | ||
| 157 | |||
| 158 | ```jinja | ||
| 159 | {% if paginator and paginator.total > 1 %} | ||
| 160 | <nav> | ||
| 161 | {% if paginator.prev_url %}<a href="{{ paginator.prev_url }}">Newer</a>{% endif %} | ||
| 162 | {% for pg in paginator.pages %} | ||
| 163 | <a href="{{ pg.url }}"{% if pg.current %} aria-current="page"{% endif %}>{{ pg.number }}</a> | ||
| 164 | {% endfor %} | ||
| 165 | {% if paginator.next_url %}<a href="{{ paginator.next_url }}">Older</a>{% endif %} | ||
| 166 | </nav> | ||
| 167 | {% endif %} | ||
| 168 | ``` | ||
| 169 | |||
| 170 | `paginator` carries `current`, `total`, `per_page`, `total_entries`, `prev_url`, | ||
| 171 | `next_url`, `first_url`, `last_url`, and `pages`. Every URL is relative to the page | ||
| 172 | carrying it, so links work from page 1 (`page/2.html`) and from page 5 (`../index.html`, | ||
| 173 | `6.html`) without the template knowing where it sits. An unpaginated collection has no | ||
| 174 | `paginator` at all, so `{% if paginator %}` is a reliable test in a shared template. | ||
| 175 | |||
| 176 | Grouping and pagination compose: each group paginates independently, which is why | ||
| 177 | `paginate_output` needs `{tag}` as well as `{n}` on a grouped collection. An empty | ||
| 178 | collection still emits page 1 — a section that exists but has nothing in it should say so | ||
| 179 | rather than 404. When the entry count shrinks, pages that no longer exist are deleted | ||
| 180 | instead of being left serving stale posts. | ||
| 181 | |||
| 182 | #### Tag pages | ||
| 183 | |||
| 184 | Add `group_by` and the collection emits one page *per group* instead of one page total, | ||
| 185 | plus an optional index of the groups: | ||
| 186 | |||
| 187 | ```toml | ||
| 188 | [[collections]] | ||
| 189 | source = "blog" | ||
| 190 | group_by = "tags" # "tags", or any #+KEYWORD: name to group by its value | ||
| 191 | output = "tags/{tag}.html" # {tag} is replaced by each group's slug | ||
| 192 | template = "tag.html" | ||
| 193 | title = "Tagged: {tag}" | ||
| 194 | index_output = "tags/index.html" # the tag index | ||
| 195 | index_template = "tags.html" | ||
| 196 | index_title = "Tags" | ||
| 197 | nav = true # adds the *index*, not every tag | ||
| 198 | ``` | ||
| 199 | |||
| 200 | A group page receives its own posts as `pages` and itself as `group` | ||
| 201 | (`.name`, `.slug`, `.url`, `.count`). The index receives `groups` — every group, sorted | ||
| 202 | by name: | ||
| 203 | |||
| 204 | ```jinja | ||
| 205 | <ul>{% for tag in groups %} | ||
| 206 | <li><a href="{{ root }}{{ tag.url }}">{{ tag.name }}</a> ({{ tag.count }})</li> | ||
| 207 | {% endfor %}</ul> | ||
| 208 | ``` | ||
| 209 | |||
| 210 | `group_by = "tags"` is multi-valued: a post appears under every tag it carries. Any other | ||
| 211 | value names a single-valued `#+KEYWORD:`, so `group_by = "category"` buckets by | ||
| 212 | `#+CATEGORY:`. | ||
| 213 | |||
| 214 | Two tags that would produce the same URL (`web_dev` and `web@dev` both slugify to | ||
| 215 | `web-dev`) are a build error rather than one page silently overwriting the other. | ||
| 216 | |||
| 217 | A tag page depends on its own posts and nothing else, so adding a post tagged `rust` | ||
| 218 | re-renders that post, its section index, `tags/rust.html`, and the tag index whose counts | ||
| 219 | changed — four pages, not one per tag. That precision is why `groups` is given to the | ||
| 220 | index and not to every group page: a page that can see every group depends on every | ||
| 221 | group. | ||
| 222 | |||
| 223 | #### Feeds and absolute URLs | ||
| 224 | |||
| 225 | **A feed is a listing page with an XML template**, not a separate feature — templates are | ||
| 226 | loaded by full filename and any extension, so `output = "feed.xml"` with | ||
| 227 | `template = "feed.xml"` is all it takes. `orgo init` writes a working RSS template. | ||
| 228 | |||
| 229 | A feed is read away from the site that served it, so relative links in one are simply | ||
| 230 | broken. Set `site.base_url` and use the `absolute` filter: | ||
| 231 | |||
| 232 | ```jinja | ||
| 233 | <link>{{ post.url | absolute }}</link> | ||
| 234 | <pubDate>{{ post.date_iso | rfc822 }}</pubDate> | ||
| 235 | ``` | ||
| 236 | |||
| 237 | | Filter | Does | | ||
| 238 | |---|---| | ||
| 239 | | `absolute` | site-root-relative path → absolute URL; already-absolute URLs pass through | | ||
| 240 | | `rfc822` | any org or ISO date → the format RSS `pubDate` requires | | ||
| 241 | | `truncate(n)` | shorten to at most `n` characters on a word boundary, with an ellipsis | | ||
| 242 | |||
| 243 | Apply `absolute` to the site-root-relative values — `page.url`, `pages[].url`, | ||
| 244 | `group.url` — and not to `nav[].url`, `paginator.*_url`, `stylesheet` or `root`, which | ||
| 245 | are relative to the page carrying them and already correct there. | ||
| 246 | |||
| 247 | With no `base_url`, `absolute` is an **error** naming the setting, rather than quietly | ||
| 248 | emitting a relative URL that would make the feed invalid everywhere while looking fine. | ||
| 249 | The default layout also emits `<link rel="canonical">` when a base URL is set. | ||
| 250 | |||
| 251 | Listing pages are cached on the entries they list, so adding a post re-renders that | ||
| 252 | section's index and nothing else. | ||
| 253 | |||
| 254 | ### Table of contents and `#+OPTIONS:` | ||
| 255 | |||
| 256 | `page.toc` is the page's headings as a **tree** — `{title, anchor, level, children}` — | ||
| 257 | because a table of contents is one, and rebuilding a tree from a flat list of levels | ||
| 258 | inside a template is what Jinja is worst at. Its anchors come from the same function the | ||
| 259 | renderer uses to emit heading `id`s, so a TOC link cannot drift from the heading it | ||
| 260 | points at. | ||
| 261 | |||
| 262 | ```jinja | ||
| 263 | {% macro toc_list(entries) %} | ||
| 264 | <ul>{% for e in entries %} | ||
| 265 | <li><a href="#{{ e.anchor }}">{{ e.title }}</a> | ||
| 266 | {%- if e.children %}{{ toc_list(e.children) }}{% endif %}</li> | ||
| 267 | {% endfor %}</ul> | ||
| 268 | {% endmacro %} | ||
| 269 | {% if page.toc %}{{ toc_list(page.toc) }}{% endif %} | ||
| 270 | ``` | ||
| 271 | |||
| 272 | Org's own per-file export switches are honoured, so a document can turn a feature off for | ||
| 273 | itself the way its author already knows: | ||
| 274 | |||
| 275 | | Switch | Effect | Site default | | ||
| 276 | |---|---|---| | ||
| 277 | | `#+OPTIONS: toc:nil` | empties `page.toc` for this page | `[html] toc = true` | | ||
| 278 | | `#+OPTIONS: num:t` | numbers headings `1.`, `1.1.`, … | `[html] section_numbers = false` | | ||
| 279 | |||
| 280 | **Section numbers default to off, which differs from Emacs on purpose.** | ||
| 281 | `org-export-with-section-numbers` is on there, so an org-published site inherits numbered | ||
| 282 | headings whether or not anyone chose them. Most sites do not want them; `num:t` or | ||
| 283 | `section_numbers = true` gets Emacs' behaviour back, with Emacs' own | ||
| 284 | `section-number-N` classes so the output stays diffable against the oracle. | ||
| 285 | |||
| 286 | ### Excerpts and drafts | ||
| 287 | |||
| 288 | `page.excerpt` is a page's `#+DESCRIPTION:` when it sets one and its first paragraph | ||
| 289 | otherwise, so a listing has something to show whether or not the author thought about | ||
| 290 | summaries. `page.word_count` and `page.reading_time` (minutes at 200 wpm) count prose | ||
| 291 | only — a post that is mostly a shell transcript should not read as an hour's work. | ||
| 292 | `truncate` exists because an excerpt is usually a whole paragraph and minijinja has no | ||
| 293 | such filter. | ||
| 294 | |||
| 295 | `#+DRAFT:` keeps a page out of the build entirely — no page, and absent from listings and | ||
| 296 | the nav rather than merely unlinked. `--drafts` includes them, which is what you want | ||
| 297 | under `watch` while writing one. A draft is out of the symbol table too, so a link *to* | ||
| 298 | one is reported as the dead link it would be once published. | ||
| 299 | |||
| 300 | The keyword is read forgivingly: `t`, `yes`, `1` and a bare `#+DRAFT:` all mean draft, | ||
| 301 | because writing the keyword at all is the signal. Only an explicit `nil`, `false`, `no`, | ||
| 302 | `0` or `off` means published. | ||
| 303 | |||
| 304 | ### `#+SLUG:` | ||
| 305 | |||
| 306 | A page's output filename comes from its `#+SLUG:` when it has one, so | ||
| 307 | `2018-11-28-aes-encryption.org` can publish as `aes-encryption.html`. Without one the | ||
| 308 | source filename is used. Slugs are sanitized to a single safe path component, and two | ||
| 309 | pages claiming one URL is a build error rather than a silently dropped page. | ||
| 310 | |||
| 311 | ## Pipeline | ||
| 312 | |||
| 313 | ``` | ||
| 314 | DISCOVER → PARSE → INDEX → RESOLVE → RENDER → TEMPLATE → EMIT | ||
| 315 | ``` | ||
| 316 | |||
| 317 | PARSE and RENDER are pure functions of their inputs (cacheable, hashable). INDEX/RESOLVE | ||
| 318 | is the only inherently global stage — it is where the link dependency graph is born. | ||
| 319 | |||
| 320 | | Stage | Module | Notes | | ||
| 321 | |---|---|---| | ||
| 322 | | config | `src/config.rs` | `orgo.toml`: site metadata, nav mode, templates, theme. A hash input. | | ||
| 323 | | PARSE | `src/parser.rs` | Hand-written recursive descent: line lexer → element builder → inline tokenizer. | | ||
| 324 | | audit | `src/audit.rs` | Phase 0 corpus audit: construct frequencies against the IN/OUT line. | | ||
| 325 | | model | `src/model.rs` | The org element tree — Elements (block) vs Objects (inline). | | ||
| 326 | | INDEX | `src/index.rs` | Collect link targets into a symbol table. | | ||
| 327 | | RESOLVE | `src/resolve.rs` | Rewrite links to URLs; return the used-target list (dependency edges). | | ||
| 328 | | RENDER | `src/render.rs` | Tree → HTML fragment; syntect highlighting; footnote two-pass. | | ||
| 329 | | TEMPLATE | `src/template.rs` | minijinja: fragment + metadata → full page. | | ||
| 330 | | incremental | `src/incremental.rs` | Content/config/template hashing, dep graph, cache manifest, invalidation. | | ||
| 331 | |||
| 332 | ## v1 scope (delivered as of v0.4; still to be reconciled against a corpus audit) | ||
| 333 | |||
| 334 | **IN — v1 must handle:** headings with nesting, at levels relative to the document's | ||
| 335 | shallowest; TODO keywords; priorities `[#A]`; tags; property drawers; plain lists | ||
| 336 | (unordered/ordered/description, checkboxes, `[@N]` counters, nesting); tables (with rule | ||
| 337 | rows and org's special marker column, no `#+TBLFM:`); source blocks with syntax | ||
| 338 | highlighting; example/quote/center/verse blocks and named special blocks; links (external, | ||
| 339 | internal `[[*Heading]]`/`[[#custom-id]]`, `id:`); footnotes (inline and referenced); `#+` | ||
| 340 | keywords/directives; inline markup (bold/italic/underline/verbatim/code/strike); org's | ||
| 341 | export-time text conversions (`--`/`---`/`...`, `x^2`, `a_{b}`, `\alpha`); timestamps | ||
| 342 | (active/inactive, ranges); paragraphs and horizontal rules; images with | ||
| 343 | `#+CAPTION`/`#+ATTR_HTML`, numbered `Figure N:`. | ||
| 344 | |||
| 345 | **OUT — explicitly not v1 (parse-and-ignore or reject loudly):** Babel execution / | ||
| 346 | `:results`; `#+TBLFM:` formulas; LaTeX / MathJax (passed through untouched, including past | ||
| 347 | the text conversions); `#+INCLUDE:` (never expanded — reported as a diagnostic, so a page | ||
| 348 | is never quietly short of content); citations; radio targets and macros; drawers other | ||
| 349 | than PROPERTIES/LOGBOOK; column view / clocking / agenda semantics; non-HTML export | ||
| 350 | blocks. | ||
| 351 | |||
| 352 | **Scope guardrail:** every IN item gets a golden-file fixture; every OUT item gets a test | ||
| 353 | asserting it degrades predictably (ignored, no crash). The IN/OUT line is enforced by | ||
| 354 | `tests/constructs.rs`, defending against the project's #1 risk: scope creep back toward | ||
| 355 | all-of-org. Phase 0 checked this line against a real 179-file corpus and found it sound | ||
| 356 | (99.9% of construct uses in scope) — but also found one thing missing from it entirely: | ||
| 357 | `#+SLUG:`. See [Phase 0](#phase-0-the-corpus-audit-and-the-emacs-oracle). | ||
| 358 | |||
| 359 | ## Phase plan | ||
| 360 | |||
| 361 | | Phase | Scope | Status | | ||
| 362 | |---|---|---| | ||
| 363 | | **M0** | **Buildable skeleton: crate layout, module stubs, deps, test harness, fixtures** | **done** | | ||
| 364 | | **v0.1** | **End-to-end core parse → render: `build` a single `.org` file to HTML** | **done** | | ||
| 365 | | **v0.2** | **Multi-file SITE build: INDEX + RESOLVE internal links, minijinja templates, `build <src-dir> <out-dir>`, tables + footnotes** | **done** | | ||
| 366 | | **v0.3** | **Incremental build layer: content/config/template hashing, dependency graph, per-page render keys, persisted cache manifest, invalidation** | **done** | | ||
| 367 | | **v0.4** | **MVP: the full v1 construct scope — heading metadata, nested/description lists, block types, timestamps, images, syntect highlighting — with the IN/OUT line under test** | **done** | | ||
| 368 | | **0** | **Corpus audit + `emacs --batch` ground-truth oracle** | **done** | | ||
| 369 | | 1 | Line lexer + heading/section skeleton | done | | ||
| 370 | | 2 | Block elements — lists, source blocks, tables, footnote defs, blocks by type, drawers | done | | ||
| 371 | | 3 | Inline objects — emphasis, links, bare URLs, footnote refs, timestamps | done | | ||
| 372 | | 4 | Rendering to HTML — tree walk, tables, footnote two-pass, minijinja templating, syntect highlighting | done | | ||
| 373 | | 5 | Link resolution + symbol table (INDEX + RESOLVE, used-target list, broken-link reporting) | done | | ||
| 374 | | 6 | Incremental build layer (hashing, dep graph, invalidation); `watch` on OS filesystem events | done | | ||
| 375 | | **7** | **Hardening: rayon parallelism, error locations in parse diagnostics** | **done** | | ||
| 376 | | **8** | **General use: config file, user templates, nav modes, `init` scaffold, safe discovery** | **done** | | ||
| 377 | | **9** | **Generated listing pages: `[[collections]]`, sorted indexes, feeds via XML templates** | **done** | | ||
| 378 | | **10** | **Grouped collections: one page per tag plus a tag index — full parity with the incumbent** | **done** | | ||
| 379 | | **11** | **Pagination: numbered pages with a `paginator` context, composing with grouping** | **done** | | ||
| 380 | | **12** | **`base_url`: `absolute`/`rfc822` filters, a valid RSS feed in the scaffold, canonical links** | **done** | | ||
| 381 | | **13** | **`watch` on OS filesystem events, debounced, with the feedback loop closed** | **done** | | ||
| 382 | | **14** | **Authoring: excerpts, word count, reading time, `truncate`, and draft pages** | **done** | | ||
| 383 | | **15** | **Table of contents, section numbers, and org's `#+OPTIONS:` per-file switches** | **done** | | ||
| 384 | | **16** | **`serve`: development server with long-poll live reload, loopback-bound** | **done** | | ||
| 385 | | **17** | **Bundled TOML and Org syntaxes, a user syntax directory, and org's comma escape** | **done** | | ||
| 386 | | **18** | **Per-page layouts: `[[pages]]` rules and `#+TEMPLATE:`** | **done** | | ||
| 387 | | **19** | **Export parity: relative heading levels, special strings, sub/superscript, caption numbering, checkbox and counter markup, table marker columns, special blocks** | **done** | | ||
| 388 | | **20** | **Correctness debt: org's entity table, table captions, a reported `#+INCLUDE:`, and an oracle that separates deliberate divergence from defects** | **done** | | ||
| 389 | | **21** | **Extra asset roots; per-template hashing so one layout edit does not re-render the site** | **done** | | ||
| 390 | | **22** | **Release engineering: CI on both platforms, a checked MSRV, release binaries, a changelog, and a written compatibility promise** | **done** | | ||
| 391 | |||
| 392 | ### v0.2 in / out | ||
| 393 | |||
| 394 | **Added in v0.2:** the INDEX stage (`SymbolTable` of `:ID:`/`:CUSTOM_ID:`/heading/`file:` | ||
| 395 | targets across a directory); the RESOLVE stage — rewrites `[[#custom-id]]`, `[[id:...]]`, | ||
| 396 | `[[*Heading]]` and `[[file:other.org]]` links to real relative output URLs, returns the | ||
| 397 | `used_targets` list (the `uses` edges, spec §4.3/R2) and reports unresolved links as | ||
| 398 | warnings rather than crashing; a minijinja base layout (title, nav, body) applied to every | ||
| 399 | page; a `build <src-dir> <out-dir>` path that walks the tree, parses + resolves + renders + | ||
| 400 | templates every `.org` into a linked static site and copies non-`.org` assets through; | ||
| 401 | plus two new constructs — pipe **tables** (with header band from the rule row) and | ||
| 402 | **footnotes** (block `[fn:1]` definitions, referenced `[fn:1]`, and inline `[fn:1:text]`, | ||
| 403 | rendered as a numbered, back-linked notes section). | ||
| 404 | |||
| 405 | **Left stubbed at v0.2, all closed in v0.4:** timestamps; TODO keywords and priorities; | ||
| 406 | generic (non-PROPERTIES) drawers; real syntect tokenizing behind the `Highlighter` trait. | ||
| 407 | |||
| 408 | ### v0.3 in / out | ||
| 409 | |||
| 410 | **Added in v0.3 — the incremental build layer (spec §4, the flagship, non-retrofittable | ||
| 411 | feature):** | ||
| 412 | |||
| 413 | - **Three hash classes (spec §4.1)** in `src/incremental.rs`: a **content hash** (blake3 | ||
| 414 | of a file's bytes), a **config hash** (blake3 of the resolved `BuildConfig`), and a | ||
| 415 | **template hash** (blake3 of the template sources). A change in any one invalidates the | ||
| 416 | pages it affects. | ||
| 417 | - **Dependency graph (spec §4.3)** built from RESOLVE's `defines`/`uses` edges: a page | ||
| 418 | depends on the targets it links to, so editing (or renaming a heading in) a file | ||
| 419 | invalidates the pages that *link into* it, not just the file itself — the load-bearing | ||
| 420 | R2 invariant. On rebuild the graph is merged with the previous build's `defines` so a | ||
| 421 | *removed* target still pulls in its linkers. | ||
| 422 | - **Per-page `render_key`** = `H(content ⊕ resolved-links ⊕ config ⊕ template)`. If a | ||
| 423 | page's render key is unchanged, its on-disk output is already correct and it is skipped. | ||
| 424 | The config component folds in a **site-structure hash** (every page's `(path, title)`), | ||
| 425 | because the shared nav bar is global chrome — a title change or a page add/remove alters | ||
| 426 | the nav on every page and so must re-render them all (otherwise byte-equivalence breaks). | ||
| 427 | - **Persisted cache manifest** (`<out>/.orgo-cache.json`, JSON), carrying per-page | ||
| 428 | records, the config/template hashes, and the serialized dependency graph, tagged with | ||
| 429 | `CACHE_FORMAT_VERSION`. A version mismatch, a missing file, or a corrupt file all fall | ||
| 430 | back to a clean full rebuild — the cache is an optimization, never a correctness | ||
| 431 | dependency. | ||
| 432 | - **Wired into `build_site`**: only pages whose render key changed (or that link into a | ||
| 433 | changed file's targets) are re-rendered; unchanged outputs are left in place. `--no-cache` | ||
| 434 | forces a full rebuild; `clean <out-dir>` removes the output directory (and its cache). | ||
| 435 | `SiteReport` now reports `rendered` vs `skipped` counts. | ||
| 436 | |||
| 437 | The hard gates are enforced by `tests/incremental.rs`: full-vs-incremental **byte | ||
| 438 | equivalence** (and a second unchanged build re-rendering **zero** pages); **edit-one-file** | ||
| 439 | re-renders exactly the changed page plus its linkers; **renamed-heading** invalidates the | ||
| 440 | linking page and updates its emitted anchor; and cache **version-bump / missing / corrupt** | ||
| 441 | all fall back to a full rebuild. | ||
| 442 | |||
| 443 | **Out of scope in v0.3:** real syntect highlighting; timestamps and TODO keywords (all | ||
| 444 | landed in v0.4). `watch` is a minimal mtime poll loop (`watch <src-dir> -o <out-dir>`), not | ||
| 445 | an OS file-watcher — the fs-notify integration is deferred. The parse-tree cache (spec §4.5, | ||
| 446 | "optionally") is not persisted: PARSE/INDEX/RESOLVE run for every file each build (cheap and | ||
| 447 | pure); the incremental win is on RENDER + EMIT. | ||
| 448 | |||
| 449 | ### v0.4 in / out — the MVP | ||
| 450 | |||
| 451 | v0.4 closes the gap between the v1 scope above and what the code actually did, so every | ||
| 452 | construct the IN list claims is now parsed, rendered, and pinned by a golden file: | ||
| 453 | |||
| 454 | - **Heading metadata** — TODO keywords (the Emacs default `TODO`/`DONE` set, matched on a | ||
| 455 | word boundary so `TODOs` is not one) and `[#A]` priority cookies, rendered with Emacs' | ||
| 456 | own export classes so the output stays diffable against an `emacs --batch` oracle. | ||
| 457 | - **Lists** — indentation-based nesting (a sub-list renders *inside* its parent `<li>`), | ||
| 458 | multi-paragraph item bodies, and `term :: definition` description lists as `<dl>`. | ||
| 459 | - **Blocks by type** — `QUOTE`, `CENTER`, `EXAMPLE`, `EXPORT` and `SRC` are now distinct | ||
| 460 | elements rather than all collapsing to a verbatim example block. Block matching is on the | ||
| 461 | specific kind, so a source block can nest inside a quote. An `html` export block passes | ||
| 462 | through; every other backend drops. | ||
| 463 | - **Timestamps** — active `<...>` and inactive `[...]`, optional times, same-day time | ||
| 464 | ranges and `--`-joined date ranges, rendered as `<time>` with a machine-readable | ||
| 465 | `datetime`. Repeater/warning cookies are recognized and discarded. | ||
| 466 | - **Images** — a description-less link to an image file renders as `<img>`; with an | ||
| 467 | affiliated `#+CAPTION:`/`#+ATTR_HTML:` it is promoted to a `<figure>` with the caption as | ||
| 468 | both `<figcaption>` and alt text. Links to non-`.org` files are now understood as asset | ||
| 469 | links: neither resolved nor reported as broken. | ||
| 470 | - **Syntax highlighting** — real syntect tokenizing to CSS classes (never inline styles, so | ||
| 471 | themes live in the stylesheet). Every build emits the matching `syntax.css` and each page | ||
| 472 | links it relative to its own depth. An unknown language degrades to escaped `<pre><code>`. | ||
| 473 | - **Diagnostics** — broken links are reported as the org syntax the author wrote | ||
| 474 | (`warning: b.org: unresolved link [[#setup]]`) rather than a Debug-printed enum. | ||
| 475 | |||
| 476 | **The OUT line is now enforced, not just asserted.** `tests/constructs.rs` pins each | ||
| 477 | excluded construct to a specific degradation: babel is never executed *and* a checked-in | ||
| 478 | `#+RESULTS:` block is dropped rather than published as if it were verified output; | ||
| 479 | `#+TBLFM:` is inert; `#+INCLUDE:` is never expanded and says so; LaTeX, macros and radio targets survive | ||
| 480 | as literal text; drawers other than PROPERTIES are captured and dropped; unmodelled block | ||
| 481 | types keep their content verbatim. | ||
| 482 | |||
| 483 | **Still out:** `#+TODO:` per-file keyword sequences; planning lines | ||
| 484 | (`SCHEDULED:`/`DEADLINE:`), which render as ordinary paragraphs; and fixed-width `: ` | ||
| 485 | lines. | ||
| 486 | |||
| 487 | ## Serving | ||
| 488 | |||
| 489 | ```bash | ||
| 490 | cargo run -- serve my-site -o _site # http://127.0.0.1:3000 | ||
| 491 | ``` | ||
| 492 | |||
| 493 | Builds, watches, serves, and reloads the browser when a rebuild lands — the loop `watch` | ||
| 494 | leaves half-open. | ||
| 495 | |||
| 496 | - **Loopback by default.** A dev server serves unreviewed drafts off your laptop, so | ||
| 497 | reaching the local network is something you ask for with `--host 0.0.0.0`, never | ||
| 498 | something you get. | ||
| 499 | - **The reload script is injected on the way out**, never written to disk. What you | ||
| 500 | deploy is the built site, and it must not carry a dev server's JavaScript. | ||
| 501 | - **Long-polling, not WebSockets or SSE.** The browser asks "anything since generation | ||
| 502 | N?" and the server holds the request until there is. Instant like a push, no protocol | ||
| 503 | beyond ordinary HTTP, and no dependency. A streamed response would have been more | ||
| 504 | elegant and does not work: tiny_http buffers a response until its body ends, so a body | ||
| 505 | that never ends never reaches the client. | ||
| 506 | - A reload only follows a **successful** rebuild. Reloading onto a stale page because the | ||
| 507 | build just failed tells you nothing; the error is already on your terminal. | ||
| 508 | |||
| 509 | URL resolution is the server's security boundary and is written as a pure function with | ||
| 510 | its own tests: `..`, percent-encoded `..`, backslashes, absolute paths and embedded NULs | ||
| 511 | all resolve to nothing rather than to somewhere outside the output directory. | ||
| 512 | |||
| 513 | ## Watching | ||
| 514 | |||
| 515 | ```bash | ||
| 516 | cargo run -- watch my-site -o _site | ||
| 517 | ``` | ||
| 518 | |||
| 519 | Rebuilds on OS filesystem events rather than polling, so it costs nothing while nothing | ||
| 520 | happens. Write bursts are debounced — an editor saving a file writes a temp file, renames | ||
| 521 | it over the original and touches the directory, which is one edit and several events. | ||
| 522 | |||
| 523 | Two rules decide what counts as a change, and they are not the same rules the build uses | ||
| 524 | to find content: | ||
| 525 | |||
| 526 | - **A build input is a change.** Editing `orgo.toml` or a template rebuilds, even | ||
| 527 | though discovery skips both as non-content. The question is "would this change the | ||
| 528 | site?", not "is this a page?". | ||
| 529 | - **Our own output is not.** `watch . -o _site` puts the output inside the source, so a | ||
| 530 | rebuild's writes raise events that would trigger a rebuild, forever. Dot-directories go | ||
| 531 | the same way — `.git` churns on every command — as do editor scratch files, including | ||
| 532 | Emacs' `file.org~` backups, which do not start with a dot. | ||
| 533 | |||
| 534 | Where native watching is unavailable (some network and container filesystems), it falls | ||
| 535 | back to polling and says so, rather than failing. | ||
| 536 | |||
| 537 | ## Phase 0: the corpus audit and the Emacs oracle | ||
| 538 | |||
| 539 | The v1 scope was, by its own admission, *recommended* — a guess about which slice of org | ||
| 540 | matters. Phase 0 replaces both halves of that guess with a measurement: an audit that asks | ||
| 541 | what a real corpus actually uses, and an oracle that asks whether we render it the way | ||
| 542 | Emacs does. | ||
| 543 | |||
| 544 | The audit runs against any corpus — point it at your own notes before trusting this tool | ||
| 545 | with them. The numbers below come from a 179-file site published today by weblorg, a | ||
| 546 | wrapper around org's own HTML exporter, which makes it both a realistic workload and a | ||
| 547 | directly comparable incumbent. With collections configured, orgo now reproduces | ||
| 548 | **all 182 of that site's URLs**. | ||
| 549 | |||
| 550 | ``` | ||
| 551 | cargo run -- audit <src-dir> # what does this corpus use, and is it in scope? | ||
| 552 | cargo test --test oracle # how does our HTML differ from Emacs' own export? | ||
| 553 | ``` | ||
| 554 | |||
| 555 | ### What the audit found | ||
| 556 | |||
| 557 | **The scope guess was sound.** 99.9% of construct uses in the corpus are in scope. The | ||
| 558 | whole out-of-scope tail is 8 uses: four `#+TBLFM:` in a post *about* org-mode, three | ||
| 559 | `\name` entities, and one `#+BEGIN_NOTE`. | ||
| 560 | |||
| 561 | **`#+SLUG:` was a hole big enough to sink the project.** 178 of 179 files set it, and the | ||
| 562 | published URL comes from it, not from the filename: `2018-11-28-aes-encryption.org` is | ||
| 563 | served at `blog/aes-encryption.html`. orgo derived output paths from source filenames, | ||
| 564 | so **169 of 179 pages would have been published at the wrong URL** — every inbound link and | ||
| 565 | every search result, broken, by a tool that reported a clean build. Output paths now come | ||
| 566 | from `#+SLUG:` when present ([`util::output_path`](src/util.rs)); slugs are sanitized so an | ||
| 567 | author-supplied `../../etc/x` cannot escape the output directory, and two pages claiming one | ||
| 568 | URL is a build error rather than a silently dropped page. Building the real corpus now | ||
| 569 | reproduces all 179 of the live site's URLs exactly. | ||
| 570 | |||
| 571 | **Some machinery is speculative.** The corpus contains no `id:`, `#custom-id` or `*Heading` | ||
| 572 | links at all — its cross-page links are hand-written relative URLs. The INDEX/RESOLVE | ||
| 573 | symbol table that v0.2 was built around is, against this corpus, unexercised. | ||
| 574 | |||
| 575 | **An audit can lie too.** The first run reported 23 uses of a custom TODO keyword sequence. | ||
| 576 | All 23 were false: the detector read the leading word of `* CSS Variables` as the keyword | ||
| 577 | `CSS`. The corpus defines no `#+TODO:` sequences at all, so the true count was zero. The | ||
| 578 | detector now matches conventional keyword names only — a tool that overstates a gap argues | ||
| 579 | for work nobody needs. | ||
| 580 | |||
| 581 | ### What the oracle found | ||
| 582 | |||
| 583 | `tests/oracle.rs` exports each fixture with org's own exporter via `emacs --batch`, reduces | ||
| 584 | both sides to a semantic skeleton (element opens, closes and text, with layout `div`s, | ||
| 585 | inline `span`s and all attributes but `href`/`src` dropped), and **snapshots the | ||
| 586 | disagreement**. Snapshotting rather than asserting is deliberate: a checked-in divergence | ||
| 587 | report gets reviewed and shows up as a diff, where a permanently red test gets ignored. | ||
| 588 | Three invariants are asserted outright, and all three hold — heading structure, list | ||
| 589 | nesting, and source-block text match Emacs exactly. | ||
| 590 | |||
| 591 | **No bugs in orgo.** Every remaining divergence is a deliberate choice to emit better | ||
| 592 | HTML than org does: | ||
| 593 | |||
| 594 | | | orgo | Emacs | why | | ||
| 595 | |---|---|---|---| | ||
| 596 | | emphasis | `<em>`/`<strong>` | `<i>`/`<b>` | semantic, not presentational | | ||
| 597 | | captioned image | `<figure>`/`<figcaption>` | `<p>` + `"Figure 1: …"` | real figure semantics | | ||
| 598 | | timestamp | `<time datetime="…">` | literal `<2024-01-15 Mon>` | machine-readable | | ||
| 599 | | footnotes | `<section><ol>` | `<h2>Footnotes:</h2>` | a list of notes is a list | | ||
| 600 | | heading anchor | slug of the text | `org1a2b3c4` | stable, and what the live site serves | | ||
| 601 | | code | `<pre><code>` | `<pre>` | the HTML5 idiom | | ||
| 602 | |||
| 603 | One genuine semantic difference: org treats a single blank line between a `1.` list and a | ||
| 604 | `-` list as *one* list and keeps the first item's bullet type, while we start a second list. | ||
| 605 | We keep ours, on measurement rather than taste — the pattern occurs **zero** times in the | ||
| 606 | corpus, so matching an org quirk would buy nothing and cost the more obvious reading. | ||
| 607 | |||
| 608 | **The oracle's best catch was three bugs in itself.** Naive normalization reported code as | ||
| 609 | corrupted (it trimmed each of syntect's per-token text runs, turning `def greet` into | ||
| 610 | `defgreet`) and reported blocks at 36% agreement (syntect's spans flooded the diff). Both | ||
| 611 | were measurement artifacts. A differential harness is a piece of software like any other, | ||
| 612 | and the first divergences it reports are usually its own. | ||
| 613 | |||
| 614 | ## Phase 7: hardening | ||
| 615 | |||
| 616 | ### Parse diagnostics (`file:line: message`) | ||
| 617 | |||
| 618 | The parser's contract is that it always returns a document — out-of-scope and malformed | ||
| 619 | constructs degrade rather than crash. The gap was that they degraded *silently*, and in the | ||
| 620 | worst cases the degradation is severe: an unterminated `#+BEGIN_SRC` reads the rest of the | ||
| 621 | file as block content, and an unterminated drawer does the same but renders to nothing, so | ||
| 622 | one missing line deletes most of a page from a build that reports success. | ||
| 623 | |||
| 624 | `parse` now returns `Document::diagnostics`, each carrying a 1-based source line, and the | ||
| 625 | build prints them as `file:line: message`. `--strict` turns them (and unresolved links) into | ||
| 626 | a non-zero exit. Line numbers are threaded as an absolute offset through every nested parse, | ||
| 627 | so a block inside a list item inside a section still reports its real file line — there is a | ||
| 628 | test for exactly that, because reconstructed and re-indented nested slices are precisely | ||
| 629 | where an off-by-N hides. The 179-file corpus produces zero diagnostics. | ||
| 630 | |||
| 631 | ### Parallelism | ||
| 632 | |||
| 633 | PARSE, RESOLVE and RENDER/EMIT run under rayon. PARSE is a pure function of one file's bytes | ||
| 634 | and RESOLVE only reads the shared symbol table, which is what makes both safe to parallelize | ||
| 635 | at all; INDEX stays sequential. | ||
| 636 | |||
| 637 | | corpus | before | after | speedup | | ||
| 638 | |---|---|---|---| | ||
| 639 | | 179 files (real) | 0.23s | 0.07s | 3.3× | | ||
| 640 | | 1,790 files (10× copy) | 3.98s | 0.82s | 4.9× | | ||
| 641 | |||
| 642 | Measured on 12 cores. `RAYON_NUM_THREADS=1` reproduces the old 3.98s exactly, so the gain is | ||
| 643 | parallelism rather than incidental change, and the output is byte-identical to the sequential | ||
| 644 | build across the whole corpus. | ||
| 645 | |||
| 646 | **Parallelism must not be observable in the result.** `par_iter().collect()` preserves input | ||
| 647 | order, so the emitted bytes are unaffected — but the build *report* is the fragile half: | ||
| 648 | pushing to `rendered`/`skipped` from inside the parallel pass would order them by thread | ||
| 649 | scheduling, producing a non-deterministic report over a deterministic site. The parallel pass | ||
| 650 | therefore returns only what was written, and the report is assembled sequentially afterwards. | ||
| 651 | `parallel_builds_are_deterministic_in_output_and_report_order` holds that line, and it was | ||
| 652 | verified by reintroducing the bug and watching it fail. | ||
| 653 | |||
| 654 | ### The real scaling limit was not the CPU | ||
| 655 | |||
| 656 | Going 10× on corpus size cost 17× in time, which parallelism improves without fixing: the | ||
| 657 | cause was the nav bar listing **every** page, so an *n*-page site emitted *n*² nav links. At | ||
| 658 | 1,790 pages each page carried 1,799 links and the output was 284 MB, against 5.5 MB for the | ||
| 659 | 179-page corpus — 52× the bytes for 10× the input. | ||
| 660 | |||
| 661 | The nav is now built from **top-level pages only** ([`is_top_level`](src/site.rs)): a nav is a | ||
| 662 | map of the site's top level, not an index of its contents, and section pages reach their | ||
| 663 | siblings through that section's landing page. Nav size becomes a function of the top level | ||
| 664 | rather than of the corpus, and the quadratic disappears. | ||
| 665 | |||
| 666 | | 1,790-page corpus (6 top-level pages) | before | after | | ||
| 667 | |---|---|---| | ||
| 668 | | full build | 0.82s | 0.39s | | ||
| 669 | | total output | 284 MB | 34 MB | | ||
| 670 | | nav links per page | 1,799 | 6 | | ||
| 671 | |||
| 672 | Scaling is now linear: 179 pages in 0.07s and 1,796 in 0.39s, where the small case is mostly | ||
| 673 | the fixed cost of loading syntect's syntax definitions. | ||
| 674 | |||
| 675 | The same rule sharpened the incremental build, which is the larger win. The site-structure | ||
| 676 | hash — the thing that forces a global re-render — now covers only the pages that appear in | ||
| 677 | the nav, because those are the only ones whose title or URL affects another page. **Adding a | ||
| 678 | blog post used to re-render the entire site; now it renders one page.** A top-level page's | ||
| 679 | title still invalidates everything, correctly, since every page displays it. | ||
| 680 | |||
| 681 | **Trade-off worth knowing:** on a site whose sections live in subdirectories, only genuinely | ||
| 682 | root-level pages appear — a site keeping its landing pages at `salary/index.org` and friends | ||
| 683 | gets a one-entry nav. That is what `nav.mode = "explicit"` is for: list the pages you want, | ||
| 684 | in the order you want them. | ||
| 685 | |||
| 686 | **From v0.1 (core subset):** headings with nesting and anchors (every heading is now | ||
| 687 | anchored — `:CUSTOM_ID:`/`:ID:` else a slug of its text) and trailing tags; paragraphs; | ||
| 688 | plain lists (unordered + ordered) with checkboxes; source blocks; inline markup (`*bold*`, | ||
| 689 | `/italic/`, `_underline_`, `+strike+`, `=verbatim=`, `~code~`); links and bare URLs. | ||
| 690 | |||
| 691 | ## Compatibility | ||
| 692 | |||
| 693 | Versions mean something as of 1.0. The **stable surface** — changing incompatibly requires | ||
| 694 | a major version — is what you actually build a site against: | ||
| 695 | |||
| 696 | | Stable | Detail | | ||
| 697 | |---|---| | ||
| 698 | | `orgo.toml` keys | Names, types and meaning. New keys are minor releases; removing one is major. | | ||
| 699 | | Template context | `page`, `site`, `nav`, `root`, `pages`, `group`, `groups`, `paginator`, `stylesheet`, and the `absolute` / `rfc822` / `truncate` filters. | | ||
| 700 | | CLI | Command names, flags, and exit codes. | | ||
| 701 | | URLs | How a source path becomes an output path, including `#+SLUG:`. A generator that moves your URLs breaks every link anyone has to you. | | ||
| 702 | |||
| 703 | Explicitly **not stable**, so that the above can be: | ||
| 704 | |||
| 705 | - **The incremental cache.** Versioned, discarded on mismatch, never a correctness | ||
| 706 | dependency. It changes whenever it needs to, in any release. | ||
| 707 | - **Rendered HTML details.** orgo tracks what Emacs exports from the same file, and | ||
| 708 | closing a gap changes markup. Changes that affect output are called out in | ||
| 709 | [CHANGELOG.md](CHANGELOG.md) — the class names the documentation names (`post-list`, | ||
| 710 | `figure-number`, `section-number-N`, `footnote-ref`) are the ones to write CSS against. | ||
| 711 | - **The Rust API.** The crate is published so the binary can be installed with | ||
| 712 | `cargo install`; the library exists to serve it, and its types move as the tool does. | ||
| 713 | |||
| 714 | The **MSRV is 1.88**, checked in CI on every change. orgo's own code compiles on | ||
| 715 | 1.82; the floor comes from dependencies. Raising it is a minor version, never a patch. | ||
| 716 | |||
| 717 | ## Dependencies | ||
| 718 | |||
| 719 | Parser is hand-written recursive descent (not `nom`/`chumsky`/`pest` — org is | ||
| 720 | line-oriented and context-sensitive, not clean CFG). Key crates: `syntect` (syntax | ||
| 721 | highlighting, behind a `Highlighter` trait so tree-sitter can be swapped in later), | ||
| 722 | `minijinja` (runtime templates), `blake3` (content/cache hashing), `rayon` (parallel | ||
| 723 | PARSE/RESOLVE/RENDER), `notify` (filesystem events for `watch`), `tiny_http` (the `serve` | ||
| 724 | development server), `toml` (config), `chrono`, `camino`, `walkdir`, `clap`, `anyhow`/`thiserror`. | ||
| 725 | `insta` for snapshot tests, and `emacs --batch` — optional, and only for the oracle. | ||
| 726 | |||
| 727 | ## Build & test | ||
| 728 | |||
| 729 | ``` | ||
| 730 | cargo build | ||
| 731 | cargo test # 191 tests | ||
| 732 | cargo run -- init my-site # scaffold a new site | ||
| 733 | cargo run -- build fixtures/minimal.org -o minimal.html # single file | ||
| 734 | cargo run -- build fixtures/site -o _site # whole site (incremental) | ||
| 735 | cargo run -- audit fixtures/site # corpus audit (Phase 0) | ||
| 736 | cargo run -- build fixtures/site -o _site --no-cache # force a full rebuild | ||
| 737 | cargo run -- watch fixtures/site -o _site # rebuild on filesystem events | ||
| 738 | cargo run -- serve fixtures/site -o _site # ... and serve with live reload | ||
| 739 | cargo run -- clean _site # remove output + cache | ||
| 740 | ``` | ||
| 741 | |||
| 742 | A second `build` of an unchanged site re-renders nothing; editing a page re-renders only | ||
| 743 | that page and the pages that link into it (watch the `rendered`/`cached` counts). | ||
| 744 | |||
| 745 | A build emits `syntax.css` next to its output (the highlighter emits CSS classes, so the | ||
| 746 | stylesheet has to come with them) and every page links it. | ||
| 747 | |||
| 748 | `fixtures/` holds tiny `.org` samples: the core ones (`minimal.org`, `core.org`, | ||
| 749 | `elements.org`, `table.org`, `footnote.org`), one per v1 construct group (`headings.org`, | ||
| 750 | `lists.org`, `blocks.org`, `timestamps.org`, `images.org`), the scope guardrail | ||
| 751 | (`outofscope.org`), and a linked multi-file site under `fixtures/site/` (`index.org`, | ||
| 752 | `guide.org`, `about.org` + a `style.css` asset). The real corpus (golden files derived from | ||
| 753 | actual documents) lands in Phase 0. `cargo test` runs `insta` snapshots of the element tree | ||
| 754 | and rendered HTML for each fixture, the two templated site pages (proving cross-file link | ||
| 755 | resolution), and the incremental gates. | ||
README.org added +726
| @@ -0,0 +1,726 @@ | |||
| 1 | * orgo | ||
| 2 | An org-mode static site generator, in Rust. Org is treated as the /source language/, | ||
| 3 | not an inconvenient input to be normalized into markdown. The org element tree — | ||
| 4 | headings, drawers, blocks, links with their org-specific semantics — *is* the | ||
| 5 | document model, and we render that tree straight to HTML. We never round-trip through | ||
| 6 | a markdown-shaped intermediate representation, because the point is to preserve what | ||
| 7 | markdown cannot express: property drawers, TODO/priority/tag metadata on headings, | ||
| 8 | =#+= directives, ID links, named/captioned blocks, footnote semantics. | ||
| 9 | |||
| 10 | The one non-obvious early commitment is *incremental builds keyed on content | ||
| 11 | hashing*, treated as a first-class architectural concern from day one. The discipline | ||
| 12 | it imposes on the data model — pure, hashable, dependency-tracked units — is the real | ||
| 13 | deliverable, even while the corpus is small enough that a full rebuild is instant. | ||
| 14 | |||
| 15 | *Full documentation is in [[file:docs/][=docs/=]]* — a site written in org and built by | ||
| 16 | orgo itself. Build and read it with: | ||
| 17 | |||
| 18 | #+begin_src sh | ||
| 19 | cargo run -- serve docs -o docs/_site | ||
| 20 | #+end_src | ||
| 21 | |||
| 22 | ** Quick start | ||
| 23 | #+begin_src sh | ||
| 24 | cargo run -- init my-site # config + an editable copy of the layout + a page | ||
| 25 | cargo run -- build my-site -o _site | ||
| 26 | #+end_src | ||
| 27 | |||
| 28 | Or skip the scaffolding entirely — point it at any directory of =.org= files: | ||
| 29 | |||
| 30 | #+begin_src sh | ||
| 31 | cargo run -- build ~/notes -o _site | ||
| 32 | #+end_src | ||
| 33 | |||
| 34 | *Zero configuration is a supported path, not a demo.* With no =orgo.toml=, no | ||
| 35 | templates and no orgo-specific markup in your files, you get a complete site: pages, | ||
| 36 | navigation, syntax-highlighted code and the stylesheet to colour it. Configuration | ||
| 37 | changes what you get; it is never what makes it work. | ||
| 38 | |||
| 39 | Discovery skips what should not be published — dot-directories such as =.git=, the config | ||
| 40 | file, the templates directory, and the output directory when it sits inside the source, so | ||
| 41 | =orgo build . -o _site= does the obvious thing. | ||
| 42 | |||
| 43 | ** Configuration | ||
| 44 | Everything is optional. =orgo init= writes a fully commented =orgo.toml=; every | ||
| 45 | value below is the default. | ||
| 46 | |||
| 47 | #+begin_src toml | ||
| 48 | [site] | ||
| 49 | title = "orgo site" | ||
| 50 | base_url = "" # absolute URL, no trailing slash; needed for feeds/canonical links | ||
| 51 | description = "" | ||
| 52 | language = "en" | ||
| 53 | |||
| 54 | [nav] | ||
| 55 | mode = "top-level" # top-level | all | explicit | none | ||
| 56 | # pages = ["index.org", "about.org"] # for mode = "explicit"; order is preserved | ||
| 57 | |||
| 58 | [templates] | ||
| 59 | dir = "templates" # base.html replaces the built-in layout | ||
| 60 | expose_page_list = false | ||
| 61 | |||
| 62 | # [[pages]] # which layout a section renders through; base.html by default | ||
| 63 | # match = "blog" # a source directory or one .org file; most specific rule wins | ||
| 64 | # template = "post.html" | ||
| 65 | |||
| 66 | [highlight] | ||
| 67 | theme = "InspiredGitHub" | ||
| 68 | |||
| 69 | [build] | ||
| 70 | drafts = false | ||
| 71 | assets = [] # extra directories copied to the site root, e.g. ["../theme/static"] | ||
| 72 | |||
| 73 | [html] | ||
| 74 | heading_offset = 1 # a level-1 org heading becomes <h2>, beneath the layout's <h1> | ||
| 75 | #+end_src | ||
| 76 | |||
| 77 | *** Templates | ||
| 78 | Drop a =base.html= into the templates directory and it replaces the built-in layout | ||
| 79 | entirely. Any other =.html= file there is available to ={% include %}= and | ||
| 80 | ={% extends %}=. Templates are [[https://docs.rs/minijinja][minijinja]] (Jinja2 syntax) and | ||
| 81 | receive: | ||
| 82 | |||
| 83 | | Variable | What it is | | ||
| 84 | |————--+————————————————————————————————————————————————--| | ||
| 85 | | =body= | the rendered page HTML — use ={{ body \| safe }}= | | ||
| 86 | | =page= | =.title=, =.url=, =.source=, =.date=, =.date_iso=, =.year=, =.tags=, =.content=, =.excerpt=, =.word_count=, =.reading_time=, =.toc=, =.keywords= | | ||
| 87 | | =site= | =.title=, =.base_url=, =.description=, =.language= | | ||
| 88 | | =nav= | list of ={title, url}=, relative to this page | | ||
| 89 | | =root= | =../=-prefix back to the site root from this page | | ||
| 90 | | =stylesheet= | URL of the generated =syntax.css= | | ||
| 91 | | =pages= | every page's metadata — only when =expose_page_list = true= | | ||
| 92 | |||
| 93 | =page.keywords= carries *every* =#+KEYWORD:= in the file under its lowercased name, so | ||
| 94 | your own metadata works without this crate knowing about it: =#+CUSTOM_THING: x= is | ||
| 95 | ={{ page.keywords.custom_thing }}=. | ||
| 96 | |||
| 97 | =base.html= is the default layout, not the only one. A =[[pages]]= rule gives a section | ||
| 98 | its own — =match = "blog"=, =template = "post.html"= — and =#+TEMPLATE: wide.html= gives | ||
| 99 | one page its own, which wins over any rule. A second layout usually starts with | ||
| 100 | ={% extends "base.html" %}=. | ||
| 101 | |||
| 102 | Editing a template re-renders the pages that use it — template sources are a hash input, | ||
| 103 | so a design change never leaves a site half-updated. | ||
| 104 | |||
| 105 | *** Generated listing pages | ||
| 106 | A blog index, an archive, a feed — output files with no source =.org= behind them. | ||
| 107 | Repeat the block for each one: | ||
| 108 | |||
| 109 | #+begin_src toml | ||
| 110 | [[collections]] | ||
| 111 | source = "blog" # directory to list; empty means every page | ||
| 112 | output = "blog/index.html" # where to write it | ||
| 113 | template = "list.html" | ||
| 114 | title = "Blog" | ||
| 115 | sort = "date" # date | title | path | ||
| 116 | order = "desc" # desc | asc | ||
| 117 | nav = true # put this listing page in the nav | ||
| 118 | #+end_src | ||
| 119 | |||
| 120 | The template gets the collection's entries as =pages=, already sorted, plus the usual | ||
| 121 | =site=/=nav=/=root=. It can ={% extends "base.html" %}= to inherit the site chrome: | ||
| 122 | |||
| 123 | #+begin_src jinja | ||
| 124 | {% extends "base.html" %} | ||
| 125 | {% block main %} | ||
| 126 | <ul>{% for p in pages %} | ||
| 127 | <li><time datetime="{{ p.date_iso }}">{{ p.date_iso }}</time> | ||
| 128 | <a href="{{ root }}{{ p.url }}">{{ p.title }}</a></li> | ||
| 129 | {% endfor %}</ul> | ||
| 130 | {% endblock %} | ||
| 131 | #+end_src | ||
| 132 | |||
| 133 | =p.date_iso= is the =YYYY-MM-DD= extracted from =#+DATE:=, whatever org syntax it was | ||
| 134 | written in — =[2025-09-05 Fri 10:21:00]=, =<2024-05-01 Wed>= or bare =2024-05-01=. It is | ||
| 135 | also the sort key; pages without a parseable date sort last, so an undated draft never | ||
| 136 | leads a dated archive. | ||
| 137 | |||
| 138 | **** Pagination | ||
| 139 | Set =paginate= to split a long listing across numbered pages: | ||
| 140 | |||
| 141 | #+begin_src toml | ||
| 142 | [[collections]] | ||
| 143 | source = "blog" | ||
| 144 | output = "blog/index.html" | ||
| 145 | paginate = 10 | ||
| 146 | paginate_output = "blog/page/{n}.html" # {n} is the 1-based page number | ||
| 147 | #+end_src | ||
| 148 | |||
| 149 | Page 1 stays at =output=, so a section's canonical URL never moves as its page count | ||
| 150 | changes; only pages 2..N are named by =paginate_output=. The template gets a =paginator=: | ||
| 151 | |||
| 152 | #+begin_src jinja | ||
| 153 | {% if paginator and paginator.total > 1 %} | ||
| 154 | <nav> | ||
| 155 | {% if paginator.prev_url %}<a href="{{ paginator.prev_url }}">Newer</a>{% endif %} | ||
| 156 | {% for pg in paginator.pages %} | ||
| 157 | <a href="{{ pg.url }}"{% if pg.current %} aria-current="page"{% endif %}>{{ pg.number }}</a> | ||
| 158 | {% endfor %} | ||
| 159 | {% if paginator.next_url %}<a href="{{ paginator.next_url }}">Older</a>{% endif %} | ||
| 160 | </nav> | ||
| 161 | {% endif %} | ||
| 162 | #+end_src | ||
| 163 | |||
| 164 | =paginator= carries =current=, =total=, =per_page=, =total_entries=, =prev_url=, | ||
| 165 | =next_url=, =first_url=, =last_url=, and =pages=. Every URL is relative to the page | ||
| 166 | carrying it, so links work from page 1 (=page/2.html=) and from page 5 (=../index.html=, | ||
| 167 | =6.html=) without the template knowing where it sits. An unpaginated collection has no | ||
| 168 | =paginator= at all, so ={% if paginator %}= is a reliable test in a shared template. | ||
| 169 | |||
| 170 | Grouping and pagination compose: each group paginates independently, which is why | ||
| 171 | =paginate_output= needs ={tag}= as well as ={n}= on a grouped collection. An empty | ||
| 172 | collection still emits page 1 — a section that exists but has nothing in it should say so | ||
| 173 | rather than 404. When the entry count shrinks, pages that no longer exist are deleted | ||
| 174 | instead of being left serving stale posts. | ||
| 175 | |||
| 176 | **** Tag pages | ||
| 177 | Add =group_by= and the collection emits one page /per group/ instead of one page total, | ||
| 178 | plus an optional index of the groups: | ||
| 179 | |||
| 180 | #+begin_src toml | ||
| 181 | [[collections]] | ||
| 182 | source = "blog" | ||
| 183 | group_by = "tags" # "tags", or any #+KEYWORD: name to group by its value | ||
| 184 | output = "tags/{tag}.html" # {tag} is replaced by each group's slug | ||
| 185 | template = "tag.html" | ||
| 186 | title = "Tagged: {tag}" | ||
| 187 | index_output = "tags/index.html" # the tag index | ||
| 188 | index_template = "tags.html" | ||
| 189 | index_title = "Tags" | ||
| 190 | nav = true # adds the *index*, not every tag | ||
| 191 | #+end_src | ||
| 192 | |||
| 193 | A group page receives its own posts as =pages= and itself as =group= | ||
| 194 | (=.name=, =.slug=, =.url=, =.count=). The index receives =groups= — every group, sorted | ||
| 195 | by name: | ||
| 196 | |||
| 197 | #+begin_src jinja | ||
| 198 | <ul>{% for tag in groups %} | ||
| 199 | <li><a href="{{ root }}{{ tag.url }}">{{ tag.name }}</a> ({{ tag.count }})</li> | ||
| 200 | {% endfor %}</ul> | ||
| 201 | #+end_src | ||
| 202 | |||
| 203 | =group_by = "tags"= is multi-valued: a post appears under every tag it carries. Any other | ||
| 204 | value names a single-valued =#+KEYWORD:=, so =group_by = "category"= buckets by | ||
| 205 | =#+CATEGORY:=. | ||
| 206 | |||
| 207 | Two tags that would produce the same URL (=web_dev= and =web@dev= both slugify to | ||
| 208 | =web-dev=) are a build error rather than one page silently overwriting the other. | ||
| 209 | |||
| 210 | A tag page depends on its own posts and nothing else, so adding a post tagged =rust= | ||
| 211 | re-renders that post, its section index, =tags/rust.html=, and the tag index whose counts | ||
| 212 | changed — four pages, not one per tag. That precision is why =groups= is given to the | ||
| 213 | index and not to every group page: a page that can see every group depends on every | ||
| 214 | group. | ||
| 215 | |||
| 216 | **** Feeds and absolute URLs | ||
| 217 | *A feed is a listing page with an XML template*, not a separate feature — templates are | ||
| 218 | loaded by full filename and any extension, so =output = "feed.xml"= with | ||
| 219 | =template = "feed.xml"= is all it takes. =orgo init= writes a working RSS template. | ||
| 220 | |||
| 221 | A feed is read away from the site that served it, so relative links in one are simply | ||
| 222 | broken. Set =site.base_url= and use the =absolute= filter: | ||
| 223 | |||
| 224 | #+begin_src jinja | ||
| 225 | <link>{{ post.url | absolute }}</link> | ||
| 226 | <pubDate>{{ post.date_iso | rfc822 }}</pubDate> | ||
| 227 | #+end_src | ||
| 228 | |||
| 229 | | Filter | Does | | ||
| 230 | |—————+—————————————————————————-| | ||
| 231 | | =absolute= | site-root-relative path → absolute URL; already-absolute URLs pass through | | ||
| 232 | | =rfc822= | any org or ISO date → the format RSS =pubDate= requires | | ||
| 233 | | =truncate(n)= | shorten to at most =n= characters on a word boundary, with an ellipsis | | ||
| 234 | |||
| 235 | Apply =absolute= to the site-root-relative values — =page.url=, =pages[].url=, | ||
| 236 | =group.url= — and not to =nav[].url=, =paginator.*_url=, =stylesheet= or =root=, which | ||
| 237 | are relative to the page carrying them and already correct there. | ||
| 238 | |||
| 239 | With no =base_url=, =absolute= is an *error* naming the setting, rather than quietly | ||
| 240 | emitting a relative URL that would make the feed invalid everywhere while looking fine. | ||
| 241 | The default layout also emits =<link rel="canonical">= when a base URL is set. | ||
| 242 | |||
| 243 | Listing pages are cached on the entries they list, so adding a post re-renders that | ||
| 244 | section's index and nothing else. | ||
| 245 | |||
| 246 | *** Table of contents and =#+OPTIONS:= | ||
| 247 | =page.toc= is the page's headings as a *tree* — ={title, anchor, level, children}= — | ||
| 248 | because a table of contents is one, and rebuilding a tree from a flat list of levels | ||
| 249 | inside a template is what Jinja is worst at. Its anchors come from the same function the | ||
| 250 | renderer uses to emit heading =id=s, so a TOC link cannot drift from the heading it | ||
| 251 | points at. | ||
| 252 | |||
| 253 | #+begin_src jinja | ||
| 254 | {% macro toc_list(entries) %} | ||
| 255 | <ul>{% for e in entries %} | ||
| 256 | <li><a href="#{{ e.anchor }}">{{ e.title }}</a> | ||
| 257 | {%- if e.children %}{{ toc_list(e.children) }}{% endif %}</li> | ||
| 258 | {% endfor %}</ul> | ||
| 259 | {% endmacro %} | ||
| 260 | {% if page.toc %}{{ toc_list(page.toc) }}{% endif %} | ||
| 261 | #+end_src | ||
| 262 | |||
| 263 | Org's own per-file export switches are honoured, so a document can turn a feature off for | ||
| 264 | itself the way its author already knows: | ||
| 265 | |||
| 266 | | Switch | Effect | Site default | | ||
| 267 | |———————-+————————————+———————————-| | ||
| 268 | | =#+OPTIONS: toc:nil= | empties =page.toc= for this page | =[html] toc = true= | | ||
| 269 | | =#+OPTIONS: num:t= | numbers headings =1.=, =1.1.=, … | =[html] section_numbers = false= | | ||
| 270 | |||
| 271 | *Section numbers default to off, which differs from Emacs on purpose.* | ||
| 272 | =org-export-with-section-numbers= is on there, so an org-published site inherits numbered | ||
| 273 | headings whether or not anyone chose them. Most sites do not want them; =num:t= or | ||
| 274 | =section_numbers = true= gets Emacs' behaviour back, with Emacs' own | ||
| 275 | =section-number-N= classes so the output stays diffable against the oracle. | ||
| 276 | |||
| 277 | *** Excerpts and drafts | ||
| 278 | =page.excerpt= is a page's =#+DESCRIPTION:= when it sets one and its first paragraph | ||
| 279 | otherwise, so a listing has something to show whether or not the author thought about | ||
| 280 | summaries. =page.word_count= and =page.reading_time= (minutes at 200 wpm) count prose | ||
| 281 | only — a post that is mostly a shell transcript should not read as an hour's work. | ||
| 282 | =truncate= exists because an excerpt is usually a whole paragraph and minijinja has no | ||
| 283 | such filter. | ||
| 284 | |||
| 285 | =#+DRAFT:= keeps a page out of the build entirely — no page, and absent from listings and | ||
| 286 | the nav rather than merely unlinked. =--drafts= includes them, which is what you want | ||
| 287 | under =watch= while writing one. A draft is out of the symbol table too, so a link /to/ | ||
| 288 | one is reported as the dead link it would be once published. | ||
| 289 | |||
| 290 | The keyword is read forgivingly: =t=, =yes=, =1= and a bare =#+DRAFT:= all mean draft, | ||
| 291 | because writing the keyword at all is the signal. Only an explicit =nil=, =false=, =no=, | ||
| 292 | =0= or =off= means published. | ||
| 293 | |||
| 294 | *** =#+SLUG:= | ||
| 295 | A page's output filename comes from its =#+SLUG:= when it has one, so | ||
| 296 | =2018-11-28-aes-encryption.org= can publish as =aes-encryption.html=. Without one the | ||
| 297 | source filename is used. Slugs are sanitized to a single safe path component, and two | ||
| 298 | pages claiming one URL is a build error rather than a silently dropped page. | ||
| 299 | |||
| 300 | ** Pipeline | ||
| 301 | #+begin_example | ||
| 302 | DISCOVER → PARSE → INDEX → RESOLVE → RENDER → TEMPLATE → EMIT | ||
| 303 | #+end_example | ||
| 304 | |||
| 305 | PARSE and RENDER are pure functions of their inputs (cacheable, hashable). INDEX/RESOLVE | ||
| 306 | is the only inherently global stage — it is where the link dependency graph is born. | ||
| 307 | |||
| 308 | | Stage | Module | Notes | | ||
| 309 | |————-+———————-+———————————————————————————-| | ||
| 310 | | config | =src/config.rs= | =orgo.toml=: site metadata, nav mode, templates, theme. A hash input. | | ||
| 311 | | PARSE | =src/parser.rs= | Hand-written recursive descent: line lexer → element builder → inline tokenizer. | | ||
| 312 | | audit | =src/audit.rs= | Phase 0 corpus audit: construct frequencies against the IN/OUT line. | | ||
| 313 | | model | =src/model.rs= | The org element tree — Elements (block) vs Objects (inline). | | ||
| 314 | | INDEX | =src/index.rs= | Collect link targets into a symbol table. | | ||
| 315 | | RESOLVE | =src/resolve.rs= | Rewrite links to URLs; return the used-target list (dependency edges). | | ||
| 316 | | RENDER | =src/render.rs= | Tree → HTML fragment; syntect highlighting; footnote two-pass. | | ||
| 317 | | TEMPLATE | =src/template.rs= | minijinja: fragment + metadata → full page. | | ||
| 318 | | incremental | =src/incremental.rs= | Content/config/template hashing, dep graph, cache manifest, invalidation. | | ||
| 319 | |||
| 320 | ** v1 scope (delivered as of v0.4; still to be reconciled against a corpus audit) | ||
| 321 | *IN — v1 must handle:* headings with nesting, at levels relative to the document's | ||
| 322 | shallowest; TODO keywords; priorities =[#A]=; tags; property drawers; plain lists | ||
| 323 | (unordered/ordered/description, checkboxes, =[@N]= counters, nesting); tables (with rule | ||
| 324 | rows and org's special marker column, no =#+TBLFM:=); source blocks with syntax | ||
| 325 | highlighting; example/quote/center/verse blocks and named special blocks; links (external, | ||
| 326 | internal =[[*Heading]]=/=[[#custom-id]]=, =id:=); footnotes (inline and referenced); =#+= | ||
| 327 | keywords/directives; inline markup (bold/italic/underline/verbatim/code/strike); org's | ||
| 328 | export-time text conversions (=--=/=---=/=...=, =x^2=, =a_{b}=, =\alpha=); timestamps | ||
| 329 | (active/inactive, ranges); paragraphs and horizontal rules; images with | ||
| 330 | =#+CAPTION=/=#+ATTR_HTML=, numbered =Figure N:=. | ||
| 331 | |||
| 332 | *OUT — explicitly not v1 (parse-and-ignore or reject loudly):* Babel execution / | ||
| 333 | =:results=; =#+TBLFM:= formulas; LaTeX / MathJax (passed through untouched, including past | ||
| 334 | the text conversions); =#+INCLUDE:= (never expanded — reported as a diagnostic, so a page | ||
| 335 | is never quietly short of content); citations; radio targets and macros; drawers other | ||
| 336 | than PROPERTIES/LOGBOOK; column view / clocking / agenda semantics; non-HTML export | ||
| 337 | blocks. | ||
| 338 | |||
| 339 | *Scope guardrail:* every IN item gets a golden-file fixture; every OUT item gets a test | ||
| 340 | asserting it degrades predictably (ignored, no crash). The IN/OUT line is enforced by | ||
| 341 | =tests/constructs.rs=, defending against the project's #1 risk: scope creep back toward | ||
| 342 | all-of-org. Phase 0 checked this line against a real 179-file corpus and found it sound | ||
| 343 | (99.9% of construct uses in scope) — but also found one thing missing from it entirely: | ||
| 344 | =#+SLUG:=. See [[#phase-0-the-corpus-audit-and-the-emacs-oracle][Phase 0]]. | ||
| 345 | |||
| 346 | ** Phase plan | ||
| 347 | | Phase | Scope | Status | | ||
| 348 | |——--+——————————————————————————————————————————————————————————+——--| | ||
| 349 | | *M0* | *Buildable skeleton: crate layout, module stubs, deps, test harness, fixtures* | *done* | | ||
| 350 | | *v0.1* | *End-to-end core parse → render: =build= a single =.org= file to HTML* | *done* | | ||
| 351 | | *v0.2* | *Multi-file SITE build: INDEX + RESOLVE internal links, minijinja templates, =build <src-dir> <out-dir>=, tables + footnotes* | *done* | | ||
| 352 | | *v0.3* | *Incremental build layer: content/config/template hashing, dependency graph, per-page render keys, persisted cache manifest, invalidation* | *done* | | ||
| 353 | | *v0.4* | *MVP: the full v1 construct scope — heading metadata, nested/description lists, block types, timestamps, images, syntect highlighting — with the IN/OUT line under test* | *done* | | ||
| 354 | | *0* | *Corpus audit + =emacs --batch= ground-truth oracle* | *done* | | ||
| 355 | | 1 | Line lexer + heading/section skeleton | done | | ||
| 356 | | 2 | Block elements — lists, source blocks, tables, footnote defs, blocks by type, drawers | done | | ||
| 357 | | 3 | Inline objects — emphasis, links, bare URLs, footnote refs, timestamps | done | | ||
| 358 | | 4 | Rendering to HTML — tree walk, tables, footnote two-pass, minijinja templating, syntect highlighting | done | | ||
| 359 | | 5 | Link resolution + symbol table (INDEX + RESOLVE, used-target list, broken-link reporting) | done | | ||
| 360 | | 6 | Incremental build layer (hashing, dep graph, invalidation); =watch= on OS filesystem events | done | | ||
| 361 | | *7* | *Hardening: rayon parallelism, error locations in parse diagnostics* | *done* | | ||
| 362 | | *8* | *General use: config file, user templates, nav modes, =init= scaffold, safe discovery* | *done* | | ||
| 363 | | *9* | *Generated listing pages: =[[collections]]=, sorted indexes, feeds via XML templates* | *done* | | ||
| 364 | | *10* | *Grouped collections: one page per tag plus a tag index — full parity with the incumbent* | *done* | | ||
| 365 | | *11* | *Pagination: numbered pages with a =paginator= context, composing with grouping* | *done* | | ||
| 366 | | *12* | *=base_url=: =absolute=/=rfc822= filters, a valid RSS feed in the scaffold, canonical links* | *done* | | ||
| 367 | | *13* | *=watch= on OS filesystem events, debounced, with the feedback loop closed* | *done* | | ||
| 368 | | *14* | *Authoring: excerpts, word count, reading time, =truncate=, and draft pages* | *done* | | ||
| 369 | | *15* | *Table of contents, section numbers, and org's =#+OPTIONS:= per-file switches* | *done* | | ||
| 370 | | *16* | *=serve=: development server with long-poll live reload, loopback-bound* | *done* | | ||
| 371 | | *17* | *Bundled TOML and Org syntaxes, a user syntax directory, and org's comma escape* | *done* | | ||
| 372 | | *18* | *Per-page layouts: =[[pages]]= rules and =#+TEMPLATE:=* | *done* | | ||
| 373 | | *19* | *Export parity: relative heading levels, special strings, sub/superscript, caption numbering, checkbox and counter markup, table marker columns, special blocks* | *done* | | ||
| 374 | | *20* | *Correctness debt: org's entity table, table captions, a reported =#+INCLUDE:=, and an oracle that separates deliberate divergence from defects* | *done* | | ||
| 375 | | *21* | *Extra asset roots; per-template hashing so one layout edit does not re-render the site* | *done* | | ||
| 376 | | *22* | *Release engineering: CI on both platforms, a checked MSRV, release binaries, a changelog, and a written compatibility promise* | *done* | | ||
| 377 | |||
| 378 | *** v0.2 in / out | ||
| 379 | *Added in v0.2:* the INDEX stage (=SymbolTable= of =:ID:=/=:CUSTOM_ID:=/heading/=file:= | ||
| 380 | targets across a directory); the RESOLVE stage — rewrites =[[#custom-id]]=, =[[id:...]]=, | ||
| 381 | =[[*Heading]]= and =[[file:other.org]]= links to real relative output URLs, returns the | ||
| 382 | =used_targets= list (the =uses= edges, spec §4.3/R2) and reports unresolved links as | ||
| 383 | warnings rather than crashing; a minijinja base layout (title, nav, body) applied to every | ||
| 384 | page; a =build <src-dir> <out-dir>= path that walks the tree, parses + resolves + renders + | ||
| 385 | templates every =.org= into a linked static site and copies non-=.org= assets through; | ||
| 386 | plus two new constructs — pipe *tables* (with header band from the rule row) and | ||
| 387 | *footnotes* (block =[fn:1]= definitions, referenced =[fn:1]=, and inline =[fn:1:text]=, | ||
| 388 | rendered as a numbered, back-linked notes section). | ||
| 389 | |||
| 390 | *Left stubbed at v0.2, all closed in v0.4:* timestamps; TODO keywords and priorities; | ||
| 391 | generic (non-PROPERTIES) drawers; real syntect tokenizing behind the =Highlighter= trait. | ||
| 392 | |||
| 393 | *** v0.3 in / out | ||
| 394 | *Added in v0.3 — the incremental build layer (spec §4, the flagship, non-retrofittable | ||
| 395 | feature):* | ||
| 396 | |||
| 397 | - *Three hash classes (spec §4.1)* in =src/incremental.rs=: a *content hash* (blake3 | ||
| 398 | of a file's bytes), a *config hash* (blake3 of the resolved =BuildConfig=), and a | ||
| 399 | *template hash* (blake3 of the template sources). A change in any one invalidates the | ||
| 400 | pages it affects. | ||
| 401 | - *Dependency graph (spec §4.3)* built from RESOLVE's =defines=/=uses= edges: a page | ||
| 402 | depends on the targets it links to, so editing (or renaming a heading in) a file | ||
| 403 | invalidates the pages that /link into/ it, not just the file itself — the load-bearing | ||
| 404 | R2 invariant. On rebuild the graph is merged with the previous build's =defines= so a | ||
| 405 | /removed/ target still pulls in its linkers. | ||
| 406 | - *Per-page =render_key=* = =H(content ⊕ resolved-links ⊕ config ⊕ template)=. If a | ||
| 407 | page's render key is unchanged, its on-disk output is already correct and it is skipped. | ||
| 408 | The config component folds in a *site-structure hash* (every page's =(path, title)=), | ||
| 409 | because the shared nav bar is global chrome — a title change or a page add/remove alters | ||
| 410 | the nav on every page and so must re-render them all (otherwise byte-equivalence breaks). | ||
| 411 | - *Persisted cache manifest* (=<out>/.orgo-cache.json=, JSON), carrying per-page | ||
| 412 | records, the config/template hashes, and the serialized dependency graph, tagged with | ||
| 413 | =CACHE_FORMAT_VERSION=. A version mismatch, a missing file, or a corrupt file all fall | ||
| 414 | back to a clean full rebuild — the cache is an optimization, never a correctness | ||
| 415 | dependency. | ||
| 416 | - *Wired into =build_site=*: only pages whose render key changed (or that link into a | ||
| 417 | changed file's targets) are re-rendered; unchanged outputs are left in place. =--no-cache= | ||
| 418 | forces a full rebuild; =clean <out-dir>= removes the output directory (and its cache). | ||
| 419 | =SiteReport= now reports =rendered= vs =skipped= counts. | ||
| 420 | |||
| 421 | The hard gates are enforced by =tests/incremental.rs=: full-vs-incremental *byte | ||
| 422 | equivalence* (and a second unchanged build re-rendering *zero* pages); *edit-one-file* | ||
| 423 | re-renders exactly the changed page plus its linkers; *renamed-heading* invalidates the | ||
| 424 | linking page and updates its emitted anchor; and cache *version-bump / missing / corrupt* | ||
| 425 | all fall back to a full rebuild. | ||
| 426 | |||
| 427 | *Out of scope in v0.3:* real syntect highlighting; timestamps and TODO keywords (all | ||
| 428 | landed in v0.4). =watch= is a minimal mtime poll loop (=watch <src-dir> -o <out-dir>=), not | ||
| 429 | an OS file-watcher — the fs-notify integration is deferred. The parse-tree cache (spec §4.5, | ||
| 430 | "optionally") is not persisted: PARSE/INDEX/RESOLVE run for every file each build (cheap and | ||
| 431 | pure); the incremental win is on RENDER + EMIT. | ||
| 432 | |||
| 433 | *** v0.4 in / out — the MVP | ||
| 434 | v0.4 closes the gap between the v1 scope above and what the code actually did, so every | ||
| 435 | construct the IN list claims is now parsed, rendered, and pinned by a golden file: | ||
| 436 | |||
| 437 | - *Heading metadata* — TODO keywords (the Emacs default =TODO=/=DONE= set, matched on a | ||
| 438 | word boundary so =TODOs= is not one) and =[#A]= priority cookies, rendered with Emacs' | ||
| 439 | own export classes so the output stays diffable against an =emacs --batch= oracle. | ||
| 440 | - *Lists* — indentation-based nesting (a sub-list renders /inside/ its parent =<li>=), | ||
| 441 | multi-paragraph item bodies, and =term :: definition= description lists as =<dl>=. | ||
| 442 | - *Blocks by type* — =QUOTE=, =CENTER=, =EXAMPLE=, =EXPORT= and =SRC= are now distinct | ||
| 443 | elements rather than all collapsing to a verbatim example block. Block matching is on the | ||
| 444 | specific kind, so a source block can nest inside a quote. An =html= export block passes | ||
| 445 | through; every other backend drops. | ||
| 446 | - *Timestamps* — active =<...>= and inactive =[...]=, optional times, same-day time | ||
| 447 | ranges and =--=-joined date ranges, rendered as =<time>= with a machine-readable | ||
| 448 | =datetime=. Repeater/warning cookies are recognized and discarded. | ||
| 449 | - *Images* — a description-less link to an image file renders as =<img>=; with an | ||
| 450 | affiliated =#+CAPTION:=/=#+ATTR_HTML:= it is promoted to a =<figure>= with the caption as | ||
| 451 | both =<figcaption>= and alt text. Links to non-=.org= files are now understood as asset | ||
| 452 | links: neither resolved nor reported as broken. | ||
| 453 | - *Syntax highlighting* — real syntect tokenizing to CSS classes (never inline styles, so | ||
| 454 | themes live in the stylesheet). Every build emits the matching =syntax.css= and each page | ||
| 455 | links it relative to its own depth. An unknown language degrades to escaped =<pre><code>=. | ||
| 456 | - *Diagnostics* — broken links are reported as the org syntax the author wrote | ||
| 457 | (=warning: b.org: unresolved link [[#setup]]=) rather than a Debug-printed enum. | ||
| 458 | |||
| 459 | *The OUT line is now enforced, not just asserted.* =tests/constructs.rs= pins each | ||
| 460 | excluded construct to a specific degradation: babel is never executed /and/ a checked-in | ||
| 461 | =#+RESULTS:= block is dropped rather than published as if it were verified output; | ||
| 462 | =#+TBLFM:= is inert; =#+INCLUDE:= is never expanded and says so; LaTeX, macros and radio targets survive | ||
| 463 | as literal text; drawers other than PROPERTIES are captured and dropped; unmodelled block | ||
| 464 | types keep their content verbatim. | ||
| 465 | |||
| 466 | *Still out:* =#+TODO:= per-file keyword sequences; planning lines | ||
| 467 | (=SCHEDULED:=/=DEADLINE:=), which render as ordinary paragraphs; and fixed-width =:= | ||
| 468 | lines. | ||
| 469 | |||
| 470 | ** Serving | ||
| 471 | #+begin_src sh | ||
| 472 | cargo run -- serve my-site -o _site # http://127.0.0.1:3000 | ||
| 473 | #+end_src | ||
| 474 | |||
| 475 | Builds, watches, serves, and reloads the browser when a rebuild lands — the loop =watch= | ||
| 476 | leaves half-open. | ||
| 477 | |||
| 478 | - *Loopback by default.* A dev server serves unreviewed drafts off your laptop, so | ||
| 479 | reaching the local network is something you ask for with =--host 0.0.0.0=, never | ||
| 480 | something you get. | ||
| 481 | - *The reload script is injected on the way out*, never written to disk. What you | ||
| 482 | deploy is the built site, and it must not carry a dev server's JavaScript. | ||
| 483 | - *Long-polling, not WebSockets or SSE.* The browser asks "anything since generation | ||
| 484 | N?" and the server holds the request until there is. Instant like a push, no protocol | ||
| 485 | beyond ordinary HTTP, and no dependency. A streamed response would have been more | ||
| 486 | elegant and does not work: tiny_http buffers a response until its body ends, so a body | ||
| 487 | that never ends never reaches the client. | ||
| 488 | - A reload only follows a *successful* rebuild. Reloading onto a stale page because the | ||
| 489 | build just failed tells you nothing; the error is already on your terminal. | ||
| 490 | |||
| 491 | URL resolution is the server's security boundary and is written as a pure function with | ||
| 492 | its own tests: =..=, percent-encoded =..=, backslashes, absolute paths and embedded NULs | ||
| 493 | all resolve to nothing rather than to somewhere outside the output directory. | ||
| 494 | |||
| 495 | ** Watching | ||
| 496 | #+begin_src sh | ||
| 497 | cargo run -- watch my-site -o _site | ||
| 498 | #+end_src | ||
| 499 | |||
| 500 | Rebuilds on OS filesystem events rather than polling, so it costs nothing while nothing | ||
| 501 | happens. Write bursts are debounced — an editor saving a file writes a temp file, renames | ||
| 502 | it over the original and touches the directory, which is one edit and several events. | ||
| 503 | |||
| 504 | Two rules decide what counts as a change, and they are not the same rules the build uses | ||
| 505 | to find content: | ||
| 506 | |||
| 507 | - *A build input is a change.* Editing =orgo.toml= or a template rebuilds, even | ||
| 508 | though discovery skips both as non-content. The question is "would this change the | ||
| 509 | site?", not "is this a page?". | ||
| 510 | - *Our own output is not.* =watch . -o _site= puts the output inside the source, so a | ||
| 511 | rebuild's writes raise events that would trigger a rebuild, forever. Dot-directories go | ||
| 512 | the same way — =.git= churns on every command — as do editor scratch files, including | ||
| 513 | Emacs' =file.org~= backups, which do not start with a dot. | ||
| 514 | |||
| 515 | Where native watching is unavailable (some network and container filesystems), it falls | ||
| 516 | back to polling and says so, rather than failing. | ||
| 517 | |||
| 518 | ** Phase 0: the corpus audit and the Emacs oracle | ||
| 519 | The v1 scope was, by its own admission, /recommended/ — a guess about which slice of org | ||
| 520 | matters. Phase 0 replaces both halves of that guess with a measurement: an audit that asks | ||
| 521 | what a real corpus actually uses, and an oracle that asks whether we render it the way | ||
| 522 | Emacs does. | ||
| 523 | |||
| 524 | The audit runs against any corpus — point it at your own notes before trusting this tool | ||
| 525 | with them. The numbers below come from a 179-file site published today by weblorg, a | ||
| 526 | wrapper around org's own HTML exporter, which makes it both a realistic workload and a | ||
| 527 | directly comparable incumbent. With collections configured, orgo now reproduces | ||
| 528 | *all 182 of that site's URLs*. | ||
| 529 | |||
| 530 | #+begin_example | ||
| 531 | cargo run -- audit <src-dir> # what does this corpus use, and is it in scope? | ||
| 532 | cargo test --test oracle # how does our HTML differ from Emacs' own export? | ||
| 533 | #+end_example | ||
| 534 | |||
| 535 | *** What the audit found | ||
| 536 | *The scope guess was sound.* 99.9% of construct uses in the corpus are in scope. The | ||
| 537 | whole out-of-scope tail is 8 uses: four =#+TBLFM:= in a post /about/ org-mode, three | ||
| 538 | =\name= entities, and one =#+BEGIN_NOTE=. | ||
| 539 | |||
| 540 | *=#+SLUG:= was a hole big enough to sink the project.* 178 of 179 files set it, and the | ||
| 541 | published URL comes from it, not from the filename: =2018-11-28-aes-encryption.org= is | ||
| 542 | served at =blog/aes-encryption.html=. orgo derived output paths from source filenames, | ||
| 543 | so *169 of 179 pages would have been published at the wrong URL* — every inbound link and | ||
| 544 | every search result, broken, by a tool that reported a clean build. Output paths now come | ||
| 545 | from =#+SLUG:= when present ([[file:src/util.rs][=util::output_path=]]); slugs are sanitized so an | ||
| 546 | author-supplied =../../etc/x= cannot escape the output directory, and two pages claiming one | ||
| 547 | URL is a build error rather than a silently dropped page. Building the real corpus now | ||
| 548 | reproduces all 179 of the live site's URLs exactly. | ||
| 549 | |||
| 550 | *Some machinery is speculative.* The corpus contains no =id:=, =#custom-id= or =*Heading= | ||
| 551 | links at all — its cross-page links are hand-written relative URLs. The INDEX/RESOLVE | ||
| 552 | symbol table that v0.2 was built around is, against this corpus, unexercised. | ||
| 553 | |||
| 554 | *An audit can lie too.* The first run reported 23 uses of a custom TODO keyword sequence. | ||
| 555 | All 23 were false: the detector read the leading word of =* CSS Variables= as the keyword | ||
| 556 | =CSS=. The corpus defines no =#+TODO:= sequences at all, so the true count was zero. The | ||
| 557 | detector now matches conventional keyword names only — a tool that overstates a gap argues | ||
| 558 | for work nobody needs. | ||
| 559 | |||
| 560 | *** What the oracle found | ||
| 561 | =tests/oracle.rs= exports each fixture with org's own exporter via =emacs --batch=, reduces | ||
| 562 | both sides to a semantic skeleton (element opens, closes and text, with layout =div=s, | ||
| 563 | inline =span=s and all attributes but =href=/=src= dropped), and *snapshots the | ||
| 564 | disagreement*. Snapshotting rather than asserting is deliberate: a checked-in divergence | ||
| 565 | report gets reviewed and shows up as a diff, where a permanently red test gets ignored. | ||
| 566 | Three invariants are asserted outright, and all three hold — heading structure, list | ||
| 567 | nesting, and source-block text match Emacs exactly. | ||
| 568 | |||
| 569 | *No bugs in orgo.* Every remaining divergence is a deliberate choice to emit better | ||
| 570 | HTML than org does: | ||
| 571 | |||
| 572 | | | orgo | Emacs | why | | ||
| 573 | |—————--+—————————+—————————-+—————————————| | ||
| 574 | | emphasis | =<em>=/=<strong>= | =<i>=/=<b>= | semantic, not presentational | | ||
| 575 | | captioned image | =<figure>=/=<figcaption>= | =<p>= + ="Figure 1: …"= | real figure semantics | | ||
| 576 | | timestamp | =<time datetime="…">= | literal =<2024-01-15 Mon>= | machine-readable | | ||
| 577 | | footnotes | =<section><ol>= | =<h2>Footnotes:</h2>= | a list of notes is a list | | ||
| 578 | | heading anchor | slug of the text | =org1a2b3c4= | stable, and what the live site serves | | ||
| 579 | | code | =<pre><code>= | =<pre>= | the HTML5 idiom | | ||
| 580 | |||
| 581 | One genuine semantic difference: org treats a single blank line between a =1.= list and a | ||
| 582 | =-= list as /one/ list and keeps the first item's bullet type, while we start a second list. | ||
| 583 | We keep ours, on measurement rather than taste — the pattern occurs *zero* times in the | ||
| 584 | corpus, so matching an org quirk would buy nothing and cost the more obvious reading. | ||
| 585 | |||
| 586 | *The oracle's best catch was three bugs in itself.* Naive normalization reported code as | ||
| 587 | corrupted (it trimmed each of syntect's per-token text runs, turning =def greet= into | ||
| 588 | =defgreet=) and reported blocks at 36% agreement (syntect's spans flooded the diff). Both | ||
| 589 | were measurement artifacts. A differential harness is a piece of software like any other, | ||
| 590 | and the first divergences it reports are usually its own. | ||
| 591 | |||
| 592 | ** Phase 7: hardening | ||
| 593 | *** Parse diagnostics (=file:line: message=) | ||
| 594 | The parser's contract is that it always returns a document — out-of-scope and malformed | ||
| 595 | constructs degrade rather than crash. The gap was that they degraded /silently/, and in the | ||
| 596 | worst cases the degradation is severe: an unterminated =#+BEGIN_SRC= reads the rest of the | ||
| 597 | file as block content, and an unterminated drawer does the same but renders to nothing, so | ||
| 598 | one missing line deletes most of a page from a build that reports success. | ||
| 599 | |||
| 600 | =parse= now returns =Document::diagnostics=, each carrying a 1-based source line, and the | ||
| 601 | build prints them as =file:line: message=. =--strict= turns them (and unresolved links) into | ||
| 602 | a non-zero exit. Line numbers are threaded as an absolute offset through every nested parse, | ||
| 603 | so a block inside a list item inside a section still reports its real file line — there is a | ||
| 604 | test for exactly that, because reconstructed and re-indented nested slices are precisely | ||
| 605 | where an off-by-N hides. The 179-file corpus produces zero diagnostics. | ||
| 606 | |||
| 607 | *** Parallelism | ||
| 608 | PARSE, RESOLVE and RENDER/EMIT run under rayon. PARSE is a pure function of one file's bytes | ||
| 609 | and RESOLVE only reads the shared symbol table, which is what makes both safe to parallelize | ||
| 610 | at all; INDEX stays sequential. | ||
| 611 | |||
| 612 | | corpus | before | after | speedup | | ||
| 613 | |————————+——--+——-+———| | ||
| 614 | | 179 files (real) | 0.23s | 0.07s | 3.3× | | ||
| 615 | | 1,790 files (10× copy) | 3.98s | 0.82s | 4.9× | | ||
| 616 | |||
| 617 | Measured on 12 cores. =RAYON_NUM_THREADS=1= reproduces the old 3.98s exactly, so the gain is | ||
| 618 | parallelism rather than incidental change, and the output is byte-identical to the sequential | ||
| 619 | build across the whole corpus. | ||
| 620 | |||
| 621 | *Parallelism must not be observable in the result.* =par_iter().collect()= preserves input | ||
| 622 | order, so the emitted bytes are unaffected — but the build /report/ is the fragile half: | ||
| 623 | pushing to =rendered=/=skipped= from inside the parallel pass would order them by thread | ||
| 624 | scheduling, producing a non-deterministic report over a deterministic site. The parallel pass | ||
| 625 | therefore returns only what was written, and the report is assembled sequentially afterwards. | ||
| 626 | =parallel_builds_are_deterministic_in_output_and_report_order= holds that line, and it was | ||
| 627 | verified by reintroducing the bug and watching it fail. | ||
| 628 | |||
| 629 | *** The real scaling limit was not the CPU | ||
| 630 | Going 10× on corpus size cost 17× in time, which parallelism improves without fixing: the | ||
| 631 | cause was the nav bar listing *every* page, so an /n/-page site emitted /n/² nav links. At | ||
| 632 | 1,790 pages each page carried 1,799 links and the output was 284 MB, against 5.5 MB for the | ||
| 633 | 179-page corpus — 52× the bytes for 10× the input. | ||
| 634 | |||
| 635 | The nav is now built from *top-level pages only* ([[file:src/site.rs][=is_top_level=]]): a nav is a | ||
| 636 | map of the site's top level, not an index of its contents, and section pages reach their | ||
| 637 | siblings through that section's landing page. Nav size becomes a function of the top level | ||
| 638 | rather than of the corpus, and the quadratic disappears. | ||
| 639 | |||
| 640 | | 1,790-page corpus (6 top-level pages) | before | after | | ||
| 641 | |—————————————+——--+——-| | ||
| 642 | | full build | 0.82s | 0.39s | | ||
| 643 | | total output | 284 MB | 34 MB | | ||
| 644 | | nav links per page | 1,799 | 6 | | ||
| 645 | |||
| 646 | Scaling is now linear: 179 pages in 0.07s and 1,796 in 0.39s, where the small case is mostly | ||
| 647 | the fixed cost of loading syntect's syntax definitions. | ||
| 648 | |||
| 649 | The same rule sharpened the incremental build, which is the larger win. The site-structure | ||
| 650 | hash — the thing that forces a global re-render — now covers only the pages that appear in | ||
| 651 | the nav, because those are the only ones whose title or URL affects another page. *Adding a | ||
| 652 | blog post used to re-render the entire site; now it renders one page.* A top-level page's | ||
| 653 | title still invalidates everything, correctly, since every page displays it. | ||
| 654 | |||
| 655 | *Trade-off worth knowing:* on a site whose sections live in subdirectories, only genuinely | ||
| 656 | root-level pages appear — a site keeping its landing pages at =salary/index.org= and friends | ||
| 657 | gets a one-entry nav. That is what =nav.mode = "explicit"= is for: list the pages you want, | ||
| 658 | in the order you want them. | ||
| 659 | |||
| 660 | *From v0.1 (core subset):* headings with nesting and anchors (every heading is now | ||
| 661 | anchored — =:CUSTOM_ID:=/=:ID:= else a slug of its text) and trailing tags; paragraphs; | ||
| 662 | plain lists (unordered + ordered) with checkboxes; source blocks; inline markup (=*bold*=, | ||
| 663 | =/italic/=, =_underline_=, =+strike+=, ==verbatim==, =~code~=); links and bare URLs. | ||
| 664 | |||
| 665 | ** Compatibility | ||
| 666 | Versions mean something as of 1.0. The *stable surface* — changing incompatibly requires | ||
| 667 | a major version — is what you actually build a site against: | ||
| 668 | |||
| 669 | | Stable | Detail | | ||
| 670 | |——————+——————————————————————————————————————————————-| | ||
| 671 | | =orgo.toml= keys | Names, types and meaning. New keys are minor releases; removing one is major. | | ||
| 672 | | Template context | =page=, =site=, =nav=, =root=, =pages=, =group=, =groups=, =paginator=, =stylesheet=, and the =absolute= / =rfc822= / =truncate= filters. | | ||
| 673 | | CLI | Command names, flags, and exit codes. | | ||
| 674 | | URLs | How a source path becomes an output path, including =#+SLUG:=. A generator that moves your URLs breaks every link anyone has to you. | | ||
| 675 | |||
| 676 | Explicitly *not stable*, so that the above can be: | ||
| 677 | |||
| 678 | - *The incremental cache.* Versioned, discarded on mismatch, never a correctness | ||
| 679 | dependency. It changes whenever it needs to, in any release. | ||
| 680 | - *Rendered HTML details.* orgo tracks what Emacs exports from the same file, and | ||
| 681 | closing a gap changes markup. Changes that affect output are called out in | ||
| 682 | [[file:CHANGELOG.org][CHANGELOG.org]] — the class names the documentation names (=post-list=, | ||
| 683 | =figure-number=, =section-number-N=, =footnote-ref=) are the ones to write CSS against. | ||
| 684 | - *The Rust API.* The crate is published so the binary can be installed with | ||
| 685 | =cargo install=; the library exists to serve it, and its types move as the tool does. | ||
| 686 | |||
| 687 | The *MSRV is 1.88*, checked in CI on every change. orgo's own code compiles on | ||
| 688 | 1.82; the floor comes from dependencies. Raising it is a minor version, never a patch. | ||
| 689 | |||
| 690 | ** Dependencies | ||
| 691 | Parser is hand-written recursive descent (not =nom=/=chumsky=/=pest= — org is | ||
| 692 | line-oriented and context-sensitive, not clean CFG). Key crates: =syntect= (syntax | ||
| 693 | highlighting, behind a =Highlighter= trait so tree-sitter can be swapped in later), | ||
| 694 | =minijinja= (runtime templates), =blake3= (content/cache hashing), =rayon= (parallel | ||
| 695 | PARSE/RESOLVE/RENDER), =notify= (filesystem events for =watch=), =tiny_http= (the =serve= | ||
| 696 | development server), =toml= (config), =chrono=, =camino=, =walkdir=, =clap=, =anyhow=/=thiserror=. | ||
| 697 | =insta= for snapshot tests, and =emacs --batch= — optional, and only for the oracle. | ||
| 698 | |||
| 699 | ** Build & test | ||
| 700 | #+begin_example | ||
| 701 | cargo build | ||
| 702 | cargo test # 191 tests | ||
| 703 | cargo run -- init my-site # scaffold a new site | ||
| 704 | cargo run -- build fixtures/minimal.org -o minimal.html # single file | ||
| 705 | cargo run -- build fixtures/site -o _site # whole site (incremental) | ||
| 706 | cargo run -- audit fixtures/site # corpus audit (Phase 0) | ||
| 707 | cargo run -- build fixtures/site -o _site --no-cache # force a full rebuild | ||
| 708 | cargo run -- watch fixtures/site -o _site # rebuild on filesystem events | ||
| 709 | cargo run -- serve fixtures/site -o _site # ... and serve with live reload | ||
| 710 | cargo run -- clean _site # remove output + cache | ||
| 711 | #+end_example | ||
| 712 | |||
| 713 | A second =build= of an unchanged site re-renders nothing; editing a page re-renders only | ||
| 714 | that page and the pages that link into it (watch the =rendered=/=cached= counts). | ||
| 715 | |||
| 716 | A build emits =syntax.css= next to its output (the highlighter emits CSS classes, so the | ||
| 717 | stylesheet has to come with them) and every page links it. | ||
| 718 | |||
| 719 | =fixtures/= holds tiny =.org= samples: the core ones (=minimal.org=, =core.org=, | ||
| 720 | =elements.org=, =table.org=, =footnote.org=), one per v1 construct group (=headings.org=, | ||
| 721 | =lists.org=, =blocks.org=, =timestamps.org=, =images.org=), the scope guardrail | ||
| 722 | (=outofscope.org=), and a linked multi-file site under =fixtures/site/= (=index.org=, | ||
| 723 | =guide.org=, =about.org= + a =style.css= asset). The real corpus (golden files derived from | ||
| 724 | actual documents) lands in Phase 0. =cargo test= runs =insta= snapshots of the element tree | ||
| 725 | and rendered HTML for each fixture, the two templated site pages (proving cross-file link | ||
| 726 | resolution), and the incremental gates. | ||
RELEASING.md → RELEASING.org renamed +31 −35
| @@ -1,79 +1,75 @@ | |||
| 1 | # Releasing | 1 | * Releasing |
| 2 | 2 | A release is three things that must agree: a version in =Cargo.toml=, a git tag, and a | |
| 3 | A release is three things that must agree: a version in `Cargo.toml`, a git tag, and a | ||
| 4 | changelog entry. The release workflow checks the first two against each other and refuses | 3 | changelog entry. The release workflow checks the first two against each other and refuses |
| 5 | to build if they differ, because a release tagged `v0.18.0` containing a binary that | 4 | to build if they differ, because a release tagged =v0.18.0= containing a binary that |
| 6 | reports `0.17.0` is the kind of mistake nobody notices for months. | 5 | reports =0.17.0= is the kind of mistake nobody notices for months. |
| 7 | |||
| 8 | ## Before the first publish | ||
| 9 | 6 | ||
| 10 | ```bash | 7 | ** Before the first publish |
| 8 | #+begin_src sh | ||
| 11 | cargo login # a crates.io token, once per machine | 9 | cargo login # a crates.io token, once per machine |
| 12 | cargo publish --dry-run | 10 | cargo publish --dry-run |
| 13 | ``` | 11 | #+end_src |
| 14 | 12 | ||
| 15 | `repository` and `homepage` in `Cargo.toml` point at GitHub and at the documentation site | 13 | =repository= and =homepage= in =Cargo.toml= point at GitHub and at the documentation site |
| 16 | on Pages. If git.krz.sh becomes the primary remote, `repository` should follow it — | 14 | on Pages. If git.krz.sh becomes the primary remote, =repository= should follow it — |
| 17 | crates.io shows that link on the crate page, and it should lead somewhere you read. | 15 | crates.io shows that link on the crate page, and it should lead somewhere you read. |
| 18 | 16 | ||
| 19 | ## Every release | 17 | ** Every release |
| 20 | 18 | 1. *Write the changelog entry first.* [[file:CHANGELOG.org][CHANGELOG.org]] names behaviour, not | |
| 21 | 1. **Write the changelog entry first.** [CHANGELOG.md](CHANGELOG.md) names behaviour, not | ||
| 22 | commits — someone reading it wants to know what their next build will do differently. | 19 | commits — someone reading it wants to know what their next build will do differently. |
| 23 | Anything that changes rendered HTML gets said out loud. | 20 | Anything that changes rendered HTML gets said out loud. |
| 24 | 21 | ||
| 25 | 2. **Bump the version** in `Cargo.toml`, and build once so `Cargo.lock` follows. | 22 | 2. *Bump the version* in =Cargo.toml=, and build once so =Cargo.lock= follows. |
| 26 | 23 | ||
| 27 | Patch for fixes that change nothing about the stable surface. Minor for new config | 24 | Patch for fixes that change nothing about the stable surface. Minor for new config |
| 28 | keys, new template variables, an MSRV bump, or output that changes to track Emacs more | 25 | keys, new template variables, an MSRV bump, or output that changes to track Emacs more |
| 29 | closely. Major for anything that breaks the promises in the README's Compatibility | 26 | closely. Major for anything that breaks the promises in the README's Compatibility |
| 30 | section — config keys, template context, CLI, or URLs. | 27 | section — config keys, template context, CLI, or URLs. |
| 31 | 28 | ||
| 32 | 3. **Check it.** | 29 | 3. *Check it.* |
| 33 | 30 | ||
| 34 | ```bash | 31 | #+begin_src sh |
| 35 | cargo test | 32 | cargo test |
| 36 | cargo clippy --all-targets -- -D warnings | 33 | cargo clippy --all-targets -- -D warnings |
| 37 | cargo run -- build docs -o docs/_site --strict | 34 | cargo run -- build docs -o docs/_site --strict |
| 38 | cargo package | 35 | cargo package |
| 39 | ``` | 36 | #+end_src |
| 40 | 37 | ||
| 41 | `cargo package` is the one people forget: it builds the crate exactly as crates.io will | 38 | =cargo package= is the one people forget: it builds the crate exactly as crates.io will |
| 42 | receive it, and catches a file the `exclude` list should not have removed. | 39 | receive it, and catches a file the =exclude= list should not have removed. |
| 43 | 40 | ||
| 44 | 4. **Verify against a real corpus.** The test suite says the code does what it did; a | 41 | 4. *Verify against a real corpus.* The test suite says the code does what it did; a |
| 45 | corpus says the *site* does. Build a site you know with `--no-cache` and diff the | 42 | corpus says the /site/ does. Build a site you know with =--no-cache= and diff the |
| 46 | output against the previous version's. A release that quietly changes 200 pages should | 43 | output against the previous version's. A release that quietly changes 200 pages should |
| 47 | do so on purpose. | 44 | do so on purpose. |
| 48 | 45 | ||
| 49 | 5. **Commit, tag, push.** | 46 | 5. *Commit, tag, push.* |
| 50 | 47 | ||
| 51 | ```bash | 48 | #+begin_src sh |
| 52 | git commit -am "0.18: <what changed>" | 49 | git commit -am "0.18: <what changed>" |
| 53 | git tag -a v0.18.0 -m "0.18.0" | 50 | git tag -a v0.18.0 -m "0.18.0" |
| 54 | git push && git push --tags | 51 | git push && git push --tags |
| 55 | ``` | 52 | #+end_src |
| 56 | 53 | ||
| 57 | 6. **Publish the crate.** | 54 | 6. *Publish the crate.* |
| 58 | 55 | ||
| 59 | ```bash | 56 | #+begin_src sh |
| 60 | cargo publish | 57 | cargo publish |
| 61 | ``` | 58 | #+end_src |
| 62 | 59 | ||
| 63 | This is irreversible: a published version can be yanked but never replaced. | 60 | This is irreversible: a published version can be yanked but never replaced. |
| 64 | 61 | ||
| 65 | 7. **Finish the GitHub release.** Pushing the tag builds binaries for macOS (arm64 and | 62 | 7. *Finish the GitHub release.* Pushing the tag builds binaries for macOS (arm64 and |
| 66 | x86_64) and Linux (gnu and musl) and opens a *draft* release with them attached. Paste | 63 | x86_64) and Linux (gnu and musl) and opens a /draft/ release with them attached. Paste |
| 67 | the changelog entry in and publish it. The draft is deliberate — a release that | 64 | the changelog entry in and publish it. The draft is deliberate — a release that |
| 68 | publishes itself before anyone has read it cannot be edited quietly. | 65 | publishes itself before anyone has read it cannot be edited quietly. |
| 69 | 66 | ||
| 70 | ## If a release goes wrong | 67 | ** If a release goes wrong |
| 71 | |||
| 72 | Yank rather than delete, and ship a fix as a new version: | 68 | Yank rather than delete, and ship a fix as a new version: |
| 73 | 69 | ||
| 74 | ```bash | 70 | #+begin_src sh |
| 75 | cargo yank --version 0.18.0 | 71 | cargo yank --version 0.18.0 |
| 76 | ``` | 72 | #+end_src |
| 77 | 73 | ||
| 78 | Yanking stops new dependents from selecting it; anyone who already has it keeps working. | 74 | Yanking stops new dependents from selecting it; anyone who already has it keeps working. |
| 79 | Then release `0.18.1` with the fix and a changelog entry that says what happened. | 75 | Then release =0.18.1= with the fix and a changelog entry that says what happened. |
SECURITY.md → SECURITY.org renamed +12 −16
| @@ -1,32 +1,28 @@ | |||
| 1 | # Security Policy | 1 | * Security Policy |
| 2 | 2 | ** Supported Versions | |
| 3 | ## Supported Versions | 3 | | Version | Supported | |
| 4 | 4 | |—————-+———--| | |
| 5 | | Version | Supported | | 5 | | latest release | yes | |
| 6 | |---------|-----------| | 6 | | anything older | no | |
| 7 | | latest release | yes | | ||
| 8 | | anything older | no | | ||
| 9 | 7 | ||
| 10 | Fixes land in a new release rather than as patches to an old one. | 8 | Fixes land in a new release rather than as patches to an old one. |
| 11 | 9 | ||
| 12 | ## Reporting | 10 | ** Reporting |
| 13 | 11 | Email [[mailto:hello@cleberg.net][hello@cleberg.net]], or open a private advisory through GitHub's /Security/ tab. | |
| 14 | Email <hello@cleberg.net>, or open a private advisory through GitHub's *Security* tab. | ||
| 15 | Please do not open a public issue for something exploitable. | 12 | Please do not open a public issue for something exploitable. |
| 16 | 13 | ||
| 17 | ## What is worth reporting | 14 | ** What is worth reporting |
| 18 | 15 | orgo reads org files and writes HTML, so the interesting cases are about what a /document/ | |
| 19 | orgo reads org files and writes HTML, so the interesting cases are about what a *document* | ||
| 20 | can make it do: | 16 | can make it do: |
| 21 | 17 | ||
| 22 | - Content from a source file escaping into HTML unescaped — a page that can inject script | 18 | - Content from a source file escaping into HTML unescaped — a page that can inject script |
| 23 | into the site it is published on. | 19 | into the site it is published on. |
| 24 | - A path in a document or config that writes outside the output directory. | 20 | - A path in a document or config that writes outside the output directory. |
| 25 | - The `serve` development server reachable, or made reachable, beyond loopback, or serving | 21 | - The =serve= development server reachable, or made reachable, beyond loopback, or serving |
| 26 | files from outside the output directory. | 22 | files from outside the output directory. |
| 27 | - A crash, hang or unbounded allocation triggered by a crafted org file. A build that | 23 | - A crash, hang or unbounded allocation triggered by a crafted org file. A build that |
| 28 | refuses a file is fine; one that never finishes is not. | 24 | refuses a file is fine; one that never finishes is not. |
| 29 | 25 | ||
| 30 | Out of scope: `--strict` not catching something, an unhandled org construct rendering | 26 | Out of scope: =--strict= not catching something, an unhandled org construct rendering |
| 31 | oddly, and anything requiring you to run orgo against files you already do not trust while | 27 | oddly, and anything requiring you to run orgo against files you already do not trust while |
| 32 | also deploying the result unread. | 28 | also deploying the result unread. |