Commit a52aee8095
Unsigned
Layout: unified · split
.github/workflows/release.yml +1 −1
| @@ -77,7 +77,7 @@ jobs: | ||
| 77 | 77 | staging="orgo-${{ github.event.inputs.tag || github.ref_name }}-${{ matrix.target }}" |
| 78 | 78 | mkdir "$staging" |
| 79 | 79 | cp "target/${{ matrix.target }}/release/orgo" "$staging/" |
| 80 | cp README.md LICENSE CHANGELOG.md "$staging/" | |
| 80 | cp README.org LICENSE CHANGELOG.org "$staging/" | |
| 81 | 81 | tar czf "$staging.tar.gz" "$staging" |
| 82 | 82 | shasum -a 256 "$staging.tar.gz" > "$staging.tar.gz.sha256" |
| 83 | 83 | |
CHANGELOG.md deleted −177
| @@ -1,177 +0,0 @@ | ||
| 1 | # Changelog | |
| 2 | ||
| 3 | What changed and why, newest first. Entries name the *behaviour* that moved, since that is | |
| 4 | what a rebuild will show you. | |
| 5 | ||
| 6 | Two conventions worth knowing before reading: | |
| 7 | ||
| 8 | - **A cache-format bump is not a change you need to act on.** The incremental cache is | |
| 9 | versioned and discards itself; a bump means the next build re-renders everything once. | |
| 10 | - **Output changes are called out.** orgo aims at what Emacs exports from the same | |
| 11 | file, so an entry that says "now renders X" means your pages will change. That is the | |
| 12 | product, not a regression — but it belongs in a changelog rather than a diff you find | |
| 13 | later. | |
| 14 | ||
| 15 | Versions follow the compatibility promise in the README: config keys, template variables, | |
| 16 | CLI flags and URLs are the stable surface. | |
| 17 | ||
| 18 | ## 0.19.1 | |
| 19 | ||
| 20 | - Footnote back-links carry `aria-label="Back to reference N"`, and the notes section is | |
| 21 | labelled. A link whose only visible content is `↩` has that glyph as its whole | |
| 22 | accessible name, so a screen reader announced "left arrow with hook" once per note with | |
| 23 | no way to tell them apart. | |
| 24 | ||
| 25 | ## 0.19.0 | |
| 26 | ||
| 27 | - **Full-content collections.** `include_content = true` gives a listing template each | |
| 28 | entry's rendered HTML as `entry.content` — a feed that carries whole posts rather than | |
| 29 | excerpts. Rendered only when the listing is actually rebuilt, so a cached feed costs | |
| 30 | nothing. | |
| 31 | - **Fixed: a listing could show a stale excerpt.** Its cache key covered a hand-picked set | |
| 32 | of fields, and the excerpt was not among them, so rewriting a post's first paragraph | |
| 33 | left the old text on the index until something unrelated invalidated it. Entries are now | |
| 34 | hashed through their serialization, which cannot drift from what a template can read. | |
| 35 | Editing a post's body now rebuilds the listings that show it. | |
| 36 | - `page.toc` entries carry `number`, so a site with section numbering on can number its | |
| 37 | contents list to match its headings. | |
| 38 | ||
| 39 | ## 0.18.0 | |
| 40 | ||
| 41 | Release engineering, so that a version number is worth reading. | |
| 42 | ||
| 43 | - **A written compatibility promise.** Config keys, template variables, CLI flags and URLs | |
| 44 | are the stable surface; the incremental cache, HTML details and the Rust API are not. | |
| 45 | In the README, and in the guide under *Versioning and upgrades*. | |
| 46 | - **CI** on Linux and macOS: build, test, clippy as an error, and the documentation site | |
| 47 | built with `--strict`. Emacs is installed on both, so the oracle suite runs for real | |
| 48 | instead of skipping. | |
| 49 | - **A checked MSRV**, 1.88 — which is how it came to be 1.88 rather than the 1.82 | |
| 50 | orgo's own code needs. The floor comes from dependencies, and nobody finds that out | |
| 51 | by reasoning about it. | |
| 52 | - **Release binaries** for macOS (arm64, x86_64) and Linux (gnu, musl), built on tag into | |
| 53 | a draft release. The tag is checked against `Cargo.toml` before anything is built. | |
| 54 | - A `LICENSE` file to go with the MIT declaration, crates.io metadata, and a release | |
| 55 | profile that produces a 5.0 MB binary rather than 6.5 MB. | |
| 56 | - This changelog, and `RELEASING.md`. | |
| 57 | ||
| 58 | ## 0.17.0 | |
| 59 | ||
| 60 | - **Asset directories outside the source.** `[build] assets = ["../theme/static"]` copies | |
| 61 | a directory's contents to the site root. A site's static files do not always live where | |
| 62 | its writing does, and copying them next to the writing is how a repository ends up with | |
| 63 | two of every stylesheet. `watch` and `serve` watch these directories too. Two files | |
| 64 | claiming one URL is a build error naming both. | |
| 65 | - **Template hashing is per template.** A page's render key covered every template, so | |
| 66 | editing a feed template re-rendered the whole site. It now covers the layout the page | |
| 67 | uses plus what that layout extends, includes or imports. On a 196-page site, editing the | |
| 68 | feed template renders one page instead of 196. | |
| 69 | - Cache format 7. | |
| 70 | ||
| 71 | ## 0.16.0 | |
| 72 | ||
| 73 | - **Org's entity table.** `\alpha`, `\rarr`, `20\deg` and the other 412 names, generated | |
| 74 | from Emacs' own `org-entities`. An unknown name stays literal; `#+OPTIONS: e:nil` turns | |
| 75 | the table off. *Output changes* for any page using entities. | |
| 76 | - **Table captions.** `#+CAPTION:` above a table becomes a numbered `<caption>`. | |
| 77 | - **`#+INCLUDE:` reports itself.** It was inert and silent, which publishes a page with | |
| 78 | content missing and nobody told. Now a diagnostic, and `--strict` makes it a failure. | |
| 79 | - The Emacs oracle separates deliberate divergence from defects. Every difference from | |
| 80 | org's exporter is named and justified, and a test asserts there are no others. | |
| 81 | ||
| 82 | ## 0.15.0 | |
| 83 | ||
| 84 | Export parity, from a page-by-page diff of a 179-file corpus against the site Emacs | |
| 85 | publishes from the same sources. **All of these change output.** | |
| 86 | ||
| 87 | - Heading levels are relative to a document's shallowest heading, as org exports them. | |
| 88 | - Org's text conversions: `--`, `---`, `...`, and `x^2` / `a_{b}`. Never inside verbatim, | |
| 89 | code, source blocks or LaTeX. `#+OPTIONS: -:nil`, `^:nil` and `^:{}` all work. | |
| 90 | - Captioned figures are numbered `Figure N:`. | |
| 91 | - A caption attaches to the element *directly* below it; a blank line between attaches to | |
| 92 | nothing. | |
| 93 | - Checkboxes render as org writes them, which keeps the `[-]` partly-done state a disabled | |
| 94 | `<input>` could not express. `[@4]` sets a list item's number. | |
| 95 | - A table's special marker column and its marker rows stay out of the output. | |
| 96 | - `#+BEGIN_NOTE` and any other unrecognised name is a special block: a div holding parsed | |
| 97 | org rather than a `<pre>` of literal text. Verse keeps its line breaks. | |
| 98 | - Emphasis borders forbid whitespace and nothing else, so `="proxied":false=` is verbatim | |
| 99 | and `~~/.config/doom/config.el~` is a path that starts with a tilde. | |
| 100 | - Listings sort on the time of day when a timestamp carries one. | |
| 101 | - Cache format 6. | |
| 102 | ||
| 103 | ## 0.14.0 | |
| 104 | ||
| 105 | - **Per-page layouts.** `[[pages]]` rules map a source path to a template, and | |
| 106 | `#+TEMPLATE:` on a page overrides any rule. A missing template fails the build naming | |
| 107 | the page, the template, and what does exist. | |
| 108 | - `page.year`, for grouping a listing by year with minijinja's `groupby`. | |
| 109 | - An explicit nav can order generated pages among authored ones. `nav.mode = "none"` now | |
| 110 | really means none. | |
| 111 | ||
| 112 | ## 0.13.0 | |
| 113 | ||
| 114 | - Bundled TOML and Org syntax definitions, a `syntaxes_dir` for your own, and org's comma | |
| 115 | escape (`,* heading` inside a block). | |
| 116 | ||
| 117 | ## 0.12.0 | |
| 118 | ||
| 119 | - `serve`: a development server with live reload, bound to loopback. | |
| 120 | - A documentation site under `docs/`, built by orgo itself. | |
| 121 | ||
| 122 | ## 0.11.0 | |
| 123 | ||
| 124 | - Table of contents as `page.toc`, section numbers, and org's `#+OPTIONS:` per-file | |
| 125 | switches. | |
| 126 | ||
| 127 | ## 0.10.0 | |
| 128 | ||
| 129 | - Excerpts, word count, reading time, a `truncate` filter, and `#+DRAFT:` pages. | |
| 130 | ||
| 131 | ## 0.9.0 | |
| 132 | ||
| 133 | - `watch`: rebuilds on OS filesystem events, debounced. | |
| 134 | ||
| 135 | ## 0.8.0 | |
| 136 | ||
| 137 | - `site.base_url`, the `absolute` and `rfc822` filters, canonical links, and an RSS feed | |
| 138 | in the scaffold that validates. | |
| 139 | ||
| 140 | ## 0.7.0 | |
| 141 | ||
| 142 | - Pagination for large listings, with a `paginator` template context that composes with | |
| 143 | grouping. | |
| 144 | ||
| 145 | ## 0.6.0 | |
| 146 | ||
| 147 | - Grouped collections: one page per tag plus a tag index. | |
| 148 | - Generated listing pages (`[[collections]]`), sorted indexes, and feeds via XML | |
| 149 | templates. | |
| 150 | - A config file, user templates, nav modes, an `init` scaffold, and discovery that will | |
| 151 | not publish `.git`. | |
| 152 | ||
| 153 | ## 0.5.0 | |
| 154 | ||
| 155 | - Parse diagnostics carry `file:line`, and pages render in parallel. | |
| 156 | - The corpus audit (`orgo audit`) and the `emacs --batch` oracle. | |
| 157 | - `#+SLUG:` decides a page's output filename — found by auditing a real corpus, where it | |
| 158 | affected 169 of 182 URLs. | |
| 159 | ||
| 160 | ## 0.4.0 | |
| 161 | ||
| 162 | - The full v1 construct scope, with the IN/OUT line under test. | |
| 163 | ||
| 164 | ## 0.3.0 | |
| 165 | ||
| 166 | - The incremental build layer: content, config and template hashing, a dependency graph, | |
| 167 | per-page render keys, and a persisted cache manifest. A full build and an incremental | |
| 168 | build produce byte-identical output. | |
| 169 | ||
| 170 | ## 0.2.0 | |
| 171 | ||
| 172 | - Multi-file site builds: a symbol table, internal link resolution, minijinja templates, | |
| 173 | tables and footnotes. | |
| 174 | ||
| 175 | ## 0.1.0 | |
| 176 | ||
| 177 | - Parse and render a single `.org` file to HTML. | |
CHANGELOG.org added +156
| @@ -0,0 +1,156 @@ | ||
| 1 | * Changelog | |
| 2 | What changed and why, newest first. Entries name the /behaviour/ that moved, since that is | |
| 3 | what a rebuild will show you. | |
| 4 | ||
| 5 | Two conventions worth knowing before reading: | |
| 6 | ||
| 7 | - *A cache-format bump is not a change you need to act on.* The incremental cache is | |
| 8 | versioned and discards itself; a bump means the next build re-renders everything once. | |
| 9 | - *Output changes are called out.* orgo aims at what Emacs exports from the same | |
| 10 | file, so an entry that says "now renders X" means your pages will change. That is the | |
| 11 | product, not a regression — but it belongs in a changelog rather than a diff you find | |
| 12 | later. | |
| 13 | ||
| 14 | Versions follow the compatibility promise in the README: config keys, template variables, | |
| 15 | CLI flags and URLs are the stable surface. | |
| 16 | ||
| 17 | ** 0.19.1 | |
| 18 | - Footnote back-links carry =aria-label="Back to reference N"=, and the notes section is | |
| 19 | labelled. A link whose only visible content is =↩= has that glyph as its whole | |
| 20 | accessible name, so a screen reader announced "left arrow with hook" once per note with | |
| 21 | no way to tell them apart. | |
| 22 | ||
| 23 | ** 0.19.0 | |
| 24 | - *Full-content collections.* =include_content = true= gives a listing template each | |
| 25 | entry's rendered HTML as =entry.content= — a feed that carries whole posts rather than | |
| 26 | excerpts. Rendered only when the listing is actually rebuilt, so a cached feed costs | |
| 27 | nothing. | |
| 28 | - *Fixed: a listing could show a stale excerpt.* Its cache key covered a hand-picked set | |
| 29 | of fields, and the excerpt was not among them, so rewriting a post's first paragraph | |
| 30 | left the old text on the index until something unrelated invalidated it. Entries are now | |
| 31 | hashed through their serialization, which cannot drift from what a template can read. | |
| 32 | Editing a post's body now rebuilds the listings that show it. | |
| 33 | - =page.toc= entries carry =number=, so a site with section numbering on can number its | |
| 34 | contents list to match its headings. | |
| 35 | ||
| 36 | ** 0.18.0 | |
| 37 | Release engineering, so that a version number is worth reading. | |
| 38 | ||
| 39 | - *A written compatibility promise.* Config keys, template variables, CLI flags and URLs | |
| 40 | are the stable surface; the incremental cache, HTML details and the Rust API are not. | |
| 41 | In the README, and in the guide under /Versioning and upgrades/. | |
| 42 | - *CI* on Linux and macOS: build, test, clippy as an error, and the documentation site | |
| 43 | built with =--strict=. Emacs is installed on both, so the oracle suite runs for real | |
| 44 | instead of skipping. | |
| 45 | - *A checked MSRV*, 1.88 — which is how it came to be 1.88 rather than the 1.82 | |
| 46 | orgo's own code needs. The floor comes from dependencies, and nobody finds that out | |
| 47 | by reasoning about it. | |
| 48 | - *Release binaries* for macOS (arm64, x86_64) and Linux (gnu, musl), built on tag into | |
| 49 | a draft release. The tag is checked against =Cargo.toml= before anything is built. | |
| 50 | - A =LICENSE= file to go with the MIT declaration, crates.io metadata, and a release | |
| 51 | profile that produces a 5.0 MB binary rather than 6.5 MB. | |
| 52 | - This changelog, and =RELEASING.org=. | |
| 53 | ||
| 54 | ** 0.17.0 | |
| 55 | - *Asset directories outside the source.* =[build] assets = ["../theme/static"]= copies | |
| 56 | a directory's contents to the site root. A site's static files do not always live where | |
| 57 | its writing does, and copying them next to the writing is how a repository ends up with | |
| 58 | two of every stylesheet. =watch= and =serve= watch these directories too. Two files | |
| 59 | claiming one URL is a build error naming both. | |
| 60 | - *Template hashing is per template.* A page's render key covered every template, so | |
| 61 | editing a feed template re-rendered the whole site. It now covers the layout the page | |
| 62 | uses plus what that layout extends, includes or imports. On a 196-page site, editing the | |
| 63 | feed template renders one page instead of 196. | |
| 64 | - Cache format 7. | |
| 65 | ||
| 66 | ** 0.16.0 | |
| 67 | - *Org's entity table.* =\alpha=, =\rarr=, =20\deg= and the other 412 names, generated | |
| 68 | from Emacs' own =org-entities=. An unknown name stays literal; =#+OPTIONS: e:nil= turns | |
| 69 | the table off. /Output changes/ for any page using entities. | |
| 70 | - *Table captions.* =#+CAPTION:= above a table becomes a numbered =<caption>=. | |
| 71 | - *=#+INCLUDE:= reports itself.* It was inert and silent, which publishes a page with | |
| 72 | content missing and nobody told. Now a diagnostic, and =--strict= makes it a failure. | |
| 73 | - The Emacs oracle separates deliberate divergence from defects. Every difference from | |
| 74 | org's exporter is named and justified, and a test asserts there are no others. | |
| 75 | ||
| 76 | ** 0.15.0 | |
| 77 | Export parity, from a page-by-page diff of a 179-file corpus against the site Emacs | |
| 78 | publishes from the same sources. *All of these change output.* | |
| 79 | ||
| 80 | - Heading levels are relative to a document's shallowest heading, as org exports them. | |
| 81 | - Org's text conversions: =--=, =---=, =...=, and =x^2= / =a_{b}=. Never inside verbatim, | |
| 82 | code, source blocks or LaTeX. =#+OPTIONS: -:nil=, =^:nil= and =^:{}= all work. | |
| 83 | - Captioned figures are numbered =Figure N:=. | |
| 84 | - A caption attaches to the element /directly/ below it; a blank line between attaches to | |
| 85 | nothing. | |
| 86 | - Checkboxes render as org writes them, which keeps the =[-]= partly-done state a disabled | |
| 87 | =<input>= could not express. =[@4]= sets a list item's number. | |
| 88 | - A table's special marker column and its marker rows stay out of the output. | |
| 89 | - =#+BEGIN_NOTE= and any other unrecognised name is a special block: a div holding parsed | |
| 90 | org rather than a =<pre>= of literal text. Verse keeps its line breaks. | |
| 91 | - Emphasis borders forbid whitespace and nothing else, so =="proxied":false== is verbatim | |
| 92 | and =~~/.config/doom/config.el~= is a path that starts with a tilde. | |
| 93 | - Listings sort on the time of day when a timestamp carries one. | |
| 94 | - Cache format 6. | |
| 95 | ||
| 96 | ** 0.14.0 | |
| 97 | - *Per-page layouts.* =[[pages]]= rules map a source path to a template, and | |
| 98 | =#+TEMPLATE:= on a page overrides any rule. A missing template fails the build naming | |
| 99 | the page, the template, and what does exist. | |
| 100 | - =page.year=, for grouping a listing by year with minijinja's =groupby=. | |
| 101 | - An explicit nav can order generated pages among authored ones. =nav.mode = "none"= now | |
| 102 | really means none. | |
| 103 | ||
| 104 | ** 0.13.0 | |
| 105 | - Bundled TOML and Org syntax definitions, a =syntaxes_dir= for your own, and org's comma | |
| 106 | escape (=,* heading= inside a block). | |
| 107 | ||
| 108 | ** 0.12.0 | |
| 109 | - =serve=: a development server with live reload, bound to loopback. | |
| 110 | - A documentation site under =docs/=, built by orgo itself. | |
| 111 | ||
| 112 | ** 0.11.0 | |
| 113 | - Table of contents as =page.toc=, section numbers, and org's =#+OPTIONS:= per-file | |
| 114 | switches. | |
| 115 | ||
| 116 | ** 0.10.0 | |
| 117 | - Excerpts, word count, reading time, a =truncate= filter, and =#+DRAFT:= pages. | |
| 118 | ||
| 119 | ** 0.9.0 | |
| 120 | - =watch=: rebuilds on OS filesystem events, debounced. | |
| 121 | ||
| 122 | ** 0.8.0 | |
| 123 | - =site.base_url=, the =absolute= and =rfc822= filters, canonical links, and an RSS feed | |
| 124 | in the scaffold that validates. | |
| 125 | ||
| 126 | ** 0.7.0 | |
| 127 | - Pagination for large listings, with a =paginator= template context that composes with | |
| 128 | grouping. | |
| 129 | ||
| 130 | ** 0.6.0 | |
| 131 | - Grouped collections: one page per tag plus a tag index. | |
| 132 | - Generated listing pages (=[[collections]]=), sorted indexes, and feeds via XML | |
| 133 | templates. | |
| 134 | - A config file, user templates, nav modes, an =init= scaffold, and discovery that will | |
| 135 | not publish =.git=. | |
| 136 | ||
| 137 | ** 0.5.0 | |
| 138 | - Parse diagnostics carry =file:line=, and pages render in parallel. | |
| 139 | - The corpus audit (=orgo audit=) and the =emacs --batch= oracle. | |
| 140 | - =#+SLUG:= decides a page's output filename — found by auditing a real corpus, where it | |
| 141 | affected 169 of 182 URLs. | |
| 142 | ||
| 143 | ** 0.4.0 | |
| 144 | - The full v1 construct scope, with the IN/OUT line under test. | |
| 145 | ||
| 146 | ** 0.3.0 | |
| 147 | - The incremental build layer: content, config and template hashing, a dependency graph, | |
| 148 | per-page render keys, and a persisted cache manifest. A full build and an incremental | |
| 149 | build produce byte-identical output. | |
| 150 | ||
| 151 | ** 0.2.0 | |
| 152 | - Multi-file site builds: a symbol table, internal link resolution, minijinja templates, | |
| 153 | tables and footnotes. | |
| 154 | ||
| 155 | ** 0.1.0 | |
| 156 | - Parse and render a single =.org= file to HTML. | |
Cargo.toml +1 −1
| @@ -4,7 +4,7 @@ version = "0.19.1" | ||
| 4 | 4 | edition = "2021" |
| 5 | 5 | description = "Org-mode static site generator that renders the org element tree straight to HTML" |
| 6 | 6 | license = "0BSD" |
| 7 | readme = "README.md" | |
| 7 | readme = "README.org" | |
| 8 | 8 | keywords = ["org-mode", "static-site-generator", "emacs", "html", "blog"] |
| 9 | 9 | categories = ["command-line-utilities", "text-processing"] |
| 10 | 10 | repository = "https://github.com/ccleberg/orgo" |
README.md deleted −755
| @@ -1,755 +0,0 @@ | ||
| 1 | # orgo | |
| 2 | ||
| 3 | An org-mode static site generator, in Rust. Org is treated as the *source language*, | |
| 4 | not an inconvenient input to be normalized into markdown. The org element tree — | |
| 5 | headings, drawers, blocks, links with their org-specific semantics — **is** the | |
| 6 | document model, and we render that tree straight to HTML. We never round-trip through | |
| 7 | a markdown-shaped intermediate representation, because the point is to preserve what | |
| 8 | markdown cannot express: property drawers, TODO/priority/tag metadata on headings, | |
| 9 | `#+` directives, ID links, named/captioned blocks, footnote semantics. | |
| 10 | ||
| 11 | The one non-obvious early commitment is **incremental builds keyed on content | |
| 12 | hashing**, treated as a first-class architectural concern from day one. The discipline | |
| 13 | it imposes on the data model — pure, hashable, dependency-tracked units — is the real | |
| 14 | deliverable, even while the corpus is small enough that a full rebuild is instant. | |
| 15 | ||
| 16 | **Full documentation is in [`docs/`](docs/)** — a site written in org and built by | |
| 17 | orgo itself. Build and read it with: | |
| 18 | ||
| 19 | ```bash | |
| 20 | cargo run -- serve docs -o docs/_site | |
| 21 | ``` | |
| 22 | ||
| 23 | ## Quick start | |
| 24 | ||
| 25 | ```bash | |
| 26 | cargo run -- init my-site # config + an editable copy of the layout + a page | |
| 27 | cargo run -- build my-site -o _site | |
| 28 | ``` | |
| 29 | ||
| 30 | Or skip the scaffolding entirely — point it at any directory of `.org` files: | |
| 31 | ||
| 32 | ```bash | |
| 33 | cargo run -- build ~/notes -o _site | |
| 34 | ``` | |
| 35 | ||
| 36 | **Zero configuration is a supported path, not a demo.** With no `orgo.toml`, no | |
| 37 | templates and no orgo-specific markup in your files, you get a complete site: pages, | |
| 38 | navigation, syntax-highlighted code and the stylesheet to colour it. Configuration | |
| 39 | changes what you get; it is never what makes it work. | |
| 40 | ||
| 41 | Discovery skips what should not be published — dot-directories such as `.git`, the config | |
| 42 | file, the templates directory, and the output directory when it sits inside the source, so | |
| 43 | `orgo build . -o _site` does the obvious thing. | |
| 44 | ||
| 45 | ## Configuration | |
| 46 | ||
| 47 | Everything is optional. `orgo init` writes a fully commented `orgo.toml`; every | |
| 48 | value below is the default. | |
| 49 | ||
| 50 | ```toml | |
| 51 | [site] | |
| 52 | title = "orgo site" | |
| 53 | base_url = "" # absolute URL, no trailing slash; needed for feeds/canonical links | |
| 54 | description = "" | |
| 55 | language = "en" | |
| 56 | ||
| 57 | [nav] | |
| 58 | mode = "top-level" # top-level | all | explicit | none | |
| 59 | # pages = ["index.org", "about.org"] # for mode = "explicit"; order is preserved | |
| 60 | ||
| 61 | [templates] | |
| 62 | dir = "templates" # base.html replaces the built-in layout | |
| 63 | expose_page_list = false | |
| 64 | ||
| 65 | # [[pages]] # which layout a section renders through; base.html by default | |
| 66 | # match = "blog" # a source directory or one .org file; most specific rule wins | |
| 67 | # template = "post.html" | |
| 68 | ||
| 69 | [highlight] | |
| 70 | theme = "InspiredGitHub" | |
| 71 | ||
| 72 | [build] | |
| 73 | drafts = false | |
| 74 | assets = [] # extra directories copied to the site root, e.g. ["../theme/static"] | |
| 75 | ||
| 76 | [html] | |
| 77 | heading_offset = 1 # a level-1 org heading becomes <h2>, beneath the layout's <h1> | |
| 78 | ``` | |
| 79 | ||
| 80 | ### Templates | |
| 81 | ||
| 82 | Drop a `base.html` into the templates directory and it replaces the built-in layout | |
| 83 | entirely. Any other `.html` file there is available to `{% include %}` and | |
| 84 | `{% extends %}`. Templates are [minijinja](https://docs.rs/minijinja) (Jinja2 syntax) and | |
| 85 | receive: | |
| 86 | ||
| 87 | | Variable | What it is | | |
| 88 | |---|---| | |
| 89 | | `body` | the rendered page HTML — use `{{ body \| safe }}` | | |
| 90 | | `page` | `.title`, `.url`, `.source`, `.date`, `.date_iso`, `.year`, `.tags`, `.content`, `.excerpt`, `.word_count`, `.reading_time`, `.toc`, `.keywords` | | |
| 91 | | `site` | `.title`, `.base_url`, `.description`, `.language` | | |
| 92 | | `nav` | list of `{title, url}`, relative to this page | | |
| 93 | | `root` | `../`-prefix back to the site root from this page | | |
| 94 | | `stylesheet` | URL of the generated `syntax.css` | | |
| 95 | | `pages` | every page's metadata — only when `expose_page_list = true` | | |
| 96 | ||
| 97 | `page.keywords` carries **every** `#+KEYWORD:` in the file under its lowercased name, so | |
| 98 | your own metadata works without this crate knowing about it: `#+CUSTOM_THING: x` is | |
| 99 | `{{ page.keywords.custom_thing }}`. | |
| 100 | ||
| 101 | `base.html` is the default layout, not the only one. A `[[pages]]` rule gives a section | |
| 102 | its own — `match = "blog"`, `template = "post.html"` — and `#+TEMPLATE: wide.html` gives | |
| 103 | one page its own, which wins over any rule. A second layout usually starts with | |
| 104 | `{% extends "base.html" %}`. | |
| 105 | ||
| 106 | Editing a template re-renders the pages that use it — template sources are a hash input, | |
| 107 | so a design change never leaves a site half-updated. | |
| 108 | ||
| 109 | ### Generated listing pages | |
| 110 | ||
| 111 | A blog index, an archive, a feed — output files with no source `.org` behind them. | |
| 112 | Repeat the block for each one: | |
| 113 | ||
| 114 | ```toml | |
| 115 | [[collections]] | |
| 116 | source = "blog" # directory to list; empty means every page | |
| 117 | output = "blog/index.html" # where to write it | |
| 118 | template = "list.html" | |
| 119 | title = "Blog" | |
| 120 | sort = "date" # date | title | path | |
| 121 | order = "desc" # desc | asc | |
| 122 | nav = true # put this listing page in the nav | |
| 123 | ``` | |
| 124 | ||
| 125 | The template gets the collection's entries as `pages`, already sorted, plus the usual | |
| 126 | `site`/`nav`/`root`. It can `{% extends "base.html" %}` to inherit the site chrome: | |
| 127 | ||
| 128 | ```jinja | |
| 129 | {% extends "base.html" %} | |
| 130 | {% block main %} | |
| 131 | <ul>{% for p in pages %} | |
| 132 | <li><time datetime="{{ p.date_iso }}">{{ p.date_iso }}</time> | |
| 133 | <a href="{{ root }}{{ p.url }}">{{ p.title }}</a></li> | |
| 134 | {% endfor %}</ul> | |
| 135 | {% endblock %} | |
| 136 | ``` | |
| 137 | ||
| 138 | `p.date_iso` is the `YYYY-MM-DD` extracted from `#+DATE:`, whatever org syntax it was | |
| 139 | written in — `[2025-09-05 Fri 10:21:00]`, `<2024-05-01 Wed>` or bare `2024-05-01`. It is | |
| 140 | also the sort key; pages without a parseable date sort last, so an undated draft never | |
| 141 | leads a dated archive. | |
| 142 | ||
| 143 | #### Pagination | |
| 144 | ||
| 145 | Set `paginate` to split a long listing across numbered pages: | |
| 146 | ||
| 147 | ```toml | |
| 148 | [[collections]] | |
| 149 | source = "blog" | |
| 150 | output = "blog/index.html" | |
| 151 | paginate = 10 | |
| 152 | paginate_output = "blog/page/{n}.html" # {n} is the 1-based page number | |
| 153 | ``` | |
| 154 | ||
| 155 | Page 1 stays at `output`, so a section's canonical URL never moves as its page count | |
| 156 | changes; only pages 2..N are named by `paginate_output`. The template gets a `paginator`: | |
| 157 | ||
| 158 | ```jinja | |
| 159 | {% if paginator and paginator.total > 1 %} | |
| 160 | <nav> | |
| 161 | {% if paginator.prev_url %}<a href="{{ paginator.prev_url }}">Newer</a>{% endif %} | |
| 162 | {% for pg in paginator.pages %} | |
| 163 | <a href="{{ pg.url }}"{% if pg.current %} aria-current="page"{% endif %}>{{ pg.number }}</a> | |
| 164 | {% endfor %} | |
| 165 | {% if paginator.next_url %}<a href="{{ paginator.next_url }}">Older</a>{% endif %} | |
| 166 | </nav> | |
| 167 | {% endif %} | |
| 168 | ``` | |
| 169 | ||
| 170 | `paginator` carries `current`, `total`, `per_page`, `total_entries`, `prev_url`, | |
| 171 | `next_url`, `first_url`, `last_url`, and `pages`. Every URL is relative to the page | |
| 172 | carrying it, so links work from page 1 (`page/2.html`) and from page 5 (`../index.html`, | |
| 173 | `6.html`) without the template knowing where it sits. An unpaginated collection has no | |
| 174 | `paginator` at all, so `{% if paginator %}` is a reliable test in a shared template. | |
| 175 | ||
| 176 | Grouping and pagination compose: each group paginates independently, which is why | |
| 177 | `paginate_output` needs `{tag}` as well as `{n}` on a grouped collection. An empty | |
| 178 | collection still emits page 1 — a section that exists but has nothing in it should say so | |
| 179 | rather than 404. When the entry count shrinks, pages that no longer exist are deleted | |
| 180 | instead of being left serving stale posts. | |
| 181 | ||
| 182 | #### Tag pages | |
| 183 | ||
| 184 | Add `group_by` and the collection emits one page *per group* instead of one page total, | |
| 185 | plus an optional index of the groups: | |
| 186 | ||
| 187 | ```toml | |
| 188 | [[collections]] | |
| 189 | source = "blog" | |
| 190 | group_by = "tags" # "tags", or any #+KEYWORD: name to group by its value | |
| 191 | output = "tags/{tag}.html" # {tag} is replaced by each group's slug | |
| 192 | template = "tag.html" | |
| 193 | title = "Tagged: {tag}" | |
| 194 | index_output = "tags/index.html" # the tag index | |
| 195 | index_template = "tags.html" | |
| 196 | index_title = "Tags" | |
| 197 | nav = true # adds the *index*, not every tag | |
| 198 | ``` | |
| 199 | ||
| 200 | A group page receives its own posts as `pages` and itself as `group` | |
| 201 | (`.name`, `.slug`, `.url`, `.count`). The index receives `groups` — every group, sorted | |
| 202 | by name: | |
| 203 | ||
| 204 | ```jinja | |
| 205 | <ul>{% for tag in groups %} | |
| 206 | <li><a href="{{ root }}{{ tag.url }}">{{ tag.name }}</a> ({{ tag.count }})</li> | |
| 207 | {% endfor %}</ul> | |
| 208 | ``` | |
| 209 | ||
| 210 | `group_by = "tags"` is multi-valued: a post appears under every tag it carries. Any other | |
| 211 | value names a single-valued `#+KEYWORD:`, so `group_by = "category"` buckets by | |
| 212 | `#+CATEGORY:`. | |
| 213 | ||
| 214 | Two tags that would produce the same URL (`web_dev` and `web@dev` both slugify to | |
| 215 | `web-dev`) are a build error rather than one page silently overwriting the other. | |
| 216 | ||
| 217 | A tag page depends on its own posts and nothing else, so adding a post tagged `rust` | |
| 218 | re-renders that post, its section index, `tags/rust.html`, and the tag index whose counts | |
| 219 | changed — four pages, not one per tag. That precision is why `groups` is given to the | |
| 220 | index and not to every group page: a page that can see every group depends on every | |
| 221 | group. | |
| 222 | ||
| 223 | #### Feeds and absolute URLs | |
| 224 | ||
| 225 | **A feed is a listing page with an XML template**, not a separate feature — templates are | |
| 226 | loaded by full filename and any extension, so `output = "feed.xml"` with | |
| 227 | `template = "feed.xml"` is all it takes. `orgo init` writes a working RSS template. | |
| 228 | ||
| 229 | A feed is read away from the site that served it, so relative links in one are simply | |
| 230 | broken. Set `site.base_url` and use the `absolute` filter: | |
| 231 | ||
| 232 | ```jinja | |
| 233 | <link>{{ post.url | absolute }}</link> | |
| 234 | <pubDate>{{ post.date_iso | rfc822 }}</pubDate> | |
| 235 | ``` | |
| 236 | ||
| 237 | | Filter | Does | | |
| 238 | |---|---| | |
| 239 | | `absolute` | site-root-relative path → absolute URL; already-absolute URLs pass through | | |
| 240 | | `rfc822` | any org or ISO date → the format RSS `pubDate` requires | | |
| 241 | | `truncate(n)` | shorten to at most `n` characters on a word boundary, with an ellipsis | | |
| 242 | ||
| 243 | Apply `absolute` to the site-root-relative values — `page.url`, `pages[].url`, | |
| 244 | `group.url` — and not to `nav[].url`, `paginator.*_url`, `stylesheet` or `root`, which | |
| 245 | are relative to the page carrying them and already correct there. | |
| 246 | ||
| 247 | With no `base_url`, `absolute` is an **error** naming the setting, rather than quietly | |
| 248 | emitting a relative URL that would make the feed invalid everywhere while looking fine. | |
| 249 | The default layout also emits `<link rel="canonical">` when a base URL is set. | |
| 250 | ||
| 251 | Listing pages are cached on the entries they list, so adding a post re-renders that | |
| 252 | section's index and nothing else. | |
| 253 | ||
| 254 | ### Table of contents and `#+OPTIONS:` | |
| 255 | ||
| 256 | `page.toc` is the page's headings as a **tree** — `{title, anchor, level, children}` — | |
| 257 | because a table of contents is one, and rebuilding a tree from a flat list of levels | |
| 258 | inside a template is what Jinja is worst at. Its anchors come from the same function the | |
| 259 | renderer uses to emit heading `id`s, so a TOC link cannot drift from the heading it | |
| 260 | points at. | |
| 261 | ||
| 262 | ```jinja | |
| 263 | {% macro toc_list(entries) %} | |
| 264 | <ul>{% for e in entries %} | |
| 265 | <li><a href="#{{ e.anchor }}">{{ e.title }}</a> | |
| 266 | {%- if e.children %}{{ toc_list(e.children) }}{% endif %}</li> | |
| 267 | {% endfor %}</ul> | |
| 268 | {% endmacro %} | |
| 269 | {% if page.toc %}{{ toc_list(page.toc) }}{% endif %} | |
| 270 | ``` | |
| 271 | ||
| 272 | Org's own per-file export switches are honoured, so a document can turn a feature off for | |
| 273 | itself the way its author already knows: | |
| 274 | ||
| 275 | | Switch | Effect | Site default | | |
| 276 | |---|---|---| | |
| 277 | | `#+OPTIONS: toc:nil` | empties `page.toc` for this page | `[html] toc = true` | | |
| 278 | | `#+OPTIONS: num:t` | numbers headings `1.`, `1.1.`, … | `[html] section_numbers = false` | | |
| 279 | ||
| 280 | **Section numbers default to off, which differs from Emacs on purpose.** | |
| 281 | `org-export-with-section-numbers` is on there, so an org-published site inherits numbered | |
| 282 | headings whether or not anyone chose them. Most sites do not want them; `num:t` or | |
| 283 | `section_numbers = true` gets Emacs' behaviour back, with Emacs' own | |
| 284 | `section-number-N` classes so the output stays diffable against the oracle. | |
| 285 | ||
| 286 | ### Excerpts and drafts | |
| 287 | ||
| 288 | `page.excerpt` is a page's `#+DESCRIPTION:` when it sets one and its first paragraph | |
| 289 | otherwise, so a listing has something to show whether or not the author thought about | |
| 290 | summaries. `page.word_count` and `page.reading_time` (minutes at 200 wpm) count prose | |
| 291 | only — a post that is mostly a shell transcript should not read as an hour's work. | |
| 292 | `truncate` exists because an excerpt is usually a whole paragraph and minijinja has no | |
| 293 | such filter. | |
| 294 | ||
| 295 | `#+DRAFT:` keeps a page out of the build entirely — no page, and absent from listings and | |
| 296 | the nav rather than merely unlinked. `--drafts` includes them, which is what you want | |
| 297 | under `watch` while writing one. A draft is out of the symbol table too, so a link *to* | |
| 298 | one is reported as the dead link it would be once published. | |
| 299 | ||
| 300 | The keyword is read forgivingly: `t`, `yes`, `1` and a bare `#+DRAFT:` all mean draft, | |
| 301 | because writing the keyword at all is the signal. Only an explicit `nil`, `false`, `no`, | |
| 302 | `0` or `off` means published. | |
| 303 | ||
| 304 | ### `#+SLUG:` | |
| 305 | ||
| 306 | A page's output filename comes from its `#+SLUG:` when it has one, so | |
| 307 | `2018-11-28-aes-encryption.org` can publish as `aes-encryption.html`. Without one the | |
| 308 | source filename is used. Slugs are sanitized to a single safe path component, and two | |
| 309 | pages claiming one URL is a build error rather than a silently dropped page. | |
| 310 | ||
| 311 | ## Pipeline | |
| 312 | ||
| 313 | ``` | |
| 314 | DISCOVER → PARSE → INDEX → RESOLVE → RENDER → TEMPLATE → EMIT | |
| 315 | ``` | |
| 316 | ||
| 317 | PARSE and RENDER are pure functions of their inputs (cacheable, hashable). INDEX/RESOLVE | |
| 318 | is the only inherently global stage — it is where the link dependency graph is born. | |
| 319 | ||
| 320 | | Stage | Module | Notes | | |
| 321 | |---|---|---| | |
| 322 | | config | `src/config.rs` | `orgo.toml`: site metadata, nav mode, templates, theme. A hash input. | | |
| 323 | | PARSE | `src/parser.rs` | Hand-written recursive descent: line lexer → element builder → inline tokenizer. | | |
| 324 | | audit | `src/audit.rs` | Phase 0 corpus audit: construct frequencies against the IN/OUT line. | | |
| 325 | | model | `src/model.rs` | The org element tree — Elements (block) vs Objects (inline). | | |
| 326 | | INDEX | `src/index.rs` | Collect link targets into a symbol table. | | |
| 327 | | RESOLVE | `src/resolve.rs` | Rewrite links to URLs; return the used-target list (dependency edges). | | |
| 328 | | RENDER | `src/render.rs` | Tree → HTML fragment; syntect highlighting; footnote two-pass. | | |
| 329 | | TEMPLATE | `src/template.rs` | minijinja: fragment + metadata → full page. | | |
| 330 | | incremental | `src/incremental.rs` | Content/config/template hashing, dep graph, cache manifest, invalidation. | | |
| 331 | ||
| 332 | ## v1 scope (delivered as of v0.4; still to be reconciled against a corpus audit) | |
| 333 | ||
| 334 | **IN — v1 must handle:** headings with nesting, at levels relative to the document's | |
| 335 | shallowest; TODO keywords; priorities `[#A]`; tags; property drawers; plain lists | |
| 336 | (unordered/ordered/description, checkboxes, `[@N]` counters, nesting); tables (with rule | |
| 337 | rows and org's special marker column, no `#+TBLFM:`); source blocks with syntax | |
| 338 | highlighting; example/quote/center/verse blocks and named special blocks; links (external, | |
| 339 | internal `[[*Heading]]`/`[[#custom-id]]`, `id:`); footnotes (inline and referenced); `#+` | |
| 340 | keywords/directives; inline markup (bold/italic/underline/verbatim/code/strike); org's | |
| 341 | export-time text conversions (`--`/`---`/`...`, `x^2`, `a_{b}`, `\alpha`); timestamps | |
| 342 | (active/inactive, ranges); paragraphs and horizontal rules; images with | |
| 343 | `#+CAPTION`/`#+ATTR_HTML`, numbered `Figure N:`. | |
| 344 | ||
| 345 | **OUT — explicitly not v1 (parse-and-ignore or reject loudly):** Babel execution / | |
| 346 | `:results`; `#+TBLFM:` formulas; LaTeX / MathJax (passed through untouched, including past | |
| 347 | the text conversions); `#+INCLUDE:` (never expanded — reported as a diagnostic, so a page | |
| 348 | is never quietly short of content); citations; radio targets and macros; drawers other | |
| 349 | than PROPERTIES/LOGBOOK; column view / clocking / agenda semantics; non-HTML export | |
| 350 | blocks. | |
| 351 | ||
| 352 | **Scope guardrail:** every IN item gets a golden-file fixture; every OUT item gets a test | |
| 353 | asserting it degrades predictably (ignored, no crash). The IN/OUT line is enforced by | |
| 354 | `tests/constructs.rs`, defending against the project's #1 risk: scope creep back toward | |
| 355 | all-of-org. Phase 0 checked this line against a real 179-file corpus and found it sound | |
| 356 | (99.9% of construct uses in scope) — but also found one thing missing from it entirely: | |
| 357 | `#+SLUG:`. See [Phase 0](#phase-0-the-corpus-audit-and-the-emacs-oracle). | |
| 358 | ||
| 359 | ## Phase plan | |
| 360 | ||
| 361 | | Phase | Scope | Status | | |
| 362 | |---|---|---| | |
| 363 | | **M0** | **Buildable skeleton: crate layout, module stubs, deps, test harness, fixtures** | **done** | | |
| 364 | | **v0.1** | **End-to-end core parse → render: `build` a single `.org` file to HTML** | **done** | | |
| 365 | | **v0.2** | **Multi-file SITE build: INDEX + RESOLVE internal links, minijinja templates, `build <src-dir> <out-dir>`, tables + footnotes** | **done** | | |
| 366 | | **v0.3** | **Incremental build layer: content/config/template hashing, dependency graph, per-page render keys, persisted cache manifest, invalidation** | **done** | | |
| 367 | | **v0.4** | **MVP: the full v1 construct scope — heading metadata, nested/description lists, block types, timestamps, images, syntect highlighting — with the IN/OUT line under test** | **done** | | |
| 368 | | **0** | **Corpus audit + `emacs --batch` ground-truth oracle** | **done** | | |
| 369 | | 1 | Line lexer + heading/section skeleton | done | | |
| 370 | | 2 | Block elements — lists, source blocks, tables, footnote defs, blocks by type, drawers | done | | |
| 371 | | 3 | Inline objects — emphasis, links, bare URLs, footnote refs, timestamps | done | | |
| 372 | | 4 | Rendering to HTML — tree walk, tables, footnote two-pass, minijinja templating, syntect highlighting | done | | |
| 373 | | 5 | Link resolution + symbol table (INDEX + RESOLVE, used-target list, broken-link reporting) | done | | |
| 374 | | 6 | Incremental build layer (hashing, dep graph, invalidation); `watch` on OS filesystem events | done | | |
| 375 | | **7** | **Hardening: rayon parallelism, error locations in parse diagnostics** | **done** | | |
| 376 | | **8** | **General use: config file, user templates, nav modes, `init` scaffold, safe discovery** | **done** | | |
| 377 | | **9** | **Generated listing pages: `[[collections]]`, sorted indexes, feeds via XML templates** | **done** | | |
| 378 | | **10** | **Grouped collections: one page per tag plus a tag index — full parity with the incumbent** | **done** | | |
| 379 | | **11** | **Pagination: numbered pages with a `paginator` context, composing with grouping** | **done** | | |
| 380 | | **12** | **`base_url`: `absolute`/`rfc822` filters, a valid RSS feed in the scaffold, canonical links** | **done** | | |
| 381 | | **13** | **`watch` on OS filesystem events, debounced, with the feedback loop closed** | **done** | | |
| 382 | | **14** | **Authoring: excerpts, word count, reading time, `truncate`, and draft pages** | **done** | | |
| 383 | | **15** | **Table of contents, section numbers, and org's `#+OPTIONS:` per-file switches** | **done** | | |
| 384 | | **16** | **`serve`: development server with long-poll live reload, loopback-bound** | **done** | | |
| 385 | | **17** | **Bundled TOML and Org syntaxes, a user syntax directory, and org's comma escape** | **done** | | |
| 386 | | **18** | **Per-page layouts: `[[pages]]` rules and `#+TEMPLATE:`** | **done** | | |
| 387 | | **19** | **Export parity: relative heading levels, special strings, sub/superscript, caption numbering, checkbox and counter markup, table marker columns, special blocks** | **done** | | |
| 388 | | **20** | **Correctness debt: org's entity table, table captions, a reported `#+INCLUDE:`, and an oracle that separates deliberate divergence from defects** | **done** | | |
| 389 | | **21** | **Extra asset roots; per-template hashing so one layout edit does not re-render the site** | **done** | | |
| 390 | | **22** | **Release engineering: CI on both platforms, a checked MSRV, release binaries, a changelog, and a written compatibility promise** | **done** | | |
| 391 | ||
| 392 | ### v0.2 in / out | |
| 393 | ||
| 394 | **Added in v0.2:** the INDEX stage (`SymbolTable` of `:ID:`/`:CUSTOM_ID:`/heading/`file:` | |
| 395 | targets across a directory); the RESOLVE stage — rewrites `[[#custom-id]]`, `[[id:...]]`, | |
| 396 | `[[*Heading]]` and `[[file:other.org]]` links to real relative output URLs, returns the | |
| 397 | `used_targets` list (the `uses` edges, spec §4.3/R2) and reports unresolved links as | |
| 398 | warnings rather than crashing; a minijinja base layout (title, nav, body) applied to every | |
| 399 | page; a `build <src-dir> <out-dir>` path that walks the tree, parses + resolves + renders + | |
| 400 | templates every `.org` into a linked static site and copies non-`.org` assets through; | |
| 401 | plus two new constructs — pipe **tables** (with header band from the rule row) and | |
| 402 | **footnotes** (block `[fn:1]` definitions, referenced `[fn:1]`, and inline `[fn:1:text]`, | |
| 403 | rendered as a numbered, back-linked notes section). | |
| 404 | ||
| 405 | **Left stubbed at v0.2, all closed in v0.4:** timestamps; TODO keywords and priorities; | |
| 406 | generic (non-PROPERTIES) drawers; real syntect tokenizing behind the `Highlighter` trait. | |
| 407 | ||
| 408 | ### v0.3 in / out | |
| 409 | ||
| 410 | **Added in v0.3 — the incremental build layer (spec §4, the flagship, non-retrofittable | |
| 411 | feature):** | |
| 412 | ||
| 413 | - **Three hash classes (spec §4.1)** in `src/incremental.rs`: a **content hash** (blake3 | |
| 414 | of a file's bytes), a **config hash** (blake3 of the resolved `BuildConfig`), and a | |
| 415 | **template hash** (blake3 of the template sources). A change in any one invalidates the | |
| 416 | pages it affects. | |
| 417 | - **Dependency graph (spec §4.3)** built from RESOLVE's `defines`/`uses` edges: a page | |
| 418 | depends on the targets it links to, so editing (or renaming a heading in) a file | |
| 419 | invalidates the pages that *link into* it, not just the file itself — the load-bearing | |
| 420 | R2 invariant. On rebuild the graph is merged with the previous build's `defines` so a | |
| 421 | *removed* target still pulls in its linkers. | |
| 422 | - **Per-page `render_key`** = `H(content ⊕ resolved-links ⊕ config ⊕ template)`. If a | |
| 423 | page's render key is unchanged, its on-disk output is already correct and it is skipped. | |
| 424 | The config component folds in a **site-structure hash** (every page's `(path, title)`), | |
| 425 | because the shared nav bar is global chrome — a title change or a page add/remove alters | |
| 426 | the nav on every page and so must re-render them all (otherwise byte-equivalence breaks). | |
| 427 | - **Persisted cache manifest** (`<out>/.orgo-cache.json`, JSON), carrying per-page | |
| 428 | records, the config/template hashes, and the serialized dependency graph, tagged with | |
| 429 | `CACHE_FORMAT_VERSION`. A version mismatch, a missing file, or a corrupt file all fall | |
| 430 | back to a clean full rebuild — the cache is an optimization, never a correctness | |
| 431 | dependency. | |
| 432 | - **Wired into `build_site`**: only pages whose render key changed (or that link into a | |
| 433 | changed file's targets) are re-rendered; unchanged outputs are left in place. `--no-cache` | |
| 434 | forces a full rebuild; `clean <out-dir>` removes the output directory (and its cache). | |
| 435 | `SiteReport` now reports `rendered` vs `skipped` counts. | |
| 436 | ||
| 437 | The hard gates are enforced by `tests/incremental.rs`: full-vs-incremental **byte | |
| 438 | equivalence** (and a second unchanged build re-rendering **zero** pages); **edit-one-file** | |
| 439 | re-renders exactly the changed page plus its linkers; **renamed-heading** invalidates the | |
| 440 | linking page and updates its emitted anchor; and cache **version-bump / missing / corrupt** | |
| 441 | all fall back to a full rebuild. | |
| 442 | ||
| 443 | **Out of scope in v0.3:** real syntect highlighting; timestamps and TODO keywords (all | |
| 444 | landed in v0.4). `watch` is a minimal mtime poll loop (`watch <src-dir> -o <out-dir>`), not | |
| 445 | an OS file-watcher — the fs-notify integration is deferred. The parse-tree cache (spec §4.5, | |
| 446 | "optionally") is not persisted: PARSE/INDEX/RESOLVE run for every file each build (cheap and | |
| 447 | pure); the incremental win is on RENDER + EMIT. | |
| 448 | ||
| 449 | ### v0.4 in / out — the MVP | |
| 450 | ||
| 451 | v0.4 closes the gap between the v1 scope above and what the code actually did, so every | |
| 452 | construct the IN list claims is now parsed, rendered, and pinned by a golden file: | |
| 453 | ||
| 454 | - **Heading metadata** — TODO keywords (the Emacs default `TODO`/`DONE` set, matched on a | |
| 455 | word boundary so `TODOs` is not one) and `[#A]` priority cookies, rendered with Emacs' | |
| 456 | own export classes so the output stays diffable against an `emacs --batch` oracle. | |
| 457 | - **Lists** — indentation-based nesting (a sub-list renders *inside* its parent `<li>`), | |
| 458 | multi-paragraph item bodies, and `term :: definition` description lists as `<dl>`. | |
| 459 | - **Blocks by type** — `QUOTE`, `CENTER`, `EXAMPLE`, `EXPORT` and `SRC` are now distinct | |
| 460 | elements rather than all collapsing to a verbatim example block. Block matching is on the | |
| 461 | specific kind, so a source block can nest inside a quote. An `html` export block passes | |
| 462 | through; every other backend drops. | |
| 463 | - **Timestamps** — active `<...>` and inactive `[...]`, optional times, same-day time | |
| 464 | ranges and `--`-joined date ranges, rendered as `<time>` with a machine-readable | |
| 465 | `datetime`. Repeater/warning cookies are recognized and discarded. | |
| 466 | - **Images** — a description-less link to an image file renders as `<img>`; with an | |
| 467 | affiliated `#+CAPTION:`/`#+ATTR_HTML:` it is promoted to a `<figure>` with the caption as | |
| 468 | both `<figcaption>` and alt text. Links to non-`.org` files are now understood as asset | |
| 469 | links: neither resolved nor reported as broken. | |
| 470 | - **Syntax highlighting** — real syntect tokenizing to CSS classes (never inline styles, so | |
| 471 | themes live in the stylesheet). Every build emits the matching `syntax.css` and each page | |
| 472 | links it relative to its own depth. An unknown language degrades to escaped `<pre><code>`. | |
| 473 | - **Diagnostics** — broken links are reported as the org syntax the author wrote | |
| 474 | (`warning: b.org: unresolved link [[#setup]]`) rather than a Debug-printed enum. | |
| 475 | ||
| 476 | **The OUT line is now enforced, not just asserted.** `tests/constructs.rs` pins each | |
| 477 | excluded construct to a specific degradation: babel is never executed *and* a checked-in | |
| 478 | `#+RESULTS:` block is dropped rather than published as if it were verified output; | |
| 479 | `#+TBLFM:` is inert; `#+INCLUDE:` is never expanded and says so; LaTeX, macros and radio targets survive | |
| 480 | as literal text; drawers other than PROPERTIES are captured and dropped; unmodelled block | |
| 481 | types keep their content verbatim. | |
| 482 | ||
| 483 | **Still out:** `#+TODO:` per-file keyword sequences; planning lines | |
| 484 | (`SCHEDULED:`/`DEADLINE:`), which render as ordinary paragraphs; and fixed-width `: ` | |
| 485 | lines. | |
| 486 | ||
| 487 | ## Serving | |
| 488 | ||
| 489 | ```bash | |
| 490 | cargo run -- serve my-site -o _site # http://127.0.0.1:3000 | |
| 491 | ``` | |
| 492 | ||
| 493 | Builds, watches, serves, and reloads the browser when a rebuild lands — the loop `watch` | |
| 494 | leaves half-open. | |
| 495 | ||
| 496 | - **Loopback by default.** A dev server serves unreviewed drafts off your laptop, so | |
| 497 | reaching the local network is something you ask for with `--host 0.0.0.0`, never | |
| 498 | something you get. | |
| 499 | - **The reload script is injected on the way out**, never written to disk. What you | |
| 500 | deploy is the built site, and it must not carry a dev server's JavaScript. | |
| 501 | - **Long-polling, not WebSockets or SSE.** The browser asks "anything since generation | |
| 502 | N?" and the server holds the request until there is. Instant like a push, no protocol | |
| 503 | beyond ordinary HTTP, and no dependency. A streamed response would have been more | |
| 504 | elegant and does not work: tiny_http buffers a response until its body ends, so a body | |
| 505 | that never ends never reaches the client. | |
| 506 | - A reload only follows a **successful** rebuild. Reloading onto a stale page because the | |
| 507 | build just failed tells you nothing; the error is already on your terminal. | |
| 508 | ||
| 509 | URL resolution is the server's security boundary and is written as a pure function with | |
| 510 | its own tests: `..`, percent-encoded `..`, backslashes, absolute paths and embedded NULs | |
| 511 | all resolve to nothing rather than to somewhere outside the output directory. | |
| 512 | ||
| 513 | ## Watching | |
| 514 | ||
| 515 | ```bash | |
| 516 | cargo run -- watch my-site -o _site | |
| 517 | ``` | |
| 518 | ||
| 519 | Rebuilds on OS filesystem events rather than polling, so it costs nothing while nothing | |
| 520 | happens. Write bursts are debounced — an editor saving a file writes a temp file, renames | |
| 521 | it over the original and touches the directory, which is one edit and several events. | |
| 522 | ||
| 523 | Two rules decide what counts as a change, and they are not the same rules the build uses | |
| 524 | to find content: | |
| 525 | ||
| 526 | - **A build input is a change.** Editing `orgo.toml` or a template rebuilds, even | |
| 527 | though discovery skips both as non-content. The question is "would this change the | |
| 528 | site?", not "is this a page?". | |
| 529 | - **Our own output is not.** `watch . -o _site` puts the output inside the source, so a | |
| 530 | rebuild's writes raise events that would trigger a rebuild, forever. Dot-directories go | |
| 531 | the same way — `.git` churns on every command — as do editor scratch files, including | |
| 532 | Emacs' `file.org~` backups, which do not start with a dot. | |
| 533 | ||
| 534 | Where native watching is unavailable (some network and container filesystems), it falls | |
| 535 | back to polling and says so, rather than failing. | |
| 536 | ||
| 537 | ## Phase 0: the corpus audit and the Emacs oracle | |
| 538 | ||
| 539 | The v1 scope was, by its own admission, *recommended* — a guess about which slice of org | |
| 540 | matters. Phase 0 replaces both halves of that guess with a measurement: an audit that asks | |
| 541 | what a real corpus actually uses, and an oracle that asks whether we render it the way | |
| 542 | Emacs does. | |
| 543 | ||
| 544 | The audit runs against any corpus — point it at your own notes before trusting this tool | |
| 545 | with them. The numbers below come from a 179-file site published today by weblorg, a | |
| 546 | wrapper around org's own HTML exporter, which makes it both a realistic workload and a | |
| 547 | directly comparable incumbent. With collections configured, orgo now reproduces | |
| 548 | **all 182 of that site's URLs**. | |
| 549 | ||
| 550 | ``` | |
| 551 | cargo run -- audit <src-dir> # what does this corpus use, and is it in scope? | |
| 552 | cargo test --test oracle # how does our HTML differ from Emacs' own export? | |
| 553 | ``` | |
| 554 | ||
| 555 | ### What the audit found | |
| 556 | ||
| 557 | **The scope guess was sound.** 99.9% of construct uses in the corpus are in scope. The | |
| 558 | whole out-of-scope tail is 8 uses: four `#+TBLFM:` in a post *about* org-mode, three | |
| 559 | `\name` entities, and one `#+BEGIN_NOTE`. | |
| 560 | ||
| 561 | **`#+SLUG:` was a hole big enough to sink the project.** 178 of 179 files set it, and the | |
| 562 | published URL comes from it, not from the filename: `2018-11-28-aes-encryption.org` is | |
| 563 | served at `blog/aes-encryption.html`. orgo derived output paths from source filenames, | |
| 564 | so **169 of 179 pages would have been published at the wrong URL** — every inbound link and | |
| 565 | every search result, broken, by a tool that reported a clean build. Output paths now come | |
| 566 | from `#+SLUG:` when present ([`util::output_path`](src/util.rs)); slugs are sanitized so an | |
| 567 | author-supplied `../../etc/x` cannot escape the output directory, and two pages claiming one | |
| 568 | URL is a build error rather than a silently dropped page. Building the real corpus now | |
| 569 | reproduces all 179 of the live site's URLs exactly. | |
| 570 | ||
| 571 | **Some machinery is speculative.** The corpus contains no `id:`, `#custom-id` or `*Heading` | |
| 572 | links at all — its cross-page links are hand-written relative URLs. The INDEX/RESOLVE | |
| 573 | symbol table that v0.2 was built around is, against this corpus, unexercised. | |
| 574 | ||
| 575 | **An audit can lie too.** The first run reported 23 uses of a custom TODO keyword sequence. | |
| 576 | All 23 were false: the detector read the leading word of `* CSS Variables` as the keyword | |
| 577 | `CSS`. The corpus defines no `#+TODO:` sequences at all, so the true count was zero. The | |
| 578 | detector now matches conventional keyword names only — a tool that overstates a gap argues | |
| 579 | for work nobody needs. | |
| 580 | ||
| 581 | ### What the oracle found | |
| 582 | ||
| 583 | `tests/oracle.rs` exports each fixture with org's own exporter via `emacs --batch`, reduces | |
| 584 | both sides to a semantic skeleton (element opens, closes and text, with layout `div`s, | |
| 585 | inline `span`s and all attributes but `href`/`src` dropped), and **snapshots the | |
| 586 | disagreement**. Snapshotting rather than asserting is deliberate: a checked-in divergence | |
| 587 | report gets reviewed and shows up as a diff, where a permanently red test gets ignored. | |
| 588 | Three invariants are asserted outright, and all three hold — heading structure, list | |
| 589 | nesting, and source-block text match Emacs exactly. | |
| 590 | ||
| 591 | **No bugs in orgo.** Every remaining divergence is a deliberate choice to emit better | |
| 592 | HTML than org does: | |
| 593 | ||
| 594 | | | orgo | Emacs | why | | |
| 595 | |---|---|---|---| | |
| 596 | | emphasis | `<em>`/`<strong>` | `<i>`/`<b>` | semantic, not presentational | | |
| 597 | | captioned image | `<figure>`/`<figcaption>` | `<p>` + `"Figure 1: …"` | real figure semantics | | |
| 598 | | timestamp | `<time datetime="…">` | literal `<2024-01-15 Mon>` | machine-readable | | |
| 599 | | footnotes | `<section><ol>` | `<h2>Footnotes:</h2>` | a list of notes is a list | | |
| 600 | | heading anchor | slug of the text | `org1a2b3c4` | stable, and what the live site serves | | |
| 601 | | code | `<pre><code>` | `<pre>` | the HTML5 idiom | | |
| 602 | ||
| 603 | One genuine semantic difference: org treats a single blank line between a `1.` list and a | |
| 604 | `-` list as *one* list and keeps the first item's bullet type, while we start a second list. | |
| 605 | We keep ours, on measurement rather than taste — the pattern occurs **zero** times in the | |
| 606 | corpus, so matching an org quirk would buy nothing and cost the more obvious reading. | |
| 607 | ||
| 608 | **The oracle's best catch was three bugs in itself.** Naive normalization reported code as | |
| 609 | corrupted (it trimmed each of syntect's per-token text runs, turning `def greet` into | |
| 610 | `defgreet`) and reported blocks at 36% agreement (syntect's spans flooded the diff). Both | |
| 611 | were measurement artifacts. A differential harness is a piece of software like any other, | |
| 612 | and the first divergences it reports are usually its own. | |
| 613 | ||
| 614 | ## Phase 7: hardening | |
| 615 | ||
| 616 | ### Parse diagnostics (`file:line: message`) | |
| 617 | ||
| 618 | The parser's contract is that it always returns a document — out-of-scope and malformed | |
| 619 | constructs degrade rather than crash. The gap was that they degraded *silently*, and in the | |
| 620 | worst cases the degradation is severe: an unterminated `#+BEGIN_SRC` reads the rest of the | |
| 621 | file as block content, and an unterminated drawer does the same but renders to nothing, so | |
| 622 | one missing line deletes most of a page from a build that reports success. | |
| 623 | ||
| 624 | `parse` now returns `Document::diagnostics`, each carrying a 1-based source line, and the | |
| 625 | build prints them as `file:line: message`. `--strict` turns them (and unresolved links) into | |
| 626 | a non-zero exit. Line numbers are threaded as an absolute offset through every nested parse, | |
| 627 | so a block inside a list item inside a section still reports its real file line — there is a | |
| 628 | test for exactly that, because reconstructed and re-indented nested slices are precisely | |
| 629 | where an off-by-N hides. The 179-file corpus produces zero diagnostics. | |
| 630 | ||
| 631 | ### Parallelism | |
| 632 | ||
| 633 | PARSE, RESOLVE and RENDER/EMIT run under rayon. PARSE is a pure function of one file's bytes | |
| 634 | and RESOLVE only reads the shared symbol table, which is what makes both safe to parallelize | |
| 635 | at all; INDEX stays sequential. | |
| 636 | ||
| 637 | | corpus | before | after | speedup | | |
| 638 | |---|---|---|---| | |
| 639 | | 179 files (real) | 0.23s | 0.07s | 3.3× | | |
| 640 | | 1,790 files (10× copy) | 3.98s | 0.82s | 4.9× | | |
| 641 | ||
| 642 | Measured on 12 cores. `RAYON_NUM_THREADS=1` reproduces the old 3.98s exactly, so the gain is | |
| 643 | parallelism rather than incidental change, and the output is byte-identical to the sequential | |
| 644 | build across the whole corpus. | |
| 645 | ||
| 646 | **Parallelism must not be observable in the result.** `par_iter().collect()` preserves input | |
| 647 | order, so the emitted bytes are unaffected — but the build *report* is the fragile half: | |
| 648 | pushing to `rendered`/`skipped` from inside the parallel pass would order them by thread | |
| 649 | scheduling, producing a non-deterministic report over a deterministic site. The parallel pass | |
| 650 | therefore returns only what was written, and the report is assembled sequentially afterwards. | |
| 651 | `parallel_builds_are_deterministic_in_output_and_report_order` holds that line, and it was | |
| 652 | verified by reintroducing the bug and watching it fail. | |
| 653 | ||
| 654 | ### The real scaling limit was not the CPU | |
| 655 | ||
| 656 | Going 10× on corpus size cost 17× in time, which parallelism improves without fixing: the | |
| 657 | cause was the nav bar listing **every** page, so an *n*-page site emitted *n*² nav links. At | |
| 658 | 1,790 pages each page carried 1,799 links and the output was 284 MB, against 5.5 MB for the | |
| 659 | 179-page corpus — 52× the bytes for 10× the input. | |
| 660 | ||
| 661 | The nav is now built from **top-level pages only** ([`is_top_level`](src/site.rs)): a nav is a | |
| 662 | map of the site's top level, not an index of its contents, and section pages reach their | |
| 663 | siblings through that section's landing page. Nav size becomes a function of the top level | |
| 664 | rather than of the corpus, and the quadratic disappears. | |
| 665 | ||
| 666 | | 1,790-page corpus (6 top-level pages) | before | after | | |
| 667 | |---|---|---| | |
| 668 | | full build | 0.82s | 0.39s | | |
| 669 | | total output | 284 MB | 34 MB | | |
| 670 | | nav links per page | 1,799 | 6 | | |
| 671 | ||
| 672 | Scaling is now linear: 179 pages in 0.07s and 1,796 in 0.39s, where the small case is mostly | |
| 673 | the fixed cost of loading syntect's syntax definitions. | |
| 674 | ||
| 675 | The same rule sharpened the incremental build, which is the larger win. The site-structure | |
| 676 | hash — the thing that forces a global re-render — now covers only the pages that appear in | |
| 677 | the nav, because those are the only ones whose title or URL affects another page. **Adding a | |
| 678 | blog post used to re-render the entire site; now it renders one page.** A top-level page's | |
| 679 | title still invalidates everything, correctly, since every page displays it. | |
| 680 | ||
| 681 | **Trade-off worth knowing:** on a site whose sections live in subdirectories, only genuinely | |
| 682 | root-level pages appear — a site keeping its landing pages at `salary/index.org` and friends | |
| 683 | gets a one-entry nav. That is what `nav.mode = "explicit"` is for: list the pages you want, | |
| 684 | in the order you want them. | |
| 685 | ||
| 686 | **From v0.1 (core subset):** headings with nesting and anchors (every heading is now | |
| 687 | anchored — `:CUSTOM_ID:`/`:ID:` else a slug of its text) and trailing tags; paragraphs; | |
| 688 | plain lists (unordered + ordered) with checkboxes; source blocks; inline markup (`*bold*`, | |
| 689 | `/italic/`, `_underline_`, `+strike+`, `=verbatim=`, `~code~`); links and bare URLs. | |
| 690 | ||
| 691 | ## Compatibility | |
| 692 | ||
| 693 | Versions mean something as of 1.0. The **stable surface** — changing incompatibly requires | |
| 694 | a major version — is what you actually build a site against: | |
| 695 | ||
| 696 | | Stable | Detail | | |
| 697 | |---|---| | |
| 698 | | `orgo.toml` keys | Names, types and meaning. New keys are minor releases; removing one is major. | | |
| 699 | | Template context | `page`, `site`, `nav`, `root`, `pages`, `group`, `groups`, `paginator`, `stylesheet`, and the `absolute` / `rfc822` / `truncate` filters. | | |
| 700 | | CLI | Command names, flags, and exit codes. | | |
| 701 | | URLs | How a source path becomes an output path, including `#+SLUG:`. A generator that moves your URLs breaks every link anyone has to you. | | |
| 702 | ||
| 703 | Explicitly **not stable**, so that the above can be: | |
| 704 | ||
| 705 | - **The incremental cache.** Versioned, discarded on mismatch, never a correctness | |
| 706 | dependency. It changes whenever it needs to, in any release. | |
| 707 | - **Rendered HTML details.** orgo tracks what Emacs exports from the same file, and | |
| 708 | closing a gap changes markup. Changes that affect output are called out in | |
| 709 | [CHANGELOG.md](CHANGELOG.md) — the class names the documentation names (`post-list`, | |
| 710 | `figure-number`, `section-number-N`, `footnote-ref`) are the ones to write CSS against. | |
| 711 | - **The Rust API.** The crate is published so the binary can be installed with | |
| 712 | `cargo install`; the library exists to serve it, and its types move as the tool does. | |
| 713 | ||
| 714 | The **MSRV is 1.88**, checked in CI on every change. orgo's own code compiles on | |
| 715 | 1.82; the floor comes from dependencies. Raising it is a minor version, never a patch. | |
| 716 | ||
| 717 | ## Dependencies | |
| 718 | ||
| 719 | Parser is hand-written recursive descent (not `nom`/`chumsky`/`pest` — org is | |
| 720 | line-oriented and context-sensitive, not clean CFG). Key crates: `syntect` (syntax | |
| 721 | highlighting, behind a `Highlighter` trait so tree-sitter can be swapped in later), | |
| 722 | `minijinja` (runtime templates), `blake3` (content/cache hashing), `rayon` (parallel | |
| 723 | PARSE/RESOLVE/RENDER), `notify` (filesystem events for `watch`), `tiny_http` (the `serve` | |
| 724 | development server), `toml` (config), `chrono`, `camino`, `walkdir`, `clap`, `anyhow`/`thiserror`. | |
| 725 | `insta` for snapshot tests, and `emacs --batch` — optional, and only for the oracle. | |
| 726 | ||
| 727 | ## Build & test | |
| 728 | ||
| 729 | ``` | |
| 730 | cargo build | |
| 731 | cargo test # 191 tests | |
| 732 | cargo run -- init my-site # scaffold a new site | |
| 733 | cargo run -- build fixtures/minimal.org -o minimal.html # single file | |
| 734 | cargo run -- build fixtures/site -o _site # whole site (incremental) | |
| 735 | cargo run -- audit fixtures/site # corpus audit (Phase 0) | |
| 736 | cargo run -- build fixtures/site -o _site --no-cache # force a full rebuild | |
| 737 | cargo run -- watch fixtures/site -o _site # rebuild on filesystem events | |
| 738 | cargo run -- serve fixtures/site -o _site # ... and serve with live reload | |
| 739 | cargo run -- clean _site # remove output + cache | |
| 740 | ``` | |
| 741 | ||
| 742 | A second `build` of an unchanged site re-renders nothing; editing a page re-renders only | |
| 743 | that page and the pages that link into it (watch the `rendered`/`cached` counts). | |
| 744 | ||
| 745 | A build emits `syntax.css` next to its output (the highlighter emits CSS classes, so the | |
| 746 | stylesheet has to come with them) and every page links it. | |
| 747 | ||
| 748 | `fixtures/` holds tiny `.org` samples: the core ones (`minimal.org`, `core.org`, | |
| 749 | `elements.org`, `table.org`, `footnote.org`), one per v1 construct group (`headings.org`, | |
| 750 | `lists.org`, `blocks.org`, `timestamps.org`, `images.org`), the scope guardrail | |
| 751 | (`outofscope.org`), and a linked multi-file site under `fixtures/site/` (`index.org`, | |
| 752 | `guide.org`, `about.org` + a `style.css` asset). The real corpus (golden files derived from | |
| 753 | actual documents) lands in Phase 0. `cargo test` runs `insta` snapshots of the element tree | |
| 754 | and rendered HTML for each fixture, the two templated site pages (proving cross-file link | |
| 755 | resolution), and the incremental gates. | |
README.org added +726
| @@ -0,0 +1,726 @@ | ||
| 1 | * orgo | |
| 2 | An org-mode static site generator, in Rust. Org is treated as the /source language/, | |
| 3 | not an inconvenient input to be normalized into markdown. The org element tree — | |
| 4 | headings, drawers, blocks, links with their org-specific semantics — *is* the | |
| 5 | document model, and we render that tree straight to HTML. We never round-trip through | |
| 6 | a markdown-shaped intermediate representation, because the point is to preserve what | |
| 7 | markdown cannot express: property drawers, TODO/priority/tag metadata on headings, | |
| 8 | =#+= directives, ID links, named/captioned blocks, footnote semantics. | |
| 9 | ||
| 10 | The one non-obvious early commitment is *incremental builds keyed on content | |
| 11 | hashing*, treated as a first-class architectural concern from day one. The discipline | |
| 12 | it imposes on the data model — pure, hashable, dependency-tracked units — is the real | |
| 13 | deliverable, even while the corpus is small enough that a full rebuild is instant. | |
| 14 | ||
| 15 | *Full documentation is in [[file:docs/][=docs/=]]* — a site written in org and built by | |
| 16 | orgo itself. Build and read it with: | |
| 17 | ||
| 18 | #+begin_src sh | |
| 19 | cargo run -- serve docs -o docs/_site | |
| 20 | #+end_src | |
| 21 | ||
| 22 | ** Quick start | |
| 23 | #+begin_src sh | |
| 24 | cargo run -- init my-site # config + an editable copy of the layout + a page | |
| 25 | cargo run -- build my-site -o _site | |
| 26 | #+end_src | |
| 27 | ||
| 28 | Or skip the scaffolding entirely — point it at any directory of =.org= files: | |
| 29 | ||
| 30 | #+begin_src sh | |
| 31 | cargo run -- build ~/notes -o _site | |
| 32 | #+end_src | |
| 33 | ||
| 34 | *Zero configuration is a supported path, not a demo.* With no =orgo.toml=, no | |
| 35 | templates and no orgo-specific markup in your files, you get a complete site: pages, | |
| 36 | navigation, syntax-highlighted code and the stylesheet to colour it. Configuration | |
| 37 | changes what you get; it is never what makes it work. | |
| 38 | ||
| 39 | Discovery skips what should not be published — dot-directories such as =.git=, the config | |
| 40 | file, the templates directory, and the output directory when it sits inside the source, so | |
| 41 | =orgo build . -o _site= does the obvious thing. | |
| 42 | ||
| 43 | ** Configuration | |
| 44 | Everything is optional. =orgo init= writes a fully commented =orgo.toml=; every | |
| 45 | value below is the default. | |
| 46 | ||
| 47 | #+begin_src toml | |
| 48 | [site] | |
| 49 | title = "orgo site" | |
| 50 | base_url = "" # absolute URL, no trailing slash; needed for feeds/canonical links | |
| 51 | description = "" | |
| 52 | language = "en" | |
| 53 | ||
| 54 | [nav] | |
| 55 | mode = "top-level" # top-level | all | explicit | none | |
| 56 | # pages = ["index.org", "about.org"] # for mode = "explicit"; order is preserved | |
| 57 | ||
| 58 | [templates] | |
| 59 | dir = "templates" # base.html replaces the built-in layout | |
| 60 | expose_page_list = false | |
| 61 | ||
| 62 | # [[pages]] # which layout a section renders through; base.html by default | |
| 63 | # match = "blog" # a source directory or one .org file; most specific rule wins | |
| 64 | # template = "post.html" | |
| 65 | ||
| 66 | [highlight] | |
| 67 | theme = "InspiredGitHub" | |
| 68 | ||
| 69 | [build] | |
| 70 | drafts = false | |
| 71 | assets = [] # extra directories copied to the site root, e.g. ["../theme/static"] | |
| 72 | ||
| 73 | [html] | |
| 74 | heading_offset = 1 # a level-1 org heading becomes <h2>, beneath the layout's <h1> | |
| 75 | #+end_src | |
| 76 | ||
| 77 | *** Templates | |
| 78 | Drop a =base.html= into the templates directory and it replaces the built-in layout | |
| 79 | entirely. Any other =.html= file there is available to ={% include %}= and | |
| 80 | ={% extends %}=. Templates are [[https://docs.rs/minijinja][minijinja]] (Jinja2 syntax) and | |
| 81 | receive: | |
| 82 | ||
| 83 | | Variable | What it is | | |
| 84 | |————--+————————————————————————————————————————————————--| | |
| 85 | | =body= | the rendered page HTML — use ={{ body \| safe }}= | | |
| 86 | | =page= | =.title=, =.url=, =.source=, =.date=, =.date_iso=, =.year=, =.tags=, =.content=, =.excerpt=, =.word_count=, =.reading_time=, =.toc=, =.keywords= | | |
| 87 | | =site= | =.title=, =.base_url=, =.description=, =.language= | | |
| 88 | | =nav= | list of ={title, url}=, relative to this page | | |
| 89 | | =root= | =../=-prefix back to the site root from this page | | |
| 90 | | =stylesheet= | URL of the generated =syntax.css= | | |
| 91 | | =pages= | every page's metadata — only when =expose_page_list = true= | | |
| 92 | ||
| 93 | =page.keywords= carries *every* =#+KEYWORD:= in the file under its lowercased name, so | |
| 94 | your own metadata works without this crate knowing about it: =#+CUSTOM_THING: x= is | |
| 95 | ={{ page.keywords.custom_thing }}=. | |
| 96 | ||
| 97 | =base.html= is the default layout, not the only one. A =[[pages]]= rule gives a section | |
| 98 | its own — =match = "blog"=, =template = "post.html"= — and =#+TEMPLATE: wide.html= gives | |
| 99 | one page its own, which wins over any rule. A second layout usually starts with | |
| 100 | ={% extends "base.html" %}=. | |
| 101 | ||
| 102 | Editing a template re-renders the pages that use it — template sources are a hash input, | |
| 103 | so a design change never leaves a site half-updated. | |
| 104 | ||
| 105 | *** Generated listing pages | |
| 106 | A blog index, an archive, a feed — output files with no source =.org= behind them. | |
| 107 | Repeat the block for each one: | |
| 108 | ||
| 109 | #+begin_src toml | |
| 110 | [[collections]] | |
| 111 | source = "blog" # directory to list; empty means every page | |
| 112 | output = "blog/index.html" # where to write it | |
| 113 | template = "list.html" | |
| 114 | title = "Blog" | |
| 115 | sort = "date" # date | title | path | |
| 116 | order = "desc" # desc | asc | |
| 117 | nav = true # put this listing page in the nav | |
| 118 | #+end_src | |
| 119 | ||
| 120 | The template gets the collection's entries as =pages=, already sorted, plus the usual | |
| 121 | =site=/=nav=/=root=. It can ={% extends "base.html" %}= to inherit the site chrome: | |
| 122 | ||
| 123 | #+begin_src jinja | |
| 124 | {% extends "base.html" %} | |
| 125 | {% block main %} | |
| 126 | <ul>{% for p in pages %} | |
| 127 | <li><time datetime="{{ p.date_iso }}">{{ p.date_iso }}</time> | |
| 128 | <a href="{{ root }}{{ p.url }}">{{ p.title }}</a></li> | |
| 129 | {% endfor %}</ul> | |
| 130 | {% endblock %} | |
| 131 | #+end_src | |
| 132 | ||
| 133 | =p.date_iso= is the =YYYY-MM-DD= extracted from =#+DATE:=, whatever org syntax it was | |
| 134 | written in — =[2025-09-05 Fri 10:21:00]=, =<2024-05-01 Wed>= or bare =2024-05-01=. It is | |
| 135 | also the sort key; pages without a parseable date sort last, so an undated draft never | |
| 136 | leads a dated archive. | |
| 137 | ||
| 138 | **** Pagination | |
| 139 | Set =paginate= to split a long listing across numbered pages: | |
| 140 | ||
| 141 | #+begin_src toml | |
| 142 | [[collections]] | |
| 143 | source = "blog" | |
| 144 | output = "blog/index.html" | |
| 145 | paginate = 10 | |
| 146 | paginate_output = "blog/page/{n}.html" # {n} is the 1-based page number | |
| 147 | #+end_src | |
| 148 | ||
| 149 | Page 1 stays at =output=, so a section's canonical URL never moves as its page count | |
| 150 | changes; only pages 2..N are named by =paginate_output=. The template gets a =paginator=: | |
| 151 | ||
| 152 | #+begin_src jinja | |
| 153 | {% if paginator and paginator.total > 1 %} | |
| 154 | <nav> | |
| 155 | {% if paginator.prev_url %}<a href="{{ paginator.prev_url }}">Newer</a>{% endif %} | |
| 156 | {% for pg in paginator.pages %} | |
| 157 | <a href="{{ pg.url }}"{% if pg.current %} aria-current="page"{% endif %}>{{ pg.number }}</a> | |
| 158 | {% endfor %} | |
| 159 | {% if paginator.next_url %}<a href="{{ paginator.next_url }}">Older</a>{% endif %} | |
| 160 | </nav> | |
| 161 | {% endif %} | |
| 162 | #+end_src | |
| 163 | ||
| 164 | =paginator= carries =current=, =total=, =per_page=, =total_entries=, =prev_url=, | |
| 165 | =next_url=, =first_url=, =last_url=, and =pages=. Every URL is relative to the page | |
| 166 | carrying it, so links work from page 1 (=page/2.html=) and from page 5 (=../index.html=, | |
| 167 | =6.html=) without the template knowing where it sits. An unpaginated collection has no | |
| 168 | =paginator= at all, so ={% if paginator %}= is a reliable test in a shared template. | |
| 169 | ||
| 170 | Grouping and pagination compose: each group paginates independently, which is why | |
| 171 | =paginate_output= needs ={tag}= as well as ={n}= on a grouped collection. An empty | |
| 172 | collection still emits page 1 — a section that exists but has nothing in it should say so | |
| 173 | rather than 404. When the entry count shrinks, pages that no longer exist are deleted | |
| 174 | instead of being left serving stale posts. | |
| 175 | ||
| 176 | **** Tag pages | |
| 177 | Add =group_by= and the collection emits one page /per group/ instead of one page total, | |
| 178 | plus an optional index of the groups: | |
| 179 | ||
| 180 | #+begin_src toml | |
| 181 | [[collections]] | |
| 182 | source = "blog" | |
| 183 | group_by = "tags" # "tags", or any #+KEYWORD: name to group by its value | |
| 184 | output = "tags/{tag}.html" # {tag} is replaced by each group's slug | |
| 185 | template = "tag.html" | |
| 186 | title = "Tagged: {tag}" | |
| 187 | index_output = "tags/index.html" # the tag index | |
| 188 | index_template = "tags.html" | |
| 189 | index_title = "Tags" | |
| 190 | nav = true # adds the *index*, not every tag | |
| 191 | #+end_src | |
| 192 | ||
| 193 | A group page receives its own posts as =pages= and itself as =group= | |
| 194 | (=.name=, =.slug=, =.url=, =.count=). The index receives =groups= — every group, sorted | |
| 195 | by name: | |
| 196 | ||
| 197 | #+begin_src jinja | |
| 198 | <ul>{% for tag in groups %} | |
| 199 | <li><a href="{{ root }}{{ tag.url }}">{{ tag.name }}</a> ({{ tag.count }})</li> | |
| 200 | {% endfor %}</ul> | |
| 201 | #+end_src | |
| 202 | ||
| 203 | =group_by = "tags"= is multi-valued: a post appears under every tag it carries. Any other | |
| 204 | value names a single-valued =#+KEYWORD:=, so =group_by = "category"= buckets by | |
| 205 | =#+CATEGORY:=. | |
| 206 | ||
| 207 | Two tags that would produce the same URL (=web_dev= and =web@dev= both slugify to | |
| 208 | =web-dev=) are a build error rather than one page silently overwriting the other. | |
| 209 | ||
| 210 | A tag page depends on its own posts and nothing else, so adding a post tagged =rust= | |
| 211 | re-renders that post, its section index, =tags/rust.html=, and the tag index whose counts | |
| 212 | changed — four pages, not one per tag. That precision is why =groups= is given to the | |
| 213 | index and not to every group page: a page that can see every group depends on every | |
| 214 | group. | |
| 215 | ||
| 216 | **** Feeds and absolute URLs | |
| 217 | *A feed is a listing page with an XML template*, not a separate feature — templates are | |
| 218 | loaded by full filename and any extension, so =output = "feed.xml"= with | |
| 219 | =template = "feed.xml"= is all it takes. =orgo init= writes a working RSS template. | |
| 220 | ||
| 221 | A feed is read away from the site that served it, so relative links in one are simply | |
| 222 | broken. Set =site.base_url= and use the =absolute= filter: | |
| 223 | ||
| 224 | #+begin_src jinja | |
| 225 | <link>{{ post.url | absolute }}</link> | |
| 226 | <pubDate>{{ post.date_iso | rfc822 }}</pubDate> | |
| 227 | #+end_src | |
| 228 | ||
| 229 | | Filter | Does | | |
| 230 | |—————+—————————————————————————-| | |
| 231 | | =absolute= | site-root-relative path → absolute URL; already-absolute URLs pass through | | |
| 232 | | =rfc822= | any org or ISO date → the format RSS =pubDate= requires | | |
| 233 | | =truncate(n)= | shorten to at most =n= characters on a word boundary, with an ellipsis | | |
| 234 | ||
| 235 | Apply =absolute= to the site-root-relative values — =page.url=, =pages[].url=, | |
| 236 | =group.url= — and not to =nav[].url=, =paginator.*_url=, =stylesheet= or =root=, which | |
| 237 | are relative to the page carrying them and already correct there. | |
| 238 | ||
| 239 | With no =base_url=, =absolute= is an *error* naming the setting, rather than quietly | |
| 240 | emitting a relative URL that would make the feed invalid everywhere while looking fine. | |
| 241 | The default layout also emits =<link rel="canonical">= when a base URL is set. | |
| 242 | ||
| 243 | Listing pages are cached on the entries they list, so adding a post re-renders that | |
| 244 | section's index and nothing else. | |
| 245 | ||
| 246 | *** Table of contents and =#+OPTIONS:= | |
| 247 | =page.toc= is the page's headings as a *tree* — ={title, anchor, level, children}= — | |
| 248 | because a table of contents is one, and rebuilding a tree from a flat list of levels | |
| 249 | inside a template is what Jinja is worst at. Its anchors come from the same function the | |
| 250 | renderer uses to emit heading =id=s, so a TOC link cannot drift from the heading it | |
| 251 | points at. | |
| 252 | ||
| 253 | #+begin_src jinja | |
| 254 | {% macro toc_list(entries) %} | |
| 255 | <ul>{% for e in entries %} | |
| 256 | <li><a href="#{{ e.anchor }}">{{ e.title }}</a> | |
| 257 | {%- if e.children %}{{ toc_list(e.children) }}{% endif %}</li> | |
| 258 | {% endfor %}</ul> | |
| 259 | {% endmacro %} | |
| 260 | {% if page.toc %}{{ toc_list(page.toc) }}{% endif %} | |
| 261 | #+end_src | |
| 262 | ||
| 263 | Org's own per-file export switches are honoured, so a document can turn a feature off for | |
| 264 | itself the way its author already knows: | |
| 265 | ||
| 266 | | Switch | Effect | Site default | | |
| 267 | |———————-+————————————+———————————-| | |
| 268 | | =#+OPTIONS: toc:nil= | empties =page.toc= for this page | =[html] toc = true= | | |
| 269 | | =#+OPTIONS: num:t= | numbers headings =1.=, =1.1.=, … | =[html] section_numbers = false= | | |
| 270 | ||
| 271 | *Section numbers default to off, which differs from Emacs on purpose.* | |
| 272 | =org-export-with-section-numbers= is on there, so an org-published site inherits numbered | |
| 273 | headings whether or not anyone chose them. Most sites do not want them; =num:t= or | |
| 274 | =section_numbers = true= gets Emacs' behaviour back, with Emacs' own | |
| 275 | =section-number-N= classes so the output stays diffable against the oracle. | |
| 276 | ||
| 277 | *** Excerpts and drafts | |
| 278 | =page.excerpt= is a page's =#+DESCRIPTION:= when it sets one and its first paragraph | |
| 279 | otherwise, so a listing has something to show whether or not the author thought about | |
| 280 | summaries. =page.word_count= and =page.reading_time= (minutes at 200 wpm) count prose | |
| 281 | only — a post that is mostly a shell transcript should not read as an hour's work. | |
| 282 | =truncate= exists because an excerpt is usually a whole paragraph and minijinja has no | |
| 283 | such filter. | |
| 284 | ||
| 285 | =#+DRAFT:= keeps a page out of the build entirely — no page, and absent from listings and | |
| 286 | the nav rather than merely unlinked. =--drafts= includes them, which is what you want | |
| 287 | under =watch= while writing one. A draft is out of the symbol table too, so a link /to/ | |
| 288 | one is reported as the dead link it would be once published. | |
| 289 | ||
| 290 | The keyword is read forgivingly: =t=, =yes=, =1= and a bare =#+DRAFT:= all mean draft, | |
| 291 | because writing the keyword at all is the signal. Only an explicit =nil=, =false=, =no=, | |
| 292 | =0= or =off= means published. | |
| 293 | ||
| 294 | *** =#+SLUG:= | |
| 295 | A page's output filename comes from its =#+SLUG:= when it has one, so | |
| 296 | =2018-11-28-aes-encryption.org= can publish as =aes-encryption.html=. Without one the | |
| 297 | source filename is used. Slugs are sanitized to a single safe path component, and two | |
| 298 | pages claiming one URL is a build error rather than a silently dropped page. | |
| 299 | ||
| 300 | ** Pipeline | |
| 301 | #+begin_example | |
| 302 | DISCOVER → PARSE → INDEX → RESOLVE → RENDER → TEMPLATE → EMIT | |
| 303 | #+end_example | |
| 304 | ||
| 305 | PARSE and RENDER are pure functions of their inputs (cacheable, hashable). INDEX/RESOLVE | |
| 306 | is the only inherently global stage — it is where the link dependency graph is born. | |
| 307 | ||
| 308 | | Stage | Module | Notes | | |
| 309 | |————-+———————-+———————————————————————————-| | |
| 310 | | config | =src/config.rs= | =orgo.toml=: site metadata, nav mode, templates, theme. A hash input. | | |
| 311 | | PARSE | =src/parser.rs= | Hand-written recursive descent: line lexer → element builder → inline tokenizer. | | |
| 312 | | audit | =src/audit.rs= | Phase 0 corpus audit: construct frequencies against the IN/OUT line. | | |
| 313 | | model | =src/model.rs= | The org element tree — Elements (block) vs Objects (inline). | | |
| 314 | | INDEX | =src/index.rs= | Collect link targets into a symbol table. | | |
| 315 | | RESOLVE | =src/resolve.rs= | Rewrite links to URLs; return the used-target list (dependency edges). | | |
| 316 | | RENDER | =src/render.rs= | Tree → HTML fragment; syntect highlighting; footnote two-pass. | | |
| 317 | | TEMPLATE | =src/template.rs= | minijinja: fragment + metadata → full page. | | |
| 318 | | incremental | =src/incremental.rs= | Content/config/template hashing, dep graph, cache manifest, invalidation. | | |
| 319 | ||
| 320 | ** v1 scope (delivered as of v0.4; still to be reconciled against a corpus audit) | |
| 321 | *IN — v1 must handle:* headings with nesting, at levels relative to the document's | |
| 322 | shallowest; TODO keywords; priorities =[#A]=; tags; property drawers; plain lists | |
| 323 | (unordered/ordered/description, checkboxes, =[@N]= counters, nesting); tables (with rule | |
| 324 | rows and org's special marker column, no =#+TBLFM:=); source blocks with syntax | |
| 325 | highlighting; example/quote/center/verse blocks and named special blocks; links (external, | |
| 326 | internal =[[*Heading]]=/=[[#custom-id]]=, =id:=); footnotes (inline and referenced); =#+= | |
| 327 | keywords/directives; inline markup (bold/italic/underline/verbatim/code/strike); org's | |
| 328 | export-time text conversions (=--=/=---=/=...=, =x^2=, =a_{b}=, =\alpha=); timestamps | |
| 329 | (active/inactive, ranges); paragraphs and horizontal rules; images with | |
| 330 | =#+CAPTION=/=#+ATTR_HTML=, numbered =Figure N:=. | |
| 331 | ||
| 332 | *OUT — explicitly not v1 (parse-and-ignore or reject loudly):* Babel execution / | |
| 333 | =:results=; =#+TBLFM:= formulas; LaTeX / MathJax (passed through untouched, including past | |
| 334 | the text conversions); =#+INCLUDE:= (never expanded — reported as a diagnostic, so a page | |
| 335 | is never quietly short of content); citations; radio targets and macros; drawers other | |
| 336 | than PROPERTIES/LOGBOOK; column view / clocking / agenda semantics; non-HTML export | |
| 337 | blocks. | |
| 338 | ||
| 339 | *Scope guardrail:* every IN item gets a golden-file fixture; every OUT item gets a test | |
| 340 | asserting it degrades predictably (ignored, no crash). The IN/OUT line is enforced by | |
| 341 | =tests/constructs.rs=, defending against the project's #1 risk: scope creep back toward | |
| 342 | all-of-org. Phase 0 checked this line against a real 179-file corpus and found it sound | |
| 343 | (99.9% of construct uses in scope) — but also found one thing missing from it entirely: | |
| 344 | =#+SLUG:=. See [[#phase-0-the-corpus-audit-and-the-emacs-oracle][Phase 0]]. | |
| 345 | ||
| 346 | ** Phase plan | |
| 347 | | Phase | Scope | Status | | |
| 348 | |——--+——————————————————————————————————————————————————————————+——--| | |
| 349 | | *M0* | *Buildable skeleton: crate layout, module stubs, deps, test harness, fixtures* | *done* | | |
| 350 | | *v0.1* | *End-to-end core parse → render: =build= a single =.org= file to HTML* | *done* | | |
| 351 | | *v0.2* | *Multi-file SITE build: INDEX + RESOLVE internal links, minijinja templates, =build <src-dir> <out-dir>=, tables + footnotes* | *done* | | |
| 352 | | *v0.3* | *Incremental build layer: content/config/template hashing, dependency graph, per-page render keys, persisted cache manifest, invalidation* | *done* | | |
| 353 | | *v0.4* | *MVP: the full v1 construct scope — heading metadata, nested/description lists, block types, timestamps, images, syntect highlighting — with the IN/OUT line under test* | *done* | | |
| 354 | | *0* | *Corpus audit + =emacs --batch= ground-truth oracle* | *done* | | |
| 355 | | 1 | Line lexer + heading/section skeleton | done | | |
| 356 | | 2 | Block elements — lists, source blocks, tables, footnote defs, blocks by type, drawers | done | | |
| 357 | | 3 | Inline objects — emphasis, links, bare URLs, footnote refs, timestamps | done | | |
| 358 | | 4 | Rendering to HTML — tree walk, tables, footnote two-pass, minijinja templating, syntect highlighting | done | | |
| 359 | | 5 | Link resolution + symbol table (INDEX + RESOLVE, used-target list, broken-link reporting) | done | | |
| 360 | | 6 | Incremental build layer (hashing, dep graph, invalidation); =watch= on OS filesystem events | done | | |
| 361 | | *7* | *Hardening: rayon parallelism, error locations in parse diagnostics* | *done* | | |
| 362 | | *8* | *General use: config file, user templates, nav modes, =init= scaffold, safe discovery* | *done* | | |
| 363 | | *9* | *Generated listing pages: =[[collections]]=, sorted indexes, feeds via XML templates* | *done* | | |
| 364 | | *10* | *Grouped collections: one page per tag plus a tag index — full parity with the incumbent* | *done* | | |
| 365 | | *11* | *Pagination: numbered pages with a =paginator= context, composing with grouping* | *done* | | |
| 366 | | *12* | *=base_url=: =absolute=/=rfc822= filters, a valid RSS feed in the scaffold, canonical links* | *done* | | |
| 367 | | *13* | *=watch= on OS filesystem events, debounced, with the feedback loop closed* | *done* | | |
| 368 | | *14* | *Authoring: excerpts, word count, reading time, =truncate=, and draft pages* | *done* | | |
| 369 | | *15* | *Table of contents, section numbers, and org's =#+OPTIONS:= per-file switches* | *done* | | |
| 370 | | *16* | *=serve=: development server with long-poll live reload, loopback-bound* | *done* | | |
| 371 | | *17* | *Bundled TOML and Org syntaxes, a user syntax directory, and org's comma escape* | *done* | | |
| 372 | | *18* | *Per-page layouts: =[[pages]]= rules and =#+TEMPLATE:=* | *done* | | |
| 373 | | *19* | *Export parity: relative heading levels, special strings, sub/superscript, caption numbering, checkbox and counter markup, table marker columns, special blocks* | *done* | | |
| 374 | | *20* | *Correctness debt: org's entity table, table captions, a reported =#+INCLUDE:=, and an oracle that separates deliberate divergence from defects* | *done* | | |
| 375 | | *21* | *Extra asset roots; per-template hashing so one layout edit does not re-render the site* | *done* | | |
| 376 | | *22* | *Release engineering: CI on both platforms, a checked MSRV, release binaries, a changelog, and a written compatibility promise* | *done* | | |
| 377 | ||
| 378 | *** v0.2 in / out | |
| 379 | *Added in v0.2:* the INDEX stage (=SymbolTable= of =:ID:=/=:CUSTOM_ID:=/heading/=file:= | |
| 380 | targets across a directory); the RESOLVE stage — rewrites =[[#custom-id]]=, =[[id:...]]=, | |
| 381 | =[[*Heading]]= and =[[file:other.org]]= links to real relative output URLs, returns the | |
| 382 | =used_targets= list (the =uses= edges, spec §4.3/R2) and reports unresolved links as | |
| 383 | warnings rather than crashing; a minijinja base layout (title, nav, body) applied to every | |
| 384 | page; a =build <src-dir> <out-dir>= path that walks the tree, parses + resolves + renders + | |
| 385 | templates every =.org= into a linked static site and copies non-=.org= assets through; | |
| 386 | plus two new constructs — pipe *tables* (with header band from the rule row) and | |
| 387 | *footnotes* (block =[fn:1]= definitions, referenced =[fn:1]=, and inline =[fn:1:text]=, | |
| 388 | rendered as a numbered, back-linked notes section). | |
| 389 | ||
| 390 | *Left stubbed at v0.2, all closed in v0.4:* timestamps; TODO keywords and priorities; | |
| 391 | generic (non-PROPERTIES) drawers; real syntect tokenizing behind the =Highlighter= trait. | |
| 392 | ||
| 393 | *** v0.3 in / out | |
| 394 | *Added in v0.3 — the incremental build layer (spec §4, the flagship, non-retrofittable | |
| 395 | feature):* | |
| 396 | ||
| 397 | - *Three hash classes (spec §4.1)* in =src/incremental.rs=: a *content hash* (blake3 | |
| 398 | of a file's bytes), a *config hash* (blake3 of the resolved =BuildConfig=), and a | |
| 399 | *template hash* (blake3 of the template sources). A change in any one invalidates the | |
| 400 | pages it affects. | |
| 401 | - *Dependency graph (spec §4.3)* built from RESOLVE's =defines=/=uses= edges: a page | |
| 402 | depends on the targets it links to, so editing (or renaming a heading in) a file | |
| 403 | invalidates the pages that /link into/ it, not just the file itself — the load-bearing | |
| 404 | R2 invariant. On rebuild the graph is merged with the previous build's =defines= so a | |
| 405 | /removed/ target still pulls in its linkers. | |
| 406 | - *Per-page =render_key=* = =H(content ⊕ resolved-links ⊕ config ⊕ template)=. If a | |
| 407 | page's render key is unchanged, its on-disk output is already correct and it is skipped. | |
| 408 | The config component folds in a *site-structure hash* (every page's =(path, title)=), | |
| 409 | because the shared nav bar is global chrome — a title change or a page add/remove alters | |
| 410 | the nav on every page and so must re-render them all (otherwise byte-equivalence breaks). | |
| 411 | - *Persisted cache manifest* (=<out>/.orgo-cache.json=, JSON), carrying per-page | |
| 412 | records, the config/template hashes, and the serialized dependency graph, tagged with | |
| 413 | =CACHE_FORMAT_VERSION=. A version mismatch, a missing file, or a corrupt file all fall | |
| 414 | back to a clean full rebuild — the cache is an optimization, never a correctness | |
| 415 | dependency. | |
| 416 | - *Wired into =build_site=*: only pages whose render key changed (or that link into a | |
| 417 | changed file's targets) are re-rendered; unchanged outputs are left in place. =--no-cache= | |
| 418 | forces a full rebuild; =clean <out-dir>= removes the output directory (and its cache). | |
| 419 | =SiteReport= now reports =rendered= vs =skipped= counts. | |
| 420 | ||
| 421 | The hard gates are enforced by =tests/incremental.rs=: full-vs-incremental *byte | |
| 422 | equivalence* (and a second unchanged build re-rendering *zero* pages); *edit-one-file* | |
| 423 | re-renders exactly the changed page plus its linkers; *renamed-heading* invalidates the | |
| 424 | linking page and updates its emitted anchor; and cache *version-bump / missing / corrupt* | |
| 425 | all fall back to a full rebuild. | |
| 426 | ||
| 427 | *Out of scope in v0.3:* real syntect highlighting; timestamps and TODO keywords (all | |
| 428 | landed in v0.4). =watch= is a minimal mtime poll loop (=watch <src-dir> -o <out-dir>=), not | |
| 429 | an OS file-watcher — the fs-notify integration is deferred. The parse-tree cache (spec §4.5, | |
| 430 | "optionally") is not persisted: PARSE/INDEX/RESOLVE run for every file each build (cheap and | |
| 431 | pure); the incremental win is on RENDER + EMIT. | |
| 432 | ||
| 433 | *** v0.4 in / out — the MVP | |
| 434 | v0.4 closes the gap between the v1 scope above and what the code actually did, so every | |
| 435 | construct the IN list claims is now parsed, rendered, and pinned by a golden file: | |
| 436 | ||
| 437 | - *Heading metadata* — TODO keywords (the Emacs default =TODO=/=DONE= set, matched on a | |
| 438 | word boundary so =TODOs= is not one) and =[#A]= priority cookies, rendered with Emacs' | |
| 439 | own export classes so the output stays diffable against an =emacs --batch= oracle. | |
| 440 | - *Lists* — indentation-based nesting (a sub-list renders /inside/ its parent =<li>=), | |
| 441 | multi-paragraph item bodies, and =term :: definition= description lists as =<dl>=. | |
| 442 | - *Blocks by type* — =QUOTE=, =CENTER=, =EXAMPLE=, =EXPORT= and =SRC= are now distinct | |
| 443 | elements rather than all collapsing to a verbatim example block. Block matching is on the | |
| 444 | specific kind, so a source block can nest inside a quote. An =html= export block passes | |
| 445 | through; every other backend drops. | |
| 446 | - *Timestamps* — active =<...>= and inactive =[...]=, optional times, same-day time | |
| 447 | ranges and =--=-joined date ranges, rendered as =<time>= with a machine-readable | |
| 448 | =datetime=. Repeater/warning cookies are recognized and discarded. | |
| 449 | - *Images* — a description-less link to an image file renders as =<img>=; with an | |
| 450 | affiliated =#+CAPTION:=/=#+ATTR_HTML:= it is promoted to a =<figure>= with the caption as | |
| 451 | both =<figcaption>= and alt text. Links to non-=.org= files are now understood as asset | |
| 452 | links: neither resolved nor reported as broken. | |
| 453 | - *Syntax highlighting* — real syntect tokenizing to CSS classes (never inline styles, so | |
| 454 | themes live in the stylesheet). Every build emits the matching =syntax.css= and each page | |
| 455 | links it relative to its own depth. An unknown language degrades to escaped =<pre><code>=. | |
| 456 | - *Diagnostics* — broken links are reported as the org syntax the author wrote | |
| 457 | (=warning: b.org: unresolved link [[#setup]]=) rather than a Debug-printed enum. | |
| 458 | ||
| 459 | *The OUT line is now enforced, not just asserted.* =tests/constructs.rs= pins each | |
| 460 | excluded construct to a specific degradation: babel is never executed /and/ a checked-in | |
| 461 | =#+RESULTS:= block is dropped rather than published as if it were verified output; | |
| 462 | =#+TBLFM:= is inert; =#+INCLUDE:= is never expanded and says so; LaTeX, macros and radio targets survive | |
| 463 | as literal text; drawers other than PROPERTIES are captured and dropped; unmodelled block | |
| 464 | types keep their content verbatim. | |
| 465 | ||
| 466 | *Still out:* =#+TODO:= per-file keyword sequences; planning lines | |
| 467 | (=SCHEDULED:=/=DEADLINE:=), which render as ordinary paragraphs; and fixed-width =:= | |
| 468 | lines. | |
| 469 | ||
| 470 | ** Serving | |
| 471 | #+begin_src sh | |
| 472 | cargo run -- serve my-site -o _site # http://127.0.0.1:3000 | |
| 473 | #+end_src | |
| 474 | ||
| 475 | Builds, watches, serves, and reloads the browser when a rebuild lands — the loop =watch= | |
| 476 | leaves half-open. | |
| 477 | ||
| 478 | - *Loopback by default.* A dev server serves unreviewed drafts off your laptop, so | |
| 479 | reaching the local network is something you ask for with =--host 0.0.0.0=, never | |
| 480 | something you get. | |
| 481 | - *The reload script is injected on the way out*, never written to disk. What you | |
| 482 | deploy is the built site, and it must not carry a dev server's JavaScript. | |
| 483 | - *Long-polling, not WebSockets or SSE.* The browser asks "anything since generation | |
| 484 | N?" and the server holds the request until there is. Instant like a push, no protocol | |
| 485 | beyond ordinary HTTP, and no dependency. A streamed response would have been more | |
| 486 | elegant and does not work: tiny_http buffers a response until its body ends, so a body | |
| 487 | that never ends never reaches the client. | |
| 488 | - A reload only follows a *successful* rebuild. Reloading onto a stale page because the | |
| 489 | build just failed tells you nothing; the error is already on your terminal. | |
| 490 | ||
| 491 | URL resolution is the server's security boundary and is written as a pure function with | |
| 492 | its own tests: =..=, percent-encoded =..=, backslashes, absolute paths and embedded NULs | |
| 493 | all resolve to nothing rather than to somewhere outside the output directory. | |
| 494 | ||
| 495 | ** Watching | |
| 496 | #+begin_src sh | |
| 497 | cargo run -- watch my-site -o _site | |
| 498 | #+end_src | |
| 499 | ||
| 500 | Rebuilds on OS filesystem events rather than polling, so it costs nothing while nothing | |
| 501 | happens. Write bursts are debounced — an editor saving a file writes a temp file, renames | |
| 502 | it over the original and touches the directory, which is one edit and several events. | |
| 503 | ||
| 504 | Two rules decide what counts as a change, and they are not the same rules the build uses | |
| 505 | to find content: | |
| 506 | ||
| 507 | - *A build input is a change.* Editing =orgo.toml= or a template rebuilds, even | |
| 508 | though discovery skips both as non-content. The question is "would this change the | |
| 509 | site?", not "is this a page?". | |
| 510 | - *Our own output is not.* =watch . -o _site= puts the output inside the source, so a | |
| 511 | rebuild's writes raise events that would trigger a rebuild, forever. Dot-directories go | |
| 512 | the same way — =.git= churns on every command — as do editor scratch files, including | |
| 513 | Emacs' =file.org~= backups, which do not start with a dot. | |
| 514 | ||
| 515 | Where native watching is unavailable (some network and container filesystems), it falls | |
| 516 | back to polling and says so, rather than failing. | |
| 517 | ||
| 518 | ** Phase 0: the corpus audit and the Emacs oracle | |
| 519 | The v1 scope was, by its own admission, /recommended/ — a guess about which slice of org | |
| 520 | matters. Phase 0 replaces both halves of that guess with a measurement: an audit that asks | |
| 521 | what a real corpus actually uses, and an oracle that asks whether we render it the way | |
| 522 | Emacs does. | |
| 523 | ||
| 524 | The audit runs against any corpus — point it at your own notes before trusting this tool | |
| 525 | with them. The numbers below come from a 179-file site published today by weblorg, a | |
| 526 | wrapper around org's own HTML exporter, which makes it both a realistic workload and a | |
| 527 | directly comparable incumbent. With collections configured, orgo now reproduces | |
| 528 | *all 182 of that site's URLs*. | |
| 529 | ||
| 530 | #+begin_example | |
| 531 | cargo run -- audit <src-dir> # what does this corpus use, and is it in scope? | |
| 532 | cargo test --test oracle # how does our HTML differ from Emacs' own export? | |
| 533 | #+end_example | |
| 534 | ||
| 535 | *** What the audit found | |
| 536 | *The scope guess was sound.* 99.9% of construct uses in the corpus are in scope. The | |
| 537 | whole out-of-scope tail is 8 uses: four =#+TBLFM:= in a post /about/ org-mode, three | |
| 538 | =\name= entities, and one =#+BEGIN_NOTE=. | |
| 539 | ||
| 540 | *=#+SLUG:= was a hole big enough to sink the project.* 178 of 179 files set it, and the | |
| 541 | published URL comes from it, not from the filename: =2018-11-28-aes-encryption.org= is | |
| 542 | served at =blog/aes-encryption.html=. orgo derived output paths from source filenames, | |
| 543 | so *169 of 179 pages would have been published at the wrong URL* — every inbound link and | |
| 544 | every search result, broken, by a tool that reported a clean build. Output paths now come | |
| 545 | from =#+SLUG:= when present ([[file:src/util.rs][=util::output_path=]]); slugs are sanitized so an | |
| 546 | author-supplied =../../etc/x= cannot escape the output directory, and two pages claiming one | |
| 547 | URL is a build error rather than a silently dropped page. Building the real corpus now | |
| 548 | reproduces all 179 of the live site's URLs exactly. | |
| 549 | ||
| 550 | *Some machinery is speculative.* The corpus contains no =id:=, =#custom-id= or =*Heading= | |
| 551 | links at all — its cross-page links are hand-written relative URLs. The INDEX/RESOLVE | |
| 552 | symbol table that v0.2 was built around is, against this corpus, unexercised. | |
| 553 | ||
| 554 | *An audit can lie too.* The first run reported 23 uses of a custom TODO keyword sequence. | |
| 555 | All 23 were false: the detector read the leading word of =* CSS Variables= as the keyword | |
| 556 | =CSS=. The corpus defines no =#+TODO:= sequences at all, so the true count was zero. The | |
| 557 | detector now matches conventional keyword names only — a tool that overstates a gap argues | |
| 558 | for work nobody needs. | |
| 559 | ||
| 560 | *** What the oracle found | |
| 561 | =tests/oracle.rs= exports each fixture with org's own exporter via =emacs --batch=, reduces | |
| 562 | both sides to a semantic skeleton (element opens, closes and text, with layout =div=s, | |
| 563 | inline =span=s and all attributes but =href=/=src= dropped), and *snapshots the | |
| 564 | disagreement*. Snapshotting rather than asserting is deliberate: a checked-in divergence | |
| 565 | report gets reviewed and shows up as a diff, where a permanently red test gets ignored. | |
| 566 | Three invariants are asserted outright, and all three hold — heading structure, list | |
| 567 | nesting, and source-block text match Emacs exactly. | |
| 568 | ||
| 569 | *No bugs in orgo.* Every remaining divergence is a deliberate choice to emit better | |
| 570 | HTML than org does: | |
| 571 | ||
| 572 | | | orgo | Emacs | why | | |
| 573 | |—————--+—————————+—————————-+—————————————| | |
| 574 | | emphasis | =<em>=/=<strong>= | =<i>=/=<b>= | semantic, not presentational | | |
| 575 | | captioned image | =<figure>=/=<figcaption>= | =<p>= + ="Figure 1: …"= | real figure semantics | | |
| 576 | | timestamp | =<time datetime="…">= | literal =<2024-01-15 Mon>= | machine-readable | | |
| 577 | | footnotes | =<section><ol>= | =<h2>Footnotes:</h2>= | a list of notes is a list | | |
| 578 | | heading anchor | slug of the text | =org1a2b3c4= | stable, and what the live site serves | | |
| 579 | | code | =<pre><code>= | =<pre>= | the HTML5 idiom | | |
| 580 | ||
| 581 | One genuine semantic difference: org treats a single blank line between a =1.= list and a | |
| 582 | =-= list as /one/ list and keeps the first item's bullet type, while we start a second list. | |
| 583 | We keep ours, on measurement rather than taste — the pattern occurs *zero* times in the | |
| 584 | corpus, so matching an org quirk would buy nothing and cost the more obvious reading. | |
| 585 | ||
| 586 | *The oracle's best catch was three bugs in itself.* Naive normalization reported code as | |
| 587 | corrupted (it trimmed each of syntect's per-token text runs, turning =def greet= into | |
| 588 | =defgreet=) and reported blocks at 36% agreement (syntect's spans flooded the diff). Both | |
| 589 | were measurement artifacts. A differential harness is a piece of software like any other, | |
| 590 | and the first divergences it reports are usually its own. | |
| 591 | ||
| 592 | ** Phase 7: hardening | |
| 593 | *** Parse diagnostics (=file:line: message=) | |
| 594 | The parser's contract is that it always returns a document — out-of-scope and malformed | |
| 595 | constructs degrade rather than crash. The gap was that they degraded /silently/, and in the | |
| 596 | worst cases the degradation is severe: an unterminated =#+BEGIN_SRC= reads the rest of the | |
| 597 | file as block content, and an unterminated drawer does the same but renders to nothing, so | |
| 598 | one missing line deletes most of a page from a build that reports success. | |
| 599 | ||
| 600 | =parse= now returns =Document::diagnostics=, each carrying a 1-based source line, and the | |
| 601 | build prints them as =file:line: message=. =--strict= turns them (and unresolved links) into | |
| 602 | a non-zero exit. Line numbers are threaded as an absolute offset through every nested parse, | |
| 603 | so a block inside a list item inside a section still reports its real file line — there is a | |
| 604 | test for exactly that, because reconstructed and re-indented nested slices are precisely | |
| 605 | where an off-by-N hides. The 179-file corpus produces zero diagnostics. | |
| 606 | ||
| 607 | *** Parallelism | |
| 608 | PARSE, RESOLVE and RENDER/EMIT run under rayon. PARSE is a pure function of one file's bytes | |
| 609 | and RESOLVE only reads the shared symbol table, which is what makes both safe to parallelize | |
| 610 | at all; INDEX stays sequential. | |
| 611 | ||
| 612 | | corpus | before | after | speedup | | |
| 613 | |————————+——--+——-+———| | |
| 614 | | 179 files (real) | 0.23s | 0.07s | 3.3× | | |
| 615 | | 1,790 files (10× copy) | 3.98s | 0.82s | 4.9× | | |
| 616 | ||
| 617 | Measured on 12 cores. =RAYON_NUM_THREADS=1= reproduces the old 3.98s exactly, so the gain is | |
| 618 | parallelism rather than incidental change, and the output is byte-identical to the sequential | |
| 619 | build across the whole corpus. | |
| 620 | ||
| 621 | *Parallelism must not be observable in the result.* =par_iter().collect()= preserves input | |
| 622 | order, so the emitted bytes are unaffected — but the build /report/ is the fragile half: | |
| 623 | pushing to =rendered=/=skipped= from inside the parallel pass would order them by thread | |
| 624 | scheduling, producing a non-deterministic report over a deterministic site. The parallel pass | |
| 625 | therefore returns only what was written, and the report is assembled sequentially afterwards. | |
| 626 | =parallel_builds_are_deterministic_in_output_and_report_order= holds that line, and it was | |
| 627 | verified by reintroducing the bug and watching it fail. | |
| 628 | ||
| 629 | *** The real scaling limit was not the CPU | |
| 630 | Going 10× on corpus size cost 17× in time, which parallelism improves without fixing: the | |
| 631 | cause was the nav bar listing *every* page, so an /n/-page site emitted /n/² nav links. At | |
| 632 | 1,790 pages each page carried 1,799 links and the output was 284 MB, against 5.5 MB for the | |
| 633 | 179-page corpus — 52× the bytes for 10× the input. | |
| 634 | ||
| 635 | The nav is now built from *top-level pages only* ([[file:src/site.rs][=is_top_level=]]): a nav is a | |
| 636 | map of the site's top level, not an index of its contents, and section pages reach their | |
| 637 | siblings through that section's landing page. Nav size becomes a function of the top level | |
| 638 | rather than of the corpus, and the quadratic disappears. | |
| 639 | ||
| 640 | | 1,790-page corpus (6 top-level pages) | before | after | | |
| 641 | |—————————————+——--+——-| | |
| 642 | | full build | 0.82s | 0.39s | | |
| 643 | | total output | 284 MB | 34 MB | | |
| 644 | | nav links per page | 1,799 | 6 | | |
| 645 | ||
| 646 | Scaling is now linear: 179 pages in 0.07s and 1,796 in 0.39s, where the small case is mostly | |
| 647 | the fixed cost of loading syntect's syntax definitions. | |
| 648 | ||
| 649 | The same rule sharpened the incremental build, which is the larger win. The site-structure | |
| 650 | hash — the thing that forces a global re-render — now covers only the pages that appear in | |
| 651 | the nav, because those are the only ones whose title or URL affects another page. *Adding a | |
| 652 | blog post used to re-render the entire site; now it renders one page.* A top-level page's | |
| 653 | title still invalidates everything, correctly, since every page displays it. | |
| 654 | ||
| 655 | *Trade-off worth knowing:* on a site whose sections live in subdirectories, only genuinely | |
| 656 | root-level pages appear — a site keeping its landing pages at =salary/index.org= and friends | |
| 657 | gets a one-entry nav. That is what =nav.mode = "explicit"= is for: list the pages you want, | |
| 658 | in the order you want them. | |
| 659 | ||
| 660 | *From v0.1 (core subset):* headings with nesting and anchors (every heading is now | |
| 661 | anchored — =:CUSTOM_ID:=/=:ID:= else a slug of its text) and trailing tags; paragraphs; | |
| 662 | plain lists (unordered + ordered) with checkboxes; source blocks; inline markup (=*bold*=, | |
| 663 | =/italic/=, =_underline_=, =+strike+=, ==verbatim==, =~code~=); links and bare URLs. | |
| 664 | ||
| 665 | ** Compatibility | |
| 666 | Versions mean something as of 1.0. The *stable surface* — changing incompatibly requires | |
| 667 | a major version — is what you actually build a site against: | |
| 668 | ||
| 669 | | Stable | Detail | | |
| 670 | |——————+——————————————————————————————————————————————-| | |
| 671 | | =orgo.toml= keys | Names, types and meaning. New keys are minor releases; removing one is major. | | |
| 672 | | Template context | =page=, =site=, =nav=, =root=, =pages=, =group=, =groups=, =paginator=, =stylesheet=, and the =absolute= / =rfc822= / =truncate= filters. | | |
| 673 | | CLI | Command names, flags, and exit codes. | | |
| 674 | | URLs | How a source path becomes an output path, including =#+SLUG:=. A generator that moves your URLs breaks every link anyone has to you. | | |
| 675 | ||
| 676 | Explicitly *not stable*, so that the above can be: | |
| 677 | ||
| 678 | - *The incremental cache.* Versioned, discarded on mismatch, never a correctness | |
| 679 | dependency. It changes whenever it needs to, in any release. | |
| 680 | - *Rendered HTML details.* orgo tracks what Emacs exports from the same file, and | |
| 681 | closing a gap changes markup. Changes that affect output are called out in | |
| 682 | [[file:CHANGELOG.org][CHANGELOG.org]] — the class names the documentation names (=post-list=, | |
| 683 | =figure-number=, =section-number-N=, =footnote-ref=) are the ones to write CSS against. | |
| 684 | - *The Rust API.* The crate is published so the binary can be installed with | |
| 685 | =cargo install=; the library exists to serve it, and its types move as the tool does. | |
| 686 | ||
| 687 | The *MSRV is 1.88*, checked in CI on every change. orgo's own code compiles on | |
| 688 | 1.82; the floor comes from dependencies. Raising it is a minor version, never a patch. | |
| 689 | ||
| 690 | ** Dependencies | |
| 691 | Parser is hand-written recursive descent (not =nom=/=chumsky=/=pest= — org is | |
| 692 | line-oriented and context-sensitive, not clean CFG). Key crates: =syntect= (syntax | |
| 693 | highlighting, behind a =Highlighter= trait so tree-sitter can be swapped in later), | |
| 694 | =minijinja= (runtime templates), =blake3= (content/cache hashing), =rayon= (parallel | |
| 695 | PARSE/RESOLVE/RENDER), =notify= (filesystem events for =watch=), =tiny_http= (the =serve= | |
| 696 | development server), =toml= (config), =chrono=, =camino=, =walkdir=, =clap=, =anyhow=/=thiserror=. | |
| 697 | =insta= for snapshot tests, and =emacs --batch= — optional, and only for the oracle. | |
| 698 | ||
| 699 | ** Build & test | |
| 700 | #+begin_example | |
| 701 | cargo build | |
| 702 | cargo test # 191 tests | |
| 703 | cargo run -- init my-site # scaffold a new site | |
| 704 | cargo run -- build fixtures/minimal.org -o minimal.html # single file | |
| 705 | cargo run -- build fixtures/site -o _site # whole site (incremental) | |
| 706 | cargo run -- audit fixtures/site # corpus audit (Phase 0) | |
| 707 | cargo run -- build fixtures/site -o _site --no-cache # force a full rebuild | |
| 708 | cargo run -- watch fixtures/site -o _site # rebuild on filesystem events | |
| 709 | cargo run -- serve fixtures/site -o _site # ... and serve with live reload | |
| 710 | cargo run -- clean _site # remove output + cache | |
| 711 | #+end_example | |
| 712 | ||
| 713 | A second =build= of an unchanged site re-renders nothing; editing a page re-renders only | |
| 714 | that page and the pages that link into it (watch the =rendered=/=cached= counts). | |
| 715 | ||
| 716 | A build emits =syntax.css= next to its output (the highlighter emits CSS classes, so the | |
| 717 | stylesheet has to come with them) and every page links it. | |
| 718 | ||
| 719 | =fixtures/= holds tiny =.org= samples: the core ones (=minimal.org=, =core.org=, | |
| 720 | =elements.org=, =table.org=, =footnote.org=), one per v1 construct group (=headings.org=, | |
| 721 | =lists.org=, =blocks.org=, =timestamps.org=, =images.org=), the scope guardrail | |
| 722 | (=outofscope.org=), and a linked multi-file site under =fixtures/site/= (=index.org=, | |
| 723 | =guide.org=, =about.org= + a =style.css= asset). The real corpus (golden files derived from | |
| 724 | actual documents) lands in Phase 0. =cargo test= runs =insta= snapshots of the element tree | |
| 725 | and rendered HTML for each fixture, the two templated site pages (proving cross-file link | |
| 726 | resolution), and the incremental gates. | |
RELEASING.md → RELEASING.org renamed +31 −35
| @@ -1,79 +1,75 @@ | ||
| 1 | # Releasing | |
| 2 | ||
| 3 | A release is three things that must agree: a version in `Cargo.toml`, a git tag, and a | |
| 1 | * Releasing | |
| 2 | A release is three things that must agree: a version in =Cargo.toml=, a git tag, and a | |
| 4 | 3 | changelog entry. The release workflow checks the first two against each other and refuses |
| 5 | to build if they differ, because a release tagged `v0.18.0` containing a binary that | |
| 6 | reports `0.17.0` is the kind of mistake nobody notices for months. | |
| 7 | ||
| 8 | ## Before the first publish | |
| 4 | to build if they differ, because a release tagged =v0.18.0= containing a binary that | |
| 5 | reports =0.17.0= is the kind of mistake nobody notices for months. | |
| 9 | 6 | |
| 10 | ```bash | |
| 7 | ** Before the first publish | |
| 8 | #+begin_src sh | |
| 11 | 9 | cargo login # a crates.io token, once per machine |
| 12 | 10 | cargo publish --dry-run |
| 13 | ``` | |
| 11 | #+end_src | |
| 14 | 12 | |
| 15 | `repository` and `homepage` in `Cargo.toml` point at GitHub and at the documentation site | |
| 16 | on Pages. If git.krz.sh becomes the primary remote, `repository` should follow it — | |
| 13 | =repository= and =homepage= in =Cargo.toml= point at GitHub and at the documentation site | |
| 14 | on Pages. If git.krz.sh becomes the primary remote, =repository= should follow it — | |
| 17 | 15 | crates.io shows that link on the crate page, and it should lead somewhere you read. |
| 18 | 16 | |
| 19 | ## Every release | |
| 20 | ||
| 21 | 1. **Write the changelog entry first.** [CHANGELOG.md](CHANGELOG.md) names behaviour, not | |
| 17 | ** Every release | |
| 18 | 1. *Write the changelog entry first.* [[file:CHANGELOG.org][CHANGELOG.org]] names behaviour, not | |
| 22 | 19 | commits — someone reading it wants to know what their next build will do differently. |
| 23 | 20 | Anything that changes rendered HTML gets said out loud. |
| 24 | 21 | |
| 25 | 2. **Bump the version** in `Cargo.toml`, and build once so `Cargo.lock` follows. | |
| 22 | 2. *Bump the version* in =Cargo.toml=, and build once so =Cargo.lock= follows. | |
| 26 | 23 | |
| 27 | 24 | Patch for fixes that change nothing about the stable surface. Minor for new config |
| 28 | 25 | keys, new template variables, an MSRV bump, or output that changes to track Emacs more |
| 29 | 26 | closely. Major for anything that breaks the promises in the README's Compatibility |
| 30 | 27 | section — config keys, template context, CLI, or URLs. |
| 31 | 28 | |
| 32 | 3. **Check it.** | |
| 29 | 3. *Check it.* | |
| 33 | 30 | |
| 34 | ```bash | |
| 31 | #+begin_src sh | |
| 35 | 32 | cargo test |
| 36 | 33 | cargo clippy --all-targets -- -D warnings |
| 37 | 34 | cargo run -- build docs -o docs/_site --strict |
| 38 | 35 | cargo package |
| 39 | ``` | |
| 36 | #+end_src | |
| 40 | 37 | |
| 41 | `cargo package` is the one people forget: it builds the crate exactly as crates.io will | |
| 42 | receive it, and catches a file the `exclude` list should not have removed. | |
| 38 | =cargo package= is the one people forget: it builds the crate exactly as crates.io will | |
| 39 | receive it, and catches a file the =exclude= list should not have removed. | |
| 43 | 40 | |
| 44 | 4. **Verify against a real corpus.** The test suite says the code does what it did; a | |
| 45 | corpus says the *site* does. Build a site you know with `--no-cache` and diff the | |
| 41 | 4. *Verify against a real corpus.* The test suite says the code does what it did; a | |
| 42 | corpus says the /site/ does. Build a site you know with =--no-cache= and diff the | |
| 46 | 43 | output against the previous version's. A release that quietly changes 200 pages should |
| 47 | 44 | do so on purpose. |
| 48 | 45 | |
| 49 | 5. **Commit, tag, push.** | |
| 46 | 5. *Commit, tag, push.* | |
| 50 | 47 | |
| 51 | ```bash | |
| 48 | #+begin_src sh | |
| 52 | 49 | git commit -am "0.18: <what changed>" |
| 53 | 50 | git tag -a v0.18.0 -m "0.18.0" |
| 54 | 51 | git push && git push --tags |
| 55 | ``` | |
| 52 | #+end_src | |
| 56 | 53 | |
| 57 | 6. **Publish the crate.** | |
| 54 | 6. *Publish the crate.* | |
| 58 | 55 | |
| 59 | ```bash | |
| 56 | #+begin_src sh | |
| 60 | 57 | cargo publish |
| 61 | ``` | |
| 58 | #+end_src | |
| 62 | 59 | |
| 63 | 60 | This is irreversible: a published version can be yanked but never replaced. |
| 64 | 61 | |
| 65 | 7. **Finish the GitHub release.** Pushing the tag builds binaries for macOS (arm64 and | |
| 66 | x86_64) and Linux (gnu and musl) and opens a *draft* release with them attached. Paste | |
| 62 | 7. *Finish the GitHub release.* Pushing the tag builds binaries for macOS (arm64 and | |
| 63 | x86_64) and Linux (gnu and musl) and opens a /draft/ release with them attached. Paste | |
| 67 | 64 | the changelog entry in and publish it. The draft is deliberate — a release that |
| 68 | 65 | publishes itself before anyone has read it cannot be edited quietly. |
| 69 | 66 | |
| 70 | ## If a release goes wrong | |
| 71 | ||
| 67 | ** If a release goes wrong | |
| 72 | 68 | Yank rather than delete, and ship a fix as a new version: |
| 73 | 69 | |
| 74 | ```bash | |
| 70 | #+begin_src sh | |
| 75 | 71 | cargo yank --version 0.18.0 |
| 76 | ``` | |
| 72 | #+end_src | |
| 77 | 73 | |
| 78 | 74 | Yanking stops new dependents from selecting it; anyone who already has it keeps working. |
| 79 | Then release `0.18.1` with the fix and a changelog entry that says what happened. | |
| 75 | Then release =0.18.1= with the fix and a changelog entry that says what happened. | |
SECURITY.md → SECURITY.org renamed +12 −16
| @@ -1,32 +1,28 @@ | ||
| 1 | # Security Policy | |
| 2 | ||
| 3 | ## Supported Versions | |
| 4 | ||
| 5 | | Version | Supported | | |
| 6 | |---------|-----------| | |
| 7 | | latest release | yes | | |
| 8 | | anything older | no | | |
| 1 | * Security Policy | |
| 2 | ** Supported Versions | |
| 3 | | Version | Supported | | |
| 4 | |—————-+———--| | |
| 5 | | latest release | yes | | |
| 6 | | anything older | no | | |
| 9 | 7 | |
| 10 | 8 | Fixes land in a new release rather than as patches to an old one. |
| 11 | 9 | |
| 12 | ## Reporting | |
| 13 | ||
| 14 | Email <hello@cleberg.net>, or open a private advisory through GitHub's *Security* tab. | |
| 10 | ** Reporting | |
| 11 | Email [[mailto:hello@cleberg.net][hello@cleberg.net]], or open a private advisory through GitHub's /Security/ tab. | |
| 15 | 12 | Please do not open a public issue for something exploitable. |
| 16 | 13 | |
| 17 | ## What is worth reporting | |
| 18 | ||
| 19 | orgo reads org files and writes HTML, so the interesting cases are about what a *document* | |
| 14 | ** What is worth reporting | |
| 15 | orgo reads org files and writes HTML, so the interesting cases are about what a /document/ | |
| 20 | 16 | can make it do: |
| 21 | 17 | |
| 22 | 18 | - Content from a source file escaping into HTML unescaped — a page that can inject script |
| 23 | 19 | into the site it is published on. |
| 24 | 20 | - A path in a document or config that writes outside the output directory. |
| 25 | - The `serve` development server reachable, or made reachable, beyond loopback, or serving | |
| 21 | - The =serve= development server reachable, or made reachable, beyond loopback, or serving | |
| 26 | 22 | files from outside the output directory. |
| 27 | 23 | - A crash, hang or unbounded allocation triggered by a crafted org file. A build that |
| 28 | 24 | refuses a file is fine; one that never finishes is not. |
| 29 | 25 | |
| 30 | Out of scope: `--strict` not catching something, an unhandled org construct rendering | |
| 26 | Out of scope: =--strict= not catching something, an unhandled org construct rendering | |
| 31 | 27 | oddly, and anything requiring you to run orgo against files you already do not trust while |
| 32 | 28 | also deploying the result unread. |