krz/orgo
Lightning fast org-mode static site generator.
clone: git clone https://gitbay.org/krz/orgo.git
8844e2b78830f4a83a77b6e3a56c5780731263c2
unsigned
author: Christian Cleberg <hello@cleberg.net> · 2026-08-11T22:36:50Z
.github/workflows/release.yml | 4 +- CHANGELOG.org | 156 --------- README.org | 765 ++++-------------------------------------- RELEASING.org | 45 +-- SECURITY.md | 2 +- docs/guide/11-versioning.org | 7 +- docs/install.org | 4 +- 7 files changed, 93 insertions(+), 890 deletions(-) @@ -77,7 +77,7 @@ jobs: staging="orgo-${{ github.event.inputs.tag || github.ref_name }}-${{ matrix.target }}" mkdir "$staging" cp "target/${{ matrix.target }}/release/orgo" "$staging/" - cp README.org LICENSE CHANGELOG.org "$staging/" + cp README.org LICENSE "$staging/" tar czf "$staging.tar.gz" "$staging" shasum -a 256 "$staging.tar.gz" > "$staging.tar.gz.sha256" @@ -101,7 +101,7 @@ jobs: with: merge-multiple: true - # A draft, deliberately. The changelog entry is written by a person, and a release + # A draft, deliberately. The release notes are written by a person, and a release # that publishes itself before anyone has read it cannot be edited quietly. - name: Create the draft release uses: softprops/action-gh-release@3d0d9888cb7fd7b750713d6e236d1fcb99157228 # v3.0.2 deleted file mode 100644 @@ -1,156 +0,0 @@ -* Changelog -What changed and why, newest first. Entries name the /behaviour/ that moved, since that is -what a rebuild will show you. - -Two conventions worth knowing before reading: - -- *A cache-format bump is not a change you need to act on.* The incremental cache is - versioned and discards itself; a bump means the next build re-renders everything once. -- *Output changes are called out.* orgo aims at what Emacs exports from the same - file, so an entry that says "now renders X" means your pages will change. That is the - product, not a regression — but it belongs in a changelog rather than a diff you find - later. - -Versions follow the compatibility promise in the README: config keys, template variables, -CLI flags and URLs are the stable surface. - -** 0.19.1 -- Footnote back-links carry =aria-label="Back to reference N"=, and the notes section is - labelled. A link whose only visible content is =↩= has that glyph as its whole - accessible name, so a screen reader announced "left arrow with hook" once per note with - no way to tell them apart. - -** 0.19.0 -- *Full-content collections.* =include_content = true= gives a listing template each - entry's rendered HTML as =entry.content= — a feed that carries whole posts rather than - excerpts. Rendered only when the listing is actually rebuilt, so a cached feed costs - nothing. -- *Fixed: a listing could show a stale excerpt.* Its cache key covered a hand-picked set - of fields, and the excerpt was not among them, so rewriting a post's first paragraph - left the old text on the index until something unrelated invalidated it. Entries are now - hashed through their serialization, which cannot drift from what a template can read. - Editing a post's body now rebuilds the listings that show it. -- =page.toc= entries carry =number=, so a site with section numbering on can number its - contents list to match its headings. - -** 0.18.0 -Release engineering, so that a version number is worth reading. - -- *A written compatibility promise.* Config keys, template variables, CLI flags and URLs - are the stable surface; the incremental cache, HTML details and the Rust API are not. - In the README, and in the guide under /Versioning and upgrades/. -- *CI* on Linux and macOS: build, test, clippy as an error, and the documentation site - built with =--strict=. Emacs is installed on both, so the oracle suite runs for real - instead of skipping. -- *A checked MSRV*, 1.88 — which is how it came to be 1.88 rather than the 1.82 - orgo's own code needs. The floor comes from dependencies, and nobody finds that out - by reasoning about it. -- *Release binaries* for macOS (arm64, x86_64) and Linux (gnu, musl), built on tag into - a draft release. The tag is checked against =Cargo.toml= before anything is built. -- A =LICENSE= file to go with the MIT declaration, crates.io metadata, and a release - profile that produces a 5.0 MB binary rather than 6.5 MB. -- This changelog, and =RELEASING.org=. - -** 0.17.0 -- *Asset directories outside the source.* =[build] assets = ["../theme/static"]= copies - a directory's contents to the site root. A site's static files do not always live where - its writing does, and copying them next to the writing is how a repository ends up with - two of every stylesheet. =watch= and =serve= watch these directories too. Two files - claiming one URL is a build error naming both. -- *Template hashing is per template.* A page's render key covered every template, so - editing a feed template re-rendered the whole site. It now covers the layout the page - uses plus what that layout extends, includes or imports. On a 196-page site, editing the - feed template renders one page instead of 196. -- Cache format 7. - -** 0.16.0 -- *Org's entity table.* =\alpha=, =\rarr=, =20\deg= and the other 412 names, generated - from Emacs' own =org-entities=. An unknown name stays literal; =#+OPTIONS: e:nil= turns - the table off. /Output changes/ for any page using entities. -- *Table captions.* =#+CAPTION:= above a table becomes a numbered =<caption>=. -- *=#+INCLUDE:= reports itself.* It was inert and silent, which publishes a page with - content missing and nobody told. Now a diagnostic, and =--strict= makes it a failure. -- The Emacs oracle separates deliberate divergence from defects. Every difference from - org's exporter is named and justified, and a test asserts there are no others. - -** 0.15.0 -Export parity, from a page-by-page diff of a 179-file corpus against the site Emacs -publishes from the same sources. *All of these change output.* - -- Heading levels are relative to a document's shallowest heading, as org exports them. -- Org's text conversions: =--=, =---=, =...=, and =x^2= / =a_{b}=. Never inside verbatim, - code, source blocks or LaTeX. =#+OPTIONS: -:nil=, =^:nil= and =^:{}= all work. -- Captioned figures are numbered =Figure N:=. -- A caption attaches to the element /directly/ below it; a blank line between attaches to - nothing. -- Checkboxes render as org writes them, which keeps the =[-]= partly-done state a disabled - =<input>= could not express. =[@4]= sets a list item's number. -- A table's special marker column and its marker rows stay out of the output. -- =#+BEGIN_NOTE= and any other unrecognised name is a special block: a div holding parsed - org rather than a =<pre>= of literal text. Verse keeps its line breaks. -- Emphasis borders forbid whitespace and nothing else, so =="proxied":false== is verbatim - and =~~/.config/doom/config.el~= is a path that starts with a tilde. -- Listings sort on the time of day when a timestamp carries one. -- Cache format 6. - -** 0.14.0 -- *Per-page layouts.* =[[pages]]= rules map a source path to a template, and - =#+TEMPLATE:= on a page overrides any rule. A missing template fails the build naming - the page, the template, and what does exist. -- =page.year=, for grouping a listing by year with minijinja's =groupby=. -- An explicit nav can order generated pages among authored ones. =nav.mode = "none"= now - really means none. - -** 0.13.0 -- Bundled TOML and Org syntax definitions, a =syntaxes_dir= for your own, and org's comma - escape (=,* heading= inside a block). - -** 0.12.0 -- =serve=: a development server with live reload, bound to loopback. -- A documentation site under =docs/=, built by orgo itself. - -** 0.11.0 -- Table of contents as =page.toc=, section numbers, and org's =#+OPTIONS:= per-file - switches. - -** 0.10.0 -- Excerpts, word count, reading time, a =truncate= filter, and =#+DRAFT:= pages. - -** 0.9.0 -- =watch=: rebuilds on OS filesystem events, debounced. - -** 0.8.0 -- =site.base_url=, the =absolute= and =rfc822= filters, canonical links, and an RSS feed - in the scaffold that validates. - -** 0.7.0 -- Pagination for large listings, with a =paginator= template context that composes with - grouping. - -** 0.6.0 -- Grouped collections: one page per tag plus a tag index. -- Generated listing pages (=[[collections]]=), sorted indexes, and feeds via XML - templates. -- A config file, user templates, nav modes, an =init= scaffold, and discovery that will - not publish =.git=. - -** 0.5.0 -- Parse diagnostics carry =file:line=, and pages render in parallel. -- The corpus audit (=orgo audit=) and the =emacs --batch= oracle. -- =#+SLUG:= decides a page's output filename — found by auditing a real corpus, where it - affected 169 of 182 URLs. - -** 0.4.0 -- The full v1 construct scope, with the IN/OUT line under test. - -** 0.3.0 -- The incremental build layer: content, config and template hashing, a dependency graph, - per-page render keys, and a persisted cache manifest. A full build and an incremental - build produce byte-identical output. - -** 0.2.0 -- Multi-file site builds: a symbol table, internal link resolution, minijinja templates, - tables and footnotes. - -** 0.1.0 -- Parse and render a single =.org= file to HTML. @@ -1,726 +1,107 @@ * orgo -An org-mode static site generator, in Rust. Org is treated as the /source language/, -not an inconvenient input to be normalized into markdown. The org element tree — -headings, drawers, blocks, links with their org-specific semantics — *is* the -document model, and we render that tree straight to HTML. We never round-trip through -a markdown-shaped intermediate representation, because the point is to preserve what -markdown cannot express: property drawers, TODO/priority/tag metadata on headings, -=#+= directives, ID links, named/captioned blocks, footnote semantics. -The one non-obvious early commitment is *incremental builds keyed on content -hashing*, treated as a first-class architectural concern from day one. The discipline -it imposes on the data model — pure, hashable, dependency-tracked units — is the real -deliverable, even while the corpus is small enough that a full rebuild is instant. +Turn a folder of org files into a website. -*Full documentation is in [[file:docs/][=docs/=]]* — a site written in org and built by -orgo itself. Build and read it with: +You write posts the way you already do — an =.org= file per page, in whatever directory +structure suits you — and orgo builds a complete site from them: pages, navigation, a blog +index, tags, an RSS feed, syntax-highlighted code. It is one binary with nothing to +install alongside it, and *you do not need Emacs to build your site*, only to write in a +format Emacs made. -#+begin_src sh -cargo run -- serve docs -o docs/_site -#+end_src - -** Quick start -#+begin_src sh -cargo run -- init my-site # config + an editable copy of the layout + a page -cargo run -- build my-site -o _site -#+end_src - -Or skip the scaffolding entirely — point it at any directory of =.org= files: - -#+begin_src sh -cargo run -- build ~/notes -o _site -#+end_src - -*Zero configuration is a supported path, not a demo.* With no =orgo.toml=, no -templates and no orgo-specific markup in your files, you get a complete site: pages, -navigation, syntax-highlighted code and the stylesheet to colour it. Configuration -changes what you get; it is never what makes it work. - -Discovery skips what should not be published — dot-directories such as =.git=, the config -file, the templates directory, and the output directory when it sits inside the source, so -=orgo build . -o _site= does the obvious thing. - -** Configuration -Everything is optional. =orgo init= writes a fully commented =orgo.toml=; every -value below is the default. - -#+begin_src toml -[site] -title = "orgo site" -base_url = "" # absolute URL, no trailing slash; needed for feeds/canonical links -description = "" -language = "en" - -[nav] -mode = "top-level" # top-level | all | explicit | none -# pages = ["index.org", "about.org"] # for mode = "explicit"; order is preserved - -[templates] -dir = "templates" # base.html replaces the built-in layout -expose_page_list = false - -# [[pages]] # which layout a section renders through; base.html by default -# match = "blog" # a source directory or one .org file; most specific rule wins -# template = "post.html" - -[highlight] -theme = "InspiredGitHub" - -[build] -drafts = false -assets = [] # extra directories copied to the site root, e.g. ["../theme/static"] - -[html] -heading_offset = 1 # a level-1 org heading becomes <h2>, beneath the layout's <h1> -#+end_src - -*** Templates -Drop a =base.html= into the templates directory and it replaces the built-in layout -entirely. Any other =.html= file there is available to ={% include %}= and -={% extends %}=. Templates are [[https://docs.rs/minijinja][minijinja]] (Jinja2 syntax) and -receive: - -| Variable | What it is | -|————--+————————————————————————————————————————————————--| -| =body= | the rendered page HTML — use ={{ body \| safe }}= | -| =page= | =.title=, =.url=, =.source=, =.date=, =.date_iso=, =.year=, =.tags=, =.content=, =.excerpt=, =.word_count=, =.reading_time=, =.toc=, =.keywords= | -| =site= | =.title=, =.base_url=, =.description=, =.language= | -| =nav= | list of ={title, url}=, relative to this page | -| =root= | =../=-prefix back to the site root from this page | -| =stylesheet= | URL of the generated =syntax.css= | -| =pages= | every page's metadata — only when =expose_page_list = true= | - -=page.keywords= carries *every* =#+KEYWORD:= in the file under its lowercased name, so -your own metadata works without this crate knowing about it: =#+CUSTOM_THING: x= is -={{ page.keywords.custom_thing }}=. - -=base.html= is the default layout, not the only one. A =[[pages]]= rule gives a section -its own — =match = "blog"=, =template = "post.html"= — and =#+TEMPLATE: wide.html= gives -one page its own, which wins over any rule. A second layout usually starts with -={% extends "base.html" %}=. - -Editing a template re-renders the pages that use it — template sources are a hash input, -so a design change never leaves a site half-updated. - -*** Generated listing pages -A blog index, an archive, a feed — output files with no source =.org= behind them. -Repeat the block for each one: - -#+begin_src toml -[[collections]] -source = "blog" # directory to list; empty means every page -output = "blog/index.html" # where to write it -template = "list.html" -title = "Blog" -sort = "date" # date | title | path -order = "desc" # desc | asc -nav = true # put this listing page in the nav -#+end_src - -The template gets the collection's entries as =pages=, already sorted, plus the usual -=site=/=nav=/=root=. It can ={% extends "base.html" %}= to inherit the site chrome: - -#+begin_src jinja -{% extends "base.html" %} -{% block main %} -<ul>{% for p in pages %} - <li><time datetime="{{ p.date_iso }}">{{ p.date_iso }}</time> - <a href="{{ root }}{{ p.url }}">{{ p.title }}</a></li> -{% endfor %}</ul> -{% endblock %} -#+end_src - -=p.date_iso= is the =YYYY-MM-DD= extracted from =#+DATE:=, whatever org syntax it was -written in — =[2025-09-05 Fri 10:21:00]=, =<2024-05-01 Wed>= or bare =2024-05-01=. It is -also the sort key; pages without a parseable date sort last, so an undated draft never -leads a dated archive. - -**** Pagination -Set =paginate= to split a long listing across numbered pages: - -#+begin_src toml -[[collections]] -source = "blog" -output = "blog/index.html" -paginate = 10 -paginate_output = "blog/page/{n}.html" # {n} is the 1-based page number -#+end_src - -Page 1 stays at =output=, so a section's canonical URL never moves as its page count -changes; only pages 2..N are named by =paginate_output=. The template gets a =paginator=: - -#+begin_src jinja -{% if paginator and paginator.total > 1 %} -<nav> - {% if paginator.prev_url %}<a href="{{ paginator.prev_url }}">Newer</a>{% endif %} - {% for pg in paginator.pages %} - <a href="{{ pg.url }}"{% if pg.current %} aria-current="page"{% endif %}>{{ pg.number }}</a> - {% endfor %} - {% if paginator.next_url %}<a href="{{ paginator.next_url }}">Older</a>{% endif %} -</nav> -{% endif %} -#+end_src - -=paginator= carries =current=, =total=, =per_page=, =total_entries=, =prev_url=, -=next_url=, =first_url=, =last_url=, and =pages=. Every URL is relative to the page -carrying it, so links work from page 1 (=page/2.html=) and from page 5 (=../index.html=, -=6.html=) without the template knowing where it sits. An unpaginated collection has no -=paginator= at all, so ={% if paginator %}= is a reliable test in a shared template. - -Grouping and pagination compose: each group paginates independently, which is why -=paginate_output= needs ={tag}= as well as ={n}= on a grouped collection. An empty -collection still emits page 1 — a section that exists but has nothing in it should say so -rather than 404. When the entry count shrinks, pages that no longer exist are deleted -instead of being left serving stale posts. - -**** Tag pages -Add =group_by= and the collection emits one page /per group/ instead of one page total, -plus an optional index of the groups: - -#+begin_src toml -[[collections]] -source = "blog" -group_by = "tags" # "tags", or any #+KEYWORD: name to group by its value -output = "tags/{tag}.html" # {tag} is replaced by each group's slug -template = "tag.html" -title = "Tagged: {tag}" -index_output = "tags/index.html" # the tag index -index_template = "tags.html" -index_title = "Tags" -nav = true # adds the *index*, not every tag -#+end_src - -A group page receives its own posts as =pages= and itself as =group= -(=.name=, =.slug=, =.url=, =.count=). The index receives =groups= — every group, sorted -by name: - -#+begin_src jinja -<ul>{% for tag in groups %} - <li><a href="{{ root }}{{ tag.url }}">{{ tag.name }}</a> ({{ tag.count }})</li> -{% endfor %}</ul> -#+end_src - -=group_by = "tags"= is multi-valued: a post appears under every tag it carries. Any other -value names a single-valued =#+KEYWORD:=, so =group_by = "category"= buckets by -=#+CATEGORY:=. - -Two tags that would produce the same URL (=web_dev= and =web@dev= both slugify to -=web-dev=) are a build error rather than one page silently overwriting the other. - -A tag page depends on its own posts and nothing else, so adding a post tagged =rust= -re-renders that post, its section index, =tags/rust.html=, and the tag index whose counts -changed — four pages, not one per tag. That precision is why =groups= is given to the -index and not to every group page: a page that can see every group depends on every -group. - -**** Feeds and absolute URLs -*A feed is a listing page with an XML template*, not a separate feature — templates are -loaded by full filename and any extension, so =output = "feed.xml"= with -=template = "feed.xml"= is all it takes. =orgo init= writes a working RSS template. - -A feed is read away from the site that served it, so relative links in one are simply -broken. Set =site.base_url= and use the =absolute= filter: - -#+begin_src jinja -<link>{{ post.url | absolute }}</link> -<pubDate>{{ post.date_iso | rfc822 }}</pubDate> -#+end_src +Org is the source language here, not something to convert away from first. Tools that +route org through markdown lose what markdown has no words for — property drawers, a +heading's TODO state and tags, =#+= keywords, ID links, captions on images. orgo keeps all +of it, and its output is checked page by page against what Emacs' own exporter produces +from the same file. -| Filter | Does | -|—————+—————————————————————————-| -| =absolute= | site-root-relative path → absolute URL; already-absolute URLs pass through | -| =rfc822= | any org or ISO date → the format RSS =pubDate= requires | -| =truncate(n)= | shorten to at most =n= characters on a word boundary, with an ellipsis | +*Documentation: https://ccleberg.github.io/orgo/* — that site is written in org and built +by orgo, so it doubles as the longest worked example available. -Apply =absolute= to the site-root-relative values — =page.url=, =pages[].url=, -=group.url= — and not to =nav[].url=, =paginator.*_url=, =stylesheet= or =root=, which -are relative to the page carrying them and already correct there. +** Install -With no =base_url=, =absolute= is an *error* naming the setting, rather than quietly -emitting a relative URL that would make the feed invalid everywhere while looking fine. -The default layout also emits =<link rel="canonical">= when a base URL is set. +You need [[https://rustup.rs][Rust]] (1.88 or newer). Nothing else — syntax highlighting +and its themes are compiled in. -Listing pages are cached on the entries they list, so adding a post re-renders that -section's index and nothing else. - -*** Table of contents and =#+OPTIONS:= -=page.toc= is the page's headings as a *tree* — ={title, anchor, level, children}= — -because a table of contents is one, and rebuilding a tree from a flat list of levels -inside a template is what Jinja is worst at. Its anchors come from the same function the -renderer uses to emit heading =id=s, so a TOC link cannot drift from the heading it -points at. - -#+begin_src jinja -{% macro toc_list(entries) %} -<ul>{% for e in entries %} - <li><a href="#{{ e.anchor }}">{{ e.title }}</a> - {%- if e.children %}{{ toc_list(e.children) }}{% endif %}</li> -{% endfor %}</ul> -{% endmacro %} -{% if page.toc %}{{ toc_list(page.toc) }}{% endif %} +#+begin_src sh +git clone https://github.com/ccleberg/orgo +cd orgo +cargo install --path . #+end_src -Org's own per-file export switches are honoured, so a document can turn a feature off for -itself the way its author already knows: - -| Switch | Effect | Site default | -|———————-+————————————+———————————-| -| =#+OPTIONS: toc:nil= | empties =page.toc= for this page | =[html] toc = true= | -| =#+OPTIONS: num:t= | numbers headings =1.=, =1.1.=, … | =[html] section_numbers = false= | - -*Section numbers default to off, which differs from Emacs on purpose.* -=org-export-with-section-numbers= is on there, so an org-published site inherits numbered -headings whether or not anyone chose them. Most sites do not want them; =num:t= or -=section_numbers = true= gets Emacs' behaviour back, with Emacs' own -=section-number-N= classes so the output stays diffable against the oracle. - -*** Excerpts and drafts -=page.excerpt= is a page's =#+DESCRIPTION:= when it sets one and its first paragraph -otherwise, so a listing has something to show whether or not the author thought about -summaries. =page.word_count= and =page.reading_time= (minutes at 200 wpm) count prose -only — a post that is mostly a shell transcript should not read as an hour's work. -=truncate= exists because an excerpt is usually a whole paragraph and minijinja has no -such filter. - -=#+DRAFT:= keeps a page out of the build entirely — no page, and absent from listings and -the nav rather than merely unlinked. =--drafts= includes them, which is what you want -under =watch= while writing one. A draft is out of the symbol table too, so a link /to/ -one is reported as the dead link it would be once published. - -The keyword is read forgivingly: =t=, =yes=, =1= and a bare =#+DRAFT:= all mean draft, -because writing the keyword at all is the signal. Only an explicit =nil=, =false=, =no=, -=0= or =off= means published. - -*** =#+SLUG:= -A page's output filename comes from its =#+SLUG:= when it has one, so -=2018-11-28-aes-encryption.org= can publish as =aes-encryption.html=. Without one the -source filename is used. Slugs are sanitized to a single safe path component, and two -pages claiming one URL is a build error rather than a silently dropped page. - -** Pipeline -#+begin_example -DISCOVER → PARSE → INDEX → RESOLVE → RENDER → TEMPLATE → EMIT -#+end_example - -PARSE and RENDER are pure functions of their inputs (cacheable, hashable). INDEX/RESOLVE -is the only inherently global stage — it is where the link dependency graph is born. - -| Stage | Module | Notes | -|————-+———————-+———————————————————————————-| -| config | =src/config.rs= | =orgo.toml=: site metadata, nav mode, templates, theme. A hash input. | -| PARSE | =src/parser.rs= | Hand-written recursive descent: line lexer → element builder → inline tokenizer. | -| audit | =src/audit.rs= | Phase 0 corpus audit: construct frequencies against the IN/OUT line. | -| model | =src/model.rs= | The org element tree — Elements (block) vs Objects (inline). | -| INDEX | =src/index.rs= | Collect link targets into a symbol table. | -| RESOLVE | =src/resolve.rs= | Rewrite links to URLs; return the used-target list (dependency edges). | -| RENDER | =src/render.rs= | Tree → HTML fragment; syntect highlighting; footnote two-pass. | -| TEMPLATE | =src/template.rs= | minijinja: fragment + metadata → full page. | -| incremental | =src/incremental.rs= | Content/config/template hashing, dep graph, cache manifest, invalidation. | - -** v1 scope (delivered as of v0.4; still to be reconciled against a corpus audit) -*IN — v1 must handle:* headings with nesting, at levels relative to the document's -shallowest; TODO keywords; priorities =[#A]=; tags; property drawers; plain lists -(unordered/ordered/description, checkboxes, =[@N]= counters, nesting); tables (with rule -rows and org's special marker column, no =#+TBLFM:=); source blocks with syntax -highlighting; example/quote/center/verse blocks and named special blocks; links (external, -internal =[[*Heading]]=/=[[#custom-id]]=, =id:=); footnotes (inline and referenced); =#+= -keywords/directives; inline markup (bold/italic/underline/verbatim/code/strike); org's -export-time text conversions (=--=/=---=/=...=, =x^2=, =a_{b}=, =\alpha=); timestamps -(active/inactive, ranges); paragraphs and horizontal rules; images with -=#+CAPTION=/=#+ATTR_HTML=, numbered =Figure N:=. +That puts an =orgo= command on your =PATH=. Full notes, including how to run it without +installing anything: https://ccleberg.github.io/orgo/install.html -*OUT — explicitly not v1 (parse-and-ignore or reject loudly):* Babel execution / -=:results=; =#+TBLFM:= formulas; LaTeX / MathJax (passed through untouched, including past -the text conversions); =#+INCLUDE:= (never expanded — reported as a diagnostic, so a page -is never quietly short of content); citations; radio targets and macros; drawers other -than PROPERTIES/LOGBOOK; column view / clocking / agenda semantics; non-HTML export -blocks. +** Your first site -*Scope guardrail:* every IN item gets a golden-file fixture; every OUT item gets a test -asserting it degrades predictably (ignored, no crash). The IN/OUT line is enforced by -=tests/constructs.rs=, defending against the project's #1 risk: scope creep back toward -all-of-org. Phase 0 checked this line against a real 179-file corpus and found it sound -(99.9% of construct uses in scope) — but also found one thing missing from it entirely: -=#+SLUG:=. See [[#phase-0-the-corpus-audit-and-the-emacs-oracle][Phase 0]]. - -** Phase plan -| Phase | Scope | Status | -|——--+——————————————————————————————————————————————————————————+——--| -| *M0* | *Buildable skeleton: crate layout, module stubs, deps, test harness, fixtures* | *done* | -| *v0.1* | *End-to-end core parse → render: =build= a single =.org= file to HTML* | *done* | -| *v0.2* | *Multi-file SITE build: INDEX + RESOLVE internal links, minijinja templates, =build <src-dir> <out-dir>=, tables + footnotes* | *done* | -| *v0.3* | *Incremental build layer: content/config/template hashing, dependency graph, per-page render keys, persisted cache manifest, invalidation* | *done* | -| *v0.4* | *MVP: the full v1 construct scope — heading metadata, nested/description lists, block types, timestamps, images, syntect highlighting — with the IN/OUT line under test* | *done* | -| *0* | *Corpus audit + =emacs --batch= ground-truth oracle* | *done* | -| 1 | Line lexer + heading/section skeleton | done | -| 2 | Block elements — lists, source blocks, tables, footnote defs, blocks by type, drawers | done | -| 3 | Inline objects — emphasis, links, bare URLs, footnote refs, timestamps | done | -| 4 | Rendering to HTML — tree walk, tables, footnote two-pass, minijinja templating, syntect highlighting | done | -| 5 | Link resolution + symbol table (INDEX + RESOLVE, used-target list, broken-link reporting) | done | -| 6 | Incremental build layer (hashing, dep graph, invalidation); =watch= on OS filesystem events | done | -| *7* | *Hardening: rayon parallelism, error locations in parse diagnostics* | *done* | -| *8* | *General use: config file, user templates, nav modes, =init= scaffold, safe discovery* | *done* | -| *9* | *Generated listing pages: =[[collections]]=, sorted indexes, feeds via XML templates* | *done* | -| *10* | *Grouped collections: one page per tag plus a tag index — full parity with the incumbent* | *done* | -| *11* | *Pagination: numbered pages with a =paginator= context, composing with grouping* | *done* | -| *12* | *=base_url=: =absolute=/=rfc822= filters, a valid RSS feed in the scaffold, canonical links* | *done* | -| *13* | *=watch= on OS filesystem events, debounced, with the feedback loop closed* | *done* | -| *14* | *Authoring: excerpts, word count, reading time, =truncate=, and draft pages* | *done* | -| *15* | *Table of contents, section numbers, and org's =#+OPTIONS:= per-file switches* | *done* | -| *16* | *=serve=: development server with long-poll live reload, loopback-bound* | *done* | -| *17* | *Bundled TOML and Org syntaxes, a user syntax directory, and org's comma escape* | *done* | -| *18* | *Per-page layouts: =[[pages]]= rules and =#+TEMPLATE:=* | *done* | -| *19* | *Export parity: relative heading levels, special strings, sub/superscript, caption numbering, checkbox and counter markup, table marker columns, special blocks* | *done* | -| *20* | *Correctness debt: org's entity table, table captions, a reported =#+INCLUDE:=, and an oracle that separates deliberate divergence from defects* | *done* | -| *21* | *Extra asset roots; per-template hashing so one layout edit does not re-render the site* | *done* | -| *22* | *Release engineering: CI on both platforms, a checked MSRV, release binaries, a changelog, and a written compatibility promise* | *done* | - -*** v0.2 in / out -*Added in v0.2:* the INDEX stage (=SymbolTable= of =:ID:=/=:CUSTOM_ID:=/heading/=file:= -targets across a directory); the RESOLVE stage — rewrites =[[#custom-id]]=, =[[id:...]]=, -=[[*Heading]]= and =[[file:other.org]]= links to real relative output URLs, returns the -=used_targets= list (the =uses= edges, spec §4.3/R2) and reports unresolved links as -warnings rather than crashing; a minijinja base layout (title, nav, body) applied to every -page; a =build <src-dir> <out-dir>= path that walks the tree, parses + resolves + renders + -templates every =.org= into a linked static site and copies non-=.org= assets through; -plus two new constructs — pipe *tables* (with header band from the rule row) and -*footnotes* (block =[fn:1]= definitions, referenced =[fn:1]=, and inline =[fn:1:text]=, -rendered as a numbered, back-linked notes section). - -*Left stubbed at v0.2, all closed in v0.4:* timestamps; TODO keywords and priorities; -generic (non-PROPERTIES) drawers; real syntect tokenizing behind the =Highlighter= trait. - -*** v0.3 in / out -*Added in v0.3 — the incremental build layer (spec §4, the flagship, non-retrofittable -feature):* - -- *Three hash classes (spec §4.1)* in =src/incremental.rs=: a *content hash* (blake3 - of a file's bytes), a *config hash* (blake3 of the resolved =BuildConfig=), and a - *template hash* (blake3 of the template sources). A change in any one invalidates the - pages it affects. -- *Dependency graph (spec §4.3)* built from RESOLVE's =defines=/=uses= edges: a page - depends on the targets it links to, so editing (or renaming a heading in) a file - invalidates the pages that /link into/ it, not just the file itself — the load-bearing - R2 invariant. On rebuild the graph is merged with the previous build's =defines= so a - /removed/ target still pulls in its linkers. -- *Per-page =render_key=* = =H(content ⊕ resolved-links ⊕ config ⊕ template)=. If a - page's render key is unchanged, its on-disk output is already correct and it is skipped. - The config component folds in a *site-structure hash* (every page's =(path, title)=), - because the shared nav bar is global chrome — a title change or a page add/remove alters - the nav on every page and so must re-render them all (otherwise byte-equivalence breaks). -- *Persisted cache manifest* (=<out>/.orgo-cache.json=, JSON), carrying per-page - records, the config/template hashes, and the serialized dependency graph, tagged with - =CACHE_FORMAT_VERSION=. A version mismatch, a missing file, or a corrupt file all fall - back to a clean full rebuild — the cache is an optimization, never a correctness - dependency. -- *Wired into =build_site=*: only pages whose render key changed (or that link into a - changed file's targets) are re-rendered; unchanged outputs are left in place. =--no-cache= - forces a full rebuild; =clean <out-dir>= removes the output directory (and its cache). - =SiteReport= now reports =rendered= vs =skipped= counts. - -The hard gates are enforced by =tests/incremental.rs=: full-vs-incremental *byte -equivalence* (and a second unchanged build re-rendering *zero* pages); *edit-one-file* -re-renders exactly the changed page plus its linkers; *renamed-heading* invalidates the -linking page and updates its emitted anchor; and cache *version-bump / missing / corrupt* -all fall back to a full rebuild. - -*Out of scope in v0.3:* real syntect highlighting; timestamps and TODO keywords (all -landed in v0.4). =watch= is a minimal mtime poll loop (=watch <src-dir> -o <out-dir>=), not -an OS file-watcher — the fs-notify integration is deferred. The parse-tree cache (spec §4.5, -"optionally") is not persisted: PARSE/INDEX/RESOLVE run for every file each build (cheap and -pure); the incremental win is on RENDER + EMIT. - -*** v0.4 in / out — the MVP -v0.4 closes the gap between the v1 scope above and what the code actually did, so every -construct the IN list claims is now parsed, rendered, and pinned by a golden file: - -- *Heading metadata* — TODO keywords (the Emacs default =TODO=/=DONE= set, matched on a - word boundary so =TODOs= is not one) and =[#A]= priority cookies, rendered with Emacs' - own export classes so the output stays diffable against an =emacs --batch= oracle. -- *Lists* — indentation-based nesting (a sub-list renders /inside/ its parent =<li>=), - multi-paragraph item bodies, and =term :: definition= description lists as =<dl>=. -- *Blocks by type* — =QUOTE=, =CENTER=, =EXAMPLE=, =EXPORT= and =SRC= are now distinct - elements rather than all collapsing to a verbatim example block. Block matching is on the - specific kind, so a source block can nest inside a quote. An =html= export block passes - through; every other backend drops. -- *Timestamps* — active =<...>= and inactive =[...]=, optional times, same-day time - ranges and =--=-joined date ranges, rendered as =<time>= with a machine-readable - =datetime=. Repeater/warning cookies are recognized and discarded. -- *Images* — a description-less link to an image file renders as =<img>=; with an - affiliated =#+CAPTION:=/=#+ATTR_HTML:= it is promoted to a =<figure>= with the caption as - both =<figcaption>= and alt text. Links to non-=.org= files are now understood as asset - links: neither resolved nor reported as broken. -- *Syntax highlighting* — real syntect tokenizing to CSS classes (never inline styles, so - themes live in the stylesheet). Every build emits the matching =syntax.css= and each page - links it relative to its own depth. An unknown language degrades to escaped =<pre><code>=. -- *Diagnostics* — broken links are reported as the org syntax the author wrote - (=warning: b.org: unresolved link [[#setup]]=) rather than a Debug-printed enum. - -*The OUT line is now enforced, not just asserted.* =tests/constructs.rs= pins each -excluded construct to a specific degradation: babel is never executed /and/ a checked-in -=#+RESULTS:= block is dropped rather than published as if it were verified output; -=#+TBLFM:= is inert; =#+INCLUDE:= is never expanded and says so; LaTeX, macros and radio targets survive -as literal text; drawers other than PROPERTIES are captured and dropped; unmodelled block -types keep their content verbatim. - -*Still out:* =#+TODO:= per-file keyword sequences; planning lines -(=SCHEDULED:=/=DEADLINE:=), which render as ordinary paragraphs; and fixed-width =:= -lines. - -** Serving #+begin_src sh -cargo run -- serve my-site -o _site # http://127.0.0.1:3000 +orgo init my-site +orgo serve my-site -o _site #+end_src -Builds, watches, serves, and reloads the browser when a rebuild lands — the loop =watch= -leaves half-open. +Open http://127.0.0.1:3000. Edit =my-site/index.org=, save, and the page reloads on its +own — that is the loop you will spend your time in. -- *Loopback by default.* A dev server serves unreviewed drafts off your laptop, so - reaching the local network is something you ask for with =--host 0.0.0.0=, never - something you get. -- *The reload script is injected on the way out*, never written to disk. What you - deploy is the built site, and it must not carry a dev server's JavaScript. -- *Long-polling, not WebSockets or SSE.* The browser asks "anything since generation - N?" and the server holds the request until there is. Instant like a push, no protocol - beyond ordinary HTTP, and no dependency. A streamed response would have been more - elegant and does not work: tiny_http buffers a response until its body ends, so a body - that never ends never reaches the client. -- A reload only follows a *successful* rebuild. Reloading onto a stale page because the - build just failed tells you nothing; the error is already on your terminal. +=init= writes a starter post, a page layout you can edit, and a config file with every +setting explained in comments. It never overwrites a file you already have. -URL resolution is the server's security boundary and is written as a pure function with -its own tests: =..=, percent-encoded =..=, backslashes, absolute paths and embedded NULs -all resolve to nothing rather than to somewhere outside the output directory. +** Or point it at writing you already have -** Watching #+begin_src sh -cargo run -- watch my-site -o _site +orgo build ~/notes -o _site #+end_src -Rebuilds on OS filesystem events rather than polling, so it costs nothing while nothing -happens. Write bursts are debounced — an editor saving a file writes a temp file, renames -it over the original and touches the directory, which is one edit and several events. - -Two rules decide what counts as a change, and they are not the same rules the build uses -to find content: +No config file, no templates, no orgo-specific markup in your files. You get a real site: +every page, links between them resolved, navigation across the top, code highlighted. That +is a supported way to use it rather than a demo — configuration changes what you get, it +is never what makes it work. -- *A build input is a change.* Editing =orgo.toml= or a template rebuilds, even - though discovery skips both as non-content. The question is "would this change the - site?", not "is this a page?". -- *Our own output is not.* =watch . -o _site= puts the output inside the source, so a - rebuild's writes raise events that would trigger a rebuild, forever. Dot-directories go - the same way — =.git= churns on every command — as do editor scratch files, including - Emacs' =file.org~= backups, which do not start with a dot. +Nothing that should stay private is published: dot-directories like =.git=, your templates +and the output folder itself are all skipped. -Where native watching is unavailable (some network and container filesystems), it falls -back to polling and says so, rather than failing. +Want to know what orgo will make of your files before trusting it with them? +=orgo audit ~/notes= reports which org constructs you use and how each one lands, with +counts and line numbers — never the text of your writing, so the report is safe to share. -** Phase 0: the corpus audit and the Emacs oracle -The v1 scope was, by its own admission, /recommended/ — a guess about which slice of org -matters. Phase 0 replaces both halves of that guess with a measurement: an audit that asks -what a real corpus actually uses, and an oracle that asks whether we render it the way -Emacs does. +** What you can add when you want it -The audit runs against any corpus — point it at your own notes before trusting this tool -with them. The numbers below come from a 179-file site published today by weblorg, a -wrapper around org's own HTML exporter, which makes it both a realistic workload and a -directly comparable incumbent. With collections configured, orgo now reproduces -*all 182 of that site's URLs*. +Each of these is a few lines of config, and each has a page in the guide: -#+begin_example -cargo run -- audit <src-dir> # what does this corpus use, and is it in scope? -cargo test --test oracle # how does our HTML differ from Emacs' own export? -#+end_example +| A blog index, newest first | [[https://ccleberg.github.io/orgo/guide/03-collections.html][Collections]] | +| Tag pages, and an index of tags | [[https://ccleberg.github.io/orgo/guide/03-collections.html][Collections]] | +| An RSS feed | [[https://ccleberg.github.io/orgo/guide/03-collections.html][Collections]] | +| Numbered pages when a list gets long | [[https://ccleberg.github.io/orgo/guide/03-collections.html][Collections]] | +| Your own design, in ordinary HTML templates | [[https://ccleberg.github.io/orgo/guide/04-templates.html][Templates]] | +| Drafts that stay unpublished until you say so | [[https://ccleberg.github.io/orgo/guide/06-authoring.html][Authoring]] | +| A table of contents on long posts | [[https://ccleberg.github.io/orgo/guide/06-authoring.html][Authoring]] | +| Clean URLs that survive a renamed file | [[https://ccleberg.github.io/orgo/guide/06-authoring.html][Authoring]] | -*** What the audit found -*The scope guess was sound.* 99.9% of construct uses in the corpus are in scope. The -whole out-of-scope tail is 8 uses: four =#+TBLFM:= in a post /about/ org-mode, three -=\name= entities, and one =#+BEGIN_NOTE=. +Rebuilds only touch the pages that actually changed, so saving a post on a site with +hundreds of them stays instant. -*=#+SLUG:= was a hole big enough to sink the project.* 178 of 179 files set it, and the -published URL comes from it, not from the filename: =2018-11-28-aes-encryption.org= is -served at =blog/aes-encryption.html=. orgo derived output paths from source filenames, -so *169 of 179 pages would have been published at the wrong URL* — every inbound link and -every search result, broken, by a tool that reported a clean build. Output paths now come -from =#+SLUG:= when present ([[file:src/util.rs][=util::output_path=]]); slugs are sanitized so an -author-supplied =../../etc/x= cannot escape the output directory, and two pages claiming one -URL is a build error rather than a silently dropped page. Building the real corpus now -reproduces all 179 of the live site's URLs exactly. +** The documentation -*Some machinery is speculative.* The corpus contains no =id:=, =#custom-id= or =*Heading= -links at all — its cross-page links are hand-written relative URLs. The INDEX/RESOLVE -symbol table that v0.2 was built around is, against this corpus, unexercised. +https://ccleberg.github.io/orgo/ -*An audit can lie too.* The first run reported 23 uses of a custom TODO keyword sequence. -All 23 were false: the detector read the leading word of =* CSS Variables= as the keyword -=CSS=. The corpus defines no =#+TODO:= sequences at all, so the true count was zero. The -detector now matches conventional keyword names only — a tool that overstates a gap argues -for work nobody needs. +| [[https://ccleberg.github.io/orgo/quickstart.html][Quick start]] | A working site in two commands, then your own writing, then your own design. | +| [[https://ccleberg.github.io/orgo/install.html][Install]] | Getting the binary, and running it without installing anything. | +| [[https://ccleberg.github.io/orgo/guide/01-cli.html][Commands]] | Every command and flag, and what each is for. | +| [[https://ccleberg.github.io/orgo/guide/02-configuration.html][Configuration]] | Every setting in =orgo.toml=, what it changes, and what it costs. | +| [[https://ccleberg.github.io/orgo/guide/03-collections.html][Collections]] | Blog indexes, tag pages, pagination and RSS feeds. | +| [[https://ccleberg.github.io/orgo/guide/04-templates.html][Templates]] | Layouts, and every variable a template can use. | +| [[https://ccleberg.github.io/orgo/guide/05-org-support.html][Org support]] | Which org syntax is handled, which is not, and how the rest degrades. | +| [[https://ccleberg.github.io/orgo/guide/06-authoring.html][Authoring]] | URLs, drafts, excerpts, tables of contents. | +| [[https://ccleberg.github.io/orgo/guide/07-incremental.html][Incremental builds]] | How it decides what to rebuild. | +| [[https://ccleberg.github.io/orgo/guide/08-workflow.html][Watching and serving]] | The write-save-see loop. | +| [[https://ccleberg.github.io/orgo/guide/09-auditing.html][Auditing]] | Reading a corpus before trusting a tool with it. | +| [[https://ccleberg.github.io/orgo/guide/10-deploying.html][Deploying]] | Producing a production build, and putting it somewhere. | -*** What the oracle found -=tests/oracle.rs= exports each fixture with org's own exporter via =emacs --batch=, reduces -both sides to a semantic skeleton (element opens, closes and text, with layout =div=s, -inline =span=s and all attributes but =href=/=src= dropped), and *snapshots the -disagreement*. Snapshotting rather than asserting is deliberate: a checked-in divergence -report gets reviewed and shows up as a diff, where a permanently red test gets ignored. -Three invariants are asserted outright, and all three hold — heading structure, list -nesting, and source-block text match Emacs exactly. +** Building from a checkout -*No bugs in orgo.* Every remaining divergence is a deliberate choice to emit better -HTML than org does: - -| | orgo | Emacs | why | -|—————--+—————————+—————————-+—————————————| -| emphasis | =<em>=/=<strong>= | =<i>=/=<b>= | semantic, not presentational | -| captioned image | =<figure>=/=<figcaption>= | =<p>= + ="Figure 1: …"= | real figure semantics | -| timestamp | =<time datetime="…">= | literal =<2024-01-15 Mon>= | machine-readable | -| footnotes | =<section><ol>= | =<h2>Footnotes:</h2>= | a list of notes is a list | -| heading anchor | slug of the text | =org1a2b3c4= | stable, and what the live site serves | -| code | =<pre><code>= | =<pre>= | the HTML5 idiom | - -One genuine semantic difference: org treats a single blank line between a =1.= list and a -=-= list as /one/ list and keeps the first item's bullet type, while we start a second list. -We keep ours, on measurement rather than taste — the pattern occurs *zero* times in the -corpus, so matching an org quirk would buy nothing and cost the more obvious reading. - -*The oracle's best catch was three bugs in itself.* Naive normalization reported code as -corrupted (it trimmed each of syntect's per-token text runs, turning =def greet= into -=defgreet=) and reported blocks at 36% agreement (syntect's spans flooded the diff). Both -were measurement artifacts. A differential harness is a piece of software like any other, -and the first divergences it reports are usually its own. - -** Phase 7: hardening -*** Parse diagnostics (=file:line: message=) -The parser's contract is that it always returns a document — out-of-scope and malformed -constructs degrade rather than crash. The gap was that they degraded /silently/, and in the -worst cases the degradation is severe: an unterminated =#+BEGIN_SRC= reads the rest of the -file as block content, and an unterminated drawer does the same but renders to nothing, so -one missing line deletes most of a page from a build that reports success. - -=parse= now returns =Document::diagnostics=, each carrying a 1-based source line, and the -build prints them as =file:line: message=. =--strict= turns them (and unresolved links) into -a non-zero exit. Line numbers are threaded as an absolute offset through every nested parse, -so a block inside a list item inside a section still reports its real file line — there is a -test for exactly that, because reconstructed and re-indented nested slices are precisely -where an off-by-N hides. The 179-file corpus produces zero diagnostics. - -*** Parallelism -PARSE, RESOLVE and RENDER/EMIT run under rayon. PARSE is a pure function of one file's bytes -and RESOLVE only reads the shared symbol table, which is what makes both safe to parallelize -at all; INDEX stays sequential. - -| corpus | before | after | speedup | -|————————+——--+——-+———| -| 179 files (real) | 0.23s | 0.07s | 3.3× | -| 1,790 files (10× copy) | 3.98s | 0.82s | 4.9× | - -Measured on 12 cores. =RAYON_NUM_THREADS=1= reproduces the old 3.98s exactly, so the gain is -parallelism rather than incidental change, and the output is byte-identical to the sequential -build across the whole corpus. - -*Parallelism must not be observable in the result.* =par_iter().collect()= preserves input -order, so the emitted bytes are unaffected — but the build /report/ is the fragile half: -pushing to =rendered=/=skipped= from inside the parallel pass would order them by thread -scheduling, producing a non-deterministic report over a deterministic site. The parallel pass -therefore returns only what was written, and the report is assembled sequentially afterwards. -=parallel_builds_are_deterministic_in_output_and_report_order= holds that line, and it was -verified by reintroducing the bug and watching it fail. - -*** The real scaling limit was not the CPU -Going 10× on corpus size cost 17× in time, which parallelism improves without fixing: the -cause was the nav bar listing *every* page, so an /n/-page site emitted /n/² nav links. At -1,790 pages each page carried 1,799 links and the output was 284 MB, against 5.5 MB for the -179-page corpus — 52× the bytes for 10× the input. - -The nav is now built from *top-level pages only* ([[file:src/site.rs][=is_top_level=]]): a nav is a -map of the site's top level, not an index of its contents, and section pages reach their -siblings through that section's landing page. Nav size becomes a function of the top level -rather than of the corpus, and the quadratic disappears. - -| 1,790-page corpus (6 top-level pages) | before | after | -|—————————————+——--+——-| -| full build | 0.82s | 0.39s | -| total output | 284 MB | 34 MB | -| nav links per page | 1,799 | 6 | - -Scaling is now linear: 179 pages in 0.07s and 1,796 in 0.39s, where the small case is mostly -the fixed cost of loading syntect's syntax definitions. - -The same rule sharpened the incremental build, which is the larger win. The site-structure -hash — the thing that forces a global re-render — now covers only the pages that appear in -the nav, because those are the only ones whose title or URL affects another page. *Adding a -blog post used to re-render the entire site; now it renders one page.* A top-level page's -title still invalidates everything, correctly, since every page displays it. - -*Trade-off worth knowing:* on a site whose sections live in subdirectories, only genuinely -root-level pages appear — a site keeping its landing pages at =salary/index.org= and friends -gets a one-entry nav. That is what =nav.mode = "explicit"= is for: list the pages you want, -in the order you want them. - -*From v0.1 (core subset):* headings with nesting and anchors (every heading is now -anchored — =:CUSTOM_ID:=/=:ID:= else a slug of its text) and trailing tags; paragraphs; -plain lists (unordered + ordered) with checkboxes; source blocks; inline markup (=*bold*=, -=/italic/=, =_underline_=, =+strike+=, ==verbatim==, =~code~=); links and bare URLs. - -** Compatibility -Versions mean something as of 1.0. The *stable surface* — changing incompatibly requires -a major version — is what you actually build a site against: - -| Stable | Detail | -|——————+——————————————————————————————————————————————-| -| =orgo.toml= keys | Names, types and meaning. New keys are minor releases; removing one is major. | -| Template context | =page=, =site=, =nav=, =root=, =pages=, =group=, =groups=, =paginator=, =stylesheet=, and the =absolute= / =rfc822= / =truncate= filters. | -| CLI | Command names, flags, and exit codes. | -| URLs | How a source path becomes an output path, including =#+SLUG:=. A generator that moves your URLs breaks every link anyone has to you. | - -Explicitly *not stable*, so that the above can be: - -- *The incremental cache.* Versioned, discarded on mismatch, never a correctness - dependency. It changes whenever it needs to, in any release. -- *Rendered HTML details.* orgo tracks what Emacs exports from the same file, and - closing a gap changes markup. Changes that affect output are called out in - [[file:CHANGELOG.org][CHANGELOG.org]] — the class names the documentation names (=post-list=, - =figure-number=, =section-number-N=, =footnote-ref=) are the ones to write CSS against. -- *The Rust API.* The crate is published so the binary can be installed with - =cargo install=; the library exists to serve it, and its types move as the tool does. - -The *MSRV is 1.88*, checked in CI on every change. orgo's own code compiles on -1.82; the floor comes from dependencies. Raising it is a minor version, never a patch. - -** Dependencies -Parser is hand-written recursive descent (not =nom=/=chumsky=/=pest= — org is -line-oriented and context-sensitive, not clean CFG). Key crates: =syntect= (syntax -highlighting, behind a =Highlighter= trait so tree-sitter can be swapped in later), -=minijinja= (runtime templates), =blake3= (content/cache hashing), =rayon= (parallel -PARSE/RESOLVE/RENDER), =notify= (filesystem events for =watch=), =tiny_http= (the =serve= -development server), =toml= (config), =chrono=, =camino=, =walkdir=, =clap=, =anyhow=/=thiserror=. -=insta= for snapshot tests, and =emacs --batch= — optional, and only for the oracle. - -** Build & test -#+begin_example -cargo build -cargo test # 191 tests -cargo run -- init my-site # scaffold a new site -cargo run -- build fixtures/minimal.org -o minimal.html # single file -cargo run -- build fixtures/site -o _site # whole site (incremental) -cargo run -- audit fixtures/site # corpus audit (Phase 0) -cargo run -- build fixtures/site -o _site --no-cache # force a full rebuild -cargo run -- watch fixtures/site -o _site # rebuild on filesystem events -cargo run -- serve fixtures/site -o _site # ... and serve with live reload -cargo run -- clean _site # remove output + cache -#+end_example - -A second =build= of an unchanged site re-renders nothing; editing a page re-renders only -that page and the pages that link into it (watch the =rendered=/=cached= counts). +#+begin_src sh +cargo test # includes a differential check against Emacs, when present +cargo run -- serve docs -o docs/_site # read the documentation locally +#+end_src -A build emits =syntax.css= next to its output (the highlighter emits CSS classes, so the -stylesheet has to come with them) and every page links it. +** Licence -=fixtures/= holds tiny =.org= samples: the core ones (=minimal.org=, =core.org=, -=elements.org=, =table.org=, =footnote.org=), one per v1 construct group (=headings.org=, -=lists.org=, =blocks.org=, =timestamps.org=, =images.org=), the scope guardrail -(=outofscope.org=), and a linked multi-file site under =fixtures/site/= (=index.org=, -=guide.org=, =about.org= + a =style.css= asset). The real corpus (golden files derived from -actual documents) lands in Phase 0. =cargo test= runs =insta= snapshots of the element tree -and rendered HTML for each fixture, the two templated site pages (proving cross-file link -resolution), and the incremental gates. +[[file:LICENSE][0BSD]]. Do what you like with it. @@ -1,8 +1,6 @@ * Releasing -A release is three things that must agree: a version in =Cargo.toml=, a git tag, and a -changelog entry. The release workflow checks the first two against each other and refuses -to build if they differ, because a release tagged =v0.18.0= containing a binary that -reports =0.17.0= is the kind of mistake nobody notices for months. + +Read the content below for the release process. ** Before the first publish #+begin_src sh @@ -11,22 +9,11 @@ cargo publish --dry-run #+end_src =repository= and =homepage= in =Cargo.toml= point at GitHub and at the documentation site -on Pages. If git.krz.sh becomes the primary remote, =repository= should follow it — -crates.io shows that link on the crate page, and it should lead somewhere you read. +on Pages. ** Every release -1. *Write the changelog entry first.* [[file:CHANGELOG.org][CHANGELOG.org]] names behaviour, not - commits — someone reading it wants to know what their next build will do differently. - Anything that changes rendered HTML gets said out loud. - -2. *Bump the version* in =Cargo.toml=, and build once so =Cargo.lock= follows. - - Patch for fixes that change nothing about the stable surface. Minor for new config - keys, new template variables, an MSRV bump, or output that changes to track Emacs more - closely. Major for anything that breaks the promises in the README's Compatibility - section — config keys, template context, CLI, or URLs. - -3. *Check it.* +1. Bump the version in =Cargo.toml=, and build once so =Cargo.lock= follows. +2. Test and package the release. #+begin_src sh cargo test @@ -35,15 +22,9 @@ crates.io shows that link on the crate page, and it should lead somewhere you re cargo package #+end_src - =cargo package= is the one people forget: it builds the crate exactly as crates.io will - receive it, and catches a file the =exclude= list should not have removed. - -4. *Verify against a real corpus.* The test suite says the code does what it did; a - corpus says the /site/ does. Build a site you know with =--no-cache= and diff the - output against the previous version's. A release that quietly changes 200 pages should - do so on purpose. - -5. *Commit, tag, push.* +3. Build a site you know with =--no-cache= and diff the output against the previous + version's. +4. Commit, tag, push. #+begin_src sh git commit -am "0.18: <what changed>" @@ -51,7 +32,7 @@ crates.io shows that link on the crate page, and it should lead somewhere you re git push && git push --tags #+end_src -6. *Publish the crate.* +5. Publish the crate. #+begin_src sh cargo publish @@ -59,10 +40,8 @@ crates.io shows that link on the crate page, and it should lead somewhere you re This is irreversible: a published version can be yanked but never replaced. -7. *Finish the GitHub release.* Pushing the tag builds binaries for macOS (arm64 and - x86_64) and Linux (gnu and musl) and opens a /draft/ release with them attached. Paste - the changelog entry in and publish it. The draft is deliberate — a release that - publishes itself before anyone has read it cannot be edited quietly. +6. Finish the GitHub release. Pushing the tag builds binaries for macOS (arm64 and + x86_64) and Linux (gnu and musl) and opens a /draft/ release with them attached. ** If a release goes wrong Yank rather than delete, and ship a fix as a new version: @@ -72,4 +51,4 @@ cargo yank --version 0.18.0 #+end_src Yanking stops new dependents from selecting it; anyone who already has it keeps working. -Then release =0.18.1= with the fix and a changelog entry that says what happened. +Then release =0.18.1= with the fix. @@ -11,7 +11,7 @@ Fixes land in a new release rather than as patches to an old one. ## Reporting -Email <hello@cleberg.net>, or open a private advisory through GitHub's *Security* tab. +Email <security@krz.sh>, or open a private advisory through GitHub's *Security* tab. Please do not open a public issue for something exploitable. ## What is worth reporting @@ -35,8 +35,8 @@ dependency: a missing, stale or corrupt cache produces exactly the same site, mo ** Rendered HTML details orgo aims at what Emacs exports from the same file, and closing a gap changes markup. -That is the product working rather than a regression — but it is called out in the -changelog every time, because your stylesheet is downstream of it. +That is the product working rather than a regression — but it is called out in the release +notes every time, because your stylesheet is downstream of it. The class names the documentation names are the ones to write CSS against: =post-list=, =post-list-item=, =figure-number=, =table-number=, =section-number-N=, @@ -68,5 +68,4 @@ new version's output to the old version's output rather than to a cache written mixture of both. =--strict= turns a link that stopped resolving into a failure. If you keep your built site in version control, the diff after that command *is* the -upgrade report — which is the most useful review a generator can give you, and the reason -the changelog names behaviour rather than commits. +upgrade report, and the most useful review a generator can give you. @@ -4,7 +4,7 @@ * Requirements -- *Rust 1.82 or newer.* Install from [[https://rustup.rs][rustup.rs]] if you do not have it. There is no other +- *Rust 1.88 or newer.* Install from [[https://rustup.rs][rustup.rs]] if you do not have it. There is no other runtime requirement: the binary is self-contained, with syntax definitions and highlighting themes compiled in. - *Emacs (optional).* Only the differential test suite uses it, to compare output against @@ -13,7 +13,7 @@ * From source #+BEGIN_SRC sh -git clone <repository-url> orgo +git clone https://github.com/ccleberg/orgo cd orgo cargo build --release #+END_SRC