krz/orgo

Lightning fast org-mode static site generator. fast go org-mode static-site-generator

Commit 19279bbf18

19279bbf18ad26b1591ff81cfc319935e9352194

parent: 4db41537b5

Verified · cmc

cmc <hello@cleberg.net> · 2026-08-11 04:00 UTC

org-ssg v0.4 (MVP): deliver the full v1 construct scope, enforce the IN/OUT line

Closes the gap between the v1 scope the README declared and what the parser and
renderer actually did. Every construct on the IN list is now parsed, rendered, and
pinned by a golden file; every construct on the OUT list has a test asserting how it
degrades.

Constructs:
- Headings carry TODO keywords and [#A] priorities. Uses Emacs' default TODO/DONE set
  and its export classes, matched on a word boundary, so output stays diffable against
  the emacs --batch oracle that Phase 0 will add.
- Lists nest by indentation, hold multi-paragraph bodies, and support `term ::
  definition` description lists. They were previously flat and single-line.
- QUOTE/CENTER/EXAMPLE/EXPORT are distinct elements rather than all collapsing to a
  verbatim example block. Block ends match their own kind, so SRC nests inside QUOTE.
  An html export block passes through; other backends drop.
- Timestamps are parsed and rendered as <time> with a machine-readable datetime:
  active and inactive, times, same-day and multi-day ranges. Repeater cookies are
  recognized and discarded. The model had the type; nothing produced it.
- Images render as <img>, and are promoted to <figure> by an affiliated #+CAPTION: or
  #+ATTR_HTML:. Links to non-.org files are understood as asset links, so they are
  neither resolved nor reported broken.
- Source blocks are highlighted by syntect into CSS classes (never inline styles).
  Each build emits the matching syntax.css and every page links it relative to its own
  depth. An unknown language degrades to escaped <pre><code>.

Two content bugs found by running malformed input through the binary:
- Preamble keywords were lifted into the metadata map and deleted from the body, which
  merged the paragraphs either side of a keyword and stranded #+CAPTION: away from the
  image below it. Only the preamble was affected, which is the one region where every
  real document has keywords. Collecting now copies instead of deleting; element trees
  gain inert Keyword nodes, and no rendered HTML changed.
- image_tag emitted a caption-derived alt and then appended #+ATTR_HTML attributes that
  could carry their own, producing two alt attributes on one tag. An explicit :alt now
  wins.

Scope guardrail (tests/constructs.rs): babel is never executed and a checked-in
#+RESULTS: block is dropped rather than published as if it were verified output;
#+TBLFM: is inert; #+INCLUDE: is never expanded; LaTeX, macros and radio targets
survive as literal text; non-PROPERTIES drawers are captured and dropped; unmodelled
block types keep their content.

Broken links are reported as the org syntax the author wrote rather than a
Debug-printed enum. Cache format version bumped to 4, since render output changed.

Still out, and recorded as such in the README: the Phase 0 corpus audit and oracle (the
fixtures are hand-written, so "matches Emacs" is asserted by construction, not
measured), rayon parallelism, parse errors carrying source locations, #+TODO: keyword
sequences, planning lines, and fixed-width lines.

Layout: unified · split

Cargo.lock +1 −1
@@ -538,7 +538,7 @@ dependencies = [
538538
539539[[package]]
540540name = "org-ssg"
541version = "0.3.0"
541version = "0.4.0"
542542dependencies = [
543543 "anyhow",
544544 "blake3",
Cargo.toml +1 −1
@@ -1,6 +1,6 @@
11[package]
22name = "org-ssg"
3version = "0.3.0"
3version = "0.4.0"
44edition = "2021"
55description = "Org-mode static site generator that renders the org element tree straight to HTML"
66license = "MIT"
README.md +67 −22
@@ -32,7 +32,7 @@ is the only inherently global stage — it is where the link dependency graph is
3232| TEMPLATE | `src/template.rs` | minijinja: fragment + metadata → full page. |
3333| incremental | `src/incremental.rs` | Content/config/template hashing, dep graph, cache manifest, invalidation. |
3434
35## v1 scope (recommended; must be reconciled against a corpus audit first)
35## v1 scope (delivered as of v0.4; still to be reconciled against a corpus audit)
3636
3737**IN — v1 must handle:** headings with nesting; TODO keywords; priorities `[#A]`; tags;
3838property drawers; plain lists (unordered/ordered/description, checkboxes, nesting);
@@ -47,10 +47,11 @@ paragraphs and horizontal rules; images with `#+CAPTION`/`#+ATTR_HTML`.
4747macros; drawers other than PROPERTIES/LOGBOOK; column view / clocking / agenda
4848semantics; non-HTML export blocks; the full Unicode entity set.
4949
50**Scope guardrail:** every IN item gets a golden-file fixture from a real document;
51every OUT item gets a test asserting it degrades predictably (ignored, no crash). The
52IN/OUT line is enforced by tests, defending against the project's #1 risk: scope creep
53back toward all-of-org.
50**Scope guardrail:** every IN item gets a golden-file fixture; every OUT item gets a test
51asserting it degrades predictably (ignored, no crash). The IN/OUT line is enforced by
52`tests/constructs.rs`, defending against the project's #1 risk: scope creep back toward
53all-of-org. The fixtures are hand-written today; deriving them from a real corpus is
54Phase 0.
5455
5556## Phase plan
5657
@@ -60,14 +61,15 @@ back toward all-of-org.
6061| **v0.1** | **End-to-end core parse → render: `build` a single `.org` file to HTML** | **done** |
6162| **v0.2** | **Multi-file SITE build: INDEX + RESOLVE internal links, minijinja templates, `build <src-dir> <out-dir>`, tables + footnotes** | **done** |
6263| **v0.3** | **Incremental build layer: content/config/template hashing, dependency graph, per-page render keys, persisted cache manifest, invalidation** | **done** |
64| **v0.4** | **MVP: the full v1 construct scope — heading metadata, nested/description lists, block types, timestamps, images, syntect highlighting — with the IN/OUT line under test** | **done** |
6365| 0 | Corpus audit + `emacs --batch` ground-truth oracle | todo |
6466| 1 | Line lexer + heading/section skeleton | done |
65| 2 | Block elements — lists, source blocks, tables, footnote defs done; generic drawers | partial |
66| 3 | Inline objects — emphasis, links, bare URLs, footnote refs done; timestamps | partial |
67| 4 | Rendering to HTML — tree walk, tables, footnote two-pass, minijinja templating done; real syntect highlighting | partial |
67| 2 | Block elements — lists, source blocks, tables, footnote defs, blocks by type, drawers | done |
68| 3 | Inline objects — emphasis, links, bare URLs, footnote refs, timestamps | done |
69| 4 | Rendering to HTML — tree walk, tables, footnote two-pass, minijinja templating, syntect highlighting | done |
6870| 5 | Link resolution + symbol table (INDEX + RESOLVE, used-target list, broken-link reporting) | done |
6971| 6 | Incremental build layer (hashing, dep graph, invalidation) done; `watch` is a simple poll loop | done |
70| 7 | Hardening: rayon parallelism, CLI polish, error locations | todo |
72| 7 | Hardening: rayon parallelism, error locations in parse diagnostics | todo |
7173
7274### v0.2 in / out
7375
@@ -82,9 +84,8 @@ plus two new constructs — pipe **tables** (with header band from the rule row)
8284**footnotes** (block `[fn:1]` definitions, referenced `[fn:1]`, and inline `[fn:1:text]`,
8385rendered as a numbered, back-linked notes section).
8486
85**Still stubbed (`todo!`):** timestamps; TODO keywords and priorities; generic
86(non-PROPERTIES) drawers. Source-block syntax highlighting remains a `<pre><code>`
87passthrough behind the `Highlighter` trait; real syntect tokenizing is deferred.
87**Left stubbed at v0.2, all closed in v0.4:** timestamps; TODO keywords and priorities;
88generic (non-PROPERTIES) drawers; real syntect tokenizing behind the `Highlighter` trait.
8889
8990### v0.3 in / out
9091
@@ -121,12 +122,52 @@ re-renders exactly the changed page plus its linkers; **renamed-heading** invali
121122linking page and updates its emitted anchor; and cache **version-bump / missing / corrupt**
122123all fall back to a full rebuild.
123124
124**Out of scope in v0.3 (unchanged from v0.2):** real syntect highlighting; timestamps and
125TODO keywords. `watch` is a minimal mtime poll loop (`watch <src-dir> -o <out-dir>`), not an
126OS file-watcher — the fs-notify integration is deferred. The parse-tree cache (spec §4.5,
125**Out of scope in v0.3:** real syntect highlighting; timestamps and TODO keywords (all
126landed in v0.4). `watch` is a minimal mtime poll loop (`watch <src-dir> -o <out-dir>`), not
127an OS file-watcher — the fs-notify integration is deferred. The parse-tree cache (spec §4.5,
127128"optionally") is not persisted: PARSE/INDEX/RESOLVE run for every file each build (cheap and
128129pure); the incremental win is on RENDER + EMIT.
129130
131### v0.4 in / out — the MVP
132
133v0.4 closes the gap between the v1 scope above and what the code actually did, so every
134construct the IN list claims is now parsed, rendered, and pinned by a golden file:
135
136- **Heading metadata** — TODO keywords (the Emacs default `TODO`/`DONE` set, matched on a
137 word boundary so `TODOs` is not one) and `[#A]` priority cookies, rendered with Emacs'
138 own export classes so the output stays diffable against an `emacs --batch` oracle.
139- **Lists** — indentation-based nesting (a sub-list renders *inside* its parent `<li>`),
140 multi-paragraph item bodies, and `term :: definition` description lists as `<dl>`.
141- **Blocks by type** — `QUOTE`, `CENTER`, `EXAMPLE`, `EXPORT` and `SRC` are now distinct
142 elements rather than all collapsing to a verbatim example block. Block matching is on the
143 specific kind, so a source block can nest inside a quote. An `html` export block passes
144 through; every other backend drops.
145- **Timestamps** — active `<...>` and inactive `[...]`, optional times, same-day time
146 ranges and `--`-joined date ranges, rendered as `<time>` with a machine-readable
147 `datetime`. Repeater/warning cookies are recognized and discarded.
148- **Images** — a description-less link to an image file renders as `<img>`; with an
149 affiliated `#+CAPTION:`/`#+ATTR_HTML:` it is promoted to a `<figure>` with the caption as
150 both `<figcaption>` and alt text. Links to non-`.org` files are now understood as asset
151 links: neither resolved nor reported as broken.
152- **Syntax highlighting** — real syntect tokenizing to CSS classes (never inline styles, so
153 themes live in the stylesheet). Every build emits the matching `syntax.css` and each page
154 links it relative to its own depth. An unknown language degrades to escaped `<pre><code>`.
155- **Diagnostics** — broken links are reported as the org syntax the author wrote
156 (`warning: b.org: unresolved link [[#setup]]`) rather than a Debug-printed enum.
157
158**The OUT line is now enforced, not just asserted.** `tests/constructs.rs` pins each
159excluded construct to a specific degradation: babel is never executed *and* a checked-in
160`#+RESULTS:` block is dropped rather than published as if it were verified output;
161`#+TBLFM:` is inert; `#+INCLUDE:` is never expanded; LaTeX, macros and radio targets survive
162as literal text; drawers other than PROPERTIES are captured and dropped; unmodelled block
163types keep their content verbatim.
164
165**Still out at v0.4:** the Phase 0 corpus audit and `emacs --batch` oracle (the fixtures are
166hand-written, so "matches Emacs" is asserted by construction, not measured); rayon
167parallelism; parse errors carrying source locations; `#+TODO:` per-file keyword sequences;
168planning lines (`SCHEDULED:`/`DEADLINE:`), which render as ordinary paragraphs; fixed-width
169`: ` lines; and the `watch` fs-notify integration.
170
130171**From v0.1 (core subset):** headings with nesting and anchors (every heading is now
131172anchored — `:CUSTOM_ID:`/`:ID:` else a slug of its text) and trailing tags; paragraphs;
132173plain lists (unordered + ordered) with checkboxes; source blocks; inline markup (`*bold*`,
@@ -155,10 +196,14 @@ cargo run -- clean _site # remove output + cach
155196A second `build` of an unchanged site re-renders nothing; editing a page re-renders only
156197that page and the pages that link into it (watch the `rendered`/`cached` counts).
157198
158`fixtures/` holds tiny `.org` samples: single-file ones (`minimal.org`, `core.org`,
159`elements.org`, `table.org`, `footnote.org`) and a linked multi-file site under
160`fixtures/site/` (`index.org`, `guide.org`, `about.org` + a `style.css` asset). The
161real corpus (golden files derived from actual documents) lands in Phase 0. `cargo test`
162includes `insta` snapshots of the element tree and rendered HTML for the single-file
163fixtures, the two templated site pages (proving cross-file link resolution), and the
164table and footnote constructs.
199A build emits `syntax.css` next to its output (the highlighter emits CSS classes, so the
200stylesheet has to come with them) and every page links it.
201
202`fixtures/` holds tiny `.org` samples: the core ones (`minimal.org`, `core.org`,
203`elements.org`, `table.org`, `footnote.org`), one per v1 construct group (`headings.org`,
204`lists.org`, `blocks.org`, `timestamps.org`, `images.org`), the scope guardrail
205(`outofscope.org`), and a linked multi-file site under `fixtures/site/` (`index.org`,
206`guide.org`, `about.org` + a `style.css` asset). The real corpus (golden files derived from
207actual documents) lands in Phase 0. `cargo test` runs `insta` snapshots of the element tree
208and rendered HTML for each fixture, the two templated site pages (proving cross-file link
209resolution), and the incremental gates.
fixtures/blocks.org added +53
@@ -0,0 +1,53 @@
1#+TITLE: Blocks
2
3* Quote
4
5#+BEGIN_QUOTE
6A quoted paragraph with /markup/.
7
8And a second paragraph.
9#+END_QUOTE
10
11* Center
12
13#+BEGIN_CENTER
14Centred text.
15#+END_CENTER
16
17* Example
18
19#+BEGIN_EXAMPLE
20Verbatim *not bold* text.
21 Indentation preserved.
22#+END_EXAMPLE
23
24* Export
25
26#+BEGIN_EXPORT html
27<aside class="raw">Raw HTML passes through.</aside>
28#+END_EXPORT
29
30#+BEGIN_EXPORT latex
31\emph{A non-HTML backend is dropped.}
32#+END_EXPORT
33
34* Source
35
36#+BEGIN_SRC python
37def greet(name):
38 return f"hello {name}"
39#+END_SRC
40
41#+BEGIN_SRC
42plain block, no language
43#+END_SRC
44
45* Nested
46
47#+BEGIN_QUOTE
48A quote containing a source block:
49
50#+BEGIN_SRC sh
51echo hi
52#+END_SRC
53#+END_QUOTE
fixtures/headings.org added +20
@@ -0,0 +1,20 @@
1#+TITLE: Heading Metadata
2
3* TODO [#A] Write the parser :work:rust:
4:PROPERTIES:
5:CUSTOM_ID: write-parser
6:OWNER: nobody
7:END:
8A heading carrying a keyword, a priority, tags and a property drawer.
9
10** DONE Nested and finished
11Sub-headings nest by star count.
12
13** [#C] Priority without a keyword
14A priority cookie can stand alone.
15
16* TODOs are not a keyword
17The word boundary matters: this heading has no TODO keyword.
18
19* DONE
20A keyword with no title at all.
fixtures/images.org added +25
@@ -0,0 +1,25 @@
1#+TITLE: Images
2
3* Bare image
4
5[[file:diagram.png]]
6
7* Captioned figure
8
9#+CAPTION: The pipeline, end to end
10#+ATTR_HTML: :width 640 :class diagram
11[[file:pipeline.svg]]
12
13* Caption with markup
14
15#+CAPTION: A /stylised/ chart
16[[file:chart.png]]
17
18* Quoted attribute values
19
20#+ATTR_HTML: :alt "a cat, sitting" :loading lazy
21[[file:cat.jpg]]
22
23* Image with a description is a link
24
25[[file:diagram.png][the diagram]]
fixtures/lists.org added +38
@@ -0,0 +1,38 @@
1#+TITLE: Lists
2
3* Nesting
4
5- outer item
6 - inner item
7 - deepest item
8 - second inner
9- second outer
10
11* Ordered
12
131. first
142. second
15 1. second point one
16 2. second point two
173. third
18
19* Checkboxes
20
21- [ ] not done
22- [X] done
23- [-] partially done
24
25* Description
26
27- term one :: the first definition
28- term two :: the second definition, which is
29 soft-wrapped across two lines
30- /marked up/ term :: definitions hold inline markup
31
32* Multi-paragraph items
33
34- an item whose body has two paragraphs
35
36 the second paragraph, indented under the bullet
37
38- a plain sibling
fixtures/outofscope.org added +55
@@ -0,0 +1,55 @@
1#+TITLE: Out of Scope
2#+INCLUDE: "other.org"
3
4Every construct here is on the README's explicit OUT list. The contract is not that we
5handle them — it is that they degrade predictably and never crash the build.
6
7* Babel
8
9#+BEGIN_SRC sh :results output :exports both
10echo "the block renders; :results is never executed"
11#+END_SRC
12
13#+RESULTS:
14: stale output from a previous evaluation
15
16* Table formulas
17
18| item | cost |
19|------+------|
20| a | 1 |
21| b | 2 |
22#+TBLFM: $2=vsum(@2..@3)
23
24* LaTeX
25
26Inline math $x^2 + y^2$ and a display block:
27
28\begin{equation}
29E = mc^2
30\end{equation}
31
32* Macros and radio targets
33
34A macro call {{{author}}} and a <<<radio target>>> stay literal.
35
36* Drawers
37
38:LOGBOOK:
39CLOCK: [2024-01-15 Mon 09:00]--[2024-01-15 Mon 10:00] => 1:00
40:END:
41
42:CUSTOM_DRAWER:
43Drawer contents are captured and dropped.
44:END:
45
46* Verse
47
48#+BEGIN_VERSE
49An unmodelled block type
50keeps its content verbatim.
51#+END_VERSE
52
53* Entities
54
55The full entity set is out of scope, so \alpha stays literal.
fixtures/timestamps.org added +21
@@ -0,0 +1,21 @@
1#+TITLE: Timestamps
2
3* Single
4
5An active date <2024-01-15 Mon> and an inactive one [2024-01-15 Mon].
6
7With a time: <2024-01-15 Mon 10:30>.
8
9* Ranges
10
11A same-day time range <2024-01-15 Mon 10:00-11:45>.
12
13A multi-day range <2024-01-15 Mon>--<2024-01-20 Sat>.
14
15* Ignored decorations
16
17A repeater is dropped: <2024-01-15 Mon +1w>.
18
19* Not timestamps
20
21Comparisons like 3 < 4 and [not a stamp] stay literal text.
src/incremental.rs +2 −2
@@ -29,7 +29,7 @@ use crate::util::output_url;
2929/// Bump whenever the `Document` type, hashing scheme, or resolution rules change.
3030/// On mismatch: discard cache, full rebuild (spec §4.5). The blake3 crate's major
3131/// version is folded in as the "hash-algo version" so a hash upgrade also busts.
32pub const CACHE_FORMAT_VERSION: u32 = 3;
32pub const CACHE_FORMAT_VERSION: u32 = 4;
3333
3434/// blake3 hex identity for a content/config/template/render-key hash class (spec §4.1).
3535pub type Hash = ContentHash;
@@ -48,7 +48,7 @@ impl Default for BuildConfig {
4848 fn default() -> Self {
4949 BuildConfig {
5050 output_extension: "html".to_string(),
51 highlighter_theme: "passthrough".to_string(),
51 highlighter_theme: crate::render::SYNTAX_THEME.to_string(),
5252 }
5353 }
5454}
src/index.rs +13
@@ -21,6 +21,19 @@ pub enum TargetId {
2121 File(Utf8PathBuf),
2222}
2323
24/// How a target is written in org source, so a broken-link warning names something the
25/// author can search for.
26impl std::fmt::Display for TargetId {
27 fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
28 match self {
29 TargetId::Id(s) => write!(f, "[[id:{s}]]"),
30 TargetId::CustomId(s) => write!(f, "[[#{s}]]"),
31 TargetId::Heading(s) => write!(f, "[[*{s}]]"),
32 TargetId::File(p) => write!(f, "[[file:{p}]]"),
33 }
34 }
35}
36
2437impl TargetId {
2538 /// A stable string form used to order targets deterministically when hashing
2639 /// (so a page's `resolved_links_hash` does not depend on `HashSet` iteration order).
src/main.rs +10 −6
@@ -7,13 +7,13 @@ use camino::{Utf8Path, Utf8PathBuf};
77use clap::{Parser, Subcommand};
88
99use org_ssg::parser::parse;
10use org_ssg::render::{render, Html, SyntectHighlighter};
10use org_ssg::render::{render, syntax_css, Html, SyntectHighlighter};
1111use org_ssg::resolve::ResolvedDoc;
12use org_ssg::site::{build_site, BuildOptions};
12use org_ssg::site::{build_site, BuildOptions, SYNTAX_STYLESHEET};
1313use org_ssg::template::Templater;
1414
1515#[derive(Parser)]
16#[command(name = "org-ssg", about = "Org-mode static site generator")]
16#[command(name = "org-ssg", version, about = "Org-mode static site generator")]
1717struct Cli {
1818 #[command(subcommand)]
1919 command: Command,
@@ -153,7 +153,8 @@ fn watch(input: &Utf8Path, output: &Utf8Path) -> Result<()> {
153153
154154/// Single-file build: read → PARSE → RENDER → TEMPLATE → write. No cross-file link
155155/// resolution (there is no corpus to resolve against); links keep their best-effort
156/// URLs. Whole-site link resolution lives in [`build_site`].
156/// URLs. Whole-site link resolution lives in [`build_site`]. The syntax stylesheet is
157/// written alongside the page, since highlighting emits CSS classes.
157158fn build_file(input: &Utf8Path, output: &Utf8Path) -> Result<()> {
158159 let source = fs::read_to_string(input)
159160 .with_context(|| format!("reading source file {input}"))?;
@@ -168,13 +169,16 @@ fn build_file(input: &Utf8Path, output: &Utf8Path) -> Result<()> {
168169 .unwrap_or_else(|| input.file_stem().unwrap_or("untitled").to_string());
169170
170171 let resolved = ResolvedDoc { document };
171 let highlighter = SyntectHighlighter;
172 let highlighter = SyntectHighlighter::new();
172173 let Html(fragment) = render(&resolved, &highlighter);
173174
174175 let templater = Templater::new();
175176 let page = templater
176 .render_page(&title, &fragment, &[])
177 .render_page(&title, &fragment, &[], SYNTAX_STYLESHEET)
177178 .with_context(|| format!("templating {input}"))?;
178179 fs::write(output, page).with_context(|| format!("writing output file {output}"))?;
180
181 let css = output.with_file_name(SYNTAX_STYLESHEET);
182 fs::write(&css, syntax_css()).with_context(|| format!("writing stylesheet {css}"))?;
179183 Ok(())
180184}
src/model.rs +8
@@ -86,6 +86,14 @@ pub enum Element {
8686 code: String,
8787 },
8888 ExampleBlock(String),
89 /// An image link carrying affiliated `#+CAPTION:`/`#+ATTR_HTML:` metadata, which
90 /// promotes it from an inline image to a block-level `<figure>`.
91 Figure {
92 link: Link,
93 caption: Vec<Object>,
94 /// Raw `#+ATTR_HTML:` attribute string, passed through to the `<img>` tag.
95 attrs: String,
96 },
8997 QuoteBlock(Vec<Element>),
9098 CenterBlock(Vec<Element>),
9199 /// html passes through; others dropped at render (spec §1 OUT).
src/parser.rs +501 −85
@@ -9,18 +9,24 @@
99//! PARSE is a pure function of a single file's bytes (spec §2.1): it never depends on
1010//! another file, which is what makes content-hash caching sound.
1111//!
12//! v0.1 scope (the CORE subset): headings + nesting, property drawers on headings,
13//! paragraphs, plain lists (unordered + ordered) with checkboxes, source blocks, and
14//! inline markup (bold/italic/underline/strike/verbatim/code, links, bare URLs).
15//! Out of scope and left graceful (parsed-and-ignored, never crashing): tables,
16//! footnotes, timestamps, TODO keywords, non-SRC blocks (kept verbatim as example
17//! blocks), generic drawers other than PROPERTIES.
12//! Scope is the v1 IN list (README §"v1 scope"): headings with nesting, TODO keywords,
13//! priorities, tags and property drawers; paragraphs; plain lists (unordered, ordered,
14//! description) with checkboxes and nesting; tables; source/example/quote/center/export
15//! blocks; footnotes; `#+` keywords; inline markup, links, timestamps; images with
16//! `#+CAPTION`/`#+ATTR_HTML`.
17//!
18//! Out-of-scope constructs are parsed-and-ignored, never fatal: babel `:results` and
19//! `#+TBLFM:` are inert keywords, unknown block types keep their content verbatim as
20//! example blocks, generic drawers are captured and dropped at render, and LaTeX,
21//! macros and radio targets survive as literal text.
1822
1923use camino::Utf8Path;
24use chrono::{NaiveDate, NaiveDateTime, NaiveTime};
2025
2126use crate::model::{
2227 BlockParams, Bullet, Checkbox, ContentHash, Document, Element, Heading, Keywords, Link,
23 LinkTarget, List, ListItem, ListKind, Object, Properties, Section, Table, TableRow,
28 LinkTarget, List, ListItem, ListKind, Object, Properties, Section, Table, TableRow, Timestamp,
29 TodoKeyword,
2430};
2531
2632#[derive(Debug, thiserror::Error)]
@@ -113,20 +119,23 @@ pub fn parse(path: &Utf8Path, source: &str) -> Result<Document, ParseError> {
113119 .collect();
114120 let first = heading_idxs.first().copied().unwrap_or(lines.len());
115121
116 // Preamble: document-level keywords are lifted into `keywords`; the remaining
117 // lines become the root section's block content.
122 // Preamble: document-level keywords are *copied* into `keywords`, which is the
123 // metadata map. They are not removed from the body — collecting is not deleting.
124 // Dropping the lines would merge the paragraphs either side of a keyword and would
125 // strand affiliated keywords (`#+CAPTION:`) away from the element they belong to;
126 // left in place, `parse_elements` handles both. Affiliated keywords are not document
127 // metadata, so they are not copied.
118128 {
119 let mut body: Vec<&str> = Vec::new();
120129 for (l, c) in lines[..first].iter().zip(&classes[..first]) {
121130 if *c == Line::Keyword {
122131 if let Some((k, v)) = keyword_kv(l) {
123 keywords.entries.push((k, v));
132 if !is_affiliated(&k) {
133 keywords.entries.push((k, v));
134 }
124135 }
125 } else {
126 body.push(l);
127136 }
128137 }
129 root.content = parse_elements(&body);
138 root.content = parse_elements(&lines[..first]);
130139 }
131140
132141 // Each heading segment runs from its own line up to (but excluding) the next heading.
@@ -206,14 +215,24 @@ fn heading_level(line: &str) -> Option<u8> {
206215 }
207216}
208217
218/// The default TODO keyword set, matching Emacs' out-of-the-box `org-todo-keywords`
219/// (`("TODO" "DONE")`) so our output can be diffed against an `emacs --batch` oracle.
220/// Per-file `#+TODO:` sequences are out of scope; the set is a documented [`BuildConfig`]
221/// slot for when it becomes configurable.
222///
223/// [`BuildConfig`]: crate::incremental::BuildConfig
224const TODO_KEYWORDS: &[(&str, bool)] = &[("TODO", false), ("DONE", true)];
225
209226fn parse_heading(line: &str) -> Heading {
210227 let level = heading_level(line).unwrap_or(1);
211228 let rest = line[level as usize..].trim();
212229 let (title_str, tags) = split_tags(rest);
230 let (todo, after_todo) = split_todo(title_str.trim());
231 let (priority, title_str) = split_priority(after_todo);
213232 Heading {
214233 level,
215 todo: None, // TODO keywords: out of scope for v0.1.
216 priority: None, // priorities: out of scope for v0.1.
234 todo,
235 priority,
217236 title: inline(title_str.trim()),
218237 tags,
219238 properties: Properties::default(),
@@ -222,6 +241,43 @@ fn parse_heading(line: &str) -> Heading {
222241 }
223242}
224243
244/// A leading TODO keyword: a bare word from the keyword set, followed by whitespace or
245/// end of the heading. `* TODOs are great` is NOT a keyword (no word boundary).
246fn split_todo(title: &str) -> (Option<TodoKeyword>, &str) {
247 let word_end = title.find(char::is_whitespace).unwrap_or(title.len());
248 let word = &title[..word_end];
249 for (name, done) in TODO_KEYWORDS {
250 if word == *name {
251 return (
252 Some(TodoKeyword {
253 name: (*name).to_string(),
254 done: *done,
255 }),
256 title[word_end..].trim_start(),
257 );
258 }
259 }
260 (None, title)
261}
262
263/// A priority cookie `[#A]` immediately after the TODO keyword.
264fn split_priority(title: &str) -> (Option<char>, &str) {
265 let Some(rest) = title.strip_prefix("[#") else {
266 return (None, title);
267 };
268 let mut chars = rest.chars();
269 let Some(c) = chars.next().filter(|c| c.is_ascii_alphanumeric()) else {
270 return (None, title);
271 };
272 match chars.next() {
273 Some(']') => (
274 Some(c.to_ascii_uppercase()),
275 rest[c.len_utf8() + 1..].trim_start(),
276 ),
277 _ => (None, title),
278 }
279}
280
225281/// Split a trailing `:tag1:tag2:` cluster off the heading text.
226282fn split_tags(rest: &str) -> (&str, Vec<String>) {
227283 let trimmed = rest.trim_end();
@@ -311,6 +367,10 @@ fn parse_property(line: &str) -> Option<(String, String)> {
311367
312368fn parse_elements(lines: &[&str]) -> Vec<Element> {
313369 let mut out = Vec::new();
370 // Affiliated keywords (`#+CAPTION:` and friends) belong to the element that follows
371 // them, so they are held aside until that element is built.
372 let mut affiliated: Vec<(String, String)> = Vec::new();
373 let mut drop_next = false;
314374 let mut i = 0;
315375 while i < lines.len() {
316376 let line = lines[i];
@@ -318,70 +378,84 @@ fn parse_elements(lines: &[&str]) -> Vec<Element> {
318378 i += 1;
319379 continue;
320380 }
321 if let Some((kind, after)) = block_begin(line) {
322 let mut j = i + 1;
323 let mut inner = Vec::new();
324 while j < lines.len() && !is_block_end(lines[j]) {
325 inner.push(lines[j]);
326 j += 1;
327 }
328 let code = inner.join("\n");
329 if kind.eq_ignore_ascii_case("SRC") {
330 let (lang, params) = parse_src_header(&after);
331 out.push(Element::SrcBlock { lang, params, code });
381 if let Some((key, value)) = keyword_kv(line) {
382 if key.eq_ignore_ascii_case("RESULTS") {
383 // Babel is never executed (README §OUT), so a checked-in `#+RESULTS:`
384 // block is output from someone else's Emacs session at some other time.
385 // Emitting it would put unverifiable content on the page dressed as
386 // real content, so the block it labels is dropped.
387 drop_next = true;
388 } else if is_affiliated(&key) {
389 affiliated.push((key, value));
332390 } else {
333 // Non-SRC blocks (quote/example/center/export) are kept verbatim for
334 // v0.1 rather than richly modeled — see module scope note.
335 out.push(Element::ExampleBlock(code));
391 out.push(Element::Keyword { key, value });
336392 }
337 i = if j < lines.len() { j + 1 } else { j };
338 continue;
339 }
340 if is_rule(line) {
341 out.push(Element::HorizontalRule);
342393 i += 1;
343394 continue;
344395 }
345 if let Some((key, value)) = keyword_kv(line) {
346 out.push(Element::Keyword { key, value });
347 i += 1;
348 continue;
349 }
350 if line.trim_start().starts_with('|') {
351 let (table, next) = parse_table(lines, i);
352 out.push(Element::Table(table));
353 i = next;
354 continue;
355 }
356 if let Some((label, first_rest)) = footnote_def_label(line) {
357 let (def, next) = parse_footnote_def(lines, i, label, first_rest);
358 out.push(def);
359 i = next;
396 let (element, next) = parse_one_element(lines, i);
397 i = next;
398 if std::mem::take(&mut drop_next) {
399 affiliated.clear();
360400 continue;
361401 }
362 if is_list_item(line.trim_start()).is_some() {
363 let (list, next) = parse_list(lines, i);
364 out.push(Element::List(list));
365 i = next;
366 continue;
367 }
368 // Paragraph: gather consecutive soft-wrapped text lines.
369 let mut para = Vec::new();
370 while i < lines.len() {
371 let l = lines[i];
372 if l.trim().is_empty() || is_structural(l) {
373 break;
374 }
375 para.push(l.trim());
376 i += 1;
377 }
378 if !para.is_empty() {
379 out.push(Element::Paragraph(inline(&para.join(" "))));
402 if let Some(element) = element {
403 out.push(attach_affiliated(element, std::mem::take(&mut affiliated)));
380404 }
381405 }
382406 out
383407}
384408
409/// Build the single element starting at `lines[start]`, returning it with the index of
410/// the first line past it. `None` means the lines were consumed without producing an
411/// element. `start` is guaranteed non-blank and not an affiliated keyword.
412fn parse_one_element(lines: &[&str], start: usize) -> (Option<Element>, usize) {
413 let line = lines[start];
414 if let Some(text) = comment_text(line) {
415 return (Some(Element::Comment(text)), start + 1);
416 }
417 if let Some((kind, after)) = block_begin(line) {
418 let (el, next) = parse_block(lines, start, &kind, &after);
419 return (Some(el), next);
420 }
421 if let Some(name) = drawer_begin_name(line) {
422 let (el, next) = parse_drawer(lines, start, name);
423 return (Some(el), next);
424 }
425 if is_rule(line) {
426 return (Some(Element::HorizontalRule), start + 1);
427 }
428 if line.trim_start().starts_with('|') {
429 let (table, next) = parse_table(lines, start);
430 return (Some(Element::Table(table)), next);
431 }
432 if let Some((label, first_rest)) = footnote_def_label(line) {
433 let (def, next) = parse_footnote_def(lines, start, label, first_rest);
434 return (Some(def), next);
435 }
436 if is_list_item(line.trim_start()).is_some() {
437 let (list, next) = parse_list(lines, start);
438 return (Some(Element::List(list)), next);
439 }
440 // Paragraph: gather consecutive soft-wrapped text lines.
441 let mut para = Vec::new();
442 let mut i = start;
443 while i < lines.len() {
444 let l = lines[i];
445 if l.trim().is_empty() || is_structural(l) {
446 break;
447 }
448 para.push(l.trim());
449 i += 1;
450 }
451 if para.is_empty() {
452 // `is_structural` said this line begins a construct that no branch above claimed
453 // (a stray `#+END_`); skip it rather than looping forever.
454 return (None, start + 1);
455 }
456 (Some(Element::Paragraph(inline(&para.join(" ")))), i)
457}
458
385459/// Is this line the start of a non-paragraph construct?
386460fn is_structural(line: &str) -> bool {
387461 let t = line.trim_start();
@@ -389,12 +463,153 @@ fn is_structural(line: &str) -> bool {
389463 || is_block_end(line)
390464 || is_rule(line)
391465 || keyword_kv(line).is_some()
466 || comment_text(line).is_some()
467 || drawer_begin_name(line).is_some()
392468 || is_list_item(t).is_some()
393469 || t.starts_with('|')
394470 || footnote_def_label(line).is_some()
395471 || heading_level(line).is_some()
396472}
397473
474// ---------------------------------------------------------------------------
475// Blocks, drawers, comments, affiliated keywords
476// ---------------------------------------------------------------------------
477
478/// Consume `#+BEGIN_<KIND> … #+END_<KIND>`. Matching is on the *specific* kind so a
479/// source block can sit inside a quote block; an unterminated block runs to end of
480/// input rather than failing.
481fn parse_block(lines: &[&str], start: usize, kind: &str, after: &str) -> (Element, usize) {
482 let mut inner: Vec<&str> = Vec::new();
483 let mut j = start + 1;
484 while j < lines.len() && !is_block_end_of(lines[j], kind) {
485 inner.push(lines[j]);
486 j += 1;
487 }
488 let next = if j < lines.len() { j + 1 } else { j };
489 let element = match kind.to_ascii_uppercase().as_str() {
490 "SRC" => {
491 let (lang, params) = parse_src_header(after);
492 Element::SrcBlock {
493 lang,
494 params,
495 code: inner.join("\n"),
496 }
497 }
498 "EXAMPLE" => Element::ExampleBlock(inner.join("\n")),
499 "QUOTE" => Element::QuoteBlock(parse_elements(&inner)),
500 "CENTER" => Element::CenterBlock(parse_elements(&inner)),
501 "EXPORT" => Element::ExportBlock {
502 backend: after.split_whitespace().next().unwrap_or("").to_string(),
503 raw: inner.join("\n"),
504 },
505 // Out-of-scope block types (verse, comment, ascii, custom) degrade to a verbatim
506 // example block: content preserved, no crash.
507 _ => Element::ExampleBlock(inner.join("\n")),
508 };
509 (element, next)
510}
511
512/// `:NAME:` … `:END:` at block level. A PROPERTIES drawer directly under a heading is
513/// consumed by [`parse_section_body`]; anything reaching here is a generic drawer,
514/// which the renderer drops (README §OUT).
515fn parse_drawer(lines: &[&str], start: usize, name: String) -> (Element, usize) {
516 let mut inner: Vec<&str> = Vec::new();
517 let mut j = start + 1;
518 while j < lines.len() && !lines[j].trim().eq_ignore_ascii_case(":END:") {
519 inner.push(lines[j]);
520 j += 1;
521 }
522 let next = if j < lines.len() { j + 1 } else { j };
523 (
524 Element::Drawer {
525 name,
526 content: parse_elements(&inner),
527 },
528 next,
529 )
530}
531
532/// The drawer name in a `:NAME:` opening line, if this line is one. `:END:` closes a
533/// drawer rather than opening one.
534fn drawer_begin_name(line: &str) -> Option<String> {
535 let t = line.trim();
536 if !is_drawer_begin(t) {
537 return None;
538 }
539 let name = &t[1..t.len() - 1];
540 if name.eq_ignore_ascii_case("END") {
541 return None;
542 }
543 Some(name.to_string())
544}
545
546/// A comment line: `#` followed by whitespace or nothing. `#+KEY:` is a keyword (checked
547/// first) and `#hashtag` is ordinary text.
548fn comment_text(line: &str) -> Option<String> {
549 let rest = line.trim_start().strip_prefix('#')?;
550 if rest.is_empty() {
551 return Some(String::new());
552 }
553 if !rest.starts_with(char::is_whitespace) {
554 return None;
555 }
556 Some(rest.trim().to_string())
557}
558
559/// Keywords that attach to the element that follows them rather than standing alone.
560fn is_affiliated(key: &str) -> bool {
561 let k = key.to_ascii_uppercase();
562 matches!(k.as_str(), "CAPTION" | "NAME" | "ATTR_HTML")
563}
564
565/// A paragraph holding nothing but an image link becomes a block-level figure when a
566/// `#+CAPTION:`/`#+ATTR_HTML:` precedes it. Affiliated keywords on anything else are
567/// parsed and dropped (README §IN covers captions for images only).
568fn attach_affiliated(element: Element, affiliated: Vec<(String, String)>) -> Element {
569 let value = |key: &str| {
570 affiliated
571 .iter()
572 .find(|(k, _)| k.eq_ignore_ascii_case(key))
573 .map(|(_, v)| v.clone())
574 };
575 let caption = value("CAPTION").unwrap_or_default();
576 let attrs = value("ATTR_HTML").unwrap_or_default();
577 if caption.is_empty() && attrs.is_empty() {
578 return element;
579 }
580 let Element::Paragraph(objs) = &element else {
581 return element;
582 };
583 let [Object::Link(link)] = objs.as_slice() else {
584 return element;
585 };
586 if !is_image_target(&link.target) {
587 return element;
588 }
589 Element::Figure {
590 link: link.clone(),
591 caption: inline(&caption),
592 attrs,
593 }
594}
595
596/// Does this link point at an image file? Drives both figure promotion and inline
597/// `<img>` rendering.
598pub fn is_image_target(target: &LinkTarget) -> bool {
599 let path = match target {
600 LinkTarget::File { path, .. } => path.as_str(),
601 LinkTarget::External(url) => url.split(['?', '#']).next().unwrap_or(url),
602 _ => return false,
603 };
604 let Some(ext) = path.rsplit('.').next() else {
605 return false;
606 };
607 matches!(
608 ext.to_ascii_lowercase().as_str(),
609 "png" | "jpg" | "jpeg" | "gif" | "svg" | "webp" | "avif"
610 )
611}
612
398613// ---------------------------------------------------------------------------
399614// Tables (spec §1 IN; `#+TBLFM:` formulas are parse-and-ignored via keyword_kv)
400615// ---------------------------------------------------------------------------
@@ -479,39 +694,142 @@ fn parse_footnote_def(
479694 (Element::FootnoteDefinition { label, content }, i)
480695}
481696
697/// Consume one plain list. Items are delimited by bullets at the list's own indent
698/// column; everything indented further is that item's body, re-parsed as block content —
699/// which is what makes lists nest. A single blank line does not end a list, but a blank
700/// line followed by anything that is not a sibling bullet does.
482701fn parse_list(lines: &[&str], start: usize) -> (List, usize) {
483 let kind = match is_list_item(lines[start].trim_start()) {
484 Some(Bullet::Ordered(_)) => ListKind::Ordered,
702 let base = indent_of(lines[start]);
703 let family = bullet_family(&is_list_item(lines[start].trim_start()).expect("list item"));
704 // A list is a description list when its FIRST item carries a `::` term separator.
705 let kind = match (&family, split_term(item_text(lines[start].trim_start()))) {
706 (ListKind::Ordered, _) => ListKind::Ordered,
707 (_, Some(_)) => ListKind::Description,
485708 _ => ListKind::Unordered,
486709 };
710
487711 let mut items = Vec::new();
488712 let mut i = start;
489 while i < lines.len() {
490 let t = lines[i].trim_start();
491 let bullet = match is_list_item(t) {
492 Some(b) => b,
493 None => break,
494 };
495 let item_kind = match bullet {
496 Bullet::Ordered(_) => ListKind::Ordered,
497 _ => ListKind::Unordered,
713 loop {
714 // Skip blank lines, but only stay in the list if a sibling bullet follows.
715 let mut j = i;
716 while j < lines.len() && lines[j].trim().is_empty() {
717 j += 1;
718 }
719 if j >= lines.len() || indent_of(lines[j]) != base {
720 break;
721 }
722 let Some(bullet) = is_list_item(lines[j].trim_start()) else {
723 break;
498724 };
499 if item_kind != kind {
725 if bullet_family(&bullet) != family {
500726 break;
501727 }
502 let rest = item_body(t, &bullet);
503 let (checkbox, text) = split_checkbox(rest);
728
729 // Body = the text after the bullet, plus every following line indented past the
730 // bullet column (blank lines included, so an item can hold several paragraphs).
731 let rest = item_body(lines[j].trim_start(), &bullet);
732 let (checkbox, rest) = split_checkbox(rest);
733 let (term, rest) = match kind {
734 ListKind::Description => match split_term(rest) {
735 Some((term, def)) => (Some(inline(term.trim())), def),
736 None => (None, rest),
737 },
738 _ => (None, rest),
739 };
740
741 let mut body: Vec<String> = vec![rest.trim().to_string()];
742 i = j + 1;
743 while i < lines.len() {
744 if lines[i].trim().is_empty() {
745 // Trailing blanks belong to the item only if more of it follows.
746 let mut k = i;
747 while k < lines.len() && lines[k].trim().is_empty() {
748 k += 1;
749 }
750 if k < lines.len() && indent_of(lines[k]) > base {
751 body.resize(body.len() + (k - i), String::new());
752 i = k;
753 continue;
754 }
755 break;
756 }
757 if indent_of(lines[i]) <= base {
758 break;
759 }
760 body.push(lines[i].to_string());
761 i += 1;
762 }
763
504764 items.push(ListItem {
505765 bullet,
506766 checkbox,
507 term: None, // description lists: out of scope for v0.1.
508 content: vec![Element::Paragraph(inline(text.trim()))],
767 term,
768 content: parse_elements(&dedent(&body)),
509769 });
510 i += 1;
511770 }
512771 (List { kind, items }, i)
513772}
514773
774/// Ordered and unordered bullets cannot share a list; description items use unordered
775/// bullets, so they are the same family.
776fn bullet_family(bullet: &Bullet) -> ListKind {
777 match bullet {
778 Bullet::Ordered(_) => ListKind::Ordered,
779 _ => ListKind::Unordered,
780 }
781}
782
783fn indent_of(line: &str) -> usize {
784 line.len() - line.trim_start().len()
785}
786
787/// Strip the common leading indent from an item's body lines so the recursive
788/// [`parse_elements`] call sees them at column zero. The first entry is already
789/// dedented (it is the text that followed the bullet), so it is excluded from the
790/// measurement.
791fn dedent(body: &[String]) -> Vec<&str> {
792 let common = body
793 .iter()
794 .skip(1)
795 .filter(|l| !l.trim().is_empty())
796 .map(|l| indent_of(l))
797 .min()
798 .unwrap_or(0);
799 body.iter()
800 .enumerate()
801 .map(|(idx, l)| {
802 if idx == 0 || l.len() < common {
803 l.as_str()
804 } else {
805 &l[common..]
806 }
807 })
808 .collect()
809}
810
811/// The text of a list item line after its bullet, for kind detection.
812fn item_text(t: &str) -> &str {
813 match is_list_item(t) {
814 Some(bullet) => item_body(t, &bullet),
815 None => t,
816 }
817}
818
819/// Split `term :: definition`. The separator must be surrounded by whitespace (or end
820/// the line) so `a::b` in code text is not mistaken for one.
821fn split_term(text: &str) -> Option<(&str, &str)> {
822 let idx = text.find(" :: ").or_else(|| {
823 text.strip_suffix(" ::")
824 .map(|before| before.len())
825 })?;
826 let term = &text[..idx];
827 if term.trim().is_empty() {
828 return None;
829 }
830 Some((term, text[idx..].trim_start_matches(" ::").trim_start()))
831}
832
515833/// Text of a list item after its bullet marker.
516834fn item_body<'a>(item: &'a str, bullet: &Bullet) -> &'a str {
517835 match bullet {
@@ -567,6 +885,15 @@ fn is_block_end(line: &str) -> bool {
567885 line.trim_start().to_ascii_uppercase().starts_with("#+END_")
568886}
569887
888/// Does this line close a block of exactly `kind`?
889fn is_block_end_of(line: &str, kind: &str) -> bool {
890 let upper = line.trim().to_ascii_uppercase();
891 match upper.strip_prefix("#+END_") {
892 Some(rest) => rest.trim() == kind.to_ascii_uppercase(),
893 None => false,
894 }
895}
896
570897fn parse_src_header(after: &str) -> (Option<String>, BlockParams) {
571898 let mut parts = after.splitn(2, char::is_whitespace);
572899 let lang = parts.next().filter(|s| !s.is_empty()).map(|s| s.to_string());
@@ -666,6 +993,14 @@ fn parse_inline_run(chars: &[char]) -> Vec<Object> {
666993 continue;
667994 }
668995 }
996 if c == '<' || c == '[' {
997 if let Some((obj, next)) = try_timestamp(chars, i) {
998 flush(&mut buf, &mut out);
999 out.push(obj);
1000 i = next;
1001 continue;
1002 }
1003 }
6691004 if is_scheme_start(chars, i) && boundary_before(chars, i) {
6701005 if let Some((obj, next)) = try_bare_url(chars, i) {
6711006 flush(&mut buf, &mut out);
@@ -818,6 +1153,87 @@ fn try_bare_url(chars: &[char], i: usize) -> Option<(Object, usize)> {
8181153 ))
8191154}
8201155
1156// ---------------------------------------------------------------------------
1157// Timestamps
1158// ---------------------------------------------------------------------------
1159
1160/// An org timestamp: `<2024-01-15 Mon>` (active) or `[2024-01-15 Mon]` (inactive), with
1161/// an optional `HH:MM` time, an optional `HH:MM-HH:MM` same-day range, and an optional
1162/// `--`-joined second stamp for a multi-day range.
1163fn try_timestamp(chars: &[char], i: usize) -> Option<(Object, usize)> {
1164 let active = chars[i] == '<';
1165 let (start, same_day_end, has_time, mut next) = parse_stamp(chars, i)?;
1166 let mut end = same_day_end;
1167 if end.is_none() && starts_with_at(chars, next, "--") {
1168 // A range's two halves must agree on activeness, or it is two adjacent stamps.
1169 if chars.get(next + 2) == Some(&chars[i]) {
1170 if let Some((stamp_end, _, _, after)) = parse_stamp(chars, next + 2) {
1171 end = Some(stamp_end);
1172 next = after;
1173 }
1174 }
1175 }
1176 Some((
1177 Object::Timestamp(Timestamp {
1178 active,
1179 start,
1180 end,
1181 has_time,
1182 }),
1183 next,
1184 ))
1185}
1186
1187/// One bracketed stamp → `(start, same-day end, has_time, index past the bracket)`.
1188/// Day names (`Mon`) and repeater/warning cookies (`+1w`, `-2d`) are recognized and
1189/// discarded — they carry no export meaning (README §OUT: agenda semantics).
1190fn parse_stamp(
1191 chars: &[char],
1192 i: usize,
1193) -> Option<(NaiveDateTime, Option<NaiveDateTime>, bool, usize)> {
1194 let open = *chars.get(i)?;
1195 let close = match open {
1196 '<' => '>',
1197 '[' => ']',
1198 _ => return None,
1199 };
1200 let end = (i + 1..chars.len()).find(|&k| chars[k] == close)?;
1201 let body: String = chars[i + 1..end].iter().collect();
1202 let mut parts = body.split_whitespace();
1203 let date = NaiveDate::parse_from_str(parts.next()?, "%Y-%m-%d").ok()?;
1204
1205 let mut has_time = false;
1206 let mut start_time = NaiveTime::MIN;
1207 let mut end_time = None;
1208 for part in parts {
1209 if let Some((from, to)) = parse_time_spec(part) {
1210 has_time = true;
1211 start_time = from;
1212 end_time = to;
1213 }
1214 }
1215 Some((
1216 date.and_time(start_time),
1217 end_time.map(|t| date.and_time(t)),
1218 has_time,
1219 end + 1,
1220 ))
1221}
1222
1223/// `HH:MM` or `HH:MM-HH:MM`.
1224fn parse_time_spec(s: &str) -> Option<(NaiveTime, Option<NaiveTime>)> {
1225 let (from, to) = match s.split_once('-') {
1226 Some((a, b)) => (a, Some(b)),
1227 None => (s, None),
1228 };
1229 let from = NaiveTime::parse_from_str(from, "%H:%M").ok()?;
1230 let to = match to {
1231 Some(b) => Some(NaiveTime::parse_from_str(b, "%H:%M").ok()?),
1232 None => None,
1233 };
1234 Some((from, to))
1235}
1236
8211237fn is_marker(c: char) -> bool {
8221238 matches!(c, '*' | '/' | '_' | '+' | '=' | '~')
8231239}
src/render.rs +338 −48
@@ -8,14 +8,23 @@
88//! and its cost must be cache-skippable (spec §4.2). Emit CSS classes, not inline
99//! styles, so themes live in the stylesheet (spec §3.2).
1010//!
11//! v0.2 renders: headings (always anchored, with tags), paragraphs, plain lists
12//! (unordered/ordered + checkboxes), source/example blocks, horizontal rules, tables
13//! (with header band from the rule row), footnotes, and inline markup. Real syntect
14//! tokenizing remains a `<pre><code>` passthrough for now (see [`SyntectHighlighter`]).
11//! Renders the v1 IN set: headings (always anchored, with TODO keyword, priority and
12//! tags), paragraphs, plain lists (unordered/ordered/description, nested, with
13//! checkboxes), tables, source blocks (syntect-highlighted), example/quote/center
14//! blocks, HTML export blocks, horizontal rules, images and captioned figures,
15//! footnotes, timestamps, and inline markup. Out-of-scope elements (generic drawers,
16//! comments, stray keywords, non-HTML export blocks) render to nothing.
1517
1618use std::collections::HashMap;
19use std::sync::OnceLock;
1720
18use crate::model::{Checkbox, Element, LinkTarget, ListKind, Object, Section, TableRow};
21use syntect::highlighting::ThemeSet;
22use syntect::html::{css_for_theme_with_class_style, ClassStyle, ClassedHTMLGenerator};
23use syntect::parsing::SyntaxSet;
24use syntect::util::LinesWithEndings;
25
26use crate::model::{Checkbox, Element, Link, LinkTarget, ListKind, Object, Section, TableRow};
27use crate::parser::is_image_target;
1928use crate::resolve::ResolvedDoc;
2029use crate::util::{plain_text, slugify};
2130
@@ -29,25 +38,94 @@ pub trait Highlighter {
2938 fn highlight(&self, code: &str, lang: Option<&str>) -> Html;
3039}
3140
32/// Default v1 highlighter. For now this is a plain `<pre><code>` passthrough that
33/// escapes the code and tags it with a `language-*` class; real syntect tokenizing
34/// to CSS-class spans is deferred (spec §3.2, §4.2).
35pub struct SyntectHighlighter;
41/// The class style used for both the emitted spans and the generated stylesheet. The
42/// two must agree or the CSS will not match the markup.
43const CLASS_STYLE: ClassStyle = ClassStyle::Spaced;
44
45/// The syntect theme whose colours become [`syntax_css`]. Mirrored in
46/// [`BuildConfig::highlighter_theme`](crate::incremental::BuildConfig) so a theme change
47/// flows into the config hash and invalidates every page.
48pub const SYNTAX_THEME: &str = "InspiredGitHub";
49
50/// Syntect's default syntax definitions, loaded once per process (loading is far more
51/// expensive than highlighting, and a site build highlights many blocks).
52fn syntax_set() -> &'static SyntaxSet {
53 static SET: OnceLock<SyntaxSet> = OnceLock::new();
54 SET.get_or_init(SyntaxSet::load_defaults_newlines)
55}
56
57/// The stylesheet the emitted highlight classes refer to. Highlighting emits CSS
58/// classes rather than inline styles (spec §3.2), so a build must also emit this.
59pub fn syntax_css() -> &'static str {
60 static CSS: OnceLock<String> = OnceLock::new();
61 CSS.get_or_init(|| {
62 let themes = ThemeSet::load_defaults();
63 themes
64 .themes
65 .get(SYNTAX_THEME)
66 .and_then(|theme| css_for_theme_with_class_style(theme, CLASS_STYLE).ok())
67 .unwrap_or_default()
68 })
69}
70
71/// The v1 highlighter: syntect tokenizing to CSS-class spans (spec §3.2, §4.2). A block
72/// whose language syntect does not know falls back to escaped `<pre><code>`.
73pub struct SyntectHighlighter {
74 syntaxes: &'static SyntaxSet,
75}
76
77impl SyntectHighlighter {
78 pub fn new() -> Self {
79 SyntectHighlighter {
80 syntaxes: syntax_set(),
81 }
82 }
83}
84
85impl Default for SyntectHighlighter {
86 fn default() -> Self {
87 Self::new()
88 }
89}
3690
3791impl Highlighter for SyntectHighlighter {
3892 fn highlight(&self, code: &str, lang: Option<&str>) -> Html {
39 let class = match lang {
40 Some(l) => format!(" class=\"language-{}\"", escape_attr(l)),
41 None => String::new(),
93 let Some(syntax) = lang.and_then(|l| self.syntaxes.find_syntax_by_token(l)) else {
94 return Html(plain_code(code, lang));
4295 };
96 let mut generator =
97 ClassedHTMLGenerator::new_with_class_style(syntax, self.syntaxes, CLASS_STYLE);
98 for line in LinesWithEndings::from(code) {
99 if generator
100 .parse_html_for_line_which_includes_newline(line)
101 .is_err()
102 {
103 return Html(plain_code(code, lang));
104 }
105 }
43106 Html(format!(
44 "<pre><code{}>{}</code></pre>\n",
45 class,
46 escape_html(code)
107 "<pre><code class=\"{} highlight\">{}</code></pre>\n",
108 language_class(lang),
109 generator.finalize()
47110 ))
48111 }
49112}
50113
114fn plain_code(code: &str, lang: Option<&str>) -> String {
115 format!(
116 "<pre><code class=\"{}\">{}</code></pre>\n",
117 language_class(lang),
118 escape_html(code)
119 )
120}
121
122fn language_class(lang: Option<&str>) -> String {
123 match lang {
124 Some(l) => format!("language-{}", escape_attr(l)),
125 None => "language-none".to_string(),
126 }
127}
128
51129/// Carries the highlighter plus the footnote collector across the tree walk (spec §2.4).
52130struct Renderer<'a> {
53131 hl: &'a dyn Highlighter,
@@ -91,7 +169,29 @@ impl Renderer<'_> {
91169 .clone()
92170 .or_else(|| h.id.clone())
93171 .unwrap_or_else(|| slugify(&plain_text(&h.title)));
94 out.push_str(&format!("<h{} id=\"{}\">", level, escape_attr(&anchor)));
172 // A heading with no title text has no meaningful slug; emit no `id` at all
173 // rather than a run of duplicate empty ones.
174 if anchor.is_empty() {
175 out.push_str(&format!("<h{}>", level));
176 } else {
177 out.push_str(&format!("<h{} id=\"{}\">", level, escape_attr(&anchor)));
178 }
179 // Keyword/priority markup mirrors Emacs' own HTML export classes, so output
180 // stays diffable against an `emacs --batch` oracle.
181 if let Some(todo) = &h.todo {
182 out.push_str(&format!(
183 "<span class=\"{} {}\">{}</span> ",
184 if todo.done { "done" } else { "todo" },
185 escape_attr(&todo.name),
186 escape_html(&todo.name)
187 ));
188 }
189 if let Some(priority) = h.priority {
190 out.push_str(&format!(
191 "<span class=\"priority\">[#{}]</span> ",
192 escape_html(&priority.to_string())
193 ));
194 }
95195 self.render_objects(&h.title, out);
96196 for tag in &h.tags {
97197 out.push_str(&format!(" <span class=\"tag\">{}</span>", escape_html(tag)));
@@ -113,33 +213,7 @@ impl Renderer<'_> {
113213 self.render_objects(objs, out);
114214 out.push_str("</p>\n");
115215 }
116 Element::List(list) => {
117 let tag = match list.kind {
118 ListKind::Ordered => "ol",
119 _ => "ul",
120 };
121 out.push_str(&format!("<{}>\n", tag));
122 for item in &list.items {
123 out.push_str("<li>");
124 if let Some(cb) = &item.checkbox {
125 let checked = matches!(cb, Checkbox::On);
126 out.push_str(&format!(
127 "<input type=\"checkbox\" disabled{}> ",
128 if checked { " checked" } else { "" }
129 ));
130 }
131 match item.content.as_slice() {
132 [Element::Paragraph(objs)] => self.render_objects(objs, out),
133 els => {
134 for el in els {
135 self.render_element(el, out);
136 }
137 }
138 }
139 out.push_str("</li>\n");
140 }
141 out.push_str(&format!("</{}>\n", tag));
142 }
216 Element::List(list) => self.render_list(list, out),
143217 Element::Table(table) => self.render_table(table, out),
144218 Element::SrcBlock { lang, code, .. } => {
145219 let Html(h) = self.hl.highlight(code, lang.as_deref());
@@ -148,12 +222,108 @@ impl Renderer<'_> {
148222 Element::ExampleBlock(code) => {
149223 out.push_str(&format!("<pre>{}</pre>\n", escape_html(code)));
150224 }
225 Element::QuoteBlock(inner) => {
226 out.push_str("<blockquote>\n");
227 for el in inner {
228 self.render_element(el, out);
229 }
230 out.push_str("</blockquote>\n");
231 }
232 Element::CenterBlock(inner) => {
233 out.push_str("<div class=\"center\">\n");
234 for el in inner {
235 self.render_element(el, out);
236 }
237 out.push_str("</div>\n");
238 }
239 // An `html` export block is verbatim output by definition; every other
240 // backend is out of scope and drops (README §OUT).
241 Element::ExportBlock { backend, raw } => {
242 if backend.eq_ignore_ascii_case("html") {
243 out.push_str(raw);
244 out.push('\n');
245 }
246 }
247 Element::Figure {
248 link,
249 caption,
250 attrs,
251 } => {
252 out.push_str("<figure>");
253 out.push_str(&image_tag(link, attrs, &plain_text(caption)));
254 if !caption.is_empty() {
255 out.push_str("<figcaption>");
256 self.render_objects(caption, out);
257 out.push_str("</figcaption>");
258 }
259 out.push_str("</figure>\n");
260 }
151261 Element::HorizontalRule => out.push_str("<hr>\n"),
152262 // Definitions are emitted in the footnotes section, not inline.
153263 Element::FootnoteDefinition { .. } => {}
154 // Out of scope (non-HTML export, generic drawers, stray keywords, comments):
155 // emitted as nothing rather than crashing.
156 _ => {}
264 // Out of scope (generic drawers, stray keywords, comments): emitted as
265 // nothing rather than crashing.
266 Element::Drawer { .. } | Element::Keyword { .. } | Element::Comment(_) => {}
267 }
268 }
269
270 fn render_list(&mut self, list: &crate::model::List, out: &mut String) {
271 if list.kind == ListKind::Description {
272 out.push_str("<dl>\n");
273 for item in &list.items {
274 out.push_str("<dt>");
275 if let Some(term) = &item.term {
276 self.render_objects(term, out);
277 }
278 out.push_str("</dt>\n<dd>");
279 self.render_item_content(&item.content, out);
280 out.push_str("</dd>\n");
281 }
282 out.push_str("</dl>\n");
283 return;
284 }
285 let tag = if list.kind == ListKind::Ordered {
286 "ol"
287 } else {
288 "ul"
289 };
290 out.push_str(&format!("<{}>\n", tag));
291 for item in &list.items {
292 out.push_str("<li>");
293 if let Some(cb) = &item.checkbox {
294 out.push_str(&format!(
295 "<input type=\"checkbox\" disabled{}> ",
296 if matches!(cb, Checkbox::On) {
297 " checked"
298 } else {
299 ""
300 }
301 ));
302 }
303 self.render_item_content(&item.content, out);
304 out.push_str("</li>\n");
305 }
306 out.push_str(&format!("</{}>\n", tag));
307 }
308
309 /// A single-paragraph item renders its text bare — `<li>text<ul>…` rather than
310 /// `<li><p>text</p><ul>…` — which is what org does and what makes a nested list read
311 /// as a continuation of its parent item. An item holding *several* paragraphs wraps
312 /// them all, so they do not run together.
313 fn render_item_content(&mut self, content: &[Element], out: &mut String) {
314 let lead_is_bare = matches!(content.first(), Some(Element::Paragraph(_)))
315 && !content[1..]
316 .iter()
317 .any(|el| matches!(el, Element::Paragraph(_)));
318 let mut rest = content;
319 if lead_is_bare {
320 if let Some((Element::Paragraph(objs), tail)) = content.split_first() {
321 self.render_objects(objs, out);
322 rest = tail;
323 }
324 }
325 for el in rest {
326 self.render_element(el, out);
157327 }
158328 }
159329
@@ -224,6 +394,10 @@ impl Renderer<'_> {
224394 out.push_str(&format!("<code class=\"verbatim\">{}</code>", escape_html(s)))
225395 }
226396 Object::Code(s) => out.push_str(&format!("<code>{}</code>", escape_html(s))),
397 // A description-less link to an image is the image itself, not a link to it.
398 Object::Link(link) if link.description.is_none() && is_image_target(&link.target) => {
399 out.push_str(&image_tag(link, "", ""));
400 }
227401 Object::Link(link) => {
228402 let href = link_href(&link.target);
229403 out.push_str(&format!("<a href=\"{}\">", escape_attr(&href)));
@@ -252,8 +426,8 @@ impl Renderer<'_> {
252426 ));
253427 }
254428 Object::LineBreak => out.push_str("<br>\n"),
255 // Timestamps, entities: out of scope for now.
256 _ => {}
429 Object::Timestamp(ts) => out.push_str(&timestamp_html(ts)),
430 Object::Entity(e) => out.push_str(&escape_html(e)),
257431 }
258432 }
259433
@@ -298,6 +472,122 @@ fn collect_defs_in(elements: &[Element], defs: &mut HashMap<String, Vec<Element>
298472 }
299473}
300474
475/// An `<img>` for an image link, carrying any `#+ATTR_HTML:` attributes and falling back
476/// to the caption for alt text — but only when the author did not write an `:alt` of
477/// their own, since two `alt` attributes on one tag is invalid HTML.
478fn image_tag(link: &Link, attrs: &str, alt: &str) -> String {
479 let mut pairs = attr_html(attrs);
480 if !pairs.iter().any(|(k, _)| k.eq_ignore_ascii_case("alt")) {
481 pairs.insert(0, ("alt".to_string(), alt.to_string()));
482 }
483 let attributes: String = pairs
484 .iter()
485 .map(|(k, v)| format!(" {}=\"{}\"", escape_attr(k), escape_attr(v)))
486 .collect();
487 format!(
488 "<img src=\"{}\"{}>",
489 escape_attr(&link_href(&link.target)),
490 attributes
491 )
492}
493
494/// `#+ATTR_HTML: :width 400 :class hero` → `[(width, 400), (class, hero)]`. Values run to
495/// the next `:key` token and may be double-quoted to include spaces. A malformed spec
496/// contributes nothing rather than emitting broken markup.
497fn attr_html(spec: &str) -> Vec<(String, String)> {
498 let mut out = Vec::new();
499 let mut key: Option<&str> = None;
500 let mut value = String::new();
501 let mut quoted: Option<String> = None;
502
503 let flush = |out: &mut Vec<(String, String)>, key: &mut Option<&str>, value: &mut String| {
504 if let Some(k) = key.take() {
505 out.push((k.to_string(), value.trim().to_string()));
506 }
507 value.clear();
508 };
509
510 for token in spec.split_whitespace() {
511 // Inside a quoted value, everything up to the closing quote is literal.
512 if let Some(buf) = &mut quoted {
513 buf.push(' ');
514 buf.push_str(token.trim_end_matches('"'));
515 if token.ends_with('"') {
516 value = quoted.take().expect("quoted value in progress");
517 }
518 continue;
519 }
520 if let Some(k) = token.strip_prefix(':') {
521 flush(&mut out, &mut key, &mut value);
522 if !k.is_empty() {
523 key = Some(k);
524 }
525 continue;
526 }
527 if key.is_none() {
528 continue;
529 }
530 if let Some(rest) = token.strip_prefix('"') {
531 if let Some(inner) = rest.strip_suffix('"') {
532 value = inner.to_string();
533 } else {
534 quoted = Some(rest.to_string());
535 }
536 continue;
537 }
538 if !value.is_empty() {
539 value.push(' ');
540 }
541 value.push_str(token);
542 }
543 if let Some(buf) = quoted {
544 value = buf;
545 }
546 flush(&mut out, &mut key, &mut value);
547 out
548}
549
550/// `<time>` markup for a timestamp. A range emits both endpoints; a same-day range
551/// abbreviates its end to just the time.
552fn timestamp_html(ts: &crate::model::Timestamp) -> String {
553 let class = if ts.active {
554 "timestamp"
555 } else {
556 "timestamp inactive"
557 };
558 let one = |dt: &chrono::NaiveDateTime, text: String| {
559 let attr = if ts.has_time {
560 dt.format("%Y-%m-%dT%H:%M").to_string()
561 } else {
562 dt.format("%Y-%m-%d").to_string()
563 };
564 format!(
565 "<time class=\"{class}\" datetime=\"{}\">{}</time>",
566 escape_attr(&attr),
567 escape_html(&text)
568 )
569 };
570 let text_of = |dt: &chrono::NaiveDateTime| {
571 if ts.has_time {
572 dt.format("%Y-%m-%d %H:%M").to_string()
573 } else {
574 dt.format("%Y-%m-%d").to_string()
575 }
576 };
577
578 let mut out = one(&ts.start, text_of(&ts.start));
579 if let Some(end) = &ts.end {
580 out.push_str("&#8211;");
581 let text = if end.date() == ts.start.date() && ts.has_time {
582 end.format("%H:%M").to_string()
583 } else {
584 text_of(end)
585 };
586 out.push_str(&one(end, text));
587 }
588 out
589}
590
301591/// Best-effort URL for a link target. After RESOLVE, internal targets have been
302592/// rewritten to `External` with their final URL; anything still internal here is an
303593/// unresolved link, rendered to a plausible anchor so the page stays self-consistent.
src/resolve.rs +6
@@ -148,6 +148,12 @@ impl Cx<'_> {
148148 LinkTarget::Id(id) => TargetId::Id(id.clone()),
149149 LinkTarget::Heading(t) => TargetId::Heading(t.clone()),
150150 LinkTarget::File { path, .. } => {
151 // Only `.org` files are pages. A link to an asset (an image, a PDF) is
152 // already a correct relative URL in the output tree, since assets are
153 // copied preserving layout — so it is neither resolved nor reported.
154 if path.extension() != Some("org") {
155 return;
156 }
151157 TargetId::File(normalize_link_path(self.from, path))
152158 }
153159 };
src/site.rs +17 −7
@@ -23,10 +23,10 @@ use crate::incremental::{
2323use crate::index::{document_targets, SymbolTable, TargetId};
2424use crate::model::{ContentHash, Document};
2525use crate::parser::parse;
26use crate::render::{render, Html, SyntectHighlighter};
26use crate::render::{render, syntax_css, Html, SyntectHighlighter};
2727use crate::resolve::resolve;
2828use crate::template::{template_sources, NavItem, Templater};
29use crate::util::output_url;
29use crate::util::{output_url, relative_root};
3030
3131/// A fully built page: source and output paths (relative to their roots) and its
3232/// final templated HTML.
@@ -142,7 +142,7 @@ fn prepare_pages(src: &Utf8Path) -> Result<(Vec<PagePrep>, SymbolTable)> {
142142/// touching the output directory. Shared by the tests (full render, every page).
143143pub fn render_site(src: &Utf8Path) -> Result<(Vec<BuiltPage>, BrokenLinks)> {
144144 let (preps, _symbols) = prepare_pages(src)?;
145 let highlighter = SyntectHighlighter;
145 let highlighter = SyntectHighlighter::new();
146146 let templater = Templater::new();
147147
148148 let mut pages = Vec::new();
@@ -169,11 +169,15 @@ fn render_page(
169169 p: &PagePrep,
170170) -> Result<String> {
171171 let Html(fragment) = render(&p.resolved, highlighter);
172 let stylesheet = format!("{}{}", relative_root(&p.source), SYNTAX_STYLESHEET);
172173 templater
173 .render_page(&p.title, &fragment, &p.nav)
174 .render_page(&p.title, &fragment, &p.nav, &stylesheet)
174175 .with_context(|| format!("templating {}", p.source))
175176}
176177
178/// Site-root-relative name of the generated syntax stylesheet. Every page links to it.
179pub const SYNTAX_STYLESHEET: &str = "syntax.css";
180
177181/// Full site build with the incremental layer (spec §4). Renders only the pages whose
178182/// `render_key` changed or that link into a changed file's targets; reuses the on-disk
179183/// output of everything else; persists an updated cache manifest.
@@ -243,7 +247,7 @@ pub fn build_site(src: &Utf8Path, out: &Utf8Path, opts: &BuildOptions) -> Result
243247 }
244248 }
245249
246 let highlighter = SyntectHighlighter;
250 let highlighter = SyntectHighlighter::new();
247251 let templater = Templater::new();
248252 let mut report = SiteReport::default();
249253
@@ -267,6 +271,12 @@ pub fn build_site(src: &Utf8Path, out: &Utf8Path, opts: &BuildOptions) -> Result
267271 }
268272 }
269273
274 // The syntax stylesheet the highlighter's CSS classes refer to. Written every build
275 // (it is a few KB and depends only on the theme, which lives in the config hash).
276 fs::create_dir_all(out).with_context(|| format!("creating {out}"))?;
277 fs::write(out.join(SYNTAX_STYLESHEET), syntax_css())
278 .with_context(|| format!("writing {SYNTAX_STYLESHEET} under {out}"))?;
279
270280 // Assets are a dumb copy in v0.3 (spec §8 Q11): copy every run. Cheap, and keeps the
271281 // full-vs-incremental byte equivalence trivially true for non-`.org` files.
272282 for rel in &assets {
@@ -295,7 +305,7 @@ pub fn build_site(src: &Utf8Path, out: &Utf8Path, opts: &BuildOptions) -> Result
295305
296306 if opts.strict && !report.broken.is_empty() {
297307 for (page, target) in &report.broken {
298 eprintln!("error: unresolved link in {page}: {target:?}");
308 eprintln!("error: {page}: unresolved link {target}");
299309 }
300310 anyhow::bail!(
301311 "{} unresolved internal link(s) under --strict",
@@ -303,7 +313,7 @@ pub fn build_site(src: &Utf8Path, out: &Utf8Path, opts: &BuildOptions) -> Result
303313 );
304314 }
305315 for (page, target) in &report.broken {
306 eprintln!("warning: unresolved link in {page}: {target:?}");
316 eprintln!("warning: {page}: unresolved link {target}");
307317 }
308318
309319 Ok(report)
src/template.rs +8 −2
@@ -22,6 +22,9 @@ const BASE_TEMPLATE: &str = r#"<!DOCTYPE html>
2222<head>
2323<meta charset="utf-8">
2424<title>{{ title }}</title>
25{%- if stylesheet %}
26<link rel="stylesheet" href="{{ stylesheet }}">
27{%- endif %}
2528</head>
2629<body>
2730<nav>
@@ -62,18 +65,21 @@ impl Templater {
6265 Templater { env }
6366 }
6467
65 /// fragment + page metadata → full HTML page.
68 /// fragment + page metadata → full HTML page. `stylesheet` is the URL of the
69 /// syntax-highlighting stylesheet relative to *this* page (highlighting emits CSS
70 /// classes, so the sheet has to come with it).
6671 pub fn render_page(
6772 &self,
6873 title: &str,
6974 body: &str,
7075 nav: &[NavItem],
76 stylesheet: &str,
7177 ) -> Result<String, TemplateError> {
7278 let tmpl = self
7379 .env
7480 .get_template("base")
7581 .map_err(|e| TemplateError::Render(e.to_string()))?;
76 tmpl.render(context! { title => title, body => body, nav => nav })
82 tmpl.render(context! { title => title, body => body, nav => nav, stylesheet => stylesheet })
7783 .map_err(|e| TemplateError::Render(e.to_string()))
7884 }
7985}
src/util.rs +10
@@ -78,6 +78,16 @@ pub fn output_url(from_rel: &Utf8Path, to_rel: &Utf8Path, anchor: Option<&str>)
7878 }
7979}
8080
81/// The `../`-prefix that reaches the site root from the page at `from_rel`. Empty for a
82/// top-level page. Used for site-global assets like the syntax stylesheet.
83pub fn relative_root(from_rel: &Utf8Path) -> String {
84 let depth = from_rel
85 .parent()
86 .map(|p| p.components().count())
87 .unwrap_or(0);
88 "../".repeat(depth)
89}
90
8191/// Relative path from `from_dir` to `to`, using `../` where needed. `/`-joined for URLs.
8292fn relative_path(from_dir: &Utf8Path, to: &Utf8Path) -> String {
8393 let from_c: Vec<&str> = from_dir.components().map(|c| c.as_str()).collect();
tests/constructs.rs added +292
@@ -0,0 +1,292 @@
1//! Golden-file coverage of the v1 scope line (README §"v1 scope").
2//!
3//! Two halves, and the second is the point:
4//!
5//! - **IN** — every construct the v1 scope claims gets an element-tree snapshot (parser
6//! correctness) and a rendered-HTML snapshot (renderer correctness).
7//! - **OUT** — every construct the v1 scope explicitly excludes gets an assertion that it
8//! *degrades predictably*: parsed and ignored, content preserved where that is the
9//! honest fallback, never a crash and never a half-rendered artifact.
10//!
11//! The OUT half is the scope guardrail: it is what defends against this project's stated
12//! #1 risk, creeping back toward all-of-org.
13
14use camino::Utf8PathBuf;
15
16use org_ssg::model::Document;
17use org_ssg::parser::parse;
18use org_ssg::render::{render, Html, SyntectHighlighter};
19use org_ssg::resolve::ResolvedDoc;
20
21fn parse_fixture(name: &str) -> Document {
22 let path = Utf8PathBuf::from(env!("CARGO_MANIFEST_DIR"))
23 .join("fixtures")
24 .join(name);
25 let source = std::fs::read_to_string(&path).expect("read fixture");
26 // A stable relative path keeps snapshots free of absolute machine paths.
27 parse(Utf8PathBuf::from("fixtures").join(name).as_path(), &source).expect("parse fixture")
28}
29
30fn render_fixture(name: &str) -> String {
31 let document = parse_fixture(name);
32 let Html(html) = render(&ResolvedDoc { document }, &SyntectHighlighter::new());
33 html
34}
35
36// ---------------------------------------------------------------------------
37// IN: the constructs v1 promises to handle
38// ---------------------------------------------------------------------------
39
40#[test]
41fn headings_element_tree() {
42 insta::assert_json_snapshot!(parse_fixture("headings.org").root);
43}
44
45#[test]
46fn headings_html() {
47 insta::assert_snapshot!(render_fixture("headings.org"));
48}
49
50#[test]
51fn lists_element_tree() {
52 insta::assert_json_snapshot!(parse_fixture("lists.org").root);
53}
54
55#[test]
56fn lists_html() {
57 insta::assert_snapshot!(render_fixture("lists.org"));
58}
59
60#[test]
61fn blocks_html() {
62 insta::assert_snapshot!(render_fixture("blocks.org"));
63}
64
65#[test]
66fn timestamps_element_tree() {
67 insta::assert_json_snapshot!(parse_fixture("timestamps.org").root);
68}
69
70#[test]
71fn timestamps_html() {
72 insta::assert_snapshot!(render_fixture("timestamps.org"));
73}
74
75#[test]
76fn images_html() {
77 insta::assert_snapshot!(render_fixture("images.org"));
78}
79
80/// A TODO keyword is a whole word from the configured set, not a prefix: `TODOs are not
81/// a keyword` is a plain title. This is the boundary rule most likely to regress.
82#[test]
83fn todo_keyword_requires_a_word_boundary() {
84 let html = render_fixture("headings.org");
85 assert!(
86 html.contains("<span class=\"todo TODO\">TODO</span> "),
87 "a real TODO keyword is marked up:\n{html}"
88 );
89 assert!(
90 !html.contains("<span class=\"todo TODO\">TODO</span> s are"),
91 "`TODOs` must not be split into a keyword plus a title:\n{html}"
92 );
93}
94
95/// Nesting is by indentation, so a nested list must land *inside* its parent `<li>`.
96#[test]
97fn nested_list_is_nested_in_the_parent_item() {
98 let html = render_fixture("lists.org");
99 assert!(
100 html.contains("<li>outer item<ul>"),
101 "an indented sub-list belongs to the item above it:\n{html}"
102 );
103}
104
105/// The caption supplies alt text, but an explicit `:alt` must win — emitting both
106/// would put two `alt` attributes on one tag.
107#[test]
108fn explicit_alt_attribute_replaces_the_caption_derived_one() {
109 let html = render_fixture("images.org");
110 assert!(
111 html.contains("<img src=\"cat.jpg\" alt=\"a cat, sitting\" loading=\"lazy\">"),
112 "a quoted `:alt` should be the only alt attribute:\n{html}"
113 );
114 for line in html.lines() {
115 assert!(
116 line.matches(" alt=").count() <= 1,
117 "no tag may carry two alt attributes:\n{line}"
118 );
119 }
120}
121
122/// Keywords in the file preamble are *copied* into the metadata map, not removed from
123/// the body. Removing them used to merge the paragraphs either side of a keyword and
124/// strand `#+CAPTION:` away from the image below it — content damage from a metadata
125/// step, in the one region of a file where every real document has keywords.
126#[test]
127fn preamble_keywords_do_not_disturb_the_content_around_them() {
128 let source = "#+TITLE: T\n\nOne.\n#+SOMEKEY: v\nTwo.\n\n#+CAPTION: shot\n[[file:a.png]]\n";
129 let document = parse(Utf8PathBuf::from("t.org").as_path(), source).expect("parse");
130 assert!(
131 document
132 .keywords
133 .entries
134 .iter()
135 .any(|(k, v)| k == "TITLE" && v == "T"),
136 "document metadata is still collected"
137 );
138 let Html(html) = render(&ResolvedDoc { document }, &SyntectHighlighter::new());
139 assert!(
140 html.contains("<p>One.</p>") && html.contains("<p>Two.</p>"),
141 "a keyword between two paragraphs must not merge them:\n{html}"
142 );
143 assert!(
144 html.contains("<figcaption>shot</figcaption>"),
145 "a preamble `#+CAPTION:` must still attach to the image below it:\n{html}"
146 );
147}
148
149/// Highlighting must emit CSS classes, never inline styles, so themes live in the
150/// stylesheet (spec §3.2) — and the stylesheet the classes refer to must exist.
151#[test]
152fn highlighting_emits_classes_not_inline_styles() {
153 let html = render_fixture("blocks.org");
154 assert!(
155 html.contains("<span class=\"storage type function python\">"),
156 "python source should be tokenized into classed spans:\n{html}"
157 );
158 assert!(
159 !html.contains("style=\""),
160 "highlighting must not emit inline styles:\n{html}"
161 );
162 assert!(
163 org_ssg::render::syntax_css().contains(".storage"),
164 "the generated stylesheet must define the emitted classes"
165 );
166}
167
168/// An unknown language is not an error: the block keeps its content, escaped.
169#[test]
170fn unknown_source_language_falls_back_to_plain_code() {
171 let doc = "#+BEGIN_SRC nosuchlang\n<not markup> & such\n#+END_SRC\n";
172 let document = parse(Utf8PathBuf::from("t.org").as_path(), doc).expect("parse");
173 let Html(html) = render(&ResolvedDoc { document }, &SyntectHighlighter::new());
174 assert_eq!(
175 html,
176 "<pre><code class=\"language-nosuchlang\">&lt;not markup&gt; &amp; such</code></pre>\n"
177 );
178}
179
180// ---------------------------------------------------------------------------
181// OUT: the constructs v1 explicitly excludes must degrade, not explode
182// ---------------------------------------------------------------------------
183
184/// The whole OUT fixture parses and renders. This is the crash gate.
185#[test]
186fn out_of_scope_fixture_renders_without_crashing() {
187 let html = render_fixture("outofscope.org");
188 assert!(!html.is_empty(), "an out-of-scope document still renders");
189}
190
191#[test]
192fn out_of_scope_html() {
193 insta::assert_snapshot!(render_fixture("outofscope.org"));
194}
195
196/// Babel is never executed and `#+RESULTS:` blocks are never trusted: the source block
197/// renders as code, and its stale results do not reach the page.
198#[test]
199fn babel_is_not_executed_and_results_are_dropped() {
200 let html = render_fixture("outofscope.org");
201 assert!(
202 html.contains("the block renders; :results is never executed"),
203 "the source block itself still renders:\n{html}"
204 );
205 assert!(
206 !html.contains("stale output from a previous evaluation"),
207 "a `#+RESULTS:` block must not be emitted:\n{html}"
208 );
209}
210
211/// `#+TBLFM:` is inert: the table renders with the values as written, and the formula
212/// is neither evaluated nor printed.
213#[test]
214fn table_formulas_are_inert() {
215 let html = render_fixture("outofscope.org");
216 assert!(html.contains("<table>"), "the table still renders:\n{html}");
217 assert!(
218 !html.contains("vsum"),
219 "the `#+TBLFM:` formula must not reach the page:\n{html}"
220 );
221}
222
223/// LaTeX, macros and radio targets have no v1 semantics, so they survive as the literal
224/// text the author typed — lossless, and obviously unhandled to a reader.
225#[test]
226fn latex_macros_and_radio_targets_stay_literal() {
227 let html = render_fixture("outofscope.org");
228 for literal in ["$x^2 + y^2$", "E = mc^2", "{{{author}}}", "\\alpha"] {
229 assert!(
230 html.contains(literal),
231 "`{literal}` should survive as literal text:\n{html}"
232 );
233 }
234}
235
236/// Drawers other than PROPERTIES are captured by the parser and dropped by the
237/// renderer — including LOGBOOK clock lines, which are agenda state, not content.
238#[test]
239fn drawers_are_parsed_and_dropped() {
240 let html = render_fixture("outofscope.org");
241 assert!(
242 !html.contains("CLOCK:"),
243 "LOGBOOK contents must not be emitted:\n{html}"
244 );
245 assert!(
246 !html.contains("Drawer contents are captured and dropped"),
247 "generic drawer contents must not be emitted:\n{html}"
248 );
249}
250
251/// A non-HTML export block is dropped whole: emitting LaTeX into an HTML page would be
252/// worse than emitting nothing.
253#[test]
254fn non_html_export_blocks_are_dropped() {
255 let html = render_fixture("blocks.org");
256 assert!(
257 html.contains("<aside class=\"raw\">Raw HTML passes through.</aside>"),
258 "an `html` export block passes through verbatim:\n{html}"
259 );
260 assert!(
261 !html.contains("\\emph"),
262 "a `latex` export block must be dropped:\n{html}"
263 );
264}
265
266/// An unmodelled block type keeps its content rather than vanishing.
267#[test]
268fn unknown_block_types_keep_their_content() {
269 let html = render_fixture("outofscope.org");
270 assert!(
271 html.contains("An unmodelled block type"),
272 "a verse block degrades to a verbatim example block:\n{html}"
273 );
274}
275
276/// `#+INCLUDE:` is not expanded — the build must not silently pull in another file.
277#[test]
278fn include_is_not_expanded() {
279 let doc = parse_fixture("outofscope.org");
280 assert!(
281 doc.keywords
282 .entries
283 .iter()
284 .any(|(k, _)| k.eq_ignore_ascii_case("INCLUDE")),
285 "`#+INCLUDE:` is captured as an inert keyword"
286 );
287 let html = render_fixture("outofscope.org");
288 assert!(
289 !html.contains("other.org"),
290 "`#+INCLUDE:` must not be expanded or echoed:\n{html}"
291 );
292}
tests/pipeline.rs +1 −1
@@ -23,7 +23,7 @@ fn parse_fixture(name: &str) -> Document {
2323fn render_fixture(name: &str) -> String {
2424 let document = parse_fixture(name);
2525 let resolved = ResolvedDoc { document };
26 let Html(html) = render(&resolved, &SyntectHighlighter);
26 let Html(html) = render(&resolved, &SyntectHighlighter::new());
2727 html
2828}
2929
tests/site.rs +1 −1
@@ -33,7 +33,7 @@ fn render_fragment(name: &str) -> String {
3333 let path = fixtures().join(name);
3434 let source = std::fs::read_to_string(&path).expect("read fixture");
3535 let document = parse(Utf8PathBuf::from(name).as_path(), &source).expect("parse");
36 let Html(html) = render(&ResolvedDoc { document }, &SyntectHighlighter);
36 let Html(html) = render(&ResolvedDoc { document }, &SyntectHighlighter::new());
3737 html
3838}
3939
tests/snapshots/constructs__blocks_html.snap added +27
@@ -0,0 +1,27 @@
1---
2source: tests/constructs.rs
3expression: "render_fixture(\"blocks.org\")"
4---
5<h1 id="quote">Quote</h1>
6<blockquote>
7<p>A quoted paragraph with <em>markup</em>.</p>
8<p>And a second paragraph.</p>
9</blockquote>
10<h1 id="center">Center</h1>
11<div class="center">
12<p>Centred text.</p>
13</div>
14<h1 id="example">Example</h1>
15<pre>Verbatim *not bold* text.
16 Indentation preserved.</pre>
17<h1 id="export">Export</h1>
18<aside class="raw">Raw HTML passes through.</aside>
19<h1 id="source">Source</h1>
20<pre><code class="language-python highlight"><span class="source python"><span class="meta function python"><span class="storage type function python">def</span> <span class="entity name function python"><span class="meta generic-name python">greet</span></span></span><span class="meta function parameters python"><span class="punctuation section parameters begin python">(</span></span><span class="meta function parameters python"><span class="variable parameter python">name</span><span class="punctuation section parameters end python">)</span></span><span class="meta function python"><span class="punctuation section function begin python">:</span></span>
21 <span class="keyword control flow return python">return</span> <span class="storage type string python">f</span><span class="meta string interpolated python"><span class="string quoted double python"><span class="punctuation definition string begin python">&quot;</span></span></span><span class="meta string interpolated python"><span class="string quoted double python">hello </span><span class="meta interpolation python"><span class="punctuation section interpolation begin python">{</span><span class="source python embedded"><span class="meta qualified-name python"><span class="meta generic-name python">name</span></span></span></span><span class="meta interpolation python"><span class="punctuation section interpolation end python">}</span></span><span class="string quoted double python"><span class="punctuation definition string end python">&quot;</span></span></span></span></code></pre>
22<pre><code class="language-none">plain block, no language</code></pre>
23<h1 id="nested">Nested</h1>
24<blockquote>
25<p>A quote containing a source block:</p>
26<pre><code class="language-sh highlight"><span class="source shell bash"><span class="meta function-call shell"><span class="support function echo shell">echo</span></span><span class="meta function-call arguments shell"> hi</span></span></code></pre>
27</blockquote>
tests/snapshots/constructs__headings_element_tree.snap added +175
@@ -0,0 +1,175 @@
1---
2source: tests/constructs.rs
3expression: "parse_fixture(\"headings.org\").root"
4---
5{
6 "heading": null,
7 "content": [
8 {
9 "Keyword": {
10 "key": "TITLE",
11 "value": "Heading Metadata"
12 }
13 }
14 ],
15 "children": [
16 {
17 "heading": {
18 "level": 1,
19 "todo": {
20 "name": "TODO",
21 "done": false
22 },
23 "priority": "A",
24 "title": [
25 {
26 "Text": "Write the parser"
27 }
28 ],
29 "tags": [
30 "work",
31 "rust"
32 ],
33 "properties": {
34 "entries": [
35 [
36 "CUSTOM_ID",
37 "write-parser"
38 ],
39 [
40 "OWNER",
41 "nobody"
42 ]
43 ]
44 },
45 "id": null,
46 "custom_id": "write-parser"
47 },
48 "content": [
49 {
50 "Paragraph": [
51 {
52 "Text": "A heading carrying a keyword, a priority, tags and a property drawer."
53 }
54 ]
55 }
56 ],
57 "children": [
58 {
59 "heading": {
60 "level": 2,
61 "todo": {
62 "name": "DONE",
63 "done": true
64 },
65 "priority": null,
66 "title": [
67 {
68 "Text": "Nested and finished"
69 }
70 ],
71 "tags": [],
72 "properties": {
73 "entries": []
74 },
75 "id": null,
76 "custom_id": null
77 },
78 "content": [
79 {
80 "Paragraph": [
81 {
82 "Text": "Sub-headings nest by star count."
83 }
84 ]
85 }
86 ],
87 "children": []
88 },
89 {
90 "heading": {
91 "level": 2,
92 "todo": null,
93 "priority": "C",
94 "title": [
95 {
96 "Text": "Priority without a keyword"
97 }
98 ],
99 "tags": [],
100 "properties": {
101 "entries": []
102 },
103 "id": null,
104 "custom_id": null
105 },
106 "content": [
107 {
108 "Paragraph": [
109 {
110 "Text": "A priority cookie can stand alone."
111 }
112 ]
113 }
114 ],
115 "children": []
116 }
117 ]
118 },
119 {
120 "heading": {
121 "level": 1,
122 "todo": null,
123 "priority": null,
124 "title": [
125 {
126 "Text": "TODOs are not a keyword"
127 }
128 ],
129 "tags": [],
130 "properties": {
131 "entries": []
132 },
133 "id": null,
134 "custom_id": null
135 },
136 "content": [
137 {
138 "Paragraph": [
139 {
140 "Text": "The word boundary matters: this heading has no TODO keyword."
141 }
142 ]
143 }
144 ],
145 "children": []
146 },
147 {
148 "heading": {
149 "level": 1,
150 "todo": {
151 "name": "DONE",
152 "done": true
153 },
154 "priority": null,
155 "title": [],
156 "tags": [],
157 "properties": {
158 "entries": []
159 },
160 "id": null,
161 "custom_id": null
162 },
163 "content": [
164 {
165 "Paragraph": [
166 {
167 "Text": "A keyword with no title at all."
168 }
169 ]
170 }
171 ],
172 "children": []
173 }
174 ]
175}
tests/snapshots/constructs__headings_html.snap added +14
@@ -0,0 +1,14 @@
1---
2source: tests/constructs.rs
3expression: "render_fixture(\"headings.org\")"
4---
5<h1 id="write-parser"><span class="todo TODO">TODO</span> <span class="priority">[#A]</span> Write the parser <span class="tag">work</span> <span class="tag">rust</span></h1>
6<p>A heading carrying a keyword, a priority, tags and a property drawer.</p>
7<h2 id="nested-and-finished"><span class="done DONE">DONE</span> Nested and finished</h2>
8<p>Sub-headings nest by star count.</p>
9<h2 id="priority-without-a-keyword"><span class="priority">[#C]</span> Priority without a keyword</h2>
10<p>A priority cookie can stand alone.</p>
11<h1 id="todos-are-not-a-keyword">TODOs are not a keyword</h1>
12<p>The word boundary matters: this heading has no TODO keyword.</p>
13<h1><span class="done DONE">DONE</span> </h1>
14<p>A keyword with no title at all.</p>
tests/snapshots/constructs__images_html.snap added +14
@@ -0,0 +1,14 @@
1---
2source: tests/constructs.rs
3expression: "render_fixture(\"images.org\")"
4---
5<h1 id="bare-image">Bare image</h1>
6<p><img src="diagram.png" alt=""></p>
7<h1 id="captioned-figure">Captioned figure</h1>
8<figure><img src="pipeline.svg" alt="The pipeline, end to end" width="640" class="diagram"><figcaption>The pipeline, end to end</figcaption></figure>
9<h1 id="caption-with-markup">Caption with markup</h1>
10<figure><img src="chart.png" alt="A stylised chart"><figcaption>A <em>stylised</em> chart</figcaption></figure>
11<h1 id="quoted-attribute-values">Quoted attribute values</h1>
12<figure><img src="cat.jpg" alt="a cat, sitting" loading="lazy"></figure>
13<h1 id="image-with-a-description-is-a-link">Image with a description is a link</h1>
14<p><a href="diagram.png">the diagram</a></p>
tests/snapshots/constructs__lists_element_tree.snap added +466
@@ -0,0 +1,466 @@
1---
2source: tests/constructs.rs
3expression: "parse_fixture(\"lists.org\").root"
4---
5{
6 "heading": null,
7 "content": [
8 {
9 "Keyword": {
10 "key": "TITLE",
11 "value": "Lists"
12 }
13 }
14 ],
15 "children": [
16 {
17 "heading": {
18 "level": 1,
19 "todo": null,
20 "priority": null,
21 "title": [
22 {
23 "Text": "Nesting"
24 }
25 ],
26 "tags": [],
27 "properties": {
28 "entries": []
29 },
30 "id": null,
31 "custom_id": null
32 },
33 "content": [
34 {
35 "List": {
36 "kind": "Unordered",
37 "items": [
38 {
39 "bullet": "Dash",
40 "checkbox": null,
41 "term": null,
42 "content": [
43 {
44 "Paragraph": [
45 {
46 "Text": "outer item"
47 }
48 ]
49 },
50 {
51 "List": {
52 "kind": "Unordered",
53 "items": [
54 {
55 "bullet": "Dash",
56 "checkbox": null,
57 "term": null,
58 "content": [
59 {
60 "Paragraph": [
61 {
62 "Text": "inner item"
63 }
64 ]
65 },
66 {
67 "List": {
68 "kind": "Unordered",
69 "items": [
70 {
71 "bullet": "Dash",
72 "checkbox": null,
73 "term": null,
74 "content": [
75 {
76 "Paragraph": [
77 {
78 "Text": "deepest item"
79 }
80 ]
81 }
82 ]
83 }
84 ]
85 }
86 }
87 ]
88 },
89 {
90 "bullet": "Dash",
91 "checkbox": null,
92 "term": null,
93 "content": [
94 {
95 "Paragraph": [
96 {
97 "Text": "second inner"
98 }
99 ]
100 }
101 ]
102 }
103 ]
104 }
105 }
106 ]
107 },
108 {
109 "bullet": "Dash",
110 "checkbox": null,
111 "term": null,
112 "content": [
113 {
114 "Paragraph": [
115 {
116 "Text": "second outer"
117 }
118 ]
119 }
120 ]
121 }
122 ]
123 }
124 }
125 ],
126 "children": []
127 },
128 {
129 "heading": {
130 "level": 1,
131 "todo": null,
132 "priority": null,
133 "title": [
134 {
135 "Text": "Ordered"
136 }
137 ],
138 "tags": [],
139 "properties": {
140 "entries": []
141 },
142 "id": null,
143 "custom_id": null
144 },
145 "content": [
146 {
147 "List": {
148 "kind": "Ordered",
149 "items": [
150 {
151 "bullet": {
152 "Ordered": 1
153 },
154 "checkbox": null,
155 "term": null,
156 "content": [
157 {
158 "Paragraph": [
159 {
160 "Text": "first"
161 }
162 ]
163 }
164 ]
165 },
166 {
167 "bullet": {
168 "Ordered": 2
169 },
170 "checkbox": null,
171 "term": null,
172 "content": [
173 {
174 "Paragraph": [
175 {
176 "Text": "second"
177 }
178 ]
179 },
180 {
181 "List": {
182 "kind": "Ordered",
183 "items": [
184 {
185 "bullet": {
186 "Ordered": 1
187 },
188 "checkbox": null,
189 "term": null,
190 "content": [
191 {
192 "Paragraph": [
193 {
194 "Text": "second point one"
195 }
196 ]
197 }
198 ]
199 },
200 {
201 "bullet": {
202 "Ordered": 2
203 },
204 "checkbox": null,
205 "term": null,
206 "content": [
207 {
208 "Paragraph": [
209 {
210 "Text": "second point two"
211 }
212 ]
213 }
214 ]
215 }
216 ]
217 }
218 }
219 ]
220 },
221 {
222 "bullet": {
223 "Ordered": 3
224 },
225 "checkbox": null,
226 "term": null,
227 "content": [
228 {
229 "Paragraph": [
230 {
231 "Text": "third"
232 }
233 ]
234 }
235 ]
236 }
237 ]
238 }
239 }
240 ],
241 "children": []
242 },
243 {
244 "heading": {
245 "level": 1,
246 "todo": null,
247 "priority": null,
248 "title": [
249 {
250 "Text": "Checkboxes"
251 }
252 ],
253 "tags": [],
254 "properties": {
255 "entries": []
256 },
257 "id": null,
258 "custom_id": null
259 },
260 "content": [
261 {
262 "List": {
263 "kind": "Unordered",
264 "items": [
265 {
266 "bullet": "Dash",
267 "checkbox": "Off",
268 "term": null,
269 "content": [
270 {
271 "Paragraph": [
272 {
273 "Text": "not done"
274 }
275 ]
276 }
277 ]
278 },
279 {
280 "bullet": "Dash",
281 "checkbox": "On",
282 "term": null,
283 "content": [
284 {
285 "Paragraph": [
286 {
287 "Text": "done"
288 }
289 ]
290 }
291 ]
292 },
293 {
294 "bullet": "Dash",
295 "checkbox": "Trans",
296 "term": null,
297 "content": [
298 {
299 "Paragraph": [
300 {
301 "Text": "partially done"
302 }
303 ]
304 }
305 ]
306 }
307 ]
308 }
309 }
310 ],
311 "children": []
312 },
313 {
314 "heading": {
315 "level": 1,
316 "todo": null,
317 "priority": null,
318 "title": [
319 {
320 "Text": "Description"
321 }
322 ],
323 "tags": [],
324 "properties": {
325 "entries": []
326 },
327 "id": null,
328 "custom_id": null
329 },
330 "content": [
331 {
332 "List": {
333 "kind": "Description",
334 "items": [
335 {
336 "bullet": "Dash",
337 "checkbox": null,
338 "term": [
339 {
340 "Text": "term one"
341 }
342 ],
343 "content": [
344 {
345 "Paragraph": [
346 {
347 "Text": "the first definition"
348 }
349 ]
350 }
351 ]
352 },
353 {
354 "bullet": "Dash",
355 "checkbox": null,
356 "term": [
357 {
358 "Text": "term two"
359 }
360 ],
361 "content": [
362 {
363 "Paragraph": [
364 {
365 "Text": "the second definition, which is soft-wrapped across two lines"
366 }
367 ]
368 }
369 ]
370 },
371 {
372 "bullet": "Dash",
373 "checkbox": null,
374 "term": [
375 {
376 "Italic": [
377 {
378 "Text": "marked up"
379 }
380 ]
381 },
382 {
383 "Text": " term"
384 }
385 ],
386 "content": [
387 {
388 "Paragraph": [
389 {
390 "Text": "definitions hold inline markup"
391 }
392 ]
393 }
394 ]
395 }
396 ]
397 }
398 }
399 ],
400 "children": []
401 },
402 {
403 "heading": {
404 "level": 1,
405 "todo": null,
406 "priority": null,
407 "title": [
408 {
409 "Text": "Multi-paragraph items"
410 }
411 ],
412 "tags": [],
413 "properties": {
414 "entries": []
415 },
416 "id": null,
417 "custom_id": null
418 },
419 "content": [
420 {
421 "List": {
422 "kind": "Unordered",
423 "items": [
424 {
425 "bullet": "Dash",
426 "checkbox": null,
427 "term": null,
428 "content": [
429 {
430 "Paragraph": [
431 {
432 "Text": "an item whose body has two paragraphs"
433 }
434 ]
435 },
436 {
437 "Paragraph": [
438 {
439 "Text": "the second paragraph, indented under the bullet"
440 }
441 ]
442 }
443 ]
444 },
445 {
446 "bullet": "Dash",
447 "checkbox": null,
448 "term": null,
449 "content": [
450 {
451 "Paragraph": [
452 {
453 "Text": "a plain sibling"
454 }
455 ]
456 }
457 ]
458 }
459 ]
460 }
461 }
462 ],
463 "children": []
464 }
465 ]
466}
tests/snapshots/constructs__lists_html.snap added +48
@@ -0,0 +1,48 @@
1---
2source: tests/constructs.rs
3expression: "render_fixture(\"lists.org\")"
4---
5<h1 id="nesting">Nesting</h1>
6<ul>
7<li>outer item<ul>
8<li>inner item<ul>
9<li>deepest item</li>
10</ul>
11</li>
12<li>second inner</li>
13</ul>
14</li>
15<li>second outer</li>
16</ul>
17<h1 id="ordered">Ordered</h1>
18<ol>
19<li>first</li>
20<li>second<ol>
21<li>second point one</li>
22<li>second point two</li>
23</ol>
24</li>
25<li>third</li>
26</ol>
27<h1 id="checkboxes">Checkboxes</h1>
28<ul>
29<li><input type="checkbox" disabled> not done</li>
30<li><input type="checkbox" disabled checked> done</li>
31<li><input type="checkbox" disabled> partially done</li>
32</ul>
33<h1 id="description">Description</h1>
34<dl>
35<dt>term one</dt>
36<dd>the first definition</dd>
37<dt>term two</dt>
38<dd>the second definition, which is soft-wrapped across two lines</dd>
39<dt><em>marked up</em> term</dt>
40<dd>definitions hold inline markup</dd>
41</dl>
42<h1 id="multi-paragraph-items">Multi-paragraph items</h1>
43<ul>
44<li><p>an item whose body has two paragraphs</p>
45<p>the second paragraph, indented under the bullet</p>
46</li>
47<li>a plain sibling</li>
48</ul>
tests/snapshots/constructs__out_of_scope_html.snap added +28
@@ -0,0 +1,28 @@
1---
2source: tests/constructs.rs
3expression: "render_fixture(\"outofscope.org\")"
4---
5<p>Every construct here is on the README's explicit OUT list. The contract is not that we handle them — it is that they degrade predictably and never crash the build.</p>
6<h1 id="babel">Babel</h1>
7<pre><code class="language-sh highlight"><span class="source shell bash"><span class="meta function-call shell"><span class="support function echo shell">echo</span></span><span class="meta function-call arguments shell"> <span class="string quoted double shell"><span class="punctuation definition string begin shell">&quot;</span>the block renders; :results is never executed<span class="punctuation definition string end shell">&quot;</span></span></span></span></code></pre>
8<h1 id="table-formulas">Table formulas</h1>
9<table>
10<thead>
11<tr><th>item</th><th>cost</th></tr>
12</thead>
13<tbody>
14<tr><td>a</td><td>1</td></tr>
15<tr><td>b</td><td>2</td></tr>
16</tbody>
17</table>
18<h1 id="latex">LaTeX</h1>
19<p>Inline math $x^2 + y^2$ and a display block:</p>
20<p>\begin{equation} E = mc^2 \end{equation}</p>
21<h1 id="macros-and-radio-targets">Macros and radio targets</h1>
22<p>A macro call {{{author}}} and a &lt;&lt;&lt;radio target&gt;&gt;&gt; stay literal.</p>
23<h1 id="drawers">Drawers</h1>
24<h1 id="verse">Verse</h1>
25<pre>An unmodelled block type
26keeps its content verbatim.</pre>
27<h1 id="entities">Entities</h1>
28<p>The full entity set is out of scope, so \alpha stays literal.</p>
tests/snapshots/constructs__timestamps_element_tree.snap added +209
@@ -0,0 +1,209 @@
1---
2source: tests/constructs.rs
3expression: "parse_fixture(\"timestamps.org\").root"
4---
5{
6 "heading": null,
7 "content": [
8 {
9 "Keyword": {
10 "key": "TITLE",
11 "value": "Timestamps"
12 }
13 }
14 ],
15 "children": [
16 {
17 "heading": {
18 "level": 1,
19 "todo": null,
20 "priority": null,
21 "title": [
22 {
23 "Text": "Single"
24 }
25 ],
26 "tags": [],
27 "properties": {
28 "entries": []
29 },
30 "id": null,
31 "custom_id": null
32 },
33 "content": [
34 {
35 "Paragraph": [
36 {
37 "Text": "An active date "
38 },
39 {
40 "Timestamp": {
41 "active": true,
42 "start": "2024-01-15T00:00:00",
43 "end": null,
44 "has_time": false
45 }
46 },
47 {
48 "Text": " and an inactive one "
49 },
50 {
51 "Timestamp": {
52 "active": false,
53 "start": "2024-01-15T00:00:00",
54 "end": null,
55 "has_time": false
56 }
57 },
58 {
59 "Text": "."
60 }
61 ]
62 },
63 {
64 "Paragraph": [
65 {
66 "Text": "With a time: "
67 },
68 {
69 "Timestamp": {
70 "active": true,
71 "start": "2024-01-15T10:30:00",
72 "end": null,
73 "has_time": true
74 }
75 },
76 {
77 "Text": "."
78 }
79 ]
80 }
81 ],
82 "children": []
83 },
84 {
85 "heading": {
86 "level": 1,
87 "todo": null,
88 "priority": null,
89 "title": [
90 {
91 "Text": "Ranges"
92 }
93 ],
94 "tags": [],
95 "properties": {
96 "entries": []
97 },
98 "id": null,
99 "custom_id": null
100 },
101 "content": [
102 {
103 "Paragraph": [
104 {
105 "Text": "A same-day time range "
106 },
107 {
108 "Timestamp": {
109 "active": true,
110 "start": "2024-01-15T10:00:00",
111 "end": "2024-01-15T11:45:00",
112 "has_time": true
113 }
114 },
115 {
116 "Text": "."
117 }
118 ]
119 },
120 {
121 "Paragraph": [
122 {
123 "Text": "A multi-day range "
124 },
125 {
126 "Timestamp": {
127 "active": true,
128 "start": "2024-01-15T00:00:00",
129 "end": "2024-01-20T00:00:00",
130 "has_time": false
131 }
132 },
133 {
134 "Text": "."
135 }
136 ]
137 }
138 ],
139 "children": []
140 },
141 {
142 "heading": {
143 "level": 1,
144 "todo": null,
145 "priority": null,
146 "title": [
147 {
148 "Text": "Ignored decorations"
149 }
150 ],
151 "tags": [],
152 "properties": {
153 "entries": []
154 },
155 "id": null,
156 "custom_id": null
157 },
158 "content": [
159 {
160 "Paragraph": [
161 {
162 "Text": "A repeater is dropped: "
163 },
164 {
165 "Timestamp": {
166 "active": true,
167 "start": "2024-01-15T00:00:00",
168 "end": null,
169 "has_time": false
170 }
171 },
172 {
173 "Text": "."
174 }
175 ]
176 }
177 ],
178 "children": []
179 },
180 {
181 "heading": {
182 "level": 1,
183 "todo": null,
184 "priority": null,
185 "title": [
186 {
187 "Text": "Not timestamps"
188 }
189 ],
190 "tags": [],
191 "properties": {
192 "entries": []
193 },
194 "id": null,
195 "custom_id": null
196 },
197 "content": [
198 {
199 "Paragraph": [
200 {
201 "Text": "Comparisons like 3 < 4 and [not a stamp] stay literal text."
202 }
203 ]
204 }
205 ],
206 "children": []
207 }
208 ]
209}
tests/snapshots/constructs__timestamps_html.snap added +14
@@ -0,0 +1,14 @@
1---
2source: tests/constructs.rs
3expression: "render_fixture(\"timestamps.org\")"
4---
5<h1 id="single">Single</h1>
6<p>An active date <time class="timestamp" datetime="2024-01-15">2024-01-15</time> and an inactive one <time class="timestamp inactive" datetime="2024-01-15">2024-01-15</time>.</p>
7<p>With a time: <time class="timestamp" datetime="2024-01-15T10:30">2024-01-15 10:30</time>.</p>
8<h1 id="ranges">Ranges</h1>
9<p>A same-day time range <time class="timestamp" datetime="2024-01-15T10:00">2024-01-15 10:00</time>&#8211;<time class="timestamp" datetime="2024-01-15T11:45">11:45</time>.</p>
10<p>A multi-day range <time class="timestamp" datetime="2024-01-15">2024-01-15</time>&#8211;<time class="timestamp" datetime="2024-01-20">2024-01-20</time>.</p>
11<h1 id="ignored-decorations">Ignored decorations</h1>
12<p>A repeater is dropped: <time class="timestamp" datetime="2024-01-15">2024-01-15</time>.</p>
13<h1 id="not-timestamps">Not timestamps</h1>
14<p>Comparisons like 3 &lt; 4 and [not a stamp] stay literal text.</p>
tests/snapshots/pipeline__core_element_tree.snap +6
@@ -5,6 +5,12 @@ expression: "parse_fixture(\"core.org\").root"
55{
66 "heading": null,
77 "content": [
8 {
9 "Keyword": {
10 "key": "TITLE",
11 "value": "Core Constructs"
12 }
13 },
814 {
915 "Paragraph": [
1016 {
tests/snapshots/pipeline__core_html.snap +3 −3
@@ -14,6 +14,6 @@ expression: "render_fixture(\"core.org\")"
1414</ul>
1515<h1 id="links-and-code">Links and code</h1>
1616<p>An external <a href="https://example.org">site</a> and a bare <a href="https://bare.example">https://bare.example</a>.</p>
17<pre><code class="language-rust">fn main() {
18 println!("hello");
19}</code></pre>
17<pre><code class="language-rust highlight"><span class="source rust"><span class="meta function rust"><span class="meta function rust"><span class="storage type function rust">fn</span> </span><span class="entity name function rust">main</span></span><span class="meta function rust"><span class="meta function parameters rust"><span class="punctuation section parameters begin rust">(</span></span><span class="meta function rust"><span class="meta function parameters rust"><span class="punctuation section parameters end rust">)</span></span></span></span><span class="meta function rust"> </span><span class="meta function rust"><span class="meta block rust"><span class="punctuation section block begin rust">{</span>
18 <span class="support macro rust">println!</span><span class="meta group rust"><span class="punctuation section group begin rust">(</span></span><span class="meta group rust"><span class="string quoted double rust"><span class="punctuation definition string begin rust">&quot;</span>hello<span class="punctuation definition string end rust">&quot;</span></span></span><span class="meta group rust"><span class="punctuation section group end rust">)</span></span><span class="punctuation terminator rust">;</span>
19</span><span class="meta block rust"><span class="punctuation section block end rust">}</span></span></span></span></code></pre>
tests/snapshots/pipeline__minimal_element_tree.snap +18
@@ -5,6 +5,24 @@ expression: "parse_fixture(\"minimal.org\").root"
55{
66 "heading": null,
77 "content": [
8 {
9 "Keyword": {
10 "key": "TITLE",
11 "value": "Minimal Fixture"
12 }
13 },
14 {
15 "Keyword": {
16 "key": "DATE",
17 "value": "2026-08-08"
18 }
19 },
20 {
21 "Keyword": {
22 "key": "AUTHOR",
23 "value": "Owner"
24 }
25 },
826 {
927 "Paragraph": [
1028 {
tests/snapshots/site__site_guide_html.snap +1
@@ -7,6 +7,7 @@ expression: "page(&pages, \"guide.org\").html"
77<head>
88<meta charset="utf-8">
99<title>Guide</title>
10<link rel="stylesheet" href="syntax.css">
1011</head>
1112<body>
1213<nav>
tests/snapshots/site__site_index_html.snap +1
@@ -7,6 +7,7 @@ expression: "page(&pages, \"index.org\").html"
77<head>
88<meta charset="utf-8">
99<title>Home</title>
10<link rel="stylesheet" href="syntax.css">
1011</head>
1112<body>
1213<nav>