krz/orgo

Lightning fast org-mode static site generator. fast go org-mode static-site-generator

Commit ba4152c75e

ba4152c75ebfc45fa02279a0d355613465d42a0a

parent: 19279bbf18

Verified · cmc

cmc <hello@cleberg.net> · 2026-08-11 04:20 UTC

Phase 0: corpus audit and emacs --batch oracle, and honor #+SLUG:

Replaces the two guesses the v1 scope rested on with measurements. The corpus is the
179 files behind cleberg.net, published today by weblorg — a wrapper around org's own
HTML exporter, so it is both the workload and the incumbent.

Corpus audit (src/audit.rs, `org-ssg audit <dir>`):
- Reports construct frequencies classified against the IN/OUT line, plus a census of
  every keyword, block type, drawer and link scheme seen, so unrecognized names
  self-report instead of hiding. Deliberately a separate line scanner rather than a
  reuse of the parser: auditing with the parser could only find constructs the parser
  already knows, which is the wrong instrument for finding blind spots.
- Reports names, counts and file:line only, never document text, so auditing private
  notes stays publishable.

What it found:
- The scope guess was sound: 99.9% of construct uses are in scope. The entire
  out-of-scope tail is 8 uses.
- #+SLUG: was missing entirely, and it decides the published URL: 178 of 179 files set
  one, and 2018-11-28-aes-encryption.org is served at blog/aes-encryption.html. Output
  paths came from source filenames, so 169 of 179 pages would have been published at
  the wrong URL by a build that reported success. Output paths now come from the slug
  (util::output_path), threaded through INDEX/RESOLVE/nav so links follow it. Slugs are
  sanitized — an author-supplied ../../etc/x cannot escape the output directory — and
  two pages claiming one URL is a build error, not a silently dropped page. Building
  the real corpus now reproduces all 179 live URLs exactly.
- INDEX/RESOLVE is speculative against this corpus: it contains no id:, #custom-id or
  *Heading links at all.
- The audit's own first run lied, reporting 23 custom TODO keyword sequences. All were
  false — it read the leading word of "* CSS Variables" as the keyword "CSS", and the
  corpus defines no #+TODO: sequences. Now matched against conventional names only.

Emacs oracle (tests/oracle.el, tests/oracle.rs):
- Exports each fixture with org's own exporter and reduces both sides to a semantic
  skeleton, dropping layout divs, inline spans and all attributes but href/src. Byte
  equality was never the goal; org wraps every section in outline-container divs keyed
  by generated ids.
- Snapshots the disagreement rather than asserting agreement: a checked-in divergence
  report gets reviewed and shows up as a diff, where a permanently red test gets
  ignored. Three invariants are asserted outright and all hold — heading structure,
  list nesting and source block text match Emacs exactly.
- Skips cleanly when emacs is absent, so it never blocks CI.

It found no bugs in org-ssg. Every divergence is a deliberate choice to emit better
HTML: <em>/<strong> over <i>/<b>, <figure>/<figcaption> over "Figure 1:", <time
datetime> over a literal timestamp, <section><ol> footnotes over an <h2>, slugged
heading anchors over org1a2b3c4, <pre><code> over bare <pre>. One real semantic
difference is kept on measurement rather than taste: org merges a 1. list and a
following - list separated by one blank line into a single list, and that pattern
occurs zero times in the corpus.

Its best catch was three bugs in itself: normalization that trimmed each of syntect's
per-token text runs reported code as corrupted (def greet -> defgreet), and keeping
syntect's spans put blocks.org at 36% agreement. Both were measurement artifacts.

Layout: unified · split

README.md +82 −8
@@ -25,6 +25,7 @@ is the only inherently global stage — it is where the link dependency graph is
2525| Stage | Module | Notes |
2626|---|---|---|
2727| PARSE | `src/parser.rs` | Hand-written recursive descent: line lexer → element builder → inline tokenizer. |
28| audit | `src/audit.rs` | Phase 0 corpus audit: construct frequencies against the IN/OUT line. |
2829| model | `src/model.rs` | The org element tree — Elements (block) vs Objects (inline). |
2930| INDEX | `src/index.rs` | Collect link targets into a symbol table. |
3031| RESOLVE | `src/resolve.rs` | Rewrite links to URLs; return the used-target list (dependency edges). |
@@ -50,8 +51,9 @@ semantics; non-HTML export blocks; the full Unicode entity set.
5051**Scope guardrail:** every IN item gets a golden-file fixture; every OUT item gets a test
5152asserting it degrades predictably (ignored, no crash). The IN/OUT line is enforced by
5253`tests/constructs.rs`, defending against the project's #1 risk: scope creep back toward
53all-of-org. The fixtures are hand-written today; deriving them from a real corpus is
54Phase 0.
54all-of-org. Phase 0 checked this line against a real 179-file corpus and found it sound
55(99.9% of construct uses in scope) — but also found one thing missing from it entirely:
56`#+SLUG:`. See [Phase 0](#phase-0-the-corpus-audit-and-the-emacs-oracle).
5557
5658## Phase plan
5759
@@ -62,7 +64,7 @@ Phase 0.
6264| **v0.2** | **Multi-file SITE build: INDEX + RESOLVE internal links, minijinja templates, `build <src-dir> <out-dir>`, tables + footnotes** | **done** |
6365| **v0.3** | **Incremental build layer: content/config/template hashing, dependency graph, per-page render keys, persisted cache manifest, invalidation** | **done** |
6466| **v0.4** | **MVP: the full v1 construct scope — heading metadata, nested/description lists, block types, timestamps, images, syntect highlighting — with the IN/OUT line under test** | **done** |
65| 0 | Corpus audit + `emacs --batch` ground-truth oracle | todo |
67| **0** | **Corpus audit + `emacs --batch` ground-truth oracle** | **done** |
6668| 1 | Line lexer + heading/section skeleton | done |
6769| 2 | Block elements — lists, source blocks, tables, footnote defs, blocks by type, drawers | done |
6870| 3 | Inline objects — emphasis, links, bare URLs, footnote refs, timestamps | done |
@@ -162,11 +164,82 @@ excluded construct to a specific degradation: babel is never executed *and* a ch
162164as literal text; drawers other than PROPERTIES are captured and dropped; unmodelled block
163165types keep their content verbatim.
164166
165**Still out at v0.4:** the Phase 0 corpus audit and `emacs --batch` oracle (the fixtures are
166hand-written, so "matches Emacs" is asserted by construction, not measured); rayon
167parallelism; parse errors carrying source locations; `#+TODO:` per-file keyword sequences;
168planning lines (`SCHEDULED:`/`DEADLINE:`), which render as ordinary paragraphs; fixed-width
169`: ` lines; and the `watch` fs-notify integration.
167**Still out at v0.4:** rayon parallelism; parse errors carrying source locations; `#+TODO:`
168per-file keyword sequences; planning lines (`SCHEDULED:`/`DEADLINE:`), which render as
169ordinary paragraphs; fixed-width `: ` lines; and the `watch` fs-notify integration.
170
171## Phase 0: the corpus audit and the Emacs oracle
172
173The v1 scope was, by its own admission, *recommended* — a guess about which slice of org
174matters. Phase 0 replaces both halves of that guess with a measurement: an audit that asks
175what a real corpus actually uses, and an oracle that asks whether we render it the way
176Emacs does. The corpus is the 179 files behind [cleberg.net](https://cleberg.net), which is
177published today by weblorg — a wrapper around org's own HTML exporter. That makes it both
178the workload and the incumbent.
179
180```
181cargo run -- audit <src-dir> # what does this corpus use, and is it in scope?
182cargo test --test oracle # how does our HTML differ from Emacs' own export?
183```
184
185### What the audit found
186
187**The scope guess was sound.** 99.9% of construct uses in the corpus are in scope. The
188whole out-of-scope tail is 8 uses: four `#+TBLFM:` in a post *about* org-mode, three
189`\name` entities, and one `#+BEGIN_NOTE`.
190
191**`#+SLUG:` was a hole big enough to sink the project.** 178 of 179 files set it, and the
192published URL comes from it, not from the filename: `2018-11-28-aes-encryption.org` is
193served at `blog/aes-encryption.html`. org-ssg derived output paths from source filenames,
194so **169 of 179 pages would have been published at the wrong URL** — every inbound link and
195every search result, broken, by a tool that reported a clean build. Output paths now come
196from `#+SLUG:` when present ([`util::output_path`](src/util.rs)); slugs are sanitized so an
197author-supplied `../../etc/x` cannot escape the output directory, and two pages claiming one
198URL is a build error rather than a silently dropped page. Building the real corpus now
199reproduces all 179 of the live site's URLs exactly.
200
201**Some machinery is speculative.** The corpus contains no `id:`, `#custom-id` or `*Heading`
202links at all — its cross-page links are hand-written relative URLs. The INDEX/RESOLVE
203symbol table that v0.2 was built around is, against this corpus, unexercised.
204
205**An audit can lie too.** The first run reported 23 uses of a custom TODO keyword sequence.
206All 23 were false: the detector read the leading word of `* CSS Variables` as the keyword
207`CSS`. The corpus defines no `#+TODO:` sequences at all, so the true count was zero. The
208detector now matches conventional keyword names only — a tool that overstates a gap argues
209for work nobody needs.
210
211### What the oracle found
212
213`tests/oracle.rs` exports each fixture with org's own exporter via `emacs --batch`, reduces
214both sides to a semantic skeleton (element opens, closes and text, with layout `div`s,
215inline `span`s and all attributes but `href`/`src` dropped), and **snapshots the
216disagreement**. Snapshotting rather than asserting is deliberate: a checked-in divergence
217report gets reviewed and shows up as a diff, where a permanently red test gets ignored.
218Three invariants are asserted outright, and all three hold — heading structure, list
219nesting, and source-block text match Emacs exactly.
220
221**No bugs in org-ssg.** Every remaining divergence is a deliberate choice to emit better
222HTML than org does:
223
224| | org-ssg | Emacs | why |
225|---|---|---|---|
226| emphasis | `<em>`/`<strong>` | `<i>`/`<b>` | semantic, not presentational |
227| captioned image | `<figure>`/`<figcaption>` | `<p>` + `"Figure 1: …"` | real figure semantics |
228| timestamp | `<time datetime="…">` | literal `<2024-01-15 Mon>` | machine-readable |
229| footnotes | `<section><ol>` | `<h2>Footnotes:</h2>` | a list of notes is a list |
230| heading anchor | slug of the text | `org1a2b3c4` | stable, and what the live site serves |
231| code | `<pre><code>` | `<pre>` | the HTML5 idiom |
232
233One genuine semantic difference: org treats a single blank line between a `1.` list and a
234`-` list as *one* list and keeps the first item's bullet type, while we start a second list.
235We keep ours, on measurement rather than taste — the pattern occurs **zero** times in the
236corpus, so matching an org quirk would buy nothing and cost the more obvious reading.
237
238**The oracle's best catch was three bugs in itself.** Naive normalization reported code as
239corrupted (it trimmed each of syntect's per-token text runs, turning `def greet` into
240`defgreet`) and reported blocks at 36% agreement (syntect's spans flooded the diff). Both
241were measurement artifacts. A differential harness is a piece of software like any other,
242and the first divergences it reports are usually its own.
170243
171244**From v0.1 (core subset):** headings with nesting and anchors (every heading is now
172245anchored — `:CUSTOM_ID:`/`:ID:` else a slug of its text) and trailing tags; paragraphs;
@@ -188,6 +261,7 @@ cargo build
188261cargo test
189262cargo run -- build fixtures/minimal.org -o minimal.html # single file
190263cargo run -- build fixtures/site -o _site # whole site (incremental)
264cargo run -- audit fixtures/site # corpus audit (Phase 0)
191265cargo run -- build fixtures/site -o _site --no-cache # force a full rebuild
192266cargo run -- watch fixtures/site -o _site # poll + rebuild on change
193267cargo run -- clean _site # remove output + cache
fixtures/slugsite/2024-02-11-long-source-name.org added +11
@@ -0,0 +1,11 @@
1#+TITLE: The Post
2#+SLUG: short-url
3
4The source filename carries a date; the published URL does not.
5
6* Setup
7:PROPERTIES:
8:CUSTOM_ID: setup
9:END:
10
11Linking into this heading must land on the slugged page, not the source name.
fixtures/slugsite/index.org added +3
@@ -0,0 +1,3 @@
1#+TITLE: Home
2
3Read [[file:2024-02-11-long-source-name.org][the post]], or jump to its [[#setup][setup section]].
src/audit.rs added +620
@@ -0,0 +1,620 @@
1//! Corpus audit (spec §5, Phase 0): measure which org constructs a real corpus actually
2//! uses, and classify each against the v1 IN/OUT line.
3//!
4//! This exists because the v1 scope was, on the README's own admission, *recommended*
5//! rather than measured — a guess about which slice of org matters. A guess about a
6//! corpus is a hypothesis, and this is the experiment. It answers two questions:
7//!
8//! 1. **Coverage** — of the constructs this corpus uses, which do we handle? A construct
9//! that is common here and out of scope is a scope bug, not a corpus quirk.
10//! 2. **Blind spots** — which constructs are here that the implementation has no opinion
11//! about at all? These are the dangerous ones: not "known unsupported" but unknown.
12//!
13//! The audit is deliberately a *separate, line-oriented scanner* rather than a reuse of
14//! [`crate::parser`]. Auditing with the parser could only ever find constructs the parser
15//! already knows about, which is precisely the wrong instrument for question 2 — it would
16//! report a blind spot as clean.
17//!
18//! Nothing here reports document *text*. Counts, construct names, and `file:line`
19//! locations only, so an audit of private notes stays publishable.
20
21use std::collections::BTreeMap;
22
23use anyhow::{Context, Result};
24use camino::{Utf8Path, Utf8PathBuf};
25use walkdir::WalkDir;
26
27/// Where a construct sits relative to the v1 scope line (README §"v1 scope").
28#[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord)]
29pub enum Scope {
30 /// v1 handles this.
31 In,
32 /// v1 deliberately excludes this; it degrades predictably.
33 Out,
34}
35
36impl Scope {
37 fn label(self) -> &'static str {
38 match self {
39 Scope::In => "IN ",
40 Scope::Out => "OUT",
41 }
42 }
43}
44
45/// One construct's tally across the corpus.
46#[derive(Debug, Default, Clone)]
47pub struct Tally {
48 pub occurrences: usize,
49 pub files: usize,
50 /// First `file:line` the construct was seen at, to make a finding actionable.
51 pub first_seen: Option<String>,
52 /// Set while scanning one file, to count each file once.
53 seen_in_current_file: bool,
54}
55
56/// The audit result: the fixed construct catalog plus the dynamic name censuses.
57#[derive(Debug, Default)]
58pub struct Audit {
59 pub files: usize,
60 pub lines: usize,
61 /// Catalogued constructs → tally.
62 pub constructs: BTreeMap<(Scope, &'static str), Tally>,
63 /// Every distinct `#+KEYWORD:` seen, by name.
64 pub keywords: BTreeMap<String, Tally>,
65 /// Every distinct `#+BEGIN_<TYPE>` seen, by type.
66 pub blocks: BTreeMap<String, Tally>,
67 /// Every distinct `:DRAWER:` seen, by name.
68 pub drawers: BTreeMap<String, Tally>,
69 /// Every distinct link scheme seen (`https`, `file`, `id`, `denote`, ...).
70 pub link_schemes: BTreeMap<String, Tally>,
71}
72
73/// Names the implementation understands, so the census can flag everything else. These
74/// are the *recognized* sets, not the supported ones: `INCLUDE` is recognized (it is
75/// deliberately inert) while an unlisted keyword is a genuine blind spot.
76const KNOWN_KEYWORDS: &[&str] = &[
77 "TITLE", "AUTHOR", "DATE", "EMAIL", "LANGUAGE", "OPTIONS", "FILETAGS", "DESCRIPTION",
78 "KEYWORDS", "CAPTION", "NAME", "ATTR_HTML", "RESULTS", "TBLFM", "INCLUDE", "TODO",
79 "STARTUP", "SUBTITLE", "SETUPFILE", "MACRO", "PROPERTY", "HTML_HEAD", "EXCLUDE_TAGS",
80];
81const KNOWN_BLOCKS: &[&str] = &["SRC", "QUOTE", "EXAMPLE", "CENTER", "EXPORT"];
82const KNOWN_DRAWERS: &[&str] = &["PROPERTIES", "LOGBOOK", "END"];
83/// Keyword names conventional enough to be worth flagging when they lead a heading.
84/// A custom sequence is only *real* if some `#+TODO:` declares it, which the census
85/// reports separately — this list keeps the heading-level signal honest.
86const CONVENTIONAL_TODO_KEYWORDS: &[&str] = &[
87 "NEXT", "WAITING", "HOLD", "CANCELLED", "CANCELED", "STARTED", "SOMEDAY", "PROJ",
88 "IN-PROGRESS", "BLOCKED", "REVIEW",
89];
90const KNOWN_SCHEMES: &[&str] = &[
91 "http", "https", "mailto", "ftp", "news", "tel", "file", "id", "custom-id", "heading",
92 "relative",
93];
94
95impl Audit {
96 /// Is this name one the implementation recognizes?
97 pub fn is_known(kind: Census, name: &str) -> bool {
98 let known = match kind {
99 Census::Keyword => KNOWN_KEYWORDS,
100 Census::Block => KNOWN_BLOCKS,
101 Census::Drawer => KNOWN_DRAWERS,
102 Census::Scheme => KNOWN_SCHEMES,
103 };
104 known.iter().any(|k| k.eq_ignore_ascii_case(name))
105 }
106}
107
108/// Which dynamic census a name belongs to.
109#[derive(Debug, Clone, Copy)]
110pub enum Census {
111 Keyword,
112 Block,
113 Drawer,
114 Scheme,
115}
116
117/// Walk `root`, auditing every `.org` file.
118pub fn audit(root: &Utf8Path) -> Result<Audit> {
119 let mut audit = Audit::default();
120 let mut paths: Vec<Utf8PathBuf> = Vec::new();
121
122 if root.is_file() {
123 paths.push(root.to_owned());
124 } else {
125 for entry in WalkDir::new(root).sort_by_file_name() {
126 let entry = entry.with_context(|| format!("walking {root}"))?;
127 if !entry.file_type().is_file() {
128 continue;
129 }
130 let path = Utf8PathBuf::from_path_buf(entry.into_path())
131 .map_err(|p| anyhow::anyhow!("non-UTF-8 path: {}", p.display()))?;
132 if path.extension() == Some("org") {
133 paths.push(path);
134 }
135 }
136 }
137
138 for path in &paths {
139 // A file that cannot be read is reported and skipped: an audit of 179 files
140 // should not be lost to one unreadable one.
141 let source = match std::fs::read_to_string(path) {
142 Ok(s) => s,
143 Err(e) => {
144 eprintln!("warning: skipping {path}: {e}");
145 continue;
146 }
147 };
148 let rel = path.strip_prefix(root).unwrap_or(path).to_owned();
149 audit.scan_file(&rel, &source);
150 audit.files += 1;
151 }
152 Ok(audit)
153}
154
155impl Audit {
156 fn scan_file(&mut self, path: &Utf8Path, source: &str) {
157 // Reset the per-file flags so each construct counts this file at most once.
158 for tally in self.constructs.values_mut() {
159 tally.seen_in_current_file = false;
160 }
161 for map in [
162 &mut self.keywords,
163 &mut self.blocks,
164 &mut self.drawers,
165 &mut self.link_schemes,
166 ] {
167 for tally in map.values_mut() {
168 tally.seen_in_current_file = false;
169 }
170 }
171
172 let mut in_block: Option<String> = None;
173 for (idx, line) in source.lines().enumerate() {
174 self.lines += 1;
175 let at = format!("{path}:{}", idx + 1);
176 let trimmed = line.trim_start();
177
178 // Inside a verbatim block only the terminator matters — a `*` in a source
179 // block is not a heading, and counting it as one would corrupt the audit.
180 if let Some(kind) = &in_block {
181 if trimmed.to_ascii_uppercase().starts_with("#+END_") {
182 in_block = None;
183 } else if kind.eq_ignore_ascii_case("SRC") || kind.eq_ignore_ascii_case("EXAMPLE") {
184 continue;
185 }
186 continue;
187 }
188 if let Some(rest) = trimmed.to_ascii_uppercase().strip_prefix("#+BEGIN_") {
189 let kind = rest.split_whitespace().next().unwrap_or("").to_string();
190 self.count_census(Census::Block, &kind, &at);
191 self.count(scope_of_block(&kind), block_construct(&kind), &at);
192 if trimmed.to_ascii_uppercase().contains(":RESULTS") {
193 self.count(Scope::Out, "babel header args (:results)", &at);
194 }
195 in_block = Some(kind);
196 continue;
197 }
198
199 self.scan_line(line, trimmed, &at);
200 }
201 }
202
203 fn scan_line(&mut self, line: &str, trimmed: &str, at: &str) {
204 // --- headings and their metadata ---
205 if let Some(stars) = heading_stars(line) {
206 self.count(Scope::In, "heading", at);
207 let rest = line[stars..].trim();
208 let word = rest.split_whitespace().next().unwrap_or("");
209 if word == "TODO" || word == "DONE" {
210 self.count(Scope::In, "TODO keyword (default set)", at);
211 } else if CONVENTIONAL_TODO_KEYWORDS.contains(&word) {
212 // Only conventional keyword names count. "Any all-caps first word" is
213 // the tempting rule and it is wrong: it reads `* CSS Variables` as the
214 // keyword `CSS`, which on this corpus produced 23 false positives and
215 // zero true ones. An audit that overstates a gap is worse than no audit,
216 // because it argues for work nobody needs.
217 self.count(Scope::Out, "TODO keyword (custom sequence)", at);
218 }
219 if rest.contains("[#") {
220 self.count(Scope::In, "priority cookie", at);
221 }
222 if rest.trim_end().ends_with(':') && rest.trim_end().matches(':').count() >= 2 {
223 self.count(Scope::In, "heading tags", at);
224 }
225 if rest.contains("[/") || rest.contains("[%") {
226 self.count(Scope::Out, "statistics cookie", at);
227 }
228 return;
229 }
230
231 // --- planning and clocking ---
232 for marker in ["SCHEDULED:", "DEADLINE:", "CLOSED:"] {
233 if trimmed.starts_with(marker) {
234 self.count(Scope::Out, "planning line", at);
235 }
236 }
237 if trimmed.starts_with("CLOCK:") {
238 self.count(Scope::Out, "clock entry", at);
239 }
240
241 // --- keywords and drawers ---
242 if let Some(rest) = trimmed.strip_prefix("#+") {
243 if let Some(colon) = rest.find(':') {
244 let key = rest[..colon].trim().to_ascii_uppercase();
245 if !key.is_empty() && !key.contains(char::is_whitespace) {
246 self.count_census(Census::Keyword, &key, at);
247 match key.as_str() {
248 "CAPTION" | "NAME" | "ATTR_HTML" => {
249 self.count(Scope::In, "affiliated keyword", at)
250 }
251 "TBLFM" => self.count(Scope::Out, "table formula (#+TBLFM:)", at),
252 "INCLUDE" => self.count(Scope::Out, "#+INCLUDE:", at),
253 "RESULTS" => self.count(Scope::Out, "babel results block", at),
254 "TODO" => self.count(Scope::Out, "#+TODO: keyword sequence", at),
255 "MACRO" => self.count(Scope::Out, "macro definition", at),
256 _ => self.count(Scope::In, "#+ keyword", at),
257 }
258 }
259 }
260 } else if is_drawer(trimmed) {
261 let name = trimmed[1..trimmed.len() - 1].to_ascii_uppercase();
262 if name != "END" {
263 self.count_census(Census::Drawer, &name, at);
264 match name.as_str() {
265 "PROPERTIES" => self.count(Scope::In, "property drawer", at),
266 _ => self.count(Scope::Out, "non-PROPERTIES drawer", at),
267 }
268 }
269 }
270
271 // --- lists, tables, rules ---
272 if let Some(bullet) = list_bullet(trimmed) {
273 self.count(Scope::In, "list item", at);
274 if bullet == Bullet::Ordered {
275 self.count(Scope::In, "ordered list", at);
276 }
277 let indent = line.len() - trimmed.len();
278 if indent > 0 {
279 self.count(Scope::In, "nested list item", at);
280 }
281 if trimmed.contains(" :: ") {
282 self.count(Scope::In, "description list", at);
283 }
284 let after = trimmed.trim_start_matches(['-', '+', '*', ' ']);
285 if after.starts_with("[ ]") || after.starts_with("[X]") || after.starts_with("[-]") {
286 self.count(Scope::In, "checkbox", at);
287 }
288 }
289 if trimmed.starts_with('|') {
290 self.count(Scope::In, "table row", at);
291 }
292 if trimmed.starts_with(':') && !is_drawer(trimmed) && trimmed.starts_with(": ") {
293 self.count(Scope::Out, "fixed-width line", at);
294 }
295
296 // --- footnotes ---
297 if trimmed.starts_with("[fn:") {
298 self.count(Scope::In, "footnote definition", at);
299 } else if line.contains("[fn:") {
300 self.count(Scope::In, "footnote reference", at);
301 }
302
303 // --- inline objects ---
304 self.scan_inline(line, at);
305 }
306
307 fn scan_inline(&mut self, line: &str, at: &str) {
308 // Links: count each `[[target]]`, censusing its scheme.
309 let mut rest = line;
310 while let Some(start) = rest.find("[[") {
311 let after = &rest[start + 2..];
312 let Some(end) = after.find("]]") else { break };
313 let inner = &after[..end];
314 let target = inner.split("][").next().unwrap_or(inner);
315 self.count(Scope::In, "link", at);
316 self.count_census(Census::Scheme, &link_scheme(target), at);
317 rest = &after[end..];
318 }
319
320 if has_timestamp(line) {
321 self.count(Scope::In, "timestamp", at);
322 }
323 if line.contains("{{{") {
324 self.count(Scope::Out, "macro call", at);
325 }
326 if line.contains("<<<") {
327 self.count(Scope::Out, "radio target", at);
328 } else if line.contains("<<") && line.contains(">>") {
329 self.count(Scope::Out, "internal target", at);
330 }
331 if line.contains("\\begin{") || latex_inline(line) {
332 self.count(Scope::Out, "LaTeX fragment", at);
333 }
334 if entity_ref(line) {
335 self.count(Scope::Out, "entity (\\name)", at);
336 }
337 for (marker, name) in [
338 ('*', "bold"),
339 ('/', "italic"),
340 ('_', "underline"),
341 ('+', "strike-through"),
342 ('=', "verbatim"),
343 ('~', "code"),
344 ] {
345 if emphasis_pair(line, marker) {
346 self.count(Scope::In, name, at);
347 }
348 }
349 }
350
351 fn count(&mut self, scope: Scope, name: &'static str, at: &str) {
352 let tally = self.constructs.entry((scope, name)).or_default();
353 bump(tally, at);
354 }
355
356 fn count_census(&mut self, kind: Census, name: &str, at: &str) {
357 let map = match kind {
358 Census::Keyword => &mut self.keywords,
359 Census::Block => &mut self.blocks,
360 Census::Drawer => &mut self.drawers,
361 Census::Scheme => &mut self.link_schemes,
362 };
363 let tally = map.entry(name.to_string()).or_default();
364 bump(tally, at);
365 }
366}
367
368fn bump(tally: &mut Tally, at: &str) {
369 tally.occurrences += 1;
370 if !tally.seen_in_current_file {
371 tally.seen_in_current_file = true;
372 tally.files += 1;
373 }
374 if tally.first_seen.is_none() {
375 tally.first_seen = Some(at.to_string());
376 }
377}
378
379// ---------------------------------------------------------------------------
380// Line-level detectors. Deliberately independent of the parser (see module docs).
381// ---------------------------------------------------------------------------
382
383fn heading_stars(line: &str) -> Option<usize> {
384 if !line.starts_with('*') {
385 return None;
386 }
387 let stars = line.chars().take_while(|c| *c == '*').count();
388 let after = &line[stars..];
389 (after.starts_with(' ') || after.is_empty()).then_some(stars)
390}
391
392#[derive(PartialEq)]
393enum Bullet {
394 Unordered,
395 Ordered,
396}
397
398fn list_bullet(trimmed: &str) -> Option<Bullet> {
399 let bytes = trimmed.as_bytes();
400 if bytes.is_empty() {
401 return None;
402 }
403 if (bytes[0] == b'-' || bytes[0] == b'+') && (bytes.len() == 1 || bytes[1] == b' ') {
404 return Some(Bullet::Unordered);
405 }
406 let digits = trimmed.chars().take_while(|c| c.is_ascii_digit()).count();
407 if digits > 0 {
408 let after = &trimmed[digits..];
409 if (after.starts_with('.') || after.starts_with(')'))
410 && (after.len() == 1 || after.as_bytes()[1] == b' ')
411 {
412 return Some(Bullet::Ordered);
413 }
414 }
415 None
416}
417
418fn is_drawer(trimmed: &str) -> bool {
419 let t = trimmed.trim_end();
420 t.len() >= 3
421 && t.starts_with(':')
422 && t.ends_with(':')
423 && t[1..t.len() - 1]
424 .chars()
425 .all(|c| c.is_ascii_alphanumeric() || c == '_' || c == '-')
426 && t.len() > 2
427}
428
429fn scope_of_block(kind: &str) -> Scope {
430 if KNOWN_BLOCKS.iter().any(|k| k.eq_ignore_ascii_case(kind)) {
431 Scope::In
432 } else {
433 Scope::Out
434 }
435}
436
437fn block_construct(kind: &str) -> &'static str {
438 match kind.to_ascii_uppercase().as_str() {
439 "SRC" => "source block",
440 "QUOTE" => "quote block",
441 "EXAMPLE" => "example block",
442 "CENTER" => "center block",
443 "EXPORT" => "export block",
444 _ => "unmodelled block type",
445 }
446}
447
448/// The scheme of a link target, normalized into the census's vocabulary.
449fn link_scheme(target: &str) -> String {
450 if let Some(rest) = target.split_once(':') {
451 let scheme = rest.0;
452 if !scheme.is_empty()
453 && scheme
454 .chars()
455 .all(|c| c.is_ascii_alphanumeric() || c == '-' || c == '+')
456 {
457 return scheme.to_ascii_lowercase();
458 }
459 }
460 if target.starts_with('#') {
461 return "custom-id".to_string();
462 }
463 if target.starts_with('*') {
464 return "heading".to_string();
465 }
466 "relative".to_string()
467}
468
469/// A `<...>`/`[...]` span opening with an ISO date is a timestamp.
470fn has_timestamp(line: &str) -> bool {
471 let bytes = line.as_bytes();
472 for (i, c) in line.char_indices() {
473 if c != '<' && c != '[' {
474 continue;
475 }
476 let rest = &bytes[i + 1..];
477 if rest.len() >= 10
478 && rest[..4].iter().all(u8::is_ascii_digit)
479 && rest[4] == b'-'
480 && rest[5..7].iter().all(u8::is_ascii_digit)
481 && rest[7] == b'-'
482 && rest[8..10].iter().all(u8::is_ascii_digit)
483 {
484 return true;
485 }
486 }
487 false
488}
489
490/// `$x$` or `\(x\)` inline math. `$` alone (a price, a shell prompt) is not math.
491fn latex_inline(line: &str) -> bool {
492 if line.contains("\\(") && line.contains("\\)") {
493 return true;
494 }
495 let dollars = line.matches('$').count();
496 dollars >= 2 && line.contains("$\\")
497}
498
499/// A `\name` entity reference such as `\alpha`, excluding LaTeX environment commands.
500fn entity_ref(line: &str) -> bool {
501 for (i, c) in line.char_indices() {
502 if c != '\\' {
503 continue;
504 }
505 let rest = &line[i + 1..];
506 let name: String = rest.chars().take_while(|c| c.is_ascii_alphabetic()).collect();
507 if name.len() >= 3 && !matches!(name.as_str(), "begin" | "end") {
508 return true;
509 }
510 }
511 false
512}
513
514/// A plausible `*bold*`-style emphasis pair: two markers on one line with non-space
515/// content between them. Approximate by design — the audit measures prevalence, and the
516/// parser owns the exact pre/post-character rules.
517fn emphasis_pair(line: &str, marker: char) -> bool {
518 let positions: Vec<usize> = line
519 .char_indices()
520 .filter(|(_, c)| *c == marker)
521 .map(|(i, _)| i)
522 .collect();
523 if positions.len() < 2 {
524 return false;
525 }
526 // A leading `*` is a heading, and `-`/`+` at line start is a bullet.
527 let trimmed = line.trim_start();
528 if trimmed.starts_with(marker) {
529 return false;
530 }
531 positions.windows(2).any(|w| w[1] > w[0] + 1)
532}
533
534// ---------------------------------------------------------------------------
535// Report
536// ---------------------------------------------------------------------------
537
538/// Render the audit as a readable report. Names, counts and locations only — never
539/// document text, so an audit of private notes is safe to paste into an issue.
540pub fn report(audit: &Audit) -> String {
541 let mut out = String::new();
542 out.push_str(&format!(
543 "corpus: {} file(s), {} line(s)\n",
544 audit.files, audit.lines
545 ));
546
547 let mut rows: Vec<(&(Scope, &str), &Tally)> = audit.constructs.iter().collect();
548 rows.sort_by(|a, b| {
549 b.1.occurrences
550 .cmp(&a.1.occurrences)
551 .then_with(|| a.0 .1.cmp(b.0 .1))
552 });
553
554 out.push_str("\nCONSTRUCTS (by frequency)\n");
555 out.push_str(&format!(
556 "{:<4} {:<32} {:>8} {:>7} {}\n",
557 "", "construct", "uses", "files", "first seen"
558 ));
559 for ((scope, name), tally) in &rows {
560 out.push_str(&format!(
561 "{:<4} {:<32} {:>8} {:>7} {}\n",
562 scope.label(),
563 name,
564 tally.occurrences,
565 tally.files,
566 tally.first_seen.as_deref().unwrap_or("")
567 ));
568 }
569
570 let in_uses: usize = rows
571 .iter()
572 .filter(|((s, _), _)| *s == Scope::In)
573 .map(|(_, t)| t.occurrences)
574 .sum();
575 let out_uses: usize = rows
576 .iter()
577 .filter(|((s, _), _)| *s == Scope::Out)
578 .map(|(_, t)| t.occurrences)
579 .sum();
580 let total = in_uses + out_uses;
581 let pct = |n: usize| {
582 if total == 0 {
583 0.0
584 } else {
585 100.0 * n as f64 / total as f64
586 }
587 };
588 out.push_str(&format!(
589 "\ncoverage: {in_uses} in-scope use(s) ({:.1}%), {out_uses} out-of-scope ({:.1}%)\n",
590 pct(in_uses),
591 pct(out_uses)
592 ));
593
594 for (title, kind, map) in [
595 ("KEYWORDS", Census::Keyword, &audit.keywords),
596 ("BLOCK TYPES", Census::Block, &audit.blocks),
597 ("DRAWERS", Census::Drawer, &audit.drawers),
598 ("LINK SCHEMES", Census::Scheme, &audit.link_schemes),
599 ] {
600 let mut names: Vec<(&String, &Tally)> = map.iter().collect();
601 names.sort_by(|a, b| b.1.occurrences.cmp(&a.1.occurrences).then(a.0.cmp(b.0)));
602 out.push_str(&format!("\n{title}\n"));
603 for (name, tally) in names {
604 let flag = if Audit::is_known(kind, name) {
605 " "
606 } else {
607 "??? "
608 };
609 out.push_str(&format!(
610 "{flag}{:<32} {:>8} {:>7} {}\n",
611 name,
612 tally.occurrences,
613 tally.files,
614 tally.first_seen.as_deref().unwrap_or("")
615 ));
616 }
617 }
618 out.push_str("\n`???` marks a name the implementation does not recognize at all.\n");
619 out
620}
src/index.rs +23 −21
@@ -7,7 +7,7 @@ use camino::{Utf8Path, Utf8PathBuf};
77use serde::{Deserialize, Serialize};
88
99use crate::model::{Document, Section};
10use crate::util::{plain_text, slugify};
10use crate::util::{output_path, plain_text, slugify};
1111
1212/// Identity of a link target. A target is owned by exactly one file (spec §4.3).
1313///
@@ -51,6 +51,9 @@ impl TargetId {
5151#[derive(Debug, Clone)]
5252pub struct TargetLocation {
5353 pub source_path: Utf8PathBuf,
54 /// The page this target is emitted into. Recorded at INDEX time because it depends
55 /// on the defining document's `#+SLUG:`, which only that document knows.
56 pub output_path: Utf8PathBuf,
5457 /// Final URL fragment/anchor for the target, filled during resolution.
5558 pub anchor: Option<String>,
5659}
@@ -71,14 +74,16 @@ impl SymbolTable {
7174 /// the renderer emits for that target's heading.
7275 pub fn index_document(&mut self, doc: &Document) {
7376 let path = &doc.source_path;
77 let out = output_path(path, &doc.keywords);
7478 self.targets.insert(
7579 TargetId::File(path.clone()),
7680 TargetLocation {
7781 source_path: path.clone(),
82 output_path: out.clone(),
7883 anchor: None,
7984 },
8085 );
81 index_section(&doc.root, path, &mut self.targets);
86 index_section(&doc.root, path, &out, &mut self.targets);
8287 }
8388}
8489
@@ -107,40 +112,37 @@ fn collect_targets(section: &Section, out: &mut Vec<TargetId>) {
107112 }
108113}
109114
110fn index_section(section: &Section, path: &Utf8Path, targets: &mut HashMap<TargetId, TargetLocation>) {
115fn index_section(
116 section: &Section,
117 path: &Utf8Path,
118 out: &Utf8Path,
119 targets: &mut HashMap<TargetId, TargetLocation>,
120) {
111121 if let Some(h) = &section.heading {
112122 let anchor = h
113123 .custom_id
114124 .clone()
115125 .or_else(|| h.id.clone())
116126 .unwrap_or_else(|| slugify(&plain_text(&h.title)));
117 if let Some(cid) = &h.custom_id {
127 let mut record = |id: TargetId, anchor: Option<String>| {
118128 targets.insert(
119 TargetId::CustomId(cid.clone()),
129 id,
120130 TargetLocation {
121131 source_path: path.to_owned(),
122 anchor: Some(cid.clone()),
132 output_path: out.to_owned(),
133 anchor,
123134 },
124135 );
136 };
137 if let Some(cid) = &h.custom_id {
138 record(TargetId::CustomId(cid.clone()), Some(cid.clone()));
125139 }
126140 if let Some(id) = &h.id {
127 targets.insert(
128 TargetId::Id(id.clone()),
129 TargetLocation {
130 source_path: path.to_owned(),
131 anchor: Some(id.clone()),
132 },
133 );
141 record(TargetId::Id(id.clone()), Some(id.clone()));
134142 }
135 targets.insert(
136 TargetId::Heading(plain_text(&h.title)),
137 TargetLocation {
138 source_path: path.to_owned(),
139 anchor: Some(anchor),
140 },
141 );
143 record(TargetId::Heading(plain_text(&h.title)), Some(anchor));
142144 }
143145 for child in &section.children {
144 index_section(child, path, targets);
146 index_section(child, path, out, targets);
145147 }
146148}
src/lib.rs +1
@@ -8,6 +8,7 @@
88//! [`render`] (RENDER) → [`template`] (TEMPLATE) → EMIT, with [`incremental`]
99//! deciding which pages actually need rewriting.
1010
11pub mod audit;
1112pub mod incremental;
1213pub mod index;
1314pub mod model;
src/main.rs +11
@@ -49,6 +49,12 @@ enum Command {
4949 /// Output directory to remove.
5050 output: Utf8PathBuf,
5151 },
52 /// Audit a corpus: report which org constructs it uses and how they land against
53 /// the v1 scope line. Reports names, counts and locations — never document text.
54 Audit {
55 /// Source directory (or single `.org` file) to audit.
56 input: Utf8PathBuf,
57 },
5258}
5359
5460fn main() -> Result<()> {
@@ -86,6 +92,11 @@ fn main() -> Result<()> {
8692 // 6 lists `watch`; the real fs-notify integration is deferred). It rebuilds
8793 // incrementally whenever any source file's mtime advances.
8894 Command::Watch { input, output } => watch(&input, &output),
95 Command::Audit { input } => {
96 let result = org_ssg::audit::audit(&input)?;
97 print!("{}", org_ssg::audit::report(&result));
98 Ok(())
99 }
89100 Command::Clean { output } => {
90101 if output.exists() {
91102 fs::remove_dir_all(&output)
src/resolve.rs +6 −2
@@ -15,7 +15,7 @@ use camino::Utf8Path;
1515
1616use crate::index::{SymbolTable, TargetId};
1717use crate::model::{Document, Element, Link, LinkTarget, Object, Section, TableRow};
18use crate::util::{normalize_link_path, output_url};
18use crate::util::{normalize_link_path, output_path, output_url};
1919
2020/// A document whose links have been rewritten to concrete URLs.
2121#[derive(Debug, Clone)]
@@ -42,10 +42,13 @@ pub struct ResolveOutput {
4242pub fn resolve(doc: &Document, symbols: &SymbolTable) -> ResolveOutput {
4343 let mut document = doc.clone();
4444 let from = doc.source_path.clone();
45 // URLs are computed between *output* paths, which `#+SLUG:` can rename.
46 let from_out = output_path(&from, &doc.keywords);
4547 let mut used = Vec::new();
4648 let mut broken = Vec::new();
4749 let mut cx = Cx {
4850 from: &from,
51 from_out: &from_out,
4952 symbols,
5053 used: &mut used,
5154 broken: &mut broken,
@@ -70,6 +73,7 @@ fn human_text(target: &LinkTarget) -> Option<String> {
7073
7174struct Cx<'a> {
7275 from: &'a Utf8Path,
76 from_out: &'a Utf8Path,
7377 symbols: &'a SymbolTable,
7478 used: &'a mut Vec<TargetId>,
7579 broken: &'a mut Vec<BrokenLink>,
@@ -167,7 +171,7 @@ impl Cx<'_> {
167171 link.description = Some(vec![Object::Text(text)]);
168172 }
169173 }
170 let url = output_url(self.from, &loc.source_path, loc.anchor.as_deref());
174 let url = output_url(self.from_out, &loc.output_path, loc.anchor.as_deref());
171175 link.target = LinkTarget::External(url);
172176 }
173177 None => {
src/site.rs +25 −6
@@ -26,7 +26,7 @@ use crate::parser::parse;
2626use crate::render::{render, syntax_css, Html, SyntectHighlighter};
2727use crate::resolve::resolve;
2828use crate::template::{template_sources, NavItem, Templater};
29use crate::util::{output_url, relative_root};
29use crate::util::{output_path, output_url, relative_root};
3030
3131/// A fully built page: source and output paths (relative to their roots) and its
3232/// final templated HTML.
@@ -100,12 +100,27 @@ fn prepare_pages(src: &Utf8Path) -> Result<(Vec<PagePrep>, SymbolTable)> {
100100 symbols.index_document(doc);
101101 }
102102
103 // Nav is global; titles come from #+TITLE (falling back to the file stem).
103 // Nav is global; titles come from #+TITLE (falling back to the file stem) and URLs
104 // from each page's output path, which `#+SLUG:` can rename.
104105 let entries: Vec<(Utf8PathBuf, String)> = docs
105106 .iter()
106 .map(|d| (d.source_path.clone(), page_title(d)))
107 .map(|d| (output_path(&d.source_path, &d.keywords), page_title(d)))
107108 .collect();
108109
110 // Two sources emitting one page would silently drop a page — and with slugs, a
111 // collision is a typo away and invisible in the source filenames.
112 let mut claimed: std::collections::HashMap<&Utf8PathBuf, &Utf8PathBuf> =
113 std::collections::HashMap::new();
114 for (doc, (out, _)) in docs.iter().zip(&entries) {
115 if let Some(other) = claimed.insert(out, &doc.source_path) {
116 anyhow::bail!(
117 "output collision: {} and {} both build to {out} (check their #+SLUG:)",
118 other,
119 doc.source_path
120 );
121 }
122 }
123
109124 let mut pages = Vec::new();
110125 for doc in &docs {
111126 let out = resolve(doc, &symbols);
@@ -113,18 +128,20 @@ fn prepare_pages(src: &Utf8Path) -> Result<(Vec<PagePrep>, SymbolTable)> {
113128 let broken: Vec<TargetId> = out.broken.iter().map(|b| b.target.clone()).collect();
114129 let defines: HashSet<TargetId> = document_targets(doc).into_iter().collect();
115130
131 let output = output_path(&doc.source_path, &doc.keywords);
132
116133 // Nav links are relative to *this* page (spec URL scheme, §8 Q3).
117134 let nav: Vec<NavItem> = entries
118135 .iter()
119136 .map(|(path, title)| NavItem {
120137 title: title.clone(),
121 url: output_url(&doc.source_path, path, None),
138 url: output_url(&output, path, None),
122139 })
123140 .collect();
124141
125142 pages.push(PagePrep {
126143 source: doc.source_path.clone(),
127 output: doc.source_path.with_extension("html"),
144 output,
128145 title: page_title(doc),
129146 content_hash: doc.content_hash,
130147 resolved: out.resolved,
@@ -190,9 +207,11 @@ pub fn build_site(src: &Utf8Path, out: &Utf8Path, opts: &BuildOptions) -> Result
190207 // chrome on every page — is built from every page's (path, title), so a title/path
191208 // change or a page add/remove must re-render every page (else stale nav on disk).
192209 let cfg = BuildConfig::default();
210 // Keyed on the *output* path: a `#+SLUG:` change moves a page's URL, which changes
211 // the nav on every other page even though no source filename moved.
193212 let nav_entries: Vec<(String, String)> = preps
194213 .iter()
195 .map(|p| (p.source.to_string(), p.title.clone()))
214 .map(|p| (p.output.to_string(), p.title.clone()))
196215 .collect();
197216 let cfg_hash = combine(config_hash(&cfg), site_structure_hash(&nav_entries));
198217 let tmpl_hash = template_hash(template_sources());
src/util.rs +55 −9
@@ -4,7 +4,50 @@
44
55use camino::{Utf8Path, Utf8PathBuf};
66
7use crate::model::Object;
7use crate::model::{Keywords, Object};
8
9/// The output path for a document, relative to the site root.
10///
11/// Normally this is the source path with `.org` swapped for `.html`, but a `#+SLUG:`
12/// keyword renames the file — which is how the target corpus works: 178 of its 179 files
13/// set one, and `2018-11-28-aes-encryption.org` publishes as `aes-encryption.html`. The
14/// slug names the *file*, never the directory, so the page stays where its source lives.
15pub fn output_path(source: &Utf8Path, keywords: &Keywords) -> Utf8PathBuf {
16 let slug = keywords
17 .entries
18 .iter()
19 .find(|(k, _)| k.eq_ignore_ascii_case("SLUG"))
20 .map(|(_, v)| sanitize_slug(v))
21 .filter(|s| !s.is_empty());
22
23 match slug {
24 Some(slug) => {
25 let dir = source.parent().unwrap_or_else(|| Utf8Path::new(""));
26 dir.join(format!("{slug}.html"))
27 }
28 None => source.with_extension("html"),
29 }
30}
31
32/// Reduce a slug to a safe single filename component.
33///
34/// A slug is author-controlled text that becomes a path we write to, so `../../etc/x`
35/// has to be impossible by construction rather than by convention: separators and dots
36/// are folded to `-`, which cannot traverse and cannot produce a hidden file.
37fn sanitize_slug(raw: &str) -> String {
38 let mut out = String::with_capacity(raw.len());
39 let mut prev_dash = false;
40 for c in raw.trim().chars() {
41 if c.is_ascii_alphanumeric() || c == '_' {
42 out.extend(c.to_lowercase());
43 prev_dash = false;
44 } else if !prev_dash {
45 out.push('-');
46 prev_dash = true;
47 }
48 }
49 out.trim_matches('-').to_string()
50}
851
952/// Flatten inline objects to their plain-text content (markup stripped). Used to
1053/// derive heading anchors and `[[*Heading]]` link identities (spec §4.3).
@@ -49,16 +92,19 @@ pub fn slugify(text: &str) -> String {
4992 out.trim_matches('-').to_string()
5093}
5194
52/// The output URL to reach `to_rel` (a source `.org` path relative to the site root)
53/// from the page at `from_rel`, honoring an optional `anchor`. Same-file links reduce
54/// to a bare `#anchor` fragment; cross-file links become a relative `.html` path.
55pub fn output_url(from_rel: &Utf8Path, to_rel: &Utf8Path, anchor: Option<&str>) -> String {
56 let path = if from_rel == to_rel {
95/// The URL to reach the page output at `to_out` from the page output at `from_out`,
96/// honoring an optional `anchor`. Same-page links reduce to a bare `#anchor` fragment;
97/// cross-page links become a relative path.
98///
99/// Both arguments are *output* paths, not source paths, because `#+SLUG:` means the two
100/// no longer correspond: deriving the URL here would reintroduce the filename assumption
101/// that [`output_path`] exists to remove.
102pub fn output_url(from_out: &Utf8Path, to_out: &Utf8Path, anchor: Option<&str>) -> String {
103 let path = if from_out == to_out {
57104 String::new()
58105 } else {
59 let to_html = to_rel.with_extension("html");
60 let from_dir = from_rel.parent().unwrap_or_else(|| Utf8Path::new(""));
61 relative_path(from_dir, &to_html)
106 let from_dir = from_out.parent().unwrap_or_else(|| Utf8Path::new(""));
107 relative_path(from_dir, to_out)
62108 };
63109 match anchor {
64110 Some(a) if !a.is_empty() => {
tests/oracle.el added +42
@@ -0,0 +1,42 @@
1;;; oracle.el --- ground-truth HTML export for the org-ssg differential tests -*- lexical-binding: t -*-
2
3;; Exports the org file named by $ORG_ORACLE_INPUT to HTML on stdout, using org's own
4;; exporter — the same one weblorg wraps to publish the corpus this project targets.
5;; Run with: ORG_ORACLE_INPUT=x.org emacs -Q --batch -l tests/oracle.el
6;;
7;; `-Q' is deliberate: no user init, so the oracle is the stock org exporter and not
8;; this machine's Emacs configuration. The path is passed by environment variable
9;; rather than as an argument because batch Emacs would otherwise try to visit it.
10
11(require 'org)
12(require 'ox-html)
13
14;; Presentation settings are normalized so the diff carries semantic divergences only.
15;; Everything that affects *content* is left at its default, because the point is to
16;; learn what stock org does — normalizing that away would be marking our own homework.
17(setq org-export-with-toc nil ; we emit no table of contents
18 org-export-with-section-numbers nil ; we do not number headings
19 org-html-toplevel-hlevel 1 ; org defaults to h2 for a level-1 heading,
20 ; because a template supplies the page <h1>.
21 ; Aligning here keeps a global +1 offset from
22 ; drowning every real finding in the diff.
23 org-html-htmlize-output-type nil ; plain <pre>, not htmlize spans: we highlight
24 ; with syntect, so comparing code *text* is
25 ; the meaningful part
26 org-html-head-include-default-style nil
27 org-html-head-include-scripts nil
28 ;; Fixtures link to ids that live in org-ssg's own symbol table, not in an
29 ;; `org-id' database. Without this, org aborts the whole export on the first one.
30 org-export-with-broken-links t
31 make-backup-files nil)
32
33(let ((input (getenv "ORG_ORACLE_INPUT")))
34 (unless input
35 (error "ORG_ORACLE_INPUT is not set"))
36 (with-temp-buffer
37 (insert-file-contents input)
38 (org-mode)
39 ;; BODY-ONLY: emit the content, not a full document with <head> chrome.
40 (princ (org-export-as 'html nil nil t nil))))
41
42;;; oracle.el ends here
tests/oracle.rs added +454
@@ -0,0 +1,454 @@
1//! The `emacs --batch` ground-truth oracle (spec §5, Phase 0).
2//!
3//! Every other test in this suite checks org-ssg against org-ssg: a snapshot says our
4//! output has not *changed*, never that it is *right*. Those two questions are different,
5//! and only one of them matters to someone whose site is currently published by Emacs.
6//! This file answers the second by exporting the same fixture with org's own HTML
7//! exporter — the exporter weblorg wraps to publish the target corpus — and diffing the
8//! two.
9//!
10//! **What is compared.** Byte equality is not a useful goal: org wraps every section in
11//! `outline-container` divs keyed by generated ids, and no amount of agreement on
12//! semantics would survive that. Both sides are reduced to a *semantic skeleton* — the
13//! sequence of element opens, closes, and text runs, with `<div>`s and all attributes
14//! except `href`/`src` dropped, whitespace collapsed, and entities decoded. What remains
15//! is the question worth asking: does org think this is a `<blockquote><p>`, and do we?
16//!
17//! **What the result means.** These tests do not assert agreement — they *snapshot the
18//! disagreement*. A divergence report that is checked in and reviewed is worth more than
19//! a red test nobody can act on, and it makes any new divergence show up as a diff in
20//! code review. A few invariants that must never break are asserted outright.
21//!
22//! The suite skips cleanly when Emacs is absent, so it never blocks a machine or CI
23//! runner that has no Emacs.
24
25use std::process::Command;
26
27use camino::Utf8PathBuf;
28
29use org_ssg::parser::parse;
30use org_ssg::render::{render, Html, SyntectHighlighter};
31use org_ssg::resolve::ResolvedDoc;
32
33fn manifest_dir() -> Utf8PathBuf {
34 Utf8PathBuf::from(env!("CARGO_MANIFEST_DIR"))
35}
36
37/// Is a usable Emacs on PATH? The oracle is a development instrument, not a build
38/// dependency, so its absence skips rather than fails.
39fn emacs_available() -> bool {
40 Command::new("emacs")
41 .arg("--version")
42 .output()
43 .map(|o| o.status.success())
44 .unwrap_or(false)
45}
46
47/// Export a fixture with org's own HTML exporter.
48fn org_export(fixture: &str) -> String {
49 let root = manifest_dir();
50 let output = Command::new("emacs")
51 .args(["-Q", "--batch", "-l"])
52 .arg(root.join("tests/oracle.el"))
53 .env("ORG_ORACLE_INPUT", root.join("fixtures").join(fixture))
54 .current_dir(&root)
55 .output()
56 .expect("run emacs");
57 assert!(
58 output.status.success(),
59 "emacs export of {fixture} failed:\n{}",
60 String::from_utf8_lossy(&output.stderr)
61 );
62 String::from_utf8(output.stdout).expect("emacs emits UTF-8")
63}
64
65/// Render a fixture with org-ssg.
66fn our_export(fixture: &str) -> String {
67 let path = manifest_dir().join("fixtures").join(fixture);
68 let source = std::fs::read_to_string(&path).expect("read fixture");
69 let document = parse(Utf8PathBuf::from(fixture).as_path(), &source).expect("parse");
70 let Html(html) = render(&ResolvedDoc { document }, &SyntectHighlighter::new());
71 html
72}
73
74// ---------------------------------------------------------------------------
75// HTML → semantic skeleton
76// ---------------------------------------------------------------------------
77
78/// Elements dropped from the skeleton entirely, because once attributes are gone they
79/// carry no meaning the two exporters could agree or disagree *about*.
80///
81/// `div` is pure layout: org wraps every section in `outline-container`/`outline-text`
82/// wrappers and we emit none. `span` is the same story at the inline level, and matters
83/// far more than it looks: syntect emits one span per code token, so keeping them made a
84/// source block contribute ~60 skeleton lines of pure noise and dragged the agreement on
85/// `blocks.org` down to 36% — a number that said nothing about whether we render blocks
86/// correctly. Text still carries the signal: a `<span class="todo">` shows up as its
87/// text, `"TODO"`, which is the part worth comparing.
88const IGNORED: &[&str] = &["div", "span"];
89
90/// Attributes kept in the skeleton. Ids and classes are generated (`org6c28c1b`) or
91/// cosmetic (`org-ul`); `href` and `src` are the content.
92const KEPT_ATTRS: &[&str] = &["href", "src"];
93
94/// HTML void elements, which never emit a close event.
95const VOID: &[&str] = &[
96 "br", "hr", "img", "input", "meta", "link", "col", "area", "base", "source", "wbr",
97];
98
99/// Reduce an HTML fragment to its semantic skeleton: one line per element open, element
100/// close, or text run.
101fn skeleton(html: &str) -> Vec<String> {
102 let mut out = Vec::new();
103 let chars: Vec<char> = html.chars().collect();
104 let mut i = 0;
105 let mut text = String::new();
106
107 while i < chars.len() {
108 if chars[i] != '<' {
109 text.push(chars[i]);
110 i += 1;
111 continue;
112 }
113
114 // Comments and doctypes carry nothing.
115 if chars[i..].starts_with(&['<', '!']) {
116 i += match find_from(&chars, i, ">") {
117 Some(end) => end - i + 1,
118 None => break,
119 };
120 continue;
121 }
122 let Some(end) = find_from(&chars, i, ">") else {
123 break;
124 };
125 let raw: String = chars[i + 1..end].iter().collect();
126 i = end + 1;
127
128 let raw = raw.trim().trim_end_matches('/').trim().to_string();
129 // Text is flushed only when a tag is actually *emitted*. Text either side of an
130 // ignored tag therefore merges into one run, which is what makes a highlighted
131 // source block compare as the one string of code it is, rather than as a
132 // token-by-token sequence that has to line up exactly.
133 if let Some(name) = raw.strip_prefix('/') {
134 let name = name.trim().to_ascii_lowercase();
135 if !IGNORED.contains(&name.as_str()) && !VOID.contains(&name.as_str()) {
136 flush_text(&mut text, &mut out);
137 out.push(format!("</{name}>"));
138 }
139 continue;
140 }
141 let mut parts = raw.splitn(2, char::is_whitespace);
142 let name = parts.next().unwrap_or("").to_ascii_lowercase();
143 if name.is_empty() || IGNORED.contains(&name.as_str()) {
144 continue;
145 }
146 let attrs = kept_attributes(parts.next().unwrap_or(""));
147 flush_text(&mut text, &mut out);
148 out.push(format!("<{name}{attrs}>"));
149 }
150 flush_text(&mut text, &mut out);
151 out
152}
153
154fn flush_text(text: &mut String, out: &mut Vec<String>) {
155 let decoded = decode_entities(text);
156 let collapsed = decoded.split_whitespace().collect::<Vec<_>>().join(" ");
157 if !collapsed.is_empty() {
158 out.push(format!("{collapsed:?}"));
159 }
160 text.clear();
161}
162
163fn find_from(chars: &[char], from: usize, needle: &str) -> Option<usize> {
164 let n: Vec<char> = needle.chars().collect();
165 (from..chars.len()).find(|&k| chars[k..].starts_with(&n[..]))
166}
167
168/// Keep only the content-bearing attributes, in a stable order.
169fn kept_attributes(rest: &str) -> String {
170 let mut kept: Vec<(String, String)> = Vec::new();
171 for attr in KEPT_ATTRS {
172 if let Some(value) = attribute_value(rest, attr) {
173 kept.push(((*attr).to_string(), value));
174 }
175 }
176 kept.iter()
177 .map(|(k, v)| format!(" {k}=\"{}\"", decode_entities(v)))
178 .collect()
179}
180
181fn attribute_value(rest: &str, name: &str) -> Option<String> {
182 let mut search = rest;
183 while let Some(pos) = search.find(name) {
184 let before_ok = pos == 0
185 || search[..pos]
186 .chars()
187 .next_back()
188 .is_some_and(char::is_whitespace);
189 let after = &search[pos + name.len()..];
190 let after_trimmed = after.trim_start();
191 if before_ok && after_trimmed.starts_with('=') {
192 let value = after_trimmed[1..].trim_start();
193 let quote = value.chars().next()?;
194 if quote == '"' || quote == '\'' {
195 let end = value[1..].find(quote)? + 1;
196 return Some(value[1..end].to_string());
197 }
198 let end = value.find(char::is_whitespace).unwrap_or(value.len());
199 return Some(value[..end].to_string());
200 }
201 search = &search[pos + name.len()..];
202 }
203 None
204}
205
206/// Decode the entities either exporter is likely to emit, so an encoding difference is
207/// never reported as a semantic one.
208fn decode_entities(s: &str) -> String {
209 let mut out = String::with_capacity(s.len());
210 let mut rest = s;
211 while let Some(amp) = rest.find('&') {
212 out.push_str(&rest[..amp]);
213 let tail = &rest[amp..];
214 let Some(semi) = tail.find(';').filter(|s| *s <= 12) else {
215 out.push('&');
216 rest = &tail[1..];
217 continue;
218 };
219 let entity = &tail[1..semi];
220 let decoded = match entity {
221 "amp" => Some('&'),
222 "lt" => Some('<'),
223 "gt" => Some('>'),
224 "quot" => Some('"'),
225 "apos" => Some('\''),
226 "nbsp" => Some(' '),
227 _ => entity
228 .strip_prefix('#')
229 .and_then(|n| match n.strip_prefix(['x', 'X']) {
230 Some(hex) => u32::from_str_radix(hex, 16).ok(),
231 None => n.parse::<u32>().ok(),
232 })
233 .and_then(char::from_u32),
234 };
235 match decoded {
236 // A non-breaking space is a space for comparison purposes.
237 Some('\u{a0}') => out.push(' '),
238 Some(c) => out.push(c),
239 None => {
240 out.push('&');
241 rest = &tail[1..];
242 continue;
243 }
244 }
245 rest = &tail[semi + 1..];
246 }
247 out.push_str(rest);
248 out
249}
250
251// ---------------------------------------------------------------------------
252// Divergence report
253// ---------------------------------------------------------------------------
254
255/// A unified diff of the two skeletons, via a longest-common-subsequence walk. `-` is
256/// org-ssg, `+` is Emacs.
257fn divergence(ours: &[String], theirs: &[String]) -> String {
258 let (n, m) = (ours.len(), theirs.len());
259 // lcs[i][j] = length of the longest common subsequence of ours[i..] and theirs[j..].
260 let mut lcs = vec![vec![0usize; m + 1]; n + 1];
261 for i in (0..n).rev() {
262 for j in (0..m).rev() {
263 lcs[i][j] = if ours[i] == theirs[j] {
264 lcs[i + 1][j + 1] + 1
265 } else {
266 lcs[i + 1][j].max(lcs[i][j + 1])
267 };
268 }
269 }
270
271 let mut out = String::new();
272 let (mut i, mut j) = (0, 0);
273 let mut agreed = 0usize;
274 while i < n && j < m {
275 if ours[i] == theirs[j] {
276 out.push_str(&format!(" {}\n", ours[i]));
277 agreed += 1;
278 i += 1;
279 j += 1;
280 } else if lcs[i + 1][j] >= lcs[i][j + 1] {
281 out.push_str(&format!("- {}\n", ours[i]));
282 i += 1;
283 } else {
284 out.push_str(&format!("+ {}\n", theirs[j]));
285 j += 1;
286 }
287 }
288 for line in &ours[i..] {
289 out.push_str(&format!("- {line}\n"));
290 }
291 for line in &theirs[j..] {
292 out.push_str(&format!("+ {line}\n"));
293 }
294
295 let total = n.max(m);
296 let pct = if total == 0 {
297 100.0
298 } else {
299 100.0 * agreed as f64 / total as f64
300 };
301 format!("agreement: {agreed}/{total} skeleton lines ({pct:.1}%)\n(- org-ssg, + emacs)\n\n{out}")
302}
303
304/// Snapshot the divergence between org-ssg and Emacs for one fixture.
305fn compare(fixture: &str) -> Option<String> {
306 if !emacs_available() {
307 eprintln!("skipping oracle comparison for {fixture}: no emacs on PATH");
308 return None;
309 }
310 let ours = skeleton(&our_export(fixture));
311 let theirs = skeleton(&org_export(fixture));
312 Some(divergence(&ours, &theirs))
313}
314
315macro_rules! oracle_test {
316 ($name:ident, $fixture:literal) => {
317 #[test]
318 fn $name() {
319 if let Some(report) = compare($fixture) {
320 insta::assert_snapshot!(report);
321 }
322 }
323 };
324}
325
326oracle_test!(oracle_minimal, "minimal.org");
327oracle_test!(oracle_core, "core.org");
328oracle_test!(oracle_headings, "headings.org");
329oracle_test!(oracle_lists, "lists.org");
330oracle_test!(oracle_blocks, "blocks.org");
331oracle_test!(oracle_table, "table.org");
332oracle_test!(oracle_footnote, "footnote.org");
333oracle_test!(oracle_timestamps, "timestamps.org");
334oracle_test!(oracle_images, "images.org");
335oracle_test!(oracle_elements, "elements.org");
336
337// ---------------------------------------------------------------------------
338// Invariants that must hold against the oracle, not merely be snapshotted
339// ---------------------------------------------------------------------------
340
341/// How many headings a document has and at what depth is the shape of the document.
342/// Getting it wrong reorganizes someone's writing, so it is asserted rather than
343/// snapshotted. Heading *decoration* (priority cookies, tag markup) is a policy
344/// difference and is left to the snapshots.
345#[test]
346fn heading_structure_matches_emacs() {
347 if !emacs_available() {
348 eprintln!("skipping: no emacs on PATH");
349 return;
350 }
351 for fixture in ["minimal.org", "core.org", "headings.org", "lists.org"] {
352 let ours = heading_levels(&skeleton(&our_export(fixture)));
353 let theirs = heading_levels(&skeleton(&org_export(fixture)));
354 assert_eq!(
355 ours, theirs,
356 "heading structure diverges from Emacs in {fixture}"
357 );
358 }
359}
360
361/// The sequence of heading open tags, e.g. `["<h1>", "<h2>", "<h1>"]`.
362fn heading_levels(skeleton: &[String]) -> Vec<String> {
363 skeleton
364 .iter()
365 .filter(|l| l.starts_with("<h") && l[2..].starts_with(|c: char| c.is_ascii_digit()))
366 .cloned()
367 .collect()
368}
369
370/// A list is the construct where nesting is easiest to get subtly wrong, and where being
371/// wrong changes the meaning of the document rather than its looks.
372#[test]
373fn list_nesting_matches_emacs() {
374 if !emacs_available() {
375 eprintln!("skipping: no emacs on PATH");
376 return;
377 }
378 let ours = list_shape(&skeleton(&our_export("lists.org")));
379 let theirs = list_shape(&skeleton(&org_export("lists.org")));
380 assert_eq!(ours, theirs, "list nesting diverges from Emacs");
381}
382
383/// The sequence of list opens/closes, ignoring content — the shape of the nesting.
384fn list_shape(skeleton: &[String]) -> Vec<String> {
385 skeleton
386 .iter()
387 .filter(|l| {
388 matches!(
389 l.as_str(),
390 "<ul>" | "</ul>" | "<ol>" | "</ol>" | "<li>" | "</li>" | "<dl>" | "</dl>"
391 | "<dt>" | "</dt>" | "<dd>" | "</dd>"
392 )
393 })
394 .cloned()
395 .collect()
396}
397
398/// Code must survive verbatim. Highlighting markup differs by construction (syntect
399/// spans vs htmlize), but if the *characters of the program* differ, we have corrupted
400/// the author's content.
401#[test]
402fn source_block_text_matches_emacs() {
403 if !emacs_available() {
404 eprintln!("skipping: no emacs on PATH");
405 return;
406 }
407 for fixture in ["blocks.org", "core.org", "elements.org"] {
408 let ours = code_text(&our_export(fixture));
409 let theirs = code_text(&org_export(fixture));
410 assert_eq!(ours, theirs, "source block text diverges from Emacs in {fixture}");
411 }
412}
413
414/// All text inside `<pre>` blocks, with tags stripped and whitespace collapsed.
415fn code_text(html: &str) -> Vec<String> {
416 let mut out = Vec::new();
417 let mut rest = html;
418 while let Some(start) = rest.find("<pre") {
419 let after = &rest[start..];
420 let Some(open_end) = after.find('>') else { break };
421 let Some(close) = after.find("</pre>") else { break };
422 let inner = &after[open_end + 1..close];
423 out.push(strip_tags(inner));
424 rest = &after[close + 6..];
425 }
426 out
427}
428
429/// All text in a fragment with tags removed and entities decoded, then whitespace
430/// collapsed once at the end.
431///
432/// [`skeleton`] cannot do this job: it trims each text run individually, which is
433/// invisible for prose (one run per paragraph) but destructive for highlighted code,
434/// where syntect splits a line into one run per token and the spaces *between* tokens
435/// live at the edges of those runs. Trimming each run turns `def greet` into `defgreet`.
436fn strip_tags(html: &str) -> String {
437 let mut text = String::new();
438 let mut rest = html;
439 while let Some(open) = rest.find('<') {
440 text.push_str(&rest[..open]);
441 match rest[open..].find('>') {
442 Some(close) => rest = &rest[open + close + 1..],
443 None => {
444 rest = "";
445 break;
446 }
447 }
448 }
449 text.push_str(rest);
450 decode_entities(&text)
451 .split_whitespace()
452 .collect::<Vec<_>>()
453 .join(" ")
454}
tests/site.rs +83
@@ -122,3 +122,86 @@ fn table_render() {
122122fn footnote_render() {
123123 insta::assert_snapshot!(render_fragment("footnote.org"));
124124}
125
126// ---------------------------------------------------------------------------
127// `#+SLUG:` output paths (Phase 0 corpus-audit finding)
128// ---------------------------------------------------------------------------
129
130/// The audit found `#+SLUG:` in 178 of the target corpus's 179 files, and the live site
131/// derives every URL from it — `2018-11-28-aes-encryption.org` publishes as
132/// `aes-encryption.html`. Deriving output paths from source filenames would therefore
133/// have rewritten every URL on the site.
134#[test]
135fn slug_renames_the_output_page() {
136 let (pages, broken) = render_site(&fixtures().join("slugsite")).expect("build site");
137 assert!(broken.is_empty(), "fixture site has no broken links: {broken:?}");
138 let post = pages
139 .iter()
140 .find(|p| p.source == "2024-02-11-long-source-name.org")
141 .expect("post page");
142 assert_eq!(
143 post.output, "short-url.html",
144 "the slug names the output file, not the source stem"
145 );
146}
147
148/// A link's URL has to follow the target's slug. If resolution kept using source paths,
149/// every cross-page link would point at a file that was never written.
150#[test]
151fn links_resolve_through_the_slug() {
152 let (pages, _) = render_site(&fixtures().join("slugsite")).expect("build site");
153 let index = &page(&pages, "index.org").html;
154 assert!(
155 index.contains("href=\"short-url.html\""),
156 "a file: link must target the slugged page:\n{index}"
157 );
158 assert!(
159 index.contains("href=\"short-url.html#setup\""),
160 "a custom-id link must target the slugged page plus the anchor:\n{index}"
161 );
162 assert!(
163 !index.contains("long-source-name"),
164 "no URL may mention the source filename:\n{index}"
165 );
166}
167
168/// A slug is author-controlled text that becomes a path we write to, so traversal has to
169/// be impossible by construction rather than by convention.
170#[test]
171fn slugs_cannot_escape_the_output_directory() {
172 use org_ssg::model::Keywords;
173 let source = Utf8PathBuf::from("blog/post.org");
174 let slugged = |value: &str| {
175 let keywords = Keywords {
176 entries: vec![("SLUG".to_string(), value.to_string())],
177 };
178 org_ssg::util::output_path(&source, &keywords).to_string()
179 };
180 assert_eq!(slugged("../../etc/passwd"), "blog/etc-passwd.html");
181 assert_eq!(slugged("/absolute"), "blog/absolute.html");
182 assert_eq!(slugged(".hidden"), "blog/hidden.html");
183 assert_eq!(slugged("Mixed Case Slug"), "blog/mixed-case-slug.html");
184 // An empty or punctuation-only slug falls back to the source stem rather than
185 // producing `.html` with no name at all.
186 assert_eq!(slugged("///"), "blog/post.html");
187}
188
189/// Two pages claiming one URL silently drops a page. With slugs that is a typo away and
190/// invisible in the source filenames, so the build refuses rather than picking a winner.
191#[test]
192fn colliding_slugs_are_a_build_error() {
193 let dir = std::env::temp_dir().join(format!("org-ssg-slug-{}", std::process::id()));
194 let dir = Utf8PathBuf::from_path_buf(dir).expect("utf-8 temp dir");
195 let _ = std::fs::remove_dir_all(&dir);
196 std::fs::create_dir_all(&dir).unwrap();
197 std::fs::write(dir.join("a.org"), "#+TITLE: A\n#+SLUG: same\n").unwrap();
198 std::fs::write(dir.join("b.org"), "#+TITLE: B\n#+SLUG: same\n").unwrap();
199
200 let err = render_site(&dir).expect_err("colliding slugs must fail the build");
201 let message = format!("{err:#}");
202 assert!(
203 message.contains("collision") && message.contains("same.html"),
204 "the error must name the collision: {message}"
205 );
206 std::fs::remove_dir_all(&dir).unwrap();
207}
tests/snapshots/oracle__oracle_blocks.snap added +68
@@ -0,0 +1,68 @@
1---
2source: tests/oracle.rs
3expression: report
4---
5agreement: 51/59 skeleton lines (86.4%)
6(- org-ssg, + emacs)
7
8 <h1>
9 "Quote"
10 </h1>
11 <blockquote>
12 <p>
13 "A quoted paragraph with"
14- <em>
15+ <i>
16 "markup"
17- </em>
18+ </i>
19 "."
20 </p>
21 <p>
22 "And a second paragraph."
23 </p>
24 </blockquote>
25 <h1>
26 "Center"
27 </h1>
28 <p>
29 "Centred text."
30 </p>
31 <h1>
32 "Example"
33 </h1>
34 <pre>
35 "Verbatim *not bold* text. Indentation preserved."
36 </pre>
37 <h1>
38 "Export"
39 </h1>
40 <aside>
41 "Raw HTML passes through."
42 </aside>
43 <h1>
44 "Source"
45 </h1>
46 <pre>
47- <code>
48 "def greet(name): return f\"hello {name}\""
49- </code>
50 </pre>
51 <pre>
52- <code>
53 "plain block, no language"
54- </code>
55 </pre>
56 <h1>
57 "Nested"
58 </h1>
59 <blockquote>
60 <p>
61 "A quote containing a source block:"
62 </p>
63 <pre>
64- <code>
65 "echo hi"
66- </code>
67 </pre>
68 </blockquote>
tests/snapshots/oracle__oracle_core.snap added +70
@@ -0,0 +1,70 @@
1---
2source: tests/oracle.rs
3expression: report
4---
5agreement: 45/54 skeleton lines (83.3%)
6(- org-ssg, + emacs)
7
8 <p>
9 "Intro paragraph with a bare URL"
10 <a href="https://example.com">
11 "https://example.com"
12 </a>
13 "and some"
14 <code>
15 "inline code"
16 </code>
17 "."
18 </p>
19 <h1>
20 "Ordered and checked"
21 </h1>
22 <ol>
23 <li>
24 "first item"
25 </li>
26 <li>
27 "second item with"
28- <em>
29+ <i>
30 "emphasis"
31- </em>
32+ </i>
33 </li>
34- </ol>
35- <ul>
36 <li>
37- <input>
38+ <code>
39+ "[ ]"
40+ </code>
41 "todo item"
42 </li>
43 <li>
44- <input>
45+ <code>
46+ "[X]"
47+ </code>
48 "done item"
49 </li>
50- </ul>
51+ </ol>
52 <h1>
53 "Links and code"
54 </h1>
55 <p>
56 "An external"
57 <a href="https://example.org">
58 "site"
59 </a>
60 "and a bare"
61 <a href="https://bare.example">
62 "https://bare.example"
63 </a>
64 "."
65 </p>
66 <pre>
67- <code>
68 "fn main() { println!(\"hello\"); }"
69- </code>
70 </pre>
tests/snapshots/oracle__oracle_elements.snap added +103
@@ -0,0 +1,103 @@
1---
2source: tests/oracle.rs
3expression: report
4---
5agreement: 64/82 skeleton lines (78.0%)
6(- org-ssg, + emacs)
7
8 <h1>
9 "Code and tables"
10 </h1>
11 <pre>
12- <code>
13 "fn main() { println!(\"hello\"); }"
14- </code>
15 </pre>
16 <table>
17+ <colgroup>
18+ <col>
19+ <col>
20+ </colgroup>
21 <thead>
22 <tr>
23 <th>
24 "Name"
25 </th>
26 <th>
27 "Score"
28 </th>
29 </tr>
30 </thead>
31 <tbody>
32 <tr>
33 <td>
34 "alpha"
35 </td>
36 <td>
37 "10"
38 </td>
39 </tr>
40 <tr>
41 <td>
42 "beta"
43 </td>
44 <td>
45 "20"
46 </td>
47 </tr>
48 </tbody>
49 </table>
50 <h1>
51 "Links and footnotes"
52 </h1>
53 <p>
54 "An external link:"
55 <a href="https://example.com">
56 "Example"
57 </a>
58- "and an id link"
59- <a href="#abc-123">
60- "abc-123"
61- </a>
62- "."
63+ "and an id link ."
64 </p>
65 <p>
66 "Text with a footnote reference."
67 <sup>
68- <a href="#fn-1">
69+ <a href="#fn.1">
70 "1"
71 </a>
72 </sup>
73 </p>
74 <h1>
75 "Blocks"
76 </h1>
77 <blockquote>
78 <p>
79 "A quoted paragraph."
80 </p>
81 </blockquote>
82 <hr>
83- <section>
84- <hr>
85- <ol>
86- <li>
87+ <h2>
88+ "Footnotes:"
89+ </h2>
90+ <sup>
91+ <a href="#fnr.1">
92+ "1"
93+ </a>
94+ </sup>
95 <p>
96 "The footnote definition."
97 </p>
98- <a href="#fnr-1">
99- "↩"
100- </a>
101- </li>
102- </ol>
103- </section>
tests/snapshots/oracle__oracle_footnote.snap added +83
@@ -0,0 +1,83 @@
1---
2source: tests/oracle.rs
3expression: report
4---
5agreement: 30/53 skeleton lines (56.6%)
6(- org-ssg, + emacs)
7
8 <p>
9 "Text with a reference."
10 <sup>
11- <a href="#fn-1">
12+ <a href="#fn.1">
13 "1"
14 </a>
15 </sup>
16 "And a second one."
17 <sup>
18- <a href="#fn-2">
19+ <a href="#fn.2">
20 "2"
21 </a>
22 </sup>
23 </p>
24 <p>
25 "An inline footnote."
26 <sup>
27- <a href="#fn-3">
28+ <a href="#fn.3">
29 "3"
30 </a>
31 </sup>
32 </p>
33- <section>
34- <hr>
35- <ol>
36- <li>
37+ <h2>
38+ "Footnotes:"
39+ </h2>
40+ <sup>
41+ <a href="#fnr.1">
42+ "1"
43+ </a>
44+ </sup>
45 <p>
46 "The first definition."
47 </p>
48- <a href="#fnr-1">
49- "↩"
50+ <sup>
51+ <a href="#fnr.2">
52+ "2"
53 </a>
54- </li>
55- <li>
56+ </sup>
57 <p>
58 "The second definition, with"
59- <em>
60+ <i>
61 "emphasis"
62- </em>
63+ </i>
64 "."
65 </p>
66- <a href="#fnr-2">
67- "↩"
68+ <sup>
69+ <a href="#fnr.3">
70+ "3"
71 </a>
72- </li>
73- <li>
74+ </sup>
75+ <p>
76 "defined right here"
77- <a href="#fnr-3">
78- "↩"
79- </a>
80- </li>
81- </ol>
82- </section>
83+ </p>
tests/snapshots/oracle__oracle_headings.snap added +39
@@ -0,0 +1,39 @@
1---
2source: tests/oracle.rs
3expression: report
4---
5agreement: 28/30 skeleton lines (93.3%)
6(- org-ssg, + emacs)
7
8 <h1>
9- "TODO [#A] Write the parser work rust"
10+ "TODO Write the parser work rust"
11 </h1>
12 <p>
13 "A heading carrying a keyword, a priority, tags and a property drawer."
14 </p>
15 <h2>
16 "DONE Nested and finished"
17 </h2>
18 <p>
19 "Sub-headings nest by star count."
20 </p>
21 <h2>
22- "[#C] Priority without a keyword"
23+ "Priority without a keyword"
24 </h2>
25 <p>
26 "A priority cookie can stand alone."
27 </p>
28 <h1>
29 "TODOs are not a keyword"
30 </h1>
31 <p>
32 "The word boundary matters: this heading has no TODO keyword."
33 </p>
34 <h1>
35 "DONE"
36 </h1>
37 <p>
38 "A keyword with no title at all."
39 </p>
tests/snapshots/oracle__oracle_images.snap added +63
@@ -0,0 +1,63 @@
1---
2source: tests/oracle.rs
3expression: report
4---
5agreement: 28/42 skeleton lines (66.7%)
6(- org-ssg, + emacs)
7
8 <h1>
9 "Bare image"
10 </h1>
11 <p>
12 <img src="diagram.png">
13 </p>
14 <h1>
15 "Captioned figure"
16 </h1>
17- <figure>
18+ <p>
19 <img src="pipeline.svg">
20- <figcaption>
21- "The pipeline, end to end"
22- </figcaption>
23- </figure>
24+ </p>
25+ <p>
26+ "Figure 1: The pipeline, end to end"
27+ </p>
28 <h1>
29 "Caption with markup"
30 </h1>
31- <figure>
32+ <p>
33 <img src="chart.png">
34- <figcaption>
35- "A"
36- <em>
37+ </p>
38+ <p>
39+ "Figure 2: A"
40+ <i>
41 "stylised"
42- </em>
43+ </i>
44 "chart"
45- </figcaption>
46- </figure>
47+ </p>
48 <h1>
49 "Quoted attribute values"
50 </h1>
51- <figure>
52+ <p>
53 <img src="cat.jpg">
54- </figure>
55+ </p>
56 <h1>
57 "Image with a description is a link"
58 </h1>
59 <p>
60 <a href="diagram.png">
61 "the diagram"
62 </a>
63 </p>
tests/snapshots/oracle__oracle_lists.snap added +123
@@ -0,0 +1,123 @@
1---
2source: tests/oracle.rs
3expression: report
4---
5agreement: 100/111 skeleton lines (90.1%)
6(- org-ssg, + emacs)
7
8 <h1>
9 "Nesting"
10 </h1>
11 <ul>
12 <li>
13 "outer item"
14 <ul>
15 <li>
16 "inner item"
17 <ul>
18 <li>
19 "deepest item"
20 </li>
21 </ul>
22 </li>
23 <li>
24 "second inner"
25 </li>
26 </ul>
27 </li>
28 <li>
29 "second outer"
30 </li>
31 </ul>
32 <h1>
33 "Ordered"
34 </h1>
35 <ol>
36 <li>
37 "first"
38 </li>
39 <li>
40 "second"
41 <ol>
42 <li>
43 "second point one"
44 </li>
45 <li>
46 "second point two"
47 </li>
48 </ol>
49 </li>
50 <li>
51 "third"
52 </li>
53 </ol>
54 <h1>
55 "Checkboxes"
56 </h1>
57 <ul>
58 <li>
59- <input>
60+ <code>
61+ "[ ]"
62+ </code>
63 "not done"
64 </li>
65 <li>
66- <input>
67+ <code>
68+ "[X]"
69+ </code>
70 "done"
71 </li>
72 <li>
73- <input>
74+ <code>
75+ "[-]"
76+ </code>
77 "partially done"
78 </li>
79 </ul>
80 <h1>
81 "Description"
82 </h1>
83 <dl>
84 <dt>
85 "term one"
86 </dt>
87 <dd>
88 "the first definition"
89 </dd>
90 <dt>
91 "term two"
92 </dt>
93 <dd>
94 "the second definition, which is soft-wrapped across two lines"
95 </dd>
96 <dt>
97- <em>
98+ <i>
99 "marked up"
100- </em>
101+ </i>
102 "term"
103 </dt>
104 <dd>
105 "definitions hold inline markup"
106 </dd>
107 </dl>
108 <h1>
109 "Multi-paragraph items"
110 </h1>
111 <ul>
112 <li>
113 <p>
114 "an item whose body has two paragraphs"
115 </p>
116 <p>
117 "the second paragraph, indented under the bullet"
118 </p>
119 </li>
120 <li>
121 "a plain sibling"
122 </li>
123 </ul>
tests/snapshots/oracle__oracle_minimal.snap added +53
@@ -0,0 +1,53 @@
1---
2source: tests/oracle.rs
3expression: report
4---
5agreement: 38/42 skeleton lines (90.5%)
6(- org-ssg, + emacs)
7
8 <p>
9 "A single paragraph of preamble text before any heading."
10 </p>
11 <h1>
12 "First Heading"
13 </h1>
14 <p>
15 "Some body text with"
16- <strong>
17+ <b>
18 "bold"
19- </strong>
20+ </b>
21 ","
22- <em>
23+ <i>
24 "italic"
25- </em>
26+ </i>
27 ", and"
28 <code>
29 "verbatim"
30 </code>
31 "."
32 </p>
33 <h2>
34 "A Subheading tag1 tag2"
35 </h2>
36 <ul>
37 <li>
38 "an unordered item"
39 </li>
40 <li>
41 "another with a checkbox [ ]"
42 </li>
43 </ul>
44 <h1>
45 "Second Heading"
46 </h1>
47 <p>
48 "See"
49 <a href="#first">
50 "the first heading"
51 </a>
52 "."
53 </p>
tests/snapshots/oracle__oracle_table.snap added +41
@@ -0,0 +1,41 @@
1---
2source: tests/oracle.rs
3expression: report
4---
5agreement: 30/34 skeleton lines (88.2%)
6(- org-ssg, + emacs)
7
8 <table>
9+ <colgroup>
10+ <col>
11+ <col>
12+ </colgroup>
13 <thead>
14 <tr>
15 <th>
16 "Name"
17 </th>
18 <th>
19 "Score"
20 </th>
21 </tr>
22 </thead>
23 <tbody>
24 <tr>
25 <td>
26 "alpha"
27 </td>
28 <td>
29 "10"
30 </td>
31 </tr>
32 <tr>
33 <td>
34 "beta"
35 </td>
36 <td>
37 "20"
38 </td>
39 </tr>
40 </tbody>
41 </table>
tests/snapshots/oracle__oracle_timestamps.snap added +74
@@ -0,0 +1,74 @@
1---
2source: tests/oracle.rs
3expression: report
4---
5agreement: 25/62 skeleton lines (40.3%)
6(- org-ssg, + emacs)
7
8 <h1>
9 "Single"
10 </h1>
11 <p>
12- "An active date"
13- <time>
14- "2024-01-15"
15- </time>
16- "and an inactive one"
17- <time>
18- "2024-01-15"
19- </time>
20- "."
21+ "An active date <2024-01-15 Mon> and an inactive one [2024-01-15 Mon]."
22 </p>
23 <p>
24- "With a time:"
25- <time>
26- "2024-01-15 10:30"
27- </time>
28- "."
29+ "With a time: <2024-01-15 Mon 10:30>."
30 </p>
31 <h1>
32 "Ranges"
33 </h1>
34 <p>
35- "A same-day time range"
36- <time>
37- "2024-01-15 10:00"
38- </time>
39- "–"
40- <time>
41- "11:45"
42- </time>
43- "."
44+ "A same-day time range <2024-01-15 Mon 10:00-11:45>."
45 </p>
46 <p>
47- "A multi-day range"
48- <time>
49- "2024-01-15"
50- </time>
51- "–"
52- <time>
53- "2024-01-20"
54- </time>
55- "."
56+ "A multi-day range <2024-01-15 Mon>–<2024-01-20 Sat>."
57 </p>
58 <h1>
59 "Ignored decorations"
60 </h1>
61 <p>
62- "A repeater is dropped:"
63- <time>
64- "2024-01-15"
65- </time>
66- "."
67+ "A repeater is dropped: <2024-01-15 Mon +1w>."
68 </p>
69 <h1>
70 "Not timestamps"
71 </h1>
72 <p>
73 "Comparisons like 3 < 4 and [not a stamp] stay literal text."
74 </p>