Commit ba4152c75e
Verified · cmc
Layout: unified · split
README.md +82 −8
| @@ -25,6 +25,7 @@ is the only inherently global stage — it is where the link dependency graph is | ||
| 25 | 25 | | Stage | Module | Notes | |
| 26 | 26 | |---|---|---| |
| 27 | 27 | | PARSE | `src/parser.rs` | Hand-written recursive descent: line lexer → element builder → inline tokenizer. | |
| 28 | | audit | `src/audit.rs` | Phase 0 corpus audit: construct frequencies against the IN/OUT line. | | |
| 28 | 29 | | model | `src/model.rs` | The org element tree — Elements (block) vs Objects (inline). | |
| 29 | 30 | | INDEX | `src/index.rs` | Collect link targets into a symbol table. | |
| 30 | 31 | | RESOLVE | `src/resolve.rs` | Rewrite links to URLs; return the used-target list (dependency edges). | |
| @@ -50,8 +51,9 @@ semantics; non-HTML export blocks; the full Unicode entity set. | ||
| 50 | 51 | **Scope guardrail:** every IN item gets a golden-file fixture; every OUT item gets a test |
| 51 | 52 | asserting it degrades predictably (ignored, no crash). The IN/OUT line is enforced by |
| 52 | 53 | `tests/constructs.rs`, defending against the project's #1 risk: scope creep back toward |
| 53 | all-of-org. The fixtures are hand-written today; deriving them from a real corpus is | |
| 54 | Phase 0. | |
| 54 | all-of-org. Phase 0 checked this line against a real 179-file corpus and found it sound | |
| 55 | (99.9% of construct uses in scope) — but also found one thing missing from it entirely: | |
| 56 | `#+SLUG:`. See [Phase 0](#phase-0-the-corpus-audit-and-the-emacs-oracle). | |
| 55 | 57 | |
| 56 | 58 | ## Phase plan |
| 57 | 59 | |
| @@ -62,7 +64,7 @@ Phase 0. | ||
| 62 | 64 | | **v0.2** | **Multi-file SITE build: INDEX + RESOLVE internal links, minijinja templates, `build <src-dir> <out-dir>`, tables + footnotes** | **done** | |
| 63 | 65 | | **v0.3** | **Incremental build layer: content/config/template hashing, dependency graph, per-page render keys, persisted cache manifest, invalidation** | **done** | |
| 64 | 66 | | **v0.4** | **MVP: the full v1 construct scope — heading metadata, nested/description lists, block types, timestamps, images, syntect highlighting — with the IN/OUT line under test** | **done** | |
| 65 | | 0 | Corpus audit + `emacs --batch` ground-truth oracle | todo | | |
| 67 | | **0** | **Corpus audit + `emacs --batch` ground-truth oracle** | **done** | | |
| 66 | 68 | | 1 | Line lexer + heading/section skeleton | done | |
| 67 | 69 | | 2 | Block elements — lists, source blocks, tables, footnote defs, blocks by type, drawers | done | |
| 68 | 70 | | 3 | Inline objects — emphasis, links, bare URLs, footnote refs, timestamps | done | |
| @@ -162,11 +164,82 @@ excluded construct to a specific degradation: babel is never executed *and* a ch | ||
| 162 | 164 | as literal text; drawers other than PROPERTIES are captured and dropped; unmodelled block |
| 163 | 165 | types keep their content verbatim. |
| 164 | 166 | |
| 165 | **Still out at v0.4:** the Phase 0 corpus audit and `emacs --batch` oracle (the fixtures are | |
| 166 | hand-written, so "matches Emacs" is asserted by construction, not measured); rayon | |
| 167 | parallelism; parse errors carrying source locations; `#+TODO:` per-file keyword sequences; | |
| 168 | planning lines (`SCHEDULED:`/`DEADLINE:`), which render as ordinary paragraphs; fixed-width | |
| 169 | `: ` lines; and the `watch` fs-notify integration. | |
| 167 | **Still out at v0.4:** rayon parallelism; parse errors carrying source locations; `#+TODO:` | |
| 168 | per-file keyword sequences; planning lines (`SCHEDULED:`/`DEADLINE:`), which render as | |
| 169 | ordinary paragraphs; fixed-width `: ` lines; and the `watch` fs-notify integration. | |
| 170 | ||
| 171 | ## Phase 0: the corpus audit and the Emacs oracle | |
| 172 | ||
| 173 | The v1 scope was, by its own admission, *recommended* — a guess about which slice of org | |
| 174 | matters. Phase 0 replaces both halves of that guess with a measurement: an audit that asks | |
| 175 | what a real corpus actually uses, and an oracle that asks whether we render it the way | |
| 176 | Emacs does. The corpus is the 179 files behind [cleberg.net](https://cleberg.net), which is | |
| 177 | published today by weblorg — a wrapper around org's own HTML exporter. That makes it both | |
| 178 | the workload and the incumbent. | |
| 179 | ||
| 180 | ``` | |
| 181 | cargo run -- audit <src-dir> # what does this corpus use, and is it in scope? | |
| 182 | cargo test --test oracle # how does our HTML differ from Emacs' own export? | |
| 183 | ``` | |
| 184 | ||
| 185 | ### What the audit found | |
| 186 | ||
| 187 | **The scope guess was sound.** 99.9% of construct uses in the corpus are in scope. The | |
| 188 | whole out-of-scope tail is 8 uses: four `#+TBLFM:` in a post *about* org-mode, three | |
| 189 | `\name` entities, and one `#+BEGIN_NOTE`. | |
| 190 | ||
| 191 | **`#+SLUG:` was a hole big enough to sink the project.** 178 of 179 files set it, and the | |
| 192 | published URL comes from it, not from the filename: `2018-11-28-aes-encryption.org` is | |
| 193 | served at `blog/aes-encryption.html`. org-ssg derived output paths from source filenames, | |
| 194 | so **169 of 179 pages would have been published at the wrong URL** — every inbound link and | |
| 195 | every search result, broken, by a tool that reported a clean build. Output paths now come | |
| 196 | from `#+SLUG:` when present ([`util::output_path`](src/util.rs)); slugs are sanitized so an | |
| 197 | author-supplied `../../etc/x` cannot escape the output directory, and two pages claiming one | |
| 198 | URL is a build error rather than a silently dropped page. Building the real corpus now | |
| 199 | reproduces all 179 of the live site's URLs exactly. | |
| 200 | ||
| 201 | **Some machinery is speculative.** The corpus contains no `id:`, `#custom-id` or `*Heading` | |
| 202 | links at all — its cross-page links are hand-written relative URLs. The INDEX/RESOLVE | |
| 203 | symbol table that v0.2 was built around is, against this corpus, unexercised. | |
| 204 | ||
| 205 | **An audit can lie too.** The first run reported 23 uses of a custom TODO keyword sequence. | |
| 206 | All 23 were false: the detector read the leading word of `* CSS Variables` as the keyword | |
| 207 | `CSS`. The corpus defines no `#+TODO:` sequences at all, so the true count was zero. The | |
| 208 | detector now matches conventional keyword names only — a tool that overstates a gap argues | |
| 209 | for work nobody needs. | |
| 210 | ||
| 211 | ### What the oracle found | |
| 212 | ||
| 213 | `tests/oracle.rs` exports each fixture with org's own exporter via `emacs --batch`, reduces | |
| 214 | both sides to a semantic skeleton (element opens, closes and text, with layout `div`s, | |
| 215 | inline `span`s and all attributes but `href`/`src` dropped), and **snapshots the | |
| 216 | disagreement**. Snapshotting rather than asserting is deliberate: a checked-in divergence | |
| 217 | report gets reviewed and shows up as a diff, where a permanently red test gets ignored. | |
| 218 | Three invariants are asserted outright, and all three hold — heading structure, list | |
| 219 | nesting, and source-block text match Emacs exactly. | |
| 220 | ||
| 221 | **No bugs in org-ssg.** Every remaining divergence is a deliberate choice to emit better | |
| 222 | HTML than org does: | |
| 223 | ||
| 224 | | | org-ssg | Emacs | why | | |
| 225 | |---|---|---|---| | |
| 226 | | emphasis | `<em>`/`<strong>` | `<i>`/`<b>` | semantic, not presentational | | |
| 227 | | captioned image | `<figure>`/`<figcaption>` | `<p>` + `"Figure 1: …"` | real figure semantics | | |
| 228 | | timestamp | `<time datetime="…">` | literal `<2024-01-15 Mon>` | machine-readable | | |
| 229 | | footnotes | `<section><ol>` | `<h2>Footnotes:</h2>` | a list of notes is a list | | |
| 230 | | heading anchor | slug of the text | `org1a2b3c4` | stable, and what the live site serves | | |
| 231 | | code | `<pre><code>` | `<pre>` | the HTML5 idiom | | |
| 232 | ||
| 233 | One genuine semantic difference: org treats a single blank line between a `1.` list and a | |
| 234 | `-` list as *one* list and keeps the first item's bullet type, while we start a second list. | |
| 235 | We keep ours, on measurement rather than taste — the pattern occurs **zero** times in the | |
| 236 | corpus, so matching an org quirk would buy nothing and cost the more obvious reading. | |
| 237 | ||
| 238 | **The oracle's best catch was three bugs in itself.** Naive normalization reported code as | |
| 239 | corrupted (it trimmed each of syntect's per-token text runs, turning `def greet` into | |
| 240 | `defgreet`) and reported blocks at 36% agreement (syntect's spans flooded the diff). Both | |
| 241 | were measurement artifacts. A differential harness is a piece of software like any other, | |
| 242 | and the first divergences it reports are usually its own. | |
| 170 | 243 | |
| 171 | 244 | **From v0.1 (core subset):** headings with nesting and anchors (every heading is now |
| 172 | 245 | anchored — `:CUSTOM_ID:`/`:ID:` else a slug of its text) and trailing tags; paragraphs; |
| @@ -188,6 +261,7 @@ cargo build | ||
| 188 | 261 | cargo test |
| 189 | 262 | cargo run -- build fixtures/minimal.org -o minimal.html # single file |
| 190 | 263 | cargo run -- build fixtures/site -o _site # whole site (incremental) |
| 264 | cargo run -- audit fixtures/site # corpus audit (Phase 0) | |
| 191 | 265 | cargo run -- build fixtures/site -o _site --no-cache # force a full rebuild |
| 192 | 266 | cargo run -- watch fixtures/site -o _site # poll + rebuild on change |
| 193 | 267 | cargo run -- clean _site # remove output + cache |
fixtures/slugsite/2024-02-11-long-source-name.org added +11
| @@ -0,0 +1,11 @@ | ||
| 1 | #+TITLE: The Post | |
| 2 | #+SLUG: short-url | |
| 3 | ||
| 4 | The source filename carries a date; the published URL does not. | |
| 5 | ||
| 6 | * Setup | |
| 7 | :PROPERTIES: | |
| 8 | :CUSTOM_ID: setup | |
| 9 | :END: | |
| 10 | ||
| 11 | Linking into this heading must land on the slugged page, not the source name. | |
fixtures/slugsite/index.org added +3
| @@ -0,0 +1,3 @@ | ||
| 1 | #+TITLE: Home | |
| 2 | ||
| 3 | Read [[file:2024-02-11-long-source-name.org][the post]], or jump to its [[#setup][setup section]]. | |
src/audit.rs added +620
| @@ -0,0 +1,620 @@ | ||
| 1 | //! Corpus audit (spec §5, Phase 0): measure which org constructs a real corpus actually | |
| 2 | //! uses, and classify each against the v1 IN/OUT line. | |
| 3 | //! | |
| 4 | //! This exists because the v1 scope was, on the README's own admission, *recommended* | |
| 5 | //! rather than measured — a guess about which slice of org matters. A guess about a | |
| 6 | //! corpus is a hypothesis, and this is the experiment. It answers two questions: | |
| 7 | //! | |
| 8 | //! 1. **Coverage** — of the constructs this corpus uses, which do we handle? A construct | |
| 9 | //! that is common here and out of scope is a scope bug, not a corpus quirk. | |
| 10 | //! 2. **Blind spots** — which constructs are here that the implementation has no opinion | |
| 11 | //! about at all? These are the dangerous ones: not "known unsupported" but unknown. | |
| 12 | //! | |
| 13 | //! The audit is deliberately a *separate, line-oriented scanner* rather than a reuse of | |
| 14 | //! [`crate::parser`]. Auditing with the parser could only ever find constructs the parser | |
| 15 | //! already knows about, which is precisely the wrong instrument for question 2 — it would | |
| 16 | //! report a blind spot as clean. | |
| 17 | //! | |
| 18 | //! Nothing here reports document *text*. Counts, construct names, and `file:line` | |
| 19 | //! locations only, so an audit of private notes stays publishable. | |
| 20 | ||
| 21 | use std::collections::BTreeMap; | |
| 22 | ||
| 23 | use anyhow::{Context, Result}; | |
| 24 | use camino::{Utf8Path, Utf8PathBuf}; | |
| 25 | use walkdir::WalkDir; | |
| 26 | ||
| 27 | /// Where a construct sits relative to the v1 scope line (README §"v1 scope"). | |
| 28 | #[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord)] | |
| 29 | pub enum Scope { | |
| 30 | /// v1 handles this. | |
| 31 | In, | |
| 32 | /// v1 deliberately excludes this; it degrades predictably. | |
| 33 | Out, | |
| 34 | } | |
| 35 | ||
| 36 | impl Scope { | |
| 37 | fn label(self) -> &'static str { | |
| 38 | match self { | |
| 39 | Scope::In => "IN ", | |
| 40 | Scope::Out => "OUT", | |
| 41 | } | |
| 42 | } | |
| 43 | } | |
| 44 | ||
| 45 | /// One construct's tally across the corpus. | |
| 46 | #[derive(Debug, Default, Clone)] | |
| 47 | pub struct Tally { | |
| 48 | pub occurrences: usize, | |
| 49 | pub files: usize, | |
| 50 | /// First `file:line` the construct was seen at, to make a finding actionable. | |
| 51 | pub first_seen: Option<String>, | |
| 52 | /// Set while scanning one file, to count each file once. | |
| 53 | seen_in_current_file: bool, | |
| 54 | } | |
| 55 | ||
| 56 | /// The audit result: the fixed construct catalog plus the dynamic name censuses. | |
| 57 | #[derive(Debug, Default)] | |
| 58 | pub struct Audit { | |
| 59 | pub files: usize, | |
| 60 | pub lines: usize, | |
| 61 | /// Catalogued constructs → tally. | |
| 62 | pub constructs: BTreeMap<(Scope, &'static str), Tally>, | |
| 63 | /// Every distinct `#+KEYWORD:` seen, by name. | |
| 64 | pub keywords: BTreeMap<String, Tally>, | |
| 65 | /// Every distinct `#+BEGIN_<TYPE>` seen, by type. | |
| 66 | pub blocks: BTreeMap<String, Tally>, | |
| 67 | /// Every distinct `:DRAWER:` seen, by name. | |
| 68 | pub drawers: BTreeMap<String, Tally>, | |
| 69 | /// Every distinct link scheme seen (`https`, `file`, `id`, `denote`, ...). | |
| 70 | pub link_schemes: BTreeMap<String, Tally>, | |
| 71 | } | |
| 72 | ||
| 73 | /// Names the implementation understands, so the census can flag everything else. These | |
| 74 | /// are the *recognized* sets, not the supported ones: `INCLUDE` is recognized (it is | |
| 75 | /// deliberately inert) while an unlisted keyword is a genuine blind spot. | |
| 76 | const KNOWN_KEYWORDS: &[&str] = &[ | |
| 77 | "TITLE", "AUTHOR", "DATE", "EMAIL", "LANGUAGE", "OPTIONS", "FILETAGS", "DESCRIPTION", | |
| 78 | "KEYWORDS", "CAPTION", "NAME", "ATTR_HTML", "RESULTS", "TBLFM", "INCLUDE", "TODO", | |
| 79 | "STARTUP", "SUBTITLE", "SETUPFILE", "MACRO", "PROPERTY", "HTML_HEAD", "EXCLUDE_TAGS", | |
| 80 | ]; | |
| 81 | const KNOWN_BLOCKS: &[&str] = &["SRC", "QUOTE", "EXAMPLE", "CENTER", "EXPORT"]; | |
| 82 | const KNOWN_DRAWERS: &[&str] = &["PROPERTIES", "LOGBOOK", "END"]; | |
| 83 | /// Keyword names conventional enough to be worth flagging when they lead a heading. | |
| 84 | /// A custom sequence is only *real* if some `#+TODO:` declares it, which the census | |
| 85 | /// reports separately — this list keeps the heading-level signal honest. | |
| 86 | const CONVENTIONAL_TODO_KEYWORDS: &[&str] = &[ | |
| 87 | "NEXT", "WAITING", "HOLD", "CANCELLED", "CANCELED", "STARTED", "SOMEDAY", "PROJ", | |
| 88 | "IN-PROGRESS", "BLOCKED", "REVIEW", | |
| 89 | ]; | |
| 90 | const KNOWN_SCHEMES: &[&str] = &[ | |
| 91 | "http", "https", "mailto", "ftp", "news", "tel", "file", "id", "custom-id", "heading", | |
| 92 | "relative", | |
| 93 | ]; | |
| 94 | ||
| 95 | impl Audit { | |
| 96 | /// Is this name one the implementation recognizes? | |
| 97 | pub fn is_known(kind: Census, name: &str) -> bool { | |
| 98 | let known = match kind { | |
| 99 | Census::Keyword => KNOWN_KEYWORDS, | |
| 100 | Census::Block => KNOWN_BLOCKS, | |
| 101 | Census::Drawer => KNOWN_DRAWERS, | |
| 102 | Census::Scheme => KNOWN_SCHEMES, | |
| 103 | }; | |
| 104 | known.iter().any(|k| k.eq_ignore_ascii_case(name)) | |
| 105 | } | |
| 106 | } | |
| 107 | ||
| 108 | /// Which dynamic census a name belongs to. | |
| 109 | #[derive(Debug, Clone, Copy)] | |
| 110 | pub enum Census { | |
| 111 | Keyword, | |
| 112 | Block, | |
| 113 | Drawer, | |
| 114 | Scheme, | |
| 115 | } | |
| 116 | ||
| 117 | /// Walk `root`, auditing every `.org` file. | |
| 118 | pub fn audit(root: &Utf8Path) -> Result<Audit> { | |
| 119 | let mut audit = Audit::default(); | |
| 120 | let mut paths: Vec<Utf8PathBuf> = Vec::new(); | |
| 121 | ||
| 122 | if root.is_file() { | |
| 123 | paths.push(root.to_owned()); | |
| 124 | } else { | |
| 125 | for entry in WalkDir::new(root).sort_by_file_name() { | |
| 126 | let entry = entry.with_context(|| format!("walking {root}"))?; | |
| 127 | if !entry.file_type().is_file() { | |
| 128 | continue; | |
| 129 | } | |
| 130 | let path = Utf8PathBuf::from_path_buf(entry.into_path()) | |
| 131 | .map_err(|p| anyhow::anyhow!("non-UTF-8 path: {}", p.display()))?; | |
| 132 | if path.extension() == Some("org") { | |
| 133 | paths.push(path); | |
| 134 | } | |
| 135 | } | |
| 136 | } | |
| 137 | ||
| 138 | for path in &paths { | |
| 139 | // A file that cannot be read is reported and skipped: an audit of 179 files | |
| 140 | // should not be lost to one unreadable one. | |
| 141 | let source = match std::fs::read_to_string(path) { | |
| 142 | Ok(s) => s, | |
| 143 | Err(e) => { | |
| 144 | eprintln!("warning: skipping {path}: {e}"); | |
| 145 | continue; | |
| 146 | } | |
| 147 | }; | |
| 148 | let rel = path.strip_prefix(root).unwrap_or(path).to_owned(); | |
| 149 | audit.scan_file(&rel, &source); | |
| 150 | audit.files += 1; | |
| 151 | } | |
| 152 | Ok(audit) | |
| 153 | } | |
| 154 | ||
| 155 | impl Audit { | |
| 156 | fn scan_file(&mut self, path: &Utf8Path, source: &str) { | |
| 157 | // Reset the per-file flags so each construct counts this file at most once. | |
| 158 | for tally in self.constructs.values_mut() { | |
| 159 | tally.seen_in_current_file = false; | |
| 160 | } | |
| 161 | for map in [ | |
| 162 | &mut self.keywords, | |
| 163 | &mut self.blocks, | |
| 164 | &mut self.drawers, | |
| 165 | &mut self.link_schemes, | |
| 166 | ] { | |
| 167 | for tally in map.values_mut() { | |
| 168 | tally.seen_in_current_file = false; | |
| 169 | } | |
| 170 | } | |
| 171 | ||
| 172 | let mut in_block: Option<String> = None; | |
| 173 | for (idx, line) in source.lines().enumerate() { | |
| 174 | self.lines += 1; | |
| 175 | let at = format!("{path}:{}", idx + 1); | |
| 176 | let trimmed = line.trim_start(); | |
| 177 | ||
| 178 | // Inside a verbatim block only the terminator matters — a `*` in a source | |
| 179 | // block is not a heading, and counting it as one would corrupt the audit. | |
| 180 | if let Some(kind) = &in_block { | |
| 181 | if trimmed.to_ascii_uppercase().starts_with("#+END_") { | |
| 182 | in_block = None; | |
| 183 | } else if kind.eq_ignore_ascii_case("SRC") || kind.eq_ignore_ascii_case("EXAMPLE") { | |
| 184 | continue; | |
| 185 | } | |
| 186 | continue; | |
| 187 | } | |
| 188 | if let Some(rest) = trimmed.to_ascii_uppercase().strip_prefix("#+BEGIN_") { | |
| 189 | let kind = rest.split_whitespace().next().unwrap_or("").to_string(); | |
| 190 | self.count_census(Census::Block, &kind, &at); | |
| 191 | self.count(scope_of_block(&kind), block_construct(&kind), &at); | |
| 192 | if trimmed.to_ascii_uppercase().contains(":RESULTS") { | |
| 193 | self.count(Scope::Out, "babel header args (:results)", &at); | |
| 194 | } | |
| 195 | in_block = Some(kind); | |
| 196 | continue; | |
| 197 | } | |
| 198 | ||
| 199 | self.scan_line(line, trimmed, &at); | |
| 200 | } | |
| 201 | } | |
| 202 | ||
| 203 | fn scan_line(&mut self, line: &str, trimmed: &str, at: &str) { | |
| 204 | // --- headings and their metadata --- | |
| 205 | if let Some(stars) = heading_stars(line) { | |
| 206 | self.count(Scope::In, "heading", at); | |
| 207 | let rest = line[stars..].trim(); | |
| 208 | let word = rest.split_whitespace().next().unwrap_or(""); | |
| 209 | if word == "TODO" || word == "DONE" { | |
| 210 | self.count(Scope::In, "TODO keyword (default set)", at); | |
| 211 | } else if CONVENTIONAL_TODO_KEYWORDS.contains(&word) { | |
| 212 | // Only conventional keyword names count. "Any all-caps first word" is | |
| 213 | // the tempting rule and it is wrong: it reads `* CSS Variables` as the | |
| 214 | // keyword `CSS`, which on this corpus produced 23 false positives and | |
| 215 | // zero true ones. An audit that overstates a gap is worse than no audit, | |
| 216 | // because it argues for work nobody needs. | |
| 217 | self.count(Scope::Out, "TODO keyword (custom sequence)", at); | |
| 218 | } | |
| 219 | if rest.contains("[#") { | |
| 220 | self.count(Scope::In, "priority cookie", at); | |
| 221 | } | |
| 222 | if rest.trim_end().ends_with(':') && rest.trim_end().matches(':').count() >= 2 { | |
| 223 | self.count(Scope::In, "heading tags", at); | |
| 224 | } | |
| 225 | if rest.contains("[/") || rest.contains("[%") { | |
| 226 | self.count(Scope::Out, "statistics cookie", at); | |
| 227 | } | |
| 228 | return; | |
| 229 | } | |
| 230 | ||
| 231 | // --- planning and clocking --- | |
| 232 | for marker in ["SCHEDULED:", "DEADLINE:", "CLOSED:"] { | |
| 233 | if trimmed.starts_with(marker) { | |
| 234 | self.count(Scope::Out, "planning line", at); | |
| 235 | } | |
| 236 | } | |
| 237 | if trimmed.starts_with("CLOCK:") { | |
| 238 | self.count(Scope::Out, "clock entry", at); | |
| 239 | } | |
| 240 | ||
| 241 | // --- keywords and drawers --- | |
| 242 | if let Some(rest) = trimmed.strip_prefix("#+") { | |
| 243 | if let Some(colon) = rest.find(':') { | |
| 244 | let key = rest[..colon].trim().to_ascii_uppercase(); | |
| 245 | if !key.is_empty() && !key.contains(char::is_whitespace) { | |
| 246 | self.count_census(Census::Keyword, &key, at); | |
| 247 | match key.as_str() { | |
| 248 | "CAPTION" | "NAME" | "ATTR_HTML" => { | |
| 249 | self.count(Scope::In, "affiliated keyword", at) | |
| 250 | } | |
| 251 | "TBLFM" => self.count(Scope::Out, "table formula (#+TBLFM:)", at), | |
| 252 | "INCLUDE" => self.count(Scope::Out, "#+INCLUDE:", at), | |
| 253 | "RESULTS" => self.count(Scope::Out, "babel results block", at), | |
| 254 | "TODO" => self.count(Scope::Out, "#+TODO: keyword sequence", at), | |
| 255 | "MACRO" => self.count(Scope::Out, "macro definition", at), | |
| 256 | _ => self.count(Scope::In, "#+ keyword", at), | |
| 257 | } | |
| 258 | } | |
| 259 | } | |
| 260 | } else if is_drawer(trimmed) { | |
| 261 | let name = trimmed[1..trimmed.len() - 1].to_ascii_uppercase(); | |
| 262 | if name != "END" { | |
| 263 | self.count_census(Census::Drawer, &name, at); | |
| 264 | match name.as_str() { | |
| 265 | "PROPERTIES" => self.count(Scope::In, "property drawer", at), | |
| 266 | _ => self.count(Scope::Out, "non-PROPERTIES drawer", at), | |
| 267 | } | |
| 268 | } | |
| 269 | } | |
| 270 | ||
| 271 | // --- lists, tables, rules --- | |
| 272 | if let Some(bullet) = list_bullet(trimmed) { | |
| 273 | self.count(Scope::In, "list item", at); | |
| 274 | if bullet == Bullet::Ordered { | |
| 275 | self.count(Scope::In, "ordered list", at); | |
| 276 | } | |
| 277 | let indent = line.len() - trimmed.len(); | |
| 278 | if indent > 0 { | |
| 279 | self.count(Scope::In, "nested list item", at); | |
| 280 | } | |
| 281 | if trimmed.contains(" :: ") { | |
| 282 | self.count(Scope::In, "description list", at); | |
| 283 | } | |
| 284 | let after = trimmed.trim_start_matches(['-', '+', '*', ' ']); | |
| 285 | if after.starts_with("[ ]") || after.starts_with("[X]") || after.starts_with("[-]") { | |
| 286 | self.count(Scope::In, "checkbox", at); | |
| 287 | } | |
| 288 | } | |
| 289 | if trimmed.starts_with('|') { | |
| 290 | self.count(Scope::In, "table row", at); | |
| 291 | } | |
| 292 | if trimmed.starts_with(':') && !is_drawer(trimmed) && trimmed.starts_with(": ") { | |
| 293 | self.count(Scope::Out, "fixed-width line", at); | |
| 294 | } | |
| 295 | ||
| 296 | // --- footnotes --- | |
| 297 | if trimmed.starts_with("[fn:") { | |
| 298 | self.count(Scope::In, "footnote definition", at); | |
| 299 | } else if line.contains("[fn:") { | |
| 300 | self.count(Scope::In, "footnote reference", at); | |
| 301 | } | |
| 302 | ||
| 303 | // --- inline objects --- | |
| 304 | self.scan_inline(line, at); | |
| 305 | } | |
| 306 | ||
| 307 | fn scan_inline(&mut self, line: &str, at: &str) { | |
| 308 | // Links: count each `[[target]]`, censusing its scheme. | |
| 309 | let mut rest = line; | |
| 310 | while let Some(start) = rest.find("[[") { | |
| 311 | let after = &rest[start + 2..]; | |
| 312 | let Some(end) = after.find("]]") else { break }; | |
| 313 | let inner = &after[..end]; | |
| 314 | let target = inner.split("][").next().unwrap_or(inner); | |
| 315 | self.count(Scope::In, "link", at); | |
| 316 | self.count_census(Census::Scheme, &link_scheme(target), at); | |
| 317 | rest = &after[end..]; | |
| 318 | } | |
| 319 | ||
| 320 | if has_timestamp(line) { | |
| 321 | self.count(Scope::In, "timestamp", at); | |
| 322 | } | |
| 323 | if line.contains("{{{") { | |
| 324 | self.count(Scope::Out, "macro call", at); | |
| 325 | } | |
| 326 | if line.contains("<<<") { | |
| 327 | self.count(Scope::Out, "radio target", at); | |
| 328 | } else if line.contains("<<") && line.contains(">>") { | |
| 329 | self.count(Scope::Out, "internal target", at); | |
| 330 | } | |
| 331 | if line.contains("\\begin{") || latex_inline(line) { | |
| 332 | self.count(Scope::Out, "LaTeX fragment", at); | |
| 333 | } | |
| 334 | if entity_ref(line) { | |
| 335 | self.count(Scope::Out, "entity (\\name)", at); | |
| 336 | } | |
| 337 | for (marker, name) in [ | |
| 338 | ('*', "bold"), | |
| 339 | ('/', "italic"), | |
| 340 | ('_', "underline"), | |
| 341 | ('+', "strike-through"), | |
| 342 | ('=', "verbatim"), | |
| 343 | ('~', "code"), | |
| 344 | ] { | |
| 345 | if emphasis_pair(line, marker) { | |
| 346 | self.count(Scope::In, name, at); | |
| 347 | } | |
| 348 | } | |
| 349 | } | |
| 350 | ||
| 351 | fn count(&mut self, scope: Scope, name: &'static str, at: &str) { | |
| 352 | let tally = self.constructs.entry((scope, name)).or_default(); | |
| 353 | bump(tally, at); | |
| 354 | } | |
| 355 | ||
| 356 | fn count_census(&mut self, kind: Census, name: &str, at: &str) { | |
| 357 | let map = match kind { | |
| 358 | Census::Keyword => &mut self.keywords, | |
| 359 | Census::Block => &mut self.blocks, | |
| 360 | Census::Drawer => &mut self.drawers, | |
| 361 | Census::Scheme => &mut self.link_schemes, | |
| 362 | }; | |
| 363 | let tally = map.entry(name.to_string()).or_default(); | |
| 364 | bump(tally, at); | |
| 365 | } | |
| 366 | } | |
| 367 | ||
| 368 | fn bump(tally: &mut Tally, at: &str) { | |
| 369 | tally.occurrences += 1; | |
| 370 | if !tally.seen_in_current_file { | |
| 371 | tally.seen_in_current_file = true; | |
| 372 | tally.files += 1; | |
| 373 | } | |
| 374 | if tally.first_seen.is_none() { | |
| 375 | tally.first_seen = Some(at.to_string()); | |
| 376 | } | |
| 377 | } | |
| 378 | ||
| 379 | // --------------------------------------------------------------------------- | |
| 380 | // Line-level detectors. Deliberately independent of the parser (see module docs). | |
| 381 | // --------------------------------------------------------------------------- | |
| 382 | ||
| 383 | fn heading_stars(line: &str) -> Option<usize> { | |
| 384 | if !line.starts_with('*') { | |
| 385 | return None; | |
| 386 | } | |
| 387 | let stars = line.chars().take_while(|c| *c == '*').count(); | |
| 388 | let after = &line[stars..]; | |
| 389 | (after.starts_with(' ') || after.is_empty()).then_some(stars) | |
| 390 | } | |
| 391 | ||
| 392 | #[derive(PartialEq)] | |
| 393 | enum Bullet { | |
| 394 | Unordered, | |
| 395 | Ordered, | |
| 396 | } | |
| 397 | ||
| 398 | fn list_bullet(trimmed: &str) -> Option<Bullet> { | |
| 399 | let bytes = trimmed.as_bytes(); | |
| 400 | if bytes.is_empty() { | |
| 401 | return None; | |
| 402 | } | |
| 403 | if (bytes[0] == b'-' || bytes[0] == b'+') && (bytes.len() == 1 || bytes[1] == b' ') { | |
| 404 | return Some(Bullet::Unordered); | |
| 405 | } | |
| 406 | let digits = trimmed.chars().take_while(|c| c.is_ascii_digit()).count(); | |
| 407 | if digits > 0 { | |
| 408 | let after = &trimmed[digits..]; | |
| 409 | if (after.starts_with('.') || after.starts_with(')')) | |
| 410 | && (after.len() == 1 || after.as_bytes()[1] == b' ') | |
| 411 | { | |
| 412 | return Some(Bullet::Ordered); | |
| 413 | } | |
| 414 | } | |
| 415 | None | |
| 416 | } | |
| 417 | ||
| 418 | fn is_drawer(trimmed: &str) -> bool { | |
| 419 | let t = trimmed.trim_end(); | |
| 420 | t.len() >= 3 | |
| 421 | && t.starts_with(':') | |
| 422 | && t.ends_with(':') | |
| 423 | && t[1..t.len() - 1] | |
| 424 | .chars() | |
| 425 | .all(|c| c.is_ascii_alphanumeric() || c == '_' || c == '-') | |
| 426 | && t.len() > 2 | |
| 427 | } | |
| 428 | ||
| 429 | fn scope_of_block(kind: &str) -> Scope { | |
| 430 | if KNOWN_BLOCKS.iter().any(|k| k.eq_ignore_ascii_case(kind)) { | |
| 431 | Scope::In | |
| 432 | } else { | |
| 433 | Scope::Out | |
| 434 | } | |
| 435 | } | |
| 436 | ||
| 437 | fn block_construct(kind: &str) -> &'static str { | |
| 438 | match kind.to_ascii_uppercase().as_str() { | |
| 439 | "SRC" => "source block", | |
| 440 | "QUOTE" => "quote block", | |
| 441 | "EXAMPLE" => "example block", | |
| 442 | "CENTER" => "center block", | |
| 443 | "EXPORT" => "export block", | |
| 444 | _ => "unmodelled block type", | |
| 445 | } | |
| 446 | } | |
| 447 | ||
| 448 | /// The scheme of a link target, normalized into the census's vocabulary. | |
| 449 | fn link_scheme(target: &str) -> String { | |
| 450 | if let Some(rest) = target.split_once(':') { | |
| 451 | let scheme = rest.0; | |
| 452 | if !scheme.is_empty() | |
| 453 | && scheme | |
| 454 | .chars() | |
| 455 | .all(|c| c.is_ascii_alphanumeric() || c == '-' || c == '+') | |
| 456 | { | |
| 457 | return scheme.to_ascii_lowercase(); | |
| 458 | } | |
| 459 | } | |
| 460 | if target.starts_with('#') { | |
| 461 | return "custom-id".to_string(); | |
| 462 | } | |
| 463 | if target.starts_with('*') { | |
| 464 | return "heading".to_string(); | |
| 465 | } | |
| 466 | "relative".to_string() | |
| 467 | } | |
| 468 | ||
| 469 | /// A `<...>`/`[...]` span opening with an ISO date is a timestamp. | |
| 470 | fn has_timestamp(line: &str) -> bool { | |
| 471 | let bytes = line.as_bytes(); | |
| 472 | for (i, c) in line.char_indices() { | |
| 473 | if c != '<' && c != '[' { | |
| 474 | continue; | |
| 475 | } | |
| 476 | let rest = &bytes[i + 1..]; | |
| 477 | if rest.len() >= 10 | |
| 478 | && rest[..4].iter().all(u8::is_ascii_digit) | |
| 479 | && rest[4] == b'-' | |
| 480 | && rest[5..7].iter().all(u8::is_ascii_digit) | |
| 481 | && rest[7] == b'-' | |
| 482 | && rest[8..10].iter().all(u8::is_ascii_digit) | |
| 483 | { | |
| 484 | return true; | |
| 485 | } | |
| 486 | } | |
| 487 | false | |
| 488 | } | |
| 489 | ||
| 490 | /// `$x$` or `\(x\)` inline math. `$` alone (a price, a shell prompt) is not math. | |
| 491 | fn latex_inline(line: &str) -> bool { | |
| 492 | if line.contains("\\(") && line.contains("\\)") { | |
| 493 | return true; | |
| 494 | } | |
| 495 | let dollars = line.matches('$').count(); | |
| 496 | dollars >= 2 && line.contains("$\\") | |
| 497 | } | |
| 498 | ||
| 499 | /// A `\name` entity reference such as `\alpha`, excluding LaTeX environment commands. | |
| 500 | fn entity_ref(line: &str) -> bool { | |
| 501 | for (i, c) in line.char_indices() { | |
| 502 | if c != '\\' { | |
| 503 | continue; | |
| 504 | } | |
| 505 | let rest = &line[i + 1..]; | |
| 506 | let name: String = rest.chars().take_while(|c| c.is_ascii_alphabetic()).collect(); | |
| 507 | if name.len() >= 3 && !matches!(name.as_str(), "begin" | "end") { | |
| 508 | return true; | |
| 509 | } | |
| 510 | } | |
| 511 | false | |
| 512 | } | |
| 513 | ||
| 514 | /// A plausible `*bold*`-style emphasis pair: two markers on one line with non-space | |
| 515 | /// content between them. Approximate by design — the audit measures prevalence, and the | |
| 516 | /// parser owns the exact pre/post-character rules. | |
| 517 | fn emphasis_pair(line: &str, marker: char) -> bool { | |
| 518 | let positions: Vec<usize> = line | |
| 519 | .char_indices() | |
| 520 | .filter(|(_, c)| *c == marker) | |
| 521 | .map(|(i, _)| i) | |
| 522 | .collect(); | |
| 523 | if positions.len() < 2 { | |
| 524 | return false; | |
| 525 | } | |
| 526 | // A leading `*` is a heading, and `-`/`+` at line start is a bullet. | |
| 527 | let trimmed = line.trim_start(); | |
| 528 | if trimmed.starts_with(marker) { | |
| 529 | return false; | |
| 530 | } | |
| 531 | positions.windows(2).any(|w| w[1] > w[0] + 1) | |
| 532 | } | |
| 533 | ||
| 534 | // --------------------------------------------------------------------------- | |
| 535 | // Report | |
| 536 | // --------------------------------------------------------------------------- | |
| 537 | ||
| 538 | /// Render the audit as a readable report. Names, counts and locations only — never | |
| 539 | /// document text, so an audit of private notes is safe to paste into an issue. | |
| 540 | pub fn report(audit: &Audit) -> String { | |
| 541 | let mut out = String::new(); | |
| 542 | out.push_str(&format!( | |
| 543 | "corpus: {} file(s), {} line(s)\n", | |
| 544 | audit.files, audit.lines | |
| 545 | )); | |
| 546 | ||
| 547 | let mut rows: Vec<(&(Scope, &str), &Tally)> = audit.constructs.iter().collect(); | |
| 548 | rows.sort_by(|a, b| { | |
| 549 | b.1.occurrences | |
| 550 | .cmp(&a.1.occurrences) | |
| 551 | .then_with(|| a.0 .1.cmp(b.0 .1)) | |
| 552 | }); | |
| 553 | ||
| 554 | out.push_str("\nCONSTRUCTS (by frequency)\n"); | |
| 555 | out.push_str(&format!( | |
| 556 | "{:<4} {:<32} {:>8} {:>7} {}\n", | |
| 557 | "", "construct", "uses", "files", "first seen" | |
| 558 | )); | |
| 559 | for ((scope, name), tally) in &rows { | |
| 560 | out.push_str(&format!( | |
| 561 | "{:<4} {:<32} {:>8} {:>7} {}\n", | |
| 562 | scope.label(), | |
| 563 | name, | |
| 564 | tally.occurrences, | |
| 565 | tally.files, | |
| 566 | tally.first_seen.as_deref().unwrap_or("") | |
| 567 | )); | |
| 568 | } | |
| 569 | ||
| 570 | let in_uses: usize = rows | |
| 571 | .iter() | |
| 572 | .filter(|((s, _), _)| *s == Scope::In) | |
| 573 | .map(|(_, t)| t.occurrences) | |
| 574 | .sum(); | |
| 575 | let out_uses: usize = rows | |
| 576 | .iter() | |
| 577 | .filter(|((s, _), _)| *s == Scope::Out) | |
| 578 | .map(|(_, t)| t.occurrences) | |
| 579 | .sum(); | |
| 580 | let total = in_uses + out_uses; | |
| 581 | let pct = |n: usize| { | |
| 582 | if total == 0 { | |
| 583 | 0.0 | |
| 584 | } else { | |
| 585 | 100.0 * n as f64 / total as f64 | |
| 586 | } | |
| 587 | }; | |
| 588 | out.push_str(&format!( | |
| 589 | "\ncoverage: {in_uses} in-scope use(s) ({:.1}%), {out_uses} out-of-scope ({:.1}%)\n", | |
| 590 | pct(in_uses), | |
| 591 | pct(out_uses) | |
| 592 | )); | |
| 593 | ||
| 594 | for (title, kind, map) in [ | |
| 595 | ("KEYWORDS", Census::Keyword, &audit.keywords), | |
| 596 | ("BLOCK TYPES", Census::Block, &audit.blocks), | |
| 597 | ("DRAWERS", Census::Drawer, &audit.drawers), | |
| 598 | ("LINK SCHEMES", Census::Scheme, &audit.link_schemes), | |
| 599 | ] { | |
| 600 | let mut names: Vec<(&String, &Tally)> = map.iter().collect(); | |
| 601 | names.sort_by(|a, b| b.1.occurrences.cmp(&a.1.occurrences).then(a.0.cmp(b.0))); | |
| 602 | out.push_str(&format!("\n{title}\n")); | |
| 603 | for (name, tally) in names { | |
| 604 | let flag = if Audit::is_known(kind, name) { | |
| 605 | " " | |
| 606 | } else { | |
| 607 | "??? " | |
| 608 | }; | |
| 609 | out.push_str(&format!( | |
| 610 | "{flag}{:<32} {:>8} {:>7} {}\n", | |
| 611 | name, | |
| 612 | tally.occurrences, | |
| 613 | tally.files, | |
| 614 | tally.first_seen.as_deref().unwrap_or("") | |
| 615 | )); | |
| 616 | } | |
| 617 | } | |
| 618 | out.push_str("\n`???` marks a name the implementation does not recognize at all.\n"); | |
| 619 | out | |
| 620 | } | |
src/index.rs +23 −21
| @@ -7,7 +7,7 @@ use camino::{Utf8Path, Utf8PathBuf}; | ||
| 7 | 7 | use serde::{Deserialize, Serialize}; |
| 8 | 8 | |
| 9 | 9 | use crate::model::{Document, Section}; |
| 10 | use crate::util::{plain_text, slugify}; | |
| 10 | use crate::util::{output_path, plain_text, slugify}; | |
| 11 | 11 | |
| 12 | 12 | /// Identity of a link target. A target is owned by exactly one file (spec §4.3). |
| 13 | 13 | /// |
| @@ -51,6 +51,9 @@ impl TargetId { | ||
| 51 | 51 | #[derive(Debug, Clone)] |
| 52 | 52 | pub struct TargetLocation { |
| 53 | 53 | pub source_path: Utf8PathBuf, |
| 54 | /// The page this target is emitted into. Recorded at INDEX time because it depends | |
| 55 | /// on the defining document's `#+SLUG:`, which only that document knows. | |
| 56 | pub output_path: Utf8PathBuf, | |
| 54 | 57 | /// Final URL fragment/anchor for the target, filled during resolution. |
| 55 | 58 | pub anchor: Option<String>, |
| 56 | 59 | } |
| @@ -71,14 +74,16 @@ impl SymbolTable { | ||
| 71 | 74 | /// the renderer emits for that target's heading. |
| 72 | 75 | pub fn index_document(&mut self, doc: &Document) { |
| 73 | 76 | let path = &doc.source_path; |
| 77 | let out = output_path(path, &doc.keywords); | |
| 74 | 78 | self.targets.insert( |
| 75 | 79 | TargetId::File(path.clone()), |
| 76 | 80 | TargetLocation { |
| 77 | 81 | source_path: path.clone(), |
| 82 | output_path: out.clone(), | |
| 78 | 83 | anchor: None, |
| 79 | 84 | }, |
| 80 | 85 | ); |
| 81 | index_section(&doc.root, path, &mut self.targets); | |
| 86 | index_section(&doc.root, path, &out, &mut self.targets); | |
| 82 | 87 | } |
| 83 | 88 | } |
| 84 | 89 | |
| @@ -107,40 +112,37 @@ fn collect_targets(section: &Section, out: &mut Vec<TargetId>) { | ||
| 107 | 112 | } |
| 108 | 113 | } |
| 109 | 114 | |
| 110 | fn index_section(section: &Section, path: &Utf8Path, targets: &mut HashMap<TargetId, TargetLocation>) { | |
| 115 | fn index_section( | |
| 116 | section: &Section, | |
| 117 | path: &Utf8Path, | |
| 118 | out: &Utf8Path, | |
| 119 | targets: &mut HashMap<TargetId, TargetLocation>, | |
| 120 | ) { | |
| 111 | 121 | if let Some(h) = §ion.heading { |
| 112 | 122 | let anchor = h |
| 113 | 123 | .custom_id |
| 114 | 124 | .clone() |
| 115 | 125 | .or_else(|| h.id.clone()) |
| 116 | 126 | .unwrap_or_else(|| slugify(&plain_text(&h.title))); |
| 117 | if let Some(cid) = &h.custom_id { | |
| 127 | let mut record = |id: TargetId, anchor: Option<String>| { | |
| 118 | 128 | targets.insert( |
| 119 | TargetId::CustomId(cid.clone()), | |
| 129 | id, | |
| 120 | 130 | TargetLocation { |
| 121 | 131 | source_path: path.to_owned(), |
| 122 | anchor: Some(cid.clone()), | |
| 132 | output_path: out.to_owned(), | |
| 133 | anchor, | |
| 123 | 134 | }, |
| 124 | 135 | ); |
| 136 | }; | |
| 137 | if let Some(cid) = &h.custom_id { | |
| 138 | record(TargetId::CustomId(cid.clone()), Some(cid.clone())); | |
| 125 | 139 | } |
| 126 | 140 | if let Some(id) = &h.id { |
| 127 | targets.insert( | |
| 128 | TargetId::Id(id.clone()), | |
| 129 | TargetLocation { | |
| 130 | source_path: path.to_owned(), | |
| 131 | anchor: Some(id.clone()), | |
| 132 | }, | |
| 133 | ); | |
| 141 | record(TargetId::Id(id.clone()), Some(id.clone())); | |
| 134 | 142 | } |
| 135 | targets.insert( | |
| 136 | TargetId::Heading(plain_text(&h.title)), | |
| 137 | TargetLocation { | |
| 138 | source_path: path.to_owned(), | |
| 139 | anchor: Some(anchor), | |
| 140 | }, | |
| 141 | ); | |
| 143 | record(TargetId::Heading(plain_text(&h.title)), Some(anchor)); | |
| 142 | 144 | } |
| 143 | 145 | for child in §ion.children { |
| 144 | index_section(child, path, targets); | |
| 146 | index_section(child, path, out, targets); | |
| 145 | 147 | } |
| 146 | 148 | } |
src/lib.rs +1
| @@ -8,6 +8,7 @@ | ||
| 8 | 8 | //! [`render`] (RENDER) → [`template`] (TEMPLATE) → EMIT, with [`incremental`] |
| 9 | 9 | //! deciding which pages actually need rewriting. |
| 10 | 10 | |
| 11 | pub mod audit; | |
| 11 | 12 | pub mod incremental; |
| 12 | 13 | pub mod index; |
| 13 | 14 | pub mod model; |
src/main.rs +11
| @@ -49,6 +49,12 @@ enum Command { | ||
| 49 | 49 | /// Output directory to remove. |
| 50 | 50 | output: Utf8PathBuf, |
| 51 | 51 | }, |
| 52 | /// Audit a corpus: report which org constructs it uses and how they land against | |
| 53 | /// the v1 scope line. Reports names, counts and locations — never document text. | |
| 54 | Audit { | |
| 55 | /// Source directory (or single `.org` file) to audit. | |
| 56 | input: Utf8PathBuf, | |
| 57 | }, | |
| 52 | 58 | } |
| 53 | 59 | |
| 54 | 60 | fn main() -> Result<()> { |
| @@ -86,6 +92,11 @@ fn main() -> Result<()> { | ||
| 86 | 92 | // 6 lists `watch`; the real fs-notify integration is deferred). It rebuilds |
| 87 | 93 | // incrementally whenever any source file's mtime advances. |
| 88 | 94 | Command::Watch { input, output } => watch(&input, &output), |
| 95 | Command::Audit { input } => { | |
| 96 | let result = org_ssg::audit::audit(&input)?; | |
| 97 | print!("{}", org_ssg::audit::report(&result)); | |
| 98 | Ok(()) | |
| 99 | } | |
| 89 | 100 | Command::Clean { output } => { |
| 90 | 101 | if output.exists() { |
| 91 | 102 | fs::remove_dir_all(&output) |
src/resolve.rs +6 −2
| @@ -15,7 +15,7 @@ use camino::Utf8Path; | ||
| 15 | 15 | |
| 16 | 16 | use crate::index::{SymbolTable, TargetId}; |
| 17 | 17 | use crate::model::{Document, Element, Link, LinkTarget, Object, Section, TableRow}; |
| 18 | use crate::util::{normalize_link_path, output_url}; | |
| 18 | use crate::util::{normalize_link_path, output_path, output_url}; | |
| 19 | 19 | |
| 20 | 20 | /// A document whose links have been rewritten to concrete URLs. |
| 21 | 21 | #[derive(Debug, Clone)] |
| @@ -42,10 +42,13 @@ pub struct ResolveOutput { | ||
| 42 | 42 | pub fn resolve(doc: &Document, symbols: &SymbolTable) -> ResolveOutput { |
| 43 | 43 | let mut document = doc.clone(); |
| 44 | 44 | let from = doc.source_path.clone(); |
| 45 | // URLs are computed between *output* paths, which `#+SLUG:` can rename. | |
| 46 | let from_out = output_path(&from, &doc.keywords); | |
| 45 | 47 | let mut used = Vec::new(); |
| 46 | 48 | let mut broken = Vec::new(); |
| 47 | 49 | let mut cx = Cx { |
| 48 | 50 | from: &from, |
| 51 | from_out: &from_out, | |
| 49 | 52 | symbols, |
| 50 | 53 | used: &mut used, |
| 51 | 54 | broken: &mut broken, |
| @@ -70,6 +73,7 @@ fn human_text(target: &LinkTarget) -> Option<String> { | ||
| 70 | 73 | |
| 71 | 74 | struct Cx<'a> { |
| 72 | 75 | from: &'a Utf8Path, |
| 76 | from_out: &'a Utf8Path, | |
| 73 | 77 | symbols: &'a SymbolTable, |
| 74 | 78 | used: &'a mut Vec<TargetId>, |
| 75 | 79 | broken: &'a mut Vec<BrokenLink>, |
| @@ -167,7 +171,7 @@ impl Cx<'_> { | ||
| 167 | 171 | link.description = Some(vec![Object::Text(text)]); |
| 168 | 172 | } |
| 169 | 173 | } |
| 170 | let url = output_url(self.from, &loc.source_path, loc.anchor.as_deref()); | |
| 174 | let url = output_url(self.from_out, &loc.output_path, loc.anchor.as_deref()); | |
| 171 | 175 | link.target = LinkTarget::External(url); |
| 172 | 176 | } |
| 173 | 177 | None => { |
src/site.rs +25 −6
| @@ -26,7 +26,7 @@ use crate::parser::parse; | ||
| 26 | 26 | use crate::render::{render, syntax_css, Html, SyntectHighlighter}; |
| 27 | 27 | use crate::resolve::resolve; |
| 28 | 28 | use crate::template::{template_sources, NavItem, Templater}; |
| 29 | use crate::util::{output_url, relative_root}; | |
| 29 | use crate::util::{output_path, output_url, relative_root}; | |
| 30 | 30 | |
| 31 | 31 | /// A fully built page: source and output paths (relative to their roots) and its |
| 32 | 32 | /// final templated HTML. |
| @@ -100,12 +100,27 @@ fn prepare_pages(src: &Utf8Path) -> Result<(Vec<PagePrep>, SymbolTable)> { | ||
| 100 | 100 | symbols.index_document(doc); |
| 101 | 101 | } |
| 102 | 102 | |
| 103 | // Nav is global; titles come from #+TITLE (falling back to the file stem). | |
| 103 | // Nav is global; titles come from #+TITLE (falling back to the file stem) and URLs | |
| 104 | // from each page's output path, which `#+SLUG:` can rename. | |
| 104 | 105 | let entries: Vec<(Utf8PathBuf, String)> = docs |
| 105 | 106 | .iter() |
| 106 | .map(|d| (d.source_path.clone(), page_title(d))) | |
| 107 | .map(|d| (output_path(&d.source_path, &d.keywords), page_title(d))) | |
| 107 | 108 | .collect(); |
| 108 | 109 | |
| 110 | // Two sources emitting one page would silently drop a page — and with slugs, a | |
| 111 | // collision is a typo away and invisible in the source filenames. | |
| 112 | let mut claimed: std::collections::HashMap<&Utf8PathBuf, &Utf8PathBuf> = | |
| 113 | std::collections::HashMap::new(); | |
| 114 | for (doc, (out, _)) in docs.iter().zip(&entries) { | |
| 115 | if let Some(other) = claimed.insert(out, &doc.source_path) { | |
| 116 | anyhow::bail!( | |
| 117 | "output collision: {} and {} both build to {out} (check their #+SLUG:)", | |
| 118 | other, | |
| 119 | doc.source_path | |
| 120 | ); | |
| 121 | } | |
| 122 | } | |
| 123 | ||
| 109 | 124 | let mut pages = Vec::new(); |
| 110 | 125 | for doc in &docs { |
| 111 | 126 | let out = resolve(doc, &symbols); |
| @@ -113,18 +128,20 @@ fn prepare_pages(src: &Utf8Path) -> Result<(Vec<PagePrep>, SymbolTable)> { | ||
| 113 | 128 | let broken: Vec<TargetId> = out.broken.iter().map(|b| b.target.clone()).collect(); |
| 114 | 129 | let defines: HashSet<TargetId> = document_targets(doc).into_iter().collect(); |
| 115 | 130 | |
| 131 | let output = output_path(&doc.source_path, &doc.keywords); | |
| 132 | ||
| 116 | 133 | // Nav links are relative to *this* page (spec URL scheme, §8 Q3). |
| 117 | 134 | let nav: Vec<NavItem> = entries |
| 118 | 135 | .iter() |
| 119 | 136 | .map(|(path, title)| NavItem { |
| 120 | 137 | title: title.clone(), |
| 121 | url: output_url(&doc.source_path, path, None), | |
| 138 | url: output_url(&output, path, None), | |
| 122 | 139 | }) |
| 123 | 140 | .collect(); |
| 124 | 141 | |
| 125 | 142 | pages.push(PagePrep { |
| 126 | 143 | source: doc.source_path.clone(), |
| 127 | output: doc.source_path.with_extension("html"), | |
| 144 | output, | |
| 128 | 145 | title: page_title(doc), |
| 129 | 146 | content_hash: doc.content_hash, |
| 130 | 147 | resolved: out.resolved, |
| @@ -190,9 +207,11 @@ pub fn build_site(src: &Utf8Path, out: &Utf8Path, opts: &BuildOptions) -> Result | ||
| 190 | 207 | // chrome on every page — is built from every page's (path, title), so a title/path |
| 191 | 208 | // change or a page add/remove must re-render every page (else stale nav on disk). |
| 192 | 209 | let cfg = BuildConfig::default(); |
| 210 | // Keyed on the *output* path: a `#+SLUG:` change moves a page's URL, which changes | |
| 211 | // the nav on every other page even though no source filename moved. | |
| 193 | 212 | let nav_entries: Vec<(String, String)> = preps |
| 194 | 213 | .iter() |
| 195 | .map(|p| (p.source.to_string(), p.title.clone())) | |
| 214 | .map(|p| (p.output.to_string(), p.title.clone())) | |
| 196 | 215 | .collect(); |
| 197 | 216 | let cfg_hash = combine(config_hash(&cfg), site_structure_hash(&nav_entries)); |
| 198 | 217 | let tmpl_hash = template_hash(template_sources()); |
src/util.rs +55 −9
| @@ -4,7 +4,50 @@ | ||
| 4 | 4 | |
| 5 | 5 | use camino::{Utf8Path, Utf8PathBuf}; |
| 6 | 6 | |
| 7 | use crate::model::Object; | |
| 7 | use crate::model::{Keywords, Object}; | |
| 8 | ||
| 9 | /// The output path for a document, relative to the site root. | |
| 10 | /// | |
| 11 | /// Normally this is the source path with `.org` swapped for `.html`, but a `#+SLUG:` | |
| 12 | /// keyword renames the file — which is how the target corpus works: 178 of its 179 files | |
| 13 | /// set one, and `2018-11-28-aes-encryption.org` publishes as `aes-encryption.html`. The | |
| 14 | /// slug names the *file*, never the directory, so the page stays where its source lives. | |
| 15 | pub fn output_path(source: &Utf8Path, keywords: &Keywords) -> Utf8PathBuf { | |
| 16 | let slug = keywords | |
| 17 | .entries | |
| 18 | .iter() | |
| 19 | .find(|(k, _)| k.eq_ignore_ascii_case("SLUG")) | |
| 20 | .map(|(_, v)| sanitize_slug(v)) | |
| 21 | .filter(|s| !s.is_empty()); | |
| 22 | ||
| 23 | match slug { | |
| 24 | Some(slug) => { | |
| 25 | let dir = source.parent().unwrap_or_else(|| Utf8Path::new("")); | |
| 26 | dir.join(format!("{slug}.html")) | |
| 27 | } | |
| 28 | None => source.with_extension("html"), | |
| 29 | } | |
| 30 | } | |
| 31 | ||
| 32 | /// Reduce a slug to a safe single filename component. | |
| 33 | /// | |
| 34 | /// A slug is author-controlled text that becomes a path we write to, so `../../etc/x` | |
| 35 | /// has to be impossible by construction rather than by convention: separators and dots | |
| 36 | /// are folded to `-`, which cannot traverse and cannot produce a hidden file. | |
| 37 | fn sanitize_slug(raw: &str) -> String { | |
| 38 | let mut out = String::with_capacity(raw.len()); | |
| 39 | let mut prev_dash = false; | |
| 40 | for c in raw.trim().chars() { | |
| 41 | if c.is_ascii_alphanumeric() || c == '_' { | |
| 42 | out.extend(c.to_lowercase()); | |
| 43 | prev_dash = false; | |
| 44 | } else if !prev_dash { | |
| 45 | out.push('-'); | |
| 46 | prev_dash = true; | |
| 47 | } | |
| 48 | } | |
| 49 | out.trim_matches('-').to_string() | |
| 50 | } | |
| 8 | 51 | |
| 9 | 52 | /// Flatten inline objects to their plain-text content (markup stripped). Used to |
| 10 | 53 | /// derive heading anchors and `[[*Heading]]` link identities (spec §4.3). |
| @@ -49,16 +92,19 @@ pub fn slugify(text: &str) -> String { | ||
| 49 | 92 | out.trim_matches('-').to_string() |
| 50 | 93 | } |
| 51 | 94 | |
| 52 | /// The output URL to reach `to_rel` (a source `.org` path relative to the site root) | |
| 53 | /// from the page at `from_rel`, honoring an optional `anchor`. Same-file links reduce | |
| 54 | /// to a bare `#anchor` fragment; cross-file links become a relative `.html` path. | |
| 55 | pub fn output_url(from_rel: &Utf8Path, to_rel: &Utf8Path, anchor: Option<&str>) -> String { | |
| 56 | let path = if from_rel == to_rel { | |
| 95 | /// The URL to reach the page output at `to_out` from the page output at `from_out`, | |
| 96 | /// honoring an optional `anchor`. Same-page links reduce to a bare `#anchor` fragment; | |
| 97 | /// cross-page links become a relative path. | |
| 98 | /// | |
| 99 | /// Both arguments are *output* paths, not source paths, because `#+SLUG:` means the two | |
| 100 | /// no longer correspond: deriving the URL here would reintroduce the filename assumption | |
| 101 | /// that [`output_path`] exists to remove. | |
| 102 | pub fn output_url(from_out: &Utf8Path, to_out: &Utf8Path, anchor: Option<&str>) -> String { | |
| 103 | let path = if from_out == to_out { | |
| 57 | 104 | String::new() |
| 58 | 105 | } else { |
| 59 | let to_html = to_rel.with_extension("html"); | |
| 60 | let from_dir = from_rel.parent().unwrap_or_else(|| Utf8Path::new("")); | |
| 61 | relative_path(from_dir, &to_html) | |
| 106 | let from_dir = from_out.parent().unwrap_or_else(|| Utf8Path::new("")); | |
| 107 | relative_path(from_dir, to_out) | |
| 62 | 108 | }; |
| 63 | 109 | match anchor { |
| 64 | 110 | Some(a) if !a.is_empty() => { |
tests/oracle.el added +42
| @@ -0,0 +1,42 @@ | ||
| 1 | ;;; oracle.el --- ground-truth HTML export for the org-ssg differential tests -*- lexical-binding: t -*- | |
| 2 | ||
| 3 | ;; Exports the org file named by $ORG_ORACLE_INPUT to HTML on stdout, using org's own | |
| 4 | ;; exporter — the same one weblorg wraps to publish the corpus this project targets. | |
| 5 | ;; Run with: ORG_ORACLE_INPUT=x.org emacs -Q --batch -l tests/oracle.el | |
| 6 | ;; | |
| 7 | ;; `-Q' is deliberate: no user init, so the oracle is the stock org exporter and not | |
| 8 | ;; this machine's Emacs configuration. The path is passed by environment variable | |
| 9 | ;; rather than as an argument because batch Emacs would otherwise try to visit it. | |
| 10 | ||
| 11 | (require 'org) | |
| 12 | (require 'ox-html) | |
| 13 | ||
| 14 | ;; Presentation settings are normalized so the diff carries semantic divergences only. | |
| 15 | ;; Everything that affects *content* is left at its default, because the point is to | |
| 16 | ;; learn what stock org does — normalizing that away would be marking our own homework. | |
| 17 | (setq org-export-with-toc nil ; we emit no table of contents | |
| 18 | org-export-with-section-numbers nil ; we do not number headings | |
| 19 | org-html-toplevel-hlevel 1 ; org defaults to h2 for a level-1 heading, | |
| 20 | ; because a template supplies the page <h1>. | |
| 21 | ; Aligning here keeps a global +1 offset from | |
| 22 | ; drowning every real finding in the diff. | |
| 23 | org-html-htmlize-output-type nil ; plain <pre>, not htmlize spans: we highlight | |
| 24 | ; with syntect, so comparing code *text* is | |
| 25 | ; the meaningful part | |
| 26 | org-html-head-include-default-style nil | |
| 27 | org-html-head-include-scripts nil | |
| 28 | ;; Fixtures link to ids that live in org-ssg's own symbol table, not in an | |
| 29 | ;; `org-id' database. Without this, org aborts the whole export on the first one. | |
| 30 | org-export-with-broken-links t | |
| 31 | make-backup-files nil) | |
| 32 | ||
| 33 | (let ((input (getenv "ORG_ORACLE_INPUT"))) | |
| 34 | (unless input | |
| 35 | (error "ORG_ORACLE_INPUT is not set")) | |
| 36 | (with-temp-buffer | |
| 37 | (insert-file-contents input) | |
| 38 | (org-mode) | |
| 39 | ;; BODY-ONLY: emit the content, not a full document with <head> chrome. | |
| 40 | (princ (org-export-as 'html nil nil t nil)))) | |
| 41 | ||
| 42 | ;;; oracle.el ends here | |
tests/oracle.rs added +454
| @@ -0,0 +1,454 @@ | ||
| 1 | //! The `emacs --batch` ground-truth oracle (spec §5, Phase 0). | |
| 2 | //! | |
| 3 | //! Every other test in this suite checks org-ssg against org-ssg: a snapshot says our | |
| 4 | //! output has not *changed*, never that it is *right*. Those two questions are different, | |
| 5 | //! and only one of them matters to someone whose site is currently published by Emacs. | |
| 6 | //! This file answers the second by exporting the same fixture with org's own HTML | |
| 7 | //! exporter — the exporter weblorg wraps to publish the target corpus — and diffing the | |
| 8 | //! two. | |
| 9 | //! | |
| 10 | //! **What is compared.** Byte equality is not a useful goal: org wraps every section in | |
| 11 | //! `outline-container` divs keyed by generated ids, and no amount of agreement on | |
| 12 | //! semantics would survive that. Both sides are reduced to a *semantic skeleton* — the | |
| 13 | //! sequence of element opens, closes, and text runs, with `<div>`s and all attributes | |
| 14 | //! except `href`/`src` dropped, whitespace collapsed, and entities decoded. What remains | |
| 15 | //! is the question worth asking: does org think this is a `<blockquote><p>`, and do we? | |
| 16 | //! | |
| 17 | //! **What the result means.** These tests do not assert agreement — they *snapshot the | |
| 18 | //! disagreement*. A divergence report that is checked in and reviewed is worth more than | |
| 19 | //! a red test nobody can act on, and it makes any new divergence show up as a diff in | |
| 20 | //! code review. A few invariants that must never break are asserted outright. | |
| 21 | //! | |
| 22 | //! The suite skips cleanly when Emacs is absent, so it never blocks a machine or CI | |
| 23 | //! runner that has no Emacs. | |
| 24 | ||
| 25 | use std::process::Command; | |
| 26 | ||
| 27 | use camino::Utf8PathBuf; | |
| 28 | ||
| 29 | use org_ssg::parser::parse; | |
| 30 | use org_ssg::render::{render, Html, SyntectHighlighter}; | |
| 31 | use org_ssg::resolve::ResolvedDoc; | |
| 32 | ||
| 33 | fn manifest_dir() -> Utf8PathBuf { | |
| 34 | Utf8PathBuf::from(env!("CARGO_MANIFEST_DIR")) | |
| 35 | } | |
| 36 | ||
| 37 | /// Is a usable Emacs on PATH? The oracle is a development instrument, not a build | |
| 38 | /// dependency, so its absence skips rather than fails. | |
| 39 | fn emacs_available() -> bool { | |
| 40 | Command::new("emacs") | |
| 41 | .arg("--version") | |
| 42 | .output() | |
| 43 | .map(|o| o.status.success()) | |
| 44 | .unwrap_or(false) | |
| 45 | } | |
| 46 | ||
| 47 | /// Export a fixture with org's own HTML exporter. | |
| 48 | fn org_export(fixture: &str) -> String { | |
| 49 | let root = manifest_dir(); | |
| 50 | let output = Command::new("emacs") | |
| 51 | .args(["-Q", "--batch", "-l"]) | |
| 52 | .arg(root.join("tests/oracle.el")) | |
| 53 | .env("ORG_ORACLE_INPUT", root.join("fixtures").join(fixture)) | |
| 54 | .current_dir(&root) | |
| 55 | .output() | |
| 56 | .expect("run emacs"); | |
| 57 | assert!( | |
| 58 | output.status.success(), | |
| 59 | "emacs export of {fixture} failed:\n{}", | |
| 60 | String::from_utf8_lossy(&output.stderr) | |
| 61 | ); | |
| 62 | String::from_utf8(output.stdout).expect("emacs emits UTF-8") | |
| 63 | } | |
| 64 | ||
| 65 | /// Render a fixture with org-ssg. | |
| 66 | fn our_export(fixture: &str) -> String { | |
| 67 | let path = manifest_dir().join("fixtures").join(fixture); | |
| 68 | let source = std::fs::read_to_string(&path).expect("read fixture"); | |
| 69 | let document = parse(Utf8PathBuf::from(fixture).as_path(), &source).expect("parse"); | |
| 70 | let Html(html) = render(&ResolvedDoc { document }, &SyntectHighlighter::new()); | |
| 71 | html | |
| 72 | } | |
| 73 | ||
| 74 | // --------------------------------------------------------------------------- | |
| 75 | // HTML → semantic skeleton | |
| 76 | // --------------------------------------------------------------------------- | |
| 77 | ||
| 78 | /// Elements dropped from the skeleton entirely, because once attributes are gone they | |
| 79 | /// carry no meaning the two exporters could agree or disagree *about*. | |
| 80 | /// | |
| 81 | /// `div` is pure layout: org wraps every section in `outline-container`/`outline-text` | |
| 82 | /// wrappers and we emit none. `span` is the same story at the inline level, and matters | |
| 83 | /// far more than it looks: syntect emits one span per code token, so keeping them made a | |
| 84 | /// source block contribute ~60 skeleton lines of pure noise and dragged the agreement on | |
| 85 | /// `blocks.org` down to 36% — a number that said nothing about whether we render blocks | |
| 86 | /// correctly. Text still carries the signal: a `<span class="todo">` shows up as its | |
| 87 | /// text, `"TODO"`, which is the part worth comparing. | |
| 88 | const IGNORED: &[&str] = &["div", "span"]; | |
| 89 | ||
| 90 | /// Attributes kept in the skeleton. Ids and classes are generated (`org6c28c1b`) or | |
| 91 | /// cosmetic (`org-ul`); `href` and `src` are the content. | |
| 92 | const KEPT_ATTRS: &[&str] = &["href", "src"]; | |
| 93 | ||
| 94 | /// HTML void elements, which never emit a close event. | |
| 95 | const VOID: &[&str] = &[ | |
| 96 | "br", "hr", "img", "input", "meta", "link", "col", "area", "base", "source", "wbr", | |
| 97 | ]; | |
| 98 | ||
| 99 | /// Reduce an HTML fragment to its semantic skeleton: one line per element open, element | |
| 100 | /// close, or text run. | |
| 101 | fn skeleton(html: &str) -> Vec<String> { | |
| 102 | let mut out = Vec::new(); | |
| 103 | let chars: Vec<char> = html.chars().collect(); | |
| 104 | let mut i = 0; | |
| 105 | let mut text = String::new(); | |
| 106 | ||
| 107 | while i < chars.len() { | |
| 108 | if chars[i] != '<' { | |
| 109 | text.push(chars[i]); | |
| 110 | i += 1; | |
| 111 | continue; | |
| 112 | } | |
| 113 | ||
| 114 | // Comments and doctypes carry nothing. | |
| 115 | if chars[i..].starts_with(&['<', '!']) { | |
| 116 | i += match find_from(&chars, i, ">") { | |
| 117 | Some(end) => end - i + 1, | |
| 118 | None => break, | |
| 119 | }; | |
| 120 | continue; | |
| 121 | } | |
| 122 | let Some(end) = find_from(&chars, i, ">") else { | |
| 123 | break; | |
| 124 | }; | |
| 125 | let raw: String = chars[i + 1..end].iter().collect(); | |
| 126 | i = end + 1; | |
| 127 | ||
| 128 | let raw = raw.trim().trim_end_matches('/').trim().to_string(); | |
| 129 | // Text is flushed only when a tag is actually *emitted*. Text either side of an | |
| 130 | // ignored tag therefore merges into one run, which is what makes a highlighted | |
| 131 | // source block compare as the one string of code it is, rather than as a | |
| 132 | // token-by-token sequence that has to line up exactly. | |
| 133 | if let Some(name) = raw.strip_prefix('/') { | |
| 134 | let name = name.trim().to_ascii_lowercase(); | |
| 135 | if !IGNORED.contains(&name.as_str()) && !VOID.contains(&name.as_str()) { | |
| 136 | flush_text(&mut text, &mut out); | |
| 137 | out.push(format!("</{name}>")); | |
| 138 | } | |
| 139 | continue; | |
| 140 | } | |
| 141 | let mut parts = raw.splitn(2, char::is_whitespace); | |
| 142 | let name = parts.next().unwrap_or("").to_ascii_lowercase(); | |
| 143 | if name.is_empty() || IGNORED.contains(&name.as_str()) { | |
| 144 | continue; | |
| 145 | } | |
| 146 | let attrs = kept_attributes(parts.next().unwrap_or("")); | |
| 147 | flush_text(&mut text, &mut out); | |
| 148 | out.push(format!("<{name}{attrs}>")); | |
| 149 | } | |
| 150 | flush_text(&mut text, &mut out); | |
| 151 | out | |
| 152 | } | |
| 153 | ||
| 154 | fn flush_text(text: &mut String, out: &mut Vec<String>) { | |
| 155 | let decoded = decode_entities(text); | |
| 156 | let collapsed = decoded.split_whitespace().collect::<Vec<_>>().join(" "); | |
| 157 | if !collapsed.is_empty() { | |
| 158 | out.push(format!("{collapsed:?}")); | |
| 159 | } | |
| 160 | text.clear(); | |
| 161 | } | |
| 162 | ||
| 163 | fn find_from(chars: &[char], from: usize, needle: &str) -> Option<usize> { | |
| 164 | let n: Vec<char> = needle.chars().collect(); | |
| 165 | (from..chars.len()).find(|&k| chars[k..].starts_with(&n[..])) | |
| 166 | } | |
| 167 | ||
| 168 | /// Keep only the content-bearing attributes, in a stable order. | |
| 169 | fn kept_attributes(rest: &str) -> String { | |
| 170 | let mut kept: Vec<(String, String)> = Vec::new(); | |
| 171 | for attr in KEPT_ATTRS { | |
| 172 | if let Some(value) = attribute_value(rest, attr) { | |
| 173 | kept.push(((*attr).to_string(), value)); | |
| 174 | } | |
| 175 | } | |
| 176 | kept.iter() | |
| 177 | .map(|(k, v)| format!(" {k}=\"{}\"", decode_entities(v))) | |
| 178 | .collect() | |
| 179 | } | |
| 180 | ||
| 181 | fn attribute_value(rest: &str, name: &str) -> Option<String> { | |
| 182 | let mut search = rest; | |
| 183 | while let Some(pos) = search.find(name) { | |
| 184 | let before_ok = pos == 0 | |
| 185 | || search[..pos] | |
| 186 | .chars() | |
| 187 | .next_back() | |
| 188 | .is_some_and(char::is_whitespace); | |
| 189 | let after = &search[pos + name.len()..]; | |
| 190 | let after_trimmed = after.trim_start(); | |
| 191 | if before_ok && after_trimmed.starts_with('=') { | |
| 192 | let value = after_trimmed[1..].trim_start(); | |
| 193 | let quote = value.chars().next()?; | |
| 194 | if quote == '"' || quote == '\'' { | |
| 195 | let end = value[1..].find(quote)? + 1; | |
| 196 | return Some(value[1..end].to_string()); | |
| 197 | } | |
| 198 | let end = value.find(char::is_whitespace).unwrap_or(value.len()); | |
| 199 | return Some(value[..end].to_string()); | |
| 200 | } | |
| 201 | search = &search[pos + name.len()..]; | |
| 202 | } | |
| 203 | None | |
| 204 | } | |
| 205 | ||
| 206 | /// Decode the entities either exporter is likely to emit, so an encoding difference is | |
| 207 | /// never reported as a semantic one. | |
| 208 | fn decode_entities(s: &str) -> String { | |
| 209 | let mut out = String::with_capacity(s.len()); | |
| 210 | let mut rest = s; | |
| 211 | while let Some(amp) = rest.find('&') { | |
| 212 | out.push_str(&rest[..amp]); | |
| 213 | let tail = &rest[amp..]; | |
| 214 | let Some(semi) = tail.find(';').filter(|s| *s <= 12) else { | |
| 215 | out.push('&'); | |
| 216 | rest = &tail[1..]; | |
| 217 | continue; | |
| 218 | }; | |
| 219 | let entity = &tail[1..semi]; | |
| 220 | let decoded = match entity { | |
| 221 | "amp" => Some('&'), | |
| 222 | "lt" => Some('<'), | |
| 223 | "gt" => Some('>'), | |
| 224 | "quot" => Some('"'), | |
| 225 | "apos" => Some('\''), | |
| 226 | "nbsp" => Some(' '), | |
| 227 | _ => entity | |
| 228 | .strip_prefix('#') | |
| 229 | .and_then(|n| match n.strip_prefix(['x', 'X']) { | |
| 230 | Some(hex) => u32::from_str_radix(hex, 16).ok(), | |
| 231 | None => n.parse::<u32>().ok(), | |
| 232 | }) | |
| 233 | .and_then(char::from_u32), | |
| 234 | }; | |
| 235 | match decoded { | |
| 236 | // A non-breaking space is a space for comparison purposes. | |
| 237 | Some('\u{a0}') => out.push(' '), | |
| 238 | Some(c) => out.push(c), | |
| 239 | None => { | |
| 240 | out.push('&'); | |
| 241 | rest = &tail[1..]; | |
| 242 | continue; | |
| 243 | } | |
| 244 | } | |
| 245 | rest = &tail[semi + 1..]; | |
| 246 | } | |
| 247 | out.push_str(rest); | |
| 248 | out | |
| 249 | } | |
| 250 | ||
| 251 | // --------------------------------------------------------------------------- | |
| 252 | // Divergence report | |
| 253 | // --------------------------------------------------------------------------- | |
| 254 | ||
| 255 | /// A unified diff of the two skeletons, via a longest-common-subsequence walk. `-` is | |
| 256 | /// org-ssg, `+` is Emacs. | |
| 257 | fn divergence(ours: &[String], theirs: &[String]) -> String { | |
| 258 | let (n, m) = (ours.len(), theirs.len()); | |
| 259 | // lcs[i][j] = length of the longest common subsequence of ours[i..] and theirs[j..]. | |
| 260 | let mut lcs = vec![vec![0usize; m + 1]; n + 1]; | |
| 261 | for i in (0..n).rev() { | |
| 262 | for j in (0..m).rev() { | |
| 263 | lcs[i][j] = if ours[i] == theirs[j] { | |
| 264 | lcs[i + 1][j + 1] + 1 | |
| 265 | } else { | |
| 266 | lcs[i + 1][j].max(lcs[i][j + 1]) | |
| 267 | }; | |
| 268 | } | |
| 269 | } | |
| 270 | ||
| 271 | let mut out = String::new(); | |
| 272 | let (mut i, mut j) = (0, 0); | |
| 273 | let mut agreed = 0usize; | |
| 274 | while i < n && j < m { | |
| 275 | if ours[i] == theirs[j] { | |
| 276 | out.push_str(&format!(" {}\n", ours[i])); | |
| 277 | agreed += 1; | |
| 278 | i += 1; | |
| 279 | j += 1; | |
| 280 | } else if lcs[i + 1][j] >= lcs[i][j + 1] { | |
| 281 | out.push_str(&format!("- {}\n", ours[i])); | |
| 282 | i += 1; | |
| 283 | } else { | |
| 284 | out.push_str(&format!("+ {}\n", theirs[j])); | |
| 285 | j += 1; | |
| 286 | } | |
| 287 | } | |
| 288 | for line in &ours[i..] { | |
| 289 | out.push_str(&format!("- {line}\n")); | |
| 290 | } | |
| 291 | for line in &theirs[j..] { | |
| 292 | out.push_str(&format!("+ {line}\n")); | |
| 293 | } | |
| 294 | ||
| 295 | let total = n.max(m); | |
| 296 | let pct = if total == 0 { | |
| 297 | 100.0 | |
| 298 | } else { | |
| 299 | 100.0 * agreed as f64 / total as f64 | |
| 300 | }; | |
| 301 | format!("agreement: {agreed}/{total} skeleton lines ({pct:.1}%)\n(- org-ssg, + emacs)\n\n{out}") | |
| 302 | } | |
| 303 | ||
| 304 | /// Snapshot the divergence between org-ssg and Emacs for one fixture. | |
| 305 | fn compare(fixture: &str) -> Option<String> { | |
| 306 | if !emacs_available() { | |
| 307 | eprintln!("skipping oracle comparison for {fixture}: no emacs on PATH"); | |
| 308 | return None; | |
| 309 | } | |
| 310 | let ours = skeleton(&our_export(fixture)); | |
| 311 | let theirs = skeleton(&org_export(fixture)); | |
| 312 | Some(divergence(&ours, &theirs)) | |
| 313 | } | |
| 314 | ||
| 315 | macro_rules! oracle_test { | |
| 316 | ($name:ident, $fixture:literal) => { | |
| 317 | #[test] | |
| 318 | fn $name() { | |
| 319 | if let Some(report) = compare($fixture) { | |
| 320 | insta::assert_snapshot!(report); | |
| 321 | } | |
| 322 | } | |
| 323 | }; | |
| 324 | } | |
| 325 | ||
| 326 | oracle_test!(oracle_minimal, "minimal.org"); | |
| 327 | oracle_test!(oracle_core, "core.org"); | |
| 328 | oracle_test!(oracle_headings, "headings.org"); | |
| 329 | oracle_test!(oracle_lists, "lists.org"); | |
| 330 | oracle_test!(oracle_blocks, "blocks.org"); | |
| 331 | oracle_test!(oracle_table, "table.org"); | |
| 332 | oracle_test!(oracle_footnote, "footnote.org"); | |
| 333 | oracle_test!(oracle_timestamps, "timestamps.org"); | |
| 334 | oracle_test!(oracle_images, "images.org"); | |
| 335 | oracle_test!(oracle_elements, "elements.org"); | |
| 336 | ||
| 337 | // --------------------------------------------------------------------------- | |
| 338 | // Invariants that must hold against the oracle, not merely be snapshotted | |
| 339 | // --------------------------------------------------------------------------- | |
| 340 | ||
| 341 | /// How many headings a document has and at what depth is the shape of the document. | |
| 342 | /// Getting it wrong reorganizes someone's writing, so it is asserted rather than | |
| 343 | /// snapshotted. Heading *decoration* (priority cookies, tag markup) is a policy | |
| 344 | /// difference and is left to the snapshots. | |
| 345 | #[test] | |
| 346 | fn heading_structure_matches_emacs() { | |
| 347 | if !emacs_available() { | |
| 348 | eprintln!("skipping: no emacs on PATH"); | |
| 349 | return; | |
| 350 | } | |
| 351 | for fixture in ["minimal.org", "core.org", "headings.org", "lists.org"] { | |
| 352 | let ours = heading_levels(&skeleton(&our_export(fixture))); | |
| 353 | let theirs = heading_levels(&skeleton(&org_export(fixture))); | |
| 354 | assert_eq!( | |
| 355 | ours, theirs, | |
| 356 | "heading structure diverges from Emacs in {fixture}" | |
| 357 | ); | |
| 358 | } | |
| 359 | } | |
| 360 | ||
| 361 | /// The sequence of heading open tags, e.g. `["<h1>", "<h2>", "<h1>"]`. | |
| 362 | fn heading_levels(skeleton: &[String]) -> Vec<String> { | |
| 363 | skeleton | |
| 364 | .iter() | |
| 365 | .filter(|l| l.starts_with("<h") && l[2..].starts_with(|c: char| c.is_ascii_digit())) | |
| 366 | .cloned() | |
| 367 | .collect() | |
| 368 | } | |
| 369 | ||
| 370 | /// A list is the construct where nesting is easiest to get subtly wrong, and where being | |
| 371 | /// wrong changes the meaning of the document rather than its looks. | |
| 372 | #[test] | |
| 373 | fn list_nesting_matches_emacs() { | |
| 374 | if !emacs_available() { | |
| 375 | eprintln!("skipping: no emacs on PATH"); | |
| 376 | return; | |
| 377 | } | |
| 378 | let ours = list_shape(&skeleton(&our_export("lists.org"))); | |
| 379 | let theirs = list_shape(&skeleton(&org_export("lists.org"))); | |
| 380 | assert_eq!(ours, theirs, "list nesting diverges from Emacs"); | |
| 381 | } | |
| 382 | ||
| 383 | /// The sequence of list opens/closes, ignoring content — the shape of the nesting. | |
| 384 | fn list_shape(skeleton: &[String]) -> Vec<String> { | |
| 385 | skeleton | |
| 386 | .iter() | |
| 387 | .filter(|l| { | |
| 388 | matches!( | |
| 389 | l.as_str(), | |
| 390 | "<ul>" | "</ul>" | "<ol>" | "</ol>" | "<li>" | "</li>" | "<dl>" | "</dl>" | |
| 391 | | "<dt>" | "</dt>" | "<dd>" | "</dd>" | |
| 392 | ) | |
| 393 | }) | |
| 394 | .cloned() | |
| 395 | .collect() | |
| 396 | } | |
| 397 | ||
| 398 | /// Code must survive verbatim. Highlighting markup differs by construction (syntect | |
| 399 | /// spans vs htmlize), but if the *characters of the program* differ, we have corrupted | |
| 400 | /// the author's content. | |
| 401 | #[test] | |
| 402 | fn source_block_text_matches_emacs() { | |
| 403 | if !emacs_available() { | |
| 404 | eprintln!("skipping: no emacs on PATH"); | |
| 405 | return; | |
| 406 | } | |
| 407 | for fixture in ["blocks.org", "core.org", "elements.org"] { | |
| 408 | let ours = code_text(&our_export(fixture)); | |
| 409 | let theirs = code_text(&org_export(fixture)); | |
| 410 | assert_eq!(ours, theirs, "source block text diverges from Emacs in {fixture}"); | |
| 411 | } | |
| 412 | } | |
| 413 | ||
| 414 | /// All text inside `<pre>` blocks, with tags stripped and whitespace collapsed. | |
| 415 | fn code_text(html: &str) -> Vec<String> { | |
| 416 | let mut out = Vec::new(); | |
| 417 | let mut rest = html; | |
| 418 | while let Some(start) = rest.find("<pre") { | |
| 419 | let after = &rest[start..]; | |
| 420 | let Some(open_end) = after.find('>') else { break }; | |
| 421 | let Some(close) = after.find("</pre>") else { break }; | |
| 422 | let inner = &after[open_end + 1..close]; | |
| 423 | out.push(strip_tags(inner)); | |
| 424 | rest = &after[close + 6..]; | |
| 425 | } | |
| 426 | out | |
| 427 | } | |
| 428 | ||
| 429 | /// All text in a fragment with tags removed and entities decoded, then whitespace | |
| 430 | /// collapsed once at the end. | |
| 431 | /// | |
| 432 | /// [`skeleton`] cannot do this job: it trims each text run individually, which is | |
| 433 | /// invisible for prose (one run per paragraph) but destructive for highlighted code, | |
| 434 | /// where syntect splits a line into one run per token and the spaces *between* tokens | |
| 435 | /// live at the edges of those runs. Trimming each run turns `def greet` into `defgreet`. | |
| 436 | fn strip_tags(html: &str) -> String { | |
| 437 | let mut text = String::new(); | |
| 438 | let mut rest = html; | |
| 439 | while let Some(open) = rest.find('<') { | |
| 440 | text.push_str(&rest[..open]); | |
| 441 | match rest[open..].find('>') { | |
| 442 | Some(close) => rest = &rest[open + close + 1..], | |
| 443 | None => { | |
| 444 | rest = ""; | |
| 445 | break; | |
| 446 | } | |
| 447 | } | |
| 448 | } | |
| 449 | text.push_str(rest); | |
| 450 | decode_entities(&text) | |
| 451 | .split_whitespace() | |
| 452 | .collect::<Vec<_>>() | |
| 453 | .join(" ") | |
| 454 | } | |
tests/site.rs +83
| @@ -122,3 +122,86 @@ fn table_render() { | ||
| 122 | 122 | fn footnote_render() { |
| 123 | 123 | insta::assert_snapshot!(render_fragment("footnote.org")); |
| 124 | 124 | } |
| 125 | ||
| 126 | // --------------------------------------------------------------------------- | |
| 127 | // `#+SLUG:` output paths (Phase 0 corpus-audit finding) | |
| 128 | // --------------------------------------------------------------------------- | |
| 129 | ||
| 130 | /// The audit found `#+SLUG:` in 178 of the target corpus's 179 files, and the live site | |
| 131 | /// derives every URL from it — `2018-11-28-aes-encryption.org` publishes as | |
| 132 | /// `aes-encryption.html`. Deriving output paths from source filenames would therefore | |
| 133 | /// have rewritten every URL on the site. | |
| 134 | #[test] | |
| 135 | fn slug_renames_the_output_page() { | |
| 136 | let (pages, broken) = render_site(&fixtures().join("slugsite")).expect("build site"); | |
| 137 | assert!(broken.is_empty(), "fixture site has no broken links: {broken:?}"); | |
| 138 | let post = pages | |
| 139 | .iter() | |
| 140 | .find(|p| p.source == "2024-02-11-long-source-name.org") | |
| 141 | .expect("post page"); | |
| 142 | assert_eq!( | |
| 143 | post.output, "short-url.html", | |
| 144 | "the slug names the output file, not the source stem" | |
| 145 | ); | |
| 146 | } | |
| 147 | ||
| 148 | /// A link's URL has to follow the target's slug. If resolution kept using source paths, | |
| 149 | /// every cross-page link would point at a file that was never written. | |
| 150 | #[test] | |
| 151 | fn links_resolve_through_the_slug() { | |
| 152 | let (pages, _) = render_site(&fixtures().join("slugsite")).expect("build site"); | |
| 153 | let index = &page(&pages, "index.org").html; | |
| 154 | assert!( | |
| 155 | index.contains("href=\"short-url.html\""), | |
| 156 | "a file: link must target the slugged page:\n{index}" | |
| 157 | ); | |
| 158 | assert!( | |
| 159 | index.contains("href=\"short-url.html#setup\""), | |
| 160 | "a custom-id link must target the slugged page plus the anchor:\n{index}" | |
| 161 | ); | |
| 162 | assert!( | |
| 163 | !index.contains("long-source-name"), | |
| 164 | "no URL may mention the source filename:\n{index}" | |
| 165 | ); | |
| 166 | } | |
| 167 | ||
| 168 | /// A slug is author-controlled text that becomes a path we write to, so traversal has to | |
| 169 | /// be impossible by construction rather than by convention. | |
| 170 | #[test] | |
| 171 | fn slugs_cannot_escape_the_output_directory() { | |
| 172 | use org_ssg::model::Keywords; | |
| 173 | let source = Utf8PathBuf::from("blog/post.org"); | |
| 174 | let slugged = |value: &str| { | |
| 175 | let keywords = Keywords { | |
| 176 | entries: vec![("SLUG".to_string(), value.to_string())], | |
| 177 | }; | |
| 178 | org_ssg::util::output_path(&source, &keywords).to_string() | |
| 179 | }; | |
| 180 | assert_eq!(slugged("../../etc/passwd"), "blog/etc-passwd.html"); | |
| 181 | assert_eq!(slugged("/absolute"), "blog/absolute.html"); | |
| 182 | assert_eq!(slugged(".hidden"), "blog/hidden.html"); | |
| 183 | assert_eq!(slugged("Mixed Case Slug"), "blog/mixed-case-slug.html"); | |
| 184 | // An empty or punctuation-only slug falls back to the source stem rather than | |
| 185 | // producing `.html` with no name at all. | |
| 186 | assert_eq!(slugged("///"), "blog/post.html"); | |
| 187 | } | |
| 188 | ||
| 189 | /// Two pages claiming one URL silently drops a page. With slugs that is a typo away and | |
| 190 | /// invisible in the source filenames, so the build refuses rather than picking a winner. | |
| 191 | #[test] | |
| 192 | fn colliding_slugs_are_a_build_error() { | |
| 193 | let dir = std::env::temp_dir().join(format!("org-ssg-slug-{}", std::process::id())); | |
| 194 | let dir = Utf8PathBuf::from_path_buf(dir).expect("utf-8 temp dir"); | |
| 195 | let _ = std::fs::remove_dir_all(&dir); | |
| 196 | std::fs::create_dir_all(&dir).unwrap(); | |
| 197 | std::fs::write(dir.join("a.org"), "#+TITLE: A\n#+SLUG: same\n").unwrap(); | |
| 198 | std::fs::write(dir.join("b.org"), "#+TITLE: B\n#+SLUG: same\n").unwrap(); | |
| 199 | ||
| 200 | let err = render_site(&dir).expect_err("colliding slugs must fail the build"); | |
| 201 | let message = format!("{err:#}"); | |
| 202 | assert!( | |
| 203 | message.contains("collision") && message.contains("same.html"), | |
| 204 | "the error must name the collision: {message}" | |
| 205 | ); | |
| 206 | std::fs::remove_dir_all(&dir).unwrap(); | |
| 207 | } | |
tests/snapshots/oracle__oracle_blocks.snap added +68
| @@ -0,0 +1,68 @@ | ||
| 1 | --- | |
| 2 | source: tests/oracle.rs | |
| 3 | expression: report | |
| 4 | --- | |
| 5 | agreement: 51/59 skeleton lines (86.4%) | |
| 6 | (- org-ssg, + emacs) | |
| 7 | ||
| 8 | <h1> | |
| 9 | "Quote" | |
| 10 | </h1> | |
| 11 | <blockquote> | |
| 12 | <p> | |
| 13 | "A quoted paragraph with" | |
| 14 | - <em> | |
| 15 | + <i> | |
| 16 | "markup" | |
| 17 | - </em> | |
| 18 | + </i> | |
| 19 | "." | |
| 20 | </p> | |
| 21 | <p> | |
| 22 | "And a second paragraph." | |
| 23 | </p> | |
| 24 | </blockquote> | |
| 25 | <h1> | |
| 26 | "Center" | |
| 27 | </h1> | |
| 28 | <p> | |
| 29 | "Centred text." | |
| 30 | </p> | |
| 31 | <h1> | |
| 32 | "Example" | |
| 33 | </h1> | |
| 34 | <pre> | |
| 35 | "Verbatim *not bold* text. Indentation preserved." | |
| 36 | </pre> | |
| 37 | <h1> | |
| 38 | "Export" | |
| 39 | </h1> | |
| 40 | <aside> | |
| 41 | "Raw HTML passes through." | |
| 42 | </aside> | |
| 43 | <h1> | |
| 44 | "Source" | |
| 45 | </h1> | |
| 46 | <pre> | |
| 47 | - <code> | |
| 48 | "def greet(name): return f\"hello {name}\"" | |
| 49 | - </code> | |
| 50 | </pre> | |
| 51 | <pre> | |
| 52 | - <code> | |
| 53 | "plain block, no language" | |
| 54 | - </code> | |
| 55 | </pre> | |
| 56 | <h1> | |
| 57 | "Nested" | |
| 58 | </h1> | |
| 59 | <blockquote> | |
| 60 | <p> | |
| 61 | "A quote containing a source block:" | |
| 62 | </p> | |
| 63 | <pre> | |
| 64 | - <code> | |
| 65 | "echo hi" | |
| 66 | - </code> | |
| 67 | </pre> | |
| 68 | </blockquote> | |
tests/snapshots/oracle__oracle_core.snap added +70
| @@ -0,0 +1,70 @@ | ||
| 1 | --- | |
| 2 | source: tests/oracle.rs | |
| 3 | expression: report | |
| 4 | --- | |
| 5 | agreement: 45/54 skeleton lines (83.3%) | |
| 6 | (- org-ssg, + emacs) | |
| 7 | ||
| 8 | <p> | |
| 9 | "Intro paragraph with a bare URL" | |
| 10 | <a href="https://example.com"> | |
| 11 | "https://example.com" | |
| 12 | </a> | |
| 13 | "and some" | |
| 14 | <code> | |
| 15 | "inline code" | |
| 16 | </code> | |
| 17 | "." | |
| 18 | </p> | |
| 19 | <h1> | |
| 20 | "Ordered and checked" | |
| 21 | </h1> | |
| 22 | <ol> | |
| 23 | <li> | |
| 24 | "first item" | |
| 25 | </li> | |
| 26 | <li> | |
| 27 | "second item with" | |
| 28 | - <em> | |
| 29 | + <i> | |
| 30 | "emphasis" | |
| 31 | - </em> | |
| 32 | + </i> | |
| 33 | </li> | |
| 34 | - </ol> | |
| 35 | - <ul> | |
| 36 | <li> | |
| 37 | - <input> | |
| 38 | + <code> | |
| 39 | + "[ ]" | |
| 40 | + </code> | |
| 41 | "todo item" | |
| 42 | </li> | |
| 43 | <li> | |
| 44 | - <input> | |
| 45 | + <code> | |
| 46 | + "[X]" | |
| 47 | + </code> | |
| 48 | "done item" | |
| 49 | </li> | |
| 50 | - </ul> | |
| 51 | + </ol> | |
| 52 | <h1> | |
| 53 | "Links and code" | |
| 54 | </h1> | |
| 55 | <p> | |
| 56 | "An external" | |
| 57 | <a href="https://example.org"> | |
| 58 | "site" | |
| 59 | </a> | |
| 60 | "and a bare" | |
| 61 | <a href="https://bare.example"> | |
| 62 | "https://bare.example" | |
| 63 | </a> | |
| 64 | "." | |
| 65 | </p> | |
| 66 | <pre> | |
| 67 | - <code> | |
| 68 | "fn main() { println!(\"hello\"); }" | |
| 69 | - </code> | |
| 70 | </pre> | |
tests/snapshots/oracle__oracle_elements.snap added +103
| @@ -0,0 +1,103 @@ | ||
| 1 | --- | |
| 2 | source: tests/oracle.rs | |
| 3 | expression: report | |
| 4 | --- | |
| 5 | agreement: 64/82 skeleton lines (78.0%) | |
| 6 | (- org-ssg, + emacs) | |
| 7 | ||
| 8 | <h1> | |
| 9 | "Code and tables" | |
| 10 | </h1> | |
| 11 | <pre> | |
| 12 | - <code> | |
| 13 | "fn main() { println!(\"hello\"); }" | |
| 14 | - </code> | |
| 15 | </pre> | |
| 16 | <table> | |
| 17 | + <colgroup> | |
| 18 | + <col> | |
| 19 | + <col> | |
| 20 | + </colgroup> | |
| 21 | <thead> | |
| 22 | <tr> | |
| 23 | <th> | |
| 24 | "Name" | |
| 25 | </th> | |
| 26 | <th> | |
| 27 | "Score" | |
| 28 | </th> | |
| 29 | </tr> | |
| 30 | </thead> | |
| 31 | <tbody> | |
| 32 | <tr> | |
| 33 | <td> | |
| 34 | "alpha" | |
| 35 | </td> | |
| 36 | <td> | |
| 37 | "10" | |
| 38 | </td> | |
| 39 | </tr> | |
| 40 | <tr> | |
| 41 | <td> | |
| 42 | "beta" | |
| 43 | </td> | |
| 44 | <td> | |
| 45 | "20" | |
| 46 | </td> | |
| 47 | </tr> | |
| 48 | </tbody> | |
| 49 | </table> | |
| 50 | <h1> | |
| 51 | "Links and footnotes" | |
| 52 | </h1> | |
| 53 | <p> | |
| 54 | "An external link:" | |
| 55 | <a href="https://example.com"> | |
| 56 | "Example" | |
| 57 | </a> | |
| 58 | - "and an id link" | |
| 59 | - <a href="#abc-123"> | |
| 60 | - "abc-123" | |
| 61 | - </a> | |
| 62 | - "." | |
| 63 | + "and an id link ." | |
| 64 | </p> | |
| 65 | <p> | |
| 66 | "Text with a footnote reference." | |
| 67 | <sup> | |
| 68 | - <a href="#fn-1"> | |
| 69 | + <a href="#fn.1"> | |
| 70 | "1" | |
| 71 | </a> | |
| 72 | </sup> | |
| 73 | </p> | |
| 74 | <h1> | |
| 75 | "Blocks" | |
| 76 | </h1> | |
| 77 | <blockquote> | |
| 78 | <p> | |
| 79 | "A quoted paragraph." | |
| 80 | </p> | |
| 81 | </blockquote> | |
| 82 | <hr> | |
| 83 | - <section> | |
| 84 | - <hr> | |
| 85 | - <ol> | |
| 86 | - <li> | |
| 87 | + <h2> | |
| 88 | + "Footnotes:" | |
| 89 | + </h2> | |
| 90 | + <sup> | |
| 91 | + <a href="#fnr.1"> | |
| 92 | + "1" | |
| 93 | + </a> | |
| 94 | + </sup> | |
| 95 | <p> | |
| 96 | "The footnote definition." | |
| 97 | </p> | |
| 98 | - <a href="#fnr-1"> | |
| 99 | - "↩" | |
| 100 | - </a> | |
| 101 | - </li> | |
| 102 | - </ol> | |
| 103 | - </section> | |
tests/snapshots/oracle__oracle_footnote.snap added +83
| @@ -0,0 +1,83 @@ | ||
| 1 | --- | |
| 2 | source: tests/oracle.rs | |
| 3 | expression: report | |
| 4 | --- | |
| 5 | agreement: 30/53 skeleton lines (56.6%) | |
| 6 | (- org-ssg, + emacs) | |
| 7 | ||
| 8 | <p> | |
| 9 | "Text with a reference." | |
| 10 | <sup> | |
| 11 | - <a href="#fn-1"> | |
| 12 | + <a href="#fn.1"> | |
| 13 | "1" | |
| 14 | </a> | |
| 15 | </sup> | |
| 16 | "And a second one." | |
| 17 | <sup> | |
| 18 | - <a href="#fn-2"> | |
| 19 | + <a href="#fn.2"> | |
| 20 | "2" | |
| 21 | </a> | |
| 22 | </sup> | |
| 23 | </p> | |
| 24 | <p> | |
| 25 | "An inline footnote." | |
| 26 | <sup> | |
| 27 | - <a href="#fn-3"> | |
| 28 | + <a href="#fn.3"> | |
| 29 | "3" | |
| 30 | </a> | |
| 31 | </sup> | |
| 32 | </p> | |
| 33 | - <section> | |
| 34 | - <hr> | |
| 35 | - <ol> | |
| 36 | - <li> | |
| 37 | + <h2> | |
| 38 | + "Footnotes:" | |
| 39 | + </h2> | |
| 40 | + <sup> | |
| 41 | + <a href="#fnr.1"> | |
| 42 | + "1" | |
| 43 | + </a> | |
| 44 | + </sup> | |
| 45 | <p> | |
| 46 | "The first definition." | |
| 47 | </p> | |
| 48 | - <a href="#fnr-1"> | |
| 49 | - "↩" | |
| 50 | + <sup> | |
| 51 | + <a href="#fnr.2"> | |
| 52 | + "2" | |
| 53 | </a> | |
| 54 | - </li> | |
| 55 | - <li> | |
| 56 | + </sup> | |
| 57 | <p> | |
| 58 | "The second definition, with" | |
| 59 | - <em> | |
| 60 | + <i> | |
| 61 | "emphasis" | |
| 62 | - </em> | |
| 63 | + </i> | |
| 64 | "." | |
| 65 | </p> | |
| 66 | - <a href="#fnr-2"> | |
| 67 | - "↩" | |
| 68 | + <sup> | |
| 69 | + <a href="#fnr.3"> | |
| 70 | + "3" | |
| 71 | </a> | |
| 72 | - </li> | |
| 73 | - <li> | |
| 74 | + </sup> | |
| 75 | + <p> | |
| 76 | "defined right here" | |
| 77 | - <a href="#fnr-3"> | |
| 78 | - "↩" | |
| 79 | - </a> | |
| 80 | - </li> | |
| 81 | - </ol> | |
| 82 | - </section> | |
| 83 | + </p> | |
tests/snapshots/oracle__oracle_headings.snap added +39
| @@ -0,0 +1,39 @@ | ||
| 1 | --- | |
| 2 | source: tests/oracle.rs | |
| 3 | expression: report | |
| 4 | --- | |
| 5 | agreement: 28/30 skeleton lines (93.3%) | |
| 6 | (- org-ssg, + emacs) | |
| 7 | ||
| 8 | <h1> | |
| 9 | - "TODO [#A] Write the parser work rust" | |
| 10 | + "TODO Write the parser work rust" | |
| 11 | </h1> | |
| 12 | <p> | |
| 13 | "A heading carrying a keyword, a priority, tags and a property drawer." | |
| 14 | </p> | |
| 15 | <h2> | |
| 16 | "DONE Nested and finished" | |
| 17 | </h2> | |
| 18 | <p> | |
| 19 | "Sub-headings nest by star count." | |
| 20 | </p> | |
| 21 | <h2> | |
| 22 | - "[#C] Priority without a keyword" | |
| 23 | + "Priority without a keyword" | |
| 24 | </h2> | |
| 25 | <p> | |
| 26 | "A priority cookie can stand alone." | |
| 27 | </p> | |
| 28 | <h1> | |
| 29 | "TODOs are not a keyword" | |
| 30 | </h1> | |
| 31 | <p> | |
| 32 | "The word boundary matters: this heading has no TODO keyword." | |
| 33 | </p> | |
| 34 | <h1> | |
| 35 | "DONE" | |
| 36 | </h1> | |
| 37 | <p> | |
| 38 | "A keyword with no title at all." | |
| 39 | </p> | |
tests/snapshots/oracle__oracle_images.snap added +63
| @@ -0,0 +1,63 @@ | ||
| 1 | --- | |
| 2 | source: tests/oracle.rs | |
| 3 | expression: report | |
| 4 | --- | |
| 5 | agreement: 28/42 skeleton lines (66.7%) | |
| 6 | (- org-ssg, + emacs) | |
| 7 | ||
| 8 | <h1> | |
| 9 | "Bare image" | |
| 10 | </h1> | |
| 11 | <p> | |
| 12 | <img src="diagram.png"> | |
| 13 | </p> | |
| 14 | <h1> | |
| 15 | "Captioned figure" | |
| 16 | </h1> | |
| 17 | - <figure> | |
| 18 | + <p> | |
| 19 | <img src="pipeline.svg"> | |
| 20 | - <figcaption> | |
| 21 | - "The pipeline, end to end" | |
| 22 | - </figcaption> | |
| 23 | - </figure> | |
| 24 | + </p> | |
| 25 | + <p> | |
| 26 | + "Figure 1: The pipeline, end to end" | |
| 27 | + </p> | |
| 28 | <h1> | |
| 29 | "Caption with markup" | |
| 30 | </h1> | |
| 31 | - <figure> | |
| 32 | + <p> | |
| 33 | <img src="chart.png"> | |
| 34 | - <figcaption> | |
| 35 | - "A" | |
| 36 | - <em> | |
| 37 | + </p> | |
| 38 | + <p> | |
| 39 | + "Figure 2: A" | |
| 40 | + <i> | |
| 41 | "stylised" | |
| 42 | - </em> | |
| 43 | + </i> | |
| 44 | "chart" | |
| 45 | - </figcaption> | |
| 46 | - </figure> | |
| 47 | + </p> | |
| 48 | <h1> | |
| 49 | "Quoted attribute values" | |
| 50 | </h1> | |
| 51 | - <figure> | |
| 52 | + <p> | |
| 53 | <img src="cat.jpg"> | |
| 54 | - </figure> | |
| 55 | + </p> | |
| 56 | <h1> | |
| 57 | "Image with a description is a link" | |
| 58 | </h1> | |
| 59 | <p> | |
| 60 | <a href="diagram.png"> | |
| 61 | "the diagram" | |
| 62 | </a> | |
| 63 | </p> | |
tests/snapshots/oracle__oracle_lists.snap added +123
| @@ -0,0 +1,123 @@ | ||
| 1 | --- | |
| 2 | source: tests/oracle.rs | |
| 3 | expression: report | |
| 4 | --- | |
| 5 | agreement: 100/111 skeleton lines (90.1%) | |
| 6 | (- org-ssg, + emacs) | |
| 7 | ||
| 8 | <h1> | |
| 9 | "Nesting" | |
| 10 | </h1> | |
| 11 | <ul> | |
| 12 | <li> | |
| 13 | "outer item" | |
| 14 | <ul> | |
| 15 | <li> | |
| 16 | "inner item" | |
| 17 | <ul> | |
| 18 | <li> | |
| 19 | "deepest item" | |
| 20 | </li> | |
| 21 | </ul> | |
| 22 | </li> | |
| 23 | <li> | |
| 24 | "second inner" | |
| 25 | </li> | |
| 26 | </ul> | |
| 27 | </li> | |
| 28 | <li> | |
| 29 | "second outer" | |
| 30 | </li> | |
| 31 | </ul> | |
| 32 | <h1> | |
| 33 | "Ordered" | |
| 34 | </h1> | |
| 35 | <ol> | |
| 36 | <li> | |
| 37 | "first" | |
| 38 | </li> | |
| 39 | <li> | |
| 40 | "second" | |
| 41 | <ol> | |
| 42 | <li> | |
| 43 | "second point one" | |
| 44 | </li> | |
| 45 | <li> | |
| 46 | "second point two" | |
| 47 | </li> | |
| 48 | </ol> | |
| 49 | </li> | |
| 50 | <li> | |
| 51 | "third" | |
| 52 | </li> | |
| 53 | </ol> | |
| 54 | <h1> | |
| 55 | "Checkboxes" | |
| 56 | </h1> | |
| 57 | <ul> | |
| 58 | <li> | |
| 59 | - <input> | |
| 60 | + <code> | |
| 61 | + "[ ]" | |
| 62 | + </code> | |
| 63 | "not done" | |
| 64 | </li> | |
| 65 | <li> | |
| 66 | - <input> | |
| 67 | + <code> | |
| 68 | + "[X]" | |
| 69 | + </code> | |
| 70 | "done" | |
| 71 | </li> | |
| 72 | <li> | |
| 73 | - <input> | |
| 74 | + <code> | |
| 75 | + "[-]" | |
| 76 | + </code> | |
| 77 | "partially done" | |
| 78 | </li> | |
| 79 | </ul> | |
| 80 | <h1> | |
| 81 | "Description" | |
| 82 | </h1> | |
| 83 | <dl> | |
| 84 | <dt> | |
| 85 | "term one" | |
| 86 | </dt> | |
| 87 | <dd> | |
| 88 | "the first definition" | |
| 89 | </dd> | |
| 90 | <dt> | |
| 91 | "term two" | |
| 92 | </dt> | |
| 93 | <dd> | |
| 94 | "the second definition, which is soft-wrapped across two lines" | |
| 95 | </dd> | |
| 96 | <dt> | |
| 97 | - <em> | |
| 98 | + <i> | |
| 99 | "marked up" | |
| 100 | - </em> | |
| 101 | + </i> | |
| 102 | "term" | |
| 103 | </dt> | |
| 104 | <dd> | |
| 105 | "definitions hold inline markup" | |
| 106 | </dd> | |
| 107 | </dl> | |
| 108 | <h1> | |
| 109 | "Multi-paragraph items" | |
| 110 | </h1> | |
| 111 | <ul> | |
| 112 | <li> | |
| 113 | <p> | |
| 114 | "an item whose body has two paragraphs" | |
| 115 | </p> | |
| 116 | <p> | |
| 117 | "the second paragraph, indented under the bullet" | |
| 118 | </p> | |
| 119 | </li> | |
| 120 | <li> | |
| 121 | "a plain sibling" | |
| 122 | </li> | |
| 123 | </ul> | |
tests/snapshots/oracle__oracle_minimal.snap added +53
| @@ -0,0 +1,53 @@ | ||
| 1 | --- | |
| 2 | source: tests/oracle.rs | |
| 3 | expression: report | |
| 4 | --- | |
| 5 | agreement: 38/42 skeleton lines (90.5%) | |
| 6 | (- org-ssg, + emacs) | |
| 7 | ||
| 8 | <p> | |
| 9 | "A single paragraph of preamble text before any heading." | |
| 10 | </p> | |
| 11 | <h1> | |
| 12 | "First Heading" | |
| 13 | </h1> | |
| 14 | <p> | |
| 15 | "Some body text with" | |
| 16 | - <strong> | |
| 17 | + <b> | |
| 18 | "bold" | |
| 19 | - </strong> | |
| 20 | + </b> | |
| 21 | "," | |
| 22 | - <em> | |
| 23 | + <i> | |
| 24 | "italic" | |
| 25 | - </em> | |
| 26 | + </i> | |
| 27 | ", and" | |
| 28 | <code> | |
| 29 | "verbatim" | |
| 30 | </code> | |
| 31 | "." | |
| 32 | </p> | |
| 33 | <h2> | |
| 34 | "A Subheading tag1 tag2" | |
| 35 | </h2> | |
| 36 | <ul> | |
| 37 | <li> | |
| 38 | "an unordered item" | |
| 39 | </li> | |
| 40 | <li> | |
| 41 | "another with a checkbox [ ]" | |
| 42 | </li> | |
| 43 | </ul> | |
| 44 | <h1> | |
| 45 | "Second Heading" | |
| 46 | </h1> | |
| 47 | <p> | |
| 48 | "See" | |
| 49 | <a href="#first"> | |
| 50 | "the first heading" | |
| 51 | </a> | |
| 52 | "." | |
| 53 | </p> | |
tests/snapshots/oracle__oracle_table.snap added +41
| @@ -0,0 +1,41 @@ | ||
| 1 | --- | |
| 2 | source: tests/oracle.rs | |
| 3 | expression: report | |
| 4 | --- | |
| 5 | agreement: 30/34 skeleton lines (88.2%) | |
| 6 | (- org-ssg, + emacs) | |
| 7 | ||
| 8 | <table> | |
| 9 | + <colgroup> | |
| 10 | + <col> | |
| 11 | + <col> | |
| 12 | + </colgroup> | |
| 13 | <thead> | |
| 14 | <tr> | |
| 15 | <th> | |
| 16 | "Name" | |
| 17 | </th> | |
| 18 | <th> | |
| 19 | "Score" | |
| 20 | </th> | |
| 21 | </tr> | |
| 22 | </thead> | |
| 23 | <tbody> | |
| 24 | <tr> | |
| 25 | <td> | |
| 26 | "alpha" | |
| 27 | </td> | |
| 28 | <td> | |
| 29 | "10" | |
| 30 | </td> | |
| 31 | </tr> | |
| 32 | <tr> | |
| 33 | <td> | |
| 34 | "beta" | |
| 35 | </td> | |
| 36 | <td> | |
| 37 | "20" | |
| 38 | </td> | |
| 39 | </tr> | |
| 40 | </tbody> | |
| 41 | </table> | |
tests/snapshots/oracle__oracle_timestamps.snap added +74
| @@ -0,0 +1,74 @@ | ||
| 1 | --- | |
| 2 | source: tests/oracle.rs | |
| 3 | expression: report | |
| 4 | --- | |
| 5 | agreement: 25/62 skeleton lines (40.3%) | |
| 6 | (- org-ssg, + emacs) | |
| 7 | ||
| 8 | <h1> | |
| 9 | "Single" | |
| 10 | </h1> | |
| 11 | <p> | |
| 12 | - "An active date" | |
| 13 | - <time> | |
| 14 | - "2024-01-15" | |
| 15 | - </time> | |
| 16 | - "and an inactive one" | |
| 17 | - <time> | |
| 18 | - "2024-01-15" | |
| 19 | - </time> | |
| 20 | - "." | |
| 21 | + "An active date <2024-01-15 Mon> and an inactive one [2024-01-15 Mon]." | |
| 22 | </p> | |
| 23 | <p> | |
| 24 | - "With a time:" | |
| 25 | - <time> | |
| 26 | - "2024-01-15 10:30" | |
| 27 | - </time> | |
| 28 | - "." | |
| 29 | + "With a time: <2024-01-15 Mon 10:30>." | |
| 30 | </p> | |
| 31 | <h1> | |
| 32 | "Ranges" | |
| 33 | </h1> | |
| 34 | <p> | |
| 35 | - "A same-day time range" | |
| 36 | - <time> | |
| 37 | - "2024-01-15 10:00" | |
| 38 | - </time> | |
| 39 | - "–" | |
| 40 | - <time> | |
| 41 | - "11:45" | |
| 42 | - </time> | |
| 43 | - "." | |
| 44 | + "A same-day time range <2024-01-15 Mon 10:00-11:45>." | |
| 45 | </p> | |
| 46 | <p> | |
| 47 | - "A multi-day range" | |
| 48 | - <time> | |
| 49 | - "2024-01-15" | |
| 50 | - </time> | |
| 51 | - "–" | |
| 52 | - <time> | |
| 53 | - "2024-01-20" | |
| 54 | - </time> | |
| 55 | - "." | |
| 56 | + "A multi-day range <2024-01-15 Mon>–<2024-01-20 Sat>." | |
| 57 | </p> | |
| 58 | <h1> | |
| 59 | "Ignored decorations" | |
| 60 | </h1> | |
| 61 | <p> | |
| 62 | - "A repeater is dropped:" | |
| 63 | - <time> | |
| 64 | - "2024-01-15" | |
| 65 | - </time> | |
| 66 | - "." | |
| 67 | + "A repeater is dropped: <2024-01-15 Mon +1w>." | |
| 68 | </p> | |
| 69 | <h1> | |
| 70 | "Not timestamps" | |
| 71 | </h1> | |
| 72 | <p> | |
| 73 | "Comparisons like 3 < 4 and [not a stamp] stay literal text." | |
| 74 | </p> | |