docs/plans/2026-10-04-orgcore-parser.md
1600 lines · 55657 bytes
11 symbols in this file
OrgCore Parser Foundation Implementation PlanGlobal ConstraintsOut of scope for this planFile structureTask 1: Package and SourceTextTask 2: Green and red syntax treeTask 3: Line splitting and classificationTask 4: In-buffer settingsTask 5: Parser — document, sections and headingsTask 6: Parser — elementsTask 7: Round-trip fuzz, encoding fixtures and corpus test
1# OrgCore Parser Foundation Implementation Plan
2
3> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
4
5**Goal:** A Swift package, `OrgCore`, that turns org file bytes into a lossless syntax tree of block-level elements, with the byte and encoding contract from the design.
6
7**Architecture:** Bytes decode into `SourceText` (UTF-8 only, BOM kept, invalid input read-only). The parser splits lines without normalizing endings, classifies each line once, matches block and drawer ends, scans in-buffer TODO settings outside blocks, then builds an immutable green tree through `GreenBuilder`. `SyntaxNode` gives offset-aware red views. Every input byte ends up in exactly one token, so the tree text always equals the source.
8
9**Tech Stack:** Swift 6.2 tools, Swift Testing, Foundation only. No third-party dependencies.
10
11**Spec:** `docs/design.md` (sections "OrgCore data model", "Testing", "Phases").
12
13## Global Constraints
14
15- Platforms: macOS 26, iOS 26.
16- `OrgCore` imports Foundation only; never AppKit, UIKit or SwiftUI.
17- Ranges and offsets exposed by the tree are UTF-16 code units.
18- Supported encoding: UTF-8 with or without BOM. Anything else is read-only and never converted.
19- Round trip: tree text equals source text for every input, and encoding unchanged text returns the original bytes.
20- License 0BSD. No attribution lines in code, commits or docs.
21
22## Out of scope for this plan
23
24Inline objects (emphasis, links, timestamps), the semantic layer, incremental reparse, conformance rendering, and the private-corpus benchmarks. Each gets its own plan.
25
26## File structure
27
28| File | Responsibility |
29| --- | --- |
30| `Package.swift` | Package with `OrgCore` library and `OrgCoreTests` |
31| `Sources/OrgCore/SourceText.swift` | Byte decoding, BOM, validity, encoding back to bytes |
32| `Sources/OrgCore/Syntax/SyntaxKind.swift` | Token and node kinds |
33| `Sources/OrgCore/Syntax/GreenTree.swift` | `GreenToken`, `GreenNode`, `GreenElement`, `GreenBuilder` |
34| `Sources/OrgCore/Syntax/SyntaxNode.swift` | Red nodes, tokens, `OrgTree` |
35| `Sources/OrgCore/Parser/Lines.swift` | Line splitting and classification |
36| `Sources/OrgCore/Parser/Settings.swift` | TODO sequences, priorities, settings scan |
37| `Sources/OrgCore/Parser/Parser.swift` | Tree construction |
38| `Tests/OrgCoreTests/*.swift` | One test file per source file, plus round-trip fuzz and corpus tests |
39
40---
41
42### Task 1: Package and SourceText
43
44**Files:**
45- Create: `Package.swift`
46- Create: `Sources/OrgCore/SourceText.swift`
47- Test: `Tests/OrgCoreTests/SourceTextTests.swift`
48
49**Interfaces:**
50- Produces: `SourceText(bytes: [UInt8])`, `SourceText(_ text: String)`, properties `originalBytes`, `hasBOM`, `isValidUTF8`, `isEditable`, `text`; `encode(_ newText: String) -> [UInt8]`.
51
52- [ ] **Step 1: Create the package**
53
54```swift
55// swift-tools-version: 6.2
56import PackageDescription
57
58let package = Package(
59 name: "Orgstar",
60 platforms: [.macOS(.v26), .iOS(.v26)],
61 products: [
62 .library(name: "OrgCore", targets: ["OrgCore"])
63 ],
64 targets: [
65 .target(name: "OrgCore"),
66 .testTarget(name: "OrgCoreTests", dependencies: ["OrgCore"])
67 ]
68)
69```
70
71- [ ] **Step 2: Write the failing tests**
72
73```swift
74import Testing
75@testable import OrgCore
76
77struct SourceTextTests {
78 @Test func plainUTF8() {
79 let source = SourceText(bytes: Array("* a\n".utf8))
80 #expect(source.text == "* a\n")
81 #expect(!source.hasBOM)
82 #expect(source.isEditable)
83 }
84
85 @Test func bomIsStrippedAndRestored() {
86 let bytes: [UInt8] = [0xEF, 0xBB, 0xBF] + Array("x\n".utf8)
87 let source = SourceText(bytes: bytes)
88 #expect(source.hasBOM)
89 #expect(source.text == "x\n")
90 #expect(source.encode(source.text) == bytes)
91 #expect(source.encode("y\n") == [0xEF, 0xBB, 0xBF] + Array("y\n".utf8))
92 }
93
94 @Test func crlfAndMixedEndingsSurvive() {
95 let bytes = Array("a\r\nb\nc\r\n".utf8)
96 let source = SourceText(bytes: bytes)
97 #expect(source.encode(source.text) == bytes)
98 }
99
100 @Test func invalidUTF8IsReadOnlyAndUnchanged() {
101 let bytes: [UInt8] = [0x61, 0xFF, 0x0A]
102 let source = SourceText(bytes: bytes)
103 #expect(!source.isValidUTF8)
104 #expect(!source.isEditable)
105 #expect(source.encode(source.text) == bytes)
106 }
107
108 @Test func nonBMPAndCombiningSurvive() {
109 let bytes = Array("😀 e\u{301}\n".utf8)
110 let source = SourceText(bytes: bytes)
111 #expect(source.encode(source.text) == bytes)
112 }
113}
114```
115
116- [ ] **Step 3: Run tests to verify they fail**
117
118Run: `swift test --filter SourceTextTests`
119Expected: build failure, `cannot find 'SourceText' in scope`.
120
121- [ ] **Step 4: Implement**
122
123```swift
124import Foundation
125
126/// A file's bytes and their decoded text. Only UTF-8 (with or without a BOM) is editable;
127/// anything else decodes with replacement characters for display and is never written back.
128public struct SourceText: Sendable {
129 public let originalBytes: [UInt8]
130 public let hasBOM: Bool
131 public let isValidUTF8: Bool
132 /// Decoded text without the BOM.
133 public let text: String
134
135 private static let bom: [UInt8] = [0xEF, 0xBB, 0xBF]
136
137 public init(bytes: [UInt8]) {
138 originalBytes = bytes
139 hasBOM = bytes.starts(with: Self.bom)
140 let body = hasBOM ? Array(bytes.dropFirst(3)) : bytes
141 if let decoded = String(validating: body, as: UTF8.self) {
142 text = decoded
143 isValidUTF8 = true
144 } else {
145 text = String(decoding: body, as: UTF8.self)
146 isValidUTF8 = false
147 }
148 }
149
150 public init(_ text: String) {
151 self.init(bytes: Array(text.utf8))
152 }
153
154 public var isEditable: Bool { isValidUTF8 }
155
156 /// Bytes to write for `newText`. Unchanged text returns the original bytes. Valid UTF-8
157 /// round-trips through `String` unchanged, so untouched spans keep their exact bytes.
158 public func encode(_ newText: String) -> [UInt8] {
159 if newText == text { return originalBytes }
160 precondition(isEditable, "a read-only document cannot be re-encoded")
161 return (hasBOM ? Self.bom : []) + Array(newText.utf8)
162 }
163}
164```
165
166- [ ] **Step 5: Run tests to verify they pass**
167
168Run: `swift test --filter SourceTextTests`
169Expected: 5 tests pass.
170
171- [ ] **Step 6: Commit**
172
173```bash
174git add Package.swift Sources Tests
175git commit -m "Add OrgCore package and SourceText byte contract"
176```
177
178---
179
180### Task 2: Green and red syntax tree
181
182**Files:**
183- Create: `Sources/OrgCore/Syntax/SyntaxKind.swift`
184- Create: `Sources/OrgCore/Syntax/GreenTree.swift`
185- Create: `Sources/OrgCore/Syntax/SyntaxNode.swift`
186- Test: `Tests/OrgCoreTests/SyntaxTreeTests.swift`
187
188**Interfaces:**
189- Produces: `SyntaxKind` (enum, `String` raw values), `GreenToken(kind:text:)`, `GreenNode(kind:children:)` with `.length`, `.text`; `GreenElement` (`.node`, `.token`); internal `GreenBuilder` with `start(_:)`, `token(_:_:)`, `finish()`, `build()`; `SyntaxNode` with `kind`, `range`, `text`, `children`, `tokens`, `descendants()`; `SyntaxToken`. `OrgTree` depends on `OrgSettings` (Task 4), so it is added in Task 5.
190
191- [ ] **Step 1: Write the failing tests**
192
193```swift
194import Testing
195@testable import OrgCore
196
197struct SyntaxTreeTests {
198 func sample() -> GreenNode {
199 var b = GreenBuilder()
200 b.start(.document)
201 b.start(.paragraph)
202 b.token(.text, "hé😀")
203 b.token(.newline, "\n")
204 b.finish()
205 b.token(.newline, "\r\n")
206 b.finish()
207 return b.build()
208 }
209
210 @Test func lengthsAreUTF16() {
211 let green = sample()
212 #expect(green.length == 4 + 1 + 2)
213 #expect(green.text == "hé😀\n\r\n")
214 }
215
216 @Test func redNodesCarryOffsets() {
217 let root = SyntaxNode(green: sample(), offset: 0, parent: nil)
218 let paragraph = root.children[0]
219 #expect(paragraph.kind == .paragraph)
220 #expect(paragraph.range == 0..<5)
221 #expect(paragraph.parent === root)
222 #expect(root.tokens.map(\.range) == [5..<7])
223 #expect(paragraph.tokens.map(\.kind) == [.text, .newline])
224 }
225
226 @Test func builderSkipsEmptyTokens() {
227 var b = GreenBuilder()
228 b.start(.document)
229 b.token(.whitespace, "")
230 b.finish()
231 #expect(b.build().children.isEmpty)
232 }
233
234 @Test func descendantsArePreorder() {
235 let root = SyntaxNode(green: sample(), offset: 0, parent: nil)
236 #expect(root.descendants().map(\.kind) == [.document, .paragraph])
237 }
238}
239```
240
241- [ ] **Step 2: Run tests to verify they fail**
242
243Run: `swift test --filter SyntaxTreeTests`
244Expected: build failure, `cannot find 'GreenBuilder' in scope`.
245
246- [ ] **Step 3: Implement `SyntaxKind.swift`**
247
248```swift
249public enum SyntaxKind: String, Sendable {
250 // Tokens
251 case text, newline, whitespace
252 case stars, todoKeyword, priority, title, tags
253
254 // Nodes
255 case document, zerothSection, section, heading
256 case planning, propertyDrawer, nodeProperty, drawer, clock
257 case paragraph, plainList, item, table, tableRow, tableFormula
258 case block, dynamicBlock, keyword, affiliatedKeyword
259 case comment, fixedWidth, horizontalRule, footnoteDefinition
260}
261```
262
263- [ ] **Step 4: Implement `GreenTree.swift`**
264
265```swift
266public struct GreenToken: Sendable, Equatable {
267 public let kind: SyntaxKind
268 public let text: String
269 /// Length in UTF-16 code units.
270 public let length: Int
271
272 public init(kind: SyntaxKind, text: String) {
273 self.kind = kind
274 self.text = text
275 self.length = text.utf16.count
276 }
277}
278
279/// An immutable node. Stores only kind, children and length, so unchanged subtrees can be
280/// shared between versions of a document.
281public final class GreenNode: Sendable, Equatable {
282 public let kind: SyntaxKind
283 public let children: [GreenElement]
284 /// Length in UTF-16 code units.
285 public let length: Int
286
287 public init(kind: SyntaxKind, children: [GreenElement]) {
288 self.kind = kind
289 self.children = children
290 self.length = children.reduce(0) { $0 + $1.length }
291 }
292
293 public static func == (lhs: GreenNode, rhs: GreenNode) -> Bool {
294 lhs === rhs || (lhs.kind == rhs.kind && lhs.children == rhs.children)
295 }
296
297 public var text: String {
298 var out = ""
299 write(to: &out)
300 return out
301 }
302
303 func write(to out: inout String) {
304 for child in children {
305 switch child {
306 case .node(let node): node.write(to: &out)
307 case .token(let token): out += token.text
308 }
309 }
310 }
311}
312
313public enum GreenElement: Sendable, Equatable {
314 case node(GreenNode)
315 case token(GreenToken)
316
317 public var length: Int {
318 switch self {
319 case .node(let node): node.length
320 case .token(let token): token.length
321 }
322 }
323}
324
325struct GreenBuilder {
326 private var stack: [(kind: SyntaxKind, children: [GreenElement])] = []
327 private var root: GreenNode?
328
329 mutating func start(_ kind: SyntaxKind) {
330 stack.append((kind, []))
331 }
332
333 mutating func token(_ kind: SyntaxKind, _ text: some StringProtocol) {
334 guard !text.isEmpty else { return }
335 stack[stack.count - 1].children.append(.token(GreenToken(kind: kind, text: String(text))))
336 }
337
338 mutating func finish() {
339 let (kind, children) = stack.removeLast()
340 let node = GreenNode(kind: kind, children: children)
341 if stack.isEmpty {
342 root = node
343 } else {
344 stack[stack.count - 1].children.append(.node(node))
345 }
346 }
347
348 func build() -> GreenNode {
349 precondition(stack.isEmpty, "unfinished nodes")
350 return root!
351 }
352}
353```
354
355- [ ] **Step 5: Implement `SyntaxNode.swift`**
356
357```swift
358/// A view of a green node at an absolute offset, with a parent link. Created on demand.
359public final class SyntaxNode: Sendable {
360 public let green: GreenNode
361 public let offset: Int
362 public let parent: SyntaxNode?
363
364 init(green: GreenNode, offset: Int, parent: SyntaxNode?) {
365 self.green = green
366 self.offset = offset
367 self.parent = parent
368 }
369
370 public var kind: SyntaxKind { green.kind }
371 public var range: Range<Int> { offset..<(offset + green.length) }
372 public var text: String { green.text }
373
374 public var children: [SyntaxNode] {
375 var result: [SyntaxNode] = []
376 var at = offset
377 for child in green.children {
378 if case .node(let node) = child {
379 result.append(SyntaxNode(green: node, offset: at, parent: self))
380 }
381 at += child.length
382 }
383 return result
384 }
385
386 public var tokens: [SyntaxToken] {
387 var result: [SyntaxToken] = []
388 var at = offset
389 for child in green.children {
390 if case .token(let token) = child {
391 result.append(SyntaxToken(kind: token.kind, text: token.text, range: at..<(at + token.length)))
392 }
393 at += child.length
394 }
395 return result
396 }
397
398 /// This node and every node below it, in document order.
399 public func descendants() -> [SyntaxNode] {
400 [self] + children.flatMap { $0.descendants() }
401 }
402}
403
404public struct SyntaxToken: Sendable, Equatable {
405 public let kind: SyntaxKind
406 public let text: String
407 public let range: Range<Int>
408}
409```
410
411- [ ] **Step 6: Run tests to verify they pass**
412
413Run: `swift test --filter SyntaxTreeTests`
414Expected: 4 tests pass.
415
416- [ ] **Step 7: Commit**
417
418```bash
419git add Sources Tests
420git commit -m "Add green and red syntax tree"
421```
422
423---
424
425### Task 3: Line splitting and classification
426
427**Files:**
428- Create: `Sources/OrgCore/Parser/Lines.swift`
429- Test: `Tests/OrgCoreTests/LinesTests.swift`
430
431**Interfaces:**
432- Produces (internal): `RawLine(content: Substring, ending: Substring)`, `splitRawLines(_ text: String) -> [RawLine]`, `LineClass` enum, `ClassifiedLine(cls:indent:)`, `classifyLine(_ line: Substring) -> ClassifiedLine`, `Substring.trimmingTrailingWhitespace`.
433
434- [ ] **Step 1: Write the failing tests**
435
436```swift
437import Testing
438@testable import OrgCore
439
440struct LinesTests {
441 @Test func splitKeepsEveryEnding() {
442 let lines = splitRawLines("a\r\nb\n\nc")
443 #expect(lines.map { String($0.content) } == ["a", "b", "", "c"])
444 #expect(lines.map { String($0.ending) } == ["\r\n", "\n", "\n", ""])
445 #expect(splitRawLines("").isEmpty)
446 #expect(splitRawLines("x\n").count == 1)
447 }
448
449 @Test func splitIsLossless() {
450 let text = "\r\n\n a\r b\r\n😀\n"
451 #expect(splitRawLines(text).map { String($0.content) + String($0.ending) }.joined() == text)
452 }
453
454 @Test(arguments: [
455 ("", LineClass.blank),
456 (" \t", .blank),
457 ("* a", .heading(level: 1)),
458 ("*** ", .heading(level: 3)),
459 ("*", .heading(level: 1)),
460 ("*bold* text", .plain),
461 (" * a", .listItem),
462 ("#+BEGIN_SRC sh :results output", .blockBegin(name: "src")),
463 ("#+end_src", .blockEnd(name: "src")),
464 ("#+BEGIN: clocktable :scope file", .dynamicBegin),
465 ("#+END:", .dynamicEnd),
466 ("#+TITLE: x", .keyword(key: "TITLE")),
467 ("#+tblfm: $2=$1", .keyword(key: "TBLFM")),
468 ("# comment", .comment),
469 ("#", .comment),
470 ("#hashtag", .plain),
471 (": fixed", .fixedWidth),
472 (":", .fixedWidth),
473 (":PROPERTIES:", .drawerBegin(name: "PROPERTIES")),
474 (" :LOGBOOK:", .drawerBegin(name: "LOGBOOK")),
475 (":END:", .drawerEnd),
476 ("| a | b |", .tableRow),
477 ("-----", .horizontalRule),
478 ("----", .plain),
479 ("[fn:1] note", .footnoteDefinition),
480 ("CLOCK: [2026-10-04 Sun 10:00]", .clock),
481 ("SCHEDULED: <2026-10-04 Sun>", .planning),
482 ("- item", .listItem),
483 ("+ item", .listItem),
484 ("1. item", .listItem),
485 ("2) item", .listItem),
486 ("-", .listItem),
487 ("-x", .plain),
488 ("1.5 apples", .plain),
489 ("plain text", .plain),
490 ])
491 func classify(line: String, expected: LineClass) {
492 #expect(classifyLine(line[...]).cls == expected)
493 }
494
495 @Test func indentCountsTabsToEight() {
496 #expect(classifyLine("\t- a").indent == 8)
497 #expect(classifyLine(" \t- a").indent == 8)
498 #expect(classifyLine(" - a").indent == 3)
499 }
500}
501```
502
503- [ ] **Step 2: Run tests to verify they fail**
504
505Run: `swift test --filter LinesTests`
506Expected: build failure, `cannot find 'splitRawLines' in scope`.
507
508- [ ] **Step 3: Implement**
509
510```swift
511/// One line of source: its content and its terminator ("", "\n" or "\r\n"), both as slices of
512/// the original text.
513struct RawLine {
514 let content: Substring
515 let ending: Substring
516}
517
518/// Splits on "\n" without normalizing anything. Works on unicode scalars, because String
519/// treats "\r\n" as a single Character.
520func splitRawLines(_ text: String) -> [RawLine] {
521 var lines: [RawLine] = []
522 let scalars = text.unicodeScalars
523 var lineStart = scalars.startIndex
524 var i = lineStart
525 while i != scalars.endIndex {
526 if scalars[i] == "\n" {
527 var contentEnd = i
528 if contentEnd > lineStart, scalars[scalars.index(before: i)] == "\r" {
529 contentEnd = scalars.index(before: i)
530 }
531 let next = scalars.index(after: i)
532 lines.append(RawLine(content: text[lineStart..<contentEnd], ending: text[contentEnd..<next]))
533 lineStart = next
534 i = next
535 } else {
536 i = scalars.index(after: i)
537 }
538 }
539 if lineStart != scalars.endIndex {
540 lines.append(RawLine(content: text[lineStart...], ending: ""))
541 }
542 return lines
543}
544
545enum LineClass: Equatable {
546 case blank
547 case heading(level: Int)
548 case blockBegin(name: String)
549 case blockEnd(name: String)
550 case dynamicBegin
551 case dynamicEnd
552 case drawerBegin(name: String)
553 case drawerEnd
554 case keyword(key: String)
555 case comment
556 case fixedWidth
557 case horizontalRule
558 case tableRow
559 case footnoteDefinition
560 case clock
561 case planning
562 case listItem
563 case plain
564}
565
566struct ClassifiedLine {
567 let cls: LineClass
568 /// Column of the first non-blank character, with tabs advancing to the next multiple of 8.
569 let indent: Int
570}
571
572func classifyLine(_ line: Substring) -> ClassifiedLine {
573 var column = 0
574 var rest = line
575 while let c = rest.first, c == " " || c == "\t" {
576 column = c == "\t" ? (column / 8 + 1) * 8 : column + 1
577 rest = rest.dropFirst()
578 }
579 if rest.isEmpty { return ClassifiedLine(cls: .blank, indent: column) }
580 return ClassifiedLine(cls: lineClass(rest, columnZero: column == 0), indent: column)
581}
582
583private func lineClass(_ rest: Substring, columnZero: Bool) -> LineClass {
584 let trimmed = rest.trimmingTrailingWhitespace
585
586 if columnZero, rest.first == "*" {
587 let stars = rest.prefix { $0 == "*" }
588 let after = rest.dropFirst(stars.count)
589 if after.isEmpty || after.first == " " || after.first == "\t" {
590 return .heading(level: stars.count)
591 }
592 }
593
594 if rest.hasPrefix("#+") {
595 let lower = trimmed.lowercased()
596 if lower.hasPrefix("#+begin_") {
597 let name = lower.dropFirst(8).prefix { !$0.isWhitespace }
598 if !name.isEmpty { return .blockBegin(name: String(name)) }
599 }
600 if lower.hasPrefix("#+end_") {
601 let name = lower.dropFirst(6)
602 if !name.isEmpty, !name.contains(where: \.isWhitespace) { return .blockEnd(name: String(name)) }
603 }
604 if lower.hasPrefix("#+begin:") { return .dynamicBegin }
605 if lower == "#+end:" { return .dynamicEnd }
606 if let colon = rest.firstIndex(of: ":") {
607 let key = rest[rest.index(rest.startIndex, offsetBy: 2)..<colon]
608 if !key.isEmpty, !key.contains(where: \.isWhitespace) { return .keyword(key: key.uppercased()) }
609 }
610 }
611
612 if trimmed == "#" || rest.hasPrefix("# ") || rest.hasPrefix("#\t") { return .comment }
613
614 if rest.first == ":" {
615 if trimmed == ":" || rest.hasPrefix(": ") || rest.hasPrefix(":\t") { return .fixedWidth }
616 if trimmed.uppercased() == ":END:" { return .drawerEnd }
617 if trimmed.count >= 3, trimmed.last == ":" {
618 let name = trimmed.dropFirst().dropLast()
619 if name.allSatisfy({ $0.isLetter || $0.isNumber || $0 == "_" || $0 == "-" }) {
620 return .drawerBegin(name: String(name))
621 }
622 }
623 }
624
625 if rest.first == "|" { return .tableRow }
626 if trimmed.count >= 5, trimmed.allSatisfy({ $0 == "-" }) { return .horizontalRule }
627
628 if columnZero, rest.hasPrefix("[fn:"), let close = rest.firstIndex(of: "]"),
629 close > rest.index(rest.startIndex, offsetBy: 4) {
630 return .footnoteDefinition
631 }
632
633 if rest.hasPrefix("CLOCK:") { return .clock }
634 if rest.hasPrefix("SCHEDULED:") || rest.hasPrefix("DEADLINE:") || rest.hasPrefix("CLOSED:") { return .planning }
635 if isListBullet(rest, indented: !columnZero) { return .listItem }
636 return .plain
637}
638
639/// `-`, `+`, `*` (indented only), `1.` or `1)`, followed by whitespace or end of line.
640/// Alphabetical bullets are off, as in org's default.
641private func isListBullet(_ rest: Substring, indented: Bool) -> Bool {
642 guard let first = rest.first else { return false }
643 let afterBullet: Substring
644 if first == "-" || first == "+" || (first == "*" && indented) {
645 afterBullet = rest.dropFirst()
646 } else if first.isASCII, first.isNumber {
647 let digits = rest.prefix { $0.isASCII && $0.isNumber }
648 let tail = rest.dropFirst(digits.count)
649 guard let separator = tail.first, separator == "." || separator == ")" else { return false }
650 afterBullet = tail.dropFirst()
651 } else {
652 return false
653 }
654 return afterBullet.isEmpty || afterBullet.first == " " || afterBullet.first == "\t"
655}
656
657extension Substring {
658 var trimmingTrailingWhitespace: Substring {
659 var s = self
660 while let last = s.last, last == " " || last == "\t" { s = s.dropLast() }
661 return s
662 }
663}
664```
665
666- [ ] **Step 4: Run tests to verify they pass**
667
668Run: `swift test --filter LinesTests`
669Expected: all pass.
670
671- [ ] **Step 5: Commit**
672
673```bash
674git add Sources Tests
675git commit -m "Add line splitting and classification"
676```
677
678---
679
680### Task 4: In-buffer settings
681
682**Files:**
683- Create: `Sources/OrgCore/Parser/Settings.swift`
684- Test: `Tests/OrgCoreTests/SettingsTests.swift`
685
686**Interfaces:**
687- Consumes: `RawLine`, `ClassifiedLine`, `LineClass` (Task 3).
688- Produces: `TodoKeyword`, `TodoSequence`, `Priorities`, `OrgSettings` with `.default`, `todoKeywordNames: Set<String>`, `isDone(_:)`; internal `SettingsScanner.scan(lines:info:blockEnds:defaults:) -> OrgSettings`, where `blockEnds: [Int: Int]` maps a block's begin line index to its end line index.
689
690- [ ] **Step 1: Write the failing tests**
691
692```swift
693import Testing
694@testable import OrgCore
695
696struct SettingsTests {
697 func scan(_ text: String, blockEnds: [Int: Int] = [:]) -> OrgSettings {
698 let lines = splitRawLines(text)
699 let info = lines.map { classifyLine($0.content) }
700 return SettingsScanner.scan(lines: lines, info: info, blockEnds: blockEnds, defaults: .default)
701 }
702
703 @Test func defaultsWithoutKeywords() {
704 let settings = scan("* TODO a\n")
705 #expect(settings.todoKeywordNames == ["TODO", "DONE"])
706 #expect(settings.isDone("DONE"))
707 }
708
709 @Test func fileKeywordsReplaceDefaults() {
710 let settings = scan("#+TODO: NEXT(n) WAIT(w@/!) | DONE(d!) CANCELED(c@)\n")
711 #expect(settings.todoKeywordNames == ["NEXT", "WAIT", "DONE", "CANCELED"])
712 let sequence = settings.todoSequences[0]
713 #expect(sequence.active.map(\.name) == ["NEXT", "WAIT"])
714 #expect(sequence.done.map(\.name) == ["DONE", "CANCELED"])
715 #expect(sequence.active[1] == TodoKeyword(name: "WAIT", fastKey: "w", logOnEnter: "@", logOnLeave: "!"))
716 #expect(sequence.done[0] == TodoKeyword(name: "DONE", fastKey: "d", logOnEnter: "!", logOnLeave: nil))
717 }
718
719 @Test func lastWordIsDoneWithoutSeparator() {
720 let settings = scan("#+SEQ_TODO: A B C\n")
721 #expect(settings.todoSequences[0].active.map(\.name) == ["A", "B"])
722 #expect(settings.todoSequences[0].done.map(\.name) == ["C"])
723 }
724
725 @Test func severalLinesMakeSeveralSequences() {
726 let settings = scan("#+TODO: A | B\n#+TYP_TODO: X | Y\n")
727 #expect(settings.todoSequences.count == 2)
728 #expect(settings.todoSequences[1].kind == .type)
729 }
730
731 @Test func keywordsInsideBlocksAreIgnored() {
732 let text = "#+begin_example\n#+TODO: X | Y\n#+end_example\n"
733 #expect(scan(text, blockEnds: [0: 2]).todoKeywordNames == ["TODO", "DONE"])
734 }
735
736 @Test func priorities() {
737 #expect(scan("#+PRIORITIES: 1 10 5\n").priorities == Priorities(highest: "1", lowest: "10", default: "5"))
738 #expect(scan("").priorities == Priorities(highest: "A", lowest: "C", default: "B"))
739 }
740}
741```
742
743- [ ] **Step 2: Run tests to verify they fail**
744
745Run: `swift test --filter SettingsTests`
746Expected: build failure, `cannot find 'SettingsScanner' in scope`.
747
748- [ ] **Step 3: Implement**
749
750```swift
751public struct TodoKeyword: Sendable, Hashable {
752 public var name: String
753 public var fastKey: Character?
754 /// Logging flag when entering the state (`!` or `@`), from `NAME(k!/@)`.
755 public var logOnEnter: String?
756 /// Logging flag when leaving the state.
757 public var logOnLeave: String?
758
759 public init(name: String, fastKey: Character? = nil, logOnEnter: String? = nil, logOnLeave: String? = nil) {
760 self.name = name
761 self.fastKey = fastKey
762 self.logOnEnter = logOnEnter
763 self.logOnLeave = logOnLeave
764 }
765}
766
767public struct TodoSequence: Sendable, Equatable {
768 public enum Kind: Sendable, Equatable { case sequence, type }
769
770 public var kind: Kind
771 public var active: [TodoKeyword]
772 public var done: [TodoKeyword]
773
774 public init(kind: Kind, active: [TodoKeyword], done: [TodoKeyword]) {
775 self.kind = kind
776 self.active = active
777 self.done = done
778 }
779}
780
781public struct Priorities: Sendable, Equatable {
782 public var highest: String
783 public var lowest: String
784 public var `default`: String
785
786 public init(highest: String, lowest: String, default: String) {
787 self.highest = highest
788 self.lowest = lowest
789 self.default = `default`
790 }
791}
792
793public struct OrgSettings: Sendable, Equatable {
794 public var todoSequences: [TodoSequence]
795 public var priorities: Priorities
796
797 public init(todoSequences: [TodoSequence], priorities: Priorities) {
798 self.todoSequences = todoSequences
799 self.priorities = priorities
800 }
801
802 public static let `default` = OrgSettings(
803 todoSequences: [TodoSequence(kind: .sequence, active: [TodoKeyword(name: "TODO")], done: [TodoKeyword(name: "DONE")])],
804 priorities: Priorities(highest: "A", lowest: "C", default: "B")
805 )
806
807 public var todoKeywordNames: Set<String> {
808 Set(todoSequences.flatMap { ($0.active + $0.done).map(\.name) })
809 }
810
811 public func isDone(_ name: String) -> Bool {
812 todoSequences.contains { $0.done.contains { $0.name == name } }
813 }
814}
815
816enum SettingsScanner {
817 /// Reads `#+TODO`, `#+SEQ_TODO`, `#+TYP_TODO` and `#+PRIORITIES` outside blocks. Any TODO
818 /// line replaces the default sequences, as in org.
819 static func scan(lines: [RawLine], info: [ClassifiedLine], blockEnds: [Int: Int], defaults: OrgSettings) -> OrgSettings {
820 var sequences: [TodoSequence] = []
821 var priorities = defaults.priorities
822 var k = 0
823 while k < lines.count {
824 switch info[k].cls {
825 case .blockBegin, .dynamicBegin:
826 if let end = blockEnds[k] { k = end }
827 case .keyword(let key):
828 let value = keywordValue(lines[k].content)
829 switch key {
830 case "TODO", "SEQ_TODO":
831 if let s = todoSequence(value, kind: .sequence) { sequences.append(s) }
832 case "TYP_TODO":
833 if let s = todoSequence(value, kind: .type) { sequences.append(s) }
834 case "PRIORITIES":
835 let words = value.split(whereSeparator: \.isWhitespace)
836 if words.count == 3 {
837 priorities = Priorities(highest: String(words[0]), lowest: String(words[1]), default: String(words[2]))
838 }
839 default:
840 break
841 }
842 default:
843 break
844 }
845 k += 1
846 }
847 return OrgSettings(todoSequences: sequences.isEmpty ? defaults.todoSequences : sequences, priorities: priorities)
848 }
849
850 static func keywordValue(_ line: Substring) -> Substring {
851 guard let colon = line.firstIndex(of: ":") else { return "" }
852 return line[line.index(after: colon)...]
853 }
854
855 static func todoSequence(_ value: Substring, kind: TodoSequence.Kind) -> TodoSequence? {
856 let words = value.split(whereSeparator: \.isWhitespace)
857 guard !words.isEmpty else { return nil }
858 if let bar = words.firstIndex(of: "|") {
859 return TodoSequence(kind: kind, active: words[..<bar].map(todoKeyword), done: words[(bar + 1)...].map(todoKeyword))
860 }
861 return TodoSequence(kind: kind, active: words.dropLast().map(todoKeyword), done: [todoKeyword(words.last!)])
862 }
863
864 /// `NAME`, or `NAME(spec)` where spec is an optional fast key followed by `enter/leave`
865 /// logging flags.
866 static func todoKeyword(_ word: Substring) -> TodoKeyword {
867 guard let open = word.firstIndex(of: "("), word.last == ")" else { return TodoKeyword(name: String(word)) }
868 var spec = word[word.index(after: open)..<word.index(before: word.endIndex)]
869 var fastKey: Character?
870 if let first = spec.first, first != "!", first != "@", first != "/" {
871 fastKey = first
872 spec = spec.dropFirst()
873 }
874 let parts = spec.split(separator: "/", omittingEmptySubsequences: false)
875 let enter = parts.first.flatMap { $0.isEmpty ? nil : String($0) }
876 let leave = parts.count > 1 && !parts[1].isEmpty ? String(parts[1]) : nil
877 return TodoKeyword(name: String(word[..<open]), fastKey: fastKey, logOnEnter: enter, logOnLeave: leave)
878 }
879}
880```
881
882- [ ] **Step 4: Run tests to verify they pass**
883
884Run: `swift test --filter SettingsTests`
885Expected: all pass.
886
887- [ ] **Step 5: Commit**
888
889```bash
890git add Sources Tests
891git commit -m "Add in-buffer TODO and priority settings"
892```
893
894---
895
896### Task 5: Parser — document, sections and headings
897
898**Files:**
899- Create: `Sources/OrgCore/Parser/Parser.swift`
900- Modify: `Sources/OrgCore/Syntax/SyntaxNode.swift` (append `OrgTree`)
901- Test: `Tests/OrgCoreTests/ParserSectionTests.swift`
902
903**Interfaces:**
904- Consumes: Tasks 2–4.
905- Produces: `public enum OrgParser { static func parse(_ text: String, defaults: OrgSettings = .default) -> OrgTree }`; `public struct OrgTree { green: GreenNode; settings: OrgSettings; root: SyntaxNode; text: String }`; internal `struct Parser` with `element(limit:floor:)` that Task 6 fills in. In this task `element` handles every class as a paragraph or a blank line.
906
907- [ ] **Step 1: Write the failing tests**
908
909```swift
910import Testing
911@testable import OrgCore
912
913func nodeKinds(_ text: String) -> [SyntaxKind] {
914 OrgParser.parse(text).root.descendants().map(\.kind)
915}
916
917func tokens(of kind: SyntaxKind, in text: String) -> [SyntaxToken] {
918 OrgParser.parse(text).root.descendants().filter { $0.kind == kind }.flatMap(\.tokens)
919}
920
921struct ParserSectionTests {
922 @Test func emptyDocument() {
923 let tree = OrgParser.parse("")
924 #expect(tree.text == "")
925 #expect(nodeKinds("") == [.document])
926 }
927
928 @Test func zerothSectionHoldsPreamble() {
929 #expect(nodeKinds("text\n* a\n") == [.document, .zerothSection, .paragraph, .section, .heading])
930 }
931
932 @Test func sectionsNestByLevel() {
933 let text = "* a\n** b\n*** c\n** d\n* e\n"
934 let root = OrgParser.parse(text).root
935 let top = root.children
936 #expect(top.map(\.kind) == [.section, .section])
937 #expect(top[0].children.map(\.kind) == [.heading, .section, .section])
938 #expect(top[0].children[1].children.map(\.kind) == [.heading, .section])
939 }
940
941 @Test func headingTokens() {
942 let parts = tokens(of: .heading, in: "** TODO [#A] Write the plan :work:urgent: \n")
943 #expect(parts.map(\.kind) == [.stars, .whitespace, .todoKeyword, .whitespace, .priority, .whitespace, .title, .whitespace, .tags, .whitespace, .newline])
944 #expect(parts.first { $0.kind == .tags }?.text == ":work:urgent:")
945 #expect(parts.first { $0.kind == .title }?.text == "Write the plan")
946 }
947
948 @Test func todoKeywordsComeFromSettings() {
949 let text = "#+TODO: NEXT | DONE\n* NEXT a\n* TODO b\n"
950 let todo = tokens(of: .heading, in: text).filter { $0.kind == .todoKeyword }.map(\.text)
951 #expect(todo == ["NEXT"])
952 }
953
954 @Test func priorityNeedsValidValueAndSpace() {
955 #expect(tokens(of: .heading, in: "* [#B] x\n").contains { $0.kind == .priority })
956 #expect(tokens(of: .heading, in: "* [#10] x\n").contains { $0.kind == .priority })
957 #expect(!tokens(of: .heading, in: "* [#AB] x\n").contains { $0.kind == .priority })
958 #expect(!tokens(of: .heading, in: "* [#A]x\n").contains { $0.kind == .priority })
959 }
960
961 @Test func tagsNeedValidCharacters() {
962 #expect(tokens(of: .heading, in: "* a :b@c_1:\n").contains { $0.kind == .tags })
963 #expect(!tokens(of: .heading, in: "* a :b c:\n").contains { $0.kind == .tags })
964 #expect(!tokens(of: .heading, in: "* a :b:c\n").contains { $0.kind == .tags })
965 }
966
967 @Test func headingWithoutNewlineAtEnd() {
968 #expect(OrgParser.parse("* a").text == "* a")
969 }
970
971 @Test(arguments: ["* a\n", "*\n", "* TODO\n", "text\r\n* a\r\n** b\r\n", "\n\n* a\n\n"])
972 func roundTrip(text: String) {
973 #expect(OrgParser.parse(text).text == text)
974 }
975}
976```
977
978- [ ] **Step 2: Run tests to verify they fail**
979
980Run: `swift test --filter ParserSectionTests`
981Expected: build failure, `cannot find 'OrgParser' in scope`.
982
983- [ ] **Step 3: Append `OrgTree` to `SyntaxNode.swift`**
984
985```swift
986public struct OrgTree: Sendable {
987 public let green: GreenNode
988 public let settings: OrgSettings
989
990 public var root: SyntaxNode { SyntaxNode(green: green, offset: 0, parent: nil) }
991 public var text: String { green.text }
992}
993```
994
995- [ ] **Step 4: Implement `Parser.swift`**
996
997```swift
998public enum OrgParser {
999 public static func parse(_ text: String, defaults: OrgSettings = .default) -> OrgTree {
1000 var parser = Parser(text: text, defaults: defaults)
1001 return parser.run()
1002 }
1003}
1004
1005struct Parser {
1006 let lines: [RawLine]
1007 let info: [ClassifiedLine]
1008 /// Begin line → end line, for blocks, dynamic blocks and drawers that are closed before the
1009 /// next heading.
1010 let blockEnds: [Int: Int]
1011 let settings: OrgSettings
1012 var builder = GreenBuilder()
1013 var i = 0
1014
1015 init(text: String, defaults: OrgSettings) {
1016 lines = splitRawLines(text)
1017 info = lines.map { classifyLine($0.content) }
1018 blockEnds = Parser.matchEnds(info)
1019 settings = SettingsScanner.scan(lines: lines, info: info, blockEnds: blockEnds, defaults: defaults)
1020 }
1021
1022 static func matchEnds(_ info: [ClassifiedLine]) -> [Int: Int] {
1023 var ends: [Int: Int] = [:]
1024 var k = 0
1025 while k < info.count {
1026 let isEnd: ((LineClass) -> Bool)?
1027 switch info[k].cls {
1028 case .blockBegin(let name): isEnd = { $0 == .blockEnd(name: name) }
1029 case .dynamicBegin: isEnd = { $0 == .dynamicEnd }
1030 case .drawerBegin: isEnd = { $0 == .drawerEnd }
1031 default: isEnd = nil
1032 }
1033 if let isEnd {
1034 var j = k + 1
1035 while j < info.count {
1036 if case .heading = info[j].cls { break }
1037 if isEnd(info[j].cls) { ends[k] = j; break }
1038 j += 1
1039 }
1040 // Block contents are verbatim, so nothing inside starts another element.
1041 if let end = ends[k], !isDrawer(info[k].cls) { k = end }
1042 }
1043 k += 1
1044 }
1045 return ends
1046 }
1047
1048 static func isDrawer(_ cls: LineClass) -> Bool {
1049 if case .drawerBegin = cls { return true }
1050 return false
1051 }
1052
1053 mutating func run() -> OrgTree {
1054 builder.start(.document)
1055 if !lines.isEmpty, !isHeading(0) {
1056 builder.start(.zerothSection)
1057 parseContent(limit: lines.count)
1058 builder.finish()
1059 }
1060 while i < lines.count, case .heading(let level) = info[i].cls {
1061 parseSection(level: level)
1062 }
1063 builder.finish()
1064 return OrgTree(green: builder.build(), settings: settings)
1065 }
1066
1067 func isHeading(_ k: Int) -> Bool {
1068 if case .heading = info[k].cls { return true }
1069 return false
1070 }
1071
1072 mutating func parseSection(level: Int) {
1073 builder.start(.section)
1074 headingLine(lines[i])
1075 i += 1
1076 if i < lines.count, info[i].cls == .planning {
1077 builder.start(.planning)
1078 line(i)
1079 i += 1
1080 builder.finish()
1081 }
1082 if i < lines.count, case .drawerBegin(let name) = info[i].cls, name.uppercased() == "PROPERTIES",
1083 let end = blockEnds[i] {
1084 propertyDrawer(end: end)
1085 }
1086 parseContent(limit: lines.count)
1087 while i < lines.count, case .heading(let child) = info[i].cls, child > level {
1088 parseSection(level: child)
1089 }
1090 builder.finish()
1091 }
1092
1093 mutating func propertyDrawer(end: Int) {
1094 builder.start(.propertyDrawer)
1095 line(i)
1096 i += 1
1097 while i < end {
1098 if info[i].cls == .blank {
1099 line(i)
1100 } else {
1101 builder.start(.nodeProperty)
1102 line(i)
1103 builder.finish()
1104 }
1105 i += 1
1106 }
1107 line(i)
1108 i += 1
1109 builder.finish()
1110 }
1111
1112 /// Elements until `limit` or the next heading.
1113 mutating func parseContent(limit: Int) {
1114 while i < limit, !isHeading(i) {
1115 element(limit: limit, floor: nil)
1116 }
1117 }
1118
1119 /// One element starting at `i`. `floor` is the indent of the enclosing list item, if any:
1120 /// non-blank lines at or left of it end the element.
1121 mutating func element(limit: Int, floor: Int?) {
1122 if info[i].cls == .blank {
1123 line(i)
1124 i += 1
1125 } else {
1126 paragraph(limit: limit, floor: floor)
1127 }
1128 }
1129
1130 mutating func paragraph(limit: Int, floor: Int?) {
1131 builder.start(.paragraph)
1132 line(i)
1133 i += 1
1134 while i < limit, within(floor, i), continuesParagraph(i) {
1135 line(i)
1136 i += 1
1137 }
1138 builder.finish()
1139 }
1140
1141 func continuesParagraph(_ k: Int) -> Bool {
1142 switch info[k].cls {
1143 case .blank, .heading: return false
1144 default: return true
1145 }
1146 }
1147
1148 func within(_ floor: Int?, _ k: Int) -> Bool {
1149 guard let floor else { return true }
1150 return info[k].indent > floor
1151 }
1152
1153 // MARK: - Tokens
1154
1155 /// A whole line as leading whitespace, content and line ending.
1156 mutating func line(_ k: Int) {
1157 let content = lines[k].content
1158 let rest = whitespace(content)
1159 builder.token(.text, rest)
1160 builder.token(.newline, lines[k].ending)
1161 }
1162
1163 mutating func whitespace(_ s: Substring) -> Substring {
1164 let ws = s.prefix { $0 == " " || $0 == "\t" }
1165 builder.token(.whitespace, ws)
1166 return s.dropFirst(ws.count)
1167 }
1168
1169 mutating func headingLine(_ raw: RawLine) {
1170 builder.start(.heading)
1171 var rest = raw.content
1172 let stars = rest.prefix { $0 == "*" }
1173 builder.token(.stars, stars)
1174 rest = whitespace(rest.dropFirst(stars.count))
1175
1176 let word = rest.prefix { $0 != " " && $0 != "\t" }
1177 if !word.isEmpty, settings.todoKeywordNames.contains(String(word)) {
1178 builder.token(.todoKeyword, word)
1179 rest = whitespace(rest.dropFirst(word.count))
1180 }
1181
1182 if let cookie = priorityCookie(rest) {
1183 builder.token(.priority, cookie)
1184 rest = whitespace(rest.dropFirst(cookie.count))
1185 }
1186
1187 let parts = splitTags(rest)
1188 builder.token(.title, parts.title)
1189 builder.token(.whitespace, parts.gap)
1190 builder.token(.tags, parts.tags)
1191 builder.token(.whitespace, parts.trailing)
1192 builder.token(.newline, raw.ending)
1193 builder.finish()
1194 }
1195
1196 /// `[#A]` or `[#10]`, followed by whitespace or end of line.
1197 func priorityCookie(_ s: Substring) -> Substring? {
1198 guard s.hasPrefix("[#"), let close = s.firstIndex(of: "]") else { return nil }
1199 let value = s[s.index(s.startIndex, offsetBy: 2)..<close]
1200 let valid = (value.count == 1 && value.first!.isLetter && value.first!.isUppercase)
1201 || (!value.isEmpty && value.allSatisfy { $0.isASCII && $0.isNumber })
1202 guard valid else { return nil }
1203 let after = s[s.index(after: close)...]
1204 guard after.isEmpty || after.first == " " || after.first == "\t" else { return nil }
1205 return s[...close]
1206 }
1207
1208 func splitTags(_ s: Substring) -> (title: Substring, gap: Substring, tags: Substring, trailing: Substring) {
1209 let trimmed = s.trimmingTrailingWhitespace
1210 let trailing = s[trimmed.endIndex...]
1211 let none = (title: trimmed, gap: Substring(), tags: Substring(), trailing: trailing)
1212 guard trimmed.last == ":" else { return none }
1213 let tagStart = trimmed.lastIndex { $0 == " " || $0 == "\t" }.map { trimmed.index(after: $0) } ?? trimmed.startIndex
1214 let tags = trimmed[tagStart...]
1215 guard tags.count >= 3, tags.first == ":", isTagString(tags) else { return none }
1216 let before = trimmed[..<tagStart]
1217 let title = before.trimmingTrailingWhitespace
1218 return (title, before[title.endIndex...], tags, trailing)
1219 }
1220
1221 func isTagString(_ tags: Substring) -> Bool {
1222 tags.dropFirst().dropLast().split(separator: ":", omittingEmptySubsequences: false).allSatisfy { tag in
1223 !tag.isEmpty && tag.allSatisfy { $0.isLetter || $0.isNumber || "_@#%".contains($0) }
1224 }
1225 }
1226}
1227```
1228
1229- [ ] **Step 5: Run tests to verify they pass**
1230
1231Run: `swift test --filter ParserSectionTests`
1232Expected: all pass.
1233
1234- [ ] **Step 6: Commit**
1235
1236```bash
1237git add Sources Tests
1238git commit -m "Parse document, sections and headings"
1239```
1240
1241---
1242
1243### Task 6: Parser — elements
1244
1245**Files:**
1246- Modify: `Sources/OrgCore/Parser/Parser.swift` (replace `element`, `continuesParagraph`; add `list`, `item`, `table`, `consecutive`, `footnoteDefinition`)
1247- Test: `Tests/OrgCoreTests/ParserElementTests.swift`
1248
1249**Interfaces:**
1250- Consumes: Task 5's `Parser`.
1251- Produces: nodes `block`, `dynamicBlock`, `drawer`, `keyword`, `affiliatedKeyword`, `comment`, `fixedWidth`, `horizontalRule`, `table`/`tableRow`/`tableFormula`, `footnoteDefinition`, `clock`, `plainList`/`item`, `planning` (after heading only), `propertyDrawer`/`nodeProperty`.
1252
1253- [ ] **Step 1: Write the failing tests**
1254
1255```swift
1256import Testing
1257@testable import OrgCore
1258
1259func childKinds(_ text: String) -> [SyntaxKind] {
1260 let root = OrgParser.parse(text).root
1261 let container = root.children.first { $0.kind == .zerothSection || $0.kind == .section }!
1262 return container.children.map(\.kind)
1263}
1264
1265struct ParserElementTests {
1266 @Test func planningAndPropertiesFollowHeading() {
1267 let text = "* a\nSCHEDULED: <2026-10-04 Sun>\n:PROPERTIES:\n:ID: x\n:END:\nbody\n"
1268 #expect(childKinds(text) == [.heading, .planning, .propertyDrawer, .paragraph])
1269 }
1270
1271 @Test func planningElsewhereIsText() {
1272 #expect(childKinds("SCHEDULED: <2026-10-04 Sun>\n") == [.paragraph])
1273 }
1274
1275 @Test func blocks() {
1276 #expect(childKinds("#+begin_src sh\n,* escaped\n:END:\n#+end_src\nafter\n") == [.block, .paragraph])
1277 #expect(childKinds("#+BEGIN: clocktable\n#+END:\n") == [.dynamicBlock])
1278 }
1279
1280 @Test func headingsEndBlocks() {
1281 #expect(childKinds("#+begin_src sh\n* heading\n#+end_src\n") == [.paragraph])
1282 }
1283
1284 @Test func unclosedBlockIsParagraph() {
1285 #expect(childKinds("#+begin_src sh\necho\n") == [.paragraph])
1286 }
1287
1288 @Test func drawersHoldElements() {
1289 let root = OrgParser.parse(":LOGBOOK:\nCLOCK: [2026-10-04 Sun 10:00]\n:END:\n").root
1290 let drawer = root.children[0].children[0]
1291 #expect(drawer.kind == .drawer)
1292 #expect(drawer.children.map(\.kind) == [.clock])
1293 }
1294
1295 @Test func keywords() {
1296 #expect(childKinds("#+TITLE: x\n#+NAME: t\n#+ATTR_HTML: :width 10\n") == [.keyword, .affiliatedKeyword, .affiliatedKeyword])
1297 }
1298
1299 @Test func commentsAndFixedWidthGroup() {
1300 #expect(childKinds("# a\n# b\n: c\n: d\n-----\n") == [.comment, .fixedWidth, .horizontalRule])
1301 }
1302
1303 @Test func tableWithFormulas() {
1304 let root = OrgParser.parse("| a |\n|---|\n| 1 |\n#+TBLFM: $1=2\n#+TBLFM: $1=3\n").root
1305 let table = root.children[0].children[0]
1306 #expect(table.kind == .table)
1307 #expect(table.children.map(\.kind) == [.tableRow, .tableRow, .tableRow, .tableFormula, .tableFormula])
1308 }
1309
1310 @Test func footnoteDefinition() {
1311 #expect(childKinds("[fn:1] note\ncontinued\n\nafter\n") == [.footnoteDefinition, .paragraph])
1312 }
1313
1314 @Test func listsNestByIndent() {
1315 let text = "- a\n more\n - b\n- c\n\nafter\n"
1316 let root = OrgParser.parse(text).root
1317 let list = root.children[0].children[0]
1318 #expect(list.kind == .plainList)
1319 #expect(list.children.map(\.kind) == [.item, .item])
1320 #expect(list.children[0].children.map(\.kind) == [.paragraph, .plainList])
1321 #expect(childKinds(text) == [.plainList, .paragraph])
1322 }
1323
1324 @Test func twoBlankLinesEndAList() {
1325 #expect(childKinds("- a\n\n\n- b\n") == [.plainList, .plainList])
1326 #expect(childKinds("- a\n\n- b\n") == [.plainList])
1327 }
1328
1329 @Test func paragraphStopsAtElementStart() {
1330 #expect(childKinds("text\n| a |\n") == [.paragraph, .table])
1331 #expect(childKinds("text\n#+begin_quote\nq\n#+end_quote\n") == [.paragraph, .block])
1332 #expect(childKinds("text\n#+begin_quote\nq\n") == [.paragraph])
1333 }
1334}
1335```
1336
1337- [ ] **Step 2: Run tests to verify they fail**
1338
1339Run: `swift test --filter ParserElementTests`
1340Expected: failures; every element parses as a paragraph.
1341
1342- [ ] **Step 3: Replace `element` and `continuesParagraph`, add the element parsers**
1343
1344```swift
1345 static let affiliatedKeys: Set<String> = ["NAME", "CAPTION", "RESULTS", "HEADER", "PLOT"]
1346
1347 mutating func element(limit: Int, floor: Int?) {
1348 switch info[i].cls {
1349 case .blank:
1350 line(i)
1351 i += 1
1352 case .blockBegin, .dynamicBegin:
1353 if let end = blockEnds[i], end < limit {
1354 builder.start(info[i].cls == .dynamicBegin ? .dynamicBlock : .block)
1355 while i <= end {
1356 line(i)
1357 i += 1
1358 }
1359 builder.finish()
1360 } else {
1361 paragraph(limit: limit, floor: floor)
1362 }
1363 case .drawerBegin:
1364 if let end = blockEnds[i], end < limit {
1365 builder.start(.drawer)
1366 line(i)
1367 i += 1
1368 parseContent(limit: end)
1369 line(i)
1370 i += 1
1371 builder.finish()
1372 } else {
1373 paragraph(limit: limit, floor: floor)
1374 }
1375 case .keyword(let key):
1376 let affiliated = Self.affiliatedKeys.contains(key) || key.hasPrefix("ATTR_")
1377 single(affiliated ? .affiliatedKeyword : .keyword)
1378 case .comment:
1379 consecutive(.comment, limit: limit, floor: floor) { $0 == .comment }
1380 case .fixedWidth:
1381 consecutive(.fixedWidth, limit: limit, floor: floor) { $0 == .fixedWidth }
1382 case .horizontalRule:
1383 single(.horizontalRule)
1384 case .clock:
1385 single(.clock)
1386 case .tableRow:
1387 table(limit: limit, floor: floor)
1388 case .footnoteDefinition:
1389 footnoteDefinition(limit: limit)
1390 case .listItem:
1391 list(limit: limit, floor: floor)
1392 default:
1393 paragraph(limit: limit, floor: floor)
1394 }
1395 }
1396
1397 /// Lines that don't start an element of their own.
1398 func continuesParagraph(_ k: Int) -> Bool {
1399 switch info[k].cls {
1400 case .plain, .planning, .blockEnd, .dynamicEnd, .drawerEnd:
1401 return true
1402 case .blockBegin, .dynamicBegin, .drawerBegin:
1403 return blockEnds[k] == nil
1404 default:
1405 return false
1406 }
1407 }
1408
1409 mutating func single(_ kind: SyntaxKind) {
1410 builder.start(kind)
1411 line(i)
1412 i += 1
1413 builder.finish()
1414 }
1415
1416 mutating func consecutive(_ kind: SyntaxKind, limit: Int, floor: Int?, matching: (LineClass) -> Bool) {
1417 builder.start(kind)
1418 repeat {
1419 line(i)
1420 i += 1
1421 } while i < limit && matching(info[i].cls) && within(floor, i)
1422 builder.finish()
1423 }
1424
1425 mutating func table(limit: Int, floor: Int?) {
1426 builder.start(.table)
1427 while i < limit, info[i].cls == .tableRow, within(floor, i) {
1428 single(.tableRow)
1429 }
1430 while i < limit, info[i].cls == .keyword(key: "TBLFM"), within(floor, i) {
1431 single(.tableFormula)
1432 }
1433 builder.finish()
1434 }
1435
1436 mutating func footnoteDefinition(limit: Int) {
1437 builder.start(.footnoteDefinition)
1438 line(i)
1439 i += 1
1440 while i < limit, info[i].cls == .plain {
1441 line(i)
1442 i += 1
1443 }
1444 builder.finish()
1445 }
1446
1447 mutating func list(limit: Int, floor: Int?) {
1448 let base = info[i].indent
1449 builder.start(.plainList)
1450 while i < limit, info[i].cls == .listItem, info[i].indent == base, within(floor, i) {
1451 item(base: base, limit: limit)
1452 }
1453 builder.finish()
1454 }
1455
1456 /// An item's first line, then everything indented past its bullet. One blank line stays
1457 /// inside the item when the item or list continues after it; two end the list.
1458 mutating func item(base: Int, limit: Int) {
1459 builder.start(.item)
1460 line(i)
1461 i += 1
1462 while i < limit, !isHeading(i) {
1463 if info[i].cls == .blank {
1464 var j = i
1465 while j < limit, info[j].cls == .blank { j += 1 }
1466 guard j - i < 2, j < limit else { break }
1467 let continuesItem = info[j].indent > base
1468 let nextSibling = info[j].cls == .listItem && info[j].indent == base
1469 guard continuesItem || nextSibling else { break }
1470 line(i)
1471 i += 1
1472 if nextSibling { break }
1473 continue
1474 }
1475 guard info[i].indent > base else { break }
1476 element(limit: limit, floor: base)
1477 }
1478 builder.finish()
1479 }
1480```
1481
1482- [ ] **Step 4: Run all tests**
1483
1484Run: `swift test`
1485Expected: all pass, including Task 5's tests.
1486
1487- [ ] **Step 5: Commit**
1488
1489```bash
1490git add Sources Tests
1491git commit -m "Parse block-level elements"
1492```
1493
1494---
1495
1496### Task 7: Round-trip fuzz, encoding fixtures and corpus test
1497
1498**Files:**
1499- Test: `Tests/OrgCoreTests/RoundTripTests.swift`
1500
1501**Interfaces:**
1502- Consumes: `SourceText`, `OrgParser`, `SyntaxNode`.
1503- Produces: a seeded fuzz test (2,000 documents), a structural invariant check, and an opt-in corpus test driven by `ORGSTAR_CORPUS`.
1504
1505- [ ] **Step 1: Write the tests**
1506
1507```swift
1508import Foundation
1509import Testing
1510@testable import OrgCore
1511
1512/// SplitMix64, so failures reproduce from the seed.
1513struct SeededGenerator: RandomNumberGenerator {
1514 var state: UInt64
1515 mutating func next() -> UInt64 {
1516 state &+= 0x9E37_79B9_7F4A_7C15
1517 var z = state
1518 z = (z ^ (z >> 30)) &* 0xBF58_476D_1CE4_E5B9
1519 z = (z ^ (z >> 27)) &* 0x94D0_49BB_1331_11EB
1520 return z ^ (z >> 31)
1521 }
1522}
1523
1524let fragments = [
1525 "* ", "** TODO [#A] title :a:b:", "*** DONE", "#+TODO: NEXT | DONE", "#+begin_src sh", "#+end_src",
1526 "#+BEGIN_QUOTE", "#+end_quote", "#+BEGIN: clocktable", "#+END:", ":PROPERTIES:", ":ID: x", ":END:",
1527 ":LOGBOOK:", "CLOCK: [2026-10-04 Sun 10:00]", "SCHEDULED: <2026-10-04 Sun>", "- item", " - nested",
1528 "\t+ tab", "1. one", "| a | b |", "|---+---|", "#+TBLFM: $2=$1", "# comment", ": fixed", "-----",
1529 "[fn:1] note", "#+NAME: x", "plain text", "é", "😀", "e\u{301}", " ", "\t", "\n", "\n", "\r\n", "\r", "",
1530]
1531
1532func randomDocument(_ rng: inout SeededGenerator) -> String {
1533 (0..<Int.random(in: 0...40, using: &rng)).map { _ in
1534 fragments.randomElement(using: &rng)! + (Bool.random(using: &rng) ? "\n" : "")
1535 }.joined()
1536}
1537
1538func checkLengths(_ node: SyntaxNode) -> Bool {
1539 let sum = node.green.children.reduce(0) { $0 + $1.length }
1540 return sum == node.green.length && node.children.allSatisfy(checkLengths)
1541}
1542
1543struct RoundTripTests {
1544 @Test func fuzzedDocumentsRoundTrip() {
1545 var rng = SeededGenerator(state: 20261004)
1546 for n in 0..<2_000 {
1547 let text = randomDocument(&rng)
1548 let tree = OrgParser.parse(text)
1549 #expect(tree.text == text, "document \(n)")
1550 #expect(checkLengths(tree.root), "document \(n)")
1551 }
1552 }
1553
1554 @Test(arguments: [
1555 [0xEF, 0xBB, 0xBF] + Array("* a\r\n".utf8),
1556 Array("* a\r\n- b\n\tc".utf8),
1557 Array("* a".utf8),
1558 Array("😀 e\u{301}\n".utf8),
1559 [0x2A, 0x20, 0xFF, 0x0A] as [UInt8],
1560 ])
1561 func encodingFixturesRoundTrip(bytes: [UInt8]) {
1562 let source = SourceText(bytes: bytes)
1563 let tree = OrgParser.parse(source.text)
1564 #expect(source.encode(tree.text) == bytes)
1565 }
1566
1567 /// Private corpus: `ORGSTAR_CORPUS=/path/to/org swift test --filter corpus`.
1568 @Test(.enabled(if: ProcessInfo.processInfo.environment["ORGSTAR_CORPUS"] != nil))
1569 func corpusRoundTrips() throws {
1570 let root = URL(fileURLWithPath: ProcessInfo.processInfo.environment["ORGSTAR_CORPUS"]!)
1571 let files = FileManager.default.enumerator(at: root, includingPropertiesForKeys: nil)!
1572 .compactMap { $0 as? URL }
1573 .filter { ["org", "org_archive"].contains($0.pathExtension) }
1574 #expect(!files.isEmpty)
1575 for file in files {
1576 let bytes = try [UInt8](Data(contentsOf: file))
1577 let source = SourceText(bytes: bytes)
1578 let tree = OrgParser.parse(source.text)
1579 #expect(source.encode(tree.text) == bytes, "\(file.path)")
1580 }
1581 }
1582}
1583```
1584
1585- [ ] **Step 2: Run the tests**
1586
1587Run: `swift test`
1588Expected: all pass. The corpus test is skipped.
1589
1590- [ ] **Step 3: Run the private corpus locally**
1591
1592Run: `ORGSTAR_CORPUS=~/org swift test --filter corpusRoundTrips`
1593Expected: pass. Any failure names the file; reduce it to a fuzz fragment or a unit test before fixing.
1594
1595- [ ] **Step 4: Commit**
1596
1597```bash
1598git add Tests
1599git commit -m "Add round-trip fuzz, encoding and corpus tests"
1600```