Commit b84ac9f2e2
Verified · cmc
Layout: unified · split
docs/design.md +1 −1
| @@ -364,7 +364,7 @@ Seven phases, each usable on its own. Phase 1 is read-only, so it can run next t | ||
| 364 | 364 | | Incremental equals full | 0 mismatches over 100,000 fuzzed edits | |
| 365 | 365 | | Open to first render | p95 under 100 ms for a 1 MB file; under 500 ms for a 10 MB file | |
| 366 | 366 | | Keystroke to restyled frame (spike) | p95 under 16 ms for a 1 MB file, including one file whose content is a single top-level section and one with a 5,000-row table | |
| 367 | | Peak memory | Under 10x file size for an open document | | |
| 367 | | Peak memory | Under 10x file size for an open typical document (prose, lists, tables); under 45x for worst-case dense markup (the synthetic gate files: a 1 MB single section measures 30x, a 5,000-row table 43x) | | |
| 368 | 368 | | Index | Full rebuild of 10,000 files under 30 s; reconciliation after restart under 2 s with no changes | |
| 369 | 369 | | Save safety | All fault-injection cases end with both versions recoverable | |
| 370 | 370 | | Editing spike | Caret, selection, IME, VoiceOver and copy/paste checks pass on folded and indented text | |
docs/plans/2026-10-04-performance.md added +1316
| @@ -0,0 +1,1316 @@ | ||
| 1 | # Phase 1 Performance Implementation Plan | |
| 2 | ||
| 3 | > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. | |
| 4 | ||
| 5 | **Goal:** Meet the phase 1 speed gates (open a 1 MB file under 100 ms; keystroke to restyled text under 16 ms p95, including a 1 MB single-section file and a 5,000-row table) and record the memory gate the user set. | |
| 6 | ||
| 7 | **Architecture:** Four changes, each found by measuring first. (1) The inline scanner works on unicode scalars, rejects characters that can't start an object before trying recognizers, and slices token text from the source; this was most of the parse time. (2) A region strategy reparses a window of elements inside the innermost section's body instead of the whole top-level section, verifying the end of the window by reparsing one element past it, and reuses unchanged table rows. (3) The editor loads text through the text storage directly, styles the first screen before the first render and the rest in later run-loop turns, and restyles only the rows or items around an edit inside large tables and lists. (4) Tree queries on hot paths create only the child nodes they need. | |
| 8 | ||
| 9 | **Tech Stack:** Swift 6.2 tools, Swift Testing, AppKit, TextKit 2. | |
| 10 | ||
| 11 | **Spec:** `docs/design.md`, "Phases" (exit gates). | |
| 12 | ||
| 13 | ## Global Constraints | |
| 14 | ||
| 15 | - Every reparse equals a full parse: 100,000 random edits per seed, several seeds, release build. | |
| 16 | - Every incremental restyle equals a fresh editor's full styling: 1,500 random edits per seed, several seeds. | |
| 17 | - Memory gate (user decision): under 10x the file size for typical documents; under 45x for worst-case dense markup. | |
| 18 | ||
| 19 | ## Results | |
| 20 | ||
| 21 | Release build, M-series Mac; the M1 Air reference machine is still to be measured. | |
| 22 | ||
| 23 | | Measure | Before | After | Gate | | |
| 24 | | --- | --- | --- | --- | | |
| 25 | | Full parse, `csg.org` (1.5 MB) | 313 ms | 62 ms | | | |
| 26 | | Full parse, `org-manual.org` (0.86 MB) | 126 ms | 64 ms | | | |
| 27 | | Open, `csg.org` | ~440 ms | 75 ms | 100 ms per MB | | |
| 28 | | Typing p95, `csg.org` | 5.8 ms | 6.5 ms | 16 ms | | |
| 29 | | Blank line p95, `csg.org` | 256 ms | 7.3 ms | 16 ms | | |
| 30 | | Synthetic 1 MB single section: open / typing / blank line | | 76 / 7.0 / 8.1 ms | 100 / 16 / 16 ms | | |
| 31 | | Synthetic 5,000-row table: open / typing / blank line | | 38 / 3.2 / 8.8 ms | 100 / 16 / 16 ms | | |
| 32 | | Memory: `csg.org` / single section / table | | 9.6x / 30.4x / 42.9x | 10x / 45x / 45x | | |
| 33 | ||
| 34 | What the measurements showed along the way: | |
| 35 | - Line splitting and classification were never the problem (12 ms for 1.5 MB); inline scanning at about 170 ns per character was. | |
| 36 | - `NSTextView.string` was not the problem either: the first text loaded in a process pays a one-time 40 ms set-up, after which loading 1 MB takes under 1 ms through the text storage. | |
| 37 | - Restyling a whole table or list for one edit inside it cost 60 ms on 5,000 rows. | |
| 38 | - Memory is mostly the tree (19x to 28x on dense markup: one token or node per piece of markup) plus about 10x for the text view. A compact tree storing token text as offsets is possible later if memory becomes a real problem. | |
| 39 | ||
| 40 | The differential tests found three bugs during this work, each fixed and covered: the element before a reparse window can absorb the window's first line when a drawer loses its `:END:`; an emptied zeroth section must disappear as in a full parse; and the initial-styling watermark must move with edits. | |
| 41 | ||
| 42 | Run the gates: `ORGSTAR_GATES=1 ORGSTAR_GATE_FILE=<a large real file> swift test -c release --filter GateTests`. | |
| 43 | ||
| 44 | --- | |
| 45 | ||
| 46 | ### Task 1: Inline scanning on unicode scalars | |
| 47 | ||
| 48 | **Files:** | |
| 49 | - Modify: `Sources/OrgCore/Parser/Inline.swift`, `Sources/OrgCore/Timestamp.swift` | |
| 50 | ||
| 51 | - [ ] **Step 1: Rewrite the scanner and the timestamp recognizer over `[Unicode.Scalar]`** | |
| 52 | ||
| 53 | The existing inline, timestamp and round-trip tests are the specification; they don't change. | |
| 54 | ||
| 55 | ```swift | |
| 56 | /// Characters allowed before an emphasis opener, besides whitespace and the start of the run. | |
| 57 | private let emphasisPre: Set<Unicode.Scalar> = ["-", "(", "{", "'", "\""] | |
| 58 | ||
| 59 | /// Characters allowed after an emphasis closer, besides whitespace and the end of the run. | |
| 60 | private let emphasisPost: Set<Unicode.Scalar> = ["-", ".", ",", ":", "!", "?", ";", "'", "\"", ")", "}", "\\", "["] | |
| 61 | ||
| 62 | private let angleLinkSchemes: Set<String> = [ | |
| 63 | "http", "https", "mailto", "file", "id", "doi", "ftp", "news", "shell", "elisp", "info", "help", "attachment", | |
| 64 | ] | |
| 65 | ||
| 66 | private let plainLinkPrefixes: [[Unicode.Scalar]] = ["https://", "http://", "mailto:", "file:"].map { Array($0.unicodeScalars) } | |
| 67 | ||
| 68 | private func emphasisKind(_ c: Unicode.Scalar) -> SyntaxKind? { | |
| 69 | switch c { | |
| 70 | case "*": .bold | |
| 71 | case "/": .italic | |
| 72 | case "_": .underline | |
| 73 | case "+": .strikeThrough | |
| 74 | case "=": .verbatim | |
| 75 | case "~": .code | |
| 76 | default: nil | |
| 77 | } | |
| 78 | } | |
| 79 | ||
| 80 | func isNewline(_ c: Unicode.Scalar) -> Bool { | |
| 81 | c == "\n" | |
| 82 | } | |
| 83 | ||
| 84 | func isSpace(_ c: Unicode.Scalar) -> Bool { | |
| 85 | c.isASCII ? (c == " " || c == "\t" || c == "\n" || c == "\r" || c == "\u{0B}" || c == "\u{0C}") : c.properties.isWhitespace | |
| 86 | } | |
| 87 | ||
| 88 | func isWordScalar(_ c: Unicode.Scalar) -> Bool { | |
| 89 | if c.isASCII { | |
| 90 | let v = c.value | |
| 91 | return (v >= 48 && v <= 57) || (v >= 65 && v <= 90) || (v >= 97 && v <= 122) | |
| 92 | } | |
| 93 | return c.properties.isAlphabetic || c.properties.numericType != nil | |
| 94 | } | |
| 95 | ||
| 96 | enum InlineMatch { | |
| 97 | case emphasis(SyntaxKind, open: Int, close: Int) | |
| 98 | case link(path: Range<Int>, description: Range<Int>?, whole: Range<Int>) | |
| 99 | case object(SyntaxKind, Range<Int>) | |
| 100 | ||
| 101 | var end: Int { | |
| 102 | switch self { | |
| 103 | case .emphasis(_, _, let close): close + 1 | |
| 104 | case .link(_, _, let whole): whole.upperBound | |
| 105 | case .object(_, let range): range.upperBound | |
| 106 | } | |
| 107 | } | |
| 108 | } | |
| 109 | ||
| 110 | /// Turns a run of text into text, newline and object tokens. Every scalar ends up in exactly | |
| 111 | /// one token. Works on unicode scalars: every recognizer starts on an ASCII character, so | |
| 112 | /// grapheme clusters never need to be formed. | |
| 113 | struct InlineScanner { | |
| 114 | let chars: [Unicode.Scalar] | |
| 115 | /// Where each scalar starts in `source`, plus its end, so token text is sliced from the | |
| 116 | /// source instead of rebuilt scalar by scalar. | |
| 117 | let source: Substring.UnicodeScalarView | |
| 118 | let starts: [String.Index] | |
| 119 | ||
| 120 | init(_ text: Substring) { | |
| 121 | source = text.unicodeScalars | |
| 122 | var chars: [Unicode.Scalar] = [] | |
| 123 | var starts: [String.Index] = [] | |
| 124 | chars.reserveCapacity(text.utf8.count) | |
| 125 | starts.reserveCapacity(text.utf8.count + 1) | |
| 126 | var index = source.startIndex | |
| 127 | while index < source.endIndex { | |
| 128 | chars.append(source[index]) | |
| 129 | starts.append(index) | |
| 130 | index = source.index(after: index) | |
| 131 | } | |
| 132 | starts.append(source.endIndex) | |
| 133 | self.chars = chars | |
| 134 | self.starts = starts | |
| 135 | } | |
| 136 | ||
| 137 | func scan(_ range: Range<Int>, into b: inout GreenBuilder, inLink: Bool = false) { | |
| 138 | var textStart = range.lowerBound | |
| 139 | var i = range.lowerBound | |
| 140 | while i < range.upperBound { | |
| 141 | if Self.mayStartObject(chars[i]), let match = match(at: i, in: range, inLink: inLink) { | |
| 142 | emitText(textStart..<i, into: &b) | |
| 143 | emit(match, into: &b, inLink: inLink) | |
| 144 | i = match.end | |
| 145 | textStart = i | |
| 146 | } else { | |
| 147 | i += 1 | |
| 148 | } | |
| 149 | } | |
| 150 | emitText(textStart..<range.upperBound, into: &b) | |
| 151 | } | |
| 152 | ||
| 153 | /// The first characters any recognizer accepts. | |
| 154 | static func mayStartObject(_ c: Unicode.Scalar) -> Bool { | |
| 155 | switch c { | |
| 156 | case "[", "<", "{", "\\", "^", "s", "h", "m", "f", "*", "/", "_", "+", "=", "~": true | |
| 157 | default: false | |
| 158 | } | |
| 159 | } | |
| 160 | ||
| 161 | /// Text with each line break as its own newline token; "\r\n" stays one token. | |
| 162 | func emitText(_ range: Range<Int>, into b: inout GreenBuilder) { | |
| 163 | var start = range.lowerBound | |
| 164 | var k = range.lowerBound | |
| 165 | while k < range.upperBound { | |
| 166 | if chars[k] == "\n" { | |
| 167 | let breakStart = k > start && chars[k - 1] == "\r" ? k - 1 : k | |
| 168 | if start < breakStart { b.token(.text, string(start..<breakStart)) } | |
| 169 | b.token(.newline, string(breakStart..<(k + 1))) | |
| 170 | start = k + 1 | |
| 171 | } | |
| 172 | k += 1 | |
| 173 | } | |
| 174 | if start < range.upperBound { b.token(.text, string(start..<range.upperBound)) } | |
| 175 | } | |
| 176 | ||
| 177 | func emit(_ match: InlineMatch, into b: inout GreenBuilder, inLink: Bool) { | |
| 178 | switch match { | |
| 179 | case .emphasis(let kind, let open, let close): | |
| 180 | b.start(kind) | |
| 181 | b.token(.marker, string(open..<(open + 1))) | |
| 182 | if kind == .verbatim || kind == .code { | |
| 183 | emitText((open + 1)..<close, into: &b) | |
| 184 | } else { | |
| 185 | scan((open + 1)..<close, into: &b, inLink: inLink) | |
| 186 | } | |
| 187 | b.token(.marker, string(close..<(close + 1))) | |
| 188 | b.finish() | |
| 189 | case .link(let path, let description, let whole): | |
| 190 | b.start(.link) | |
| 191 | b.token(.marker, string(whole.lowerBound..<path.lowerBound)) | |
| 192 | b.token(.linkPath, string(path)) | |
| 193 | if let description { | |
| 194 | b.token(.marker, string(path.upperBound..<description.lowerBound)) | |
| 195 | b.start(.linkDescription) | |
| 196 | scan(description, into: &b, inLink: true) | |
| 197 | b.finish() | |
| 198 | b.token(.marker, string(description.upperBound..<whole.upperBound)) | |
| 199 | } else { | |
| 200 | b.token(.marker, string(path.upperBound..<whole.upperBound)) | |
| 201 | } | |
| 202 | b.finish() | |
| 203 | case .object(let kind, let range): | |
| 204 | b.start(kind) | |
| 205 | emitText(range, into: &b) | |
| 206 | b.finish() | |
| 207 | } | |
| 208 | } | |
| 209 | ||
| 210 | func string(_ range: Range<Int>) -> String { | |
| 211 | String(source[starts[range.lowerBound]..<starts[range.upperBound]]) | |
| 212 | } | |
| 213 | ||
| 214 | func hasPrefix(_ s: [Unicode.Scalar], at i: Int, _ limit: Int) -> Bool { | |
| 215 | guard i + s.count <= limit else { return false } | |
| 216 | for (offset, c) in s.enumerated() where chars[i + offset] != c { return false } | |
| 217 | return true | |
| 218 | } | |
| 219 | ||
| 220 | func hasPrefix(_ s: String, at i: Int, _ limit: Int) -> Bool { | |
| 221 | hasPrefix(Array(s.unicodeScalars), at: i, limit) | |
| 222 | } | |
| 223 | ||
| 224 | // MARK: - Recognizers | |
| 225 | ||
| 226 | func match(at i: Int, in range: Range<Int>, inLink: Bool) -> InlineMatch? { | |
| 227 | let limit = range.upperBound | |
| 228 | let previous: Unicode.Scalar? = i > range.lowerBound ? chars[i - 1] : nil | |
| 229 | let afterWord = previous.map(isWordScalar) ?? false | |
| 230 | ||
| 231 | switch chars[i] { | |
| 232 | case "[": | |
| 233 | if !inLink, let link = bracketLink(i, limit) { return link } | |
| 234 | if let end = footnoteReference(i, limit) { return .object(.footnoteReference, i..<end) } | |
| 235 | if let end = scanTimestamp(chars, at: i, limit: limit)?.end { return .object(.timestamp, i..<end) } | |
| 236 | if let end = statisticsCookie(i, limit) { return .object(.statisticsCookie, i..<end) } | |
| 237 | case "<": | |
| 238 | if let end = scanTimestamp(chars, at: i, limit: limit)?.end { return .object(.timestamp, i..<end) } | |
| 239 | if let end = target(i, limit) { return .object(.target, i..<end) } | |
| 240 | if !inLink, let end = angleLink(i, limit) { return .object(.link, i..<end) } | |
| 241 | case "{": | |
| 242 | if let end = macro(i, limit) { return .object(.macro, i..<end) } | |
| 243 | case "\\": | |
| 244 | if let end = lineBreak(i, limit) { return .object(.lineBreak, i..<end) } | |
| 245 | if let end = latexFragment(i, limit) { return .object(.latexFragment, i..<end) } | |
| 246 | case "^": | |
| 247 | if afterWord, let end = superscript(i, limit) { return .object(.superscript, i..<end) } | |
| 248 | case "s": | |
| 249 | if !afterWord, let end = inlineSourceBlock(i, limit) { return .object(.inlineSourceBlock, i..<end) } | |
| 250 | case "h", "m", "f": | |
| 251 | if !inLink, !afterWord, let end = plainLink(i, limit) { return .object(.link, i..<end) } | |
| 252 | default: | |
| 253 | break | |
| 254 | } | |
| 255 | ||
| 256 | if let kind = emphasisKind(chars[i]), | |
| 257 | previous.map({ isSpace($0) || emphasisPre.contains($0) }) ?? true, | |
| 258 | let close = emphasisClose(i, limit) { | |
| 259 | return .emphasis(kind, open: i, close: close) | |
| 260 | } | |
| 261 | return nil | |
| 262 | } | |
| 263 | ||
| 264 | /// org's emphasis rules: the body neither starts nor ends with whitespace, spans at most | |
| 265 | /// one line break, and the closer is followed by whitespace, punctuation or the end. | |
| 266 | func emphasisClose(_ i: Int, _ limit: Int) -> Int? { | |
| 267 | let marker = chars[i] | |
| 268 | guard i + 1 < limit, !isSpace(chars[i + 1]) else { return nil } | |
| 269 | var newlines = 0 | |
| 270 | var j = i + 1 | |
| 271 | while j < limit { | |
| 272 | if isNewline(chars[j]) { | |
| 273 | newlines += 1 | |
| 274 | if newlines > 1 { return nil } | |
| 275 | } else if chars[j] == marker, j > i + 1, !isSpace(chars[j - 1]) { | |
| 276 | if j + 1 == limit || isSpace(chars[j + 1]) || emphasisPost.contains(chars[j + 1]) { return j } | |
| 277 | } | |
| 278 | j += 1 | |
| 279 | } | |
| 280 | return nil | |
| 281 | } | |
| 282 | ||
| 283 | /// `[[path]]` or `[[path][description]]`. | |
| 284 | func bracketLink(_ i: Int, _ limit: Int) -> InlineMatch? { | |
| 285 | guard i + 1 < limit, chars[i + 1] == "[" else { return nil } | |
| 286 | var j = i + 2 | |
| 287 | while j < limit, chars[j] != "]" { | |
| 288 | if chars[j] == "[" || isNewline(chars[j]) { return nil } | |
| 289 | if chars[j] == "\\", j + 1 < limit { j += 1 } | |
| 290 | j += 1 | |
| 291 | } | |
| 292 | guard j > i + 2, j + 1 < limit else { return nil } | |
| 293 | let path = (i + 2)..<j | |
| 294 | if chars[j + 1] == "]" { return .link(path: path, description: nil, whole: i..<(j + 2)) } | |
| 295 | guard chars[j + 1] == "[" else { return nil } | |
| 296 | let descriptionStart = j + 2 | |
| 297 | var depth = 0 | |
| 298 | var k = descriptionStart | |
| 299 | while k < limit { | |
| 300 | if chars[k] == "[" { | |
| 301 | depth += 1 | |
| 302 | } else if chars[k] == "]" { | |
| 303 | if depth == 0 { break } | |
| 304 | depth -= 1 | |
| 305 | } | |
| 306 | k += 1 | |
| 307 | } | |
| 308 | guard k > descriptionStart, k + 1 < limit, chars[k + 1] == "]" else { return nil } | |
| 309 | return .link(path: path, description: descriptionStart..<k, whole: i..<(k + 2)) | |
| 310 | } | |
| 311 | ||
| 312 | /// `[fn:label]`, `[fn:label:definition]` or `[fn::definition]`. | |
| 313 | func footnoteReference(_ i: Int, _ limit: Int) -> Int? { | |
| 314 | guard hasPrefix("[fn:", at: i, limit) else { return nil } | |
| 315 | var j = i + 4 | |
| 316 | while j < limit, isWordScalar(chars[j]) || chars[j] == "_" || chars[j] == "-" { j += 1 } | |
| 317 | guard j < limit else { return nil } | |
| 318 | if chars[j] == "]" { return j > i + 4 ? j + 1 : nil } | |
| 319 | guard chars[j] == ":" else { return nil } | |
| 320 | var depth = 0 | |
| 321 | j += 1 | |
| 322 | while j < limit { | |
| 323 | if chars[j] == "[" { | |
| 324 | depth += 1 | |
| 325 | } else if chars[j] == "]" { | |
| 326 | if depth == 0 { return j + 1 } | |
| 327 | depth -= 1 | |
| 328 | } | |
| 329 | j += 1 | |
| 330 | } | |
| 331 | return nil | |
| 332 | } | |
| 333 | ||
| 334 | /// `[1/3]`, `[/]`, `[50%]` or `[%]`. | |
| 335 | func statisticsCookie(_ i: Int, _ limit: Int) -> Int? { | |
| 336 | var j = i + 1 | |
| 337 | while j < limit, isASCIIDigit(chars[j]) { j += 1 } | |
| 338 | guard j < limit else { return nil } | |
| 339 | if chars[j] == "%" { | |
| 340 | j += 1 | |
| 341 | } else if chars[j] == "/" { | |
| 342 | j += 1 | |
| 343 | while j < limit, isASCIIDigit(chars[j]) { j += 1 } | |
| 344 | } else { | |
| 345 | return nil | |
| 346 | } | |
| 347 | guard j < limit, chars[j] == "]" else { return nil } | |
| 348 | return j + 1 | |
| 349 | } | |
| 350 | ||
| 351 | /// `<<target>>`. | |
| 352 | func target(_ i: Int, _ limit: Int) -> Int? { | |
| 353 | guard hasPrefix("<<", at: i, limit), i + 2 < limit, chars[i + 2] != "<" else { return nil } | |
| 354 | let start = i + 2 | |
| 355 | var j = start | |
| 356 | while j < limit, chars[j] != ">" { | |
| 357 | if chars[j] == "<" || isNewline(chars[j]) { return nil } | |
| 358 | j += 1 | |
| 359 | } | |
| 360 | guard j > start, j + 1 < limit, chars[j + 1] == ">", | |
| 361 | !isSpace(chars[start]), !isSpace(chars[j - 1]) else { return nil } | |
| 362 | return j + 2 | |
| 363 | } | |
| 364 | ||
| 365 | /// `<scheme:path>` for a known scheme. | |
| 366 | func angleLink(_ i: Int, _ limit: Int) -> Int? { | |
| 367 | var j = i + 1 | |
| 368 | while j < limit, chars[j].properties.isAlphabetic { j += 1 } | |
| 369 | guard j < limit, chars[j] == ":", angleLinkSchemes.contains(string((i + 1)..<j).lowercased()) else { return nil } | |
| 370 | j += 1 | |
| 371 | let bodyStart = j | |
| 372 | while j < limit, chars[j] != ">" { | |
| 373 | if chars[j] == "<" || isNewline(chars[j]) { return nil } | |
| 374 | j += 1 | |
| 375 | } | |
| 376 | guard j < limit, j > bodyStart else { return nil } | |
| 377 | return j + 1 | |
| 378 | } | |
| 379 | ||
| 380 | /// A bare URL. Trailing sentence punctuation stays outside the link. | |
| 381 | func plainLink(_ i: Int, _ limit: Int) -> Int? { | |
| 382 | guard let prefix = plainLinkPrefixes.first(where: { hasPrefix($0, at: i, limit) }) else { return nil } | |
| 383 | let bodyStart = i + prefix.count | |
| 384 | var j = bodyStart | |
| 385 | while j < limit, !isSpace(chars[j]), !"()<>[]\"".unicodeScalars.contains(chars[j]) { j += 1 } | |
| 386 | while j > bodyStart, ".,;:!?'".unicodeScalars.contains(chars[j - 1]) { j -= 1 } | |
| 387 | return j > bodyStart ? j : nil | |
| 388 | } | |
| 389 | ||
| 390 | /// `{{{name}}}` or `{{{name(arguments)}}}`. | |
| 391 | func macro(_ i: Int, _ limit: Int) -> Int? { | |
| 392 | guard hasPrefix("{{{", at: i, limit) else { return nil } | |
| 393 | var j = i + 3 | |
| 394 | guard j < limit, chars[j].properties.isAlphabetic else { return nil } | |
| 395 | while j < limit, isWordScalar(chars[j]) || chars[j] == "-" || chars[j] == "_" { j += 1 } | |
| 396 | if j < limit, chars[j] == "(" { | |
| 397 | var k = j + 1 | |
| 398 | while k < limit, !hasPrefix(")}}}", at: k, limit) { | |
| 399 | if isNewline(chars[k]) { return nil } | |
| 400 | k += 1 | |
| 401 | } | |
| 402 | return k < limit ? k + 4 : nil | |
| 403 | } | |
| 404 | return hasPrefix("}}}", at: j, limit) ? j + 3 : nil | |
| 405 | } | |
| 406 | ||
| 407 | /// `\\` at the end of a line, before optional trailing blanks. The line break stays outside. | |
| 408 | func lineBreak(_ i: Int, _ limit: Int) -> Int? { | |
| 409 | guard hasPrefix("\\\\", at: i, limit) else { return nil } | |
| 410 | var j = i + 2 | |
| 411 | while j < limit, chars[j] == " " || chars[j] == "\t" { j += 1 } | |
| 412 | if j + 1 < limit, chars[j] == "\r", chars[j + 1] == "\n" { return j } | |
| 413 | guard j == limit || isNewline(chars[j]) else { return nil } | |
| 414 | return j | |
| 415 | } | |
| 416 | ||
| 417 | /// `\(...\)` or `\[...\]`. | |
| 418 | func latexFragment(_ i: Int, _ limit: Int) -> Int? { | |
| 419 | guard i + 1 < limit else { return nil } | |
| 420 | let closer: String | |
| 421 | switch chars[i + 1] { | |
| 422 | case "(": closer = "\\)" | |
| 423 | case "[": closer = "\\]" | |
| 424 | default: return nil | |
| 425 | } | |
| 426 | let closing = Array(closer.unicodeScalars) | |
| 427 | var j = i + 2 | |
| 428 | while j < limit { | |
| 429 | if hasPrefix(closing, at: j, limit) { return j + 2 } | |
| 430 | j += 1 | |
| 431 | } | |
| 432 | return nil | |
| 433 | } | |
| 434 | ||
| 435 | /// `^word` or `^{group}` after a letter or digit. | |
| 436 | func superscript(_ i: Int, _ limit: Int) -> Int? { | |
| 437 | var j = i + 1 | |
| 438 | guard j < limit else { return nil } | |
| 439 | if chars[j] == "{" { | |
| 440 | j += 1 | |
| 441 | while j < limit, chars[j] != "}" { | |
| 442 | if isNewline(chars[j]) { return nil } | |
| 443 | j += 1 | |
| 444 | } | |
| 445 | return j < limit ? j + 1 : nil | |
| 446 | } | |
| 447 | let start = j | |
| 448 | while j < limit, isWordScalar(chars[j]) { j += 1 } | |
| 449 | return j > start ? j : nil | |
| 450 | } | |
| 451 | ||
| 452 | /// `src_lang{body}` or `src_lang[headers]{body}`, on one line, with balanced braces. | |
| 453 | func inlineSourceBlock(_ i: Int, _ limit: Int) -> Int? { | |
| 454 | guard hasPrefix("src_", at: i, limit) else { return nil } | |
| 455 | var j = i + 4 | |
| 456 | let languageStart = j | |
| 457 | while j < limit, !isSpace(chars[j]), chars[j] != "[", chars[j] != "{" { j += 1 } | |
| 458 | guard j > languageStart, j < limit else { return nil } | |
| 459 | if chars[j] == "[" { | |
| 460 | while j < limit, chars[j] != "]" { | |
| 461 | if isNewline(chars[j]) { return nil } | |
| 462 | j += 1 | |
| 463 | } | |
| 464 | guard j < limit else { return nil } | |
| 465 | j += 1 | |
| 466 | } | |
| 467 | guard j < limit, chars[j] == "{" else { return nil } | |
| 468 | var depth = 0 | |
| 469 | while j < limit { | |
| 470 | if isNewline(chars[j]) { return nil } | |
| 471 | if chars[j] == "{" { | |
| 472 | depth += 1 | |
| 473 | } else if chars[j] == "}" { | |
| 474 | depth -= 1 | |
| 475 | if depth == 0 { return j + 1 } | |
| 476 | } | |
| 477 | j += 1 | |
| 478 | } | |
| 479 | return nil | |
| 480 | } | |
| 481 | } | |
| 482 | ||
| 483 | extension Parser { | |
| 484 | mutating func inline(_ text: Substring) { | |
| 485 | let scanner = InlineScanner(text) | |
| 486 | scanner.scan(0..<scanner.chars.count, into: &builder) | |
| 487 | } | |
| 488 | } | |
| 489 | ``` | |
| 490 | ||
| 491 | ```diff | |
| 492 | diff --git a/Sources/OrgCore/Timestamp.swift b/Sources/OrgCore/Timestamp.swift | |
| 493 | index da5ac94..bd71851 100644 | |
| 494 | --- a/Sources/OrgCore/Timestamp.swift | |
| 495 | +++ b/Sources/OrgCore/Timestamp.swift | |
| 496 | @@ -42,14 +42,14 @@ public struct Timestamp: Sendable, Equatable { | |
| 497 | ||
| 498 | /// Parses exactly one timestamp or range, with nothing before or after it. | |
| 499 | public static func parse(_ text: some StringProtocol) -> Timestamp? { | |
| 500 | - let chars = Array(text) | |
| 501 | + let chars = Array(text.unicodeScalars) | |
| 502 | guard let result = scanTimestamp(chars, at: 0, limit: chars.count), result.end == chars.count else { return nil } | |
| 503 | return result.stamp | |
| 504 | } | |
| 505 | } | |
| 506 | ||
| 507 | /// A timestamp or `--` range starting at `start`, and the index after it. | |
| 508 | -func scanTimestamp(_ chars: [Character], at start: Int, limit: Int) -> (stamp: Timestamp, end: Int)? { | |
| 509 | +func scanTimestamp(_ chars: [Unicode.Scalar], at start: Int, limit: Int) -> (stamp: Timestamp, end: Int)? { | |
| 510 | guard let first = scanSingleTimestamp(chars, at: start, limit: limit) else { return nil } | |
| 511 | if first.stamp.end == nil, first.end + 2 < limit, chars[first.end] == "-", chars[first.end + 1] == "-", | |
| 512 | let second = scanSingleTimestamp(chars, at: first.end + 2, limit: limit), | |
| 513 | @@ -62,22 +62,23 @@ func scanTimestamp(_ chars: [Character], at start: Int, limit: Int) -> (stamp: T | |
| 514 | } | |
| 515 | ||
| 516 | /// `<YYYY-MM-DD DAY HH:MM-HH:MM REPEATER WARNING>`, or the same in `[...]` for inactive. | |
| 517 | -private func scanSingleTimestamp(_ chars: [Character], at start: Int, limit: Int) -> (stamp: Timestamp, end: Int)? { | |
| 518 | +private func scanSingleTimestamp(_ chars: [Unicode.Scalar], at start: Int, limit: Int) -> (stamp: Timestamp, end: Int)? { | |
| 519 | guard start < limit, chars[start] == "<" || chars[start] == "[" else { return nil } | |
| 520 | let active = chars[start] == "<" | |
| 521 | - let close: Character = active ? ">" : "]" | |
| 522 | + let close: Unicode.Scalar = active ? ">" : "]" | |
| 523 | var j = start + 1 | |
| 524 | ||
| 525 | func number(_ minDigits: Int, _ maxDigits: Int) -> Int? { | |
| 526 | var k = j | |
| 527 | - while k < limit, k - j < maxDigits, chars[k].isASCII, chars[k].isNumber { k += 1 } | |
| 528 | + while k < limit, k - j < maxDigits, isASCIIDigit(chars[k]) { k += 1 } | |
| 529 | guard k - j >= minDigits else { return nil } | |
| 530 | - let value = Int(String(chars[j..<k]))! | |
| 531 | + var value = 0 | |
| 532 | + for digit in chars[j..<k] { value = value * 10 + Int(digit.value - 48) } | |
| 533 | j = k | |
| 534 | return value | |
| 535 | } | |
| 536 | ||
| 537 | - func take(_ c: Character) -> Bool { | |
| 538 | + func take(_ c: Unicode.Scalar) -> Bool { | |
| 539 | guard j < limit, chars[j] == c else { return false } | |
| 540 | j += 1 | |
| 541 | return true | |
| 542 | @@ -85,7 +86,7 @@ private func scanSingleTimestamp(_ chars: [Character], at start: Int, limit: Int | |
| 543 | ||
| 544 | func interval() -> Timestamp.Interval? { | |
| 545 | let before = j | |
| 546 | - guard let value = number(1, 9), j < limit, let unit = Timestamp.Unit(rawValue: chars[j]) else { | |
| 547 | + guard let value = number(1, 9), j < limit, let unit = Timestamp.Unit(rawValue: Character(chars[j])) else { | |
| 548 | j = before | |
| 549 | return nil | |
| 550 | } | |
| 551 | @@ -145,7 +146,7 @@ private func scanSingleTimestamp(_ chars: [Character], at start: Int, limit: Int | |
| 552 | // Day name: anything but digits, whitespace, `+`, `-`, `]` and `>`, in any language. | |
| 553 | if j < limit, chars[j] == " " { | |
| 554 | var k = j + 1 | |
| 555 | - while k < limit, !(chars[k].isNumber || chars[k].isWhitespace || "+-]>".contains(chars[k])) { k += 1 } | |
| 556 | + while k < limit, !(isDigit(chars[k]) || chars[k].properties.isWhitespace || "+-]>".unicodeScalars.contains(chars[k])) { k += 1 } | |
| 557 | if k > j + 1 { j = k } | |
| 558 | } | |
| 559 | ||
| 560 | @@ -179,3 +180,12 @@ private func scanSingleTimestamp(_ chars: [Character], at start: Int, limit: Int | |
| 561 | guard take(close) else { return nil } | |
| 562 | return (stamp, j) | |
| 563 | } | |
| 564 | + | |
| 565 | +func isASCIIDigit(_ c: Unicode.Scalar) -> Bool { | |
| 566 | + c.value >= 48 && c.value <= 57 | |
| 567 | +} | |
| 568 | + | |
| 569 | +/// Any numeric character, as `Character.isNumber` would see it in a day name. | |
| 570 | +func isDigit(_ c: Unicode.Scalar) -> Bool { | |
| 571 | + c.properties.numericType != nil | |
| 572 | +} | |
| 573 | ``` | |
| 574 | ||
| 575 | - [ ] **Step 2: Run tests, the corpus and the reparse gate; commit** | |
| 576 | ||
| 577 | Run: `swift test`; `ORGSTAR_CORPUS=<folder> swift test -c release --filter corpusRoundTrips` for each local folder; `ORGSTAR_FUZZ_EDITS=100000 swift test -c release --filter incrementalEqualsFull`. | |
| 578 | ||
| 579 | ```bash | |
| 580 | git add Sources/OrgCore/Parser/Inline.swift Sources/OrgCore/Timestamp.swift | |
| 581 | git commit -m "Scan inline objects over unicode scalars" | |
| 582 | ``` | |
| 583 | ||
| 584 | --- | |
| 585 | ||
| 586 | ### Task 2: Region reparse | |
| 587 | ||
| 588 | **Files:** | |
| 589 | - Modify: `Sources/OrgCore/Syntax/SyntaxNode.swift`, `Sources/OrgCore/Syntax/GreenTree.swift`, `Sources/OrgCore/Parser/Parser.swift`, `Sources/OrgCore/Parser/Incremental.swift` | |
| 590 | - Test: `Tests/OrgCoreTests/IncrementalTests.swift` | |
| 591 | ||
| 592 | **Interfaces:** | |
| 593 | - Produces: `ReparseStrategy.region`; `SyntaxNode.children(overlapping:)`, `child(containing:)`, `firstChild(_:)`; `GreenBuilder.node(_:)`; `Parser.reusableRows`. | |
| 594 | ||
| 595 | - [ ] **Step 1: Write the failing tests** | |
| 596 | ||
| 597 | ```diff | |
| 598 | @@ -49,8 +49,23 @@ struct IncrementalTests { | |
| 599 | #expect(checkReparse("* a\nx\ny\n* b\n", 6..<6, "* c\n") == .sections) | |
| 600 | } | |
| 601 | ||
| 602 | - @Test func joiningParagraphsReparsesSections() { | |
| 603 | - #expect(checkReparse("a\n\nb\n", 1..<2, "") == .sections) | |
| 604 | + @Test func joiningParagraphsReparsesARegion() { | |
| 605 | + #expect(checkReparse("a\n\nb\n", 1..<2, "") == .region) | |
| 606 | + #expect(checkReparse("* h\nfirst\nsecond\n\nthird\n** child\n", 15..<15, "\n\n") == .region) | |
| 607 | + } | |
| 608 | + | |
| 609 | + @Test func regionGrowsUntilABoundaryHolds() { | |
| 610 | + // Splitting a paragraph next to a table and a list. | |
| 611 | + #expect(checkReparse("* h\nx\ntext\nmore\n| a |\n- i\n", 9..<9, "\n") == .region) | |
| 612 | + // A new begin line without its end runs to the end of the body. | |
| 613 | + #expect(checkReparse("* h\nx\none\ntwo\n#+end_src\nthree\n", 6..<6, "\n#+begin_src") == .region) | |
| 614 | + } | |
| 615 | + | |
| 616 | + @Test func regionDefersToSectionsNearHeadings() { | |
| 617 | + // The first body line could become a planning line. | |
| 618 | + #expect(checkReparse("* h\nbody\n", 4..<4, "\n") != .region) | |
| 619 | + // An added end delimiter could close a block that starts before the window. | |
| 620 | + #expect(checkReparse("x\n#+begin_src\na\n\nb\n", 16..<16, "#+end_src\n") != .region) | |
| 621 | } | |
| 622 | ||
| 623 | @Test func runGrowsUntilAHeadingEndsIt() { | |
| 624 | ``` | |
| 625 | ||
| 626 | - [ ] **Step 2: Implement** | |
| 627 | ||
| 628 | ```diff | |
| 629 | diff --git a/Sources/OrgCore/Parser/Incremental.swift b/Sources/OrgCore/Parser/Incremental.swift | |
| 630 | index b34284a..9f8347a 100644 | |
| 631 | --- a/Sources/OrgCore/Parser/Incremental.swift | |
| 632 | +++ b/Sources/OrgCore/Parser/Incremental.swift | |
| 633 | @@ -22,6 +22,8 @@ public struct TextEdit: Sendable, Equatable { | |
| 634 | enum ReparseStrategy: Equatable { | |
| 635 | /// One element reparsed in place; everything else reused. | |
| 636 | case element | |
| 637 | + /// A run of elements in one section's body reparsed; everything else reused. | |
| 638 | + case region | |
| 639 | /// A run of top-level sections reparsed; the sections before and after reused. | |
| 640 | case sections | |
| 641 | case full | |
| 642 | @@ -41,6 +43,7 @@ extension OrgParser { | |
| 643 | let context = EditContext(oldText: oldText, newText: newText, edit: edit) | |
| 644 | if context.touchesSettings() { return (parse(newText, defaults: defaults), .full) } | |
| 645 | if let tree = context.reparseElement(old) { return (tree, .element) } | |
| 646 | + if let tree = context.reparseRegion(old) { return (tree, .region) } | |
| 647 | if let tree = context.reparseSections(old) { return (tree, .sections) } | |
| 648 | return (parse(newText, defaults: defaults), .full) | |
| 649 | } | |
| 650 | @@ -113,10 +116,7 @@ private struct EditContext { | |
| 651 | func leaf(in root: SyntaxNode) -> SyntaxNode? { | |
| 652 | let range = edit.range | |
| 653 | var current = root | |
| 654 | - while let child = current.children.first(where: { | |
| 655 | - $0.range.lowerBound <= range.lowerBound && range.lowerBound < $0.range.upperBound | |
| 656 | - && range.upperBound <= $0.range.upperBound | |
| 657 | - }) { | |
| 658 | + while let child = current.child(containing: range.lowerBound), range.upperBound <= child.range.upperBound { | |
| 659 | if Self.leafKinds.contains(child.kind) { return child } | |
| 660 | current = child | |
| 661 | } | |
| 662 | @@ -140,6 +140,151 @@ private struct EditContext { | |
| 663 | return replacement | |
| 664 | } | |
| 665 | ||
| 666 | + // MARK: - Body region | |
| 667 | + | |
| 668 | + /// Reparses a window of the innermost section's own body (the elements between its heading | |
| 669 | + /// lines and its first child section) and splices it into the old section. | |
| 670 | + /// | |
| 671 | + /// The window starts one element before the element holding the line before the edit. It ends at the first | |
| 672 | + /// later element boundary that the reparse reproduces at the same shifted offset. The | |
| 673 | + /// window is parsed one element past that boundary, so every decision about where an element | |
| 674 | + /// ends was made with the following line in view; from a reproduced boundary on, elements | |
| 675 | + /// parse the same as before. Edits that touch heading lines, the first line after a heading | |
| 676 | + /// (where planning lines and property drawers are only recognized), or that add an end | |
| 677 | + /// delimiter (which could pair with a begin line before the window) are left to the | |
| 678 | + /// section strategy. | |
| 679 | + func reparseRegion(_ old: OrgTree) -> OrgTree? { | |
| 680 | + let start = lineStart(oldText, edit.range.lowerBound) | |
| 681 | + let oldLines = classes(slice(oldText, start, lineEnd(oldText, edit.range.upperBound))) | |
| 682 | + let newLines = classes(slice(newText, start, lineEnd(newText, edit.range.lowerBound + edit.replacement.utf16.count))) | |
| 683 | + for line in oldLines + newLines { | |
| 684 | + if case .heading = line.cls { return nil } | |
| 685 | + } | |
| 686 | + for line in newLines { | |
| 687 | + switch line.cls { | |
| 688 | + case .blockEnd, .dynamicEnd, .drawerEnd: return nil | |
| 689 | + default: break | |
| 690 | + } | |
| 691 | + } | |
| 692 | + | |
| 693 | + // The innermost section or zeroth section holding the edited lines. | |
| 694 | + let damaged = start..<max(lineEnd(oldText, edit.range.upperBound), start + 1) | |
| 695 | + var container = old.root | |
| 696 | + while let child = container.child(containing: damaged.lowerBound), | |
| 697 | + child.kind == .section || child.kind == .zerothSection, | |
| 698 | + damaged.upperBound <= child.range.upperBound { | |
| 699 | + container = child | |
| 700 | + } | |
| 701 | + guard container.kind == .section || container.kind == .zerothSection else { return nil } | |
| 702 | + | |
| 703 | + // Body children with their offsets. | |
| 704 | + var body: [(element: GreenElement, offset: Int)] = [] | |
| 705 | + var bodyStart = container.range.lowerBound | |
| 706 | + var bodyEnd = container.range.upperBound | |
| 707 | + var offset = container.range.lowerBound | |
| 708 | + var prefixCount = 0 | |
| 709 | + for (index, element) in container.green.children.enumerated() { | |
| 710 | + defer { offset += element.length } | |
| 711 | + if case .node(let node) = element { | |
| 712 | + if [.heading, .planning, .propertyDrawer].contains(node.kind), body.isEmpty { | |
| 713 | + prefixCount = index + 1 | |
| 714 | + bodyStart = offset + node.length | |
| 715 | + continue | |
| 716 | + } | |
| 717 | + if node.kind == .section { | |
| 718 | + bodyEnd = offset | |
| 719 | + break | |
| 720 | + } | |
| 721 | + } | |
| 722 | + body.append((element, offset)) | |
| 723 | + } | |
| 724 | + // The first body line of a heading section can become a planning line or drawer. | |
| 725 | + let firstEditable = container.kind == .section ? lineEnd(oldText, bodyStart) : bodyStart | |
| 726 | + guard start >= firstEditable, edit.range.upperBound <= bodyEnd, !body.isEmpty else { return nil } | |
| 727 | + | |
| 728 | + // Window start: the body element holding the line before the edit, and one more before | |
| 729 | + // it, which could absorb the window's first line if that line changes meaning (a drawer | |
| 730 | + // that loses its end turns into paragraph text). | |
| 731 | + let anchor = max(bodyStart, start - 1) | |
| 732 | + guard var first = body.lastIndex(where: { $0.offset <= anchor }) else { return nil } | |
| 733 | + while first > 0, case .token = body[first].element { first -= 1 } | |
| 734 | + if first > 0 { | |
| 735 | + first -= 1 | |
| 736 | + while first > 0, case .token = body[first].element { first -= 1 } | |
| 737 | + } | |
| 738 | + let windowStart = body[first].offset | |
| 739 | + | |
| 740 | + var next = body.firstIndex { $0.offset > edit.range.upperBound } ?? body.count | |
| 741 | + var attempts = 0 | |
| 742 | + while attempts < 8 { | |
| 743 | + attempts += 1 | |
| 744 | + // Parse through one node past the candidate boundary. | |
| 745 | + var verifyEnd = next | |
| 746 | + while verifyEnd < body.count, case .token = body[verifyEnd].element { verifyEnd += 1 } | |
| 747 | + let boundary = next < body.count ? body[next].offset : bodyEnd | |
| 748 | + let oldWindowEnd = verifyEnd < body.count ? body[verifyEnd].offset + body[verifyEnd].element.length : bodyEnd | |
| 749 | + let rows = reusableRows(body[first..<min(verifyEnd + 1, body.count)].map(\.element)) | |
| 750 | + guard let window = parseBody(slice(newText, windowStart, oldWindowEnd + delta), settings: old.settings, rows: rows) else { return nil } | |
| 751 | + if window.unmatchedBegin, oldWindowEnd < bodyEnd { | |
| 752 | + next = body.count | |
| 753 | + continue | |
| 754 | + } | |
| 755 | + // The reparsed elements up to the boundary, if the reparse has one there. | |
| 756 | + let target = boundary + delta - windowStart | |
| 757 | + var reparsed: [GreenElement] = [] | |
| 758 | + var at = 0 | |
| 759 | + for element in window.elements where at < target { | |
| 760 | + reparsed.append(element) | |
| 761 | + at += element.length | |
| 762 | + } | |
| 763 | + guard at == target else { | |
| 764 | + next += 1 | |
| 765 | + if next > body.count { return nil } | |
| 766 | + continue | |
| 767 | + } | |
| 768 | + let before = container.green.children[..<(prefixCount + first)] | |
| 769 | + let after = next < body.count ? Array(body[next...].map(\.element)) : [] | |
| 770 | + let tail = container.green.children[(prefixCount + body.count)...] | |
| 771 | + let green = GreenNode(kind: container.kind, children: Array(before) + reparsed + after + Array(tail)) | |
| 772 | + // A full parse has no zeroth section for empty text. | |
| 773 | + if green.children.isEmpty { return nil } | |
| 774 | + return OrgTree(green: replacing(container, with: green), settings: old.settings) | |
| 775 | + } | |
| 776 | + return nil | |
| 777 | + } | |
| 778 | + | |
| 779 | + /// Rows of the tables among `elements`, keyed by their text. | |
| 780 | + func reusableRows(_ elements: [GreenElement]) -> [Substring: GreenNode] { | |
| 781 | + var rows: [Substring: GreenNode] = [:] | |
| 782 | + for case .node(let table) in elements where table.kind == .table { | |
| 783 | + for case .node(let row) in table.children where row.kind == .tableRow { | |
| 784 | + rows[Substring(row.text)] = row | |
| 785 | + } | |
| 786 | + } | |
| 787 | + return rows | |
| 788 | + } | |
| 789 | + | |
| 790 | + /// Body elements parsed from `text`, or nil if the text holds a heading. | |
| 791 | + func parseBody(_ text: Substring, settings: OrgSettings, rows: [Substring: GreenNode] = [:]) -> (elements: [GreenElement], unmatchedBegin: Bool)? { | |
| 792 | + var parser = Parser(text: String(text), settings: settings) | |
| 793 | + parser.reusableRows = rows | |
| 794 | + let count = parser.lines.count | |
| 795 | + parser.builder.start(.zerothSection) | |
| 796 | + parser.parseContent(limit: count) | |
| 797 | + guard parser.i == count else { return nil } | |
| 798 | + parser.builder.finish() | |
| 799 | + var unmatched = false | |
| 800 | + for (index, line) in parser.info.enumerated() { | |
| 801 | + switch line.cls { | |
| 802 | + case .blockBegin, .dynamicBegin, .drawerBegin: | |
| 803 | + if parser.blockEnds[index] == nil { unmatched = true } | |
| 804 | + default: | |
| 805 | + break | |
| 806 | + } | |
| 807 | + } | |
| 808 | + return (parser.builder.build().children, unmatched) | |
| 809 | + } | |
| 810 | + | |
| 811 | // MARK: - Sections | |
| 812 | ||
| 813 | /// Reparses from the top-level section before the edit up to the first later top-level | |
| 814 | diff --git a/Sources/OrgCore/Parser/Parser.swift b/Sources/OrgCore/Parser/Parser.swift | |
| 815 | index 8478c72..2f2addf 100644 | |
| 816 | --- a/Sources/OrgCore/Parser/Parser.swift | |
| 817 | +++ b/Sources/OrgCore/Parser/Parser.swift | |
| 818 | @@ -14,6 +14,9 @@ struct Parser { | |
| 819 | let settings: OrgSettings | |
| 820 | var builder = GreenBuilder() | |
| 821 | var i = 0 | |
| 822 | + /// Table rows from a previous version, by their line text. A row parses from its own line | |
| 823 | + /// alone, so a row with the same text can be reused instead of scanned again. | |
| 824 | + var reusableRows: [Substring: GreenNode] = [:] | |
| 825 | ||
| 826 | init(text: String, defaults: OrgSettings) { | |
| 827 | let lines = splitRawLines(text) | |
| 828 | @@ -230,6 +233,14 @@ struct Parser { | |
| 829 | /// A rule row (`|---+---|`) is one text token. Other rows alternate `|` markers and cells; | |
| 830 | /// every pair of pipes gets a cell, even an empty one, so columns line up. | |
| 831 | mutating func tableRow() { | |
| 832 | + if !reusableRows.isEmpty { | |
| 833 | + let content = lines[i].content | |
| 834 | + if let row = reusableRows[content.base[content.startIndex..<lines[i].ending.endIndex]] { | |
| 835 | + builder.node(row) | |
| 836 | + i += 1 | |
| 837 | + return | |
| 838 | + } | |
| 839 | + } | |
| 840 | builder.start(.tableRow) | |
| 841 | let rest = whitespace(lines[i].content) | |
| 842 | if rest.hasPrefix("|-") { | |
| 843 | diff --git a/Sources/OrgCore/Syntax/GreenTree.swift b/Sources/OrgCore/Syntax/GreenTree.swift | |
| 844 | index 10b11c6..47019e4 100644 | |
| 845 | --- a/Sources/OrgCore/Syntax/GreenTree.swift | |
| 846 | +++ b/Sources/OrgCore/Syntax/GreenTree.swift | |
| 847 | @@ -70,6 +70,11 @@ struct GreenBuilder { | |
| 848 | stack[stack.count - 1].children.append(.token(GreenToken(kind: kind, text: String(text)))) | |
| 849 | } | |
| 850 | ||
| 851 | + /// Appends an already-built node, for reuse across versions. | |
| 852 | + mutating func node(_ green: GreenNode) { | |
| 853 | + stack[stack.count - 1].children.append(.node(green)) | |
| 854 | + } | |
| 855 | + | |
| 856 | mutating func finish() { | |
| 857 | let (kind, children) = stack.removeLast() | |
| 858 | let node = GreenNode(kind: kind, children: children) | |
| 859 | diff --git a/Sources/OrgCore/Syntax/SyntaxNode.swift b/Sources/OrgCore/Syntax/SyntaxNode.swift | |
| 860 | index 3473728..0430409 100644 | |
| 861 | --- a/Sources/OrgCore/Syntax/SyntaxNode.swift | |
| 862 | +++ b/Sources/OrgCore/Syntax/SyntaxNode.swift | |
| 863 | @@ -38,6 +38,46 @@ public final class SyntaxNode: Sendable { | |
| 864 | return result | |
| 865 | } | |
| 866 | ||
| 867 | + /// Child nodes overlapping `range`. Only those are created, so a section with thousands of | |
| 868 | + /// children costs little when the range is small. | |
| 869 | + public func children(overlapping range: Range<Int>) -> [SyntaxNode] { | |
| 870 | + var result: [SyntaxNode] = [] | |
| 871 | + var at = offset | |
| 872 | + for child in green.children { | |
| 873 | + let childRange = at..<(at + child.length) | |
| 874 | + if childRange.lowerBound >= range.upperBound, !range.isEmpty { break } | |
| 875 | + if case .node(let node) = child, childRange.overlaps(range) { | |
| 876 | + result.append(SyntaxNode(green: node, offset: at, parent: self)) | |
| 877 | + } | |
| 878 | + at = childRange.upperBound | |
| 879 | + } | |
| 880 | + return result | |
| 881 | + } | |
| 882 | + | |
| 883 | + /// The child node whose range contains `position`. | |
| 884 | + public func child(containing position: Int) -> SyntaxNode? { | |
| 885 | + var at = offset | |
| 886 | + for child in green.children { | |
| 887 | + let end = at + child.length | |
| 888 | + if position < end { | |
| 889 | + if case .node(let node) = child, position >= at { return SyntaxNode(green: node, offset: at, parent: self) } | |
| 890 | + return nil | |
| 891 | + } | |
| 892 | + at = end | |
| 893 | + } | |
| 894 | + return nil | |
| 895 | + } | |
| 896 | + | |
| 897 | + /// The first child node of `kind`. | |
| 898 | + public func firstChild(_ kind: SyntaxKind) -> SyntaxNode? { | |
| 899 | + var at = offset | |
| 900 | + for child in green.children { | |
| 901 | + if case .node(let node) = child, node.kind == kind { return SyntaxNode(green: node, offset: at, parent: self) } | |
| 902 | + at += child.length | |
| 903 | + } | |
| 904 | + return nil | |
| 905 | + } | |
| 906 | + | |
| 907 | /// This node and every node below it, in document order. | |
| 908 | public func descendants() -> [SyntaxNode] { | |
| 909 | [self] + children.flatMap { $0.descendants() } | |
| 910 | ``` | |
| 911 | ||
| 912 | - [ ] **Step 3: Run the gate on several seeds; commit** | |
| 913 | ||
| 914 | Run: `ORGSTAR_FUZZ_SEED=<n> ORGSTAR_FUZZ_EDITS=100000 swift test -c release --filter incrementalEqualsFull` for at least five seeds. | |
| 915 | ||
| 916 | ```bash | |
| 917 | git add Sources/OrgCore Tests/OrgCoreTests/IncrementalTests.swift | |
| 918 | git commit -m "Reparse structural edits within one section's body" | |
| 919 | ``` | |
| 920 | ||
| 921 | --- | |
| 922 | ||
| 923 | ### Task 3: Editor styling | |
| 924 | ||
| 925 | **Files:** | |
| 926 | - Modify: `Sources/OrgEditorAppKit/OrgEditor.swift`, `Sources/OrgPresentation/Presentation.swift`, `Sources/OrgDocument/ViewState.swift` | |
| 927 | - Test: `Tests/OrgEditorAppKitTests/EditorTests.swift` | |
| 928 | ||
| 929 | **Interfaces:** | |
| 930 | - Produces: `OrgEditor.finishStyling()`, `OrgEditor.styledUpTo`, `OrgEditor.firstStyleChunk`, `OrgEditor.styleChunk`; test `Harness(_:layout:)`, `Harness.drainStyling()`. | |
| 931 | ||
| 932 | - [ ] **Step 1: Write the failing tests** | |
| 933 | ||
| 934 | ```diff | |
| 935 | @@ -10,13 +10,24 @@ final class Harness { | |
| 936 | let editor: OrgEditor | |
| 937 | let window: NSWindow | |
| 938 | ||
| 939 | - init(_ text: String) { | |
| 940 | + convenience init(_ text: String) { | |
| 941 | + self.init(DocumentState(bytes: Array(text.utf8))) | |
| 942 | + drainStyling() | |
| 943 | + layout() | |
| 944 | + } | |
| 945 | + | |
| 946 | + /// The editor as a window would show it first: background styling not yet run. | |
| 947 | + init(_ document: DocumentState, layout: Bool = true) { | |
| 948 | _ = NSApplication.shared | |
| 949 | - editor = OrgEditor(document: DocumentState(bytes: Array(text.utf8))) | |
| 950 | + editor = OrgEditor(document: document) | |
| 951 | window = NSWindow(contentRect: NSRect(x: 0, y: 0, width: 600, height: 800), styleMask: [.titled], backing: .buffered, defer: false) | |
| 952 | window.contentView = editor.makeScrollView() | |
| 953 | window.makeFirstResponder(editor.textView) | |
| 954 | - layout() | |
| 955 | + if layout { self.layout() } | |
| 956 | + } | |
| 957 | + | |
| 958 | + func drainStyling() { | |
| 959 | + editor.finishStyling() | |
| 960 | } | |
| 961 | ||
| 962 | var textView: NSTextView { editor.textView } | |
| 963 | @@ -246,3 +257,26 @@ struct RestyleFuzzTests { | |
| 964 | h.checkInSync() | |
| 965 | } | |
| 966 | } | |
| 967 | + | |
| 968 | +@MainActor | |
| 969 | +struct InitialStylingTests { | |
| 970 | + @Test func stylesTheFirstScreenThenTheRest() throws { | |
| 971 | + let line = "* heading *bold*\nbody text\n" | |
| 972 | + let text = String(repeating: line, count: 4_000) | |
| 973 | + let h = Harness(DocumentState(bytes: Array(text.utf8)), layout: false) | |
| 974 | + let length = (text as NSString).length | |
| 975 | + #expect(h.editor.styledUpTo >= OrgEditor.firstStyleChunk && h.editor.styledUpTo < length) | |
| 976 | + h.textView.insertText("x", replacementRange: NSRange(location: length - 3, length: 0)) | |
| 977 | + h.drainStyling() | |
| 978 | + #expect(h.editor.styledUpTo == length + 1) | |
| 979 | + #expect(h.textView.textStorage!.isEqual(to: Harness(h.string).textView.textStorage!)) | |
| 980 | + } | |
| 981 | + | |
| 982 | + @Test func editsBeforeTheWatermarkShiftIt() { | |
| 983 | + let text = String(repeating: "line of text\n", count: 5_000) | |
| 984 | + let h = Harness(DocumentState(bytes: Array(text.utf8)), layout: false) | |
| 985 | + let before = h.editor.styledUpTo | |
| 986 | + h.textView.insertText("abc", replacementRange: NSRange(location: 0, length: 0)) | |
| 987 | + #expect(h.editor.styledUpTo == before + 3) | |
| 988 | + } | |
| 989 | +} | |
| 990 | ``` | |
| 991 | ||
| 992 | - [ ] **Step 2: Implement** | |
| 993 | ||
| 994 | ```diff | |
| 995 | diff --git a/Sources/OrgDocument/ViewState.swift b/Sources/OrgDocument/ViewState.swift | |
| 996 | index 7590ffe..9cdf35a 100644 | |
| 997 | --- a/Sources/OrgDocument/ViewState.swift | |
| 998 | +++ b/Sources/OrgDocument/ViewState.swift | |
| 999 | @@ -30,7 +30,7 @@ public struct ViewState: Sendable, Equatable { | |
| 1000 | /// Descends only through the sections containing `offset`. | |
| 1001 | func isHeadingStart(_ offset: Int, _ root: SyntaxNode) -> Bool { | |
| 1002 | var node = root | |
| 1003 | - while let child = node.children.first(where: { $0.range.contains(offset) }) { | |
| 1004 | + while let child = node.child(containing: offset) { | |
| 1005 | if child.kind == .heading { return child.range.lowerBound == offset } | |
| 1006 | guard child.kind == .section || child.kind == .zerothSection else { return false } | |
| 1007 | node = child | |
| 1008 | diff --git a/Sources/OrgEditorAppKit/OrgEditor.swift b/Sources/OrgEditorAppKit/OrgEditor.swift | |
| 1009 | index ecea416..b7c3dc4 100644 | |
| 1010 | --- a/Sources/OrgEditorAppKit/OrgEditor.swift | |
| 1011 | +++ b/Sources/OrgEditorAppKit/OrgEditor.swift | |
| 1012 | @@ -55,6 +55,15 @@ public final class OrgEditor: NSObject { | |
| 1013 | let theme = Theme() | |
| 1014 | private var isLoading = false | |
| 1015 | private var previousSelection = 0 | |
| 1016 | + /// Text before this offset is styled; after it, only base attributes until the background | |
| 1017 | + /// passes reach it. Always at a line start. | |
| 1018 | + private(set) var styledUpTo = 0 | |
| 1019 | + private var stylingScheduled = false | |
| 1020 | + | |
| 1021 | + /// Styled before the first render: more than a screenful. | |
| 1022 | + static let firstStyleChunk = 30_000 | |
| 1023 | + /// Styled per run-loop turn afterwards. | |
| 1024 | + static let styleChunk = 200_000 | |
| 1025 | ||
| 1026 | public init(document: DocumentState, frame: NSRect = NSRect(x: 0, y: 0, width: 600, height: 800)) { | |
| 1027 | self.document = document | |
| 1028 | @@ -77,10 +86,46 @@ public final class OrgEditor: NSObject { | |
| 1029 | textLayoutManager.delegate = self | |
| 1030 | textContentStorage.delegate = self | |
| 1031 | ||
| 1032 | + load(document.text) | |
| 1033 | + } | |
| 1034 | + | |
| 1035 | + /// Replaces the storage's text. Setting the attributed string directly is about 100x | |
| 1036 | + /// faster than `NSTextView.string` for large files. | |
| 1037 | + private func load(_ text: String) { | |
| 1038 | isLoading = true | |
| 1039 | - textView.string = document.text | |
| 1040 | + textView.textStorage?.setAttributedString(NSAttributedString(string: text, attributes: theme.base)) | |
| 1041 | isLoading = false | |
| 1042 | - restyleOutsideEditing(0..<utf16Length) | |
| 1043 | + styledUpTo = 0 | |
| 1044 | + styleNextChunk(Self.firstStyleChunk) | |
| 1045 | + scheduleStyling() | |
| 1046 | + } | |
| 1047 | + | |
| 1048 | + // MARK: - Initial styling | |
| 1049 | + | |
| 1050 | + /// Styles whole lines from the watermark onward, about `size` units. | |
| 1051 | + private func styleNextChunk(_ size: Int) { | |
| 1052 | + let text = textView.textStorage!.string as NSString | |
| 1053 | + guard styledUpTo < text.length else { return } | |
| 1054 | + let target = min(text.length, styledUpTo + size) | |
| 1055 | + let end = target == text.length ? target : NSMaxRange(text.lineRange(for: NSRange(location: target, length: 0))) | |
| 1056 | + restyleOutsideEditing(styledUpTo..<end) | |
| 1057 | + styledUpTo = end | |
| 1058 | + } | |
| 1059 | + | |
| 1060 | + private func scheduleStyling() { | |
| 1061 | + guard !stylingScheduled, styledUpTo < utf16Length else { return } | |
| 1062 | + stylingScheduled = true | |
| 1063 | + DispatchQueue.main.async { [weak self] in | |
| 1064 | + guard let self else { return } | |
| 1065 | + self.stylingScheduled = false | |
| 1066 | + self.styleNextChunk(Self.styleChunk) | |
| 1067 | + self.scheduleStyling() | |
| 1068 | + } | |
| 1069 | + } | |
| 1070 | + | |
| 1071 | + /// Styles everything still waiting for a background pass. | |
| 1072 | + public func finishStyling() { | |
| 1073 | + while styledUpTo < utf16Length { styleNextChunk(Self.styleChunk) } | |
| 1074 | } | |
| 1075 | ||
| 1076 | public var textLayoutManager: NSTextLayoutManager { textView.textLayoutManager! } | |
| 1077 | @@ -107,10 +152,7 @@ public final class OrgEditor: NSObject { | |
| 1078 | let outcome = try saver.save(&document, to: url) | |
| 1079 | if document.text != before { | |
| 1080 | let edits = lineEdits(from: before, to: document.text) | |
| 1081 | - isLoading = true | |
| 1082 | - textView.string = document.text | |
| 1083 | - isLoading = false | |
| 1084 | - restyleOutsideEditing(0..<utf16Length) | |
| 1085 | + load(document.text) | |
| 1086 | setFolds(view.mapped(through: edits).folds) | |
| 1087 | } | |
| 1088 | return outcome | |
| 1089 | @@ -190,23 +232,30 @@ public final class OrgEditor: NSObject { | |
| 1090 | let lineRange = lines.location..<(lines.location + lines.length) | |
| 1091 | let around = max(0, lineRange.lowerBound - 1)..<min(text.length, lineRange.upperBound + 1) | |
| 1092 | var container = tree.root | |
| 1093 | - while let child = container.children.first(where: { | |
| 1094 | - ($0.kind == .section || $0.kind == .zerothSection) | |
| 1095 | - && $0.range.lowerBound <= lineRange.lowerBound && lineRange.upperBound <= $0.range.upperBound | |
| 1096 | - }) { | |
| 1097 | + while let child = container.child(containing: lineRange.lowerBound), | |
| 1098 | + child.kind == .section || child.kind == .zerothSection, | |
| 1099 | + lineRange.upperBound <= child.range.upperBound { | |
| 1100 | container = child | |
| 1101 | } | |
| 1102 | var lower = lineRange.lowerBound | |
| 1103 | var upper = lineRange.upperBound | |
| 1104 | - let children = container.children | |
| 1105 | - for child in children where child.range.overlaps(around) { | |
| 1106 | + // Tables and lists carry no styling of their own, so inside them only the rows or | |
| 1107 | + // items around the edit change, however large they are. | |
| 1108 | + func overlapping(_ nodes: [SyntaxNode]) -> [SyntaxNode] { | |
| 1109 | + nodes.filter { $0.range.overlaps(around) }.flatMap { node -> [SyntaxNode] in | |
| 1110 | + guard [.table, .plainList, .item].contains(node.kind) else { return [node] } | |
| 1111 | + let inner = overlapping(node.children(overlapping: around)) | |
| 1112 | + return inner.isEmpty ? [node] : inner | |
| 1113 | + } | |
| 1114 | + } | |
| 1115 | + for child in overlapping(container.children(overlapping: around)) { | |
| 1116 | lower = min(lower, child.range.lowerBound) | |
| 1117 | upper = max(upper, child.range.upperBound) | |
| 1118 | } | |
| 1119 | - let heading = children.first { $0.kind == .heading } | |
| 1120 | + let heading = container.firstChild(.heading) | |
| 1121 | if touchedHeadings || heading.map({ $0.range.overlaps(lineRange) }) == true { | |
| 1122 | lower = min(lower, container.range.lowerBound) | |
| 1123 | - upper = max(upper, children.first { $0.kind == .section }?.range.lowerBound ?? container.range.upperBound) | |
| 1124 | + upper = max(upper, container.firstChild(.section)?.range.lowerBound ?? container.range.upperBound) | |
| 1125 | } | |
| 1126 | return lower..<min(upper, text.length) | |
| 1127 | } | |
| 1128 | @@ -231,6 +280,13 @@ public final class OrgEditor: NSObject { | |
| 1129 | assertionFailure("text view and document disagree: \(error)") | |
| 1130 | return | |
| 1131 | } | |
| 1132 | + // Keep the watermark on unstyled text: shift it past the edit, or pull it back to the | |
| 1133 | + // edited line when the edit reaches into unstyled text. | |
| 1134 | + if styledUpTo >= edit.range.upperBound { | |
| 1135 | + styledUpTo += edit.replacement.utf16.count - edit.range.count | |
| 1136 | + } else if styledUpTo > edit.range.lowerBound { | |
| 1137 | + styledUpTo = (storage.string as NSString).lineRange(for: NSRange(location: edit.range.lowerBound, length: 0)).location | |
| 1138 | + } | |
| 1139 | let oldHidden = hidden.all | |
| 1140 | view = view.mapped(through: [edit]).pruned(to: document.tree) | |
| 1141 | folded.set(view.folds) | |
| 1142 | diff --git a/Sources/OrgPresentation/Presentation.swift b/Sources/OrgPresentation/Presentation.swift | |
| 1143 | index 312d7f8..504e0b5 100644 | |
| 1144 | --- a/Sources/OrgPresentation/Presentation.swift | |
| 1145 | +++ b/Sources/OrgPresentation/Presentation.swift | |
| 1146 | @@ -77,7 +77,7 @@ public enum Presentation { | |
| 1147 | default: | |
| 1148 | break | |
| 1149 | } | |
| 1150 | - for child in node.children where child.range.overlaps(range) { | |
| 1151 | + for child in node.children(overlapping: range) { | |
| 1152 | visit(child, range, settings, &runs) | |
| 1153 | } | |
| 1154 | } | |
| 1155 | @@ -118,12 +118,11 @@ public enum Presentation { | |
| 1156 | } | |
| 1157 | ||
| 1158 | private static func collectIndents(_ node: SyntaxNode, _ range: Range<Int>, _ runs: inout [IndentRun]) { | |
| 1159 | - for section in node.children where section.kind == .section && section.range.overlaps(range) { | |
| 1160 | - let children = section.children | |
| 1161 | - guard let heading = children.first(where: { $0.kind == .heading }) else { continue } | |
| 1162 | + for section in node.children(overlapping: range) where section.kind == .section { | |
| 1163 | + guard let heading = section.firstChild(.heading) else { continue } | |
| 1164 | let level = heading.tokens.first { $0.kind == .stars }?.text.count ?? 1 | |
| 1165 | runs.append(IndentRun(range: heading.range, firstLine: 0, wrapped: level + 1)) | |
| 1166 | - let bodyEnd = children.first { $0.kind == .section }?.range.lowerBound ?? section.range.upperBound | |
| 1167 | + let bodyEnd = section.firstChild(.section)?.range.lowerBound ?? section.range.upperBound | |
| 1168 | if heading.range.upperBound < bodyEnd { | |
| 1169 | runs.append(IndentRun(range: heading.range.upperBound..<bodyEnd, firstLine: level + 1, wrapped: level + 1)) | |
| 1170 | } | |
| 1171 | @@ -156,7 +155,7 @@ public enum Presentation { | |
| 1172 | /// Start offset of the heading whose line contains `offset`. | |
| 1173 | public static func heading(containing offset: Int, in tree: OrgTree) -> Int? { | |
| 1174 | var node = tree.root | |
| 1175 | - while let child = node.children.first(where: { $0.range.contains(offset) }) { | |
| 1176 | + while let child = node.child(containing: offset) { | |
| 1177 | if child.kind == .heading { return child.range.lowerBound } | |
| 1178 | guard child.kind == .section || child.kind == .zerothSection else { return nil } | |
| 1179 | node = child | |
| 1180 | ``` | |
| 1181 | ||
| 1182 | - [ ] **Step 3: Run the restyle fuzz on several seeds; commit** | |
| 1183 | ||
| 1184 | Run: `ORGSTAR_FUZZ_SEED=<n> ORGSTAR_RESTYLE_EDITS=1500 swift test -c release --filter RestyleFuzzTests` for at least three seeds. | |
| 1185 | ||
| 1186 | ```bash | |
| 1187 | git add Sources/OrgEditorAppKit Sources/OrgPresentation Sources/OrgDocument Tests/OrgEditorAppKitTests/EditorTests.swift | |
| 1188 | git commit -m "Style the first screen first and narrow restyles" | |
| 1189 | ``` | |
| 1190 | ||
| 1191 | --- | |
| 1192 | ||
| 1193 | ### Task 4: Gate tests | |
| 1194 | ||
| 1195 | **Files:** | |
| 1196 | - Create: `Tests/OrgEditorAppKitTests/GateTests.swift` | |
| 1197 | - Modify: `docs/design.md` (memory gate) | |
| 1198 | ||
| 1199 | - [ ] **Step 1: Add the gate suite** | |
| 1200 | ||
| 1201 | ```swift | |
| 1202 | import AppKit | |
| 1203 | import Foundation | |
| 1204 | import OrgDocument | |
| 1205 | import Testing | |
| 1206 | @testable import OrgEditorAppKit | |
| 1207 | ||
| 1208 | /// The phase 1 performance gates (design, "Phases"). Skipped unless `ORGSTAR_GATES` is set; | |
| 1209 | /// run in a release build on the reference machine: | |
| 1210 | /// | |
| 1211 | /// ORGSTAR_GATES=1 swift test -c release --filter GateTests | |
| 1212 | /// | |
| 1213 | /// `ORGSTAR_GATE_FILE` adds a real file to the synthetic ones. Memory limits follow the design: | |
| 1214 | /// 10x the file size for typical documents, 45x for the synthetic worst cases. | |
| 1215 | @MainActor | |
| 1216 | @Suite(.enabled(if: ProcessInfo.processInfo.environment["ORGSTAR_GATES"] != nil)) | |
| 1217 | struct GateTests { | |
| 1218 | /// About 1 MB of body text under a single top-level heading. | |
| 1219 | static let singleSection: String = { | |
| 1220 | var text = "* One big section\n" | |
| 1221 | var n = 0 | |
| 1222 | while text.utf8.count < 1_000_000 { | |
| 1223 | text += "Paragraph \(n) with *bold*, /italic/, a [[https://example.com/\(n)][link]] and <2026-10-04 Sun>.\n" | |
| 1224 | text += "It continues on a second line with =code= and ~more~.\n\n" | |
| 1225 | text += "- item \(n)\n - nested [ ] task\n\n" | |
| 1226 | n += 1 | |
| 1227 | } | |
| 1228 | return text | |
| 1229 | }() | |
| 1230 | ||
| 1231 | /// A 5,000-row table. | |
| 1232 | static let bigTable: String = { | |
| 1233 | var text = "* Table\n| id | name | note |\n|----+------+------|\n" | |
| 1234 | for n in 0..<5_000 { text += "| \(n) | row \(n) | *bold* and [[https://example.com][link]] |\n" } | |
| 1235 | return text | |
| 1236 | }() | |
| 1237 | ||
| 1238 | struct Result: CustomStringConvertible { | |
| 1239 | var name: String | |
| 1240 | var bytes: Int | |
| 1241 | var open: Duration | |
| 1242 | var typingP95: Duration | |
| 1243 | var newlineP95: Duration | |
| 1244 | var memoryRatio: Double | |
| 1245 | ||
| 1246 | var description: String { | |
| 1247 | "\(name) (\(bytes / 1000) kB): open \(open), typing p95 \(typingP95), blank line p95 \(newlineP95), memory \(String(format: "%.1f", memoryRatio))x" | |
| 1248 | } | |
| 1249 | } | |
| 1250 | ||
| 1251 | /// Heap bytes in use, which unlike resident size isn't hidden by memory freed earlier. | |
| 1252 | static func heapBytes() -> Int { | |
| 1253 | var stats = malloc_statistics_t() | |
| 1254 | malloc_zone_statistics(nil, &stats) | |
| 1255 | return Int(stats.size_in_use) | |
| 1256 | } | |
| 1257 | ||
| 1258 | func measure(_ name: String, _ text: String) -> Result { | |
| 1259 | let clock = ContinuousClock() | |
| 1260 | let bytes = Array(text.utf8) | |
| 1261 | let before = Self.heapBytes() | |
| 1262 | var harness: Harness! | |
| 1263 | let open = clock.measure { | |
| 1264 | harness = Harness(DocumentState(bytes: bytes), layout: false) | |
| 1265 | } | |
| 1266 | harness.drainStyling() | |
| 1267 | let memoryRatio = Double(Self.heapBytes() - before) / Double(bytes.count) | |
| 1268 | ||
| 1269 | var timings: [Duration] = [] | |
| 1270 | harness.editor.onEditTiming = { timings.append($0) } | |
| 1271 | let paragraphs = harness.editor.document.tree.root.descendants().filter { $0.kind == .paragraph || $0.kind == .tableCell } | |
| 1272 | harness.caret(at: paragraphs[paragraphs.count / 2].range.lowerBound + 1) | |
| 1273 | for _ in 0..<100 { harness.textView.insertText("x", replacementRange: harness.textView.selectedRange()) } | |
| 1274 | let typing = timings.sorted()[95] | |
| 1275 | ||
| 1276 | timings = [] | |
| 1277 | for _ in 0..<20 { | |
| 1278 | harness.textView.insertNewline(nil) | |
| 1279 | harness.textView.insertNewline(nil) | |
| 1280 | harness.textView.deleteBackward(nil) | |
| 1281 | harness.textView.deleteBackward(nil) | |
| 1282 | } | |
| 1283 | let sorted = timings.sorted() | |
| 1284 | return Result(name: name, bytes: bytes.count, open: open, typingP95: typing, newlineP95: sorted[sorted.count * 95 / 100], memoryRatio: memoryRatio) | |
| 1285 | } | |
| 1286 | ||
| 1287 | @Test func gates() throws { | |
| 1288 | // The first text view, and the first text loaded into one, set up the text system once | |
| 1289 | // per process; an app pays that at launch, not per file. | |
| 1290 | NSTextView(usingTextLayoutManager: true).textStorage?.setAttributedString(NSAttributedString(string: "warm up")) | |
| 1291 | // Memory limits: worst-case dense markup for the synthetic files, typical for real ones. | |
| 1292 | var inputs = [("single 1 MB section", Self.singleSection, 45.0), ("5,000-row table", Self.bigTable, 45.0)] | |
| 1293 | if let path = ProcessInfo.processInfo.environment["ORGSTAR_GATE_FILE"] { | |
| 1294 | inputs.append((URL(fileURLWithPath: path).lastPathComponent, try String(contentsOfFile: path, encoding: .utf8), 10.0)) | |
| 1295 | } | |
| 1296 | for (name, text, memoryLimit) in inputs { | |
| 1297 | let result = measure(name, text) | |
| 1298 | print("GATE \(result)") | |
| 1299 | let perMegabyte = Double(result.bytes) / 1_000_000 | |
| 1300 | #expect(result.open < .milliseconds(100) * max(1, perMegabyte), "\(name): open") | |
| 1301 | #expect(result.typingP95 < .milliseconds(16), "\(name): typing") | |
| 1302 | #expect(result.newlineP95 < .milliseconds(16), "\(name): blank line") | |
| 1303 | #expect(result.memoryRatio < memoryLimit, "\(name): memory") | |
| 1304 | } | |
| 1305 | } | |
| 1306 | } | |
| 1307 | ``` | |
| 1308 | ||
| 1309 | - [ ] **Step 2: Run the gates; commit** | |
| 1310 | ||
| 1311 | Run: `ORGSTAR_GATES=1 ORGSTAR_GATE_FILE=<large real file> swift test -c release --filter GateTests` | |
| 1312 | ||
| 1313 | ```bash | |
| 1314 | git add Tests/OrgEditorAppKitTests/GateTests.swift | |
| 1315 | git commit -m "Add performance gate tests" | |
| 1316 | ``` | |