krz/omaha-metro-blotter

Archive of police activity and ALPR surveillance across the Omaha metro. alpr archive omaha police surveillance

Commit 137b3438e9

137b3438e9abacc7fc404054ced30663a5747dec

parent: 63f759d70a

Verified · cmc

cmc <hello@cleberg.net> · 2026-08-23 18:22 UTC

Keep raw feed payloads; sweep in full weekly, pull twice daily

raw_records stores the JSON each feed served for every version of
every record, keyed like incident_amendments. A wrong parse or a
field added later can only be applied to history if the bytes were
kept, and the rolling feeds give no second chance. Costs ~0.7 MB
gzipped a day; the published archive goes from 20 MB to 52 MB.

A 30-day window cannot see an agency amending a record filed months
ago, which Omaha does, so Sundays sweep every feed in full.

Twice-daily pulls halve what the rolling feeds drop before capture.

Layout: unified · split

.github/workflows/daily-pull.yml +22 −5
@@ -5,9 +5,11 @@ name: daily pull
55
66on:
77 schedule:
8 # 06:00 America/Chicago in summer, 05:00 in winter. Offset from the hour
9 # because GitHub drops on-the-hour scheduled runs under load.
8 # Twice a day, because the cadence sets how much the rolling feeds drop
9 # before it is captured: a five-hour gap cost 13 Sarpy records once.
10 # Offset from the hour, GitHub drops on-the-hour runs under load.
1011 - cron: "17 11 * * *"
12 - cron: "17 23 * * *"
1113 workflow_dispatch:
1214 inputs:
1315 bootstrap:
@@ -34,7 +36,7 @@ env:
3436jobs:
3537 pull:
3638 runs-on: ubuntu-latest
37 timeout-minutes: 30
39 timeout-minutes: 45
3840
3941 steps:
4042 - uses: actions/checkout@v4
@@ -73,12 +75,22 @@ jobs:
7375 sqlite3 -noheader -separator ' ' "$DB" \
7476 "SELECT source || '+amend', COUNT(*) FROM incident_amendments
7577 GROUP BY source ORDER BY source" >> before.txt 2>/dev/null || true
78 sqlite3 -noheader -separator ' ' "$DB" \
79 "SELECT source || '+raw', COUNT(*) FROM raw_records
80 GROUP BY source ORDER BY source" >> before.txt 2>/dev/null || true
81 sqlite3 -noheader -separator ' ' "$DB" \
82 "SELECT source || '+raw', COUNT(*) FROM raw_records
83 GROUP BY source ORDER BY source" >> before.txt 2>/dev/null || true
7684 fi
7785 cat before.txt
7886
7987 - name: Pull feeds
8088 run: |
81 if [ "${{ inputs.full }}" = "true" ] || [ "${{ inputs.bootstrap }}" = "true" ]; then
89 # A 30-day window cannot see an agency amending a record filed months
90 # ago, and OPD does exactly that, so sweep the whole feed on Sundays.
91 if [ "${{ inputs.full }}" = "true" ] || [ "${{ inputs.bootstrap }}" = "true" ] \
92 || [ "$(date -u +%u)" = "7" ]; then
93 echo "full sweep"
8294 python ingest.py --full
8395 else
8496 python ingest.py
@@ -92,6 +104,9 @@ jobs:
92104 sqlite3 -noheader -separator ' ' "$DB" \
93105 "SELECT source || '+amend', COUNT(*) FROM incident_amendments
94106 GROUP BY source ORDER BY source" >> after.txt
107 sqlite3 -noheader -separator ' ' "$DB" \
108 "SELECT source || '+raw', COUNT(*) FROM raw_records
109 GROUP BY source ORDER BY source" >> after.txt
95110 cat after.txt
96111 test -s after.txt || { echo "::error::archive is empty"; exit 1; }
97112 # Keyed on FILENAME, not NR == FNR: before.txt is empty on a bootstrap
@@ -141,7 +156,9 @@ jobs:
141156 echo "incidents holds each record as first published; every later"
142157 echo "version the feed served is a row in incident_amendments"
143158 echo "($(sqlite3 -noheader "$DB" 'SELECT COUNT(*) FROM incident_amendments') so far)."
144 echo "incidents_current is the newest version of each."
159 echo "incidents_current is the newest version of each, and"
160 echo "raw_records keeps the feed's own JSON for every version"
161 echo "so a parse can be redone against what actually arrived."
145162 echo
146163 echo '```'
147164 sqlite3 -header -column "$DB" \
README.nfo +18 −1
@@ -40,13 +40,17 @@ USE
4040 .venv/bin/python app.py
4141
4242ARCHIVE
43 .github/workflows/daily-pull.yml runs the pull at 11:17 utc and
43 .github/workflows/daily-pull.yml runs at 11:17 and 23:17 utc and
4444 keeps the database as metro.db.gz on the "archive" release, so the
4545 archive does not depend on any one machine. each run restores that
4646 asset, pulls, refuses to publish if any source came back with fewer
4747 rows than it started with, then uploads and fails loudly if a feed
4848 has not moved in seven days.
4949
50 sundays it sweeps every feed in full instead of the last 30 days,
51 because a 30-day window cannot see an agency amending a record it
52 filed months ago, and omaha does that.
53
5054 first run: trigger it manually with bootstrap enabled, which pulls
5155 every feed in full and creates the release. after that the restore
5256 step is mandatory -- a bootstrap over a live archive throws away
@@ -61,6 +65,19 @@ ARCHIVE
6165
6266 0 6 * * * cd /path/to/omaha-incidents && .venv/bin/python ingest.py
6367
68RAW
69 raw_records keeps the feed's own json for every version of every
70 record, keyed the same way amendments are. a parse that turns out
71 wrong, or a field a feed adds later, can only be applied to history
72 if the bytes were kept, and the rolling feeds mean there is no
73 second chance to fetch them. the payloads already carry fields
74 ingest does not map: council bluffs response times and priority,
75 sarpy case status.
76
77 it costs about 0.7 mb gzipped a day and roughly triples the
78 database: 20 mb published without it, 52 mb with. 319 records
79 predate it and their raw is gone; the feeds no longer serve them.
80
6481AMENDMENTS
6582 agencies edit records after publishing them: a disposition changes,
6683 a case reopens, a record is withdrawn. nothing in the incidents
ingest.py +39 −24
@@ -194,15 +194,18 @@ def digest(values):
194194def upsert(conn, rows):
195195 """Insert records not seen before; file a changed record as an amendment.
196196
197 Nothing in incidents is ever updated. A record whose payload differs from
198 the one on file is appended to incident_amendments, so the version the
199 agency published first stays readable next to what it published later."""
197 Takes (values, raw) pairs, where raw is the feature exactly as the feed
198 served it. Nothing in incidents is ever updated: a record whose payload
199 differs from the one on file is appended to incident_amendments, so the
200 version the agency published first stays readable next to what it published
201 later. Every version's raw payload is kept too, so a parse can be redone
202 against what actually arrived."""
200203 now = datetime.now(LOCAL).strftime("%Y-%m-%dT%H:%M:%S")
201 marks = ",".join("?" * (len(COLUMNS) + 1))
202 staged = [r + (digest(r[2:]),) for r in rows]
204 marks = ",".join("?" * (len(COLUMNS) + 2))
205 staged = [r + (digest(r[2:]), raw) for r, raw in rows]
203206
204207 conn.execute("DROP TABLE IF EXISTS temp.incoming")
205 conn.execute(f"CREATE TEMP TABLE incoming ({','.join(COLUMNS)}, digest)")
208 conn.execute(f"CREATE TEMP TABLE incoming ({','.join(COLUMNS)}, digest, raw)")
206209 conn.executemany(f"INSERT INTO temp.incoming VALUES ({marks})", staged)
207210 conn.execute("CREATE INDEX temp.incoming_key ON incoming (source, source_key)")
208211
@@ -218,6 +221,14 @@ def upsert(conn, rows):
218221 ON o.source = i.source AND o.source_key = i.source_key
219222 WHERE o.digest <> i.digest""", (now,)).rowcount
220223
224 # OR IGNORE keyed on the version, so a run that re-serves a known record
225 # stores nothing and the first full run backfills whatever is still served.
226 conn.execute(
227 """INSERT OR IGNORE INTO raw_records (source, source_key, digest,
228 fetched_at, payload)
229 SELECT source, source_key, digest, ?, raw FROM temp.incoming
230 WHERE raw IS NOT NULL""", (now,))
231
221232 conn.execute("DROP TABLE temp.incoming")
222233 return len(rows), amended
223234
@@ -229,9 +240,10 @@ def ingest_opd(conn, since):
229240 occurred = local_iso(a["dteMidpoint"])
230241 if occurred is None:
231242 continue
232 rows.append(("opd", str(a["PK"]), "Omaha PD", a.get("RB"), occurred,
233 a.get("NIBRSCategory"), None, None, None, 0,
234 a.get("AddressBlock"), a.get("LatBlock"), a.get("LonBlock")))
243 rows.append((("opd", str(a["PK"]), "Omaha PD", a.get("RB"), occurred,
244 a.get("NIBRSCategory"), None, None, None, 0,
245 a.get("AddressBlock"), a.get("LatBlock"), a.get("LonBlock")),
246 json.dumps(f, sort_keys=True)))
235247 return upsert(conn, rows)
236248
237249
@@ -249,11 +261,12 @@ def ingest_sarpy(conn, since):
249261 unmapped.add(prefix)
250262 agency = f"Unmapped {prefix}"
251263 g = f.get("geometry") or {}
252 rows.append(("sarpy", iid, agency, iid, occurred, a.get("Category"),
253 a.get("CadTypeDesc"), a.get("CadDisposition"),
254 a.get("StatuteDesc"),
255 int(a.get("Category") == SARPY_STOP_CATEGORY),
256 a.get("BlkAddress"), g.get("y"), g.get("x")))
264 rows.append((("sarpy", iid, agency, iid, occurred, a.get("Category"),
265 a.get("CadTypeDesc"), a.get("CadDisposition"),
266 a.get("StatuteDesc"),
267 int(a.get("Category") == SARPY_STOP_CATEGORY),
268 a.get("BlkAddress"), g.get("y"), g.get("x")),
269 json.dumps(f, sort_keys=True)))
257270 result = upsert(conn, rows)
258271 if unmapped:
259272 print(f" unmapped IncidentId prefixes: {sorted(unmapped)}")
@@ -270,11 +283,12 @@ def ingest_cbpd(conn, since):
270283 g = f.get("geometry") or {}
271284 code = a.get("incident_code")
272285 # Council Bluffs withholds the street address; the point is still exact.
273 rows.append(("cbpd", a["cfs_number"], "Council Bluffs PD",
274 a.get("case_number") or a.get("cfs_number"), occurred,
275 a.get("incident_category"), code, a.get("disp_code"), None,
276 int(code == CBPD_STOP_CODE),
277 None, g.get("y"), g.get("x")))
286 rows.append((("cbpd", a["cfs_number"], "Council Bluffs PD",
287 a.get("case_number") or a.get("cfs_number"), occurred,
288 a.get("incident_category"), code, a.get("disp_code"), None,
289 int(code == CBPD_STOP_CODE),
290 None, g.get("y"), g.get("x")),
291 json.dumps(f, sort_keys=True)))
278292 return upsert(conn, rows)
279293
280294
@@ -296,11 +310,12 @@ def ingest_opd_csv(conn, _since):
296310 when = datetime.strptime(f"{date} {tm}", "%m/%d/%Y %H:%M")
297311 except ValueError:
298312 continue
299 rows.append(("opd_csv", f"{path.stem}:{i}", "Omaha PD", rb,
300 when.strftime("%Y-%m-%dT%H:%M:%S"), None, None, None,
301 desc, 0, loc,
302 float(lat) if lat else None,
303 float(lon) if lon else None))
313 rows.append((("opd_csv", f"{path.stem}:{i}", "Omaha PD", rb,
314 when.strftime("%Y-%m-%dT%H:%M:%S"), None, None,
315 None, desc, 0, loc,
316 float(lat) if lat else None,
317 float(lon) if lon else None),
318 json.dumps(r)))
304319 return upsert(conn, rows)
305320
306321
schema.sql +13
@@ -76,6 +76,19 @@ FROM (
7676)
7777WHERE rn = 1;
7878
79-- What the feed actually served, one row per version, keyed the same way
80-- incident_amendments is. A parse that turns out wrong or a field a feed adds
81-- later can only be applied to history if the bytes were kept, and the feeds
82-- age out, so there is no second chance to fetch them.
83CREATE TABLE IF NOT EXISTS raw_records (
84 source TEXT NOT NULL,
85 source_key TEXT NOT NULL,
86 digest TEXT NOT NULL, -- the version of the record this payload produced
87 fetched_at TEXT NOT NULL,
88 payload TEXT NOT NULL, -- the feature object as served, JSON
89 PRIMARY KEY (source, source_key, digest)
90);
91
7992-- ALPR cameras from OpenStreetMap (ODbL). first_seen/last_seen track when a node
8093-- entered and was last present in the Overpass result, so cameras that appear or
8194-- are removed are visible over time.