krz/omaha-metro-blotter

Archive of police activity and ALPR surveillance across the Omaha metro. alpr archive omaha police surveillance

Commit c5db763af6

c5db763af6af53a4c1f83217e44b3d0bc363bd8e

parent: ec96b33550

Verified · cmc

cmc <hello@cleberg.net> · 2026-08-23 21:18 UTC

Backfill OPD 2015-2021 from the yearly CSVs; fix the camera baseline

OPD publishes a CSV per year at a predictable path, which is where this
repo's original 2015-2023 data came from. The ArcGIS view only reaches
back to 2022-01-01, so opd_archive takes the years before that: 317,858
incidents, with the statute description rather than a NIBRS category.
Only pre-2022 years, because 31,865 of 2022's RB numbers are already in
the live feed and ingesting both would double-count them.

Keyed on a hash of the row, not its position: a single row inserted
upstream would otherwise shift every key below it. No raw payloads
kept — unlike the rolling feeds, these year files stay downloadable,
and the only field not carried into a column is Occurred District.

The new rows exposed a flaw in camera_proximity. Comparing stops with
every incident in the archive made stops look 2.6x more likely to sit
within 200 m of a camera, but the baseline was mostly Omaha, which
reports no stops, so it measured geography. Restricted to the five
agencies that do report stops it is 1.19x, with median distance 805 m
against 780 m — no meaningful separation. The site now says so.

These CSVs are reported crimes only: no stops, no dispositions, so
Omaha stays blank on the map.

Layout: unified · split

.github/workflows/daily-pull.yml +3 −1
@@ -99,7 +99,9 @@ jobs:
9999 if [ "${{ inputs.full }}" = "true" ] || [ "${{ inputs.bootstrap }}" = "true" ] \
100100 || [ "$(date -u +%u)" = "7" ]; then
101101 echo "full sweep"
102 python ingest.py --full
102 # opd_archive is closed years that never change; the weekly sweep is
103 # often enough to notice if OPD ever restates one.
104 python ingest.py --full opd sarpy cbpd alpr flock opd_archive
103105 else
104106 python ingest.py
105107 fi
README.nfo +8 −3
@@ -180,9 +180,14 @@ NOTES
180180 fact does not work: swapping a maplibre basemap at runtime leaves
181181 it rebuilding with no data layers.
182182
183 the camera-proximity panel compares stops against a non-stop
184 baseline. cameras and stops both concentrate on arterials, so a
185 gap between the curves is a starting point, not a finding.
183 the camera-proximity panel compares stops against other calls from
184 the same agencies. the baseline has to be restricted that way: run
185 against the whole archive it shows stops 2.6x more likely to be
186 within 200m of a camera, but most of the archive is omaha, which
187 reports no stops, so that number measures geography rather than
188 enforcement. like for like it is 1.19x, and median distance is
189 805m for stops against 780m for everything else -- no meaningful
190 separation.
186191
187192 raw_data/ingress.db is the old 2015-2023 sqlite build. nothing
188193 reads it any more.
analysis.py +13 −3
@@ -113,11 +113,21 @@ def stop_outcomes(df):
113113def camera_proximity(df, cameras, bin_m=200, max_m=2000):
114114 """Share of stops vs other incidents falling in each distance band.
115115
116 Both series are normalised, so a gap between them means stops cluster
117 differently around cameras than the rest of the call volume does. It is not
118 evidence of causation: cameras and stops both concentrate on arterials."""
116 The baseline is drawn only from agencies that report stops. Comparing stops
117 against every incident in the archive instead compares Sarpy and Council
118 Bluffs stops with a baseline that is mostly Omaha, a city reporting no stops
119 at all, and the geography alone then makes stops look far closer to cameras
120 than they are: 2.6x within 200 m across all agencies, 1.2x within the ones
121 actually being measured.
122
123 Even restricted this way it is not evidence of causation. Cameras and
124 enforcement both concentrate on arterials."""
119125 if df.empty or cameras.empty:
120126 return pd.DataFrame(columns=["distance_m", "kind", "share"])
127 reporting = df.loc[df["is_stop"] == 1, "agency"].unique()
128 df = df[df["agency"].isin(reporting)]
129 if df.empty:
130 return pd.DataFrame(columns=["distance_m", "kind", "share"])
121131 d = df.assign(dist=nearest_camera_m(df, cameras))
122132 d = d[d["dist"] <= max_m]
123133 if d.empty:
ingest.py +56 −27
@@ -14,8 +14,11 @@ Sources:
1414 raw_data/flock/<portal-slug>_<date>.csv and this reads what is there.
1515 The portals keep a rolling 30 days, so a gap longer than that is
1616 permanent.
17 opd_csv One-time backfill of raw_data/Incidents_*.csv (2015-2023). Statute text
18 only, no NIBRS category.
17 opd_archive
18 OPD's own yearly incident CSVs, 2015-2021. The ArcGIS view only goes
19 back to 2022-01-01, so this is the only route to the earlier years,
20 and it carries the statute description rather than a NIBRS category.
21 Reported crimes only: no stops, no dispositions.
1922
2023All three ArcGIS services return UTC epochs. Their WHERE literals do not agree:
2124OPD and Council Bluffs use UTC, Sarpy uses America/Chicago, so each source
@@ -97,6 +100,16 @@ SARPY_AGENCIES = {
97100# calls -- reactive, not officer-initiated.
98101SARPY_STOP_CATEGORY = "Proactive Policing - Vehicle Stop"
99102
103# OPD publishes a CSV per year at a predictable path, updated daily. The ArcGIS
104# view starts 2022-01-01, so only the years before that are taken from here --
105# ingesting the overlap would double-count every Omaha incident since 2022.
106OPD_ARCHIVE = "https://police-static.cityofomaha.org/crime-data/{y}/Incidents_{y}.csv"
107OPD_ARCHIVE_YEARS = range(2015, 2022)
108OPD_ARCHIVE_COLUMNS = ("RB Number", "Reported Date", "Reported Time",
109 "Statute/Ordinance Description", "Occurred Location",
110 "Occurred District", "Occurred Block LAT",
111 "Occurred Block LON")
112
100113OVERPASS = "https://overpass-api.de/api/interpreter"
101114# Douglas and Sarpy counties in Nebraska plus Council Bluffs across the river.
102115BBOX = (40.95, -96.35, 41.45, -95.65)
@@ -298,30 +311,46 @@ def ingest_cbpd(conn, since):
298311 return upsert(conn, rows)
299312
300313
301def ingest_opd_csv(conn, _since):
302 """Backfill the 2015-2023 CSV archive. Statute text goes to offense_desc;
303 category stays NULL because it is not a NIBRS category."""
304 rows = []
305 for path in sorted((ROOT / "raw_data").glob("Incidents_*.csv")):
306 with path.open(newline="") as fh:
307 if fh.readline().startswith("version https://git-lfs"):
308 print(f" {path.name}: git-lfs pointer, run 'git lfs pull'")
314def ingest_opd_archive(conn, _since):
315 """Load OPD's yearly incident CSVs for the years the ArcGIS view predates.
316
317 Keyed on a hash of the row rather than its position in the file: these are
318 closed years and should be stable, but a single row inserted upstream would
319 otherwise shift every key below it and file fifty thousand false
320 amendments."""
321 rows, seen = [], {}
322 for year in OPD_ARCHIVE_YEARS:
323 url = OPD_ARCHIVE.format(y=year)
324 req = urllib.request.Request(url, headers={"User-Agent": "omaha-incidents/1.0"})
325 with urllib.request.urlopen(req, timeout=180, context=SSL_CTX) as r:
326 text = r.read().decode("utf-8-sig", "replace")
327 reader = csv.DictReader(text.splitlines())
328 if tuple(reader.fieldnames or ()) != OPD_ARCHIVE_COLUMNS:
329 print(f" {year}: unexpected columns {reader.fieldnames}")
330 continue
331 n = 0
332 for rec in reader:
333 when = f"{rec['Reported Date']} {rec['Reported Time']}"
334 try:
335 occurred = datetime.strptime(when, "%m/%d/%Y %H:%M:%S")
336 except ValueError:
309337 continue
310 fh.seek(0)
311 for i, r in enumerate(csv.reader(fh)):
312 if len(r) < 8 or r[0] == "RB Number":
313 continue
314 rb, date, tm, desc, loc, district, lat, lon = r[:8]
315 try:
316 when = datetime.strptime(f"{date} {tm}", "%m/%d/%Y %H:%M")
317 except ValueError:
318 continue
319 rows.append((("opd_csv", f"{path.stem}:{i}", "Omaha PD", rb,
320 when.strftime("%Y-%m-%dT%H:%M:%S"), None, None,
321 None, desc, 0, loc,
322 float(lat) if lat else None,
323 float(lon) if lon else None),
324 json.dumps(r)))
338 raw = json.dumps(rec, sort_keys=True)
339 # a few rows repeat verbatim; number them so each keeps its own key
340 h = hashlib.blake2b(raw.encode(), digest_size=8).hexdigest()
341 # No raw kept: unlike the rolling feeds, whose aged-out records can
342 # never be fetched again, these year files stay downloadable. The
343 # one field not carried into a column is Occurred District.
344 seen[h] = seen.get(h, 0) + 1
345 lat, lon = rec["Occurred Block LAT"], rec["Occurred Block LON"]
346 rows.append(((
347 "opd_archive", f"{h}:{seen[h]}", "Omaha PD", rec["RB Number"],
348 occurred.strftime("%Y-%m-%dT%H:%M:%S"), None, None, None,
349 rec["Statute/Ordinance Description"], 0,
350 rec["Occurred Location"],
351 float(lat) if lat else None, float(lon) if lon else None), None))
352 n += 1
353 print(f" {year}: {n}", end="\r", file=sys.stderr, flush=True)
325354 return upsert(conn, rows)
326355
327356
@@ -443,7 +472,7 @@ SOURCES = {
443472 "cbpd": ingest_cbpd,
444473 "alpr": ingest_alpr,
445474 "flock": ingest_flock,
446 "opd_csv": ingest_opd_csv,
475 "opd_archive": ingest_opd_archive,
447476}
448477
449478
@@ -454,7 +483,7 @@ def main():
454483 help="default: opd sarpy cbpd alpr flock")
455484 p.add_argument("--since-days", type=int, default=30,
456485 help="only pull incidents this recent (default 30); "
457 "ignored by alpr, flock and opd_csv")
486 "ignored by alpr, flock and opd_archive")
458487 p.add_argument("--full", action="store_true",
459488 help="pull the complete feed instead of --since-days")
460489 p.add_argument("--import-flock", metavar="CSV",
site_template.html +4 −2
@@ -112,12 +112,14 @@ footer a { color: var(--text-secondary); }
112112
113113<section>
114114 <h2>Distance to the nearest ALPR camera</h2>
115 <p>Where vehicle stops happen relative to automated licence plate readers, against every other kind of incident as a baseline. Both lines are shares of their own series, so the gap is what matters.</p>
115 <p>Where vehicle stops happen relative to automated licence plate readers, against every other call handled by <em>the same agencies</em> as a baseline. Both lines are shares of their own series, so the gap between them is what matters.</p>
116116 <div class="card">
117117 <div class="legend" id="lg-prox"></div>
118118 <div id="c-prox"></div>
119119 </div>
120 <p class="note"><strong>This is not evidence that cameras cause stops.</strong> Cameras get mounted on arterial roads and arterial roads are where traffic enforcement happens, so the two concentrate together for reasons that have nothing to do with each other. It is a starting point for asking the agencies a question, not an answer. That question has a stated standard: all three metro Flock transparency portals list <em>traffic enforcement</em> under prohibited uses.</p>
120 <p class="note"><strong>The two lines nearly overlap, and that is the finding.</strong> Stops sit marginally closer to cameras than the same agencies' other calls &mdash; 14.7% against 12.4% within 200&nbsp;m &mdash; but the median distance is 805&nbsp;m for stops and 780&nbsp;m for everything else. There is no meaningful separation here.</p>
121 <p class="note">Comparing stops against <em>every</em> incident in the archive instead produces a much larger gap, 2.6&times; within 200&nbsp;m. That gap is an artifact: most of the archive is Omaha, a city that reports no stops at all, so the comparison was measuring the difference between two geographies rather than between stops and other calls. Restricting the baseline to the five agencies that actually report stops removes it.</p>
122 <p class="note">All three metro Flock transparency portals list <em>traffic enforcement</em> under prohibited uses. On this measure, at this resolution, nothing here contradicts that. It is a weak instrument &mdash; cameras and enforcement both concentrate on arterial roads, and camera locations are volunteer-mapped and incomplete &mdash; so it cannot clear an agency either.</p>
121123 <details><summary>Show the numbers</summary><div id="t-prox"></div></details>
122124</section>
123125