krz/omaha-metro-blotter

Archive of police activity and ALPR surveillance across the Omaha metro. alpr archive omaha police surveillance

Commit 752a483562

752a4835626d3138238de6d2f896be840a2896ac

parent: 055f5ddc3b

Verified · cmc

cmc <hello@cleberg.net> · 2024-01-24 19:34 UTC

clean up notebook

Layout: unified · split

notebooks/db_exploration.ipynb +119 −19
@@ -16,14 +16,24 @@
1616 "Let\"s explore the data a little bit to see what kind of analysis and visualizations we want to implement."
1717 ]
1818 },
19 {
20 "cell_type": "markdown",
21 "metadata": {},
22 "source": [
23 "### Set up environment\n",
24 "\n",
25 "Start by installating and importing the necessary packages. "
26 ]
27 },
1928 {
2029 "cell_type": "code",
2130 "execution_count": null,
2231 "metadata": {},
2332 "outputs": [],
2433 "source": [
25 "!pip3 install ipykernel\n",
26 "!pip3 install --upgrade pandas plotly dash \"nbformat>=4.2.0\""
34 "# Install packages, if needed\n",
35 "# !pip3 install ipykernel\n",
36 "# !pip3 install --upgrade pandas plotly dash \"nbformat>=4.2.0\""
2737 ]
2838 },
2939 {
@@ -32,8 +42,22 @@
3242 "metadata": {},
3343 "outputs": [],
3444 "source": [
45 "# Import packages\n",
3546 "import pandas as pd\n",
36 "import sqlite3"
47 "import numpy as np\n",
48 "import sqlite3\n",
49 "import plotly.express as px\n",
50 "import plotly.graph_objects as go\n",
51 "import plotly.io as pio"
52 ]
53 },
54 {
55 "cell_type": "markdown",
56 "metadata": {},
57 "source": [
58 "### Load Data\n",
59 "\n",
60 "To load the data, we need to connect to the SQLite3 database file and query it for the data we want."
3761 ]
3862 },
3963 {
@@ -42,7 +66,9 @@
4266 "metadata": {},
4367 "outputs": [],
4468 "source": [
45 "connection = sqlite3.connect(\"../raw_data/ingress.db\")"
69 "# Connect to the database\n",
70 "connection = sqlite3.connect(\"../raw_data/ingress.db\")\n",
71 "cursor = connection.cursor()"
4672 ]
4773 },
4874 {
@@ -51,7 +77,9 @@
5177 "metadata": {},
5278 "outputs": [],
5379 "source": [
54 "cursor = connection.cursor()"
80 "# If exists, delete extra header rows\n",
81 "# delete_headers = \"DELETE FROM incidents WHERE rb = 'RB Number'\"\n",
82 "# cursor.execute(delete_headers)"
5583 ]
5684 },
5785 {
@@ -60,8 +88,19 @@
6088 "metadata": {},
6189 "outputs": [],
6290 "source": [
63 "# Test query to see if the data loaded\n",
64 "select_all = \"SELECT * FROM incidents;\""
91 "# Grab all data\n",
92 "select_all = \"SELECT * FROM incidents\"\n",
93 "df = pd.read_sql_query(select_all, connection)\n",
94 "df.head()"
95 ]
96 },
97 {
98 "cell_type": "markdown",
99 "metadata": {},
100 "source": [
101 "### Data Cleaning\n",
102 "\n",
103 "We will clean up the data before we use: inserting NaN, converting types, etc."
65104 ]
66105 },
67106 {
@@ -70,10 +109,31 @@
70109 "metadata": {},
71110 "outputs": [],
72111 "source": [
73 "df = pd.read_sql_query(select_all, connection)\n",
112 "# Replace empty cells in [lat, lon] with NaN\n",
113 "df = df.replace(r'^\\s*$', np.nan, regex=True)\n",
74114 "df.head()"
75115 ]
76116 },
117 {
118 "cell_type": "code",
119 "execution_count": null,
120 "metadata": {},
121 "outputs": [],
122 "source": [
123 "# Convert date col to datetime format\n",
124 "df[\"date\"] = pd.to_datetime(df[\"date\"])\n",
125 "df"
126 ]
127 },
128 {
129 "cell_type": "markdown",
130 "metadata": {},
131 "source": [
132 "### Plotting\n",
133 "\n",
134 "Let's test a plot that will show us the top categories of incidents."
135 ]
136 },
77137 {
78138 "cell_type": "code",
79139 "execution_count": null,
@@ -81,20 +141,43 @@
81141 "outputs": [],
82142 "source": [
83143 "# test plotting by sorting & plotting top 5 crime categories\n",
84 "s = df.value_counts(subset=[\"description\"])\n",
144 "s = dff.value_counts(subset=[\"description\"])\n",
85145 "t = s.nlargest(5)\n",
86146 "t.head()\n",
87147 "t.plot(kind=\"bar\", title=\"Top 5 Incident Categories\")"
88148 ]
89149 },
150 {
151 "cell_type": "markdown",
152 "metadata": {},
153 "source": [
154 "### Data Filtering\n",
155 "\n",
156 "To reduce the workload in this rest of this notebook, I am filtering just for one description and a range of dates.\n",
157 "\n",
158 "If you are doing a lot of analysis, I recommend modifying the query at the beginning to only the pull the data you need instead of filtering after querying."
159 ]
160 },
90161 {
91162 "cell_type": "code",
92163 "execution_count": null,
93164 "metadata": {},
94165 "outputs": [],
95166 "source": [
96 "import plotly.express as px\n",
97 "import plotly.graph_objects as go"
167 "# Create a smaller dataframe based on a selected date and description\n",
168 "start_date = \"2023-01-01\"\n",
169 "end_date = \"2023-12-31\"\n",
170 "description = \"INJURY\"\n",
171 "\n",
172 "dff = df[(df['date'] > start_date) & (df['date'] < end_date)]\n",
173 "dff = dff.reset_index()\n",
174 "dff = dff[dff.description == description]\n",
175 "\n",
176 "dff_grouped = dff.groupby(by=\"date\").count()\n",
177 "dff_grouped = dff_grouped.reset_index()\n",
178 "\n",
179 "print(dff.head())\n",
180 "print(dff_grouped.head())"
98181 ]
99182 },
100183 {
@@ -103,7 +186,7 @@
103186 "metadata": {},
104187 "outputs": [],
105188 "source": [
106 "s.head(10)"
189 "dff.size"
107190 ]
108191 },
109192 {
@@ -112,8 +195,16 @@
112195 "metadata": {},
113196 "outputs": [],
114197 "source": [
115 "filtered_df = df[(df['date'] > '01/01/2023') & (df['date'] < '12/31/2023')]\n",
116 "filtered_df.head()"
198 "dff.info()"
199 ]
200 },
201 {
202 "cell_type": "markdown",
203 "metadata": {},
204 "source": [
205 "### Mapping\n",
206 "\n",
207 "Let's create a geo map of the crime data."
117208 ]
118209 },
119210 {
@@ -123,7 +214,7 @@
123214 "outputs": [],
124215 "source": [
125216 "fig = px.scatter_mapbox(\n",
126 " filtered_df,\n",
217 " dff,\n",
127218 " lat=\"lat\",\n",
128219 " lon=\"lon\",\n",
129220 " color=\"description\",\n",
@@ -138,7 +229,7 @@
138229 "fig.update_layout(mapbox_style=\"open-street-map\")\n",
139230 "fig.update_layout(margin={\"r\": 0, \"t\": 0, \"l\": 0, \"b\": 0})\n",
140231 "fig.update_layout(mapbox_bounds={\"west\": -180, \"east\": -50, \"south\": 20, \"north\": 90})\n",
141 "# fig.show()"
232 "fig.show()"
142233 ]
143234 },
144235 {
@@ -147,8 +238,17 @@
147238 "metadata": {},
148239 "outputs": [],
149240 "source": [
150 "import plotly.io as pio\n",
151 "pio.write_html(fig, file=\"test.html\", auto_open=False)"
241 "# Optionally, save the figure to an HTML file\n",
242 "# pio.write_html(fig, file=\"test.html\", auto_open=True)"
243 ]
244 },
245 {
246 "cell_type": "markdown",
247 "metadata": {},
248 "source": [
249 "## Wrapping Up\n",
250 "\n",
251 "To finish, remember to close your database connections and save any data you need."
152252 ]
153253 },
154254 {
@@ -157,7 +257,7 @@
157257 "metadata": {},
158258 "outputs": [],
159259 "source": [
160 "# clean up and close it out\n",
260 "# clean up and close out the database\n",
161261 "connection.commit()\n",
162262 "connection.close()"
163263 ]