audit-labs/audit-tools

A collection of scripts, queries, and other goodies you can use in an audit. audit automation compliance evidence scripts

Commit 982249b041

982249b041600983b0d233d6c93f74597ad49234

parent: a70cee9b1e

Unsigned

cmc <hello@cleberg.net> · 2026-08-07 02:13 UTC

Make sampling/sample.py reproducible and point to the real tool

Seed the example draw and note that audit_sample.py / sampling_tool records
population hash, seed, method, and version for defensible sampling.

Layout: unified · split

sampling/sample.py +12 −2
@@ -1,5 +1,11 @@
1""" 1"""
2Creates a sample from a CSV or Excel file based on user-defined SAMPLE_SIZE. 2Creates a sample from a CSV or Excel file based on user-defined SAMPLE_SIZE.
3
4NOTE: This is a minimal teaching snippet. For real fieldwork use the
5`audit_sample.py` CLI (or the `sampling_tool` package), which records the
6population hash, seed, method, and tool version in a manifest so the sample is
7reproducible and defensible. This file fixes a SEED only so the example itself is
8repeatable; it does not emit that provenance.
3""" 9"""
4 10
5# Import packages 11# Import packages
@@ -8,6 +14,10 @@ import pandas as pd
8# Define the sample size 14# Define the sample size
9SAMPLE_SIZE = 25 15SAMPLE_SIZE = 25
10 16
17# A fixed seed makes the draw reproducible: same population + same seed => same
18# rows. Record the seed alongside any sample you rely on.
19SEED = 20260707
20
11# Import the data to a pandas DataFrame 21# Import the data to a pandas DataFrame
12df = pd.read_csv("FILENAME_GOES_HERE.csv") 22df = pd.read_csv("FILENAME_GOES_HERE.csv")
13 23
@@ -19,7 +29,7 @@ df = pd.read_csv("FILENAME_GOES_HERE.csv")
19print("Dataframe size (rows, columns): ", df.shape) 29print("Dataframe size (rows, columns): ", df.shape)
20 30
21# Sample 31# Sample
22sample = df.sample(SAMPLE_SIZE) 32sample = df.sample(SAMPLE_SIZE, random_state=SEED)
23print("Sample size: ", SAMPLE_SIZE) 33print("Sample size: ", SAMPLE_SIZE)
24print("Sample:\n", sample) 34print("Sample:\n", sample)
25 35
@@ -31,4 +41,4 @@ print("Sample:\n", sample)
31# 41#
32# # Sample Size: 25 + 5 replacement samples 42# # Sample Size: 25 + 5 replacement samples
33# SAMPLE_SIZE = 30 43# SAMPLE_SIZE = 30
34# sample = df.sample(SAMPLE_SIZE, replace=True) 44# sample = df.sample(SAMPLE_SIZE, replace=True, random_state=SEED)