Sampling design
Does each design deliver the mix it promises?
The generator can draw addresses three ways: uniformly over suburbs (what the 2025 code did), weighted by remoteness, or stratified with fixed quotas. This page puts each one through the real generator with fixed seeds and asks two questions: does the sample hit its target mix, with honest uncertainty, and do the coordinates land where they should?
Every number is computed at build time by the same code that runs on /generate, from the seeds shown, so a rebuild gives the same numbers. How the checks were designed is on Methods.
200 seeds each
Three designs, one target
The target is the remoteness mix in config.py that the 2025 README promised (40% Major Cities, 25% Inner Regional, 20% Outer Regional, 10% Remote, 5% Very Remote). Each design generated 1,000 addresses with seeds 1 to 200. A good weighted design should miss the target only by chance: its 95% Wilson intervals should cover the target about 95% of the time, and a goodness-of-fit test at the 5% level should reject it about 5% of the time.
Uniform (2025)
Every suburb equally likely, as the 2025 code actually behaved.
- Seeds where the test rejects the target at 5%
- 200 of 200, 100.0% (95% CI 98.1% to 100.0%)
- Distance from the target (Cohen's w)
- mean 0.445 (medium), middle 95% 0.379 to 0.514
Weighted
Each address independently picks a remoteness area with the config.py weights, then a suburb.
- Seeds where the test rejects the target at 5%
- 6 of 200, 3.0% (95% CI 1.4% to 6.4%)
- Distance from the target (Cohen's w)
- mean 0.060 (negligible), middle 95% 0.025 to 0.100
Stratified
Fixed quotas per remoteness area (largest remainder), then suburbs uniformly inside each.
- Goodness-of-fit test
- Not tested: the quotas fix every share (200 of 200 seeds exactly on target)
- Distance from the target (Cohen's w)
- 0.000 in every seed (no sampling variation)
Hover or focus an area for the numbers behind each line.
Uniform misses badly and consistently: Major Cities gets about 25% instead of 40%, because only 25% of suburbs are in a major city. The test catches it in every seed, and w around 0.45 (a medium effect, close to large) says the gap is substantial, not merely detectable.
Weighted behaves as theory says. The spread across seeds matches the multinomial standard deviation (for Major Cities 1.52% observed against 1.55% expected), the Wilson intervals cover the target in 93.5% to 97.0% of seeds per area, and the test rejects in 3.0% (95% CI 1.4% to 6.4%), consistent with its 5% level.
Stratified has nothing to test: the quotas fix every share, so all 200 of 200 seeds give exactly the target, and it is reported as fixed rather than given a test result or coverage interval. The suburbs, street numbers and names inside each area are still random. Use it when every area must be represented in a small sample; use weighted when you want the realistic variation of independent draws.
Show every number as a table
| Design | Area | Target | Design share | Mean share | Middle 95% of seeds | SD (theory) | Wilson CI covers target |
|---|---|---|---|---|---|---|---|
| Uniform (2025) | Major Cities | 40% | 25.0% | 25.20% | 22.4% to 27.6% | 1.29% (1.37%) | 0/200 (0.0% to 1.9%) |
| Uniform (2025) | Inner Regional | 25% | 20.2% | 20.18% | 18.1% to 22.7% | 1.15% (1.27%) | 6/200 (1.4% to 6.4%) |
| Uniform (2025) | Outer Regional | 20% | 34.3% | 34.18% | 31.7% to 36.5% | 1.38% (1.50%) | 0/200 (0.0% to 1.9%) |
| Uniform (2025) | Remote | 10% | 11.8% | 11.75% | 9.8% to 13.6% | 0.98% (1.02%) | 116/200 (51.1% to 64.6%) |
| Uniform (2025) | Very Remote | 5% | 8.6% | 8.70% | 7.0% to 10.3% | 0.89% (0.89%) | 1/200 (0.1% to 2.8%) |
| Weighted | Major Cities | 40% | 40.0% | 40.07% | 36.9% to 42.6% | 1.52% (1.55%) | 190/200 (91.0% to 97.3%) |
| Weighted | Inner Regional | 25% | 25.0% | 24.99% | 22.0% to 27.6% | 1.39% (1.37%) | 188/200 (89.8% to 96.5%) |
| Weighted | Outer Regional | 20% | 20.0% | 19.89% | 17.7% to 22.2% | 1.20% (1.26%) | 194/200 (93.6% to 98.6%) |
| Weighted | Remote | 10% | 10.0% | 10.03% | 8.2% to 11.8% | 0.95% (0.95%) | 191/200 (91.7% to 97.6%) |
| Weighted | Very Remote | 5% | 5.0% | 5.01% | 3.6% to 6.4% | 0.70% (0.69%) | 187/200 (89.2% to 96.2%) |
| Stratified | Major Cities | 40% | 40.0% | 40.00% | 40.0% to 40.0% | 0.00% (0.00%) | n/a (fixed) |
| Stratified | Inner Regional | 25% | 25.0% | 25.00% | 25.0% to 25.0% | 0.00% (0.00%) | n/a (fixed) |
| Stratified | Outer Regional | 20% | 20.0% | 20.00% | 20.0% to 20.0% | 0.00% (0.00%) | n/a (fixed) |
| Stratified | Remote | 10% | 10.0% | 10.00% | 10.0% to 10.0% | 0.00% (0.00%) | n/a (fixed) |
| Stratified | Very Remote | 5% | 5.0% | 5.00% | 5.0% to 5.0% | 0.00% (0.00%) | n/a (fixed) |
“Design share” is what each design draws in expectation; the theoretical SD is √(p(1 − p)/n) at that share (zero for fixed quotas). Coverage intervals are 95% Wilson intervals over the 200 seeds. Coverage is not reported for the stratified design: its shares are fixed, so an interval around them would describe no real uncertainty.
Seed 2025, n = 2,000
One sample, read properly
The README's key result, recomputed: the same seed and size through the uniform and the weighted design. Each share comes with its 95% Wilson interval; the residual (O − E)/√E shows which areas drive the test statistic.
| Remoteness area | Count | Share | 95% Wilson CI | Target | Residual |
|---|---|---|---|---|---|
| Major Cities | 473 | 23.6% | 21.8% to 25.6% | 40% | -11.6 |
| Inner Regional | 391 | 19.6% | 17.9% to 21.3% | 25% | -4.9 |
| Outer Regional | 722 | 36.1% | 34.0% to 38.2% | 20% | +16.1 |
| Remote | 226 | 11.3% | 10.0% to 12.8% | 10% | +1.8 |
| Very Remote | 188 | 9.4% | 8.2% to 10.8% | 5% | +8.8 |
Pearson chi-square test: χ²(4) = 497.45, p < 0.001; Cohen's w = 0.499 (medium).
| Remoteness area | Count | Share | 95% Wilson CI | Target | Residual |
|---|---|---|---|---|---|
| Major Cities | 804 | 40.2% | 38.1% to 42.4% | 40% | +0.1 |
| Inner Regional | 486 | 24.3% | 22.5% to 26.2% | 25% | -0.6 |
| Outer Regional | 391 | 19.6% | 17.9% to 21.3% | 20% | -0.5 |
| Remote | 221 | 11.1% | 9.7% to 12.5% | 10% | +1.5 |
| Very Remote | 98 | 4.9% | 4.0% to 5.9% | 5% | -0.2 |
Pearson chi-square test: χ²(4) = 2.86, p = 0.58; Cohen's w = 0.038 (negligible).
With 2,000 addresses even a trivial gap can reach significance, so the effect size matters as much as the p-value: w = 0.499 for the uniform sample sits at the boundary of a large effect by Cohen's conventions, while w = 0.038 for the weighted one is noise.
Exact multinomial or chi-square
Small samples: which test?
Test fixtures are often small. With 30 weighted addresses, Very Remote expects only 1.5, below the usual “at least 5 per category” rule for the chi-square approximation. The target check on /generate therefore uses the exact multinomial test whenever the sample is small enough to enumerate every possible count vector, the chi-square test when every expected count is at least 5, and a seeded Monte Carlo p-value in between, and it names the test it used.
Is the exact test worth it? The table gives each test's true false-alarm rate at the 5% level when the samples really do follow the target. It is computed exactly, by adding up the probability of every count vector a test would reject: no simulation and no seed.
| Sample size | Count vectors | Exact test | Chi-square | Tests disagree |
|---|---|---|---|---|
| n = 10 | 1,001 | 4.89% | 4.86% | 2.19% |
| n = 20 | 10,626 | 4.97% | 5.28% | 2.39% |
| n = 30 | 46,376 | 4.96% | 5.04% | 1.94% |
| n = 40 | 135,751 | 5.00% | 4.89% | 1.68% |
The honest answer for this target: the chi-square approximation's overall false-alarm rate is close to 5% even at these sizes, a known robustness of Pearson's statistic. Where it matters is the individual decision: the two tests disagree on about 2% of samples, which is roughly two in every five rejections. The exact test never exceeds its level by construction, so it is the one reported when it can be computed. For seed 2025 with n = 30 the two agree: exact p = 0.29, chi-square p = 0.29.
Planning
How many addresses do you need?
Three planning questions that come up when generating test data. Each comes down to the precision of a proportion. Share precision asks where a share with a known target lands, so it uses the exact binomial distribution; the per-area estimate is about a rate nobody knows yet, so it uses the Wilson interval the rest of the lab reports; zero failures uses the exact Clopper-Pearson bound. The normal approximation sits alongside for comparison.
Share precision
How many addresses so each remoteness area's realised share lands within ±E of its target?
Confidence
Targets
| Area | Target | Normal | Exact |
|---|---|---|---|
| Major Cities | 40% | 2,305 | 2,320 |
| Inner Regional | 25% | 1,801 | 1,823 |
| Outer Regional | 20% | 1,537 | 1,569 |
| Remote | 10% | 865 | 892 |
| Very Remote | 5% | 457 | 472 |
Generate at least 2,320 addresses with a weighted design. A stratified design needs no margin: its shares are exact.
Exact is the smallest n from which the binomial chance that the share lands within ±E is at least the confidence, for that n and every larger one. It runs a little above the normal approximation because the window can only hold whole addresses.
Per-area estimate
How many addresses in each area to estimate a rate there (say, how often your parser rejects an address) within ±E?
Confidence
381 per area, 1,905 in total. Use the stratified design with equal quota shares to get exactly that many in each area.
50% is the worst case: if you expect a rate near 0% or 100%, fewer addresses give the same margin. Sizes use the Wilson interval, so they hold where the normal approximation does not.
Zero failures
If none of n addresses breaks your system, how large must n be to say the failure rate is below a bound?
Confidence
Test 299 addresses with no failure: the one-sided exact (Clopper-Pearson) upper bound is then 0.997%. The “rule of three” gives about 300 at 95%.
This bounds the failure rate for addresses drawn the way you drew them. A uniform sample says little about rare places; stratify if every area matters.
Coordinates
Spatial checks
Every point inside its own suburb
Each mock coordinate is drawn by rejection sampling inside the suburb's simplified ABS boundary. The check looks every published point up again against all 1,696 boundaries (the 1,695 suburbs plus the non-addressable SA Remainder), not only the one it was drawn in, so a point in an overlap or across a border fails. The required rate is 100%; the Wilson interval says how sure that makes us.
That lookup uses the same point-in-polygon routine as the sampler, so a bug in the routine could pass both. The routine is therefore checked on its own against Shapely (GEOS), which shares no code with it: for 3,000 seeded points, half spread over the state's bounding box and half placed 10 cm either side of a boundary edge, the two agree on the suburb (or on no suburb) every time.
| Check | Inside | Rate (95% Wilson CI) |
|---|---|---|
| Generator output5,000 addresses, uniform design, seed 2025 | 5,000 | 100.00% (99.92% to 100.00%) |
| Generator output5,000 addresses, population design, seed 2025 | 5,000 | 100.00% (99.92% to 100.00%) |
| Census of every suburb5 points in each of 1,695 suburbs, seed 2025 | 8,475 | 100.00% (99.95% to 100.00%) |
What the check found. Its first run failed one point in 5,000 (seed 2025): a point drawn inside Mobilong, a few centimetres from its edge, crossed into the neighbouring suburb when it was rounded to six decimal places for output. The sampler now rounds before it tests, so a published point is always one that passed. The unit tests keep that exact point as a regression case.
Uniform inside the suburb
Inside a suburb the points should show complete spatial randomness: no clumping, no regular spacing. The Clark-Evans ratio R compares the mean distance from each point to its nearest neighbour with what randomness predicts for that area: R ≈ 1 for random, below 1 for clustered, above 1 for evenly spaced. Because a suburb's edge cuts neighbours off, the expectation uses Donnelly's edge correction with the suburb's perimeter.
| Suburb | Area | Compactness | R, seed 2025 | Mean R, 100 seeds (95% bootstrap CI) | Middle 95% of R | Rejects randomness at 5% |
|---|---|---|---|---|---|---|
| Adelaide | 10.5 km² | 0.51 | 1.006 (p = 0.88) | 0.996 (0.988 to 1.004) | 0.91 to 1.08 | 6/100 (3% to 12%) |
| Glenelg | 0.94 km² | 0.68 | 0.950 (p = 0.20) | 0.999 (0.991 to 1.007) | 0.92 to 1.08 | 6/100 (3% to 12%) |
| Port Adelaide | 4.80 km² | 0.29 | 1.048 (p = 0.23) | 1.003 (0.996 to 1.010) | 0.94 to 1.08 | 3/100 (1% to 8%) |
| Kingscotetwo separate parts | 17.6 km² | 0.47 | 0.963 (p = 0.34) | 1.021 (1.011 to 1.030) | 0.92 to 1.10 | 13/100 (8% to 21%) |
| Mount Gambier | 26.8 km² | 0.51 | 0.976 (p = 0.54) | 0.998 (0.989 to 1.006) | 0.92 to 1.08 | 6/100 (3% to 12%) |
| Coober Pedy | 786 km² | 0.49 | 1.014 (p = 0.72) | 1.000 (0.993 to 1.008) | 0.93 to 1.08 | 5/100 (2% to 11%) |
| 2025 approach (control)Adelaide, 200 points | 10.5 km² | 0.51 | 0.000 (z = -25.4) | not repeated | - | flagged |
| Clustered (control)Adelaide, 200 points | 10.5 km² | 0.51 | 0.340 (z = -16.8) | not repeated | - | flagged |
Each suburb gets 200 points per seed, projected to kilometres around the suburb's centre; for these six the areas are within about 2% of the ABS figures, a gap that comes from simplifying the boundaries, not from the projection. Compactness is 4πA/P² (1 for a circle). For single-part suburbs the mean R sits at 1 and the test rejects in about 5% of seeds, as a calibrated test should. Kingscote, which has two separate parts, sits slightly above 1 (1.021) and is rejected in 13 of 100 seeds: Donnelly's correction was derived for a single rectangle, and a multi-part outline stretches it. Rejection sampling is uniform over the whole outline by construction, so the excess points at the reference value rather than the points; the controls show what real clustering looks like.