Skip to content
Mock Address LabSouth Australia

Sampling design

Does each design deliver the mix it promises?

The generator can draw addresses three ways: uniformly over suburbs (what the 2025 code did), weighted by remoteness, or stratified with fixed quotas. This page puts each one through the real generator with fixed seeds and asks two questions: does the sample hit its target mix, with honest uncertainty, and do the coordinates land where they should?

Every number is computed at build time by the same code that runs on /generate, from the seeds shown, so a rebuild gives the same numbers. How the checks were designed is on Methods.

200 seeds each

Three designs, one target

The target is the remoteness mix in config.py that the 2025 README promised (40% Major Cities, 25% Inner Regional, 20% Outer Regional, 10% Remote, 5% Very Remote). Each design generated 1,000 addresses with seeds 1 to 200. A good weighted design should miss the target only by chance: its 95% Wilson intervals should cover the target about 95% of the time, and a goodness-of-fit test at the 5% level should reject it about 5% of the time.

  • Uniform (2025)

    Every suburb equally likely, as the 2025 code actually behaved.

    Seeds where the test rejects the target at 5%
    200 of 200, 100.0% (95% CI 98.1% to 100.0%)
    Distance from the target (Cohen's w)
    mean 0.445 (medium), middle 95% 0.379 to 0.514
  • Weighted

    Each address independently picks a remoteness area with the config.py weights, then a suburb.

    Seeds where the test rejects the target at 5%
    6 of 200, 3.0% (95% CI 1.4% to 6.4%)
    Distance from the target (Cohen's w)
    mean 0.060 (negligible), middle 95% 0.025 to 0.100
  • Stratified

    Fixed quotas per remoteness area (largest remainder), then suburbs uniformly inside each.

    Goodness-of-fit test
    Not tested: the quotas fix every share (200 of 200 seeds exactly on target)
    Distance from the target (Cohen's w)
    0.000 in every seed (no sampling variation)
Realised share of each remoteness area across 200 seeds of 1,000 addresses, for the three designs, against the config.py target.
Uniform (2025)WeightedStratifiedTarget (config.py)Line: middle 95% of seeds · dot: mean
Major Cities
Inner Regional
Outer Regional
Remote
Very Remote

Hover or focus an area for the numbers behind each line.

Uniform misses badly and consistently: Major Cities gets about 25% instead of 40%, because only 25% of suburbs are in a major city. The test catches it in every seed, and w around 0.45 (a medium effect, close to large) says the gap is substantial, not merely detectable.

Weighted behaves as theory says. The spread across seeds matches the multinomial standard deviation (for Major Cities 1.52% observed against 1.55% expected), the Wilson intervals cover the target in 93.5% to 97.0% of seeds per area, and the test rejects in 3.0% (95% CI 1.4% to 6.4%), consistent with its 5% level.

Stratified has nothing to test: the quotas fix every share, so all 200 of 200 seeds give exactly the target, and it is reported as fixed rather than given a test result or coverage interval. The suburbs, street numbers and names inside each area are still random. Use it when every area must be represented in a small sample; use weighted when you want the realistic variation of independent draws.

Show every number as a table
DesignAreaTargetDesign shareMean shareMiddle 95% of seedsSD (theory)Wilson CI covers target
Uniform (2025)Major Cities40%25.0%25.20%22.4% to 27.6%1.29% (1.37%)0/200 (0.0% to 1.9%)
Uniform (2025)Inner Regional25%20.2%20.18%18.1% to 22.7%1.15% (1.27%)6/200 (1.4% to 6.4%)
Uniform (2025)Outer Regional20%34.3%34.18%31.7% to 36.5%1.38% (1.50%)0/200 (0.0% to 1.9%)
Uniform (2025)Remote10%11.8%11.75%9.8% to 13.6%0.98% (1.02%)116/200 (51.1% to 64.6%)
Uniform (2025)Very Remote5%8.6%8.70%7.0% to 10.3%0.89% (0.89%)1/200 (0.1% to 2.8%)
WeightedMajor Cities40%40.0%40.07%36.9% to 42.6%1.52% (1.55%)190/200 (91.0% to 97.3%)
WeightedInner Regional25%25.0%24.99%22.0% to 27.6%1.39% (1.37%)188/200 (89.8% to 96.5%)
WeightedOuter Regional20%20.0%19.89%17.7% to 22.2%1.20% (1.26%)194/200 (93.6% to 98.6%)
WeightedRemote10%10.0%10.03%8.2% to 11.8%0.95% (0.95%)191/200 (91.7% to 97.6%)
WeightedVery Remote5%5.0%5.01%3.6% to 6.4%0.70% (0.69%)187/200 (89.2% to 96.2%)
StratifiedMajor Cities40%40.0%40.00%40.0% to 40.0%0.00% (0.00%)n/a (fixed)
StratifiedInner Regional25%25.0%25.00%25.0% to 25.0%0.00% (0.00%)n/a (fixed)
StratifiedOuter Regional20%20.0%20.00%20.0% to 20.0%0.00% (0.00%)n/a (fixed)
StratifiedRemote10%10.0%10.00%10.0% to 10.0%0.00% (0.00%)n/a (fixed)
StratifiedVery Remote5%5.0%5.00%5.0% to 5.0%0.00% (0.00%)n/a (fixed)

“Design share” is what each design draws in expectation; the theoretical SD is √(p(1 − p)/n) at that share (zero for fixed quotas). Coverage intervals are 95% Wilson intervals over the 200 seeds. Coverage is not reported for the stratified design: its shares are fixed, so an interval around them would describe no real uncertainty.

Seed 2025, n = 2,000

One sample, read properly

The README's key result, recomputed: the same seed and size through the uniform and the weighted design. Each share comes with its 95% Wilson interval; the residual (O − E)/√E shows which areas drive the test statistic.

Uniform (2025 behaviour)
Remoteness areaCountShare95% Wilson CITargetResidual
Major Cities47323.6%21.8% to 25.6%40%-11.6
Inner Regional39119.6%17.9% to 21.3%25%-4.9
Outer Regional72236.1%34.0% to 38.2%20%+16.1
Remote22611.3%10.0% to 12.8%10%+1.8
Very Remote1889.4%8.2% to 10.8%5%+8.8

Pearson chi-square test: χ²(4) = 497.45, p < 0.001; Cohen's w = 0.499 (medium).

Weighted (config.py remoteness)
Remoteness areaCountShare95% Wilson CITargetResidual
Major Cities80440.2%38.1% to 42.4%40%+0.1
Inner Regional48624.3%22.5% to 26.2%25%-0.6
Outer Regional39119.6%17.9% to 21.3%20%-0.5
Remote22111.1%9.7% to 12.5%10%+1.5
Very Remote984.9%4.0% to 5.9%5%-0.2

Pearson chi-square test: χ²(4) = 2.86, p = 0.58; Cohen's w = 0.038 (negligible).

With 2,000 addresses even a trivial gap can reach significance, so the effect size matters as much as the p-value: w = 0.499 for the uniform sample sits at the boundary of a large effect by Cohen's conventions, while w = 0.038 for the weighted one is noise.

Exact multinomial or chi-square

Small samples: which test?

Test fixtures are often small. With 30 weighted addresses, Very Remote expects only 1.5, below the usual “at least 5 per category” rule for the chi-square approximation. The target check on /generate therefore uses the exact multinomial test whenever the sample is small enough to enumerate every possible count vector, the chi-square test when every expected count is at least 5, and a seeded Monte Carlo p-value in between, and it names the test it used.

Is the exact test worth it? The table gives each test's true false-alarm rate at the 5% level when the samples really do follow the target. It is computed exactly, by adding up the probability of every count vector a test would reject: no simulation and no seed.

Sample sizeCount vectorsExact testChi-squareTests disagree
n = 101,0014.89%4.86%2.19%
n = 2010,6264.97%5.28%2.39%
n = 3046,3764.96%5.04%1.94%
n = 40135,7515.00%4.89%1.68%

The honest answer for this target: the chi-square approximation's overall false-alarm rate is close to 5% even at these sizes, a known robustness of Pearson's statistic. Where it matters is the individual decision: the two tests disagree on about 2% of samples, which is roughly two in every five rejections. The exact test never exceeds its level by construction, so it is the one reported when it can be computed. For seed 2025 with n = 30 the two agree: exact p = 0.29, chi-square p = 0.29.

Planning

How many addresses do you need?

Three planning questions that come up when generating test data. Each comes down to the precision of a proportion. Share precision asks where a share with a known target lands, so it uses the exact binomial distribution; the per-area estimate is about a rate nobody knows yet, so it uses the Wilson interval the rest of the lab reports; zero failures uses the exact Clopper-Pearson bound. The normal approximation sits alongside for comparison.

Share precision

How many addresses so each remoteness area's realised share lands within ±E of its target?

Confidence

Targets

Addresses needed per remoteness area
AreaTargetNormalExact
Major Cities40%2,3052,320
Inner Regional25%1,8011,823
Outer Regional20%1,5371,569
Remote10%865892
Very Remote5%457472

Generate at least 2,320 addresses with a weighted design. A stratified design needs no margin: its shares are exact.

Exact is the smallest n from which the binomial chance that the share lands within ±E is at least the confidence, for that n and every larger one. It runs a little above the normal approximation because the window can only hold whole addresses.

Per-area estimate

How many addresses in each area to estimate a rate there (say, how often your parser rejects an address) within ±E?

Confidence

381 per area, 1,905 in total. Use the stratified design with equal quota shares to get exactly that many in each area.

50% is the worst case: if you expect a rate near 0% or 100%, fewer addresses give the same margin. Sizes use the Wilson interval, so they hold where the normal approximation does not.

Zero failures

If none of n addresses breaks your system, how large must n be to say the failure rate is below a bound?

Confidence

Test 299 addresses with no failure: the one-sided exact (Clopper-Pearson) upper bound is then 0.997%. The “rule of three” gives about 300 at 95%.

This bounds the failure rate for addresses drawn the way you drew them. A uniform sample says little about rare places; stratify if every area matters.

Coordinates

Spatial checks

Every point inside its own suburb

Each mock coordinate is drawn by rejection sampling inside the suburb's simplified ABS boundary. The check looks every published point up again against all 1,696 boundaries (the 1,695 suburbs plus the non-addressable SA Remainder), not only the one it was drawn in, so a point in an overlap or across a border fails. The required rate is 100%; the Wilson interval says how sure that makes us.

That lookup uses the same point-in-polygon routine as the sampler, so a bug in the routine could pass both. The routine is therefore checked on its own against Shapely (GEOS), which shares no code with it: for 3,000 seeded points, half spread over the state's bounding box and half placed 10 cm either side of a boundary edge, the two agree on the suburb (or on no suburb) every time.

CheckInsideRate (95% Wilson CI)
Generator output5,000 addresses, uniform design, seed 20255,000100.00% (99.92% to 100.00%)
Generator output5,000 addresses, population design, seed 20255,000100.00% (99.92% to 100.00%)
Census of every suburb5 points in each of 1,695 suburbs, seed 20258,475100.00% (99.95% to 100.00%)

What the check found. Its first run failed one point in 5,000 (seed 2025): a point drawn inside Mobilong, a few centimetres from its edge, crossed into the neighbouring suburb when it was rounded to six decimal places for output. The sampler now rounds before it tests, so a published point is always one that passed. The unit tests keep that exact point as a regression case.

Uniform inside the suburb

Inside a suburb the points should show complete spatial randomness: no clumping, no regular spacing. The Clark-Evans ratio R compares the mean distance from each point to its nearest neighbour with what randomness predicts for that area: R ≈ 1 for random, below 1 for clustered, above 1 for evenly spaced. Because a suburb's edge cuts neighbours off, the expectation uses Donnelly's edge correction with the suburb's perimeter.

This generator (seed 2025)200 points, R = 1.01: no sign of clustering or regular spacing.
2025 approachAll 200 addresses on one geocoded point, R = 0.00.
Clustered control200 points around five centres, R = 0.34: flagged, as it should be.
SuburbAreaCompactnessR, seed 2025Mean R, 100 seeds (95% bootstrap CI)Middle 95% of RRejects randomness at 5%
Adelaide10.5 km²0.511.006 (p = 0.88)0.996 (0.988 to 1.004)0.91 to 1.086/100 (3% to 12%)
Glenelg0.94 km²0.680.950 (p = 0.20)0.999 (0.991 to 1.007)0.92 to 1.086/100 (3% to 12%)
Port Adelaide4.80 km²0.291.048 (p = 0.23)1.003 (0.996 to 1.010)0.94 to 1.083/100 (1% to 8%)
Kingscotetwo separate parts17.6 km²0.470.963 (p = 0.34)1.021 (1.011 to 1.030)0.92 to 1.1013/100 (8% to 21%)
Mount Gambier26.8 km²0.510.976 (p = 0.54)0.998 (0.989 to 1.006)0.92 to 1.086/100 (3% to 12%)
Coober Pedy786 km²0.491.014 (p = 0.72)1.000 (0.993 to 1.008)0.93 to 1.085/100 (2% to 11%)
2025 approach (control)Adelaide, 200 points10.5 km²0.510.000 (z = -25.4)not repeated-flagged
Clustered (control)Adelaide, 200 points10.5 km²0.510.340 (z = -16.8)not repeated-flagged

Each suburb gets 200 points per seed, projected to kilometres around the suburb's centre; for these six the areas are within about 2% of the ABS figures, a gap that comes from simplifying the boundaries, not from the projection. Compactness is 4πA/P² (1 for a circle). For single-part suburbs the mean R sits at 1 and the test rejects in about 5% of seeds, as a calibrated test should. Kingscote, which has two separate parts, sits slightly above 1 (1.021) and is rejected in 13 of 100 seeds: Donnelly's correction was derived for a single rectangle, and a multi-part outline stretches it. Rejection sampling is uniform over the whole outline by construction, so the excess points at the reference value rather than the points; the controls show what real clustering looks like.