Skip to content
Mock Address LabSouth Australia

Methods

How it works, how it is checked, and what it cannot tell you

The data and where it comes from, the generator and its sampling designs, how each claim on the site is checked, the assumptions and limits, the decisions behind them, a data card for the reference table, and what the optional AI feature does.

01 · data

Data provenance

The reference table is rebuilt from Australian Bureau of Statistics open data (ASGS Edition 3, the 2021 Census mesh block counts and SEIFA 2021, all CC BY 4.0) by scripts/build_data.py, which anyone can rerun. The 2025 table it replaces had no recorded source; it stays in original/, and the Data page compares the two row by row. The data card below lists every source, how each field is built and the limits of each one.

No personal information is used anywhere: every input is published, area-level statistics, and every output is synthetic and stamped “MOCK: synthetic test data”.

02 · method

Method

  • The address recipe is the 2025 one, ported line by line: a suburb, a street number from 1 to 999 and one of 49 street names, drawn with a reimplementation of Python's Mersenne Twister so a seed gives the same addresses as the original code would.
  • Five designs choose the suburb: uniform (the 2025 behaviour), remoteness or SEIFA weights (a two-stage draw: category, then suburb), population, and stratified with fixed quotas per remoteness area (DR-002, DR-005).
  • Coordinates are drawn uniformly inside the suburb's ABS boundary by rejection sampling, from a separate seeded stream, rounded to 6 decimals before the inside test (DR-003).
  • The target check on /generate reports each category's realised share with a 95% Wilson interval and a goodness-of-fit test chosen for the sample (exact multinomial, chi-square or Monte Carlo), with n and Cohen's w.
  • The lookup asks Photon for coordinates through a cached, rate-limited route, then finds the suburb by point-in-polygon in the browser (DR-001).

03 · evaluation

Evaluation design

Each claim the site makes has a check that could fail. Statistical claims are checked over many seeds, with the uncertainty of the check itself reported (a rejection rate over 200 seeds is a proportion, so it gets a Wilson interval too). Seeds are fixed and shown next to every result.

ClaimHow it is checkedWhere
The port reproduces the 2025 Python20 recorded runs (11 CLI invocations, 9 generator calls) of the original code, with both random generators seeded, must print the same bytesoriginal.test.ts
The statistics are rightNormal and chi-square functions, Wilson intervals, the exact multinomial test (against brute-force enumeration), the true size of each test, sample sizes (Wilson and exact binomial, against scipy.stats.binom), Clopper-Pearson bounds and the Clark-Evans ratio (against a SciPy k-d tree) checked against SciPy and statsmodelsstats.reference.test.ts
Each design delivers its mix200 seeds of 1,000 addresses per design through the real generator: spread against multinomial theory, Wilson coverage of the target and the test's rejection rate, each with a Wilson interval/sampling#designs
The reported test fits the sampleExact multinomial when the sample can be enumerated, chi-square when every expected count is at least 5, a seeded Monte Carlo p-value otherwise; the false-alarm rate of both tests computed exactly at small n/sampling#small-samples
Every coordinate is inside its suburbEach point looked up again against all 1,696 boundaries (the 1,695 suburbs plus the non-addressable SA Remainder): 5,000 generated points per design and a census of 5 points in every suburb, required rate 100%, with a Wilson lower bound. The point-in-polygon routine itself is checked against Shapely (GEOS) on 3,000 seeded points (pip.reference.test.ts)/sampling#spatial
Coordinates are uniform within a suburbClark-Evans ratio with Donnelly's edge correction in six differently shaped suburbs, over 100 seeds (mean with a bootstrap interval, rejection rate with a Wilson interval), plus two negative controls/sampling#spatial
The numbers in the docs are the code'sThe figures quoted in the README, the decision records and on /sampling are recomputed and comparedclaims.test.ts
The AI client is safe with a keyAdapters, errors, key storage, redaction, the audit log and the proposal review, with the network mockedai.test.ts

04 · assumptions

Assumptions

  • The config.py weights are the target the 2025 README meant. Its six socio-economic bands were never defined; mapping them onto ten IRSAD deciles is my reading (DR-002).
  • A suburb is represented by its ABS Suburb and Locality, and takes the postcode, council and remoteness area of most of its 2021 residents.
  • Uniform within a suburb is the right null model for mock coordinates: they are not meant to look like dwellings.
  • The Mersenne Twister streams behave as independent uniform draws for different seeds (the standard assumption behind seeded simulation).
  • Donnelly's edge correction is adequate for single-part suburbs; the calibration over 100 seeds tests this rather than assuming it.

05 · limitations

Limitations

  • Mock addresses can coincide with real ones: the street names are 49 real Adelaide names used in every suburb. Never use them for mail, identity checks or to stand in for a person.
  • SEIFA describes areas, not people. A mock address in a low-decile suburb says nothing about anyone, and must not be used to represent a disadvantaged person.
  • Postcodes are ABS Postal Areas (approximations of Australia Post postcodes); councils and remoteness are majority rules for suburbs that straddle boundaries.
  • Boundaries are simplified, so a point near an edge can sit in the real neighbouring suburb, and points can fall in parks, lakes or reserves.
  • The Clark-Evans test is miscalibrated for multi-part suburbs (Kingscote is rejected in about 13% of 100 seeds against a nominal 5%); the reference value, not the sampler, is the weak point.
  • The 2021 data are five years old; growth areas have changed since.

06 · next

What I'd change

  • Sample coordinates within residential mesh blocks weighted by dwellings, and use G-NAF to warn when a mock address coincides with a real one.
  • Replace the Clark-Evans reference with a Monte Carlo envelope from an independent uniform sampler, which handles multi-part suburbs properly.
  • Two-way stratification (remoteness by SEIFA decile) with raking, and a bootstrap interval for Cohen's w.
  • An evaluation set of test scenarios with expected settings, to score the AI assistant per model with paired comparisons before recommending one.
  • Rebuild on ASGS Edition 4 and 2026 Census counts when they are published, recorded in a new decision record.

07 · decisions

Decision records

Each record states the decision first, then the context, the options, the reasons, what actually happened (weak numbers included) and what I would change. Records are never edited after the fact; a new record supersedes an old one. The sources are in docs/decisions.

  1. DR-001Open ABS data and Photon instead of Mapbox
  2. DR-002The weighting scheme
  3. DR-003Synthetic coordinates instead of geocoding
  4. DR-004"Describe a test scenario": optional, bring your own key, reviewed and audited
  5. DR-005A stratified design, and which goodness-of-fit test to report

DR-001 · Accepted · 2026-10-09

Open ABS data and Photon instead of Mapbox

Applies to: scripts/build_data.py, web/public/data, /lookup, /api/geocode, /data

Decision

Rebuild the suburb table from Australian Bureau of Statistics open data with a script anyone can rerun, and replace the Mapbox geocoder with the free, keyless Photon geocoder (through a cached, rate-limited route) plus point-in-polygon against ABS boundaries.

Context

The 2025 tool had two external dependencies. Its suburb table (original/data/sa_suburbs_data.csv, 1,894 rows) had no recorded source, and porting it showed why that matters: socio-economic status was 0 in every row, 997 rows had the remoteness level "Not Applicable", and 19 postcodes had lost their leading zero. Its lookup and its coordinates called the Mapbox Geocoding v5 API, which needs a personal access token, and whose terms restrict storing and republishing results. A public demo cannot ship a key, and I have no budget for one.

Decision

  • Rebuild the table from ABS ASGS Edition 3 allocation files and boundaries, 2021 Census mesh block counts and SEIFA 2021 (all CC BY 4.0, direct downloads). Each suburb takes the postcode, council and remoteness area holding most of its residents. scripts/build_data.py does it end to end and writes provenance.json, a row-by-row comparison with the 2025 table.
  • Look up real places with Photon (komoot's OpenStreetMap geocoder), called by a server route with a polite User-Agent, a one-day cache per normalised query and a per-IP limit of 30 searches a minute. The suburb, postcode, council, remoteness and decile then come from point-in-polygon against the ABS boundaries in the browser, so clicking the map works with no geocoder at all.
  • Keep the 2025 table and the stored Mapbox results in original/ as a historical record; the site never reads or serves the Mapbox file.

Options considered

  1. Keep Mapbox with a server-side key. Best geocoding quality, but a cost and abuse risk on a public demo, and storing results is restricted.
  2. Keep the 2025 table and patch it by hand. Quick, but the source would still be unknown and every fix a judgement call.
  3. Google or another commercial geocoder. Same key and cost problem as Mapbox.
  4. ABS open data plus Photon (chosen).

Why

Every field now has a named, licensed, re-downloadable source and a deterministic rule, and the whole table can be rebuilt with one command. Photon needs no key, and caching plus rate limiting keep the site within its fair-use expectations. The governance gain is provenance: a reviewer can trace any value on the site back to an ABS file.

What happened

  • 1,694 of the 1,894 names in the 2025 table match an ABS suburb. Of those, 99.3% (1,682) agree on the postcode and 98.2% (1,664) on the council after a council-name crosswalk. The 200 unmatched names are pastoral stations and outback places that the ABS folds into larger localities; all 200 were "Not Applicable" for remoteness.
  • Majority rules have a cost: a suburb split between postcodes, councils or remoteness areas gets one value. The data card lists this and the other limits (a SAL is not always the gazetted suburb; SEIFA is area-level).
  • Photon's quality is OpenStreetMap's: good for towns and streets, weaker for rural addresses and new estates, and the public instance can be slow or unavailable. The lookup degrades to clicking the map when it is.

What I'd change

  • Use the Geocoded National Address File (G-NAF, open) to check that generated street names exist in a suburb, or to warn when a mock address coincides with a real one.
  • Host a Photon instance (or use a paid tier) if the lookup ever carries real traffic.
  • Rebuild on ASGS Edition 4 and 2026 Census counts when they are published, and record the change in a new decision record rather than editing this one.

DR-002 · Accepted · 2026-10-09

The weighting scheme

Applies to: web/src/lib/generator/weights.ts, /generate, /sampling

Decision

Keep uniform sampling as the faithful default, and add the weighting the 2025 README promised as a two-stage draw (first a category with the config.py weight, then a suburb uniformly inside it), plus population weighting, with the weights renormalised over whatever categories survive the filters.

Context

The 2025 README said generation followed remoteness and socio-economic weights. config.py defined them (DEFAULT_REMOTENESS_WEIGHTS: 0.40, 0.25, 0.20, 0.10, 0.05; DEFAULT_SOCIOECONOMIC_WEIGHTS over six bands 0 to 5), but the import was commented out, so every suburb was equally likely. The six socio-economic bands were never defined, and the table's socio-economic column was 0 in every row.

Decision

  • Uniform (the default, labelled "as built in 2025") keeps the original behaviour, so the replay and the generator can be compared.
  • Remoteness weights: pick a remoteness area with probability equal to its weight, then a suburb uniformly within it. The realised area mix then matches the weights exactly in expectation, whatever the number of suburbs per area.
  • SEIFA weights: the six undefined bands are spread over the ten IRSAD deciles. Decile d belongs to band round((d − 1) × 5 / 9) and each band's weight is split evenly across its deciles, so band totals are preserved. Suburbs with no published decile get zero weight in this mode, and the page says how many.
  • Population: suburbs in proportion to 2021 usual residents.
  • Filters first, then renormalise: filters remove suburbs; the category weights are renormalised over the categories that still have suburbs, and a mode whose weights are all zero returns an error instead of silently falling back (the 2025 code's worst habit).

Options considered

  1. Single-stage weights per suburb (each suburb gets its area's weight). Simpler, but the area mix would then depend on how many suburbs each area has: Outer Regional has 582 suburbs and Major Cities 424, so Outer Regional would be over-drawn relative to the promise.
  2. Two-stage: category, then suburb (chosen). The promise is about the category mix, so the design should deliver exactly that mix.
  3. Raking or iterative proportional fitting to several margins at once (remoteness and SEIFA together). More expressive, but the 2025 tool never combined them and it is harder to explain.

Why

The two-stage design makes the target explicit and checkable: the expected share of each area equals its weight, so a goodness-of-fit test against the weights is a test of the implementation. Keeping uniform as the default keeps the revival honest about what the 2025 code actually did.

What happened

  • With seed 2025 and 2,000 addresses, the weighted design gives χ²(4) = 2.86, p = 0.58 against the target; the uniform design gives χ²(4) = 497, p < 0.001, Cohen's w = 0.50.
  • Over 200 seeds of 1,000 addresses (/sampling), the goodness-of-fit test rejects the weighted design in 6 of 200 seeds (3.0%, 95% CI 1.4% to 6.4%), consistent with its 5% level, and the 95% Wilson intervals cover the target in 93.5% to 97% of seeds per area. The uniform design is rejected in all 200.
  • The decile mapping is my reading of six undefined bands, not a documented intent. It is shown in the UI next to the editable weights.

What I'd change

  • Weight by dwellings rather than residents for address-like realism (the Census mesh block counts include dwellings).
  • Allow joint targets (remoteness by SEIFA) with raking, and test the joint mix.
  • Let users load a target mix from a file, with the realised mix and its test exported alongside the addresses.

DR-003 · Accepted · 2026-10-09

Synthetic coordinates instead of geocoding

Applies to: web/src/lib/geo.ts, web/src/lib/generator/generate.ts, /generate, /sampling#spatial

Decision

Give each mock address a point drawn uniformly at random inside its suburb's ABS boundary, from its own seeded random stream, instead of geocoding anything.

Context

The 2025 generator geocoded the suburb name through Mapbox for every address (_get_suburb_coordinates), so all addresses in a suburb shared one point, and it needed an API key. The street addresses themselves are invented (a random number from 1 to 999 and one of 49 Adelaide street names), so geocoding them would either fail or snap to a real street that happens to share the name.

Decision

  • Rejection sampling inside the simplified ABS polygon: draw a point uniformly in the bounding box until it falls inside; fall back to the suburb's label point after 2,000 failed tries (it has never happened).
  • A separate random stream for coordinates, seeded from the address seed, so turning coordinates on or off never changes the addresses.
  • Publish coordinates to 6 decimal places, and test the rounded point for being inside (see What happened).
  • Stamp every output "MOCK: synthetic test data".

Options considered

  1. Geocode each mock address. Needs a key and a network call per address, mostly fails, and where it succeeds it places a mock address on a real street, which makes it look more real than it is.
  2. One point per suburb (the 2025 behaviour). Every address in a suburb on the same spot: useless for anything spatial.
  3. Uniform inside the suburb boundary (chosen). Keyless, instant, reproducible, and visibly synthetic.
  4. Inside residential mesh blocks, weighted by dwellings. More realistic placement, at the cost of shipping mesh block geometry (much larger).

Why

Test data needs coordinates that are plausible (in the right suburb), reproducible (seeded) and unmistakably not real. Uniform sampling inside the boundary is the simplest design that meets all three, and it is checkable: a point-in-polygon validation and a uniformity statistic can confirm the implementation does what it says.

What happened

  • The point-in-polygon validation (every point looked up again against all 1,696 boundaries) failed one point in 5,000 on its first run (seed 2025): a point drawn inside Mobilong a few centimetres from the edge was rounded to 6 decimals after the inside test and crossed into the neighbouring suburb. The sampler now rounds before testing. After the fix: 5,000 of 5,000 for the uniform design, 5,000 of 5,000 for the population design and 8,475 of 8,475 in a census of 5 points per suburb (95% Wilson lower bound 99.92% and 99.95%). The failing point is a regression test.
  • The Clark-Evans ratio with Donnelly's edge correction sits at 1 for single-part suburbs (mean over 100 seeds between about 0.99 and 1.01) and the z-test rejects randomness in about 5% of seeds, as it should. For Kingscote, which has two parts, the mean is about 1.02 and the test rejects in 13 of 100 seeds: the edge correction assumes one rectangle, so the reference value, not the sampler, is off there. Two negative controls are flagged clearly: all points on one geocoded spot gives R = 0, and a clustered pattern gives R ≈ 0.34.
  • Uniform inside a suburb means points can fall in parks, lakes, reserves and airports. That is stated on the data card: the points are mock locations, not dwellings.

What I'd change

  • Sample within residential mesh blocks weighted by dwellings (option 4), and measure how much more realistic the placement is against G-NAF density.
  • Replace the Clark-Evans reference with a Monte Carlo envelope from an independent uniform sampler (for example, triangulating the polygon), which handles multi-part and irregular suburbs properly.

DR-004 · Accepted · 2026-10-09

"Describe a test scenario": optional, bring your own key, reviewed and audited

Applies to: /generate, /ai-log, /methods#ai-use, web/src/lib/ai

Decision

An optional assistant that turns a plain-language description of the test data someone needs into a proposed generator configuration, called from the visitor's browser with their own key, validated against a schema and the reference table, applied only field by field after the visitor reviews it, and recorded in a local audit log.

Context

The generator has many settings (count, seed, five designs, weights, four filters, coordinates, format), and testers think in scenarios ("a small fixture that covers every remoteness area"), not settings. Translating one into the other is a reasonable job for a language model. It is also a small, realistic test of governing generative AI: the model's output changes what software does, so it must be constrained, checked and reviewed. There is no budget for an API key and no server for AI, and nothing on the site may depend on AI.

Decision

  • Bring your own key, browser only. The visitor pastes an Anthropic (default, Claude Haiku 4.5; Claude Sonnet 5.5 optional) or OpenAI key in AI settings. It stays in sessionStorage unless they choose "remember on this device" (localStorage), and "Forget key" removes it. Calls go straight from the browser to the provider (anthropic-dangerous-direct-browser-access: true for Anthropic). The key never reaches this site, a log or the repository, and anything key-like is redacted before it is stored or shown.
  • Minimal input. The request carries the scenario (at most 1,000 characters), the current settings, and the catalogue of allowed values (remoteness areas with suburb counts, decile meaning, the 71 council names). No suburb list, no generated addresses, nothing about the visitor.
  • Structured output, checked twice. The reply must match a JSON schema (provider structured-output modes) and is validated again with zod against exactly the same contract. A second, domain check runs against the reference table: unknown suburbs or councils, out-of-range deciles or seeds, and malformed weights are flagged and cannot be applied; a count above 5,000 is clamped with a note. The model must also list anything the generator cannot do (unit numbers, PO boxes, real addresses, people's names) instead of pretending.
  • Human in the loop. The proposal is labelled "AI-generated" and shown as a table of changes (setting, now, proposed). The visitor ticks which to apply; nothing is generated until they press Generate. Applying every recommended change is recorded as "accepted", a different selection as "edited" (with exactly what was applied), and "Reject" as "rejected".
  • Audit. Every call, including failures, refusals and replies that fail validation, is written to an IndexedDB audit log without the key: time, feature, provider, model, input, raw output, latency, token usage, the checks and the human decision. It is viewable and exportable (JSON, CSV) at /ai-log.
  • Refusals. Claude Sonnet 5.5 requests opt into Anthropic's server-side refusal fallback (fallbacks: "default", beta server-side-fallback-2026-07-01); a refusal that still comes back is shown plainly and logged.

Options considered

  1. No AI. Safe, but misses a genuinely useful shortcut and a governance showcase.
  2. A server proxy with my key. Costs money, invites abuse and puts a secret on a server.
  3. Let the model generate addresses directly. Unreproducible, unverifiable and likely to produce real-looking or real addresses: the opposite of what a mock generator is for.
  4. Model proposes configuration only; deterministic code generates (chosen). The model never touches the data, only the knobs, and every knob it turns is visible and reversible.

Why

Restricting the model to proposing configuration keeps the part that matters (the seeded, tested generator) deterministic, and makes the AI's contribution small enough to review at a glance. The schema, the reference-table check and the field-by-field review are three independent controls, and the audit log makes each decision traceable. The design is informed by the Australian Government's policy for the responsible use of AI in government, the EU AI Act's transparency principles and the NIST AI Risk Management Framework; it is not a compliance claim.

What happened

  • The provider adapters, error handling, key storage, redaction, audit log and the proposal review are unit-tested with the network mocked (ai.test.ts): the key only ever travels in a request header, and invented suburbs or councils are flagged and cannot be applied even if ticked.
  • No live call was made during development or CI, because the project has no key. The request shapes follow the providers' documented APIs; a real call is the first thing to check after deployment.
  • The domain check catches names that do not exist, not names that exist but are wrong for the scenario (a real council that is not the one the visitor meant). That is what the review table is for.

What I'd change

  • Build a small evaluation set of scenarios with expected settings and score each model on it (field-level agreement with Wilson intervals, paired between models) before recommending one.
  • Offer an in-browser model as a no-key option once small models handle structured output reliably.
  • Let the visitor save a reviewed configuration as a named preset, so a good proposal is reused without another call.

DR-005 · Accepted · 2026-10-09

A stratified design, and which goodness-of-fit test to report

Applies to: web/src/lib/generator, web/src/lib/stats, /generate (target check), /sampling

Decision

Add a stratified design with fixed remoteness quotas (largest remainder), and have the target check report the exact multinomial test when the sample can be enumerated, Pearson's chi-square when every expected count is at least 5, and a seeded Monte Carlo p-value otherwise, always with the test's name, n and Cohen's w.

Context

The weighted design (DR-002) draws each address independently, so the area counts are random: a small fixture of 30 addresses can easily contain no Very Remote address even though the target is 5%. Testers who need every area represented want fixed counts. Separately, the 2025-era target check used the chi-square approximation for every sample, and small samples break its rule of thumb (expected counts of at least 5).

Decision

  • Stratified design: quotas n·w_h rounded by the largest-remainder method (ties to the earlier area), then suburbs uniformly within each area. The order of areas is a Fisher-Yates shuffle with the same random stream (identical to Python's random.shuffle), so output is still seeded and reproducible. The target check says "fixed by design" for the remoteness breakdown instead of running a meaningless test.
  • Test selection: exact multinomial (probability ordering, as in R's EMT package) when there are at most 200,000 possible count vectors; otherwise chi-square if the smallest expected count is at least 5; otherwise a chi-square statistic with a Monte Carlo p-value (as R's chisq.test(simulate.p.value = TRUE), seeded, up to 10,000 draws). The p-value is never reported without the test's name, n and Cohen's w.
  • Statistics live in web/src/lib/stats/ and are checked against SciPy and statsmodels values recorded by scripts/make_stats_reference.py.

Options considered

  1. Chi-square only, with a warning (the earlier behaviour). Simple, but the warning puts the burden on the reader.
  2. Exact test only. Correct, but enumeration explodes: 5,000 addresses over 11 decile categories is far beyond reach.
  3. Monte Carlo for everything. Always available, but adds simulation noise to results that could be exact or asymptotic.
  4. Choose by feasibility and expected counts (chosen).

Why

The design should state its guarantee: weighted draws are random with a known distribution, stratified draws are fixed. The test should match the situation and say which one it is, because a p-value without its test is not reproducible. Cohen's w is reported because with thousands of addresses a trivial gap reaches significance.

What happened

  • The exact false-alarm rate of each test at 5%, computed by enumerating every count vector for the config.py target: n = 10, exact 4.89%, chi-square 4.86%; n = 20, 4.97% against 5.28%; n = 30, 4.96% against 5.04%. The chi-square approximation is better behaved for this target than the rule of thumb suggests, which is worth saying plainly. The two tests still disagree on about 2% of samples, roughly two in five rejections, so the choice changes individual verdicts even though the overall rate barely moves.
  • The stratified design gives exact shares in all 200 seeds of the replicate study, and it has one cost: a shorter run is no longer a prefix of a longer run with the same seed, because the quotas and the shuffle depend on the count. The uniform and weighted designs keep that property.

What I'd change

  • Offer stratification by SEIFA decile and by remoteness and decile together (a two-way quota table).
  • Report a confidence interval for Cohen's w (by bootstrap over the counts) instead of the point estimate alone.

08 · data card

Data card: the reference table

The reference table behind every generated address, map and lookup on the site. It replaces the 2025 table (original/data/sa_suburbs_data.csv, 1,894 rows, source not recorded), which stays in the repository unchanged.

At a glance

ItemDetail
Filesweb/public/data/suburbs.json (table), sal-sa.geojson (boundaries), sa-context.geojson (state outline), provenance.json (comparison with 2025)
UnitOne row per ABS Suburb and Locality (SAL) 2021 in South Australia
Rows1,696 (1,695 addressable; "SA Remainder" is never used for addresses)
Built byscripts/build_data.py (uv, PEP 723), from direct ABS downloads, no login
GeographyAustralian Statistical Geography Standard (ASGS) Edition 3, July 2021 to June 2026, GDA2020
Population2021 Census usual residents, by mesh block (1,777,698 in total)
Socio-economicSEIFA 2021, Index of Relative Socio-economic Advantage and Disadvantage (IRSAD)
LicenceContains ABS data, © Commonwealth of Australia, CC BY 4.0
Personal informationNone. Every input is published area-level statistics

Sources

SourcePublisherLicenceUsed for
ASGS Ed. 3 allocation files: mesh block to SAL, LGA, POA, SA1; SA1 to Remoteness AreaABSCC BY 4.0Joining every mesh block to its suburb, council, postal area and remoteness area
ASGS Ed. 3 digital boundaries: SAL 2021 and states (GDA2020)ABSCC BY 4.0Suburb polygons, label points, the state outline
Census 2021 mesh block countsABSCC BY 4.0Usual residents per mesh block (the weights for every majority rule)
SEIFA 2021, Suburbs and LocalitiesABSCC BY 4.0IRSAD score and deciles

Exact file URLs are listed in provenance.json and on the site's Data page.

How each field is built

FieldMeaningRule
code, official, nameSAL 2021 code and namename is upper case without the " (SA)" disambiguator
postcode, postcodesABS Postal AreaThe Postal Area holding most of the suburb's 2021 residents (four digits, leading zero kept); postcodes lists every one that overlaps
council, lgaCodeLocal Government Area 2021The LGA holding most residents; land area breaks ties for empty localities
ra, raShareABS Remoteness Area 2021The Remoteness Area holding most residents, and the share of residents in it
decileSa, decileAus, irsadSEIFA 2021 IRSADDeciles ranked within South Australia and nationally; null where the ABS publishes none (83 rows)
popUsual residents, Census 2021Sum over the suburb's mesh blocks
areaKm2AreaABS SAL area
labelA point guaranteed inside the simplified boundaryUsed when a coordinate cannot be sampled (it never has been)

Boundaries are simplified with mapshaper 0.7.81 (Visvalingam, keeping 10% of removable vertices, shared edges preserved) and rounded to 0.0001 degrees (about 11 m).

Quality checks

  • 1,694 of the 1,894 names in the 2025 table match an ABS suburb; of those, 99.3% agree on the postcode and 98.2% on the council once council names are crosswalked. The 200 names only in the 2025 table are pastoral stations and outback places the ABS folds into larger localities.
  • Every label point falls inside its own polygon, and every one of 5,000 generated coordinates (and a census of 5 points in each of 1,695 suburbs) falls inside the suburb its address names, looked up again against all 1,696 boundaries (the 1,695 suburbs plus the non-addressable SA Remainder), not only its own. That lookup shares its point-in-polygon routine with the sampler, so the routine is also checked against Shapely (GEOS) on 3,000 seeded points, half of them 10 cm from a boundary: they agree on every point (see the Sampling page).
  • Projected polygon areas agree with the ABS areas to within about 1% for 90% of suburbs larger than half a square kilometre.

Known limits (read before relying on a field)

  • A SAL is not always the gazetted suburb. ABS Suburbs and Localities are built from mesh blocks to approximate the officially gazetted suburbs and localities. Edges can differ, a few names differ, and small localities can be absorbed into larger ones.
  • A postcode is the majority Postal Area, not the delivery postcode. ABS Postal Areas approximate Australia Post postcodes; they are not the official delivery list. A suburb split between postcodes gets the one most of its residents live in.
  • Council and remoteness are majority rules. Suburbs that straddle a council or a remoteness boundary carry one value for the whole suburb.
  • SEIFA is area-level, not individual. An IRSAD decile describes the residents of an area in aggregate. It says nothing about any person or household, and a mock address in a low-decile suburb must never be used to stand in for a disadvantaged person (the ecological fallacy). The ABS publishes no SEIFA for very small populations.
  • Population is from 2021. Growth areas have changed since; the 2026 Census will need a rebuild once its counts and the next ASGS edition are published.
  • Simplified boundaries. Simplification moves edges by up to tens of metres, so a mock point near an edge can sit in the real neighbouring suburb, and a point can land in a park, a lake or a reserve inside the suburb. Coordinates are mock locations, not dwellings.

Intended use

Generating synthetic, clearly stamped test addresses for software testing; illustrating sampling designs; looking up which suburb a real point is in.

Out of scope: postal delivery, identity verification, eligibility or service decisions, any inference about real people or households, and pairing mock addresses with names to create realistic-looking personal records.

Maintenance

Rebuild with uv run scripts/build_data.py (downloads about 235 MB into scripts/.cache/). The script rewrites every file above and the comparison in provenance.json. The decision to rebuild from open data rather than keep the 2025 table is recorded in DR-001.

09 · AI use

AI use statement

What the AI does

One optional feature, “Describe a test scenario” on /generate: it reads a plain-language description of the test data you need and proposes generator settings. The proposal is labelled AI-generated, checked against a schema and the reference table, and shown as a table of changes that you review and apply (or not) field by field.

What it never does

It never generates addresses (the seeded, tested generator does), never changes a setting without your review, never runs without your own key, and is never needed: every page works without it. It is not used for any statistic on the site.

What is sent, and where

From your browser directly to the provider you choose (api.anthropic.com or api.openai.com), never through this site: your scenario (at most 1,000 characters), the current settings and the list of remoteness areas, deciles and councils. No suburb list, no generated addresses, nothing about you. The provider's own terms and retention apply to what you send. Do not put personal information in a scenario.

Your key

Bring your own: Anthropic (Claude Haiku 4.5 by default, or Claude Sonnet 5.5) or OpenAI (gpt-5-mini by default, editable). It stays in this browser (session storage, or local storage if you tick “remember”), travels only in the request header to the provider, and “Forget key” removes it. It is never sent to this site, logged or stored in the audit log. The site's Content-Security-Policy backs this up in the browser: pages may only connect to this site, the two providers and the map tile server.

Human in the loop, and the record

Every call is written to an audit log in this browser (IndexedDB), with the time, provider, model, the exact input, the raw output, latency, token usage, which checks the proposal failed, and your decision: accepted, edited (with what you applied) or rejected. Failed calls are logged too. View and export it as JSON or CSV at /ai-log. The design is in DR-004.

The design is informed by the Australian Government's policy for the responsible use of AI in government (Digital Transformation Agency), the EU AI Act's transparency principles and the NIST AI Risk Management Framework. It is a personal project and makes no claim of compliance with any of them.