How do we know it works?
What we checked before publishing, how close the generated population comes to the Census tables, what we did for privacy, and where the data are weakest.
Generated people and households, not real people, families, or addresses.
Did this release pass its gates?
Before reading any figure
This version publishes a single model run: there is no across-run range, and no directional comparison between cells is permitted.
In practice: read each figure on its own, without ranking it or comparing it with another as if the difference were certain.
Scorecard notes (original wording)
Written for the scorecard’s general format, which expects several model runs. In this release, with a single run, there are no intervals; “R=1” means one run, and the “ok” status means the evaluation and the privacy audit passed.
- Intervals show variation across model runs, not calibrated confidence intervals.
- Status changes to ok only after both the R=1 evaluation and national privacy audit pass.
Release gates
Passed
Model runs
A single run
Privacy audit
Passed
The gates the release depends on
From the model card, release 1.0.3
- Coverage complete: 21 scored tables in all 3,092 parishes.
- Structural violations: 0 (impossible records; see the glossary).
- The child share is never more than 10% below the published one, in any region or size band. None fell below it: the worst case is 0.01% above (the limit was −10%).
- The exact under-15 total in all 3,092 parishes, in both the private and the resident universe.
- An integrity audit of every parish, with 0 errors.
How close do they come to INE’s tables?
Fit error against INE’s tables (SRMSE): 0 would be identical, and 0.10 means each cell is off by roughly 10% of a typical cell’s size. For each parish it is measured in one go over all the cells of the 12 person tables used in the fit, on the resident population; the chart by size shows the median across the parishes of each band. It is reported, not used as a publication gate, and it is not the “parish’s typical error” that decides the tiers (see the glossary). Because these tables were in the fit, the error measures how close the fit came, not how well the model predicts what it did not see.
Fit error, by parish size
Median across parishes, all cells of the 12 person tables; bands by INE’s residents; same scale as the chart below
- Fewer than 500 residents882 parishes0.104
- 500 to 2,000 residents1,215 parishes0.056
- 2,000 to 10,000 residents755 parishes0.017
- More than 10,000 residents240 parishes0.008
View as table: Fit error against INE’s tables, by parish size
| Size | Parishes | Fit error |
|---|---|---|
| Fewer than 500 residents | 882 parishes | 0.104 |
| 500 to 2,000 residents | 1,215 parishes | 0.056 |
| 2,000 to 10,000 residents | 755 parishes | 0.017 |
| More than 10,000 residents | 240 parishes | 0.008 |
Typical error, by table
Median across parishes of the error of each of the 12 fitted person tables (the household tables have no published median); same scale as the chart above
- Age (5-year bands)0.017
- Marital status0.063
- Education0.086
- Labour-force status0.049
- Main source of livelihood0.077
- Labour × education0.086
- Labour × source of livelihood0.035
- Nationality0.001
- Religion0.002
- Status in employment0.047
- Activity sector (four groups)0.035
- De facto union0.017
View as table: Typical error against INE’s tables, by table
| Table | Typical error |
|---|---|
| Age (5-year bands) | 0.017 |
| Marital status | 0.063 |
| Education | 0.086 |
| Labour-force status | 0.049 |
| Main source of livelihood | 0.077 |
| Labour × education | 0.086 |
| Labour × source of livelihood | 0.035 |
| Nationality | 0.001 |
| Religion | 0.002 |
| Status in employment | 0.047 |
| Activity sector (four groups) | 0.035 |
| De facto union | 0.017 |
With few people, each one weighs more in every table: that is why every parish carries a quality tier, shown at the top of its page.
Which tables were used?
Fitted and scored
- People per household and family nuclei per household
- Dwelling wheelchair accessibility (BGRI)
- The 12 person tables in the chart above: age (5-year bands), marital status, education, labour-force status, main source of livelihood, labour × education, labour × source of livelihood, nationality, religion, status in employment, activity sector (four groups) and de facto union
Scored only, not fitted
- Place of work or study
- Means of transport
- Industry (CAE section)
- Occupation (CPP major group)
- Households by number of employed people, and by active and dependent people
The scorecard publishes a median only for the 12 person tables in the chart. Single-year age is also measured and counts towards the “worst table” that decides the tier (in most parishes it is the worst one), but it is not one of the 21 coverage tables nor of the 12 fitted person tables. The quality file gives, for each parish, its worst table and that table’s error (worst_constraint); the data page says which table each code is.
The errors in the charts are measured on the tables used in the fit: they say how close the fit came. The model card mentions an independent check, with a household table held out of the fit, but does not publish its error (see “Does it get right what it did not see?”).
What do quality A, B and C mean?
Every parish is published, and every parish answers with its own figures; each parish page shows its tier at the top. The tier combines the fit to INE’s tables and the number of residents, and says how carefully to read the numbers, but it hides nothing. The counts are the published release’s.
776
parishes
A parish of 2,000 or more residents where the generated population closely reproduces the tables INE publishes.
Read the figures as a portrait close to INE’s tables for the parish.
705
parishes
A parish of 500 or more residents with a close fit to INE’s tables. Under 2,000 residents a parish sits in B even when its fit is as close as tier A’s.
Also close to the tables; in a parish of under 2,000 residents, B may reflect only its size, not a weaker fit.
1,611
parishes
A parish of under 500 residents, or one whose typical error or worst table (in the larger ones, almost always single-year age, which the site’s answers do not use) is past the tier B thresholds: read the numbers with more care.
Tier C holds two kinds of parish, read in different ways: see below.
- Under 500 on the publication count (883): they are C for their size. Each question counts only its own group (for example, those who live alone), which can be a handful of people: one or two shift a share, and 100% can be one or two. The top of each parish page says how many residents it has.
- 500 or more on the publication count (728): they are C for their worst table (almost always single-year age) or, in a few, for their typical error. Single-year age sits outside the 12 fitted person tables and the site’s answers do not use it: in those parishes the caution applies mostly to anyone using that column of the microdata. The quality file names each one’s worst table, and each parish page says what put it in tier C.
Each tier’s thresholds: A, typical error up to 0.10, worst table up to 0.18 and 2,000 or more residents; B, up to 0.15 and 0.26 with 500 or more residents; C, the rest. A parish under 500 residents sits in tier C. The 500 and 2,000 thresholds use the publication count, the smaller of INE’s residents and the generated people: a parish where INE counted 500 people can fall below 500 on that count.
“0.0%”: no generated person or household in that category, or so few that the share rounds to zero.
This is release 1.0.3, which replaced three releases dated 5 October 2026 (1.0.2 never reached GitHub); the generated population is the same in all of them. What changed between them
What does each term mean?
- Typical error (SRMSE)
- Compares, cell by cell, the generated count with the one INE published: it is the root-mean-square deviation between the two, divided by the average size of a cell in that table. It has no unit: 0 means identical; 0.10 means that, roughly, each cell is off by about 10% of a typical cell’s size. The lower, the closer. In a table with one very large category (nationality, religion) the error stays near zero even when the small categories are well off: do not read a low value as a guarantee for the small groups.
- Fit error (all cells)
- A parish’s typical error computed in one go over all the cells of the 12 fitted person tables, on the resident population. The chart by parish size shows the median of this error across the parishes of each band; it is the model card’s figure.
- A parish’s typical error
- The median of the parish’s 12 typical errors, one per fitted person table (person_srmse_median in the quality file). It is the first criterion of tiers A and B. It is not the figure in the chart by parish size, which pools the cells of the 12 tables into one error: the two are close but do not coincide.
- Worst table
- The largest typical error among the parish’s person tables, single-year age included, which is scored but is not one of the 12 fitted person tables. It is the tiers’ second criterion; in most parishes the worst table is single-year age (worst_constraint in the quality file).
- Structural violation
- A record the logic of the data does not allow, checked record by record before release: for example a child under 15 who is employed or has a source of livelihood, a university degree before 18, an age below zero or above 115, a person without a household, or a household whose size is not the number of its people. This release has none. It does not count the people who sit in a combination INE publishes as zero for their parish: those combinations are possible, they just do not occur in that parish in INE’s tables, and they are listed under the limitations.
- Maximum-entropy fit
- The way the generated candidates are brought close to INE’s tables while changing each one’s weight as little as possible.
- AUC (membership inference)
- Measures, from 0.5 (chance) to 1 (certainty), whether a test can tell who was in the sample. The published figure is the excess over chance: near zero means the data do not help.
- SA/CO replay
- A reference method that builds the population by copying sample records. It is the yardstick for matches: by design it copies almost everything. The audit itself flags it as not ready for use (not_ready); it is a reference only.
Do these data protect the people who answered the Census?
The records are generated: they carry no names or addresses, and there is no record-to-record link to respondents. The national privacy audit, run on the published population, passed. The sample referred to is the Census 2021 Public Use File, from which the model learned how attributes relate.
- Exact matches
- 10.5% of people and 0.8% of households match a sample record on 13 attributes. For comparison, a replay that reuses sample records (the SA/CO benchmark, which the audit itself flags as not ready for use) reaches 99.1% of people and 99.9% of households.
- Common combinations
- On the 11 attributes the model generates, 80% of synthetic persons share an attribute combination with some sample person. Those are common profiles, and the distance to the closest record on those 11 attributes is still larger than between the sample’s own people.
- Distance to the closest record
- In a sample of 5,000 records, synthetic persons are further from the sample than sample persons are from each other: 10.8% of the synthetic ones exactly match their closest record, against 74.3% real-to-real. It is a different measure from the exact match on 13 attributes (every person).
- Membership inference
- No excess signal: telling whether someone was in the sample does not get easier with these data (−0.002 AUC for persons, −0.007 for households).
- Attribute inference
- No advantage: guessing someone’s attribute from the others does not get easier than with the real data themselves (−0.002).
Synthetic does not mean that accidental attribute matches are impossible.
Where are the data weakest?
Known limitations of the generated population (the same from release 1.0.0 to 1.0.3), declared rather than hidden. The model card leaves the plan to remove them to the next release (1.1, a corrected synthetic population for 2021); every correction is recorded in the errata.
Workplace and commuting are the weakest attributes
Work or study location, transport mode, industry (CAE section) and occupation (CPP major group) sit much further from the published tables than the demographic and household tables do. They are scored but, by design, not fitted: the model generates them, conditioned on the region. Of the work fields, labour-force status, the four-group activity sector and status in employment are fitted (see “Which tables were used?”).
Some combinations INE publishes as zero
About 22,000 people sit in a combination that INE publishes as zero for their parish. Most are in the household-activity table, which is held out of the fit as an independent check. In the fitted person tables there are about 7,900 such entries (at most 0.08% of people; one person can count in more than one table), mostly in labour × education, education, labour-force status and marital status. The model card links them to combinations that do not occur in the training sample.
Small published cells may be missing
About 47,000 small published cells are not reproduced; 87% of them hold a single person.
A few unusual families
Well under 1% of family units have a shape that real households in the sample do not have: for example a family unit without a member aged 15 or over, or a mother–child gap above 50 years.
Clock-bound optimisation
In 116 large parishes the final search stopped at its time budget rather than at convergence. These results depend on machine speed; the producer’s internal record names how each one ended (it is not part of the published files).
Residents of collective quarters are partial records
People living in care homes and other collective quarters are appended after the fit, from INE’s published counts. Sex and age come from those counts and education is imputed from similar records; employment status, marital status and nationality are filled. By design, the family nucleus and the work fields (status in employment, sector, occupation, industry, place of work and means of transport) are left empty.
Housing and family nuclei
Tenure (owner or tenant) and the number of rooms are published in the microdata but not fitted per parish: in villages the generated population has more renters than the real one. And some generated family nuclei have a single member, where INE’s nucleus always has at least two.
Does it get right what it did not see?
Not yet known: comparing the generated population with figures the model did not use (retrodiction) awaits a separate tabulation, and the scorecard marks it “not yet available”. When it is done, it comes with a new release of the population.
Before generating this population, the producer registered error ranges for that out-of-fit check, with the earlier engine. The errors on this page are measured on the fitted tables, so they do not compare with those ranges, and the page does not say whether they landed “inside” or “below” them.