This page explains, for a general reader, how the Synthetic Population of Portugal (release 1.0.3) was generated, what it can and cannot be used for, and where its limits are. The full technical account is the model card.

What is a synthetic population?

It is a set of computer-generated people and households, one record per person, whose totals follow the tables INE published from the 2021 Census for each parish. No record is a real person.

Generated people and households, not real people, families, or addresses.

What is it for, then? The published tables give, for each parish, how many people there are by age, by education or by household type, but at such a small scale they rarely cross those variables with one another. A synthetic population makes those cross-tabulations possible, within the quality rules this page describes. It can serve as a foundation layer for demographic analysis, but it does not replace official statistics.

To the best of our knowledge, the first open-access synthetic population to cover every parish in Portugal, generated from the 2021 Census.

How was it generated?

The result does not come straight out of a neural network: a model invents possible people and households, and then the ones whose totals come as close as possible to INE's tables are chosen (a constrained generative pipeline). There are five steps:

  1. Learn how attributes go together. The computer learns, from the public-use sample of the 2021 Census (the Public Use File), which combinations of age, education, work and family are common in a person and their household. That model (an autoregressive generator) is trained once and then frozen. Single-year ages and the type of workplace are generated by it, not derived afterwards.
  2. Fit each parish to its tables. For each parish the model invents many possible households (the candidates), avoiding the combinations INE publishes as zero in that parish (a mask). The mask restricts the candidates, but the result can still leave some people in those combinations: in the fitted person tables, about 7,900 entries (see "Where are the data weakest?"). It then gives more weight to the candidates that bring the totals closer to INE's tables for the parish, changing each one's weight as little as possible (a maximum-entropy tilt). Before that, a check would refuse any parish whose tables imply totals that are inconsistent with one another; none of the 3,092 was refused.
  3. Draw whole households. A person cannot be split: whole households are drawn from the candidates, as close as possible to the counts the parish's tables ask for (an integer allocation). In parishes of under 500 people that draw is solved to the optimum (an exact selection, with no approximation), which still does not guarantee INE's total. The generated total is not always INE’s: in 762 parishes it differs, from 119 people fewer to 30 more; nationally, there are 2,625 people fewer (10,340,441 generated, against 10,343,066 residents according to INE). The quality tiers use the smaller of the two counts, the publication count. Parish by parish, in the quality file.
  4. Add the residents of collective quarters. People living in care homes and other collective quarters are appended after the fit, from INE's published counts, as partial records: sex and age come from those counts, education is imputed from similar records, and the family nucleus and the work fields are left empty (the column dictionary says which). Fine occupation codes are derived at this stage.
  5. Publish with a declared schema and a quality tier. Every published column has a declared meaning, and the column dictionary explains each one; everything else is dropped, with the reason recorded. The release carries no link between partners. Each parish comes out with a quality tier, A, B or C (see "When does a figure appear?").

What is fitted and what is derived?

The figures on this site use nine fields. Six are groupings of a generated field whose INE table is in the fit (five-year age bands group the generated single-year age, and the age bands join those groups); two are derived by a fixed rule and no table in the fit checks them (for example, "lives alone" means the household has one person); and one is a flag. The table gives, for each, the microdata column or the rule that computes it from the columns, with the same letters as the column dictionary.

AgeGrouping of a fitted field
age_5y (D)Five-year band of the generated age. INE’s sex × five-year-age table is in the fit.
EducationGrouping of a fitted field
education_level_coarse5 (D)Five levels from the generated education code (11 levels). The education table is in the fit.
Employment statusGrouping of a fitted field
employment_status_coarse3 (D)Employed, unemployed or inactive, from the generated code (7 categories). Labour-force status is in the fit.
People in the householdGrouping of a fitted field
hh_size_bin (D)The household’s number of people, 5 or more together. The household-size table is in the fit.
Family nucleiGrouping of a fitted field
hh_type_top (D)The number of family nuclei, counted from the household’s people. The nuclei-per-household table is in the fit.The site’s answer (“What families do they form?”): This answer was computed before the microdata’s final packaging and may differ slightly from what they give; the producer will correct it.
Collective living quartersFlag
is_institutional (flag)1 for a collective living quarter appended from INE’s counts, 0 for a private household.
Lives aloneDerived, outside the fit
Not a column:“Yes” when the person’s household has one person (hh_size = 1). No table in the fit checks it.
Child and someone 65+ in the householdDerived, outside the fit
Not a column:“Yes” when the household has at least one person under 15 (age < 15) and another aged 65 or over (age ≥ 65). No table in the fit checks it.
Age bandGrouping of a fitted field
Not a column:Under 45, 45 to 64, 65 or over, from age: each band joins whole five-year groups, and the five-year table is in the fit.

The household questions (people per household, family nuclei, several generations), "who lives alone" and "people aged 65 or over who live alone" count private households only, INE's household universe: people living in a care home or other collective living quarters are not part of them.

In the downloadable microdata, each column also carries its provenance: generated by the model (G), derived from another column (D), made consistent after generation (C), or assigned from the geographic code (X). The column dictionary explains each one, and the quality page says which tables were fitted and which were only scored.

To see the idea with invented people, in an imagined neighbourhood, there is an illustrated example. It uses no data from this population.

When does a figure appear?

Always. All 3,092 parishes are published, and every one answers every question with its own figures. Each has a quality tier, which combines how close the generated population comes to INE's tables and the number of residents. The tier is a reading guide, not a filter: it hides no figure, it says how carefully to read it. A small parish has few people in each table, and each person weighs more.

Parishes with fewer than 500 residents always sit in tier C. Among the tier C parishes of 500 or more, the worst table is almost always single-year age, which is scored but is not one of the fitted person tables, and which the site's answers do not use: in those parishes the tier C caution applies mostly to anyone using single-year age from the microdata. Each parish page says what put it in its tier. The terms used to measure the fit (typical error, worst table) are explained, with an example, in the quality page's glossary.

  • Quality A776 parishes

    A parish of 2,000 or more residents where the generated population closely reproduces the tables INE publishes.

  • Quality B705 parishes

    A parish of 500 or more residents with a close fit to INE’s tables. Under 2,000 residents a parish sits in B even when its fit is as close as tier A’s.

  • Quality C1,611 parishes

    A parish of under 500 residents, or one whose typical error or worst table (in the larger ones, almost always single-year age, which the site’s answers do not use) is past the tier B thresholds: read the numbers with more care.

In this release, every question is answered with the parish’s own figures, and each parish page states its quality tier.

No cell is suppressed: every category of a question appears, empty ones included. “0.0%”: no generated person or household in that category, or so few that the share rounds to zero.

This is release 1.0.3, which replaced three releases dated 5 October 2026 (1.0.2 never reached GitHub); the generated population is the same in all of them. What changed between them.

What does a single run change?

This version publishes a single model run: there is no across-run range, and no directional comparison between cells is permitted.

That is why the site's figures carry no intervals, no rankings of parishes and no "more than" comparisons between results. A second run will follow publication, as a measurement; from two runs onward, the variation between them can be shown again. Even then, we will only call it a confidence interval if its coverage is demonstrated.

Do these data protect the people who answered the Census?

The records are generated and are not intended to represent identifiable people: they carry no names or addresses, and the internal pointers to sample records have been dropped.

In the national audit, 10.5% of people and 0.8% of households match a sample record on 13 attributes. A replay that reuses sample records (the SA/CO benchmark, which the audit itself flags as not ready for use) reaches 99.1% of people and 99.9% of households. On the 11 attributes the model generates, 80% of synthetic persons share an attribute combination with some sample person: those are common profiles, and the distance to the closest record is still larger than between the sample’s own people. There is no excess membership-inference signal (guessing whether someone was in the sample used for training) and no attribute-inference advantage (guessing someone’s attribute from the others). The audit passed.

Synthetic does not mean that accidental attribute matches are impossible.

The details are on the quality page.

What is it for, and what is it not for?

Suitable for

  • Descriptive demographic exploration.
  • Small-area household and population analysis within the published quality rules.
  • Journalism and civic-data applications.
  • Education and reproducible research.
  • Aggregate scenario and poststratification work (reweighting a survey so it matches the population). This release publishes no uncertainty (a single run): treat the figures as point values.
  • Testing tools that require realistic but non-identifying population records.

Not suitable for

  • Identifying, locating, or making decisions about real people.
  • Treating a synthetic row as an individual, family, or address.
  • Unrestricted cross-tabulation in tiny or weak-quality parishes.
  • Parish-level claims based on fields that are not fitted per parish: industry, occupation, place of work, means of transport, tenure and number of rooms (the column dictionary says which they are).
  • Causal conclusions about policy effects.
  • Behavioural prediction or agent-based simulation without a separate validated model.
  • Replacing official Census statistics.
  • Legal, credit, insurance, employment, policing, or eligibility decisions.

Where are the data weakest?

These are the limitations of the generated population, the same from release 1.0.0 to 1.0.3, declared rather than hidden:

  1. Workplace and commuting are the weakest attributes. Work or study location, transport mode, industry (CAE section) and occupation (CPP major group) sit much further from the published tables than the demographic and household tables do. They are scored but, by design, not fitted: the model generates them, conditioned on the region. Of the work fields, labour-force status, the four-group activity sector and status in employment are fitted (see “Which tables were used?”).
  2. Some combinations INE publishes as zero. About 22,000 people sit in a combination that INE publishes as zero for their parish. Most are in the household-activity table, which is held out of the fit as an independent check. In the fitted person tables there are about 7,900 such entries (at most 0.08% of people; one person can count in more than one table), mostly in labour × education, education, labour-force status and marital status. The model card links them to combinations that do not occur in the training sample.
  3. Small published cells may be missing. About 47,000 small published cells are not reproduced; 87% of them hold a single person.
  4. A few unusual families. Well under 1% of family units have a shape that real households in the sample do not have: for example a family unit without a member aged 15 or over, or a mother–child gap above 50 years.
  5. Clock-bound optimisation. In 116 large parishes the final search stopped at its time budget rather than at convergence. These results depend on machine speed; the producer’s internal record names how each one ended (it is not part of the published files).
  6. Residents of collective quarters are partial records. People living in care homes and other collective quarters are appended after the fit, from INE’s published counts. Sex and age come from those counts and education is imputed from similar records; employment status, marital status and nationality are filled. By design, the family nucleus and the work fields (status in employment, sector, occupation, industry, place of work and means of transport) are left empty.
  7. Housing and family nuclei. Tenure (owner or tenant) and the number of rooms are published in the microdata but not fitted per parish: in villages the generated population has more renters than the real one. And some generated family nuclei have a single member, where INE’s nucleus always has at least two.

How is it corrected and updated?

  • The releases published on GitHub (1.0.0, 1.0.1 and 1.0.3) stay available and citable; 1.0.2 was never published. Each answer's permanent link names the release it refers to; a parish page always shows the current release, so cite it with the date you read it.
  • Corrections are recorded in a public errata log.
  • A revision of INE's sources triggers a documented patch or a new evaluation.
  • The next release (1.1) will be a corrected synthetic population for 2021. A later release, with 2025 data, will say which fields were updated to 2025, which were carried forward from 2021, and which are newly modelled.
  • The model weights are not published: that would require its own memorisation, privacy and licence review.

How do I cite it?

estimador.pt publishes the synthetic population under the CC BY-NC 4.0 licence (Attribution-NonCommercial 4.0 International): commercial use is not covered. The INE data it is based on are reused under INE’s CC BY 4.0 licence. When you redistribute the data (the files, or tables taken from them), include this attribution:

Source: Instituto Nacional de Estatística, IP – Portugal (2021 Population and Housing Census; reference period: 2021). Modified information: this is a synthetic population produced by estimador.pt from the Census 2021 published marginals and Public Use File (FUP), INE information reused under the CC BY 4.0 licence; it is not official INE microdata and INE is not responsible for its content. The synthetic population is published by estimador.pt under the CC BY-NC 4.0 licence (Attribution-NonCommercial 4.0 International).

In a news story, a chart or a card, the short form is enough, with a link to the data page: Source: INE, 2021 Census (CC BY 4.0) · information modified by estimador.pt (CC BY-NC 4.0)

Cite as: estimador.pt, População Sintética de Portugal v1.0.3 (2026), CC BY-NC 4.0.

The files, the column dictionary, the checksums and the citation file are on the data page.