Sustainability Language
Representativeness
The extent to which evidence reflects the population, places, periods and conditions about which a conclusion is being made.
Expert review openNo editor-accepted expert review yetDefinition
The extent to which evidence reflects the population, places, periods and conditions about which a conclusion is being made.
Overview
“A sample can be large, precise and still describe the wrong people. ”
Representativeness is the bridge between what was observed and what is claimed. A survey may contain thousands of records, a satellite model may cover millions of hectares and a programme database may list every registered participant. None of that establishes that the evidence reflects the wider population to which the conclusion is applied. The concept begins with a target population.
Who or what is the evidence intended to represent? All coffee farmers in a country, members of participating cooperatives, farms within mapped districts, workers present during an audit, or households reachable by mobile telephone? Each is a different population. A conclusion becomes unreliable when the population described in the claim is broader than the population that had a realistic chance of entering the data.
The 1936 Literary Digest poll remains a useful warning.
It received more than two million responses and confidently predicted that Alf Landon would defeat Franklin D. Roosevelt in the United States presidential election. Roosevelt won decisively. The problem was not sample size. The magazine's lists overrepresented wealthier citizens, and those who chose to respond differed from those who did not. A very large sample amplified confidence without repairing selection.
Sustainability systems reproduce the same problem in quieter ways. Farmers registered with a cooperative are easier to locate than isolated producers. Households with smartphones are easier to survey than households without connectivity. Farms beside roads are easier to visit than farms several hours away. Workers on permanent contracts are easier to interview than seasonal or migrant labour.
The people most exposed to risk may be precisely those least likely to appear in the system. Representativeness has several dimensions. Geographic coverage matters where environmental and livelihood conditions vary across landscapes. Seasonal coverage matters where labour, income or pesticide use changes during harvest. Demographic coverage matters where gender, age, migration status or land tenure shapes experience.
Enterprise coverage matters where small suppliers operate differently from large ones. A sample may be representative in one dimension and distorted in another. Probability sampling provides a defensible basis for generalisation because units have known or calculable chances of selection. It does not guarantee a perfect result.
Coverage gaps, non-response, inaccurate frames and measurement error can still distort estimates.
Non-probability samples can also be useful for rapid learning, qualitative insight or finding rare cases, but the claim must remain proportionate to the design. Convenience data do not become population evidence because they are stored in a dashboard. Weighting can reduce known imbalances when reliable information about the population exists.
It cannot recreate groups that were never observed or correct characteristics that were not measured. A dataset of accessible farms cannot be made fully representative of inaccessible farms by multiplying records from the accessible group. Statistical adjustment is a tool, not a substitute for field design. Time also changes representativeness.
A farmer registry assembled three years ago may no longer reflect current production. A baseline collected before a drought may not describe the population after migration, crop failure or conflict. Repeated measurement must account for who leaves, who enters and whose data disappear.
Attrition is not only a technical issue; it can remove the households experiencing the greatest difficulty. The practical discipline is to make the inference visible. State the target population, sampling frame, selection process, response rate, exclusions and known differences between participants and non-participants. Disaggregate results where averages conceal important groups.
Test whether conclusions change under alternative weights or assumptions. Where representativeness cannot be established, narrow the wording: “among surveyed participants” is more credible than “farmers in the region. ” Representativeness is not achieved once. It is maintained through updated frames, deliberate inclusion, field quality control and honest limits on generalisation.
The objective is not to make every dataset universal.
It is to ensure that the reach of the claim never exceeds the reach of the evidence.
Practical application
Define the target population before designing collection. Build or evaluate the sampling frame, identify groups with weak coverage and choose a selection method suited to the intended inference. Track contact, eligibility, refusal and non-response separately, and compare the achieved sample with known population characteristics. Report the boundaries of the evidence alongside the result.
Use weighting and sensitivity analysis where justified, but document assumptions. Commission supplementary qualitative or purposive work for groups likely to remain invisible, without presenting those findings as statistically representative.
Why it matters
Sustainability decisions distribute resources, scrutiny and opportunity. Unrepresentative evidence can direct support towards those already visible while overlooking remote producers, informal workers or vulnerable communities. It can also create false confidence in averages that do not describe the people most affected.
Common misconception
Representativeness is often confused with sample size. A larger sample reduces random sampling error under an appropriate design; it does not automatically correct selection, coverage or non-response bias. A small, well-designed probability sample can support stronger inference than millions of self-selected records.
Connections
Sampling determines how observations enter a dataset. Representativeness asks whether those observations support the intended generalisation. Bias explains systematic distortion within the design or measurement process. Data Quality addresses the wider fitness of information for use, while Inclusion asks whose experience is missing and why.
A question worth asking
Who had little or no chance of appearing in this evidence, and how would the conclusion change if their experience were visible?
Selected references
United Nations Statistics Division. 2005. Household Sample Surveys in Developing and Transition Countries: Design, Implementation and Analysis. Kish, L. 1965. Survey Sampling. Groves, R. M. et al. 2009. Survey Methodology, Second Edition. Bethlehem, J. 2010. Selection Bias in Web Surveys. International Statistical Review 78(2): 161-188. ISO 20252:2019.
Market, Opinion and Social Research, Including Insights and Data Analytics - Vocabulary and Service Requirements.
Review
Public comments appear only after editor acceptance. Draft comments stay in the review queue.
Reviewers choose the definition or an overview paragraph, leave a comment or replacement, and attach evidence or a source link.
Editors compare reviewer cards side by side. AI may help find agreement, conflicts, unsupported claims and possible source issues.
Only an editor-accepted synthesis changes the public page. Reviewer identities are shown only with consent and verification.
Submitted reviews stay private until accepted.
Loading verified endorsements… Endorsements are not votes and never determine publication.
Endorse this definition
Endorse the exact version shown here. This is not a vote, and publication remains an editorial decision.
Review board
Comment on a specific line. Each reviewer stays separate until an editor accepts a merged draft.
Each person comments on the definition or overview in their own draft card, with role, evidence and suggested wording kept together.
AI can compare comments against the current text, flag conflicting claims, surface missing evidence and identify where reviewers agree.
An editor merges compatible suggestions into a draft change, checks sources, records disagreements and decides what can be published.