Data, Technology & Verification Systems
Sampling methodology
The disciplined selection of a subset from a defined population and frame so that evidence can support estimates about the wider population with known design and uncertainty.
Expert review openNo editor-accepted expert review yetDefinition
The disciplined selection of a subset from a defined population and frame so that evidence can support estimates about the wider population with known design and uncertainty.
Overview
“A large sample can be precisely wrong when the wrong people had the chance to be selected. ”
Sampling allows organisations to learn about a population without observing every unit. It is essential in surveys, audits, field verification and data-quality review. The cost advantage is obvious. The harder issue is whether the selected evidence supports the conclusion being made. The target population comes first.
It is the group about which the organisation intends to draw conclusions: all coffee households in a region, all workers employed during harvest or all plots supplying a buyer in a season. A sample cannot be judged without knowing the population it is meant to represent. The sampling frame is the operational route to that population.
It may be a farmer register, employee list, map, transaction database or set of communities. The United Nations survey guidance emphasises that frame quality is fundamental.
If remote farmers, informal workers or unregistered plots are absent, no increase in sample size can give them a chance of selection. Probability sampling gives units known, non-zero probabilities of selection. This allows design-based estimation and quantification of sampling error. Simple random sampling is only one method. Stratification can ensure important groups are covered.
Cluster and multi-stage designs can reduce travel cost. Each choice affects precision, weighting and analysis. Sample size matters less than is often assumed. A very large convenience sample can produce stable estimates of the reachable group and biased estimates of the population. Surveying thousands of cooperative members does not represent farmers outside cooperatives.
Precision around the wrong answer is not representativeness. Non-response creates another selection process.
People who are absent, distrustful, busy or afraid may differ systematically from those who participate. Repeated call-backs, safe interview conditions, translated materials and adjusted weights can reduce bias. Reporting the response rate alone does not show whether non-respondents differed on the outcome. Design effects should be recognised. Farmers in the same village may have similar prices, climate and services.
Treating clustered observations as independent can make uncertainty appear smaller than it is. Analysis should reflect the actual selection design and weights rather than use formulas intended for simple random samples. Sampling in audits has a related but different purpose. Auditors often use risk-based or judgemental selection to test controls and find significant exceptions.
This can be appropriate for assurance but does not automatically support statistical estimates of prevalence. A sample designed to find risk should not be presented as a representative survey. Ethics and burden matter. The same accessible households are repeatedly surveyed because they are known to projects. Others remain invisible.
Sampling plans should distribute burden, compensate appropriately where justified and avoid collecting data that will not be used. Inclusion is not improved by repeatedly asking marginalised people questions without response. Documentation should make inference possible.
Users need the target population, frame, method, strata, stages, sample size, inclusion probabilities, response, weights, substitutions and limitations. Without these, representative becomes a marketing adjective rather than a methodological claim.
The discipline is to trace every conclusion back to selection. Who could be selected, who was selected, who responded and how were differences handled? Sampling quality is not the number interviewed. It is the credibility of the path from observed units to the population claim.
Practical application
Define the target population and evaluate frame coverage before calculating sample size. Choose probability designs where population estimates are required, and stratify for groups whose outcomes may differ or whose exclusion would matter. Track selection and response at every stage, calculate weights and account for clustering. Separate risk-based audit samples from representative surveys.
Publish design and limitations with results, including groups not covered by the frame.
Why it matters
Sampling determines whose experience becomes evidence. Weak design can make excluded people statistically invisible and give false confidence to programme results, risk estimates and public claims.
Common misconception
A sample is often called representative because it is large or geographically dispersed. Representativeness depends on the target population, frame, selection probabilities, response and analysis, not size alone.
Connections
Representativeness examines whether the evidence supports inference to the target population. Bias identifies systematic error, and Uncertainty expresses what remains unknown. Data Quality and Materiality determine whether the sample is adequate for the decision and consequence.
A question worth asking
Who had no realistic chance of entering your sample, and how would the conclusion change if their experience differs from the people you reached?
Selected references
United Nations Statistics Division. 2008. Designing Household Survey Samples: Practical Guidelines. United Nations Statistics Division. 2026. Handbook of Surveys on Households and Individuals: Foundations and Emerging Approaches. Cochran, W. G. 1977. Sampling Techniques, Third Edition. Groves, R. M. et al. 2009. Survey Methodology, Second Edition. Lohr, S. L. 2021. Sampling: Design and Analysis, Third Edition.
Review
Public comments appear only after editor acceptance. Draft comments stay in the review queue.
Reviewers choose the definition or an overview paragraph, leave a comment or replacement, and attach evidence or a source link.
Editors compare reviewer cards side by side. AI may help find agreement, conflicts, unsupported claims and possible source issues.
Only an editor-accepted synthesis changes the public page. Reviewer identities are shown only with consent and verification.
Submitted reviews stay private until accepted.
Loading verified endorsements… Endorsements are not votes and never determine publication.
Endorse this definition
Endorse the exact version shown here. This is not a vote, and publication remains an editorial decision.
Review board
Comment on a specific line. Each reviewer stays separate until an editor accepts a merged draft.
Each person comments on the definition or overview in their own draft card, with role, evidence and suggested wording kept together.
AI can compare comments against the current text, flag conflicting claims, surface missing evidence and identify where reviewers agree.
An editor merges compatible suggestions into a draft change, checks sources, records disagreements and decides what can be published.