Sustainability Language
Data Quality
The degree to which data are fit for their intended use across dimensions such as relevance, accuracy, completeness, consistency, timeliness, coherence and interpretability.
Expert review openNo editor-accepted expert review yetDefinition
The degree to which data are fit for their intended use across dimensions such as relevance, accuracy, completeness, consistency, timeliness, coherence and interpretability.
Overview
“Data can be perfectly accurate and still be useless for the decision being made. ”
Data quality is often reduced to accuracy. A value is checked against a source, passes validation and is declared good. Yet the value may be outdated, incomplete, defined differently from comparable records or irrelevant to the question. Quality is multidimensional because data serve decisions, not databases.
The United Nations Statistical Quality Assurance Framework describes good statistical outputs as fit for purpose and identifies dimensions including relevance, accuracy, reliability, coherence, timeliness, accessibility and interpretability. Other frameworks use slightly different lists, but the central idea is consistent: quality depends on use and context. Relevance asks whether the data address the decision.
A precise count of farmers trained does not answer whether practice changed. Accuracy asks how close the value is to reality.
Completeness asks whether required records and fields are present. Timeliness asks whether information is current enough. Coherence and consistency ask whether values can be combined and compared without contradiction. Consider a traceability database containing three records for the same farmer, each with a different spelling and plot size. Every record may reflect a real submission.
Together they inflate reach, production and risk coverage. Duplicate control and entity resolution are therefore quality functions with direct consequences for claims. Location data create another example. A polygon may be geometrically valid and mapped to the wrong farm. A coordinate may be accurate to five decimal places because a device generated it while the enumerator stood at the cooperative office.
Technical precision is not evidence of correct provenance. Quality begins at design.
Vague questions, poorly defined indicators and impossible response categories create defects that cleaning cannot repair. Enumerators may interpret a household, worker, farm or adoption differently. A data dictionary, tested instrument and training aligned with real decisions prevent inconsistency before collection. Incentives matter.
Teams rewarded for the number of complete records may invent values or copy defaults. Suppliers may report optimistic data when access to a market depends on the result. Quality controls should examine plausibility, independent evidence and the conditions under which data were produced, not only format. Uncertainty should be part of quality.
Estimates, imputation, remote-sensing classification and sampling all contain error. Hiding uncertainty does not improve the dataset; it encourages overconfident use. Metadata should identify source, method, date, coverage and confidence sufficient for the user to judge fitness. Correction requires versioning.
If an organisation changes a deforestation status or farmer count, historical reports and downstream users may need updating. Silent overwrite destroys the evidence trail. Quality governance should record what changed, why, who approved it and which decisions were affected. Quality has cost and trade-offs. Faster data may be less complete; more precise data may create privacy risk; broad coverage may reduce depth.
The appropriate balance depends on the decision and consequence. High-stakes exclusion requires stronger evidence than exploratory planning. The discipline is to make fitness for use explicit.
Which decision will use this field, what quality dimensions matter, what error is tolerable and what happens if the value is wrong? A dataset is not high quality in the abstract. It is adequate or inadequate for a defined purpose.
Practical application
Define critical data elements and quality requirements by use. Establish dictionaries, validation, provenance, duplicate control, review, correction and versioning. Monitor quality by source, geography, collector and time rather than only at aggregate level. Give users metadata on coverage, age, method and uncertainty. Investigate incentives behind recurring defects.
Set stricter thresholds and human review where data affect rights, payments, market access or public claims.
Why it matters
Sustainability programmes increasingly automate decisions and claims from distributed data. Poor quality can exclude people, misdirect resources and produce false assurance at scale. Fit-for-purpose data are a governance requirement, not a technical preference.
Common misconception
Data quality is often equated with clean, complete or accurate records. A dataset can satisfy those tests and remain irrelevant, outdated, incoherent or unsuitable for the decision. Quality must be assessed across dimensions and intended use.
Connections
Data Governance assigns responsibility for quality. Traceability Systems and Interoperability depend on stable identifiers and meaning. Sampling affects coverage and inference, while Uncertainty should communicate the limits of estimates and classification.
A question worth asking
Which decision in your sustainability programme would cause the greatest harm if the underlying data were complete, precise and wrong?
Selected references
United Nations Statistics Division. 2018. UN Statistics Quality Assurance Framework. United Nations Statistics Division. 2026. Handbook of Surveys on Households and Individuals: Foundations and Emerging Approaches. Wang, R. Y. and Strong, D. M. 1996. Beyond Accuracy: What Data Quality Means to Data Consumers. Journal of Management Information Systems 12(4): 5-33. ISO 8000-2:2022. Data Quality - Part 2: Vocabulary.
European Statistical System. 2019. Quality Assurance Framework of the European Statistical System.
Review
Public comments appear only after editor acceptance. Draft comments stay in the review queue.
Reviewers choose the definition or an overview paragraph, leave a comment or replacement, and attach evidence or a source link.
Editors compare reviewer cards side by side. AI may help find agreement, conflicts, unsupported claims and possible source issues.
Only an editor-accepted synthesis changes the public page. Reviewer identities are shown only with consent and verification.
Submitted reviews stay private until accepted.
Loading verified endorsements… Endorsements are not votes and never determine publication.
Endorse this definition
Endorse the exact version shown here. This is not a vote, and publication remains an editorial decision.
Review board
Comment on a specific line. Each reviewer stays separate until an editor accepts a merged draft.
Each person comments on the definition or overview in their own draft card, with role, evidence and suggested wording kept together.
AI can compare comments against the current text, flag conflicting claims, surface missing evidence and identify where reviewers agree.
An editor merges compatible suggestions into a draft change, checks sources, records disagreements and decides what can be published.