Sustainability & AI
Synthetic data
Artificially generated data used to train AI models where real data are scarce, as in rare-species detection or extreme-climate events.
Definition
Artificially generated data that mimics the statistical properties of real data, used to train or test AI models where real observations are scarce, expensive or sensitive.
Quick reference
At a glance
- Subject
- Sustainability & AI
- Editorial status
- Editorial draft
- Definition status
- Established
- Last updated
- 21 August 2026
References
This source provides part of the technical or institutional basis for the definition.
This source supports the explanation of how the term is applied, measured or governed in practice.
Overview
What it means
Synthetic data can come from simulators, statistical models or generative AI. It fills gaps — rare species sightings, extreme weather events, equipment failures — but its fidelity to reality must be validated, since models trained on flawed synthetic data inherit those flaws.
How it is used
Applications include augmenting rare-event training sets for ecological and climate models, stress-testing financial and grid models, and sharing privacy-preserving versions of personal energy-use data — a use case covered in UK ICO guidance on privacy-enhancing technologies.
Why it matters
Many sustainability problems are data-poor exactly where stakes are high. Synthetic data can unblock model development, but treating generated data as equivalent to observation is an analytical and governance risk that must be managed explicitly.
Review
Help keep this definition useful and accurate.
Submitted reviews stay private until an editor decides whether to accept and attribute them.
Endorse this definition
Confirm what works
Endorse the current wording when it is accurate and useful. Suggested changes use the separate editorial form below.