Sustainability & AI
Training data
The data a model learns from; its quality, representativeness and provenance shape model behaviour and bias.
Definition
The dataset from which a model learns its behaviour — its composition, provenance and quality fundamentally determining what the model can and cannot do.
Quick reference
At a glance
- Subject
- Sustainability & AI
- Editorial status
- Editorial draft
- Definition status
- Established
- Last updated
- 21 August 2026
References
This source provides part of the technical or institutional basis for the definition.
This source supports the explanation of how the term is applied, measured or governed in practice.
Overview
What it means
A model encodes the statistical patterns of its training data, including its gaps and biases. Documentation practices such as 'Datasheets for Datasets' (Gebru et al. 2021) exist because opaque data makes model behaviour unauditable; the EU AI Act now requires data-governance documentation for regulated systems.
How it is used
In sustainability applications, training-data questions are concrete: which forests were photographed, which reports were parsed, whose languages and regions are represented. Scarce environmental data drives interest in augmentation and synthetic data.
Why it matters
Every model is a crystallisation of its data. Assessing an AI system for sustainability use without asking about its training data is like reviewing a study without asking about its sample.
Review
Help keep this definition useful and accurate.
Submitted reviews stay private until an editor decides whether to accept and attribute them.
Endorse this definition
Confirm what works
Endorse the current wording when it is accurate and useful. Suggested changes use the separate editorial form below.