Sustainability & AI
Model distillation
Compressing a large model into a smaller one that keeps much of its performance while cutting compute and energy demand.
Definition
Training a compact 'student' model to reproduce the behaviour of a larger 'teacher' model, preserving much of its performance with far fewer parameters.
Quick reference
At a glance
- Subject
- Sustainability & AI
- Editorial status
- Editorial draft
- Definition status
- Established
- Last updated
- 21 August 2026
References
This source provides part of the technical or institutional basis for the definition.
This source supports the explanation of how the term is applied, measured or governed in practice.
Overview
What it means
Named by Hinton and colleagues in 2015, distillation transfers the soft, probabilistic outputs of a large model into a small one. It is now a standard efficiency technique, alongside quantisation and pruning, for deploying capable models on limited hardware.
How it is used
Distilled models run on modest infrastructure or at the edge — enabling on-device environmental sensors, cost-effective document processing, and lower-energy inference at scale.
Why it matters
Distillation is a practical lever of frugal AI: it decouples the capability of large models from their operating footprint, and its existence undercuts the assumption that useful AI must be maximal AI.
Review
Help keep this definition useful and accurate.
Submitted reviews stay private until an editor decides whether to accept and attribute them.
Endorse this definition
Confirm what works
Endorse the current wording when it is accurate and useful. Suggested changes use the separate editorial form below.