Sustainability & AI
Reinforcement learning
A machine-learning approach where an agent learns by trial and error, used in energy-system optimisation and grid control.
Definition
A machine-learning paradigm in which an agent learns what actions to take by trial and error, receiving rewards or penalties rather than labelled examples.
Quick reference
At a glance
- Subject
- Sustainability & AI
- Editorial status
- Editorial draft
- Definition status
- Established
- Last updated
- 21 August 2026
References
This source provides part of the technical or institutional basis for the definition.
This source supports the explanation of how the term is applied, measured or governed in practice.
Overview
What it means
Instead of learning from a fixed dataset, the agent interacts with an environment and optimises long-term reward. Sutton and Barto's textbook defines the field. Because training requires many interaction cycles, agents are usually trained in simulation before deployment.
How it is used
Energy is the flagship sustainability application: reinforcement-learning controllers have been used for data-centre cooling optimisation — Google reported cutting cooling energy substantially after handing control to such a system — and for grid management, battery dispatch and building control.
Why it matters
Wherever a physical system must be continuously steered toward efficiency — a grid, a chiller plant, a battery — reinforcement learning is a candidate tool. The same autonomy raises verification and safety requirements before real-world control is granted.
Review
Help keep this definition useful and accurate.
Submitted reviews stay private until an editor decides whether to accept and attribute them.
Endorse this definition
Confirm what works
Endorse the current wording when it is accurate and useful. Suggested changes use the separate editorial form below.