Sustainability Language
Metric
A defined method of quantification or calculation, including its unit, formula, boundary and data rules, used to express an aspect of performance or condition.
Expert review openNo editor-accepted expert review yetDefinition
A defined method of quantification or calculation, including its unit, formula, boundary and data rules, used to express an aspect of performance or condition.
Overview
“A metric gives a number its grammar: what is counted, how it is combined and what the result can mean. ”
Metric and indicator are often used interchangeably. In practice, the distinction is useful. A metric is the measurement rule: tonnes of carbon dioxide equivalent per tonne of product, percentage of workers receiving at least a living wage or median days to resolve a grievance. An indicator uses one or more metrics to signal progress or condition against an objective.
The difference matters because two organisations can report the same indicator label using different metrics. Both may disclose water use intensity, but one divides withdrawals by product mass and another divides consumption by revenue. The numbers are not comparable even though the title is identical.
A credible metric specifies the construct, unit, formula, population, boundary, period, exclusions, source and treatment of missing values.
Without those elements, a number can travel through dashboards and reports while changing meaning at each stage. Denominators deserve particular attention. An injury rate per 200,000 hours and the number of injuries tell different stories. An emissions-intensity metric can fall while absolute emissions rise as production expands.
A percentage of traceable volume can improve because untraceable suppliers were removed from the denominator rather than brought into the system. Aggregation can hide distribution. Average farmer income may rise because a small group improved substantially while most households did not. Median, percentile and subgroup metrics can reveal a different pattern.
The choice is not a technical detail; it determines whose experience becomes visible. Composite metrics create further risk.
A sustainability score may combine soil health, income, biodiversity and governance into one value. Weighting allows comparison, but it also embeds normative choices. A strong score in one dimension can compensate mathematically for severe failure in another. Users should be able to see the components and understand why weights were chosen. Normalisation can support comparison across size or output.
It can also detach performance from absolute limits. Water use per kilogram may improve while withdrawals exceed watershed capacity. Carbon intensity can decline while cumulative emissions continue to increase. Relative and absolute metrics should often be reported together. Precision is not the same as validity.
A metric calculated to three decimal places can rest on estimated farm areas, imputed yields or uncertain emission factors.
Reported precision should reflect the quality of the underlying evidence rather than the capability of the spreadsheet.
Metrics can change behaviour. When payment depends on a score, participants learn the formula. They may improve the underlying condition, optimise the measured components or manipulate the data. Robust metric design anticipates gaming, checks for unintended incentives and uses independent evidence where stakes are high. Comparability requires stability, but stability should not preserve a flawed method.
When definitions or science improve, organisations may need to revise metrics and restate history. Transparent versioning is more credible than silent continuity. The GHG Protocol demonstrates why detailed accounting rules matter. Scope, organisational boundary, emission factors, base-year recalculation and treatment of market instruments can materially change reported emissions.
The metric is not merely tonnes; it is the governed method that produces them.
Denominators deserve particular scrutiny. Emissions per tonne can improve while total emissions rise because production expands. Income per hectare can rise while household income falls after land loss. Audit findings per supplier can fall because higher-risk suppliers left the programme rather than because conditions improved.
A metric can be mathematically correct and still support the wrong interpretation when its denominator changes the story. Aggregation creates similar risk. Combining locations, commodities or groups can conceal extremes and transfer. A global average water-intensity metric may improve even as abstraction increases in a water-stressed basin.
Decision-useful metrics should preserve the level at which consequences occur and disclose both absolute and intensity measures where each answers a different question.
The formula should follow the decision, not the convenience of available data. The discipline is to make the calculation inspectable. A user should be able to reconstruct the number, identify assumptions and understand what the metric excludes. A metric is valuable not because it simplifies reality, but because it simplifies without disguising the choices made.
Practical application
Create a metric specification covering definition, formula, unit, population, boundary, period, source, quality rules, uncertainty and owner. Test alternative denominators and aggregation methods before choosing the one that best supports the decision.
Report absolute and relative values where each reveals a different risk. Keep composite components visible, version methodology changes and review gaming incentives when metrics affect reward, access or public claims. Create a metric dictionary that records the formula, units, conversion factors, boundary, treatment of missing values, aggregation method and version history.
Test calculations with worked examples and reconcile them to source systems. Where a methodology changes, quantify the effect separately from real performance so users can distinguish a better method from a better result.
Why it matters
Metrics determine how performance becomes comparable, manageable and claimable. Weak definitions can turn methodological differences into apparent progress and allow organisations to improve a score without improving the underlying condition.
Common misconception
A metric is often treated as a neutral number. Every metric contains choices about boundaries, units, aggregation and value. Those choices should be explicit and defensible.
Connections
Indicator uses metrics to signal progress. Baseline and Target give the metric a reference and desired direction. Data Quality and Uncertainty determine how much confidence the number can support, while Benchmark compares it with an external or internal standard.
A question worth asking
Which choice in your metric's denominator, boundary or weighting most strongly shapes the story the number tells?
Selected references
OECD. 2023. Glossary of Key Terms in Evaluation and Results-Based Management for Sustainable Development, Second Edition. United Nations Statistics Division. 2019. United Nations National Quality Assurance Frameworks Manual for Official Statistics.
Greenhouse Gas Protocol. 2004. A Corporate Accounting and Reporting Standard, Revised Edition. Hand, D. J. 2004. Measurement Theory and Practice: The World Through Quantification. Stevens, S. S. 1946. On the Theory of Scales of Measurement. Science 103(2684): 677-680.
Review
Public comments appear only after editor acceptance. Draft comments stay in the review queue.
Reviewers choose the definition or an overview paragraph, leave a comment or replacement, and attach evidence or a source link.
Editors compare reviewer cards side by side. AI may help find agreement, conflicts, unsupported claims and possible source issues.
Only an editor-accepted synthesis changes the public page. Reviewer identities are shown only with consent and verification.
Submitted reviews stay private until accepted.
Loading verified endorsements… Endorsements are not votes and never determine publication.
Endorse this definition
Endorse the exact version shown here. This is not a vote, and publication remains an editorial decision.
Review board
Comment on a specific line. Each reviewer stays separate until an editor accepts a merged draft.
Each person comments on the definition or overview in their own draft card, with role, evidence and suggested wording kept together.
AI can compare comments against the current text, flag conflicting claims, surface missing evidence and identify where reviewers agree.
An editor merges compatible suggestions into a draft change, checks sources, records disagreements and decides what can be published.