Sustainability Language
Evaluation
A systematic and objective assessment of an intervention, policy, programme or strategy to understand its design, implementation, results, value, explanation and lessons for decision-making.
Expert review openNo editor-accepted expert review yetDefinition
A systematic and objective assessment of an intervention, policy, programme or strategy to understand its design, implementation, results, value, explanation and lessons for decision-making.
Overview
“Evaluation is not the final report on whether a programme succeeded; it is an inquiry into what happened, why and for whom. ”
Evaluation is often scheduled at the end of a programme, after budgets are spent and major decisions are irreversible. Consultants collect data, rate performance and produce recommendations. The exercise can satisfy accountability while arriving too late to improve the work.
The OECD defines evaluation as the systematic and objective assessment of a planned, ongoing or completed intervention, its design, implementation and results. Evaluation examines relevance, coherence, effectiveness, efficiency, impact and sustainability according to the questions and context. These criteria are not a checklist to score mechanically.
Relevance asks whether the intervention responds to priorities and needs. Coherence examines how it fits with other actions. Effectiveness considers achievement of objectives and results. Efficiency examines how resources produce results.
Impact addresses significant higher-level effects. Sustainability considers whether benefits are likely to continue. A programme can perform differently across criteria. It may deliver efficiently but address the wrong problem. It may reach its target while excluding the people with greatest need. It may create positive outcomes that do not last after funding ends. A single success rating hides these distinctions.
Evaluation differs from monitoring. Monitoring tracks indicators and implementation over time. Evaluation asks interpretive and causal questions using monitoring data alongside interviews, comparison, observation, documents and contextual evidence. Monitoring may show that yields rose; evaluation examines the programme's contribution, distribution, cost and unintended effects. It also differs from audit.
Audit assesses evidence against defined criteria.
Evaluation can examine whether the criteria, intervention logic and objectives were appropriate. A programme can comply with its plan and still be ineffective; evaluation has permission to question the plan. Timing and purpose shape method. Formative evaluation supports design and improvement. Process evaluation examines implementation. Outcome and impact evaluation assess change.
Developmental evaluation supports innovation in complex environments where the intervention itself evolves. Ex post evaluation examines durability after completion. Independence matters, but distance alone does not create quality. Evaluators need access, competence and freedom to reach unfavourable findings. They also need understanding of context and meaningful participation by affected people.
A detached method can be independent and still miss local reality. Use should be designed from the start. Who will make which decision? What evidence will be credible to them and to affected stakeholders? When must findings arrive? An elegant study that answers no live decision can become ceremonial learning. Power influences evaluation. Funders often define success, commission the study and control publication.
Communities supply data without shaping questions or interpreting findings. Participatory approaches can redistribute some authority, but participation should be substantive rather than a consultation annex. Negative and null findings are valuable. Pressure to demonstrate impact can encourage selective methods, optimistic interpretation or unpublished failure.
A credible learning system protects evaluators and teams when evidence challenges the programme's theory. Timing affects what evaluation can see. Assessing a livelihood programme immediately after training may capture knowledge but miss adoption, income and resilience. Waiting too long can weaken records and make alternative explanations harder to test.
The evaluation horizon should follow the causal pathway rather than the funding calendar. Where outcomes mature at different speeds, staged evaluation can examine implementation, early change and durable effects separately. Use should be planned before the study begins. Who will decide whether to continue, redesign or stop the intervention? What evidence threshold would change that decision?
An evaluation commissioned only to validate a programme after renewal has already been approved has limited independence.
Evaluation becomes a management discipline when uncomfortable findings have a route into budgets, contracts and programme design. The discipline is to begin with a question worth deciding, not a method worth showcasing. Evaluation should make uncertainty visible, test alternative explanations, examine distribution and state what the evidence cannot support.
Its value is not the rating at the end; it is the quality of judgement it enables.
Practical application
Define purpose, users, decisions and questions before selecting methods. Use an evaluation design proportionate to the claim and context, combining quantitative and qualitative evidence and examining unintended and differential effects. Protect independence and access, involve affected stakeholders safely and agree publication and management-response arrangements in advance.
Track whether recommendations are acted upon and whether later evidence confirms the interpretation. Evaluation terms of reference should state the users, decisions, questions, scope, independence safeguards, methods, limitations, publication expectations and process for management response. Require a response that accepts, rejects or qualifies each recommendation with reasons and deadlines.
Track implementation afterwards; otherwise the evaluation may be rigorous and still have no consequence.
Publish the core findings and limitations wherever confidentiality permits. Selective circulation protects programmes from scrutiny and prevents others from learning from evidence already paid for.
Why it matters
Evaluation tests whether sustainability activity is relevant, effective and durable rather than merely complete. It creates structured challenge to programme assumptions and supports learning beyond individual indicators.
Common misconception
Evaluation is often treated as an endline measurement or external score. It is a broader inquiry into design, implementation, results, causality, value and learning throughout the life of an intervention.
Connections
Monitoring supplies ongoing evidence. Attribution and Contribution address causal explanation. Counterfactual supports comparison, while Impact asks about significant higher-level effects. Continuous Improvement turns evaluation findings into change.
A question worth asking
What decision will change because of this evaluation, and will the evidence arrive before that decision has already been made?
Selected references
OECD. 2023. Glossary of Key Terms in Evaluation and Results-Based Management for Sustainable Development, Second Edition. OECD. 2019. Better Criteria for Better Evaluation: Revised Evaluation Criteria Definitions and Principles for Use. OECD. 2022. Recommendation of the Council on Public Policy Evaluation. Patton, M. Q. 2011. Developmental Evaluation. Scriven, M. 1991. Evaluation Thesaurus, Fourth Edition.
Review
Public comments appear only after editor acceptance. Draft comments stay in the review queue.
Reviewers choose the definition or an overview paragraph, leave a comment or replacement, and attach evidence or a source link.
Editors compare reviewer cards side by side. AI may help find agreement, conflicts, unsupported claims and possible source issues.
Only an editor-accepted synthesis changes the public page. Reviewer identities are shown only with consent and verification.
Submitted reviews stay private until accepted.
Loading verified endorsements… Endorsements are not votes and never determine publication.
Endorse this definition
Endorse the exact version shown here. This is not a vote, and publication remains an editorial decision.
Review board
Comment on a specific line. Each reviewer stays separate until an editor accepts a merged draft.
Each person comments on the definition or overview in their own draft card, with role, evidence and suggested wording kept together.
AI can compare comments against the current text, flag conflicting claims, surface missing evidence and identify where reviewers agree.
An editor merges compatible suggestions into a draft change, checks sources, records disagreements and decides what can be published.