Evaluation, Synthesis and Judgement
What evaluation does
Evaluation tests how strong, relevant or reliable an economic conclusion is in the stated context. It asks whether a mechanism operates, how large the result is, who is affected, what trade-off arises, and whether another option ranks higher.
Evaluation is not:
- adding “however” to an unrelated sentence;
- listing one advantage and one disadvantage;
- naming a condition without its direction of effect;
- declaring that every policy has government failure;
- refusing to decide because outcomes are uncertain.
Five distinct operations
| Operation | Purpose | Example opening |
|---|---|---|
| qualify | state a condition that changes magnitude, direction or confidence | “This effect is larger when…” |
| counter | develop a competing mechanism or relevant cost | “However, the same policy may…” |
| compare | judge alternatives against a common criterion | “Relative to policy B, policy A…” |
| synthesise | combine arguments through sequence, complementarity or contingency | “In the short run…, while in the long run…” |
| judge | rank the arguments and answer the exact question | “Therefore, policy A is preferable when…” |
A strong answer may use all five, but there is no compulsory quota.
From evaluation inputs to judgement
Caption: Start with the question’s decision criterion in the central comparison band. The surrounding cards supply possible tests—evidence, assumptions, time, stakeholders, magnitude, trade-offs, alternatives and feasibility. Arrows converge only after each selected test has altered the strength of an argument. The final ranked judgement is therefore a consequence of analysis, not an isolated concluding opinion.
Productive evaluation lenses
Assumptions and elasticities
Ask which assumption is carrying the conclusion.
A tax reduces consumption more when demand is price-elastic because consumers can switch to substitutes. With inelastic demand, quantity falls little and tax revenue may be larger, though the burden on consumers may also be greater.
The elasticity changes the magnitude, revenue outcome and incidence.
Short run and long run
Time matters only when the mechanism changes:
- supply and demand may become more elastic;
- wages and contracts may adjust;
- productive capacity may expand;
- expectations may change;
- implementation and behavioural responses may take time.
Do not write “in the long run it may be different” without explaining why.
Stakeholders and distribution
Aggregate net benefit can rise while some groups lose. Identify the relevant groups and mechanism:
- consumers may pay a higher price;
- workers may face displacement or gain new opportunities;
- producers may bear compliance cost;
- taxpayers finance subsidies;
- future generations may receive benefits or liabilities;
- trading partners may respond.
Then decide whether the distribution is relevant to the question’s criterion and whether compensation is feasible.
Magnitude, likelihood and context
Direction is not enough. Consider:
- scale and persistence of the shock;
- spare capacity;
- market structure;
- openness and import dependence;
- consumer or firm responsiveness;
- confidence and expectations;
- administrative capacity;
- probability rather than mere possibility.
A logically possible outcome should not outweigh a highly likely and large effect without evidence.
Trade-offs and opportunity cost
Ask what is forgone:
- another use of government revenue;
- private consumption or investment;
- environmental quality;
- policy credibility;
- administrative attention;
- progress toward another macroeconomic objective.
Opportunity cost is the value or net benefit of the next-best feasible alternative forgone, not every imaginable alternative added together.
Alternatives and government failure
Compare the proposed policy with a feasible counterfactual:
- no intervention;
- a differently targeted instrument;
- regulation, information provision or market-based incentives;
- a sequence or complementary policy.
Government failure is not proof that markets perform better. Compare the likely imperfection of intervention with the likely imperfection of the realistic alternative.
The reversal test
Ask:
If this condition changed, would my predicted magnitude, ranking or recommendation change?
If not, the point may be background rather than evaluation.
Weak:
The policy depends on elasticity.
Strong:
A congestion charge reduces peak-hour traffic more when car demand is price-elastic because commuters can change route, time or transport mode. Where alternatives are poor, quantity changes little and the charge raises revenue more than it reduces congestion, strengthening the case for complementary public-transport investment.
Compare policies using the same criterion
| Criterion | Policy A | Policy B |
|---|---|---|
| targeting of root cause | ||
| speed and time lag | ||
| magnitude and reliability | ||
| effect on other objectives | ||
| distribution | ||
| fiscal and administrative cost | ||
| implementation risk | ||
| long-run capacity |
Do not compare policy A’s speed with policy B’s equity and then call one “better”. First decide which objectives matter most in the context.
A worked policy-ranking example
Question:
Assess whether a consumer subsidy is the best policy to improve access to childcare.
Argument for the subsidy
A subsidy lowers the effective price paid by households. Demand rises, and access may improve for families previously unable to afford childcare.
Material qualification
If short-run supply is inelastic because qualified workers and facilities are limited, providers cannot expand places quickly. Much of the subsidy may raise market price rather than increase quantity, so the fiscal cost produces limited additional access.
Alternative
Training grants and capacity support target the supply constraint but take time. Means-tested support may target affordability more precisely than a universal subsidy but creates eligibility and administrative costs.
Synthesis
Use targeted household support during the transition, while expanding training and capacity for the longer run. The policies are complements because they address different constraints over different horizons.
Judgement
A broad consumer subsidy is not the best stand-alone policy where provider capacity is the immediate binding constraint. Targeted affordability support combined with time-limited capacity expansion is more likely to increase access per dollar spent, although the balance should shift toward demand support once supply becomes more responsive.
The decision uses a common criterion—access per dollar—and identifies the decisive condition.
Synthesis is more than “use both”
Useful forms include:
- sequence: stabilise demand now, raise productive capacity later;
- targeting: use a broad incentive, then compensate a vulnerable group;
- complementarity: combine instruments that address separate causes;
- contingency: use policy A in recession but policy B near full employment;
- division of roles: one policy changes incentives while another improves information or enforcement.
Check compatibility. Two instruments may offset one another or impose conflicting signals. A policy mix must have explained roles and net transmission.
Intermediate and summative evaluation
Intermediate evaluation tests a mechanism where it appears:
The subsidy increases consumption only if providers can expand capacity; with inelastic short-run supply, much of the benefit may be captured as higher price.
Summative evaluation compares the developed mechanisms:
Because capacity is the immediate constraint, supply expansion is likely to improve access more than a broad consumer subsidy in the short run.
These are descriptive labels, not required examination headings.
Evaluating evidence
Evidence is rarely perfect. The task is to state how a limitation changes confidence.
| Evidence feature | What it can weaken |
|---|---|
| short time period | confidence that a pattern will persist |
| nominal values | claims about real output or purchasing power |
| totals rather than per-capita values | claims about average living standards |
| forecast | certainty about realised outcomes |
| two variables moving together | causal claim |
| one stakeholder quotation | representativeness |
| index without underlying level | claim about absolute magnitude |
| national average | claim about distribution across groups or regions |
Do not discard evidence automatically. A forecast may still be useful if its assumptions are explicit; a short series may still establish what happened over that period.
Decision criteria by question type
| Question asks about… | Useful criteria |
|---|---|
| effectiveness | size, reliability and speed of effect |
| best policy | targeting, trade-offs, feasibility and alternatives |
| impact | direction, magnitude, distribution and duration |
| extent | importance relative to other causes or constraints |
| desirability | net benefit, equity, opportunity cost and risk |
Choose criteria from the question and context. Do not force every criterion into every answer.
Construct a conditional but decisive judgement
A robust judgement contains:
- Decision: which conclusion or option is stronger?
- Criterion: stronger for which objective?
- Condition: under what circumstance?
- Perspective: for whom?
- Time: over what horizon?
- Reason: which mechanism dominates?
Template:
On balance, ___ is more likely to achieve ___ because ___. This conclusion is strongest when ___. For ___, the main cost is ___; therefore ___ should complement or replace the policy if ___.
Use the template to check completeness, not as a memorised sentence.
Prioritise evaluation
Rank points by:
- direct relevance to the command;
- likelihood in the stated context;
- expected magnitude;
- ability to reverse the conclusion;
- feasibility of addressing the limitation.
One decisive condition, fully explained, can be more valuable than five generic caveats.
Common failures
- “It depends” with no direction.
- Short-run/long-run labels without an adjustment mechanism.
- Stakeholder names without a welfare comparison.
- Government failure added to every policy.
- A policy mix claimed without role, sequence or compatibility.
- Arguments compared using different criteria.
- An improbable possibility treated as equally important as a likely effect.
- A major new argument introduced only in the conclusion.
- Evidence rejected merely because it has limitations.
- Uncertainty used to avoid a decision.
Final self-check
- Did the evaluation change magnitude, direction, confidence or ranking?
- Did competing arguments use the same criterion?
- Did I distinguish qualification, counterargument and synthesis?
- Did I prioritise rather than list?
- Did I acknowledge uncertainty without surrendering judgement?
- Does the final decision follow from the developed analysis?
Return to Case Study and Essay Skills.