Effect Sizes in Education Research: What Hattie's Hinge Point and the Formative-Assessment Numbers Actually Show
A widely repeated claim that formative assessment "doubles the speed of learning" traces back to small early studies; larger, later trials and statisticians who checked the method tell a more modest and more interesting story.
Ask a teacher, a school leader or an education researcher to name one number from the learning-science literature, and a good share will say "0.4." It is the "hinge point" from John Hattie's Visible Learning synthesis: the effect size above which, Hattie argued, an intervention is doing more good than an average year of schooling does on its own. A close second is the claim that formative assessment can "double the speed of student learning," a phrase that has circulated in teacher training for two decades. Both numbers get repeated as settled facts. Neither is as settled as the repetition suggests.
This matters because effect sizes are the currency researchers use to compare interventions that look nothing alike — a feedback routine, a tutoring programme, a curriculum redesign — on a single scale. If that currency is unstable, then rankings built on it, and the decisions built on the rankings, inherit the instability. Tracing one specific, well-documented case — the effect size claimed for formative assessment and feedback — through the original studies, the meta-analyses that followed, the statisticians who checked the arithmetic, and the large trials that tested the claim directly, shows what happens when a plausible number is checked rather than repeated.
The story starts with a paper that never actually reported a single effect size for its central claim.

Where the claim started
In 1998, Paul Black and Dylan Wiliam published "Inside the Black Box: Raising Standards Through Classroom Assessment" in Phi Delta Kappan, reviewing around 250 studies on classroom assessment practice. Their conclusion was that improving formative assessment — the ongoing, in-lesson feedback loop between what a student currently understands and what a teacher does next — was among the most effective levers available to raise achievement, and a more reliable one than most externally imposed reform programmes. The review did not reduce that conclusion to a single standardised effect size; it was a narrative synthesis of a mixed literature.
The number most people now associate with the claim comes from the same authors' fuller technical review, published the same year. In "Assessment and Classroom Learning" (Assessment in Education, 1998), Black and Wiliam reported effect sizes clustering between 0.4 and 0.7 across the best-controlled studies in their sample — a set that was mostly small-scale and short-duration, several run by the same researchers who had trained the participating teachers. In professional-development circles that range compressed into a single vivid claim, that formative assessment could "double the speed of student learning," and the phrase spread through teacher training for two decades, detached from the qualifications the underlying studies actually carried.
What happens when someone re-runs the numbers
In 2011, Neal Kingston and Brooke Nash published a dedicated meta-analysis in Educational Measurement: Issues and Practice, built specifically to test whether the 0.4–0.7 range held up. They started from more than 300 published studies claiming to bear on formative assessment. Only 13 of them reported enough statistical detail to compute a usable effect size, yielding 42 independent estimates. The median effect size across those estimates was 0.25; a random-effects weighted mean put it at 0.20 — well below the figure in general circulation, and the effect varied by subject, from about 0.32 in English language arts down to 0.09 in science.
A decade later, a large field trial went further than a meta-analysis of existing studies: it tested the claim directly, at scale, with random assignment. Anders and colleagues (2022), evaluating a formative-assessment training programme across 140 secondary schools in England for the Education Endowment Foundation, found a pre-registered effect on externally examined attainment of 0.09 standard deviations, rising to about 0.11 in sensitivity and complier analyses. That is not a rounding difference from "doubles the speed of learning" — it is roughly an eighth to a sixth of the originally cited figure, produced by a design (randomised, pre-registered, multi-site) that carries more evidential weight than the small, often self-selected studies behind the original number.
None of this means formative feedback does nothing. The Education Endowment Foundation's Teaching and Learning Toolkit, synthesising 155 studies as of its 2021 update, still lists feedback as one of the higher-impact strands it tracks, at an average of six months' additional progress. The pattern across all three sources — the original narrative review, the dedicated meta-analysis, and the large trial — is not "formative assessment doesn't work." It is that the size of the effect shrinks, often sharply, every time a more rigorous design is applied to the question, a pattern worth expecting by default rather than treating as a scandal each time it recurs.
A harder question underneath the number
Kingston and Nash's finding — that most claimed formative-assessment studies could not even supply a computable effect size — points at a bigger problem than any single number. It is the question of whether effect sizes computed from different outcome measures, different study designs and different comparison groups can be legitimately averaged or ranked against each other at all.
This is the question Hattie's Visible Learning method raises at scale. The synthesis pools well over 800 meta-analyses covering tens of thousands of underlying studies into a single ranked list, with the 0.40 hinge point marking the boundary between "teacher effects" and the "zone of desired effects." Two independent statistical critiques examined whether that pooling is defensible. Pierre-Jérôme Bergeron, writing in the McGill Journal of Education in 2017, argued from a statistician's standpoint that Cohen's d is not a portable, universal unit — an effect size computed from a short researcher-made quiz measuring days of learning is not commensurable with one computed from a standardised test measuring years of learning, even though both get reported as "d = 0.4" and averaged as if they were the same thing.
Separately, Adrian Simpson, in the Journal of Education Policy the same year, identified specific mechanical ways effect sizes get inflated or deflated independent of an intervention's real impact: unequal comparison groups, restricted score ranges in the sample, and — echoing Bergeron — the choice between a standardised assessment and a researcher-designed one, which tends to produce systematically larger effects because it is more closely aligned with what was actually taught. Simpson's broader point is that a single fixed threshold like 0.40, applied uniformly across outcome types, age groups and study designs, cannot do the job Hattie's framework asks of it.

A different yardstick
If 0.40 is not a reliable universal threshold, what should replace it? Matthew Kraft's 2020 paper in Educational Researcher proposes an answer grounded in what education interventions actually achieve, rather than in the general social-science conventions Jacob Cohen proposed roughly fifty years earlier for a different kind of research. Kraft's recalibrated benchmarks for causal education studies with standardised achievement outcomes classify an effect under 0.05 as small, 0.05–0.19 as medium, and 0.20 or larger as large.
Set the two systems side by side and the mismatch is immediate:
| Framework | Small | Medium | Large |
|---|---|---|---|
| Cohen's general convention | 0.20 | 0.50 | 0.80 |
| Kraft (2020), education RCTs | < 0.05 | 0.05–0.19 | ≥ 0.20 |
By Kraft's scale, the Anders et al. (2022) trial's 0.09–0.11 effect on formative assessment is a medium-to-large result for a low-cost, system-wide intervention — not the disappointing footnote it looks like against the "doubles the speed of learning" framing. And Hattie's 0.40 hinge, which sounds like a modest bar, is above Kraft's "large" threshold by a factor of two. Neither framework is wrong on its own terms; they answer different questions — general behavioural-science comparability versus realistic education-intervention impact — and confusing them is exactly the kind of category error Bergeron's and Simpson's critiques describe.
Sorting the replicated from the contested from the myth
Laid out this way, the formative-assessment case separates into three tiers that are useful to keep distinct when reading any similar claim in education research:
- Replicated, with a smaller effect than first reported: formative feedback and assessment practices produce a real, positive, but modest effect on measured achievement — roughly 0.1 to 0.3 standard deviations depending on subject and design — confirmed independently by Kingston and Nash's meta-analysis (2011) and by the large randomised trial of Anders et al. (2022).
- Single-study or early-adopter result, not replicated at scale: the 0.7 effect size behind "doubles the speed of learning" came from a small set of well-controlled early studies, largely with committed volunteer teachers — a profile known to produce larger effects than system-wide rollouts, and one that later, larger trials did not reproduce.
- Popular framing the underlying statistics do not support: a single fixed hinge point, used as a universal pass/fail threshold across different outcome measures, subjects and age groups. Bergeron (2017) and Simpson (2017) both show, from different statistical angles, that this specific use of effect size is not defensible as stated, independent of whether any particular intervention above or below the line actually works.
What this means when reading the next claimed effect
None of this is an argument against using effect sizes, or against synthesising evidence across studies — both remain necessary. It is an argument for reading the fine print before repeating the headline number. A few checks travel well beyond this one case: whether the effect size comes from a study designed to test the claim directly, such as a pre-registered randomised trial, or from a narrative review retrofitted with a number after the fact; whether the outcome measure was a standardised assessment external to the intervention or one built by the people running it; whether the reported figure is a single early estimate or one checked by an independent replication at larger scale; and whether the benchmark used to call an effect "large" or "small" was built for the kind of study in front of you, the way Kraft's benchmarks were built specifically for education RCTs rather than borrowed from a different field. A number that survives all four checks is worth taking seriously. Many of the most famous numbers in education research have not yet been asked to survive them.
- Paul Black, Dylan Wiliam. Assessment and Classroom Learning. Assessment in Education: Principles, Policy & Practice, 1998. doi:10.1080/0969595980050102
- Neal Kingston, Brooke Nash. Formative Assessment: A Meta-Analysis and a Call for Research. Educational Measurement: Issues and Practice, 2011. doi:10.1111/j.1745-3992.2011.00220.x
- Adrian Simpson. The Misdirection of Public Policy: Comparing and Combining Standardised Effect Sizes. Journal of Education Policy, 2017. doi:10.1080/02680939.2017.1280183
- Pierre-Jérôme Bergeron. How to Engage in Pseudoscience with Real Data: A Criticism of John Hattie's Arguments in Visible Learning from the Perspective of a Statistician. McGill Journal of Education, 2017. link
- Matthew A. Kraft. Interpreting Effect Sizes of Education Interventions. Educational Researcher, 2020. doi:10.3102/0013189X20912798
- Jake Anders, Francesca Foliano, Matthew Bursnall, Richard Dorsett, Nikki Hudson, Johnny Runge, Stefan Speckesser. The Effect of Embedding Formative Assessment on Pupil Attainment. Journal of Research on Educational Effectiveness, 2022. doi:10.1080/19345747.2021.2018746
- Education Endowment Foundation. Teaching and Learning Toolkit: Feedback. Education Endowment Foundation, 2021. link