Cognitive Load Theory and Worked Examples: What the Evidence Actually Supports

A look at one of instructional design's best-replicated findings, where it reverses, and why researchers still argue about how to measure it.

Mustafa Pat · 14 September 2026 · 8 min read · 2 views

Ask a group of instructors why they show students a fully worked-out example before asking them to solve a similar problem alone, and most will say it seems obviously helpful. Ask why, and the explanations get vaguer. Cognitive load theory is the closest thing this intuition has to a formal, tested account — and, per Sweller, van Merriënboer and Paas's 2019 retrospective in Educational Psychology Review, it is one of the more heavily replicated bodies of work in educational psychology, with a genuine track record of failed predictions and revisions along the way. That mixture of solid ground and open argument is what makes it worth examining carefully rather than citing as settled fact.

The theory starts from a narrow claim about human cognitive architecture: working memory, the system that holds and manipulates information during active thought, can handle only a small number of novel elements at once, while long-term memory, organized into structures called schemas, has no such practical limit. John Sweller's 1988 paper in Cognitive Science proposed that when a learner is forced to use most of their working memory capacity on the mechanics of solving a problem — trial and error, means-ends search — little capacity is left over to notice and encode the underlying structure of the problem type. The practical implication he drew was counterintuitive at the time: for a novice, solving problems from scratch can be a worse way to learn the underlying method than studying someone else's fully worked solution.

The best-replicated finding: the worked-example effect

That specific claim — novices learn more from studying worked examples than from solving equivalent problems unaided — is usually called the worked-example effect, and it has held up under scrutiny better than most single findings in the field. A 2023 meta-analysis in Educational Psychology Review by Barbieri, Miller-Cotto, Clerjuste and Chawla pooled 55 studies and 181 effect sizes on worked examples in mathematics specifically, using robust variance estimation to account for the fact that many studies contributed several effect sizes at once. They report a medium overall effect, g = 0.48, favoring worked-example study over unsupported problem solving for learning outcomes. That is a real, if moderate, effect — not the kind of number that should be read as "worked examples double learning," but a consistent advantage across a genuinely varied set of studies, age groups, and mathematical topics.

The mechanism Sweller proposed for extraneous load — cognitive effort spent on things that do not build the target skill — also produced a second, independently tested prediction: the split-attention effect. Chandler and Sweller's 1991 study in Cognition and Instruction found that when instructional text and a diagram are physically separated on a page, forcing learners to hold one in mind while consulting the other, learning suffers compared with an integrated format where labels sit directly on the diagram. Across six experiments, the integrated format consistently outperformed the split one whenever the two sources of information were unintelligible on their own and had to be mentally combined. This is one of the theory's more directly actionable findings for anyone producing instructional materials: a diagram with a separate legend, or a worked derivation next to an unlabeled figure, imposes cognitive cost that has nothing to do with the content being taught.

Where the effect reverses

Here is where the theory gets more interesting than a simple "add worked examples" rule would suggest. Kalyuga, Ayres, Chandler and Sweller's 2003 review in Educational Psychologist documents what they call the expertise reversal effect: instructional techniques that help novices — worked examples, integrated diagrams, explicit step-by-step guidance — become neutral or actively harmful once learners have enough domain knowledge to generate their own solution steps. For a more advanced learner, working through a fully explained example can mean re-reading information they already hold in long-term memory, which is redundant rather than helpful, and redundant material has its own documented cost. The practical consequence is that "worked examples are good for learning" is true only as a claim about novices; the same material handed to someone past that stage can slow them down. This is not a minor footnote — it means any blanket recommendation to add more scaffolding, worked steps, or explanatory detail needs a qualifier about who the material is for and where they are in the sequence.

The contested edge: how much guidance, and for what

Kirschner, Sweller and Clark's 2006 paper in Educational Psychologist extended the underlying cognitive-load logic into a much broader and more contested argument: that minimally guided approaches to instruction — pure discovery learning, unstructured problem-based learning, unguided inquiry — are less effective than approaches with strong explicit guidance, because novices lack the schemas needed to direct their own exploration productively. This is one of the most cited papers in the field, and also one of the most disputed. Hmelo-Silver, Duncan and Chinn's 2007 response in the same journal argued that Kirschner et al. had collapsed a wide range of instructional approaches — including problem-based and inquiry learning as actually practiced, which are typically heavily scaffolded — into a single category of "unguided discovery" that few serious practitioners of those methods would recognize. Their point was not that cognitive load theory is wrong, but that its application to a sweeping verdict on entire pedagogical traditions overstated what the underlying evidence, mostly drawn from tightly controlled laboratory tasks, could support. This exchange is worth naming explicitly because it illustrates a real distinction: the narrow, mechanism-level claims (worked examples beat unaided problem solving for novices; split attention hurts learning; guidance helps more when prior knowledge is low) rest on a large base of converging experiments, while the broad policy-level claim (guided instruction is superior to problem-based and inquiry learning generally) remains an argument, not a settled result.

Can cognitive load even be measured reliably?

A theory built on the idea of limited mental capacity needs some way to measure that capacity being used, and this is the part of cognitive load theory that has drawn the sharpest internal criticism. Ton de Jong's 2010 paper in Instructional Science raised a problem that the field has not fully resolved: most studies infer cognitive load from single-item subjective ratings collected after a task ("how much mental effort did that require?"), a method that cannot distinguish between the theory's three proposed load types — intrinsic load from the material's inherent complexity, extraneous load from poor presentation, and germane load from the effort of actually building a schema. De Jong also questioned whether these three types can be meaningfully summed at all, given that intrinsic load is arguably a property of the task while the other two depend on the learner and the instruction. A more recent effort to take stock of the measurement problem, Krieglstein, Beege, Rey, Ginns, Krell and Schneider's 2022 meta-analysis in Educational Psychology Review, examined the reliability and validity of the most commonly used subjective cognitive load questionnaires and found reasonable support for a three-factor structure overall, but also flagged that instruments vary in how well they hold up across settings and that germane load in particular remains the hardest of the three to measure cleanly. In practice, this means that when a study reports "cognitive load was reduced," a careful reader should check which instrument was used and whether it distinguishes load type — the finding is weaker evidence than a controlled learning-outcome measure, and considerably weaker than a replicated effect size like the worked-example meta-analysis above.

What doesn't hold up, and what genuinely does

It is worth being explicit about the three tiers of evidence involved here, since instructional advice online routinely blurs them. The worked-example effect and the split-attention effect are replicated across many independent labs, multiple decades, and (for worked examples specifically) a formal meta-analysis with a stated effect size — this is the closest the theory gets to settled ground. The expertise reversal effect is well-supported by a substantial and consistent experimental literature, though it is a narrower, more recent finding with fewer independent replications than the original worked-example studies. The broad claim that guided instruction beats problem-based and inquiry learning as pedagogical philosophies is a single research team's interpretation, contested in the same journal by researchers who read the underlying evidence differently — it belongs in the "actively debated" category, not the "established finding" one. This kind of layering matters generally in education research: Makel and Plucker's 2014 analysis in Educational Researcher found that direct replications made up only a small fraction of published education research, and that replication success dropped sharply when the replicating team had no overlap with the original authors — a reminder that a single influential paper, however well cited, is not the same evidentiary weight as an independently reproduced effect.

What this means in practice

For someone designing instructional material rather than running experiments, the defensible takeaways are narrower than the theory's popularity suggests. Presenting new material to genuine novices as a worked example, rather than as an unsupported problem, has a real and replicated advantage on learning outcomes in the domains studied, mathematics chief among them. Diagrams and their explanatory text should be physically integrated rather than split across a page or screen when neither makes sense without the other. That same worked-example scaffolding should be faded out as learners demonstrate they can generate steps themselves, since the expertise reversal effect predicts it will stop helping and may start to hurt. And claims about instructional methods that rest only on self-reported "mental effort" scores, or that generalize a narrow laboratory finding into a verdict on an entire teaching philosophy, deserve more skepticism than the confidence with which they are often repeated.


References
  1. John Sweller. Cognitive Load During Problem Solving: Effects on Learning. Cognitive Science, 1988. doi:10.1207/s15516709cog1202_4
  2. Paul Chandler, John Sweller. Cognitive Load Theory and the Format of Instruction. Cognition and Instruction, 1991. doi:10.1207/s1532690xci0804_2
  3. Paul A. Kirschner, John Sweller, Richard E. Clark. Why Minimal Guidance During Instruction Does Not Work: An Analysis of the Failure of Constructivist, Discovery, Problem-Based, Experiential, and Inquiry-Based Teaching. Educational Psychologist, 2006. doi:10.1207/s15326985ep4102_1
  4. Cindy E. Hmelo-Silver, Ravit Golan Duncan, Clark A. Chinn. Scaffolding and Achievement in Problem-Based and Inquiry Learning: A Response to Kirschner, Sweller, and Clark (2006). Educational Psychologist, 2007. doi:10.1080/00461520701263368
  5. Slava Kalyuga, Paul Ayres, Paul Chandler, John Sweller. The Expertise Reversal Effect. Educational Psychologist, 2003. doi:10.1207/S15326985EP3801_4
  6. John Sweller, Jeroen J. G. van Merriënboer, Fred Paas. Cognitive Architecture and Instructional Design: 20 Years Later. Educational Psychology Review, 2019. doi:10.1007/s10648-019-09465-5
  7. Felix Krieglstein, Maik Beege, Günter Daniel Rey, Paul Ginns, Moritz Krell, Sascha Schneider. A Systematic Meta-analysis of the Reliability and Validity of Subjective Cognitive Load Questionnaires in Experimental Multimedia Learning Research. Educational Psychology Review, 2022. doi:10.1007/s10648-022-09683-4
  8. Christina A. Barbieri, Dana Miller-Cotto, Sabrina N. Clerjuste, Kanika Chawla. A Meta-analysis of the Worked Examples Effect on Mathematics Performance. Educational Psychology Review, 2023. doi:10.1007/s10648-023-09745-1
  9. Ton de Jong. Cognitive Load Theory, Educational Research, and Instructional Design: Some Food for Thought. Instructional Science, 2010. doi:10.1007/s11251-009-9110-0
  10. Matthew C. Makel, Jonathan A. Plucker. Facts Are More Important Than Novelty: Replication in the Education Sciences. Educational Researcher, 2014. doi:10.3102/0013189X14545513

Cite this

Mustafa Pat. “Cognitive Load Theory and Worked Examples: What the Evidence Actually Supports.” EduFabTech, 14 September 2026. https://edufabtech.com/blog/cognitive-load-theory-worked-examples-instructional-design