How People Actually Learn: What the Evidence on Memory, Practice and Assessment Does and Doesn't Show

A look at which learning techniques are backed by replicated experiments, which rest on a single well-known study, and which popular claims the evidence contradicts outright.

Mustafa Pat · 13 September 2026 · 8 min read · 4 views

Ask a room of teachers what "good learning" looks like and you will get answers that sound right and often are not. Some claims about memory and instruction have survived decades of adversarial testing across labs, age groups and subject areas. Others rest on a single striking study that never replicated. And some — learning styles chief among them — have been tested directly and found wanting, yet remain in circulation because they are intuitive and flattering to teach. This piece tries to keep those three categories separate: what is replicated, what is a single study, and what is a popular claim the evidence does not support.

The testing effect: the most replicated finding in the field

The clearest result in learning science is that retrieving information from memory strengthens it more than re-reading or re-studying the same material. Henry Roediger and Jeffrey Karpicke's 2006 experiments had students either restudy a prose passage repeatedly or take practice tests on it, holding total study time constant. On a test given five minutes later, restudying looked better. On tests given two days or a week later, the group that had practiced retrieval outperformed the restudy group by a wide margin — a reversal the authors called counterintuitive because it means the condition that feels easier and produces better short-term performance is worse for durable learning (Roediger & Karpicke, 2006).

This is not a single-lab curiosity. A 2017 meta-analysis pooling 272 independent effect sizes from 118 articles found that practice testing produced a moderate-to-large advantage over restudying, and a still larger advantage over doing nothing at all, across age groups, subjects and test formats (Adesope, Trevisan & Sundararajan, 2017). The effect holds for short-answer and multiple-choice formats, for classroom quizzes and for low-stakes self-testing with flashcards, which is why "retrieval practice" rather than "testing" is now the preferred term — the mechanism is the retrieval itself, not the grading.

Spacing and interleaving: real effects, easy to misapply

The second well-replicated finding concerns timing. Distributing study sessions across days or weeks produces better long-term retention than massing the same amount of study into one sitting — the "spacing effect." A synthesis of 317 experiments spanning nearly a century of research confirmed the basic effect and showed that the optimal gap between study sessions depends on how long the material needs to be retained: for retention tested a year later, the best gap between initial sessions was measured in months, not days (Cepeda, Pashler, Vul, Wixted & Rohrer, 2006). There is no single "ideal" spacing interval that applies regardless of how long you need to remember something — a detail that popular summaries of spacing often drop.

Interleaving — mixing different problem types or topics within a single study session rather than practising one type in a block — is a related but separate technique, and researchers took care to disentangle it from spacing. Kelli Taylor and Doug Rohrer had students practice calculating the volume of different geometric solids either blocked by shape or interleaved, with total practice time and spacing held equal. Interleaved practice produced markedly worse performance during the practice session itself but roughly double the accuracy on a test given a day later, because mixed practice forces students to learn to discriminate which procedure a problem calls for, not just how to execute a procedure they already know is coming (Taylor & Rohrer, 2010). Both spacing and interleaving belong to what Robert Bjork termed "desirable difficulties" — conditions that slow acquisition but improve retention, which is precisely why they feel unproductive to learners while they are happening.

Cognitive load: a well-established constraint, not a technique

Working memory can hold only a small amount of new information at once. John Sweller's 1988 formulation of cognitive load theory argued that instructional materials fail when they impose "extraneous" load — effort spent on the format of the material rather than the material itself — and that instruction should be designed to keep working memory free for the actual learning task (Sweller, 1988). This has held up well as a design constraint: badly organised worked examples, split-attention formats (a diagram on one page, its explanation on another) and unnecessary decorative detail measurably impair learning compared with cleaner formats, even when the underlying content is identical.

Cognitive load theory also underpins one of the more contested debates in the field: how much guidance novice learners need. Paul Kirschner, John Sweller and Richard Clark reviewed decades of comparisons between minimally guided approaches — pure discovery learning, unstructured problem-based learning — and more explicit, guided instruction. Their conclusion was that minimal guidance imposes heavy cognitive load on novices who lack the prior knowledge to structure their own exploration, and that explicit guidance consistently outperforms discovery-based approaches for learners new to a topic, with the gap narrowing as learners gain expertise (Kirschner, Sweller & Clark, 2006). This paper is widely cited and has strong theoretical grounding, but it is worth naming as a position paper synthesising other people's data rather than a single new experiment — the "guidance versus discovery" debate it describes is still argued over, particularly around how "guidance" and "discovery" get defined in any given study.

Learning styles: tested directly, and found unsupported

The idea that students learn best when instruction is matched to their preferred modality — visual, auditory, kinaesthetic — is probably the most widely believed claim in this field that the evidence contradicts. Harold Pashler, Mark McDaniel, Doug Rohrer and Robert Bjork set out the exact experimental design needed to validate it: learners must be classified into style groups, randomly assigned to instructional methods that either match or mismatch their classified style, and then show that matched instruction produces better outcomes than mismatched instruction. Reviewing the literature against that standard, they found only one study offering even partial support and several that contradicted the hypothesis outright, concluding that "if classification of students' learning styles has practical utility, it remains to be demonstrated" (Pashler, McDaniel, Rohrer & Bjork, 2008). The theory is not merely unproven; it has been tested against a clear falsification criterion and has not met it. What is real is that people have preferences about how they like to receive information — that is different from a claim that matching instruction to those preferences improves learning.

Growth mindset: a genuine effect, but a much smaller one than the popular version claims

Carol Dweck's mindset theory — that believing intelligence is malleable rather than fixed affects motivation and achievement — is a useful case study in what happens between an original finding and its popular reception. Two large meta-analyses give a more measured picture than the popular version of the theory suggests. Victoria Sisk and colleagues pooled 273 studies (over 365,000 participants) on the mindset–achievement relationship and 43 intervention studies (over 57,000 participants), finding the correlation between mindset and achievement weak overall, and the effect of mindset interventions on achievement smaller still, though somewhat larger for academically at-risk students and those from lower-income backgrounds (Sisk et al., 2018). A subsequent large-scale randomised trial, the National Study of Learning Mindsets, delivered a short online mindset intervention to a nationally representative sample of ninth-grade students in the United States and found it raised grades among lower-achieving students, but only in schools where peer norms already supported a growth mindset — a context-dependence the original theory did not predict (Yeager et al., 2019). Read together, these are real but modest, conditional effects — a considerable distance from the popular claim that a brief mindset talk substantially raises achievement on its own.

Formative assessment: strong observational evidence, less certainty about mechanism

On the assessment side of learning, Paul Black and Dylan Wiliam's 1998 review of around 250 studies argued that formative assessment — using ongoing evidence of student understanding to adjust teaching in real time, rather than only grading work after the fact — was associated with substantial learning gains, particularly for lower-attaining students (Black & Wiliam, 1998). Their recommendation that feedback should describe the specific qualities of a piece of work and what to do next, rather than compare the student to peers or attach a grade that overshadows the comments, has held up in subsequent classroom research. It is worth being precise about the evidence base here: this was a synthesis of a very heterogeneous set of studies using different assessment practices, class sizes and subjects, which is different in kind from a set of tightly matched laboratory experiments like the retrieval-practice literature above. The direction of the finding is well supported; the exact size of the effect one should expect in a specific classroom is not something the review can pin down precisely.

What this means for study and teaching practice

Put the levels of evidence side by side and a practical hierarchy falls out:

  • Replicated across many independent experiments: retrieval practice/testing, distributed (spaced) practice, interleaving of related problem types, and the basic working-memory constraints described by cognitive load theory.
  • Directionally well supported but harder to quantify precisely: formative, descriptive feedback over comparative grading; guided instruction over minimally guided discovery for novices.
  • Real but smaller than the popular version claims: growth mindset interventions — worth doing in the right context, not a substitute for instruction or curriculum quality.
  • Tested and not supported: matching instruction to a student's preferred "learning style."

None of this requires exotic tools. Reworking a re-reading session into a closed-book self-quiz, spacing that quiz out over days rather than cramming it into one sitting, and mixing problem types instead of blocking them are changes any student or teacher can make without new materials, budget or technology. The harder part is behavioural, not informational: the techniques with the strongest evidence — retrieval and spacing — are also the ones that feel worse while they are happening, since interrupted, effortful recall is more effortful in the moment than smooth re-reading. That mismatch between what feels productive and what is productive is arguably the single most robust and most ignored finding in this literature.

"If classification of students' learning styles has practical utility, it remains to be demonstrated." — Pashler, McDaniel, Rohrer & Bjork, 2008

The broader lesson for anyone reading education research casually is to ask, for any claim, which of the four categories above it falls into — and to distrust particularly the ones that made it into a TED talk before they made it through a meta-analysis.


References
  1. Henry L. Roediger III, Jeffrey D. Karpicke. Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention. Psychological Science, 2006. doi:10.1111/j.1467-9280.2006.01693.x
  2. John Dunlosky, Katherine A. Rawson, Elizabeth J. Marsh, Mitchell J. Nathan, Daniel T. Willingham. Improving Students' Learning With Effective Learning Techniques: Promising Directions From Cognitive and Educational Psychology. Psychological Science in the Public Interest, 2013. doi:10.1177/1529100612453266
  3. Nicholas J. Cepeda, Harold Pashler, Edward Vul, John T. Wixted, Doug Rohrer. Distributed Practice in Verbal Recall Tasks: A Review and Quantitative Synthesis. Psychological Bulletin, 2006. doi:10.1037/0033-2909.132.3.354
  4. Harold Pashler, Mark McDaniel, Doug Rohrer, Robert Bjork. Learning Styles: Concepts and Evidence. Psychological Science in the Public Interest, 2008. doi:10.1111/j.1539-6053.2009.01038.x
  5. Kelli Taylor, Doug Rohrer. The Effects of Interleaved Practice. Applied Cognitive Psychology, 2010. doi:10.1002/acp.1598
  6. Olusola O. Adesope, Dominic A. Trevisan, Narayankripa Sundararajan. Rethinking the Use of Tests: A Meta-Analysis of Practice Testing. Review of Educational Research, 2017. doi:10.3102/0034654316689306
  7. John Sweller. Cognitive Load During Problem Solving: Effects on Learning. Cognitive Science, 1988. doi:10.1207/s15516709cog1202_4
  8. Victoria F. Sisk, Alexander P. Burgoyne, Jingze Sun, Jennifer L. Butler, Brooke N. Macnamara. To What Extent and Under Which Circumstances Are Growth Mind-Sets Important to Academic Achievement? Two Meta-Analyses. Psychological Science, 2018. doi:10.1177/0956797617739704
  9. David S. Yeager et al.. A National Experiment Reveals Where a Growth Mindset Improves Achievement. Nature, 2019. doi:10.1038/s41586-019-1466-y
  10. Paul Black, Dylan Wiliam. Inside the Black Box: Raising Standards Through Classroom Assessment. Phi Delta Kappan, 1998. link
  11. Paul A. Kirschner, John Sweller, Richard E. Clark. Why Minimal Guidance During Instruction Does Not Work: An Analysis of the Failure of Constructivist, Discovery, Problem-Based, Experiential, and Inquiry-Based Teaching. Educational Psychologist, 2006. doi:10.1207/s15326985ep4102_1

Cite this

Mustafa Pat. “How People Actually Learn: What the Evidence on Memory, Practice and Assessment Does and Doesn't Show.” EduFabTech, 13 September 2026. https://edufabtech.com/blog/how-people-actually-learn-evidence-memory-instruction