How to Read a Systematic Review: What the 2023 Cochrane Mask Trial Analysis Actually Showed
A widely misquoted review of mask-promotion trials is a precise, reusable lesson in reading risk ratios, confidence intervals and GRADE certainty ratings before accepting a headline.
In early 2023, an update to a long-running Cochrane review of physical interventions against respiratory virus spread moved from a specialist database into general newspaper columns within days. The headline that stuck, in dozens of outlets, was that a leading evidence-synthesis organisation had found masks "don't work." The review's own editor-in-chief publicly disagreed with that reading weeks later. Both the review and the correction are still online, and together they make an unusually clean case study in what a systematic review actually licenses you to say β and what it doesn't.
The review, led by Tom Jefferson and colleagues, pooled randomised and cluster-randomised trials of interventions meant to interrupt the spread of influenza, SARS-CoV-2 and other respiratory viruses: hand hygiene, physical distancing and mask promotion among them. The mask comparisons rested on a small number of trials β most of them cluster-randomised, meaning whole communities or workplaces were assigned to an intervention rather than individuals β comparing medical or surgical masks against no masks, and separately comparing N95/P2 respirators against medical masks.
That design detail matters more than it sounds like it should. A cluster-randomised trial of a mask programme measures whether encouraging a population to wear masks changes an outcome; it does not measure, under controlled conditions, what masks do when consistently worn. The review's own summary says as much, and the distinction is the hinge the entire controversy turns on.

What the numbers actually say
Readers who only saw the words "no evidence" never saw the numbers behind them. The review reports a risk ratio for each comparison β the ratio of event rates between the mask and no-mask groups β alongside a 95% confidence interval and a GRADE certainty rating, which Cochrane uses to describe how much the true effect might differ from the pooled estimate.
| Comparison | Trials / participants | Risk ratio (95% CI) | Certainty |
|---|---|---|---|
| Medical/surgical masks vs none β influenza-like or COVID-like illness | 9 trials, 276,917 participants | 0.95 (0.84β1.09) | Moderate |
| Medical/surgical masks vs none β lab-confirmed influenza/SARS-CoV-2 | 6 trials, 13,919 participants | 1.01 (0.72β1.42) | Moderate |
| N95/P2 vs medical/surgical masks β influenza-like illness | 5 trials, 8,407 participants | 0.82 (0.66β1.03) | Low |
| N95/P2 vs medical/surgical masks β lab-confirmed influenza | 5 trials, 8,407 participants | 1.10 (0.90β1.34) | Moderate |
Every one of those confidence intervals straddles 1.0, the line of no effect. That is what "probably makes little or no difference" means in the review's own wording β a point estimate close to parity, with an interval wide enough to be compatible with a moderate protective effect, a moderate harmful one, or nothing at all. A "moderate-certainty" or "low-certainty" rating in Cochrane's GRADE framework is a statement about how much confidence to place in that estimate, downgraded for reasons such as risk of bias, inconsistency across trials, indirectness or imprecision β not a verdict that the intervention has no effect. Imprecision alone, driven by wide confidence intervals, is grounds for downgrading even when the point estimate looks neutral.
Absence of evidence, evidence of absence
This is the distinction the review's critics kept returning to: a wide, uncertain confidence interval is "absence of evidence" about the size of an effect, not "evidence of absence" of one. The review itself flagged a mechanism for why its interval was so wide. In the most heavily weighted trial feeding the pooled estimate β a large cluster-randomised study in Bangladesh β only 42.3% of people in villages assigned to a mask-promotion intervention actually wore masks, against 13.3% in control villages, a gap cited in Cochrane's own statement on the controversy. A trial measuring a 29-point difference in actual mask-wearing, not a difference between "masked" and "unmasked" populations, is answering a narrower question than the one the headlines assumed.

The correction, and its limits
On 10 March 2023, Cochrane's then editor-in-chief, Karla Soares-Weiser, issued a statement conceding that "many commentators have claimed that a recently updated Cochrane review shows that 'masks don't work,' which is an inaccurate and misleading interpretation," and that "the review examined whether interventions to promote mask wearing help to slow the spread of respiratory viruses, and that the results were inconclusive." She added that the plain-language summary's wording "was open to misinterpretation, for which we apologize," and committed to revising it.
That revision did not straightforwardly happen. By June 2024, after further exchanges with the review's authors, Cochrane announced it was no longer seeking changes to the plain-language summary or abstract, judging that the disputed wording did not affect the review's "scientific integrity." The episode is itself worth noting methodologically: an organisation's own leadership can flag a misreading, and the underlying document can still end up unrevised, because the authors and the editors reasonably disagreed about where imprecise language ends and incorrect language begins.
A second disagreement: what should count as evidence at all
Independent scrutiny of the review split roughly along a fault line familiar to anyone who reads evidence hierarchies closely. Writing in the American Journal of Public Health, CDC researchers Brian Gurbaxani, Andrew Hill and Pragna Patel catalogued trial-level problems with the two COVID-era mask trials the review leaned on: the Danish DANMASK-19 trial was underpowered and ran during a period of low transmission, relying on antibody testing rather than acute infection markers, and produced a confidence interval running from a 46% reduction to a 23% increase in infections β nowhere near resolving the question either way. They argued that restricting synthesis to a handful of randomised trials, while treating modelling and laboratory studies as inadmissible, can produce a narrower and less reliable picture than combining evidence types.
A separate group of epidemiologists, writing in Antimicrobial Stewardship & Healthcare Epidemiology, reached the opposite methodological conclusion from a similar starting point: Daniel Halperin, Shira Doron and colleagues argued that randomised trials remain the strongest available evidence for population-level mask policy precisely because observational and mechanistic studies are more exposed to confounding, and that the pooled RCT evidence, even including the newer COVID-era trials, still had not demonstrated a measurable population-level benefit from mask mandates. Both groups read the same risk ratios. They disagreed about what weight a randomised design should carry against its own practical limitations β a disagreement about evidentiary philosophy, not arithmetic.
What to check before repeating a review's headline
None of this requires taking a side on masks to be useful. It requires treating the episode as a worked example of five checks that apply to any systematic review or meta-analysis before its topline conclusion gets repeated:
- Read the risk ratio or effect size and its confidence interval, not just the certainty label or the plain-language summary β a wide interval straddling the null is a different finding from a narrow one centred on it.
- Check which GRADE domain caused any downgrade β risk of bias, inconsistency, indirectness, imprecision or publication bias point to different problems, and imprecision in particular often means "underpowered," not "null."
- Identify exactly what was randomised and what was measured β a trial of a promotion programme, with partial uptake, answers a narrower question than a trial of the thing itself under full compliance.
- Look for the authors' own stated limitations before accepting a secondhand summary of them β Cochrane's reviewers noted the high or unclear risk of bias in their included trials in the review itself.
- When a review becomes a public flashpoint, read the subsequent commentary from people who disagree with each other, not just with the review β as the AJPH and ASHE exchanges show, informed readers can extract opposite policy conclusions from the same pooled numbers once they weigh evidence types differently.
The Cochrane mask review was not wrong about its own numbers; the fight was almost entirely about what those numbers were entitled to say beyond the confidence interval printed next to them. That is the recurring failure mode in how published evidence reaches the public β not fabricated statistics, but real ones detached from the certainty rating and the population they were measured in. A risk ratio of 0.95 with a confidence interval from 0.84 to 1.09, carried from the abstract into a headline without the interval, can be made to say almost anything. Carried with it, it says something much more modest and much more durable: this particular bundle of trials could not distinguish a protective effect from no effect at all, in a specific population, under specific and often poor adherence. Reading that sentence instead of the headline is the entire skill.
A quick question for readers
Which study habit do you think is most overrated?
Some popular habits have little evidence behind them. Which one do you think gets too much credit?
Pick an answer, create a free account in a minute, and your vote counts. Already a member? Sign in
- Jefferson T, Dooley L, Ferroni E, Al-Ansary LA, van Driel ML, Bawazeer GA, et al.. Physical interventions to interrupt or reduce the spread of respiratory viruses. Cochrane Database of Systematic Reviews, 2023. doi:10.1002/14651858.CD006207.pub6
- Cochrane. Statement on "Physical interventions to interrupt or reduce the spread of respiratory viruses" review. Cochrane, 2023. link
- Gurbaxani BM, Hill AN, Patel P. Unpacking Cochrane's Update on Masks and COVID-19. American Journal of Public Health, 2023. doi:10.2105/AJPH.2023.307377
- Halperin DT, Doron S, Hodgins S, Bailey RC, Baral S, Bhatia R, Noble J, Gandhi M, Hearst N. Masking for COVID-19 and other respiratory viral infections: implications of the available evidence. Antimicrobial Stewardship & Healthcare Epidemiology, 2024. doi:10.1017/ash.2024.67
- FactCheck.org (Annenberg Public Policy Center). SciCheck: What the Cochrane Review Says About Masks for COVID-19 β and What It Doesn't. FactCheck.org, 2023. link
- Cochrane. Cochrane Handbook for Systematic Reviews of Interventions, Chapter 14: Completing 'Summary of findings' tables and grading the certainty of the evidence. Cochrane, 2024. link