A New Tool Automates Checking Papers Against Their Preregistrations
RegCheck, posted in January 2026, flags where a published paper strays from its registered plan, building on evidence that such drift is common and rarely disclosed.
A study registration is a promise: before collecting data, researchers state which outcomes they will measure, how many participants they will recruit, and which statistical test will decide the result. The published paper is supposed to keep that promise. Checking whether it does means opening two documents side by side — the registration and the paper — and reading them against each other line by line. Almost nobody does this routinely, because it is slow, and a tool posted to arXiv in January 2026 was built specifically to change that.
The tool, called RegCheck, comes from Jamie Cummins at the University of Oxford's Bennett Institute for Applied Data Science, together with Beth Clarke, Ian Hussey and Malte Elson at the University of Bern. It does not decide whether a paper has misreported its results. It pulls the registration and the paper apart, finds the passages that correspond to a chosen comparison point — a primary outcome, an eligibility threshold, a sample size target — and hands a reviewer the matched excerpts along with a provisional judgment: deviation, no deviation, or insufficient evidence to tell.
The reason this kind of tool is worth building is not hypothetical. A still-cited audit gives the base rate. The COMPare project, led by Ben Goldacre and published in the journal Trials in 2019, assessed every trial published over a six-week window in five journals that endorse the CONSORT reporting standard — the New England Journal of Medicine, The Lancet, JAMA, the BMJ and Annals of Internal Medicine. Of 67 trials assessed, 58 (87%, 95% CI 78–95%) had a discrepancy between what the registration specified and what the paper reported, serious enough that the team sent a correction letter to the journal. On average, trials reported only 76.3% of their pre-specified primary outcomes correctly, with the figure ranging from 25% to 96% depending on the journal, and only 55.1% of secondary outcomes (range 31–72%). Trials also added a mean of 5.4 outcomes that were never pre-specified (range 2.9–8.3 by journal). Only 23 of the 58 correction letters were published, a rate of 40% (95% CI 27–53%), and where journals did publish them the median delay was 99 days.

Those numbers describe outcome switching: reporting a different primary endpoint than the one that was registered, usually because it produces a more favorable result. It is one of several ways a paper can drift from its plan without saying so, and it is the kind of drift that a registration-versus-paper comparison is built to catch. The COMPare team did this by hand, trial by trial, over the course of a year. RegCheck's authors argue that this labor cost is exactly why the check is so rarely performed outside a dedicated audit project — reviewers are not generally asked to re-read a registration against a 4,000-word results section, so they usually don't.
How the comparison actually works
RegCheck runs on what its authors call the IDEA pipeline, named for its four stages: ingestion, definition, extraction and adjudication. Ingestion parses a registration — pulled from ClinicalTrials.gov, the Open Science Framework, or uploaded directly — alongside the paper itself, stripping reference sections and normalizing formatting so the two documents can be compared on equal footing, and for papers that report more than one study, removing material unrelated to the one being checked. In the definition stage, a user chooses what to compare: RegCheck ships with preset dimensions for clinical trials (primary and secondary outcomes, eligibility criteria, sample size, randomization method, recruitment dates, ethical approval), for psychology studies (preregistered hypotheses, planned analyses, outcomes), and for economics and preclinical research, with room to add custom dimensions. Extraction then splits both documents into sentence-aware chunks, embeds them, and retrieves the passages in each document that are actually relevant to a given comparison point rather than searching the whole text.
Adjudication is where a large language model is used, but narrowly: it summarizes the retrieved excerpts and issues one of three judgments — deviation, no deviation, or insufficient evidence — and every judgment is tied to a verbatim quotation from each document, so a human reviewer can check the model's reasoning against the actual sentence it is pointing at rather than taking the summary on faith. The final report is interactive and shareable, carrying a unique identifier so a judgment can be checked by someone other than the reviewer who ran the comparison.
The paper's own worked illustration, a hypothetical type 2 diabetes trial built to demonstrate the pipeline rather than a real misconduct case, shows what the output looks like in practice. Five comparison points came back as no deviation: eligibility criteria, intervention, ethics committee approval, start date and randomization method all matched the registration. Three came back flagged: the renal-function exclusion threshold had tightened from 30 to 45 mL/min between registration and publication, enrollment fell short at 212 of a planned 300 participants, and the primary outcome had switched from a continuous measure to a dichotomous responder endpoint — the exact kind of switch the COMPare audit found in practice. Three more points were marked insufficient evidence, including a missing ethics approval number and a secondary outcome the system could not locate anywhere in the manuscript text.
What the tool does not yet claim
We do not yet provide guarantees about the fidelity of RegCheck's deviation judgements.
That line, from the RegCheck paper itself, is worth sitting with before treating any output as settled. The authors describe their validation plan rather than a validation result: they intend to measure how closely RegCheck's judgments agree with human raters, and separately to test it against artificially generated registration-paper pairs with inconsistencies deliberately inserted, rather than report a single accuracy figure. Their stated bar is that the tool should agree with human reviewers at least as often as human reviewers agree with each other — not that it should be infallible. They are explicit that false negatives are possible, that the system can also flag changes a research team would consider immaterial, and that there is no agreed, field-wide standard for what counts as a meaningful deviation in the first place. RegCheck is built to surface candidates for a reviewer's attention, not to replace the reviewer's judgment about whether a given difference matters.

Other checks a reader can run alone
Registration comparison is one branch of a broader set of self-checking tools that a reader, not just a specialist auditor, can apply to a paper. A 2025 review in Theory & Psychology by Gabriel Crone and Christopher D. Green at York University surveys several of them, grouped under what the authors call statistical forensics: methods that exploit the fact that reported numbers are constrained by the underlying data in ways that make some combinations of statistics mathematically impossible.
- GRIM (granularity-related inconsistency of means) checks whether a reported mean could have been produced by any set of whole numbers, such as Likert-scale ratings, given the stated sample size. If it could not, something in the data or the write-up is wrong. The Theory & Psychology review notes it stops being useful once the per-cell sample size passes about 100.
- GRIMMER extends the same logic to reported standard deviations, checking whether a given SD is achievable from integer data at a given N.
- SPRITE goes further, attempting to reconstruct plausible underlying datasets that could generate a reported mean and standard deviation together, and works at larger sample sizes than GRIM or GRIMMER.
Crone and Green's review frames these tools by their comprehensiveness and rigor, while also flagging their practical limits: several require conditions — small, bounded, integer-based data — that do not hold for most published statistics, which keeps them from being applied as a blanket screen across a literature.
A separate, widely used check works on a different kind of number. A 2016 study in Behavior Research Methods by Michèle Nuijten and colleagues at Tilburg University introduced statcheck, software that independently recomputes a p-value from the test statistic and degrees of freedom a paper reports, then flags any mismatch against the value the authors printed. Applied to more than 250,000 p-values published across eight psychology journals, it found that about half of papers using null-hypothesis testing contained at least one statistically inconsistent p-value. Like GRIM and SPRITE, it works best as a targeted check once a specific number already looks unusual, not as something a reader runs automatically on everything they read.
A practical sequence for reading a paper against its plan
Put together, the COMPare audit, RegCheck and the statistical forensics tools suggest a rough order of operations for a reader who wants to check a paper rather than simply trust it.
- Find the registration. For trials, this usually means ClinicalTrials.gov or an equivalent national registry; for other study types, the Open Science Framework is common. The paper should name the registry and ID directly.
- Compare the primary outcome named in the registration against the primary outcome reported in the abstract and results. A switch here is the single discrepancy the COMPare audit found most often and considered most consequential.
- Check the planned sample size against the achieved sample size, and read whatever explanation the paper gives for any shortfall.
- Check whether the analysis described in the registration — which test, which comparison, which covariates — matches the analysis actually reported, and whether any change is disclosed and justified rather than silently made.
- If a specific number still looks off after that reading, tools like GRIM or statcheck can test whether it is internally consistent with the sample size and test statistics the paper reports elsewhere.
None of this guarantees a verdict on its own. A deviation is not automatically evidence of misconduct — plans change for defensible reasons, and both RegCheck's authors and the COMPare team are careful to say that a flagged difference is the start of a question, not the end of one. What the COMPare numbers established is that these differences are common enough, and go undisclosed often enough, that not asking the question is itself a choice. What RegCheck adds is an attempt to make asking it cheap enough that reviewers, and not just dedicated audit teams willing to spend a year on 67 trials, might actually do it.
A quick question for readers
Which study habit do you think is most overrated?
Some popular habits have little evidence behind them. Which one do you think gets too much credit?
Pick an answer, create a free account in a minute, and your vote counts. Already a member? Sign in
Comments 0
CommentNo comments yet. Members start the conversation.
Source: arXiv preprint (Cummins, Clarke, Hussey & Elson)
Sources (4)
- Jamie Cummins, Beth Clarke, Ian Hussey, Malte Elson. RegCheck: A tool for structured comparisons between study registrations and papers. arXiv preprint, 2026. arxiv.org ↗ · checked 9 Oct 2026
- Ben Goldacre et al.. COMPare: a prospective cohort study correcting and monitoring 58 misreported trials in real time. Trials, 2019. pmc.ncbi.nlm.nih.gov ↗ · checked 9 Oct 2026
- Gabriel Crone, Christopher D. Green. Tools of the data detective: A review of statistical methods to detect data and result anomalies in psychology. Theory & Psychology, 2025. doi.org ↗ · checked 9 Oct 2026
- Michèle B. Nuijten, Chris H. J. Hartgerink, Marcel A. L. M. van Assen, Sacha Epskamp, Jelte M. Wicherts. The prevalence of statistical reporting errors in psychology (1985–2013). Behavior Research Methods, 2016. link.springer.com ↗ · checked 9 Oct 2026