How AI enforces CONSORT, STROBE, and PRISMA in peer review
How our methodology agent checks RCTs against CONSORT 2010, observational studies against STROBE, and systematic reviews against PRISMA, calibrated on 69,000 real peer reviews.
Reporting checklists exist because peer reviewers are inconsistent. A seasoned clinical trialist will spot a missing allocation concealment statement on page one; a reviewer from a neighbouring field may miss it entirely. CONSORT 2010 was designed to fix that inconsistency — to give every reader, and every reviewer, the same checklist. In practice, compliance is still voluntary and manual.
Our methodology agent applies that checklist automatically, for every paper, before a human editor reads a word.
Why reporting standards still fail in practice
CONSORT 2010 — the Consolidated Standards of Reporting Trials — was published in the BMJ after two decades of evidence that randomised controlled trials were systematically under-reported (Schulz KF et al., BMJ 2010, doi:10.1136/bmj.c332). The 25-item checklist covers everything from sequence generation to outcome switching. Studies since have shown that checklist adoption correlates with a measurable reduction in bias in published trials.
Yet manual reviewer compliance remains inconsistent. A 2016 analysis of 268 RCTs in high-impact journals found that fewer than half reported all 25 CONSORT items, despite editor instructions requiring the checklist. The problem is not that reviewers disagree with the standard — it is that reviewing is cognitively demanding, and granular checklists are easy to skip when the scientific argument is compelling.
An AI agent does not make that trade-off. It checks the same 25 items on paper 1 and paper 1,000.
What the methodology agent checks for CONSORT
When our methodology agent identifies a randomised controlled trial, it evaluates the paper against CONSORT 2010 item by item. The items that generate the most revision requests in our corpus of 69,068 real peer reviews are:
Allocation concealment (item 9). The paper must describe the mechanism that concealed the allocation sequence — sealed envelopes, a central randomisation service, permuted-block via a pharmacy. Stating "patients were randomised 1:1" without naming the concealment mechanism fails item 9. Our BMJ Open–calibrated training examples flag this as a systematic, near-universal omission in trials from authors unfamiliar with CONSORT.
Blinding (item 11). Three groups must be addressed: participants, those administering the intervention, and outcome assessors. Many papers blind participants and providers but say nothing about whether outcome assessors were blinded. The agent flags each gap separately.
Outcomes and estimation (item 17). Both a p-value and a confidence interval for the primary outcome are required. A p-value alone is insufficient under CONSORT item 17, and our methodology agent reports it as a mandatory revision item.
Outcome switching. The pre-specified primary outcome in the trial registration must match the outcome foregrounded in the Results section. Post-hoc elevation of a secondary outcome to primary status is flagged as a critical integrity issue — not a minor revision, but grounds for a near-reject score in the 3–4 range.
The agent quotes the exact manuscript text it examines and names the CONSORT item number, so authors know precisely what is missing, not just that something is wrong.
STROBE: what the agent checks for observational studies
The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement covers cohort, case-control, and cross-sectional studies. Three items from the STROBE checklist drive the most revision requests in our calibration corpus:
Exposure definition. The exposure must be operationally defined — what was measured, how, and when — not just labelled. "Smoking status" is a label. "Self-reported current smoker, defined as ≥1 cigarette per day for the 12 months prior to enrolment" is a definition.
Missing data mechanism. MCAR (missing completely at random), MAR (missing at random), or MNAR (missing not at random) must be stated, with the handling method justified. "We used complete-case analysis" is flagged unless the authors demonstrate the data are MCAR. Multiple imputation is expected when data are MAR.
Loss to follow-up. Loss exceeding 20% in a cohort study triggers a critical flag regardless of the reported results, because the completers may be systematically different from those lost. This threshold is drawn directly from our PLOS ONE and BMJ Open calibration examples.
PRISMA: the agent's checks for systematic reviews
For systematic reviews and meta-analyses, the agent checks against the PRISMA 2020 checklist. Three checks are structurally critical:
PROSPERO registration. A systematic review without a PROSPERO registration number raises HARKing risk — the scope may have been adjusted after results were visible. The agent flags absent registration as a mandatory disclosure item.
Search strategy completeness. Searches must cover PubMed and EMBASE at minimum. Field-specific databases — PsycINFO for psychology, CINAHL for nursing, Cochrane CENTRAL for trials — are expected where the review scope demands them. The agent notes any database that the search appears to have omitted based on the review's stated topic.
GRADE certainty assessment. PRISMA 2020 requires GRADE certainty ratings for each main finding. A systematic review that reports pooled estimates without GRADE is flagged for major revision.
How calibration makes the check field-specific
The distinction between an agent that knows the CONSORT items and one that applies them as a real BMJ Open reviewer would is calibration. Before each review, our methodology agent receives up to 50 examples drawn from a corpus of 69,068 real peer reviews using BM25 full-text search across our FTS5 index. The examples come from BMJ Open, PLOS ONE, F1000Research, Open Research Europe, and PCI networks — all platforms with explicit CONSORT, STROBE, or PRISMA compliance requirements.
The retrieval is content-aware: a paper about a randomised drug trial surfaces BMJ Open and F1000Research examples; an ecological cohort study surfaces PCI Ecology and Royal Society Open Science examples. That matching is what produces item-level specificity — "CONSORT item 9 (allocation concealment) is not addressed" rather than "methodology is incomplete."
The decision distribution of retrieved examples is also controlled. At least 40% of examples across the corpus are "reject" or "major revision" decisions; no more than 40% come from any single platform. This prevents the agent from being over-calibrated to one venue's idiosyncratic standards. You can read more about how the full calibration corpus was built in our post on how our 8 AI agents were calibrated.
Score thresholds and what happens at each level
A CONSORT, STROBE, or PRISMA failure triggers a score in the 3–4 band on the methodology agent's 1–10 scale — the tier that indicates a fundamental flaw that is near-rejection on methodology grounds alone. The score is not punitive: it is calibrated against the same 69,068 reviews that human editors use to set acceptance thresholds at journals publishing under these standards.
A 5–6 score covers significant but correctable issues — missing a power analysis, or using complete-case analysis without stating the missingness assumption. A 7–8 score covers minor gaps, such as an absent confidence interval for a secondary outcome. A 9–10 requires fully compliant reporting with all items addressed.
The full report accompanying any score names the missing item, quotes the relevant manuscript section, and specifies exactly what the authors need to supply. For authors who want to see this before submission, the pre-submission scorer runs the full methodology check in under 15 minutes.
What this means for authors
If you are submitting a clinical trial, cohort study, or systematic review:
- Run your paper against the relevant checklist before submission. CONSORT, STROBE, and PRISMA are free at equator-network.org.
- Confirm your trial registration number appears in the Methods section — not just the abstract. The agent checks both.
- State explicitly whether outcome assessors were blinded. One sentence. It is the most commonly missing CONSORT item in our 69,068-review corpus, and it takes 30 seconds to add.
- For systematic reviews: if PROSPERO registration was omitted, explain when the protocol was fixed and why registration was deferred. The agent will flag absence; reviewers are more lenient when authors acknowledge the gap proactively.
The pre-submission scorer runs the full methodology check, including CONSORT/STROBE/PRISMA compliance, in under 15 minutes and returns a report with item-level specificity. If you are still developing your research question, the research-gap finder can help position your study against the existing literature before you design your methods.
FAQ
Does the methodology agent check CONSORT for all study types?
No. The agent identifies the study type from the manuscript before applying a checklist. CONSORT applies to randomised trials. STROBE applies to cohort, case-control, and cross-sectional designs. PRISMA applies to systematic reviews and meta-analyses. For ML papers, the agent checks against NeurIPS/ICLR reproducibility standards. For theoretical physics or mathematics, it checks SciPost-style proof completeness.
What if my trial was registered but after enrolment began?
Retrospective registration is flagged as a significant methodological concern, but not an automatic rejection. The agent will note the gap between registration date and enrolment start date if you disclose it. Concealing a retrospective registration — stating only the registration number without the date — is the pattern our calibration corpus flags as a critical integrity issue.
Can I appeal a score that I believe is wrong?
The methodology agent's report is one input into the editorial process. Human editors review every report before any decision is communicated. If you believe the agent misidentified your study type or misapplied a checklist item, the author response process allows you to address that directly with the handling editor.
How many clinical papers are in the 69,000-review training corpus?
Of the 69,068 reviews, the BMJ Open and Open Research Europe platforms — the most clinically oriented in our corpus — together contribute roughly 2,600 reviews. Additional clinical content is embedded in eLife, PLOS ONE, and Nature Communications reviews that address trial design and observational methodology.
Does the agent check outcome switching against the trial registration record?
The agent checks for the internal consistency signal: whether the primary outcome stated in the Methods matches the one emphasised in the Results. It cannot access external trial registries (ClinicalTrials.gov, ISRCTN) to verify the original registered outcome. For that cross-reference, the prior-publication agent and the human editor's desk check are the relevant checkpoints.
Reviews are performed by AI agents. Editorial decisions are made by human editors.
Tools mentioned in this post
Related posts
- The complete guide to AI peer review in 2026How 8 specialized AI agents, calibrated on 69,000 real peer reviews from 19+ platforms, deliver rigorous, discipline-specific feedback in under 15 minutes.
- How Long Does Peer Review Take? (And How to Get Editor-Ready Feedback in 15 Minutes)Peer review typically takes 1 to 6 months to the first decision, and often longer. Here is what drives the timeline, realistic ranges by field, and how to catch the problems reviewers will flag before you submit.
- How to Check if a Journal Is Predatory: A 5-Minute ChecklistA predatory journal charges publication fees while skipping real peer review. Here is a fast, evidence-based checklist to vet any journal before you submit — plus a free tool that flags them automatically.