How AI Review's Agents Are Calibrated on 69,000 Real Peer Reviews
AI Review's seven language-model agents read real peer-review examples from a 69,000+ record corpus collected from 19 open-review platforms. Here is how the examples are chosen, and what they cannot do.
An AI reviewer that has never read a real peer review is guessing at what editors and referees care about. The reviewers inside Science AI Journal's AI Review do not have to guess: before each language-model agent writes a word of its report, it is shown real review reports drawn from a corpus of more than 69,000 records collected from 19 open peer-review platforms. This post explains where those reviews come from, how they reach each agent at review time, and what that kind of calibration can and cannot do.
Why calibration is the right word
Most discussion of AI reviewers focuses on the underlying language model. That misses much of the problem. A capable model that has never seen a real eLife review has no instinct for the biological-replicate rule (n = 3 technical replicates is not n = 3 biological replicates), for what machine-learning reviewers mean by "seed averaging", or for the point where a SciPost referee says "you conjectured, not proved". Those are domain norms, and they live in actual review records.
Calibration, as we use the term, is not fine-tuning. The model's weights are never changed. Instead, at the moment an agent runs, its prompt is prefixed with real review excerpts chosen for that agent's job. The examples shape what the agent notices, what it calls a major concern, and how it anchors its 1–10 score. This is a form of retrieval-augmented generation (RAG) as described by Lewis et al. in "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks" (arXiv:2005.11401, 2020), with one difference: the retrieved documents are not background facts but worked examples of expert judgement.
Where the corpus comes from
The corpus is built from review reports that the platforms themselves publish openly. Counts as of 27 September 2026:
| Source | Records |
|---|---|
| OpenReview (NeurIPS, ICLR, ICML, AISTATS, TMLR, EMNLP, ACL, COLM, MIDL, CoRL, CVPR) | 33,469 |
| SciPost | 11,433 |
| PLOS ONE | 5,043 |
| Copernicus Publications | 4,553 |
| PREreview | 3,328 |
| JOSS | 2,712 |
| PeerJ | 1,925 |
| Open Research Europe | 1,633 |
| PCI (Ecology, Evolutionary Biology, Neuroscience, Mathematical & Computational Biology) | 1,076 |
| Nature Communications | 1,057 |
| BMJ Open | 985 |
| eLife | 733 |
| ScienceOpen | 398 |
| F1000Research | 229 |
| Qeios | 180 |
| PubPeer, Review Commons, EMBO Press and the Royal Society | 237 |
| Review files extracted from papers in our local library (for example Nature Communications' peer-review files) | 2,086 |
| A small curated calibration set | 14 |
| Total | 71,091 |
Each record is stored with its source platform, its review type and section (a full report, or its weaknesses, strengths, questions or summary), a quality tier, and the editorial decision where the platform published one: accept, revise or reject. About 55,000 records carry such a decision; the rest are marked unknown. The corpus leans towards accepted papers — about 53% of records carry an accept decision — which matters for how examples are chosen.
How an agent's examples are chosen
Seven of AI Review's eight agents are language-model agents: Methodology, Formulas & Equations, Originality, Literature Coverage, Reproducibility, Clarity & Language, and Figures & Tables. They run in parallel. The eighth, Prior Publication, is a deterministic lookup and uses no examples at all (more on it below).
For each language-model agent, the server searches the corpus with a full-text index (SQLite FTS5, ranked by BM25). The search is built from that agent's own review vocabulary, between about 100 and 185 terms per agent: the methodology agent searches for terms such as "power analysis", "allocation concealment", "pseudoreplication" and "HARKing"; the figures agent for error bars, axis labels and colour-blind-safe palettes. The examples are chosen per agent, not per manuscript: your paper is what the agent reviews, and the examples show it the standard to review against.
From the candidates the search returns, a selection step then:
- prefers the "weaknesses" sections of high-quality reviews for agents whose job is to find flaws, and restricts the Methodology, Formulas, Originality and Figures agents to reviews rated high, curated or standard quality;
- aims for roughly 40% reject, 30% revise, 20% accept and 10% unknown decisions, so an agent sees what gets papers rejected, not only what gets them accepted;
- caps how many examples any single platform can contribute in that pass, so OpenReview, the largest source, does not crowd out the rest;
- drops excerpts too short to teach anything.
When the agents run on Claude, as they do in production, each agent call carries up to 60 examples, each trimmed to 3,000 characters.
What each agent is asked to check
The examples calibrate; the agent's own prompt sets the checklist. Each language-model agent returns a score from 1 to 10 with strengths, concerns, questions for the authors and a written justification that ties the score to a rubric.
Methodology carries venue-specific checklists — machine-learning venues (baselines, data leakage, seeds, ablations), eLife and the life sciences (biological replicates, blinding, randomisation), BMJ Open and clinical research (CONSORT, STROBE, PRISMA, GRADE), SciPost for physics and mathematics, Copernicus for earth and climate science, the PCI communities for ecology and evolution, and more — plus general red flags such as causation claimed from correlation, pseudoreplication and HARKing. Its rubric puts rigorous, well-controlled, reproducible designs at 9–10 and fundamental flaws such as data leakage or HARKing at 3–4.
Formulas & Equations checks derivations, dimensional consistency, conflicting notation and undefined variables, and proposes corrections.
Originality checks the manuscript's novelty claims against papers in our 4.5M-paper local library, and looks for text-level integrity signals, including statistics that are arithmetically impossible, such as GRIM-test failures.
Literature Coverage evaluates citation completeness, missing seminal references and self-citation balance (it flags self-citation above 30%), with a live OpenAlex snapshot of the field for context.
Reproducibility asks whether code and data are available and versioned, whether random seeds and compute are reported, and whether the methods are detailed enough for an independent lab to repeat the work.
Clarity & Language checks structure, abstract completeness (background, objective, methods, key results with numbers, and implications), undefined acronyms, overclaiming, and conventions such as qualifying every use of "significant".
Figures & Tables checks captions, defined error bars, readable axes, colour-blind-safe palettes, and whether small samples are shown as individual data points.
Prior Publication is not a language model. It fans out in parallel to seven external sources — CrossRef, PubMed, arXiv, bioRxiv, medRxiv, Europe PMC and Unpaywall — and to our local library, with a 12-second timeout per source, looking for title and abstract overlap. Our prior publication detection post explains how.
A final synthesis step reads the agents' reports and returns a verdict — Accept, Minor Revision, Major Revision or Reject — with the issues that must be addressed, alongside an overall score built from the agents' 1–10 scores.
What calibration cannot fix
Calibration on past reviews does not confer knowledge the corpus does not contain.
The Originality agent compares against our local library, not the whole literature, so it cannot see verbatim copying from a paper the library does not hold. The Prior Publication check is limited to what its sources and the library have indexed.
Coverage is uneven. OpenReview supplies about 47% of the corpus, reflecting the culture of open review in machine learning. Clinical research is thin by comparison — BMJ Open contributes 985 of the records — and small fields may have review norms that no platform in the corpus captures.
And calibration cannot fix a manuscript. An agent can tell you that Section 3.2 has no power analysis; it cannot run one for you. The report is a diagnosis, not a rewrite.
What this means for authors
If you are deciding whether to trust an AI review, the useful question is not "is it an AI?" but "what was it calibrated against, and how?" The sources above, the counts, and the way examples reach each agent are all documented here and on our dataset page.
When a full manuscript is ready, AI Review runs all eight agents over the PDF; it is paid in credits (see pricing). For a quicker read before that, Pre-Check is a separate, credit-metered tool: from a title and abstract it estimates acceptance odds for top, mid and open/emerging journals, based on where the most similar papers in our local library were published — it does not use this review corpus. For exploratory work, the Research Gap Finder is free to search, and unlocking a topic writes its research gaps from our 4.5M-paper local library.
Frequently asked questions
Is the model trained on these reviews? No. The model's weights are not changed. Real review excerpts are placed in an agent's prompt each time it runs, which is why this is called calibration by retrieval rather than training.
Does an agent see different examples for different manuscripts? Not today. Each agent's search is built from its own review vocabulary, not from your manuscript, so the methodology agent is calibrated on methodology critiques whatever the paper's field. The examples change as the corpus grows.
How does the reject/accept balance affect the scores? The corpus itself leans towards accepted papers, so the selection deliberately over-samples rejections. An agent calibrated only on accepted papers would miss the patterns that get papers rejected; one calibrated only on rejections would over-flag solid work.
Can I see which examples influenced a review of my paper? Not in the report today. The report shows each agent's own findings, score and justification.
Where do the reviews come from, and who owns them? From review reports the platforms above publish openly. Every record keeps its source platform, and we do not present the reviews as our own writing. Licence terms differ from platform to platform; our dataset page describes what we share and on what terms.
AI Review reports are written by AI agents. They are advice to the author, not an editorial decision, and nothing is sent to a journal on your behalf.
Tools mentioned in this post
Related posts
- When not to use AI peer review: honest limits and edge casesOur 8 AI agents are calibrated on 69,000 real peer reviews, but some papers fall outside their scope. Here is what to watch for before running a review.
- How AI enforces CONSORT, STROBE, and PRISMA in peer reviewHow our methodology agent checks RCTs against CONSORT 2010, observational studies against STROBE, and systematic reviews against PRISMA, calibrated on 69,000 real peer reviews.
- The complete guide to AI peer review in 2026How 8 specialized AI agents, calibrated on 69,000 real peer reviews from 19+ platforms, deliver rigorous, discipline-specific feedback in under 15 minutes.