Dataset card

Science AI Journal Training Corpus

69,000+ real peer reviews from 19+ academic platforms. The calibration basis for the AI Review agents. CC BY 4.0.

69,000+
peer reviews
19+
source platforms
10
disciplines
CC BY 4.0
licence

Sources

The corpus is aggregated from public open-review pages across 19+ scholarly platforms. Approximate counts, floored, as of 2026-09-27 — exact numbers shift as scrapers catch up.

PlatformReviewsFocus
OpenReview~33,400ML, CS, AI conferences (NeurIPS, ICLR, ICML and more)
SciPost~11,400Physics, CS, mathematics
PLOS ONE~5,000Multidisciplinary
Copernicus journals~4,500Earth & environmental sciences (open discussion)
PREreview~3,300Open reviews of preprints
Journal of Open Source Software~2,700Research software — open review on GitHub
PeerJ~1,900Life, environmental and computer sciences
Open Research Europe~1,600Multidisciplinary (EU-funded research)
Nature Communications~1,100Multidisciplinary (transparent review)
Peer Community In~1,000Recommendation-based review of preprints
BMJ Open~900Medicine, public health
eLife~800Life sciences — full-length open review
F1000Research~700Life sciences — open post-publication review
Other open-review venues~2,200ScienceOpen, Qeios, PubPeer, Review Commons, EMBO Press, Royal Society Open Science, plus reviews extracted from library PDFs

How the corpus maps onto the 8 agents

Reviews are sliced by section (methodology, figures, literature, etc.) and indexed with SQLite FTS5. At review time, each language-model agent pulls the examples that match its own concern as Retrieval-Augmented-Generation context (the Prior Publication agent is a deterministic lookup and uses none). This is what makes a "methodology" agent behave like a methodology reviewer and not a generalist.

Methodology
Audits study design, statistical power, and analytical choices against field-specific rigour standards (CONSORT, STROBE, PRISMA).
Formulas & Equations
Verifies mathematical derivations, checks dimensional analysis, and flags algebraic errors.
Originality
Checks the manuscript's novelty claims against papers in our 4.5M-paper local library and flags possible overlap and self-plagiarism signals.
Literature Coverage
Evaluates citation completeness, missing seminal references and self-citation balance, with a live OpenAlex snapshot of the field (volume, top venues, peak year) for context.
Reproducibility
Inspects code availability, dataset accessibility, and sufficiency of methods detail for independent replication.
Clarity & Language
Assesses readability, structural flow, and adherence to scholarly writing norms.
Figures & Tables
Checks figure quality, caption completeness, and appropriateness of visual encodings.
Prior Publication
A deterministic lookup, not a language model: fans out in parallel to 7 external sources — CrossRef, PubMed, arXiv, bioRxiv, medRxiv, Europe PMC and Unpaywall — and our local library to detect prior publication and duplicate submission.

Discipline coverage

EngineeringLife SciencesPhysical SciencesComputer ScienceMedicineMathematicsEnvironmental ScienceSocial ScienceHumanitiesEconomics

Calibration method

Each language-model agent is prompted with real peer reviews drawn from this corpus and matched to its own concern, so it reasons against how referees actually write about that concern rather than against a generic prior. At the last count (2026-09-27), 55,213 of the reviews carry the human editorial decision that accompanied them — 37,660 accept, 11,438 reject, 6,115 revise. We have not published a measured agreement rate between our agents and human editorial decisions, and we are not going to quote one we have not run. The numbers we have actually measured, with their method and their failures, are on our benchmarks page.

A weekly calibration report checks each agent’s training mix — score-range coverage, decision balance, platform diversity — and flags gaps for a person to act on; nothing is adjusted automatically. Reach out if you want the raw numbers for a research paper.

Frequently asked questions

Is the dataset downloadable?
A summary aggregate is available at /api/dataset/summary as JSON. The full reviews are derived from public open-review pages; we publish parsed JSONL tranches on request for academic collaborators, under the terms of each source platform's licence.
How is the data used?
Exclusively for Retrieval-Augmented-Generation (RAG) calibration of the AI Review agents. Each language-model agent's prompt is prefixed with real peer-review examples that match its own concern — weighted towards higher-quality reviews and a mix of accept, revise and reject decisions — so the agent's rubric matches human reviewer expectations.
Do you re-host reviewer comments that were posted under pseudonyms?
Only where the source platform's licence and policy explicitly allow it. OpenReview review text is CC BY; eLife and PLOS transparent reviews are CC BY; other sources are handled case-by-case. Where we cannot republish, we extract rubric patterns (not verbatim text) into the calibration corpus.
Does the dataset include the papers themselves?
No — the dataset is the reviews, not the manuscripts. For manuscript context we query the relevant open-access paper at runtime from the platform's API.
How do you keep calibration fresh?
Scrapers re-run on rolling schedules. A weekly calibration report (scripts/training/calibrate_agents.py, Fridays) checks the per-agent training mix — score-range coverage, whether the score-to-decision mapping is consistent with the human labels, platform diversity, and section-type balance — and flags gaps for a human to act on. It reads static data only, makes no model calls, and does not adjust anything automatically.
Can my institution contribute a tranche?
Yes — if your journal or conference has transparent review records you want included, email [email protected]. We credit source and preserve licence terms.
Dataset summary (JSON)How the agents use itResearch collaboration

Command palette

Jump anywhere, run any action.