Dataset card
Science AI Journal Training Corpus
69,000+ real peer reviews from 19+ academic platforms. The calibration basis for the AI Review agents. CC BY 4.0.
Sources
The corpus is aggregated from public open-review pages across 19+ scholarly platforms. Approximate counts, floored, as of 2026-09-27 — exact numbers shift as scrapers catch up.
| Platform | Reviews | Focus |
|---|---|---|
| OpenReview | ~33,400 | ML, CS, AI conferences (NeurIPS, ICLR, ICML and more) |
| SciPost | ~11,400 | Physics, CS, mathematics |
| PLOS ONE | ~5,000 | Multidisciplinary |
| Copernicus journals | ~4,500 | Earth & environmental sciences (open discussion) |
| PREreview | ~3,300 | Open reviews of preprints |
| Journal of Open Source Software | ~2,700 | Research software — open review on GitHub |
| PeerJ | ~1,900 | Life, environmental and computer sciences |
| Open Research Europe | ~1,600 | Multidisciplinary (EU-funded research) |
| Nature Communications | ~1,100 | Multidisciplinary (transparent review) |
| Peer Community In | ~1,000 | Recommendation-based review of preprints |
| BMJ Open | ~900 | Medicine, public health |
| eLife | ~800 | Life sciences — full-length open review |
| F1000Research | ~700 | Life sciences — open post-publication review |
| Other open-review venues | ~2,200 | ScienceOpen, Qeios, PubPeer, Review Commons, EMBO Press, Royal Society Open Science, plus reviews extracted from library PDFs |
How the corpus maps onto the 8 agents
Reviews are sliced by section (methodology, figures, literature, etc.) and indexed with SQLite FTS5. At review time, each language-model agent pulls the examples that match its own concern as Retrieval-Augmented-Generation context (the Prior Publication agent is a deterministic lookup and uses none). This is what makes a "methodology" agent behave like a methodology reviewer and not a generalist.
Discipline coverage
Calibration method
Each language-model agent is prompted with real peer reviews drawn from this corpus and matched to its own concern, so it reasons against how referees actually write about that concern rather than against a generic prior. At the last count (2026-09-27), 55,213 of the reviews carry the human editorial decision that accompanied them — 37,660 accept, 11,438 reject, 6,115 revise. We have not published a measured agreement rate between our agents and human editorial decisions, and we are not going to quote one we have not run. The numbers we have actually measured, with their method and their failures, are on our benchmarks page.
A weekly calibration report checks each agent’s training mix — score-range coverage, decision balance, platform diversity — and flags gaps for a person to act on; nothing is adjusted automatically. Reach out if you want the raw numbers for a research paper.
Frequently asked questions
- Is the dataset downloadable?
- A summary aggregate is available at /api/dataset/summary as JSON. The full reviews are derived from public open-review pages; we publish parsed JSONL tranches on request for academic collaborators, under the terms of each source platform's licence.
- How is the data used?
- Exclusively for Retrieval-Augmented-Generation (RAG) calibration of the AI Review agents. Each language-model agent's prompt is prefixed with real peer-review examples that match its own concern — weighted towards higher-quality reviews and a mix of accept, revise and reject decisions — so the agent's rubric matches human reviewer expectations.
- Do you re-host reviewer comments that were posted under pseudonyms?
- Only where the source platform's licence and policy explicitly allow it. OpenReview review text is CC BY; eLife and PLOS transparent reviews are CC BY; other sources are handled case-by-case. Where we cannot republish, we extract rubric patterns (not verbatim text) into the calibration corpus.
- Does the dataset include the papers themselves?
- No — the dataset is the reviews, not the manuscripts. For manuscript context we query the relevant open-access paper at runtime from the platform's API.
- How do you keep calibration fresh?
- Scrapers re-run on rolling schedules. A weekly calibration report (scripts/training/calibrate_agents.py, Fridays) checks the per-agent training mix — score-range coverage, whether the score-to-decision mapping is consistent with the human labels, platform diversity, and section-type balance — and flags gaps for a human to act on. It reads static data only, makes no model calls, and does not adjust anything automatically.
- Can my institution contribute a tranche?
- Yes — if your journal or conference has transparent review records you want included, email [email protected]. We credit source and preserve licence terms.