Methodology
How Science AI Journal performs rigorous peer review in under 15 minutes — training data, agent calibration, and the limits of what we claim.
The 5-step pipeline
- 1
Submission & intake
You upload a manuscript as PDF, Word (.doc/.docx) or LaTeX (.tex). The intake step extracts the title, abstract, body text and references (plus any images embedded in a Word file), and flags a low-quality PDF extraction so you can re-upload a cleaner file.
- 2
Prior-publication fan-out
Before any agent runs, the title and abstract fan out in parallel to 7 external sources — CrossRef, PubMed, arXiv, bioRxiv, medRxiv, Europe PMC and Unpaywall — and to our 4.5M-paper local library, with a 12-second timeout per source. Matches are surfaced with a confidence level and a link, so you can judge them yourself.
- 3
Calibration retrieval
For each language-model agent, real peer-review examples are retrieved from a SQLite FTS5 index of 69,000+ real reviews collected from 19+ open-review platforms (OpenReview, SciPost, PLOS ONE, Copernicus, eLife, and others). The retrieval is concern-aware: each agent pulls examples that match its own rubric, weighted towards higher-quality reviews and a mix of accept, revise and reject decisions across platforms.
- 4
Parallel agent review
The 8 agents — methodology, formulas, originality, literature coverage, reproducibility, clarity, figures, and prior publication — start at the same time. Each produces a score, a qualitative summary, and a structured report. The 7 language-model agents run on Claude Sonnet; the Prior Publication agent is a deterministic lookup, not a language model.
- 5
Synthesis & editorial decision
A synthesis pass reconciles the 8 agent reports into a single recommendation (accept, minor revision, major revision, reject) with a numeric score. Every agent's full report stays visible next to it. A journal publication route — with the review report published alongside the paper — is planned; it is not open yet.
Training data provenance
Every language-model agent is calibrated against real peer reviews collected from publicly accessible sources — not synthetic data and not proprietary publisher corpora. The training set totals 69,000+ reviews across 19+ platforms, largest first:
- OpenReview (ICLR, NeurIPS, ICML and more)
- SciPost
- PLOS ONE
- Copernicus (open discussion)
- PREreview
- Journal of Open Source Software
- PeerJ
- Open Research Europe
- Nature Communications
- Peer Community In
- BMJ Open
- eLife
- F1000Research
- ScienceOpen
- Qeios
- PubPeer
- Review Commons
- EMBO Press
- Royal Society Open Science
Aggregated per-agent JSONL files live under training-data/by-agent/ and back the FTS5 retrieval index used during review.
Agent specialisations
8 specialist agents. The language-model agents are each calibrated on the slice of the training corpus that matches their concern; the Prior Publication agent is a deterministic lookup.
Methodology
Audits study design, statistical power, and analytical choices against field-specific rigour standards (CONSORT, STROBE, PRISMA).
Formulas & Equations
Verifies mathematical derivations, checks dimensional analysis, and flags algebraic errors.
Originality
Checks the manuscript's novelty claims against papers in our 4.5M-paper local library and flags possible overlap and self-plagiarism signals.
Literature Coverage
Evaluates citation completeness, missing seminal references and self-citation balance, with a live OpenAlex snapshot of the field (volume, top venues, peak year) for context.
Reproducibility
Inspects code availability, dataset accessibility, and sufficiency of methods detail for independent replication.
Clarity & Language
Assesses readability, structural flow, and adherence to scholarly writing norms.
Figures & Tables
Checks figure quality, caption completeness, and appropriateness of visual encodings.
Prior Publication
A deterministic lookup, not a language model: fans out in parallel to 7 external sources — CrossRef, PubMed, arXiv, bioRxiv, medRxiv, Europe PMC and Unpaywall — and our local library to detect prior publication and duplicate submission.
What we don't claim
AI peer review is not a replacement for domain-expert human review in high-stakes settings (clinical trials, safety-critical systems, paradigm-shift claims). We're transparent about the limits:
- Agents can miss subtle methodological flaws that require cutting-edge domain knowledge.
- Prior-publication detection is excellent for exact overlap but weaker for paraphrased or translated duplicates.
- Originality scoring depends on index coverage — niche non-English work may be under-indexed.
- The synthesis recommendation is a starting point, not a verdict: the decision stays with you and the journal you submit to.