Peer review in 15 minutes: how Science AI Journal works
An inside look at our 8-agent review engine: what each agent checks, why eight narrow reviewers beat one broad prompt, and what the report does not claim.
Traditional peer review takes between 2 and 18 months, depending on field. For a PhD student with a tight submission deadline — or a clinician with a finding that might change practice — that timeline is incompatible with how science actually moves today.
AI Review runs a manuscript through 8 specialised AI reviewers calibrated on 69,000+ real peer reviews scraped from OpenReview, eLife, SciPost, PLOS ONE, BMJ Open, Nature Communications, and a dozen other open review platforms. A manuscript uploaded in the morning comes back well before lunch with a report from every agent and a synthesis recommendation — feedback to act on before you submit, not an editorial decision.
What each agent checks
| Agent | Focus | How it works |
|---|---|---|
| Methodology | Study design, sample-size justification, CONSORT / STROBE / PRISMA compliance, seeds and baselines for ML work, causal validity | Language-model agent; the rubric is keyed to the study design it detects |
| Formulas & Equations | Derivations, notation, dimensional consistency | Language-model agent; every concern must quote the equation |
| Originality | Novelty claims against the prior work the paper cites, self-plagiarism signals | Language-model agent |
| Literature Coverage | Whether the references present are representative, current, and support the claims they are attached to | Language-model agent; a deterministic check then flags any author-year, DOI or cite-key it names that is not in the manuscript |
| Reproducibility | Code and data availability, seeds, hyperparameters, protocol detail | Language-model agent |
| Clarity & Language | Whether the argument is followable and the claims are precisely stated | Language-model agent |
| Figures & Tables | Axis labels, error bars, caption sufficiency, figure–claim match | Language-model agent; reads the figure images when they are supplied |
| Prior Publication | Title and abstract matched against CrossRef, PubMed, Europe PMC, Unpaywall, arXiv, medRxiv, bioRxiv and our local 4.5M-paper FTS5 index, in parallel | Database lookup — no language model |
| Synthesis | Weighs all eight reports and writes the overall recommendation and score | Language-model pass over the agent reports |
Why 8 and not 1?
A single monolithic prompt hits two walls. First, it hallucinates — it wants to find issues everywhere, so it invents them. Second, it can't hold all the relevant context: methodology rubrics for every study design, plus statistics guidance, plus figure-reading heuristics, plus calibration examples for each, exceed any single prompt's useful attention.
The agent pattern lets each reviewer carry only the rubric it needs and only the calibration examples matching its domain. We have not published a measured agreement rate between the agents and human editorial decisions, and we will not quote a number we have not run; what we have measured is on the benchmarks page.
What you get back
Every agent's full report, not just a score: what each one flagged, the manuscript passage it points to, and the synthesis that weighs them. It is yours to act on before you submit to any journal. A journal publication route, with the review report published beside each paper, is planned; it is not open for submissions yet, and we have not published any papers. Open access without open review is transparency theatre; when that route opens, we think the two have to ship together.
What we won't claim
- We do not replace human peer review for stakes where it genuinely matters — drug trials, regulatory submissions, grant panels.
- We do not outperform a careful, well-resourced human reviewer on nuanced theoretical work.
- We do not generate novel scientific insight. We review.
What we do claim: for most manuscripts, a competent, fast, transparent first pass before a journal's referees see them is a better default than waiting months to learn about problems you could have fixed.
Tools mentioned in this post
Related posts
- Detecting prior publication across 8 sources in under 12 secondsHow we fan out across CrossRef, PubMed, Europe PMC, Unpaywall, arXiv, medRxiv, bioRxiv, and a local 4.5M-paper FTS5 index to catch prior publication early.
- The complete guide to AI peer review in 2026How 8 specialized AI agents, calibrated on 69,000 real peer reviews from 19+ platforms, deliver rigorous, discipline-specific feedback in under 15 minutes.
- Hallucination Guardrails for an AI Reviewer: What We Do, What We Still Can't FixHow Science AI Journal's 8 AI agents guard against fabricated review claims -- and where those guards have real limits. A transparent technical account.