Methodology

How Science AI Journal performs rigorous peer review in under 15 minutes — training data, agent calibration, and the limits of what we claim.

The 5-step pipeline

  1. 1

    Submission & intake

    You upload a manuscript as PDF, Word (.doc/.docx) or LaTeX (.tex). The intake step extracts the title, abstract, body text and references (plus any images embedded in a Word file), and flags a low-quality PDF extraction so you can re-upload a cleaner file.

  2. 2

    Prior-publication fan-out

    Before any agent runs, the title and abstract fan out in parallel to 7 external sources — CrossRef, PubMed, arXiv, bioRxiv, medRxiv, Europe PMC and Unpaywall — and to our 4.5M-paper local library, with a 12-second timeout per source. Matches are surfaced with a confidence level and a link, so you can judge them yourself.

  3. 3

    Calibration retrieval

    For each language-model agent, real peer-review examples are retrieved from a SQLite FTS5 index of 69,000+ real reviews collected from 19+ open-review platforms (OpenReview, SciPost, PLOS ONE, Copernicus, eLife, and others). The retrieval is concern-aware: each agent pulls examples that match its own rubric, weighted towards higher-quality reviews and a mix of accept, revise and reject decisions across platforms.

  4. 4

    Parallel agent review

    The 8 agents — methodology, formulas, originality, literature coverage, reproducibility, clarity, figures, and prior publication — start at the same time. Each produces a score, a qualitative summary, and a structured report. The 7 language-model agents run on Claude Sonnet; the Prior Publication agent is a deterministic lookup, not a language model.

  5. 5

    Synthesis & editorial decision

    A synthesis pass reconciles the 8 agent reports into a single recommendation (accept, minor revision, major revision, reject) with a numeric score. Every agent's full report stays visible next to it. A journal publication route — with the review report published alongside the paper — is planned; it is not open yet.

Training data provenance

Every language-model agent is calibrated against real peer reviews collected from publicly accessible sources — not synthetic data and not proprietary publisher corpora. The training set totals 69,000+ reviews across 19+ platforms, largest first:

  • OpenReview (ICLR, NeurIPS, ICML and more)
  • SciPost
  • PLOS ONE
  • Copernicus (open discussion)
  • PREreview
  • Journal of Open Source Software
  • PeerJ
  • Open Research Europe
  • Nature Communications
  • Peer Community In
  • BMJ Open
  • eLife
  • F1000Research
  • ScienceOpen
  • Qeios
  • PubPeer
  • Review Commons
  • EMBO Press
  • Royal Society Open Science

Aggregated per-agent JSONL files live under training-data/by-agent/ and back the FTS5 retrieval index used during review.

Agent specialisations

8 specialist agents. The language-model agents are each calibrated on the slice of the training corpus that matches their concern; the Prior Publication agent is a deterministic lookup.

Methodology

Audits study design, statistical power, and analytical choices against field-specific rigour standards (CONSORT, STROBE, PRISMA).

Formulas & Equations

Verifies mathematical derivations, checks dimensional analysis, and flags algebraic errors.

Originality

Checks the manuscript's novelty claims against papers in our 4.5M-paper local library and flags possible overlap and self-plagiarism signals.

Literature Coverage

Evaluates citation completeness, missing seminal references and self-citation balance, with a live OpenAlex snapshot of the field (volume, top venues, peak year) for context.

Reproducibility

Inspects code availability, dataset accessibility, and sufficiency of methods detail for independent replication.

Clarity & Language

Assesses readability, structural flow, and adherence to scholarly writing norms.

Figures & Tables

Checks figure quality, caption completeness, and appropriateness of visual encodings.

Prior Publication

A deterministic lookup, not a language model: fans out in parallel to 7 external sources — CrossRef, PubMed, arXiv, bioRxiv, medRxiv, Europe PMC and Unpaywall — and our local library to detect prior publication and duplicate submission.

What we don't claim

AI peer review is not a replacement for domain-expert human review in high-stakes settings (clinical trials, safety-critical systems, paradigm-shift claims). We're transparent about the limits:

  • Agents can miss subtle methodological flaws that require cutting-edge domain knowledge.
  • Prior-publication detection is excellent for exact overlap but weaker for paraphrased or translated duplicates.
  • Originality scoring depends on index coverage — niche non-English work may be under-indexed.
  • The synthesis recommendation is a starting point, not a verdict: the decision stays with you and the journal you submit to.

Want the full picture?

Command palette

Jump anywhere, run any action.