AI Review, agent by agent

The 8 agents behind AI Review

AI Review splits a manuscript review into 8 specialist checks that run at the same time. Of those, 7 are language-model reviewers, each calibrated on real peer reviews for its own specialty, and Prior Publication is a deterministic lookup across 7 external sources and our 4.5M-paper local library. A synthesis step then writes one report. A full review costs 30 credits.

8
specialist agents, started together
7
language-model reviewers
69K+
real peer reviews to calibrate them
7
external prior-publication sources, plus our library
30
credits for a full review

What each one checks

One narrow job per agent

Each language-model agent carries only its own rubric, so it can point at specifics — a section, a table, a sentence — instead of generalising. The Prior Publication agent uses no language model at all.

  • Methodology

    Language model · calibrated on real reviewsWeight ×2.0 in the overall score

    Audits study design, statistical power, and analytical choices against field-specific rigour standards (CONSORT, STROBE, PRISMA).

    What it looks for

    • Study design against reporting standards such as CONSORT, STROBE and PRISMA
    • Sample-size justification, statistical power, and whether the test suits the data
    • Red flags: data leakage, pseudoreplication, HARKing, causal claims from correlational data
  • Formulas & Equations

    Language model · calibrated on real reviewsWeight ×1.2 in the overall score

    Verifies mathematical derivations, checks dimensional analysis, and flags algebraic errors.

    What it looks for

    • Dimensional consistency of each equation's terms
    • Undefined variables and one symbol used for two quantities
    • Values in the text that disagree with the equations or tables, with a suggested correction
  • Originality

    Language model · calibrated on real reviewsWeight ×1.5 in the overall score

    Checks the manuscript's novelty claims against papers in our 4.5M-paper local library and flags possible overlap and self-plagiarism signals.

    What it looks for

    • Novelty claims set against similar papers in the local library
    • Incremental versus genuinely new contribution, and overclaiming such as "first ever"
    • Self-plagiarism, salami-slicing and duplicate-submission signals in the text
  • Literature Coverage

    Language model · calibrated on real reviewsWeight ×1.2 in the overall score

    Evaluates citation completeness, missing seminal references and self-citation balance, with a live OpenAlex snapshot of the field (volume, top venues, peak year) for context.

    What it looks for

    • Missing foundational and directly competing work
    • Self-citation balance and recency for the field
    • Miscitation and weak sources for key claims

    Reads the reference list it can see. With LaTeX, attach your .bib or .bbl so it reads that too.

  • Reproducibility

    Language model · calibrated on real reviewsWeight ×1.0 in the overall score

    Inspects code availability, dataset accessibility, and sufficiency of methods detail for independent replication.

    What it looks for

    • Where the data and code live: a repository, DOI or accession number, not "available on request"
    • Software versions, hardware, random seeds and dataset splits
    • Trial registration, protocol and analysis plan for clinical work
  • Clarity & Language

    Language model · calibrated on real reviewsWeight ×0.8 in the overall score

    Assesses readability, structural flow, and adherence to scholarly writing norms.

    What it looks for

    • An abstract with background, objective, methods, results with numbers, and a conclusion
    • Hedging: overclaims such as "proves", and needless over-hedging
    • Consistent terminology, over-long sentences, undefined acronyms and section flow
  • Figures & Tables

    Language model · calibrated on real reviewsWeight ×0.8 in the overall score

    Checks figure quality, caption completeness, and appropriateness of visual encodings.

    What it looks for

    • Self-contained captions: units, error-bar type, n per group, the test used
    • Chart type for the data, colour-blind-safe palettes, resolution
    • Numbers in the text that disagree with a figure or table, and image-integrity flags

    Sees the figure images embedded in a Word file, or the figure files you attach to LaTeX; from a PDF it works from the captions and text.

  • Prior Publication

    Deterministic lookup · no language modelReported on its own, not averaged in

    A deterministic lookup, not a language model: fans out in parallel to 7 external sources — CrossRef, PubMed, arXiv, bioRxiv, medRxiv, Europe PMC and Unpaywall — and our local library to detect prior publication and duplicate submission.

    What it looks for

    • An existing record whose title and abstract closely match yours
    • Searched in parallel: 7 external sources and our 4.5M-paper local library
    • A likely or possible prior publication is flagged in the report, with a link to the matching record when the source gives one

    Reads the title and abstract only.

From 8 reports to one

How the agents become one report

  1. 1

    They start together

    All 8 agents begin at the same time, so this stage takes about as long as the slowest single agent, not the sum of all of them.

  2. 2

    Each language-model agent is calibrated first

    Before it reads your paper, each of the 7 is shown real peer reviews retrieved for its own specialty from a corpus of 69K+ reviews collected from 19+ open-review platforms — reviewers' critiques first, with a mix of decisions and sources — and told to hold its concerns to that standard.

  3. 3

    It gets context on the field

    When the local library has them, each also receives research gaps that similar papers have already stated in its area, plus a live OpenAlex snapshot of the field: how much has been published, the top venues and the peak year.

  4. 4

    Scores become one weighted score

    Methodology counts ×2.0 and Originality ×1.5, the two heaviest weights; Prior Publication is reported on its own and not averaged in. Any language-model agent scoring 2 or below caps the overall at 4.0. The verdict follows the score: Accept from 8.5, Minor Revision from 6.5, Major Revision from 4.0, Reject below that.

  5. 5

    A synthesis step writes one report

    One report: the verdict, a recommendation that cites what the agents found, the issues to address first, the paper's strengths, any formula corrections and the reproducibility gaps.

Run all 8 on your manuscript

Upload a PDF, Word or LaTeX file. A full review costs 30 credits, charged once for every agent and the synthesis. Credits come in one-time packs of 50 and never expire; there is no subscription.

Questions

About the agents

No. The reviewers are these 8 software agents, not people. AI Review is a pre-submission check: the report is for you, and it is not a publication decision anywhere.

Command palette

Jump anywhere, run any action.