Detecting prior publication across 8 sources in under 12 seconds
How we fan out across CrossRef, PubMed, Europe PMC, Unpaywall, arXiv, medRxiv, bioRxiv, and a local 4.5M-paper FTS5 index to catch prior publication early.
Duplicate submissions are the cheapest editorial reject a journal can make, and also the one most prone to false negatives. A paper already living on arXiv under a different title, or translated from a regional journal, looks novel to a tired editor but is trivially discoverable for a machine that's willing to query eight databases in parallel.
We built our prior-publication detector to run first, before any of the review agents reads a manuscript. Here's what's inside it.
Architecture: fan out, 12-second budget
submission
│
├── title + abstract
│
├── CrossRef (DOI + title fuzzy)
├── PubMed (biomedical literature)
├── Europe PMC (life-science literature)
├── Unpaywall (OA version search)
├── arXiv (preprint server)
├── medRxiv (biomedical preprint)
├── bioRxiv (biology preprint)
└── local library (FTS5 over 4.5M papers)
│
▼
score + confidence bands → report
Every upstream call runs with a 12-second timeout. If a source is slow or down, we log the failure and return what we have. The local FTS5 index backstops all remote failures: 4.5 million papers harvested from EBSCO EDS, CrossRef, KKU, and Unpaywall, their titles and abstracts indexed with SQLite's full-text engine.
Why a local index matters
Remote APIs have rate limits, go down, and have coverage gaps. A local
FTS5 index — queried with match on title || abstract — returns a
top-20 candidate set without a single network call. We then apply a
word-overlap score with stop-word filtering and flag ≥ 60% overlap as
high-confidence duplicate.
Daily harvest jobs keep adding papers to it, so the local index grows with its upstream sources instead of freezing as a one-off snapshot.
Fuzzy matching that isn't dumb
Two real failure modes naive title comparison hits:
- Punctuation drift. "Towards efficient…" vs "Towards Efficient…" is
trivial; "COVID-19 outcomes in type-2 diabetics" vs "Covid 19 outcomes
in type 2 diabetics" fools
LOWER(title) = LOWER(title)but overlaps 100%. - Translation shift. A Turkish-language paper reappearing in English shares almost no title words with its original.
We normalise before comparing — lower-case, dashes turned into spaces, punctuation stripped — so the first case collapses to an exact match. Beyond that, the score is the share of the shorter title's meaningful words (after a stop-word list drops "a, an, the, and, of, in, on, with, for" and similar fillers) found in the other title, blended 60/40 with a similar word overlap across the two abstracts when both are available. The second case stays hard: word overlap only helps when both versions carry an abstract in the same language, so a clean result means nothing was found in these sources, not that nothing exists.
How to read a flag
A match is a lead, not a verdict: open the matched record and compare it with your manuscript before you conclude anything. A preprint of your own work is reported as a preprint, not as prior publication — posting a preprint and then seeking a journal is normal practice. We have not published measured false-positive or false-negative rates for this detector, so this post quotes none.
→ Run a Pre-Check (it includes this scan) · How the review engine works
Tools mentioned in this post
Related posts
- Peer review in 15 minutes: how Science AI Journal worksAn inside look at our 8-agent review engine: what each agent checks, why eight narrow reviewers beat one broad prompt, and what the report does not claim.
- How we mined 120,000+ research gaps from the literatureHow our gap finder mines the limitations authors write about their own work, and why most 'AI gap-finders' hallucinate.