Updated 2 min readengineering

Detecting prior publication across 8 sources in under 12 seconds

How we fan out across CrossRef, PubMed, Europe PMC, Unpaywall, arXiv, medRxiv, bioRxiv, and a local 4.5M-paper FTS5 index to catch prior publication early.

By Science AI Journal Editorial

Duplicate submissions are the cheapest editorial reject a journal can make, and also the one most prone to false negatives. A paper already living on arXiv under a different title, or translated from a regional journal, looks novel to a tired editor but is trivially discoverable for a machine that's willing to query eight databases in parallel.

We built our prior-publication detector to run first, before any of the review agents reads a manuscript. Here's what's inside it.

Architecture: fan out, 12-second budget

submission
    │
    ├── title + abstract
    │
    ├── CrossRef       (DOI + title fuzzy)
    ├── PubMed         (biomedical literature)
    ├── Europe PMC     (life-science literature)
    ├── Unpaywall      (OA version search)
    ├── arXiv          (preprint server)
    ├── medRxiv        (biomedical preprint)
    ├── bioRxiv        (biology preprint)
    └── local library  (FTS5 over 4.5M papers)
          │
          ▼
     score + confidence bands → report

Every upstream call runs with a 12-second timeout. If a source is slow or down, we log the failure and return what we have. The local FTS5 index backstops all remote failures: 4.5 million papers harvested from EBSCO EDS, CrossRef, KKU, and Unpaywall, their titles and abstracts indexed with SQLite's full-text engine.

Why a local index matters

Remote APIs have rate limits, go down, and have coverage gaps. A local FTS5 index — queried with match on title || abstract — returns a top-20 candidate set without a single network call. We then apply a word-overlap score with stop-word filtering and flag ≥ 60% overlap as high-confidence duplicate.

Daily harvest jobs keep adding papers to it, so the local index grows with its upstream sources instead of freezing as a one-off snapshot.

Fuzzy matching that isn't dumb

Two real failure modes naive title comparison hits:

  1. Punctuation drift. "Towards efficient…" vs "Towards Efficient…" is trivial; "COVID-19 outcomes in type-2 diabetics" vs "Covid 19 outcomes in type 2 diabetics" fools LOWER(title) = LOWER(title) but overlaps 100%.
  2. Translation shift. A Turkish-language paper reappearing in English shares almost no title words with its original.

We normalise before comparing — lower-case, dashes turned into spaces, punctuation stripped — so the first case collapses to an exact match. Beyond that, the score is the share of the shorter title's meaningful words (after a stop-word list drops "a, an, the, and, of, in, on, with, for" and similar fillers) found in the other title, blended 60/40 with a similar word overlap across the two abstracts when both are available. The second case stays hard: word overlap only helps when both versions carry an abstract in the same language, so a clean result means nothing was found in these sources, not that nothing exists.

How to read a flag

A match is a lead, not a verdict: open the matched record and compare it with your manuscript before you conclude anything. A preprint of your own work is reported as a preprint, not as prior publication — posting a preprint and then seeking a journal is normal practice. We have not published measured false-positive or false-negative rates for this detector, so this post quotes none.

→ Run a Pre-Check (it includes this scan) · How the review engine works

#prior-publication#openalex#engineering

Related posts

Command palette

Jump anywhere, run any action.