Benchmarks & methodology
The numbers we quote elsewhere on this site, with the method behind them, the sample size, the caveats, and the runs where our own measuring instrument broke. Last measured 2026-07-26.
Journal Finder — venue match accuracy
46 real published papers across 10 fields. The recommender sees only the title and abstract; a “hit@k” means the paper’s actual publication venue appeared in the top k of the ranked shortlist. Ground truth is where the paper was really published.
On Turkish and other regional venues (6 papers) the top-5 hit rate is 66.7% — the long tail that larger indexes tend to miss, and the reason we keep those venues in the index.
By field — where we are weak
| Field | Papers | Hit@5 | MRR |
|---|---|---|---|
| Chemistry | 4 | 100% | 0.8 |
| Engineering | 4 | 75% | 0.292 |
| Physics | 4 | 75% | 0.258 |
| Biology | 4 | 50% | 0.365 |
| Computer science | 4 | 50% | 0.528 |
| Economics | 4 | 50% | 0.331 |
| Social science | 4 | 50% | 0.335 |
| Medicine | 9 | 44.4% | 0.145 |
| Psychology | 5 | 40% | 0.149 |
| Earth science | 4 | 25% | 0.082 |
Each field holds only 4–9 papers, so one paper swings a cell by 11–25 points. Treat this as a rough signal about where we are weak (earth science, psychology, medicine), not a precise score.
Run history, including the broken runs
Measured MRR has been stable across runs. Separately, three weekly runs produced no measurement at all because the harness could not reach the recommender — those were briefly recorded as scores of zero, which twice triggered a false internal regression alert. They are listed here rather than quietly dropped.
- 2026-07-06MRR 0.312
- 2026-07-10MRR 0.302
- 2026-07-13MRR 0.303
- 2026-07-26MRR 0.305
- 2026-07-06Recommender unreachable (connection refused) — not scored.
- 2026-07-10Recommender unreachable (connection refused) — not scored.
- 2026-07-20Recommender unreachable (connection refused) — not scored.
How often AI assistants cite us
We run a fixed set of 38 questions weekly against a web-search-enabled model and count how often the answer cites this site. In the 2026-07-19 run: 2 of 38. Both citations were brand-name queries ("Science AI Journal", "SAIJ journal"). Every generic-intent query — "best AI tool to find research gaps", "free AI peer review tool" — returned answers that did not cite us. The result was identical for four consecutive weeks.
Corpus figures
The counts quoted across the site. All are floored — the live database numbers are equal or slightly higher — and are re-checked by an automated drift audit.
- 4.5M papers indexed in the library
- 120,000+ open research gaps mined from 100,000+ papers
- 1,100+ gaps with a permanent citable analysis page
- 17,500+ submittable venues in the journal index
- 69,000+ real peer reviews used for calibration
What these numbers do not tell you
- 46 fixtures is a small sample. Per-field cells hold 4–9 papers each, so a single paper moves a field's score by 11–25 points. Read the per-field table as a rough signal, not a precise measurement.
- The venue index grows over time, so an older run is not strictly comparable to a newer one.
- This is our tool measured against ground truth, not a head-to-head against another product. We have captured no competitor rankings, so we publish no comparison.
- A paper's actual venue is only one acceptable answer. A recommendation we score as a miss may still be a good venue for that paper.
- We do not publish a measured quality score for AI Review. Review quality is genuinely hard to measure and we do not have a defensible number, so we do not quote one.
Frequently asked
How accurate is the Science AI Journal journal finder?
On a fixed set of 46 real published papers, the paper's actual venue appeared in our top 5 recommendations 54.3% of the time and in the top 10 67.4% of the time (run of 2026-07-26, MRR 0.305). The tool sees only the title and abstract. Accuracy varies widely by field — chemistry scored 100% at top-5 on 4 papers, earth science 25%.
Is this a comparison against other journal finders?
No. This measures our tool against ground truth — the venue each paper was actually published in. Our harness has slots for competitor rankings captured by an operator and they are currently empty, so we publish no head-to-head claim. Anyone presenting a cross-tool accuracy league table for this category is not measuring, they are guessing.
Why publish your failed benchmark runs?
Because three of them were previously recorded as genuine scores of zero, and that twice triggered an internal 'quality regression' alert that pointed at the wrong problem. The recommender was fine; the harness could not reach it. A benchmark page that hides its instrument failures is not a benchmark page.
How often do AI assistants cite Science AI Journal?
Rarely, and we measure it. In the 2026-07-19 run of our citation probe, 2 of 38 queries produced an answer citing us — and both were brand-name searches. Every generic question, like "best AI tool to find research gaps", returned an answer that did not mention us.