How We Test Screening Tools: Datasets, Metrics, Assumptions, and Reproducibility
This page is our pre-registration. It is published before we run anything, and it does not change after we see results.
If you only read one sentence, read this one: we tell you exactly what we did not measure. Most “best screening software” pages compare marketing pages. We would rather publish a small table with an honest boundary around it than a large table that quietly guesses.
Why this page exists
Choosing screening software is a decision with a real cost. Pick a tool that misses relevant records and you damage the review. Pick a tool your budget cannot renew and you lose your project mid-screening.
The information you would need to make that decision well is scattered: vendor pages describe features but not performance, academic papers report performance but on datasets and metrics that differ from each other, and pricing changes without announcement.
We try to hold those three things in one place — and, critically, to keep them separated rather than blended into a single misleading score.
Our commitment on independence
Benchmark results and recommendations are independent of monetization. A free or non-affiliate product must rank first whenever the prespecified evidence shows it performs best.
Concretely:
- Affiliate availability, commission rates, and revenue potential are excluded from every score we compute — the performance results and the tool selector alike.
- ASReview is free and open source. If it wins, it is published as the winner, and we earn nothing from that page. That is the intended behaviour of this site, not an accident.
- We do not join an affiliate program for a tool before its evaluation is complete and frozen. We do not want a financial position in a result we have not yet produced.
- Any page containing affiliate links says so at the top.
The three layers, and why we never merge them
Every claim on this site belongs to exactly one of three layers. Blending them is the single most common way tool comparisons become fiction.
Layer A — Directly tested
Tools we could run ourselves, on the same dataset, under the same protocol, with the same metrics. Only these appear in the quantitative leaderboard.
Layer B — Published evidence
Tools we could not run ourselves. We summarise and re-analyse verifiable published studies, recording for each: the tools and versions evaluated, dataset, the study’s own metric definitions, sample, whether the vendor was involved in the evaluation, and stated limitations.
These results are reported separately. They are never placed in the Layer A leaderboard, because a number produced under a different protocol is not comparable, however similar its name.
Layer C — Feature and price verified
Pricing, plan limits, and trial restrictions, taken only from the vendor’s own pages, with the date we checked. Third-party review sites are not used as a source for prices.
The empty cell rule
If a tool is not in Layer A, its performance cell stays empty and is marked NOT DIRECTLY TESTED, with the reason.
We will not estimate. For example, Covidence’s individual trial is limited in the number of records it accepts, which makes a full-corpus run impossible for us. The honest output of that situation is an empty cell and an explanation — not an inferred number.
Dataset
Layer A runs use SYNERGY (doi:10.34894/HE6NAQ), an open dataset of study selection in systematic reviews: 26 reviews, 169,288 records, 2,834 of them labelled as included (1.67%). It is public, citable, and re-runnable by anyone who disagrees with us.
Subset selection. We do not run all 26 reviews and then choose which to report. We specify the subset in advance, and we report every review we specified — including the ones where results are unflattering or ambiguous. Selection criteria, fixed before running:
- spread across inclusion rates (SYNERGY ranges from roughly 0.2% to 22%), because tools behave very differently on sparse versus dense labels
- spread across corpus sizes, from a few hundred records to tens of thousands
- more than one field, so results are not an artefact of medical abstracts alone
The exact review list, with the reason each was chosen, is published alongside the results.
Metrics
| Metric | Definition |
|---|---|
| Recall @ 5% / 10% / 20% screened | Share of all relevant records found after screening that fraction of the corpus |
| Records to 95% recall | How many records you must read to find 95% of the relevant ones |
| Records to 100% recall | Same, for every relevant record |
| WSS @ 95% | Work Saved over Sampling: 0.95 − (records to 95% recall ÷ total records) |
| Setup cost | Wall-clock time and steps required to get from zero to a running screen |
Why not recall at 50%? We dropped it. Simulations stop once the last relevant record is found, so recall at high screening fractions is almost always 1.0 and separates nothing. It looks like a measurement and carries no information.
Why “records to recall” matters more than a percentage. “WSS 0.83” is abstract. “You read 247 records instead of 2,000” is the number that changes whether you can finish the review this month.
Runs, seeds, and variance
We never publish a single run.
In our pipeline test, changing only the random seed moved recall at 10% screened from 0.767 to 0.833 on identical data with an identical tool — a difference of 8.6 percentage points produced by nothing but chance. A single-run leaderboard would be reporting luck as if it were performance.
Therefore, for every tool × review combination:
- a minimum of 10 runs with pre-specified seeds
- we report the median and the full range, never a lone value
- seeds are published, so the exact runs can be repeated
Where the range for two tools overlaps substantially, we say they are not distinguishable on this evidence. We do not rank inside noise.
Stopping rules, exclusions, missing data
Stopping. Simulation runs until the last relevant record is found, or until the corpus is exhausted. Records after the stopping point count as unscreened.
Tool inclusion. A tool enters Layer A only if it can be run non-interactively on our corpus with a documented, reproducible configuration. Tools requiring manual clicking through a web interface for tens of thousands of records are not excluded from the site — they are documented in Layers B and C.
Exclusions we declare in advance. A run is excluded only for a mechanical failure — crash, timeout, corrupted output — and every exclusion is listed with its cause and the failing seed. Results are never excluded for being unexpected.
Missing data. Records lacking a title or an abstract are retained and counted, because they exist in real reviews and how a tool copes with them is part of its performance. The proportion of such records per corpus is reported.
Reproducibility
Everything needed to repeat our work is public: run scripts, configuration, tool and model versions, seeds, and raw results as CSV. The command to reproduce a result sits at the top of the results page.
If we cannot reproduce a result ourselves, we do not publish it.
Conflicts of interest
This site is run by a practising clinician-researcher who has conducted systematic reviews and has a direct interest in these tools working well. That is the source of whatever judgement this site offers, and it is also a bias worth naming.
The site intends to earn revenue through affiliate links to some of the tools discussed. The rules above exist to keep that revenue from touching the results. Every commercial link is disclosed. Where a top-ranked tool pays us nothing, we say so on the page.
What would change our conclusions
We would revise a published result if: a tool releases a version that materially changes its screening behaviour; someone reproduces our runs and gets different numbers; we discover an error in our own configuration; or a vendor demonstrates that our setup misused their tool.
Corrections are published as corrections, dated, with the previous value visible. We do not silently edit results.
Limitations we already know about
- Simulation on labelled historical reviews is not the same as prospective screening by a human team. It measures ranking quality, not the full workflow.
- SYNERGY is weighted toward medicine and psychology. Results may not transfer to fields far from those.
- We cannot directly test tools whose trial limits prevent a full run. That is a boundary of this site, and it is stated on every affected row.
- Setup cost is measured by us, once, on one machine. Treat it as an order of magnitude, not a benchmark.