How the benchmark works
What is public
- Instructions and graders for five tasks.
- The code and curation decisions used to rebuild MetaPsy reference answers.
- Aggregate results for title/abstract screening, full-text extraction, and Risk of Bias.
- Dataset manifests, file hashes, run accounting, and error analyses that do not reproduce private or copyrighted source material.
- DOI manifests pointing to the source papers.
What is not redistributed
- Copyrighted abstracts as worked examples.
- Copyrighted full-text papers or OCR.
- Model trajectories that reproduce copyrighted source text.
- Private-dataset tasks, answers, scores, costs, and per-question performance.
- Licensed assessment handbooks or other restricted methodology material.
The majority of full-text papers are copyrighted. Openness therefore means open instructions, graders, reference answers, provenance, and score receipts; it does not mean republishing every source document.
What we reconstructed
The five-task framework is broader than the current MetaPsy dataset. Three tasks currently have reconstructed MetaPsy inputs and public model results.
| Task | Current MetaPsy dataset | What the reference represents |
|---|---|---|
| Title/abstract screening | 15 research questions; 4,384 candidate papers | Expert include and exclude decisions reconstructed from replicated PubMed search results |
| Full-text screening | No MetaPsy dataset yet | Instructions and grader only |
| Full-text extraction | 13 research questions; 40 trial reports; 194 datapoints | Researcher-audited datapoints traceable to pinned MetaPsy data and selected OCR-backed reports |
| Risk of Bias | 10 research questions; 35 report/outcome records; 178 judgments | Expert labels reconstructed from pinned MetaPsy releases for selected OCR-backed reports |
| Meta-analysis | No MetaPsy dataset yet | Instructions and grader only |
Title/abstract screening
The title/abstract benchmark covers 15 research questions and 4,384 candidate papers.
We replayed the published PubMed search strategies with their original date limits and preserved each query, PubMed's translation, unavailable record IDs, retrieval time, and file hashes. We used the 15 research questions whose search strategies could be reconstructed faithfully from PubMed. Two others relied on database-specific searches that could not be replayed from an explicit PubMed query, so they were not included.
The complete reconstructed searches returned 609,918 records. The scored benchmark keeps all 1,045 unique papers included by the expert researchers and intentionally samples 3,339 excluded papers from the same replicated search results at several difficulty levels. Because the reconstructed pool matches the search universe reviewed by the original researchers, absence from their final inclusion set supplies the expert exclusion decision.
Full-text extraction
The extraction benchmark covers 13 research questions, 40 OCR-backed trial reports, and 194 datapoints. Candidate datapoints were checked against the papers during benchmark authoring. Every scored reported value also traces to a cited datapoint in the pinned MetaPsy dataset release, or to a declared calculation such as deriving a standard deviation from a reported standard error.
The public rebuild verifies those recorded decisions, hashes, source rows, report identities, and declared calculations. It does not reread the OCR and repeat the scientific audit from scratch on every rebuild. The facts were OCR-audited during authoring; that is not a claim that every number appears literally in the OCR.
Risk of Bias
The Risk of Bias benchmark covers 10 research questions, 35 selected report/outcome records, and 178 expert judgments. The labels are rebuilt from the relevant columns in pinned MetaPsy releases and attached to the report the agent reads.
The papers are OCR-backed inputs for the agent. The reconstruction proves where the expert labels came from, but it does not independently prove that every label can be recovered unambiguously from the available OCR.
Full-text screening and meta-analysis
Both tasks have public instructions and graders, but neither currently has a materialized MetaPsy dataset or public model score. They are supported stages; the current MetaPsy release does not benchmark them.
Reconstruction exclusions
| Task | Before exclusions | Explicit exclusions | Released population | Reason / boundary |
|---|---|---|---|---|
| Title/abstract screening | 17 | 2 excluded: metapsy-psychosis-psyctr-23-1-4metapsy-ptsd-psyctr-23-0-3 | 15 | The source protocol had no explicit PubMed search block to reproduce; its Ovid-only search was outside the public PubMed reconstruction method. |
| Full-text extraction | 17 | 4 excluded: metapsy-depression-perinatal-psyctr-26-0-1metapsy-depression-psyctr-24-0-2metapsy-depression-selfguided-psyctr-24-0-3metapsy-panic-psyctr-23-0-1 | 13 | Removed from the validated extraction participation scope after benchmark-quality review; it contributes no released extraction result or score. |
| Risk of Bias | 10 | None declared | 10 | No task-specific exclusion is declared in run accounting. |
The two title/abstract exclusions had Ovid-only searches with no explicit PubMed block to reproduce. The four extraction exclusions were removed from the validated participation scope after benchmark-quality review, before materialization. Run accounting declares no Risk of Bias exclusion. "Before exclusions" is task-local released population plus its explicit exclusion rows; it is never added across tasks.
All five task contracts still have deterministic, inspectable rewards with no LLM-as-judge. This table covers only the three reconstructed MetaPsy populations with released scores.
Where MetaPsy comes from
The first dataset is built from the open MetaPsy databases, which cover psychotherapy research in depression, anxiety, post-traumatic stress disorder, obsessive-compulsive disorder, gambling, and related conditions. MetaPsy is maintained by an international collaboration led by Vrije Universiteit Amsterdam. Its infrastructure is embedded in the WHO Collaborating Centre for Research and Dissemination of Psychological Interventions.
MetaPsy reports that its living databases are maintained by research groups from more than 25 international universities and research institutes. Its current nine-person core team includes six documented doctorate holders. A separate 42-person investigator roster includes at least ten people explicitly described as clinicians or licensed mental-health professionals. These counts describe the MetaPsy collaboration rather than the authorship of every benchmark dataset.
Institutions represented across the collaboration include Vrije Universiteit Amsterdam, the University of Amsterdam, Amsterdam UMC, Dartmouth and the US National Center for PTSD, the University of Pennsylvania, the University of Cape Town, Stellenbosch University, the University of Verona, the Technical University of Munich, and the University of Tokyo.
The benchmark is not limited to psychotherapy. Selecting papers, extracting data, assessing methods, and running statistical analysis are also core parts of clinical guidelines, health-technology assessment, and pharmaceutical evidence programs.
How we test the graders
| Task | Reward metric | Oracle | Observed degenerate / shortcut probes |
|---|---|---|---|
| Title/abstract screening | Asymmetric Cohen’s κ | 1.00 | Advance no candidates = 0.00 Advance every candidate = 0.00 |
| Full-text screening | Asymmetric Cohen’s κ | 1.00 | Advance no candidates = 0.00 Advance every candidate = 0.00 |
| Full-text extraction | F1 | 1.00 | Emit no extracted facts = 0.00 Emit 40 non-matching rows against 10 references = 0.00 |
| Risk of Bias | Linearly weighted Cohen’s κ | 1.00 | Emit no judgments against a majority-low reference = 0.00 Predict low for every component domain = 0.00 Predict unclear for every component domain = 0.00 |
| Meta-analysis | Point estimate and 95% CI within tolerance | 1.00 | Match direction and significance but not the pooled numbers = 0.00 |
The integrity suite detected four of five attempted grader breaks. Extraction's sole attempted break was a stale sed no-op, and title/abstract screening has no grader-break check, so these artifacts do not claim complete coverage. A 1.00 oracle establishes agreement with the versioned reference under the versioned grader, not undisputed scientific truth.
Obvious shortcuts must score poorly
Include everything. Extract nothing. Assign the same Risk of Bias judgment to every paper. Submit the wrong statistical result. The test suite runs these shortcuts against the real graders and requires poor or zero scores.
The oracle must score exactly 1.0
Every reference answer is submitted through the same staged grader used for model answers. The oracle passes with a perfect score for each task's own released population: 15 title/abstract research questions, 13 extraction research questions, and 10 Risk of Bias research questions. These populations overlap but differ and are never added into one benchmark total. A perfect oracle score proves the dataset and grader agree with each other; it does not prove that every scientific judgment is beyond dispute.
Every public fact needs a trail
The rebuild fails when a selected extraction value cannot be resolved uniquely to its cited MetaPsy rows, when a paper identity is ambiguous, when a declared calculation is missing, or when an accepted fact disappears from the compiled reference.
Limits
- Some disagreements concern whether a related publication, rather than the underlying study, answers the research question.
- Expert Risk of Bias judgments can contain legitimate boundary cases.
- A source-grounded extraction that falls outside the research question is not the same error as a fabricated value.
- Due to budget constraints, we ran each research question, model, and thinking setting once (
k=1), so these comparisons are descriptive and do not measure repeat-run stability.