Help train agents that can do scientific evidence synthesis
The benchmark is useful only if it improves the systems that researchers and institutions actually deploy. We are looking for collaborators who can help run, stress-test, extend, and train against it.
For model labs and post-training teams
- Get in touch to run your agent-model stack on the full dataset.
- Compare retrieval, scientific judgment, scope control, and calibration rather than one aggregate score.
- Train on the private dataset.
- Help fund repeated runs so we can distinguish reliable improvements from one-off variation.
For evidence-synthesis researchers
- Contribute dataset families with clear provenance and task rights.
- Audit ambiguous reference answers and decisions about related publications.
- Help distinguish model mistakes from missing context, disputed references, grader problems, and incomplete run artifacts.
For infrastructure partners
- Provide model credits for open-source baselines.
- Support reproducible containers and long-running agent evaluation.