Can an AI agent synthesize scientific evidence?

Can agents reliably do the work that powers clinical guidelines, health-technology assessment, and pharmaceutical evidence programs?

5 tasks with deterministic, verifiable rewards
View tasks
1 Title/abstract screening
Classify which papers to include from titles and abstracts.
2 Full-text screening
Classify eligibility against the complete paper.
3 Full-text extraction
Extract analysis-relevant data from each paper as JSON.
4 Risk of Bias
Judge each source of bias low, unclear, or high.
5 Meta-analysis
Run the statistical analysis on the extracted data.

How models fail

Models retrieve the evidence, then miss the scientific judgment

The recurring failure happens after reading. The model finds the paper and the numbers, then substitutes a familiar convention for the expert decision the task requires—“use the primary outcome,” “take the longest follow-up,” “combine the controls”—or attaches a correct value to the wrong outcome, treatment group, or measurement point.

More thinking changed the policy, not the quality

Four of five model families changed direction on extraction as thinking increased instead of improving. More thinking shifts what a model will include, exclude, and assert: it makes a sound approach more thorough—and a bad shortcut more systematic. The sharpest case turned uncertainty into false reassurance: psychotherapy trials cannot blind participants or therapists, and at a higher thinking setting the model marked that bias low risk where experts required high.

Read the failure taxonomy and ten trajectory analyses →

What this signal moves

Deep-research agents
Cover a large paper collection, keep a worklist, and reconcile many partial reads into one answer.
Scientific-literature agents
Identify the papers that answer a question, extract what they report, and preserve the link between evidence and claim.
Grounded extraction and abstention
Return supported datapoints and leave unsupported fields empty instead of producing plausible numbers.
Long-horizon coordination
Delegate document reading without losing inclusion criteria, treatment groups, or unresolved questions during the final merge.
Calibrated scientific judgment
Distinguish strong evidence from missing, ambiguous, or high-risk evidence without defaulting to confident labels.
Post-training with verifiable rewards
Train against reference answers and deterministic graders without using another model as the judge. This reward contract exists for all five tasks, including the two without current MetaPsy leaderboard scores.

These are expected transfer targets, not measured downstream gains. The public benchmark directly measures the five meta-analysis tasks described above.

Data

The first public benchmark is derived from the open MetaPsy databases. MetaPsy is maintained by an international collaboration led by Vrije Universiteit Amsterdam; its infrastructure is embedded in the WHO Collaborating Centre for Research and Dissemination of Psychological Interventions. The living databases are maintained by research groups across more than 25 international universities and research institutes.

The reference answers were created by research teams while producing real meta-analyses. MetaPsy’s current nine-person core team includes six documented doctorate holders. Its separate 42-person investigator roster includes at least ten people explicitly described as clinicians or licensed mental-health professionals. These counts describe the collaboration behind MetaPsy, not the authorship of every dataset.

  • Vrije Universiteit Amsterdam

    Netherlands

    MetaPsy is led by Vrije Universiteit Amsterdam.

  • University of Pennsylvania

    United States

    Penn is an affiliated institution for MetaPsy's Sypres psilocybin-depression database.

  • Technical University of Munich

    Germany

    Technical University of Munich is an affiliated institution for MetaPsy's Depression: Inpatients database.

  • The University of Tokyo

    Japan

    VU Amsterdam lists the University of Tokyo among the international research institutions collaborating in MetaPsy.

  • Massachusetts General Hospital

    United States

    Massachusetts General Hospital is represented through current MetaPsy principal investigator Samuel Acuff.

  • Brown University

    United States

    VU Amsterdam lists Brown University among the international research institutions collaborating in MetaPsy.

  • Dartmouth College

    United States

    VU Amsterdam lists Dartmouth among the international research institutions collaborating in MetaPsy.

  • Erasmus University Rotterdam

    Netherlands

    Erasmus University Rotterdam is represented through current MetaPsy principal investigator Mieke Schulte.

  • Dalhousie University

    Canada

    Dalhousie University is an affiliated institution for MetaPsy's eating-disorders database.

We want collaborators

Post-training teams, research-agent builders, evidence-synthesis groups, and labs that want to run models, contribute task families, or train against deterministic scientific rewards.

Work with us →