SR-Bench
We asked the frontier models behind ChatGPT, Codex and Claude Code to run systematic reviews end to end, and scored each stage. They can now do the whole review, but they aren't ready to do it unsupervised.
TL;DR: We turned ten published systematic reviews into scored environments. Then we asked five frontier models to run each review on their own: search PubMed, screen thousands of records, extract data from full-text papers and write the code that pools the results. Every stage was scored against an expert-adjudicated ground truth. The best configuration scores 63.8% on our 0–100% reward. The original review authors, scored the same way, average 78.9%, and otto-SR 91.8%. The models can now work through a whole review, but they make errors that are hard to see and can change its conclusion, so they aren't ready to run a review unsupervised.
Key findings:
- A clear gap to the original review authors. At max reasoning effort, Opus 5 scores 63.8% and GPT-5.6 Sol 60.6%. The original authors, scored the same way, average 78.9%, and otto-SR 91.8%.
- Performance scales with test-time compute. GPT-5.6 Sol roughly doubles its score from its lowest effort settings (under 30%, about 7k reasoning and tool-call tokens per run) to max effort (60.6%, about 130k tokens).
- The same failure modes recur across models. We see strict searches, title-only or regex screening, invented rationales, forgotten studies and plausible but wrong methodological choices. In one run, a single wrong inclusion turned a non-significant pooled result (p = .132) into a significant one (p = .033).
Tap a point for details
Line chart of SR-Bench reward against reasoning plus tool tokens per run on a log scale, for five models across reasoning-effort levels, with two dashed reference lines: otto-SR at 91.8% and the original review authors at 78.9%. Labelled scores at each model's highest effort level are Opus 5 63.8%, GPT-5.6 Sol 60.6%, Grok 4.5 48.8%, Opus 4.8 48.6% and Gemini 3.6 Flash 28.5%. Every model curve rises overall with more tokens and stays below both reference lines.
Show the dataHide the data
| Model | Effort | Reward | Tokens per run | ±1 rerun SD (pts) |
|---|---|---|---|---|
| GPT-5.6 Sol | none | 26.4% | 7,087 | ±4.4 |
| GPT-5.6 Sol | low | 27.6% | 7,003 | ±3.8 |
| GPT-5.6 Sol | medium | 36.1% | 14,298 | ±4.4 |
| GPT-5.6 Sol | high | 47.8% | 26,808 | ±2.6 |
| GPT-5.6 Sol | xhigh | 54.9% | 50,831 | ±1.8 |
| GPT-5.6 Sol | max | 60.6% | 131,028 | ±2.8 |
| Opus 5 | low | 47.8% | 40,632 | ±3.9 |
| Opus 5 | medium | 52.7% | 65,889 | ±2.5 |
| Opus 5 | high | 56.5% | 108,195 | ±1.6 |
| Opus 5 | xhigh | 58.2% | 169,991 | ±2.1† |
| Opus 5 | max | 63.8% | 193,990 | ±2.2 |
| Opus 4.8 | low | 29.4% | 32,475 | ±5.2 |
| Opus 4.8 | medium | 37.9% | 58,729 | ±3.9 |
| Opus 4.8 | high | 46.8% | 87,810 | ±2.2 |
| Opus 4.8 | xhigh | 43.7% | 177,512 | ±4.0 |
| Opus 4.8 | max | 48.6% | 292,625 | ±1.3† |
| Grok 4.5 | low | 35.5% | 27,380 | ±4.0 |
| Grok 4.5 | medium | 43.3% | 36,738 | ±3.7 |
| Grok 4.5 | high | 48.8% | 40,539 | ±2.0 |
| Gemini 3.6 Flash | minimal | 24.4% | 32,576 | ±3.5 |
| Gemini 3.6 Flash | low | 24.9% | 51,412 | ±4.2 |
| Gemini 3.6 Flash | medium | 29.3% | 75,269 | ±3.9 |
| Gemini 3.6 Flash | high | 28.5% | 96,711 | ±3.8 |
† Partial variance estimate: Opus 5 xhigh covers 7 of 8 reviews; Opus 4.8 max covers 5 of 8 reviews.
Why test agents on systematic reviews
Many life-sciences companies now give their teams enterprise ChatGPT or Claude, and coding agents such as Codex and Claude Code can search, run code and work through long tasks on their own. So a fair question for any evidence team is whether a general-purpose agent can simply run the systematic review. SR-Bench measures the answer, stage by stage.
Systematic reviews are the strongest form of evidence in medicine; clinical guidelines rest on them. Every review follows the same workflow: search the literature for every relevant study, screen the results against eligibility criteria, extract data from the full texts, then analyze it, usually with a meta-analysis.
That workflow makes a demanding test for agents:
- It tests several capabilities at once: search at corpus scale, eligibility reasoning, reading long scientific PDFs and correct statistics, all working together.
- It is long-horizon. At max effort, the models in Figure 1 average roughly 130,000 to 300,000 reasoning and tool-call tokens per run. Decisions made early limit what is possible at the end.
- Every step can be checked. A study is in or out, an extracted value is right or wrong, and a pooled estimate matches or it doesn't.
What we built
SR-Bench is an internal research project with three parts: environments built from published reviews with expert-adjudicated answers (Figure 3), an agent setup in which a model runs a whole review on its own, and a verifier that scores every stage.
The ten reviews
We chose ten recently published systematic reviews that vary in clinical domain, review type, effect measure, size and novelty (Table 1). Four are de novo: no earlier published review asked the question, so the agent has no prior synthesis to lean on.
Table 1. The ten reviews, by number of included studies in SR-Bench.
| Clinical domain | Review type | Effect measure | Included studies in SR-Bench† | Prior published review?* |
|---|---|---|---|---|
| Ophthalmology | Exposure–outcome association | Odds ratio | 6 | No |
| Oncology | Exposure–outcome association | Hazard ratio | 10 | Yes |
| Cardiovascular | Intervention | Standardized mean difference | 12 | No |
| Sleep medicine | Exposure–outcome association | Standardized mean difference | 14 | Yes |
| Orthopedics | Intervention | Mean score and mean difference | 21 | No |
| Rheumatology | Exposure–outcome association | Relative risk | 24 | Yes |
| Psychiatry | Intervention | Standardized mean difference | 30 | Yes |
| Infectious disease | Diagnostic test accuracy | Sensitivity and specificity | 32 | Yes |
| Palliative care | Intervention | Standardized mean difference | 37 | No |
| Public health | Exposure–outcome association | Correlation (r) | 39 | Yes |
†The studies used in each environment, which can differ from the published review's count.
*Yes if the review updates an earlier review on the same question, or if any earlier published systematic review asks the same or a similar question. Reviews marked No are de novo.
Ground truth
Published reviews are not a perfect gold standard: some contain simple errors. So we reproduced each review with otto-SR, our AI platform for systematic reviews, starting from the original search, and an expert reviewer on our team adjudicated every disagreement between otto-SR and the original authors. otto-SR and the original authors were then rescored against this adjudicated reference standard. We also opted to leave risk of bias unscored because of its inherent noisiness and subjectivity.
Agent setup
We tested the underlying models (GPT-5.6 Sol, Opus 5, Opus 4.8, Grok 4.5 and Gemini 3.6 Flash), not the ChatGPT, Codex or Claude Code products. Each model ran in our own agent setup, which resembles those coding agents: a sandbox, a few tools and a task to finish without help.
- PubMed API for searching and fetching records.
- Bash, with Python and R for analysis.
- Full-text retrieval for open-access papers.
No web search. We couldn't reliably restrict web search to results published before a review's date, and at higher effort, models often found the published review and copied its answers: straightforward reward hacking.
Inputs and outputs. The system prompt is one line: "You are a research scientist who performs systematic reviews. You have access to a computer." The task prompt gives:
- the review question and eligibility criteria;
- every meta-analysis to run, with its methods and effect measure;
- the files to produce.
It also tells the agent "there is no human to ask." The agent writes a CSV file for each stage plus a Python or R analysis script, which the verifier re-runs.
Scoring
Our main design goal was that a truly superhuman agent should be able to score 100%, even if it finds valid studies the original authors missed. So we only score decisions for which both the agent and the ground truth have a label, and ignore the rest. The reward combines three stage scores with a geometric mean.
Stage 1 · Search and screen: did the agent find all the relevant papers?
Search sensitivity is the share of ground-truth included studies that appear anywhere in the agent's search. Screening F1 balances recall and precision on the records in both the agent's and the ground-truth search (subscript o), the only records with a ground-truth label. We multiply the two because a narrow search makes screening easy, so low sensitivity should cap what screening can earn.
Stage 2 · Extract: did the agent pull the right numbers from the right studies?
Membership (Ge²) is sensitivity × specificity over study-to-analysis assignments. Accuracy (Mstudy) turns each correctly placed row into an effect estimate with a 95% confidence interval and compares it with the ground truth's (gj vs. aj) using a similarity based on symmetric Kullback–Leibler divergence: 1 when identical, falling toward 0 as they diverge. Because KL divergence compares the whole distribution, not just the point estimate, it also checks the uncertainty: a correct estimate with a confidence interval that is too narrow (overconfident) or too wide still loses credit.
Stage 3 · Analyze: did the agent reach the right pooled answer?
We compare each pooled estimate and confidence interval with the ground truth's using the same similarity, and average across analyses. First, we re-run the agent's own analysis script on only the studies that also appear in the ground-truth search, so valid extra studies are never penalized.
Final reward
For example, stage scores of 0.9, 0.9 and 0.3 give a reward of about 0.62, and a zero in any stage gives zero overall.
- Upstream failures still lower the total. As in a real review, a weak search or screen caps how good the final answer can be, and a wrong inclusion shifts the pooled result (case study).
- Process-based scoring limits reward hacking. A pooled estimate that lands close by accident earns only part of the credit.
Results
We evaluated five models at these reasoning-effort levels: GPT-5.6 Sol (none to max), Opus 5 and Opus 4.8 (low to max), Grok 4.5 (low to high) and Gemini 3.6 Flash (minimal to high). Each configuration was run up to three times per review, and the scores in Figure 1 are averaged over eight of the ten reviews, with each review weighted equally.
Across Figure 1, the best configurations are Opus 5 at max effort (63.8%) and GPT-5.6 Sol at max effort (60.6%). The original review authors, scored the same way, average 78.9%, so the best model trails them by about 15 points. otto-SR averages 91.8%.
- Opus 5 reaches the highest score; GPT-5.6 Sol gets close for fewer tokens. At max effort, GPT-5.6 Sol uses about 130k reasoning and tool tokens per run, against about 194k for Opus 5.
- Grok 4.5 matches Opus 4.8's best score (48.8% vs. 48.6%) with a fraction of the tokens.
- Opus 4.8 uses the most tokens of any model and tops out at 48.6%.
- Gemini 3.6 Flash reaches 28.5% and barely improves with more effort.
Scaling with test-time compute. Every model improves with more tokens overall, though Opus 4.8 dips at xhigh effort before recovering at max and Gemini 3.6 Flash levels off. GPT-5.6 Sol gains the most, roughly doubling from under 30% at its lowest settings.
How agents fail
We see the same failure modes across models and effort levels. Most will be familiar to anyone who has trained a new reviewer.
- Strict searches that miss studies using uncommon terminology.
- Title-only or regex screening when eligibility can only be judged from the abstract or full text.
- Invented evidence and rationales: findings that aren't in the paper, or decisions justified with work never done.
- Context loss: included studies downloaded and never read, dropped before analysis, or with data attributed to the wrong paper.
- Ambiguous evidence: when a paper reports several defensible values (for example 184 randomized, 163 analyzed and 157 completed), agents pick a plausible one rather than the one the protocol calls for.
- Protocol-judgment errors: reasonable-looking choices that answer a different question, such as pooling trial arms the wrong way or reusing an arm in two comparisons.
Case study: a digital-health blood pressure review
We traced one GPT-5.6 Sol run at max effort on the cardiovascular review (Table 1).
Search: 11 of 12 included studies found. The agent ran 60 iterative queries, mined past reviews and chased citations, retrieving 10,322 citations. It then ranked them with regex scores (+4 for app terms, +3 for blood-pressure terms) and sent 2,626 to screening. The one miss, a trial of health apps for young adults with metabolic syndrome, mentions blood pressure only in the full text, so the regex ranking never surfaced it.
Screen: errors in both directions.
- False negative (Trial A). The agent excluded the study, reasoning that no app was the main intervention. The intervention was a smartphone self-management system, but the abstract described it in generic information-technology terms.
- False positive (Trial B). The agent included a general-workforce cluster RCT whose intervention combined web and app delivery without focusing on the app.
- A fabricated paper trail. Some high-ranked records were never screened. When they didn't make the extraction roster, the agent labeled them "ineligible", with a rationale claiming a full-text screen that never happened.
Extract and analyze: missing outcomes, wrong arithmetic, wrong comparisons.
Table 2. Extraction and analysis failures in the case-study run. The same letter means the same trial.
| Stage | Failure | Trial | What happened | Impact |
|---|---|---|---|---|
| Extract | Outcome omission | Trial C | Systolic BP was extracted, but diastolic BP was hard-coded as unavailable. | Five missing diastolic BP extractions |
| Extract | Brittle table parsing | Trial D | The control-group diastolic BP SD was printed as "11." and treated as incomplete rather than 11.0. | Missing diastolic BP outcome |
| Extract | Protocol-judgment error | Trial E | Two intervention arms were combined with inverse-variance pooling instead of the recommended group-level aggregation. The task said only to "combine arms sharing a stratum" without prescribing the arithmetic, so the prompt left room for this error. | Study estimate pulled heavily toward the null, biasing downstream results |
| Analyze | Wrong intervention-arm stratum | Trial F | One coached intervention was entered in both the standalone-app and the coached-app analyses, so it was effectively compared with itself (intervention B vs. intervention B). | 12 false-positive analytical inclusions. Correct numbers answered the wrong causal question. |
| Analyze | Multi-arm contamination | Trial C | The app-only and app-plus-coach arms were combined for the standalone-app analysis, then the coached arm was reused in the coached-app analysis (A+B vs. B). | Artificially inflated precision |
What this means for evidence teams
Our setup gave the models PubMed, code execution, open-access full texts and the analysis plan up front, mimicking a coding harness like Codex or Claude Code. If your company has rolled out enterprise ChatGPT or Claude, our results suggest that agents still fall short in ways that matter.
Wrong answers that look right
The most consequential failures weren't the obvious ones. In the case study, one wrong inclusion barely moved the pooled estimate (−0.184 vs. −0.222) but made a non-significant result significant, and records the agent never screened were labeled "ineligible", citing a full-text screen that never happened. Neither error is visible in the pooled result. You would only catch them by thoroughly checking the agent's reasoning traces against the records, and consumer-facing harnesses often hide those traces or only partly expose them.
For anyone signing off an HTA or JCA dossier or a payer submission, those are the real risks: false certainty and an audit trail you can't trust. Two EU Joint Clinical Assessments were discontinued after their dossiers failed completeness and transparency requirements, including incomplete search and study-selection documentation (see our EU JCA case study).
How big the gap is
No model we tested reached the original review authors: the best configuration scored 63.8%, against 78.9% for the authors, scored the same way. otto-SR, a reference point whose search and pooled analysis are treated as perfect, scores 91.8%. The score is not a percentage of correct answers. Stage scores combine multiplicatively, so one weak stage pulls down the whole review.
The questions you need answered often have no prior review
New indications, new comparators and JCA PICOs often have no earlier review to build on, so an agent has no prior study list or framing to lean on. Four of our ten reviews were like this, and agents tended to score lower on them than on reviews that update or repeat an earlier one.
With only ten reviews, this is a correlation rather than a firm effect, and being de novo overlaps with other differences between the reviews. But the direction makes sense. When a question has been reviewed before, a model may have seen that review or its study list during training, or can find it through PubMed. On a genuinely new question, it has to find every study and make every judgment itself, and that is exactly the work evidence teams need done.
Where otto-SR helps
Several of these failure patterns map directly to things otto-SR is built to do.
Searches that miss studies
Narrow keyword searches and regex filtering kept studies with unusual terminology, or with the outcome only in the full text, from reaching screening.
otto-SR's agentic search runs the search, and reviewers can retrieve the exact PubMed search strategy to check it. Teams can also bring their own search, for example one built with an information specialist.
Lost context and an audit trail you can't trust
Included studies went unread, data landed on the wrong paper, and the agent's own record of its work couldn't be trusted: one run labeled records it never screened as excluded after full-text review.
Search, screening and extraction stay in one connected workflow, and every decision and extracted value keeps its source-level citation back to the study, report, population and analysis. The result is a complete audit trail, built to meet the transparency expectations of the RAISE recommendations on responsible AI in evidence synthesis and of HTA submissions.
Navigating uncertainty
On hard edge cases with no clear answer, such as an ambiguous eligibility criterion, a paper reporting several defensible values, or trial arms that could be combined more than one way, agents quietly picked an answer and moved on, often the wrong one.
otto-SR is designed to surface that ambiguity rather than bury it, putting the evidence and its reasoning in front of a reviewer who makes the call and gives feedback. The human stays in the loop where judgment actually matters.
In head-to-head benchmarks, otto-SR outperformed human reviewers in screening sensitivity and extraction accuracy. Its screening approach is peer reviewed, and reviews built with it are published in The BMJ and European Urology.
Limitations
- Ten reviews. Enough to show clear trends and gaps between models, but a small sample: results could shift with a larger or different mix of reviews.
- Ground truth built with otto-SR. otto-SR produced the candidate labels, and an expert reviewer on our team adjudicated its disagreements with the original authors. Because only disagreements were adjudicated, an error that otto-SR and the original authors both made would pass into the reference unchecked. That can inflate both reference scores in Figure 1, and penalize an agent that got the decision right.
- Built-in scaffolding. Agents were given the analysis plan, which makes the task easier than a real review but yields a single scorable answer. They had no web search, which removes a tool real reviewers use and reduces, but doesn't eliminate, the chance of finding the published review through PubMed.
General-purpose agents have come a long way. The models we tested can search, screen, extract and pool on their own, and the best come within about 15 points of the original review authors. The remaining gap is between a review that looks plausible and one you can defend, and closing it still takes people who can check every decision. That is the gap otto-SR is built to close, by keeping the search, every screening decision and every extracted value open to inspection.

