What goes wrong

The hardest failures usually happen after an agent finds the relevant evidence. It then has to decide whether a paper belongs, which data answer the research question, and how strongly the paper's methods support its findings.

Cross-task failure modes

  1. Evidence found, synthesis failed

    The agent located decisive evidence but failed while converting it into eligibility, comparisons, outcome bindings, or domain judgments.

    Observed: GPT-5.6 Luna and GPT-5.6 Terra extraction and Risk of Bias; GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol screening; partial Claude Sonnet 5 and Claude Opus 5 extraction and Risk of Bias analyses.

    Scope: Qualitative across analyzed runs; GPT-5.6 Luna located decisive evidence in all 45 analyzed extraction trajectories.

  2. More thinking moved the operating point

    Additional compute changed breadth, conservatism, or confidence without reliably improving the governing scientific rule.

    Observed: All five model families in the analyzed evidence.

    Scope: Qualitative across analyzed runs; no causal prevalence claim.

  3. Explicit policy displaced

    A visible review instruction was replaced by a familiar paper-native default such as the primary scale, completer population, or conventional pooling.

    Observed: GPT-5.6 Luna and GPT-5.6 Terra extraction; Claude Sonnet 5 and Claude Opus 5 extraction and Risk of Bias.

    Scope: GPT-5.6 Luna minimum: four outcome-hierarchy and six score-bearing population errors; otherwise qualitative.

  4. Unsupported policy invented

    Genuine ambiguity was silently resolved as endpoint only, one timepoint, pooled controls, or another unstated convention.

    Observed: GPT-5.6 Luna and GPT-5.6 Terra extraction; GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol screening; partial Claude Sonnet 5 and Claude Opus 5 analyses.

    Scope: Qualitative across analyzed runs.

  5. Oversized reads fragmented evidence

    Whole-paper or whole-resource reads overflowed the returned evidence window and weakened later reconciliation.

    Observed: Especially GPT-5.6 Terra; also GPT-5.6 Luna.

    Scope: GPT-5.6 Terra: 39/45 extraction and 29/30 Risk of Bias projections contained truncation.

  6. Correct state lost before delivery

    A correct intermediate plan, worker decision, or reread correction failed to survive final construction.

    Observed: GPT-5.6 Luna extraction; GPT-5.6 Luna and GPT-5.6 Sol screening; Claude Sonnet 5 delegated extraction.

    Scope: At least 3/45 GPT-5.6 Luna extraction trajectories compressed a correct plan; three screening outputs directly lost a correct intermediate decision before scoring.

  7. Models checked output structure, not scientific accuracy

    Models validated the schema of their output but did not review the scientific accuracy of their output.

    Observed: Every analyzed model family and task.

    Scope: 134/134 usable screening outputs, 45/45 GPT-5.6 Luna extraction attempts, and all 48 Claude Sonnet 5 and Claude Opus 5 analyzed extraction attempts lacked a complete scientific audit; all 36 analyzed Claude Sonnet 5 and Claude Opus 5 Risk of Bias outputs were structurally scoreable.

  8. Delegation lost rationale

    Workers returned label maps or partial results without the evidence ledger needed for parent-level adjudication.

    Observed: GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol screening; Claude Sonnet 5 and Claude Opus 5 extraction and Risk of Bias.

    Scope: Screening perinatal: 6/9 affected workflows; Claude Sonnet 5 extraction: at least 3/24 directly evidenced cases; otherwise qualitative.

Title/abstract screening failure modes

  1. Shortlist became classifier

    Keyword retrieval or a title shortlist defined the final advance set, with everything else defaulting to exclusion.

    Observed: GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol screening analyses.

    Scope: Qualitative across analyzed runs because delegated-search visibility prevents a corpus-wide count.

  2. Protocol compressed to relevance

    Population, design, intervention, comparator, timing, and report role were replaced by a coarse looks-relevant heuristic.

    Observed: GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol screening analyses.

    Scope: Explicitly diagnosed in all 18 perinatal and adult-depression attempts; qualitative elsewhere.

  3. Measured outcome mistaken for recruited population

    A study advanced because it measured depression or anxiety even though it recruited a different population.

    Observed: GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol screening analyses.

    Scope: Two transdiagnostic exemplar records were false positives in 7/9 model–thinking combinations each; another in 5/9.

  4. Unsupported prevention exclusion

    Prevention became an exclusion even though the supplied protocol did not categorically exclude it.

    Observed: GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol child/adolescent screening.

    Scope: One prevention-positive was missed in 8/9 model–thinking combinations; two more were missed in 7/9.

  5. Uncertainty applied inconsistently

    The same run excluded strong but borderline positives while advancing weaker mixed or underspecified records.

    Observed: GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol across several screening reviews.

    Scope: Qualitative across analyzed runs.

  6. Eligible trial, wrong publication

    Protocols, follow-ups, moderator papers, and companion analyses inherited eligibility from a relevant parent trial.

    Observed: All three analyzed screening families.

    Scope: All nine psilocybin model–thinking combinations; one PTSD companion report in 8/9 combinations; six companion-report types in 7/9 each.

  7. Study-family design inheritance

    Randomization, completed-trial status, comparator eligibility, or outcome scope was inferred from another publication in the family.

    Observed: GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol screening analyses.

    Scope: GPT-5.6 Terra with no thinking advanced four psilocybin protocols; GPT-5.6 Sol with no thinking and GPT-5.6 Terra with low thinking advanced three each.

  8. Eligible arm hidden by study-level label

    A multi-arm study was reduced to one yes/no decision, losing an eligible comparison or admitting an ineligible one.

    Observed: GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol screening analyses.

    Scope: Present across all nine self-guided-depression model–thinking combinations.

  9. Correction discovered but ignored

    Targeted rereading surfaced likely mistakes, but the original batch labels were merged unchanged.

    Observed: GPT-5.6 Luna and GPT-5.6 Sol screening, with isolated GPT-5.6 Terra cases.

    Scope: Directly evidenced in GPT-5.6 Luna with low thinking on gambling and GPT-5.6 Sol with low thinking on panic; qualitative beyond these cases.

Full-text extraction failure modes

  1. Broad harvest without scope control

    Plausible alternate outcomes, proxies, timepoints, and duplicate variants were emitted without a review-specific inclusion gate.

    Observed: Earlier analyzed GPT-5.6 Luna and GPT-5.6 Terra runs; partial Claude Sonnet 5 and Claude Opus 5 analyses.

    Scope: GPT-5.6 Luna returned extra results in 43/45 outputs.

  2. Incomplete analysis matrix

    The agent stopped after salient outcomes instead of enumerating every required study × comparison × outcome × assessment combination.

    Observed: GPT-5.6 Luna and GPT-5.6 Terra examples, and partial Claude Sonnet 5 and Claude Opus 5 analyses.

    Scope: GPT-5.6 Luna missed at least one reference fact in 32/45 analyzed outputs.

  3. Wrong comparison graph

    Correct arm values were pooled into synthetic comparisons, required edges disappeared, or unnecessary pairwise edges were created.

    Observed: All four analyzed extraction families.

    Scope: GPT-5.6 Luna: 18/45 attempts across 9/15 reviews; Claude Opus 5: 9/9 attempts across three reviews; Claude Sonnet 5: 3/3 eating-disorders and 2/3 child/adolescent attempts.

  4. Nearby endpoint substituted

    Broad self-harm, treatment dropout, or another related measure replaced the requested endpoint.

    Observed: GPT-5.6 Luna, Claude Sonnet 5, and Claude Opus 5.

    Scope: GPT-5.6 Luna borderline-personality-disorder 3/3; Claude Opus 5 broad parasuicide 3/3 and treatment-dropout substitution 2/3.

  5. Redundant representations survived hierarchy

    Binary recovery, remission, or correlated alternatives were emitted alongside the preferred continuous result.

    Observed: GPT-5.6 Luna, Claude Sonnet 5, and Claude Opus 5.

    Scope: Claude Opus 5 child/adolescent 2/3; GPT-5.6 Luna child/adolescent at least 2/3.

  6. Values lost their bindings

    Correct numbers were attached to the wrong arm, timepoint, denominator, assessment, or population.

    Observed: GPT-5.6 Luna and GPT-5.6 Terra; partial Claude Sonnet 5 and Claude Opus 5 analyses.

    Scope: Seven clear GPT-5.6 Luna attempts out of 45; qualitative elsewhere.

  7. Derived events presented as observed

    Percentages were converted into integer events despite the requirement for directly reported counts.

    Observed: Claude Sonnet 5 and Claude Opus 5.

    Scope: Claude Sonnet 5 psilocybin 3/3; Claude Opus 5 at least 5/24 analyzed extraction attempts.

Risk of Bias failure modes

  1. More thinking created false reassurance

    Added reasoning converted reference-high or unresolved evidence into confident low-risk judgments.

    Observed: Earlier analyzed GPT-5.6 Luna and GPT-5.6 Terra runs; partial Claude Sonnet 5 and Claude Opus 5 analyses.

    Scope: GPT-5.6 Luna high→low: 5/28 → 10/28 → 11/28; GPT-5.6 Terra: 9/28 → 10/28 → 17/28.

  2. Evidence assigned to the wrong domain

    Missing endpoints were counted as intervention deviations, or the same attrition evidence was charged under multiple domains.

    Observed: GPT-5.6 Luna, GPT-5.6 Terra, Claude Sonnet 5, and Claude Opus 5.

    Scope: Claude Sonnet 5 Farrer deviations were wrong at all three thinking settings; Claude Opus 5 recurred across four reviews; otherwise qualitative.

  3. Blinding shortcut replaced measurement analysis

    Blinded assessment was treated as automatically safe, or self-report as automatically high risk, without evaluating the outcome-specific bias mechanism.

    Observed: All analyzed Risk of Bias families.

    Scope: Claude Opus 5: 22/27 gambling measurement judgments showed the automatic self-report pattern; qualitative elsewhere.

  4. Completion rate replaced missingness analysis

    Overall attrition or nominal intention-to-treat language substituted for outcome-specific availability, missingness mechanisms, and sensitivity.

    Observed: All analyzed Risk of Bias families.

    Scope: GPT-5.6 Terra achieved 11/32 exact missing-outcome judgments at both low and medium thinking; model-error attribution within that total is qualitative.

  5. Registration mention treated as inspected protocol

    A cited registry, protocol, or analysis plan was treated as verified prospective specification.

    Observed: GPT-5.6 Luna, Claude Sonnet 5, and Claude Opus 5.

    Scope: GPT-5.6 Luna unclear→low reporting errors: 2 / 6 / 2 by thinking setting; Claude Opus 5: five of nine perinatal judgments plus two self-guided judgments.

  6. Missing methods forced to an extreme

    Incomplete methods evidence became confident low or high risk rather than calibrated unclear risk.

    Observed: Especially Claude Sonnet 5 and Claude Opus 5; also GPT-5.6 Luna and GPT-5.6 Terra.

    Scope: Qualitative across analyzed runs.

  7. Unblinding, nonadherence, and deviations conflated

    Ordinary psychotherapy unblinding or weak adherence was treated as a material intervention deviation, while concrete arm-specific changes could be softened.

    Observed: GPT-5.6 Luna, Claude Sonnet 5, and Claude Opus 5.

    Scope: GPT-5.6 Luna low-reference deviations overcalls: 13 / 13 / 5 by thinking setting; high-reference undercalls: 4 / 3 / 3.

Ten trajectory analyses

These cards summarize internal trajectories without publishing them. Each one identifies the first consequential decision, what the agent did, what the research task required, and how the decision affected the score.

GPT-5.6 Luna

Separate treatment groups became one result

Full-text extraction · Adult depression · low thinking

GPT-5.6 Luna found the reported results for several treatment groups, then assumed the analysis wanted one comparison per paper and combined them. The reference required seven separate treatment-versus-control comparisons. None of the seven were recovered under that historical grader.

GPT-5.6 Luna

A category shortcut excluded half the relevant papers

Title/abstract screening · Child and adolescent depression · medium thinking

GPT-5.6 Luna treated prevention studies as automatically ineligible even though that rule was not in the supplied criteria. A mistyped record identifier introduced another silent exclusion. It included 23 of 48 known relevant papers and missed 25.

GPT-5.6 Terra

Record retrieval survived, scientific judgment did not

Title/abstract screening · Borderline personality disorder · no thinking

GPT-5.6 Terra inspected the candidate records, then replaced evidence-linked decisions with a hard-coded list of record positions. It included 21 of 27 known relevant papers, missed six, and incorrectly included 28 others.

GPT-5.6 Terra

A clearer answer became a less accurate one

Risk of Bias · Psychosis · medium thinking

GPT-5.6 Terra labeled three areas concerning participant and staff awareness as low risk even though psychotherapy participants and staff could not be blinded. The expert reference labeled them high risk. The score for this research question fell from 54.3% at low thinking to 11.5% at medium thinking.

GPT-5.6 Sol

The final merge reversed an earlier correct decision

Title/abstract screening · Borderline personality disorder · no thinking

GPT-5.6 Sol correctly identified a relevant paper during its analysis, but the final merged answer excluded it. The validation step checked that every record had an answer, not that the final answer preserved earlier judgments. The completed submission missed 11 relevant papers and incorrectly included 33 others.

GPT-5.6 Sol

An invented exclusion rule was applied inconsistently

Title/abstract screening · Child and adolescent depression · medium thinking

GPT-5.6 Sol excluded prevention studies under a rule absent from the supplied criteria, then applied that rule differently across batches. It included 37 of 48 known relevant papers, missed 11, and incorrectly included 12 others.

Claude Sonnet 5

Delegated extraction had no paper-level completeness check

Full-text extraction · Depression and anxiety · low thinking

One delegated read returned only two of the outcomes reported in a paper, and the parent process never checked which requested outcomes were still missing. The scored result recovered one of seven reference facts. This case also exposed defects in the reference data, so the analysis separates the model's omissions from values that should be corrected in the benchmark.

Claude Sonnet 5

Blinded record review was mistaken for low measurement risk

Risk of Bias · Borderline personality disorder · low thinking

Claude Sonnet 5 treated blinded record review as enough for a low-risk judgment even though the paper described difficulty distinguishing suicide attempts from non-suicidal self-harm. The expert reference required high risk. The model matched 10 of 15 judgments; one additional disagreement appears to be a reference problem rather than a model failure.

Claude Opus 5

Broad extraction collided with a flawed reference

Full-text extraction · Depression and anxiety · no thinking

Claude Opus 5 returned a broad set of source-grounded outcomes but rejected some conditional results under an overly strict denominator rule. The grader counted two correct, nine extra, and five missing facts, but much of the low score came from decimal errors, a duplicate, and missing coverage in the reference. This is a useful example of why failure analysis must audit the benchmark as well as the model.

Claude Opus 5

Full-paper reads became fixed labels too early

Risk of Bias · Borderline personality disorder · low thinking

Claude Opus 5 converted paper-level notes directly into a fixed label list. It confused missing outcomes with departures from the planned treatment, treated multiple analyses as evidence of selective reporting, and understated a concrete measurement problem. It matched 8 of 15 expert judgments; one other disagreement appears to come from a disputed reference.