Shortlist became classifier
Keyword retrieval or a title shortlist defined the final advance set, with everything else defaulting to exclusion.
Observed: GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol screening analyses.
Scope: Qualitative across analyzed runs because delegated-search visibility prevents a corpus-wide count.
Protocol compressed to relevance
Population, design, intervention, comparator, timing, and report role were replaced by a coarse looks-relevant heuristic.
Observed: GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol screening analyses.
Scope: Explicitly diagnosed in all 18 perinatal and adult-depression attempts; qualitative elsewhere.
Measured outcome mistaken for recruited population
A study advanced because it measured depression or anxiety even though it recruited a different population.
Observed: GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol screening analyses.
Scope: Two transdiagnostic exemplar records were false positives in 7/9 model–thinking combinations each; another in 5/9.
Unsupported prevention exclusion
Prevention became an exclusion even though the supplied protocol did not categorically exclude it.
Observed: GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol child/adolescent screening.
Scope: One prevention-positive was missed in 8/9 model–thinking combinations; two more were missed in 7/9.
Uncertainty applied inconsistently
The same run excluded strong but borderline positives while advancing weaker mixed or underspecified records.
Observed: GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol across several screening reviews.
Scope: Qualitative across analyzed runs.
Eligible trial, wrong publication
Protocols, follow-ups, moderator papers, and companion analyses inherited eligibility from a relevant parent trial.
Observed: All three analyzed screening families.
Scope: All nine psilocybin model–thinking combinations; one PTSD companion report in 8/9 combinations; six companion-report types in 7/9 each.
Study-family design inheritance
Randomization, completed-trial status, comparator eligibility, or outcome scope was inferred from another publication in the family.
Observed: GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol screening analyses.
Scope: GPT-5.6 Terra with no thinking advanced four psilocybin protocols; GPT-5.6 Sol with no thinking and GPT-5.6 Terra with low thinking advanced three each.
Eligible arm hidden by study-level label
A multi-arm study was reduced to one yes/no decision, losing an eligible comparison or admitting an ineligible one.
Observed: GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol screening analyses.
Scope: Present across all nine self-guided-depression model–thinking combinations.
Correction discovered but ignored
Targeted rereading surfaced likely mistakes, but the original batch labels were merged unchanged.
Observed: GPT-5.6 Luna and GPT-5.6 Sol screening, with isolated GPT-5.6 Terra cases.
Scope: Directly evidenced in GPT-5.6 Luna with low thinking on gambling and GPT-5.6 Sol with low thinking on panic; qualitative beyond these cases.