The Same-Model Ceiling
Testing multi-model AI review in strategic analysis
Published August 2026
Abstract
This study examines whether three reviewers from different AI model families identify decision-relevant considerations beyond those raised by three fresh reviews from the model that produced the original analysis. It then tests whether incorporating those additional considerations produces analyses preferred under blind assessment.
Claude Opus 5 completed twelve strategic-analysis tasks based on one constructed acquisition case. Each answer received six critiques under a shared brief: three from fresh Claude Opus 5 sessions and three from GPT-5.6 Sol, Muse Spark 1.2 and DeepSeek V4 Pro. Eligible critique points were stripped of source labels and assigned shuffled identifiers. Two further models grouped equivalent points independently, with a third model resolving disagreements. Claude assessed the decision effect of considerations unique to the multi-model panel and revised each analysis where qualifying additions appeared. Two model scorers then compared each original and revised analysis under blind labels, with adjudication where their preferences differed.
The reviewers produced 216 candidate considerations. The eligibility screen retained 152, comprising 74 from same-model reviews and 78 from multi-model reviews. Blinded grouping produced 80 groups of equivalent considerations: 25 were shared across both conditions, 18 were unique to the same-model condition and 37 were unique to the multi-model condition. Every task contained at least one consideration unique to the multi-model panel. Claude rated 35 of these 37 at Level 2, indicating a change to proposed terms, evidence requirements or commitment conditions; the remaining two received Level 1 ratings. Together, qualifying additions changed the stated decision route in three revised analyses. Blind assessment preferred the revision in eight of twelve tasks and the original in one, while three comparisons ended in ties.
The evidence concerns one constructed case, one originating model and one fixed review panel. Within that design, changing reviewer identity expanded analytical coverage while reviewer count and instructions remained constant. The extent of the effect across other models, panels and decision settings remains open.
Suggested citation
Foster-Fletcher, R. (2026). The Same-Model Ceiling: Testing Multi-Model AI Review in Strategic Analysis. MKAI Inquiry Working Paper.
PDF size: 484 KB
The Separation Study varied the written instruction around one fixed model. This study fixes the instruction and varies the reviewing model. The Separation Study
Record metadata
- Type
- Inquiry working paper
- Theme
- Review arrangement and model diversity
- Status
- Published
Methods
- Matched two-condition review design on one constructed acquisition case (12 tasks, 6 critiques per answer, 216 candidate considerations, 152 retained by the eligibility screen)
- Blinded grouping of equivalent considerations by two independent models with a third resolving disagreements, producing 80 groups of which 37 were unique to the multi-model condition
- Decision-effect rating by the originating model, revision where qualifying additions appeared, and blind paired comparison by two model scorers with adjudication on disagreement