MKAI ยท Working paper

The Same-Model Ceiling

Testing multi-model AI review in strategic analysis

Published August 2026

Abstract

This study examines whether three reviewers from different AI model families identify decision-relevant considerations beyond those raised by three fresh reviews from the model that produced the original analysis. It then tests whether incorporating those additional considerations produces analyses preferred under blind assessment.

Claude Opus 5 completed twelve strategic-analysis tasks based on one constructed acquisition case. Each answer received six critiques under a shared brief: three from fresh Claude Opus 5 sessions and three from GPT-5.6 Sol, Muse Spark 1.2 and DeepSeek V4 Pro. Eligible critique points were stripped of source labels and assigned shuffled identifiers. Two further models grouped equivalent points independently, with a third model resolving disagreements. Claude assessed the decision effect of considerations unique to the multi-model panel and revised each analysis where qualifying additions appeared. Two model scorers then compared each original and revised analysis under blind labels, with adjudication where their preferences differed.

The reviewers produced 216 candidate considerations. The eligibility screen retained 152, comprising 74 from same-model reviews and 78 from multi-model reviews. Blinded grouping produced 80 groups of equivalent considerations: 25 were shared across both conditions, 18 were unique to the same-model condition and 37 were unique to the multi-model condition. Every task contained at least one consideration unique to the multi-model panel. Claude rated 35 of these 37 at Level 2, indicating a change to proposed terms, evidence requirements or commitment conditions; the remaining two received Level 1 ratings. Together, qualifying additions changed the stated decision route in three revised analyses. Blind assessment preferred the revision in eight of twelve tasks and the original in one, while three comparisons ended in ties.

The evidence concerns one constructed case, one originating model and one fixed review panel. Within that design, changing reviewer identity expanded analytical coverage while reviewer count and instructions remained constant. The extent of the effect across other models, panels and decision settings remains open.

Suggested citation

Foster-Fletcher, R. (2026). The Same-Model Ceiling: Testing Multi-Model AI Review in Strategic Analysis. MKAI Inquiry Working Paper.

Download PDF

PDF size: 484 KB

The Separation Study varied the written instruction around one fixed model. This study fixes the instruction and varies the reviewing model. The Separation Study

Record metadata

Type
Inquiry working paper
Theme
Review arrangement and model diversity
Status
Published

Methods

  • Matched two-condition review design on one constructed acquisition case (12 tasks, 6 critiques per answer, 216 candidate considerations, 152 retained by the eligibility screen)
  • Blinded grouping of equivalent considerations by two independent models with a third resolving disagreements, producing 80 groups of which 37 were unique to the multi-model condition
  • Decision-effect rating by the originating model, revision where qualifying additions appeared, and blind paired comparison by two model scorers with adjudication on disagreement

Keywords

same-model reviewmulti-model reviewmodel diversitylarge language modelsblind assessmentstrategic analysisenterprise AImodel evaluation