Can Large Language Model Multi-Agent Simulations Reproduce Expert Interdisciplinary Discussions? A Four-Layer Fidelity Framework with Behavior-Anchored Agents

Sungjin Choi
MASDS, 2026
WU, YINGNIAN
Controlled experiments on how interdisciplinary research teams discuss and converge are difficult to run. An interdisciplinary cohort takes years to assemble, and once a team finishes its work, the same discussion cannot be replayed under altered conditions. Large language model (LLM) multi-agent simulation has been proposed as a way to bypass this constraint, but only if simulated discussions can serve as a limited stand-in for real human teams. This thesis tests that condition using five interdisciplinary teams from an annual mobile-health research training program, with about 24,000 utterances across the five teams and six sessions per team. The thesis makes two contributions. First, it proposes a behavior-anchored agent setup that injects observed participation and functional tendencies into agent personas, in contrast to role-only personas. The Anchor condition consistently moves simulated outcomes closer to the human teams than the Naive condition. Second, it introduces a four-layer evaluation framework, transcript-level process metrics, an LLM process judge, an LLM outcome judge, and statistical inference, which scores process and outcome separately. Under this framework, process and outcome layers do not always move together: some simulations score lower on process yet higher on outcome, a pattern this thesis terms process-outcome divergence. The experiments produce three findings. At the transcript surface, anchoring moves the simulation closer to the human teams on a majority of metrics, but not uniformly. At the judge layers, anchoring closes roughly 39% of the process gap and 64% of the outcome gap between Naive and Human under matched-length evaluation. A coverage asymmetry in the original fixed-length setup created an apparent process advantage that disappeared once the judge read comparable portions of each transcript. Step 2 produced the sharpest case of the divergence: the no-broadcast condition produced the lowest process scores but the highest outcome scores. Together, these results argue that simulation as a research instrument depends on measuring fidelity layer by layer, with both the simulation side and the judge side validated separately.
2026