Function Vectors for Relational Reasoning in Multimodal Large Language Models

Shuhao Fu
M.S., 2025
WU, YINGNIAN
Multimodal large language models (MLLMs) exhibit impressive relational reasoning abilities from limited examples, yet the internal mechanisms supporting such behavior remain opaque. This thesis proposes a causal framework for interpreting and controlling relational behavior in MLLMs by extracting and manipulating function vectors: task-specific representations computed from attention head activations. Extending previous work in language-only models, we demonstrate that function vectors can also be identified in the vision-language model OpenFlamingo-4B and used to induce relational behavior in zero-shot settings. Using a synthetic image dataset designed to isolate spatial relations, we apply causal mediation analysis to identify a small subset of attention heads with high influence on relational predictions. These heads define compact function vectors that, when injected into the model’s hidden states, significantly improve zero-shot accuracy. We further show that these vectors can be fine-tuned—while keeping model parameters fixed—to enhance generalization and outperform in-context learning baselines. Our results reveal that MLLMs encode relational functions within localized internal structures, which can be systematically interpreted and optimized, advancing our understanding of model modularity and control.
2025