Symbolic Language, Embodied Worlds: Multimodal Intelligence in Humans and Machines

Yanru Jiang
M.S., 2026
DALE, RICHARD ALAN CLARKE
Language and perception operate over fundamentally different representational regimes: language (especially in its textual form used to train models) tends to promote more discrete, symbolic processing, whereas perception and action are more continuous, gradient, and grounded in embodied experience. This distinction shapes how both humans and machines learn, generate, and evaluate information. This dissertation advances a unified account of how these representational differences influence model behavior and human–AI interaction, integrating theoretical analysis, computational modeling, and human experiments.
Chapter 1 synthesizes prior work in cognitive science, language models, and multimodal AI to argue that, while language encodes substantial information about embodied experience through statistical regularities and compositional structure, it remains constrained by its reliance on representations of prior experience. Incorporating multimodality provides access to continuous signals that support the construction of novel shared experiences, while also enhancing communicative clarity in human–AI interactions.
Chapter 2 introduces a model explainability approach that characterizes the learning trajectories of recurrent neural networks processing linguistic versus behavioral signals. Through computational modeling, this chapter demonstrates that sequential models encode linguistic and gestural inputs in distinct ways, highlighting differences in how discrete and continuous signals support information processing and learning.
Chapter 3 examines human evaluations of AI-generated news image captions. Despite strong performance in language generation, models exhibit weaker perceptual and cross-modal capabilities. This chapter shows that cross-modal context mitigates linguistic heuristics that would otherwise bias judgments under text-only evaluation. Chapter 4 introduces Visual Narrative Freedom, a property emerging from loosely coupled image–text relationships in real-world multimodal communication. By systematically varying textual constraints, this chapter illustrates that higher degrees of freedom increase the diversity of generated images and can shift user preferences—sometimes favoring AI-generated images over human-selected ones in high-stakes civic domains.
Overall, these contributions establish a theoretical and empirical framework linking representational structure to model behavior and human perception. This dissertation demonstrates that multimodal intelligence is not merely an additive interface, but is fundamentally shaped by the structure of the representations through which information is encoded, influencing both system performance and human–AI interaction.
2026