Large Language Models Predict Human Social Behavior via Interpretable Mechanisms
The development of large language models (LLMs) offers promising opportunities for predicting human behavior across diverse contexts. However, most prior work has emphasized behavioral imitation, with limited attention to transparent or interpretable models of the cognitive mechanisms underlying human decisions. In this study, we introduce MindEvolve, an autonomous workflow designed to predict behavior in social interactions by generating interpretable symbolic models of cognition. We systematically evaluate the modeling capabilities of multiple LLMs across a battery of socioeconomic games covering four core domains of social cognition: economic preferences, social preferences, social reasoning (theory of mind), and recursive planning. The symbolic models proposed by the LLMs are assessed by expert human evaluators for interpretability and theoretical coherence. Our results show that while most LLMs robustly capture economic and social preferences in relatively simple strategic settings, and in some cases generate novel models that integrate broader knowledge than those proposed by human experts, their capacity to model more complex psychological processes remains limited. However, a subset of state-of-the-art models demonstrates promising performance in capturing higher-order reasoning processes such as theory of mind and recursive planning. Together, these findings highlight the emerging potential of LLMs not only to imitate behavior, but to generate interpretable, mechanistic accounts of human social decision-making. Our results provide a roadmap for advancing LLM-based cognitive modeling toward human-expert-level theory construction.