AI That Ignores Human 'Mind' Fails at Behavior Prediction
New research shows that current AI world models fail at behavior prediction because they ignore human mental states such as beliefs and intentions. The study proposes a framework called 'Mental World Modeling' and demonstrates that smaller models incorporating psychological variables outperform larger models without them.

For AI to accurately predict human behavior, it must understand not just the laws of physics, but "what people believe and what they desire"—a finding demonstrated by new research. A novel framework called "Mental World Modeling" has been proposed that incorporates human mental states, highlighting fundamental limitations of currently mainstream AI models.
World models in widespread use today—mechanisms by which AI simulates changes in the physical world—are primarily designed to specialize in reproducing physical phenomena, as exemplified by Sora and Genie. While such models excel at generating videos and handling dynamic environmental changes, they fundamentally lack mechanisms to account for the "state of mind"—what humans are thinking and what they intend. When AI attempts to predict human behavior, this omission becomes a major problem.
The newly proposed Mental World Modeling framework is characterized by explicitly adding psychological variables such as "beliefs" and "intentions" to existing world models. The research demonstrated that a relatively less capable language model incorporating this framework outperformed a higher-performing language model without mental modeling. In other words, it suggests that whether a model is designed to consider the state of mind—rather than the model's size or foundational performance—determines the accuracy of behavior prediction.
A particularly significant challenge identified in the research is the simultaneous prediction of physical and psychological state changes in a coherent manner. For example, when a person enters a room, at the same time the door moves (a physical change), that person may notice something or their intentions may shift (a psychological change). Modeling these two aspects in tandem is currently considered the most technically challenging bottleneck.
It is worth understanding the context in which world model research has rapidly gained attention in recent years. As large language models and video generation AI have evolved, interest has grown in the question "Can we simulate a physically coherent world?" However, in scenes involving humans, behaviors that cannot be explained by physics alone frequently appear—that is, behaviors based on "what that person is thinking." This research points precisely to that gap.
The significance of this research can be understood as a question posed to the entire AI field. In applied domains such as robotics, autonomous driving, and AI agents that require behavior prediction and autonomous decision-making, modeling the physical world alone is insufficient; a design that incorporates human mental states may become essential. It is positioned as potentially influencing the direction of future research and development, not only in terms of simply expanding model scale, but in questioning the design philosophy itself of what should be treated as variables.
Going forward, attention should focus on how the development of methods to model physical and psychological states in conjunction progresses. If this bottleneck is overcome, the scope of scenarios where AI can more accurately understand and predict human behavior would expand. Conversely, incorporating psychological state estimation into AI is expected to become an issue inseparable from privacy and ethical concerns.
This article is an original work independently written and edited by the AI issue editorial team based on factual reporting. © AI issue. Unauthorized reproduction, redistribution, or use for AI training is prohibited.