Why AI Keeps Misreading Your Instructions — New Research Explains
Scientists say goals work like language — and that could make AI finally get what you mean.
In a paper to appear in Topics in Cognitive Science, David M. Abel and Mark K. Ho draw attention to goals as representations and their content. In both cognitive science and computer science, they note, goals are conceptualized as cognitive states that flexibly combine with world knowledge to organize and specify purposeful behavior — making goals compositional representations whose content relates to rational behavior. That framing, the authors write, highlights a parallel with other areas of cognitive science, in particular the syntax-semantics interface in linguistics and logic, while foregrounding foundational questions about the expressivity, design, and efficiency of different goal representations. Goals are typically taken as fixed and as imposing constraints on desirable behavior, but the authors point to constraints on goal representations themselves — for example, whether a particular goal language is expressive enough to capture behaviors of interest, or whether different goal representations capture the same behavior. They synthesize work that aims to characterize the properties of different goal representations, suggesting these are points in a broader design space. They close by discussing how distinguishing the form and meaning of goals can elucidate the implicit assumptions we make about goals, inform the study of interactions between higher-level cognition and motivation, and isolate axes of variation for different conceptions of goals.
- Goals work like sentences: they have a structure and a meaning, and researchers say you can study the two separately.
- Authors David Abel and Mark Ho argue AI designers should ask whether a goal format is even expressive enough for the job at hand.
- No new product — this is a framework paper in Topics in Cognitive Science, aimed at other researchers, not users.
Why It Matters
Clearer theories of goals could mean AI assistants that actually do what you asked — and fail more predictably.