[2601.01828] Emergent Introspective Awareness in Large Language Models
![]()
This paper investigates introspective awareness in large language models by injecting known concept representations into model activations and measuring the models' ability to detect and report on these manipulations. The researchers find that advanced models like Claude Opus 4/4.1 can identify injected concepts, recall prior internal representations, and distinguish their own outputs from artificial inputs, though this capability is inconsistent and highly dependent on model architecture and training methods. The work demonstrates that current language models possess functional but unreliable introspective awareness of their internal states, with potential for further development as model capabilities improve.
Was this useful?