Skip to content
Quantum Fax Machine

[2601.01828] Emergent Introspective Awareness in Large Language Models

[2601.01828] Emergent Introspective Awareness in Large Language Models

This paper investigates introspective awareness in large language models by injecting known concept representations into model activations and measuring the models' ability to detect and report on these manipulations. The researchers find that advanced models like Claude Opus 4/4.1 can identify injected concepts, recall prior internal representations, and distinguish their own outputs from artificial inputs, though this capability is inconsistent and highly dependent on model architecture and training methods. The work demonstrates that current language models possess functional but unreliable introspective awareness of their internal states, with potential for further development as model capabilities improve.

⌘K

Start typing to search...

Search across content, newsletters, and subscribers