Tag
1 article
A new study finds that AI models' written reasoning steps correspond to distinct internal patterns, especially in middle layers, with important implications for AI safety and interpretability.