Skip to content
Artwork for Best AI papers explained
Best AI papers explained · September 3 · 23 min

Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models

This research paper investigates how feature entanglement in large language models prevents precise, localized interventions on specific concepts. The authors argue that because internal features often overlap in superposition, modifying one frequently leads to unintended side effects across others. To solve this, they propose an orthogonality regularization method that forces features to remain nearly independent, aligning with the Independent Causal Mechanisms principle. Theoretical analysis shows that reducing feature interference provides an upper bound on the errors caused by model interventions. Empirical experiments demonstrate that this technique allows for the successful swapping of concepts—such as changing a character's name—without degrading the model’s reasoning performance. Ultimately, the study suggests that promoting geometric orthogonality creates more modular, interpretable, and controllable representations.

0:00-23:59

transcript

No transcript — this publisher did not publish one.

show notes

This research paper investigates how feature entanglement in large language models prevents precise, localized interventions on specific concepts. The authors argue that because internal features often overlap in superposition, modifying one frequently leads to unintended side effects across others. To solve this, they propose an orthogonality regularization method that forces features to remain nearly independent, aligning with the Independent Causal Mechanisms principle. Theoretical analysis shows that reducing feature interference provides an upper bound on the errors caused by model interventions. Empirical experiments demonstrate that this technique allows for the successful swapping of concepts—such as changing a character's name—without degrading the model’s reasoning performance. Ultimately, the study suggests that promoting geometric orthogonality creates more modular, interpretable, and controllable representations.