Reproducible toy-model experiments in mechanistic interpretability. Includes a test of whether mixed (superposed) coding is what lets a linear readout see feature interactions — headline result plus a reported null.
pytorchcomputational-neurosciencesuperpositionai-safetyinterpretabilitytoy-modelsneuroaimechanistic-interpretabilitypolysemanticitymixed-selectivity
-
Updated
Aug 9, 2026 - Python