Configurable character-level transformer training suite with built-in mechanistic interpretability toolkit — scale to 150M+ parameters and beyond, no ceilings, only hardware limits. Inspect attention weights, hidden states, and head specialisation across all layers. Documented circuit findings included.
nlpdeep-learningpytorchtransformergptlanguage-modelattention-mechanismcircuit-analysisinterpretabilitycharacter-level-language-modelattention-visualizationattention-headstransformer-interpretabilitymechanistic-interpretabilityresidual-streamhidden-state-analysis
-
Updated
Jul 4, 2026 - Python