Interactive explainer: BDH's attention has no softmax, so it is exactly a Hebbian synaptic memory — one fixed-size matrix written once per token. Run both forms, watch them agree to 1e-16, then break it with one toggle. DataForge 2026, Pathway track.
machine-learning attention-mechanism hebbian-learning associative-memory pathway fast-weights interpretability explorable-explanations neurips linear-attention bdh dragon-hatchling interactive-explainer post-transformer sparse-activations
-
Updated
Sep 8, 2026 - HTML