Training-time defense that redistributes LLM refusal via mean/covariance matching + KD, raising linear-ablation attack rank from K=1 to K≥16 (Llama-3.2-1B-Instruct)
pytorchllamaai-alignmentadversarial-robustnessmechanistic-interpretabilityllm-safetyrefusalrepresentation-engineeringleace
-
Updated
Jun 12, 2026 - Python