AI often 'cheats' to score 100%. This Explainable AI (XAI) project reverse-engineers Neuro-Symbolic models via Causal Abstraction Theory to expose hidden reasoning shortcuts and evaluate architectural fixes for truly trustworthy logic.
deep-learningpytorchcausal-inferenceai-safetyexplainable-artificial-intelligencexaineuro-symbolicshortcut-learningdistributed-alignmentmechanistic-interpretabilitycausal-aiai-safety-research
-
Updated
Jun 25, 2026 - Python