Pivotal Token Search
-
Updated
Aug 5, 2026 - Python
Pivotal Token Search
Adversarial Manipulation of CoT
Analysed determinism, faithfulness, reasoning patterns, & steering. Developed and tested methods to enhance control and fail-safes
mech-interp suite for Granite4 models that use Mamba-2 architecture
Detect safety degradation during LLM fine-tuning before it becomes behavioral
Implementation and analysis of Sparse Autoencoders for neural network interpretability research. Features interactive visualization dashboard and W&B integration.
All code, stimuli, and results for a mechanistic interpretability study investigating how large language models internally represent emotional content
Collection and learnings of my journey in Artificial Intelligence
Official implementation of the 'Uncovering Competency Gaps in Large Language Models and Their Benchmarks' paper
Investigated JEPA-WM predicted futures, specifically how activation steering can impact robotics task planning and outcomes. Performed a series of ablations to test hypotheses about CEM action planning in latent space for world models.
Unofficial implementation to reproduce the experiments from "Superposition as a Phase Change" of "Toy Models of Superposition".
Local agent-driven mechanistic interpretability research platform for Apple Silicon
Cross-architecture mechanistic interpretability toolkit — first OSS Mamba SSM state extraction. Works on transformer + SSM + hybrid models with unified API.
A re-implementation of the "Refusal in Language Models is Meditated by a Single Direction" to understand mech interp
To associate your repository with the mech-interp topic, visit your repo's landing page and select "manage topics."