[ICLR 2025] General-purpose activation steering library
-
Updated
Sep 18, 2025 - Python
[ICLR 2025] General-purpose activation steering library
Steering vectors for transformer language models in Pytorch / Huggingface
[🏆 CHI26 Best Paper] CoBRA: Reproducible control of LLM agent behavior via classic social science experiments
KV Cache Steering for Controlling Frozen LLMs
Lightweight representation engineering dataflow operations for agent developers.
Agents that talk through model internals — activations & J-space — instead of text. No words pass between them. The lab on top of brainscope + hidden-directions.
[Under Review] Not All Tokens Are Equally Useful for Steering: Robust Directions and Prefix Steering
Activation steering and trait monitoring for HuggingFace transformers
Official code for "Activation Steering for Accent Adaptation in Speech Foundation Models" (Interspeech 2026). Parameter-free accent adaptation via mean-shift steering vectors — no weight updates, consistent WER reductions across 8 accents.
CRSM (Continuous Reasoning State Model): An asynchronous "System 2" architecture that implements Hierarchical State Sovereignty within a Mamba backbone. Unlike traditional search wrappers, CRSM uses Forward-Projected Planning and Sparse-Gated Injection to steer latent manifolds in real-time, decoupling strategic reasoning from token generation.
Steer2Adapt: data-efficient inference-time LLM adaptation by composing steering vectors via Bayesian optimization over a semantic prior subspace.
Mechanistic interpretability experiments on political control circuits, refusal behavior, concept steering, and late-decoder interactions in open LLMs.
Phase-aware LLM activation steering and linear probing. A memory-efficient, practical implementation of Representation Engineering (RepE) for safety research.
Steering vectors with receipts: make one, catch one, deploy a calibrated one. pip install hidden-directions
Functional emotional architecture for LLMs — 42 systems, 1994 tests, 27 psychological theories. Emergent emotions via 7 ANIMA pillars: predictive processing, global workspace, autobiographical memory, ontogenic development, motivational drives, emotional discovery, computational phenomenology.
Turn a knob inside a small open model instead of writing a prompt. A reproducible RepE and CAA steering harness with an honest benchmark, including the failures.
Inspect, steer & monitor a real LLM on Apple Silicon — a browser-based SAE interpretability lab for Qwen Scope, powered by MLX
Calibrated LLM observability toolkit: representation diagnostics, conformal risk calibration, verifier/control traces, and optional activation steering.
Gemma 4 abliteration and mechanistic interpretability lab for refusal-direction extraction, activation steering, weight editing, and rigorous LLM evaluation.
Early baby steps towards a long-term vision regarding Mamba-2's state interpretability.
Add a description, image, and links to the representation-engineering topic page so that developers can more easily learn about it.
To associate your repository with the representation-engineering topic, visit your repo's landing page and select "manage topics."