[ICLR 2025] General-purpose activation steering library
-
Updated
Sep 18, 2025 - Python
[ICLR 2025] General-purpose activation steering library
We study whether categorical refusal tokens enable controllable and interpretable safety behavior in language models.
Reproducible, evergreen benchmark for LLM refusal on biological research prompts — 19 models, 141 prompts, 13,389 adjudicated trials
🔓 Ablate — directional ablation (abliteration) toolkit for open-source LLMs. Automatic censorship/refusal removal via residual-stream direction ablation, with KL-guided search, an LLM-judge harness, and one-call push to the Hub. pip install ablate-llm
Public Driftmap harness: public-safe CSV suites + rubrics + run logs for drift detection, refusal integrity, injection resistance, and uncertainty tracking.
RAG with verifiable citations and measured refusal — retrieval scored separately (TF-IDF beats embeddings here), citations validated against chunks actually retrieved.
RefusalScope
Unified CLI/TUI to abliterate any (V)LLM (Heretic / OBLITERATUS / ErisForge) and benchmark the methods on one schema — G/P/S composite + Pareto front. Pure-stdlib core, no GPU for the core.
Training-time defense that redistributes LLM refusal via mean/covariance matching + KD, raising linear-ablation attack rank from K=1 to K≥16 (Llama-3.2-1B-Instruct)
Locating and editing refusal in the J-space workspace with the Jacobian lens: refusal is legible ~10 layers before the first token, and only ~1/3 lives in the verbalizable workspace.
An open reproduction of feature-level activation steering with the prompt set released, showing the capability tax that behavioural metrics miss
Bank and card statements into a reconciled ledger, with a plain-English refusal for the ones that do not add up: silently wrong money $62,832.40 to $0.00.
Public reference interfaces for proof-gated AI action, refusal, authority, and evidence boundaries.
A document copilot that cites what it says and refuses when the evidence is not there: 33 questions, baseline 12/33 to harness 32/33, control 8/8 both ways.
中文企业公开报告 Hybrid RAG:Docling 解析 · Qdrant 稠密/稀疏检索 · 查询理解硬过滤 · 带引用生成与拒答 · 文档生命周期与评测看板
Reproducible refusal-rate evaluation harness for open-weight LLMs — adversarial-prompt benchmarks with byte-identical reruns.
Hybrid RAG: dense + BM25 + RRF, citations-or-refuse, fail-open rerank, allowlist filters, free-path evals. No API key.
Every RAG eval reports faithfulness. None reports over-refusal beside it. An eval harness and policy engine for refusal correctness in regulated RAG: synthetic KYC/AML corpus, deterministic verifier, and a refusal policy you can commit, hash and audit.
Cross-architecture refusal-direction ablation study: Qwen 2.5 + Gemma 2/3/4. Mechanistic explanation for why Gemma 3 specifically admits single-layer jailbreaks.
Add a description, image, and links to the refusal topic page so that developers can more easily learn about it.
To associate your repository with the refusal topic, visit your repo's landing page and select "manage topics."