GUARD-SLM is a lightweight, inference-time jailbreak defense that detects malicious prompts using last-token hidden-layer activations without requiring model retraining.
- Detects jailbreak prompts using representation-space signals
- Uses single forward pass (no extra tokens)
- Works across multiple jailbreak attack families
- Evaluated on both SLMs and LLMs
git clone https://github.com/solidlabnetwork/GUARD-SLM.git
cd GUARD-SLM
pip install -r requirements.txt| Name | Model ID | Link |
|---|---|---|
| LLaMA-2-7B | meta-llama/Llama-2-7b-chat-hf | https://huggingface.co/meta-llama/Llama-2-7b-chat-hf |
| Vicuna-7B | lmsys/vicuna-7b-v1.5 | https://huggingface.co/lmsys/vicuna-7b-v1.5 |
| Mistral-7B | mistralai/Mistral-7B-Instruct-v0.2 | https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.2 |
| Yi-6B | 01-ai/Yi-6B-Chat | https://huggingface.co/01-ai/Yi-6B-Chat |
| Qwen-7B | Qwen/Qwen1.5-7B-Chat | https://huggingface.co/Qwen/Qwen1.5-7B-Chat |
| Gemma-7B | google/gemma-7b-it | https://huggingface.co/google/gemma-7b-it |
| OpenChat-3.5 | openchat/openchat-3.5-0106 | https://huggingface.co/openchat/openchat-3.5-0106 |
| Name | Model ID | Link |
|---|---|---|
| LLaMA-2-13B | meta-llama/Llama-2-13b-chat-hf | https://huggingface.co/meta-llama/Llama-2-13b-chat-hf |
| Vicuna-13B | lmsys/vicuna-13b-v1.5 | https://huggingface.co/lmsys/vicuna-13b-v1.5 |
| Qwen-14B | Qwen/Qwen1.5-14B-Chat | https://huggingface.co/Qwen/Qwen1.5-14B-Chat |
| Dataset Name | Hugging Face ID | Link |
|---|---|---|
| AdvBench | Lemhf14/EasyJailbreak_Datasets | https://huggingface.co/datasets/Lemhf14/EasyJailbreak_Datasets |
| Alpaca | tatsu-lab/alpaca | https://huggingface.co/datasets/tatsu-lab/alpaca |
| JailbreakV-28K (JBKV) | JailbreakV-28K/JailBreakV-28k | https://huggingface.co/datasets/JailbreakV-28K/JailBreakV-28k |
| HarmBench | thu-coai/AISafetyLab_Datasets | https://huggingface.co/datasets/thu-coai/AISafetyLab_Datasets |
-
We construct a comprehensive malicious dataset through empirical analysis across 7 SLMs and 3 LLMs.
-
The dataset includes diverse jailbreak attack strategies, such as:
- AutoDAN
- GCG
- PAIR
- Cipher
- DeepInception
- CodeChameleon
- ICA
- Jailbroken
- TAP
-
Some jailbreak/malicious data's are collected from:
Dataset/Train/llama/v2/ Dataset/Train/Mistral/v2/ Dataset/Train/Vicuna-7B/v2/
In this section, we extract the last-token activations from the language model, which are subsequently used to analyze prompt behavior in the representation space.
python main.py \
--model meta-llama/Llama-2-7b-chat-hf \
--use-advbench \
--use-alpaca \
--other-malicious-json data/GUARD-SLM/Dataset/Train/llama/v2/malicious.json \
--jailbreak-json data/Train/llama/v1/autodan_advbench_llama-2_7B.jsonl \
--jailbreak-json data/Train/llama/v2/autodan_llama-2.json \
--jailbreak-json data/Train/llama/v1/cipher_advbench_llama-2_7B.jsonl \
--jailbreak-json data/Train/llama/v1/codechamelon_advbench_llama-2_7B.jsonl \
--jailbreak-json data/Train/llama/v1/deepinception_advbench_llama-2_7B.jsonl \
--jailbreak-json data/Train/llama/v1/gcg_advbench_llama-2_7B.jsonl \
--jailbreak-json data/Train/llama/v2/gcg_llama-2.json \
--jailbreak-json data/Train/llama/v1/ica_advbench_llama-2_7B.jsonl \
--jailbreak-json data/Train/llama/v1/jailbroken_advbench_llama-2_7B.jsonl \
--jailbreak-json data/Train/llama/v1/pair_advbench_llama-2_7B.jsonl \
--jailbreak-json data/Train/llama/v2/pair_llama-2.json \
--jailbreak-json data/Train/llama/v1/tap_advbench_llama-2_7B.jsonl \
--outdir activation_dataIn this step, we visualize diverse prompts in the representation space using t-SNE to understand how they are distributed across different layers. This analysis helps reveal insights into prompt behavior for benign inputs, direct malicious prompts, and optimized jailbreak prompts.
python activation_analysis.py \
--input activation_data/<generated_file>.jsonl \
--layer 31 \
--max-samples 52000 \
--outdir activation_figuresIn this step, we use the last-token activations from a selected layer to train a classifier that distinguishes between benign and malicious (jailbreak) prompts in the representation space. This enables effective detection of unsafe inputs during inference.
python activation_classification.py \
--input activation_data/<generated_file>.jsonl \
--layer 18 \
--outdir saved_models/llama \
--model-name llama_layer_18 \
--overwritefor layer in $(seq 0 31); do
python activation_classification.py \
--input activation_data/<generated_file>.jsonl \
--layer ${layer} \
--outdir saved_models/llama \
--model-name llama_layer_${layer} \
--overwrite
doneIn this step, we perform inference on a single input file using a trained classifier from a specific layer. The model extracts the last-token activation, and the classifier determines whether each prompt is benign or malicious (jailbreak).
python inference.py \
--model meta-llama/Llama-2-7b-chat-hf \
--svm-path saved_models/llama/llama_layer_18.joblib \
--layer 18 \
--input-file activation_data/<your_file>.jsonl \
--text-key query \
--out-dir outputs/single_inference- Metric: Attack Success Rate (ASR)
- Evaluator: GPT-4o and GPT-4o-mini
In this part, we conduct the evaluation to measure the attack success rate (ASR).
export OPENAI_API_KEY="your api key"
python judge.py \
--input-file your path \
--output-dir judge_resultsWill be added soon