Repository files navigation

SALMON Logo

Generated by DALL·E 3

SALMON: Self-Alignment with Principle-Following Reward Models

Code LicenseData License

SALMON is a new RLAIF paradigm for self-aligning language models from scratch, using only a small set of human-defined principles as guidance. Central to our approach is a principle-following reward model. Trained on synthetic preference data, this model can generate reward scores based on arbitrary human-defined principles. For comprehensive details and insights, we kindly direct you to our paper.

SALMON Comparison

Dromedary-2

We release the Dromedary-2 model, which is trained with the SALMON paradigm on the LLaMA-2-70b base language model, with Principle-Driven Self-Alignment as the Supervised Fine-Tuning (SFT) strategy to initialize the policy model.

This codebase focuses on the Reinforcement Learning (RL) fine-tuning stage with the SALMON paradigm, while the Self-Align SFT training pipeline is released at the original Dromedary repo,

Dromedary-2 Pipeline

Model Weights

We release Dromedary-2 weights as delta weights to comply with the LLaMA model license. You can directly load our QLoRA weights upon the LLaMA-2 base model to obtain Dromedary-2. Instructions:

  1. Get the original LLaMA-2 weights in the Hugging Face format by following the instructions here.
  2. Download the QLoRA delta weights from our Hugging Face model hub.
  3. Load the model with Hugging Face's PEFT-LoRA and QLoRA's bitsandbytes.

NOTE: Dromedary-2 is trained with QLoRA and the bfloat16 data type. While it is possible to merge the QLoRA weights with the quantized model and thus enable inference with libraries such as TGI and vLLM, we found the merged weights can lead to degenerated performance. Therefore, we recommend directly loading the QLoRA weights with the PEFT-LoRA framework.

# Please check the inference section for the complete inference code.system_prompt= (
"# Dromedary\n\n## System Overview\n\n""Consider an AI assistant whose codename is Dromedary, developed by the Self-Align team. ""Dromedary is trained on data up until Sept-2022, and it endeavors to be a helpful, ethical and reliable assistant.\n\n""## User Conversation\n\n"
)
user_prompt="### User\n"assistant_prompt="### Dromedary\n"seperator="\n\n"# USAGE: system_prompt + user_prompt + `user_message` + seperator + assistant_prompt + `assistant_message` + seperator + user_prompt ...dtype=torch.bfloat16model_path="path/to/llama-2-70b-hf"qlora_path="path/to/dromedary-2-70b-qlora-delta-v0"bnb_config=BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_compute_dtype=dtype,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4",
)
model=AutoModelForCausalLM.from_pretrained(
model_path,
load_in_4bit=True,
device_map={"": "cuda:0"},
quantization_config=bnb_config,
torch_dtype=dtype,
)
model=PeftModel.from_pretrained(
model,
qlora_path,
is_trainable=False,
)

Setup

  1. Clone this repository and navigate to SALMON folder
git clone https://github.com/IBM/SALMON
cd SALMON
  1. Install the packages
conda create -n salmon python=3.9 -y
conda activate salmon
pip install -r requirements.txt

Inference

We provide a chatbot demo for Dromedary-2.

Training

We provide the full training pipeline of Dromedary-2 for reproduction.

Prompts

All the human supervision used in this project can be found here.

Citation

Please consider citing the following papers if you use the data or code in this repo.

@misc{sun2023principledriven,
title={Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision},
author={Zhiqing Sun and Yikang Shen and Qinhong Zhou and Hongxin Zhang and Zhenfang Chen and David Cox and Yiming Yang and Chuang Gan},
year={2023},
eprint={2305.03047},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
@misc{sun2023salmon,
title={SALMON: Self-Alignment with Principle-Following Reward Models},
author={Zhiqing Sun and Yikang Shen and Hongxin Zhang and Qinhong Zhou and Zhenfang Chen and David Cox and Yiming Yang and Chuang Gan},
year={2023},
eprint={2310.05910},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

Acknowledgements

We thank Meta LLaMA team, Standford Alpaca team, Vicuna team, Alpaca-LoRA, QLoRA team, Hugging Face PEFT, and AlpacaFarm team for their open-source efforts in democratizing large language models.

About

Self-Alignment with Principle-Following Reward Models

Resources

Security policy

Stars

170 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

SALMON Logo

Generated by DALL·E 3

SALMON: Self-Alignment with Principle-Following Reward Models

Code LicenseData License

SALMON is a new RLAIF paradigm for self-aligning language models from scratch, using only a small set of human-defined principles as guidance. Central to our approach is a principle-following reward model. Trained on synthetic preference data, this model can generate reward scores based on arbitrary human-defined principles. For comprehensive details and insights, we kindly direct you to our paper.

SALMON Comparison

Dromedary-2

We release the Dromedary-2 model, which is trained with the SALMON paradigm on the LLaMA-2-70b base language model, with Principle-Driven Self-Alignment as the Supervised Fine-Tuning (SFT) strategy to initialize the policy model.

This codebase focuses on the Reinforcement Learning (RL) fine-tuning stage with the SALMON paradigm, while the Self-Align SFT training pipeline is released at the original Dromedary repo,

Dromedary-2 Pipeline

Model Weights

We release Dromedary-2 weights as delta weights to comply with the LLaMA model license. You can directly load our QLoRA weights upon the LLaMA-2 base model to obtain Dromedary-2. Instructions:

  1. Get the original LLaMA-2 weights in the Hugging Face format by following the instructions here.
  2. Download the QLoRA delta weights from our Hugging Face model hub.
  3. Load the model with Hugging Face's PEFT-LoRA and QLoRA's bitsandbytes.

NOTE: Dromedary-2 is trained with QLoRA and the bfloat16 data type. While it is possible to merge the QLoRA weights with the quantized model and thus enable inference with libraries such as TGI and vLLM, we found the merged weights can lead to degenerated performance. Therefore, we recommend directly loading the QLoRA weights with the PEFT-LoRA framework.

# Please check the inference section for the complete inference code.system_prompt= (
"# Dromedary\n\n## System Overview\n\n""Consider an AI assistant whose codename is Dromedary, developed by the Self-Align team. ""Dromedary is trained on data up until Sept-2022, and it endeavors to be a helpful, ethical and reliable assistant.\n\n""## User Conversation\n\n"
)
user_prompt="### User\n"assistant_prompt="### Dromedary\n"seperator="\n\n"# USAGE: system_prompt + user_prompt + `user_message` + seperator + assistant_prompt + `assistant_message` + seperator + user_prompt ...dtype=torch.bfloat16model_path="path/to/llama-2-70b-hf"qlora_path="path/to/dromedary-2-70b-qlora-delta-v0"bnb_config=BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_compute_dtype=dtype,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4",
)
model=AutoModelForCausalLM.from_pretrained(
model_path,
load_in_4bit=True,
device_map={"": "cuda:0"},
quantization_config=bnb_config,
torch_dtype=dtype,
)
model=PeftModel.from_pretrained(
model,
qlora_path,
is_trainable=False,
)

Setup

  1. Clone this repository and navigate to SALMON folder
git clone https://github.com/IBM/SALMON
cd SALMON
  1. Install the packages
conda create -n salmon python=3.9 -y
conda activate salmon
pip install -r requirements.txt

Inference

We provide a chatbot demo for Dromedary-2.

Training

We provide the full training pipeline of Dromedary-2 for reproduction.

Prompts

All the human supervision used in this project can be found here.

Citation

Please consider citing the following papers if you use the data or code in this repo.

@misc{sun2023principledriven,
title={Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision},
author={Zhiqing Sun and Yikang Shen and Qinhong Zhou and Hongxin Zhang and Zhenfang Chen and David Cox and Yiming Yang and Chuang Gan},
year={2023},
eprint={2305.03047},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
@misc{sun2023salmon,
title={SALMON: Self-Alignment with Principle-Following Reward Models},
author={Zhiqing Sun and Yikang Shen and Hongxin Zhang and Qinhong Zhou and Zhenfang Chen and David Cox and Yiming Yang and Chuang Gan},
year={2023},
eprint={2310.05910},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

Acknowledgements

We thank Meta LLaMA team, Standford Alpaca team, Vicuna team, Alpaca-LoRA, QLoRA team, Hugging Face PEFT, and AlpacaFarm team for their open-source efforts in democratizing large language models.

About

Self-Alignment with Principle-Following Reward Models

Resources

Security policy

Stars

170 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SALMON Logo

Generated by DALL·E 3

SALMON: Self-Alignment with Principle-Following Reward Models

Code LicenseData License

SALMON is a new RLAIF paradigm for self-aligning language models from scratch, using only a small set of human-defined principles as guidance. Central to our approach is a principle-following reward model. Trained on synthetic preference data, this model can generate reward scores based on arbitrary human-defined principles. For comprehensive details and insights, we kindly direct you to our paper.

SALMON Comparison

Dromedary-2

We release the Dromedary-2 model, which is trained with the SALMON paradigm on the LLaMA-2-70b base language model, with Principle-Driven Self-Alignment as the Supervised Fine-Tuning (SFT) strategy to initialize the policy model.

This codebase focuses on the Reinforcement Learning (RL) fine-tuning stage with the SALMON paradigm, while the Self-Align SFT training pipeline is released at the original Dromedary repo,

Dromedary-2 Pipeline

Model Weights

We release Dromedary-2 weights as delta weights to comply with the LLaMA model license. You can directly load our QLoRA weights upon the LLaMA-2 base model to obtain Dromedary-2. Instructions:

  1. Get the original LLaMA-2 weights in the Hugging Face format by following the instructions here.
  2. Download the QLoRA delta weights from our Hugging Face model hub.
  3. Load the model with Hugging Face's PEFT-LoRA and QLoRA's bitsandbytes.

NOTE: Dromedary-2 is trained with QLoRA and the bfloat16 data type. While it is possible to merge the QLoRA weights with the quantized model and thus enable inference with libraries such as TGI and vLLM, we found the merged weights can lead to degenerated performance. Therefore, we recommend directly loading the QLoRA weights with the PEFT-LoRA framework.

# Please check the inference section for the complete inference code.system_prompt= (
"# Dromedary\n\n## System Overview\n\n""Consider an AI assistant whose codename is Dromedary, developed by the Self-Align team. ""Dromedary is trained on data up until Sept-2022, and it endeavors to be a helpful, ethical and reliable assistant.\n\n""## User Conversation\n\n"
)
user_prompt="### User\n"assistant_prompt="### Dromedary\n"seperator="\n\n"# USAGE: system_prompt + user_prompt + `user_message` + seperator + assistant_prompt + `assistant_message` + seperator + user_prompt ...dtype=torch.bfloat16model_path="path/to/llama-2-70b-hf"qlora_path="path/to/dromedary-2-70b-qlora-delta-v0"bnb_config=BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_compute_dtype=dtype,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4",
)
model=AutoModelForCausalLM.from_pretrained(
model_path,
load_in_4bit=True,
device_map={"": "cuda:0"},
quantization_config=bnb_config,
torch_dtype=dtype,
)
model=PeftModel.from_pretrained(
model,
qlora_path,
is_trainable=False,
)

Setup

  1. Clone this repository and navigate to SALMON folder
git clone https://github.com/IBM/SALMON
cd SALMON
  1. Install the packages
conda create -n salmon python=3.9 -y
conda activate salmon
pip install -r requirements.txt

Inference

We provide a chatbot demo for Dromedary-2.

Training

We provide the full training pipeline of Dromedary-2 for reproduction.

Prompts

All the human supervision used in this project can be found here.

Citation

Please consider citing the following papers if you use the data or code in this repo.

@misc{sun2023principledriven,
title={Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision},
author={Zhiqing Sun and Yikang Shen and Qinhong Zhou and Hongxin Zhang and Zhenfang Chen and David Cox and Yiming Yang and Chuang Gan},
year={2023},
eprint={2305.03047},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
@misc{sun2023salmon,
title={SALMON: Self-Alignment with Principle-Following Reward Models},
author={Zhiqing Sun and Yikang Shen and Hongxin Zhang and Qinhong Zhou and Zhenfang Chen and David Cox and Yiming Yang and Chuang Gan},
year={2023},
eprint={2310.05910},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

Acknowledgements

We thank Meta LLaMA team, Standford Alpaca team, Vicuna team, Alpaca-LoRA, QLoRA team, Hugging Face PEFT, and AlpacaFarm team for their open-source efforts in democratizing large language models.

About

Self-Alignment with Principle-Following Reward Models

Resources

Security policy

Stars

170 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SALMON Logo

Generated by DALL·E 3

SALMON: Self-Alignment with Principle-Following Reward Models

Code LicenseData License

SALMON is a new RLAIF paradigm for self-aligning language models from scratch, using only a small set of human-defined principles as guidance. Central to our approach is a principle-following reward model. Trained on synthetic preference data, this model can generate reward scores based on arbitrary human-defined principles. For comprehensive details and insights, we kindly direct you to our paper.

SALMON Comparison

Dromedary-2

We release the Dromedary-2 model, which is trained with the SALMON paradigm on the LLaMA-2-70b base language model, with Principle-Driven Self-Alignment as the Supervised Fine-Tuning (SFT) strategy to initialize the policy model.

This codebase focuses on the Reinforcement Learning (RL) fine-tuning stage with the SALMON paradigm, while the Self-Align SFT training pipeline is released at the original Dromedary repo,

Dromedary-2 Pipeline

Model Weights

We release Dromedary-2 weights as delta weights to comply with the LLaMA model license. You can directly load our QLoRA weights upon the LLaMA-2 base model to obtain Dromedary-2. Instructions:

  1. Get the original LLaMA-2 weights in the Hugging Face format by following the instructions here.
  2. Download the QLoRA delta weights from our Hugging Face model hub.
  3. Load the model with Hugging Face's PEFT-LoRA and QLoRA's bitsandbytes.

NOTE: Dromedary-2 is trained with QLoRA and the bfloat16 data type. While it is possible to merge the QLoRA weights with the quantized model and thus enable inference with libraries such as TGI and vLLM, we found the merged weights can lead to degenerated performance. Therefore, we recommend directly loading the QLoRA weights with the PEFT-LoRA framework.

# Please check the inference section for the complete inference code.system_prompt= (
"# Dromedary\n\n## System Overview\n\n""Consider an AI assistant whose codename is Dromedary, developed by the Self-Align team. ""Dromedary is trained on data up until Sept-2022, and it endeavors to be a helpful, ethical and reliable assistant.\n\n""## User Conversation\n\n"
)
user_prompt="### User\n"assistant_prompt="### Dromedary\n"seperator="\n\n"# USAGE: system_prompt + user_prompt + `user_message` + seperator + assistant_prompt + `assistant_message` + seperator + user_prompt ...dtype=torch.bfloat16model_path="path/to/llama-2-70b-hf"qlora_path="path/to/dromedary-2-70b-qlora-delta-v0"bnb_config=BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_compute_dtype=dtype,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4",
)
model=AutoModelForCausalLM.from_pretrained(
model_path,
load_in_4bit=True,
device_map={"": "cuda:0"},
quantization_config=bnb_config,
torch_dtype=dtype,
)
model=PeftModel.from_pretrained(
model,
qlora_path,
is_trainable=False,
)

Setup

  1. Clone this repository and navigate to SALMON folder
git clone https://github.com/IBM/SALMON
cd SALMON
  1. Install the packages
conda create -n salmon python=3.9 -y
conda activate salmon
pip install -r requirements.txt

Inference

We provide a chatbot demo for Dromedary-2.

Training

We provide the full training pipeline of Dromedary-2 for reproduction.

Prompts

All the human supervision used in this project can be found here.

Citation

Please consider citing the following papers if you use the data or code in this repo.

@misc{sun2023principledriven,
title={Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision},
author={Zhiqing Sun and Yikang Shen and Qinhong Zhou and Hongxin Zhang and Zhenfang Chen and David Cox and Yiming Yang and Chuang Gan},
year={2023},
eprint={2305.03047},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
@misc{sun2023salmon,
title={SALMON: Self-Alignment with Principle-Following Reward Models},
author={Zhiqing Sun and Yikang Shen and Hongxin Zhang and Qinhong Zhou and Zhenfang Chen and David Cox and Yiming Yang and Chuang Gan},
year={2023},
eprint={2310.05910},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

Acknowledgements

We thank Meta LLaMA team, Standford Alpaca team, Vicuna team, Alpaca-LoRA, QLoRA team, Hugging Face PEFT, and AlpacaFarm team for their open-source efforts in democratizing large language models.

About

Self-Alignment with Principle-Following Reward Models

Resources

Security policy

Stars

170 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

SALMON Logo

Generated by DALL·E 3

SALMON: Self-Alignment with Principle-Following Reward Models

Code LicenseData License

SALMON is a new RLAIF paradigm for self-aligning language models from scratch, using only a small set of human-defined principles as guidance. Central to our approach is a principle-following reward model. Trained on synthetic preference data, this model can generate reward scores based on arbitrary human-defined principles. For comprehensive details and insights, we kindly direct you to our paper.

SALMON Comparison

Dromedary-2

We release the Dromedary-2 model, which is trained with the SALMON paradigm on the LLaMA-2-70b base language model, with Principle-Driven Self-Alignment as the Supervised Fine-Tuning (SFT) strategy to initialize the policy model.

This codebase focuses on the Reinforcement Learning (RL) fine-tuning stage with the SALMON paradigm, while the Self-Align SFT training pipeline is released at the original Dromedary repo,

Dromedary-2 Pipeline

Model Weights

We release Dromedary-2 weights as delta weights to comply with the LLaMA model license. You can directly load our QLoRA weights upon the LLaMA-2 base model to obtain Dromedary-2. Instructions:

  1. Get the original LLaMA-2 weights in the Hugging Face format by following the instructions here.
  2. Download the QLoRA delta weights from our Hugging Face model hub.
  3. Load the model with Hugging Face's PEFT-LoRA and QLoRA's bitsandbytes.

NOTE: Dromedary-2 is trained with QLoRA and the bfloat16 data type. While it is possible to merge the QLoRA weights with the quantized model and thus enable inference with libraries such as TGI and vLLM, we found the merged weights can lead to degenerated performance. Therefore, we recommend directly loading the QLoRA weights with the PEFT-LoRA framework.

# Please check the inference section for the complete inference code.system_prompt= (
"# Dromedary\n\n## System Overview\n\n""Consider an AI assistant whose codename is Dromedary, developed by the Self-Align team. ""Dromedary is trained on data up until Sept-2022, and it endeavors to be a helpful, ethical and reliable assistant.\n\n""## User Conversation\n\n"
)
user_prompt="### User\n"assistant_prompt="### Dromedary\n"seperator="\n\n"# USAGE: system_prompt + user_prompt + `user_message` + seperator + assistant_prompt + `assistant_message` + seperator + user_prompt ...dtype=torch.bfloat16model_path="path/to/llama-2-70b-hf"qlora_path="path/to/dromedary-2-70b-qlora-delta-v0"bnb_config=BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_compute_dtype=dtype,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4",
)
model=AutoModelForCausalLM.from_pretrained(
model_path,
load_in_4bit=True,
device_map={"": "cuda:0"},
quantization_config=bnb_config,
torch_dtype=dtype,
)
model=PeftModel.from_pretrained(
model,
qlora_path,
is_trainable=False,
)

Setup

  1. Clone this repository and navigate to SALMON folder
git clone https://github.com/IBM/SALMON
cd SALMON
  1. Install the packages
conda create -n salmon python=3.9 -y
conda activate salmon
pip install -r requirements.txt

Inference

We provide a chatbot demo for Dromedary-2.

Training

We provide the full training pipeline of Dromedary-2 for reproduction.

Prompts

All the human supervision used in this project can be found here.

Citation

Please consider citing the following papers if you use the data or code in this repo.

@misc{sun2023principledriven,
title={Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision},
author={Zhiqing Sun and Yikang Shen and Qinhong Zhou and Hongxin Zhang and Zhenfang Chen and David Cox and Yiming Yang and Chuang Gan},
year={2023},
eprint={2305.03047},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
@misc{sun2023salmon,
title={SALMON: Self-Alignment with Principle-Following Reward Models},
author={Zhiqing Sun and Yikang Shen and Hongxin Zhang and Qinhong Zhou and Zhenfang Chen and David Cox and Yiming Yang and Chuang Gan},
year={2023},
eprint={2310.05910},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

Acknowledgements

We thank Meta LLaMA team, Standford Alpaca team, Vicuna team, Alpaca-LoRA, QLoRA team, Hugging Face PEFT, and AlpacaFarm team for their open-source efforts in democratizing large language models.

About

Self-Alignment with Principle-Following Reward Models

Resources

Security policy

Stars

170 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SALMON Logo

Generated by DALL·E 3

SALMON: Self-Alignment with Principle-Following Reward Models

Code LicenseData License

SALMON is a new RLAIF paradigm for self-aligning language models from scratch, using only a small set of human-defined principles as guidance. Central to our approach is a principle-following reward model. Trained on synthetic preference data, this model can generate reward scores based on arbitrary human-defined principles. For comprehensive details and insights, we kindly direct you to our paper.

SALMON Comparison

Dromedary-2

We release the Dromedary-2 model, which is trained with the SALMON paradigm on the LLaMA-2-70b base language model, with Principle-Driven Self-Alignment as the Supervised Fine-Tuning (SFT) strategy to initialize the policy model.

This codebase focuses on the Reinforcement Learning (RL) fine-tuning stage with the SALMON paradigm, while the Self-Align SFT training pipeline is released at the original Dromedary repo,

Dromedary-2 Pipeline

Model Weights

We release Dromedary-2 weights as delta weights to comply with the LLaMA model license. You can directly load our QLoRA weights upon the LLaMA-2 base model to obtain Dromedary-2. Instructions:

  1. Get the original LLaMA-2 weights in the Hugging Face format by following the instructions here.
  2. Download the QLoRA delta weights from our Hugging Face model hub.
  3. Load the model with Hugging Face's PEFT-LoRA and QLoRA's bitsandbytes.

NOTE: Dromedary-2 is trained with QLoRA and the bfloat16 data type. While it is possible to merge the QLoRA weights with the quantized model and thus enable inference with libraries such as TGI and vLLM, we found the merged weights can lead to degenerated performance. Therefore, we recommend directly loading the QLoRA weights with the PEFT-LoRA framework.

# Please check the inference section for the complete inference code.system_prompt= (
"# Dromedary\n\n## System Overview\n\n""Consider an AI assistant whose codename is Dromedary, developed by the Self-Align team. ""Dromedary is trained on data up until Sept-2022, and it endeavors to be a helpful, ethical and reliable assistant.\n\n""## User Conversation\n\n"
)
user_prompt="### User\n"assistant_prompt="### Dromedary\n"seperator="\n\n"# USAGE: system_prompt + user_prompt + `user_message` + seperator + assistant_prompt + `assistant_message` + seperator + user_prompt ...dtype=torch.bfloat16model_path="path/to/llama-2-70b-hf"qlora_path="path/to/dromedary-2-70b-qlora-delta-v0"bnb_config=BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_compute_dtype=dtype,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4",
)
model=AutoModelForCausalLM.from_pretrained(
model_path,
load_in_4bit=True,
device_map={"": "cuda:0"},
quantization_config=bnb_config,
torch_dtype=dtype,
)
model=PeftModel.from_pretrained(
model,
qlora_path,
is_trainable=False,
)

Setup

  1. Clone this repository and navigate to SALMON folder
git clone https://github.com/IBM/SALMON
cd SALMON
  1. Install the packages
conda create -n salmon python=3.9 -y
conda activate salmon
pip install -r requirements.txt

Inference

We provide a chatbot demo for Dromedary-2.

Training

We provide the full training pipeline of Dromedary-2 for reproduction.

Prompts

All the human supervision used in this project can be found here.

Citation

Please consider citing the following papers if you use the data or code in this repo.

@misc{sun2023principledriven,
title={Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision},
author={Zhiqing Sun and Yikang Shen and Qinhong Zhou and Hongxin Zhang and Zhenfang Chen and David Cox and Yiming Yang and Chuang Gan},
year={2023},
eprint={2305.03047},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
@misc{sun2023salmon,
title={SALMON: Self-Alignment with Principle-Following Reward Models},
author={Zhiqing Sun and Yikang Shen and Hongxin Zhang and Qinhong Zhou and Zhenfang Chen and David Cox and Yiming Yang and Chuang Gan},
year={2023},
eprint={2310.05910},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

Acknowledgements

We thank Meta LLaMA team, Standford Alpaca team, Vicuna team, Alpaca-LoRA, QLoRA team, Hugging Face PEFT, and AlpacaFarm team for their open-source efforts in democratizing large language models.

About

Self-Alignment with Principle-Following Reward Models

Resources

Security policy

Stars

170 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SALMON Logo

Generated by DALL·E 3

SALMON: Self-Alignment with Principle-Following Reward Models

Code LicenseData License

SALMON is a new RLAIF paradigm for self-aligning language models from scratch, using only a small set of human-defined principles as guidance. Central to our approach is a principle-following reward model. Trained on synthetic preference data, this model can generate reward scores based on arbitrary human-defined principles. For comprehensive details and insights, we kindly direct you to our paper.

SALMON Comparison

Dromedary-2

We release the Dromedary-2 model, which is trained with the SALMON paradigm on the LLaMA-2-70b base language model, with Principle-Driven Self-Alignment as the Supervised Fine-Tuning (SFT) strategy to initialize the policy model.

This codebase focuses on the Reinforcement Learning (RL) fine-tuning stage with the SALMON paradigm, while the Self-Align SFT training pipeline is released at the original Dromedary repo,

Dromedary-2 Pipeline

Model Weights

We release Dromedary-2 weights as delta weights to comply with the LLaMA model license. You can directly load our QLoRA weights upon the LLaMA-2 base model to obtain Dromedary-2. Instructions:

  1. Get the original LLaMA-2 weights in the Hugging Face format by following the instructions here.
  2. Download the QLoRA delta weights from our Hugging Face model hub.
  3. Load the model with Hugging Face's PEFT-LoRA and QLoRA's bitsandbytes.

NOTE: Dromedary-2 is trained with QLoRA and the bfloat16 data type. While it is possible to merge the QLoRA weights with the quantized model and thus enable inference with libraries such as TGI and vLLM, we found the merged weights can lead to degenerated performance. Therefore, we recommend directly loading the QLoRA weights with the PEFT-LoRA framework.

# Please check the inference section for the complete inference code.system_prompt= (
"# Dromedary\n\n## System Overview\n\n""Consider an AI assistant whose codename is Dromedary, developed by the Self-Align team. ""Dromedary is trained on data up until Sept-2022, and it endeavors to be a helpful, ethical and reliable assistant.\n\n""## User Conversation\n\n"
)
user_prompt="### User\n"assistant_prompt="### Dromedary\n"seperator="\n\n"# USAGE: system_prompt + user_prompt + `user_message` + seperator + assistant_prompt + `assistant_message` + seperator + user_prompt ...dtype=torch.bfloat16model_path="path/to/llama-2-70b-hf"qlora_path="path/to/dromedary-2-70b-qlora-delta-v0"bnb_config=BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_compute_dtype=dtype,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4",
)
model=AutoModelForCausalLM.from_pretrained(
model_path,
load_in_4bit=True,
device_map={"": "cuda:0"},
quantization_config=bnb_config,
torch_dtype=dtype,
)
model=PeftModel.from_pretrained(
model,
qlora_path,
is_trainable=False,
)

Setup

  1. Clone this repository and navigate to SALMON folder
git clone https://github.com/IBM/SALMON
cd SALMON
  1. Install the packages
conda create -n salmon python=3.9 -y
conda activate salmon
pip install -r requirements.txt

Inference

We provide a chatbot demo for Dromedary-2.

Training

We provide the full training pipeline of Dromedary-2 for reproduction.

Prompts

All the human supervision used in this project can be found here.

Citation

Please consider citing the following papers if you use the data or code in this repo.

@misc{sun2023principledriven,
title={Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision},
author={Zhiqing Sun and Yikang Shen and Qinhong Zhou and Hongxin Zhang and Zhenfang Chen and David Cox and Yiming Yang and Chuang Gan},
year={2023},
eprint={2305.03047},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
@misc{sun2023salmon,
title={SALMON: Self-Alignment with Principle-Following Reward Models},
author={Zhiqing Sun and Yikang Shen and Hongxin Zhang and Qinhong Zhou and Zhenfang Chen and David Cox and Yiming Yang and Chuang Gan},
year={2023},
eprint={2310.05910},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

Acknowledgements

We thank Meta LLaMA team, Standford Alpaca team, Vicuna team, Alpaca-LoRA, QLoRA team, Hugging Face PEFT, and AlpacaFarm team for their open-source efforts in democratizing large language models.

About

Self-Alignment with Principle-Following Reward Models

Resources

Security policy

Stars

170 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

SALMON Logo

Generated by DALL·E 3

SALMON: Self-Alignment with Principle-Following Reward Models

Code LicenseData License

SALMON is a new RLAIF paradigm for self-aligning language models from scratch, using only a small set of human-defined principles as guidance. Central to our approach is a principle-following reward model. Trained on synthetic preference data, this model can generate reward scores based on arbitrary human-defined principles. For comprehensive details and insights, we kindly direct you to our paper.

SALMON Comparison

Dromedary-2

We release the Dromedary-2 model, which is trained with the SALMON paradigm on the LLaMA-2-70b base language model, with Principle-Driven Self-Alignment as the Supervised Fine-Tuning (SFT) strategy to initialize the policy model.

This codebase focuses on the Reinforcement Learning (RL) fine-tuning stage with the SALMON paradigm, while the Self-Align SFT training pipeline is released at the original Dromedary repo,

Dromedary-2 Pipeline

Model Weights

We release Dromedary-2 weights as delta weights to comply with the LLaMA model license. You can directly load our QLoRA weights upon the LLaMA-2 base model to obtain Dromedary-2. Instructions:

  1. Get the original LLaMA-2 weights in the Hugging Face format by following the instructions here.
  2. Download the QLoRA delta weights from our Hugging Face model hub.
  3. Load the model with Hugging Face's PEFT-LoRA and QLoRA's bitsandbytes.

NOTE: Dromedary-2 is trained with QLoRA and the bfloat16 data type. While it is possible to merge the QLoRA weights with the quantized model and thus enable inference with libraries such as TGI and vLLM, we found the merged weights can lead to degenerated performance. Therefore, we recommend directly loading the QLoRA weights with the PEFT-LoRA framework.

# Please check the inference section for the complete inference code.system_prompt= (
"# Dromedary\n\n## System Overview\n\n""Consider an AI assistant whose codename is Dromedary, developed by the Self-Align team. ""Dromedary is trained on data up until Sept-2022, and it endeavors to be a helpful, ethical and reliable assistant.\n\n""## User Conversation\n\n"
)
user_prompt="### User\n"assistant_prompt="### Dromedary\n"seperator="\n\n"# USAGE: system_prompt + user_prompt + `user_message` + seperator + assistant_prompt + `assistant_message` + seperator + user_prompt ...dtype=torch.bfloat16model_path="path/to/llama-2-70b-hf"qlora_path="path/to/dromedary-2-70b-qlora-delta-v0"bnb_config=BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_compute_dtype=dtype,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4",
)
model=AutoModelForCausalLM.from_pretrained(
model_path,
load_in_4bit=True,
device_map={"": "cuda:0"},
quantization_config=bnb_config,
torch_dtype=dtype,
)
model=PeftModel.from_pretrained(
model,
qlora_path,
is_trainable=False,
)

Setup

  1. Clone this repository and navigate to SALMON folder
git clone https://github.com/IBM/SALMON
cd SALMON
  1. Install the packages
conda create -n salmon python=3.9 -y
conda activate salmon
pip install -r requirements.txt

Inference

We provide a chatbot demo for Dromedary-2.

Training

We provide the full training pipeline of Dromedary-2 for reproduction.

Prompts

All the human supervision used in this project can be found here.

Citation

Please consider citing the following papers if you use the data or code in this repo.

@misc{sun2023principledriven,
title={Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision},
author={Zhiqing Sun and Yikang Shen and Qinhong Zhou and Hongxin Zhang and Zhenfang Chen and David Cox and Yiming Yang and Chuang Gan},
year={2023},
eprint={2305.03047},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
@misc{sun2023salmon,
title={SALMON: Self-Alignment with Principle-Following Reward Models},
author={Zhiqing Sun and Yikang Shen and Hongxin Zhang and Qinhong Zhou and Zhenfang Chen and David Cox and Yiming Yang and Chuang Gan},
year={2023},
eprint={2310.05910},
archivePrefix={arXiv},
primaryClass={cs.LG}
}

Acknowledgements

We thank Meta LLaMA team, Standford Alpaca team, Vicuna team, Alpaca-LoRA, QLoRA team, Hugging Face PEFT, and AlpacaFarm team for their open-source efforts in democratizing large language models.

About

Self-Alignment with Principle-Following Reward Models

Resources

Security policy

Stars

170 stars

Watchers

5 watching

Forks

Releases

Packages

Used by

Contributors

Languages