Repository files navigation

MACA: Multi-Agent Consensus Alignment

Internalizing Self-Consistency in Language Models through Multi-Agent Debate

Policy

📄 Paper: arXiv:2509.15172

Overview

MACA trains language models to be more consistent reasoners through multi-agent debate and consensus-based reinforcement learning.

Key Features:

  • 🤖 Multi-Agent Debate: Orchestrate debates between agents for improved reasoning
  • 🎯 Consensus Training: Post-train on debate outputs using agreement patterns as rewards
  • Distributed Processing: Multi-GPU parallel training with QLoRA adapters
  • 📊 Analysis Tools: Built-in performance tracking and visualization

Quick Start

Installation

conda env create -f env.yml
pip install -e .

Multi-Agent Training

python main.py --model qwen2b --dataset gsm8k --gpus_per_model 1 --max_concurrent_tasks 4 --train_size 1500 --test_size 500 --lora_r 128 --lora_alpha 128 --dpo --epoch_dpo 3 --batch_dpo 6 --lr_dpo 1e-5 --beta_dpo 0.1 --gradient_accumulation_steps_dpo 4 --seed 1 --wandb

Single-Agent Training

python maca_single_agent.py --output_dir q2b_sa_runs --model qwen2b --phase kto --kto --train_datasets math gsm8k mathqa --test_datasets math gsm8k mathqa svamp gpqa csqa --use_full_test --lora_r_range 64 --lora_alpha_range 64 --lr_kto 1e-5 --evaluation_batch_size 24 --wandb

Key Arguments

  • --model: Quantized base model to use (llama1b/3b/8b, phi4b, qwen2b/7b, gemma4b, mistral7b)
  • --dataset: Dataset (gsm8k, math, mathqa, gpqa, svamp, csqa)
  • --agents: Number of agents in the debate (default: 3)
  • --finetune: Enable Majority Vote Supervised Fine-Tuning (SFT)
  • --post_train: Enable Majority Vote Group-Relative Policy Optimization training (GRPO)
  • --dpo: Enable Majority Vote Direct Preference Optimization training
  • --kto: Enable Majority Vote Kahneman-Tversky Optimization training
  • --use_consensus_reward: Enable consensus-based rewards

Wandb logging

  • --wandb: Enable Weights & Biases logging
  • --project_name: W&B project name (default: llm-marl, requires setting --wandb)
  • --entity_name: W&B entity/team name (default: llm-marl, requires setting --wandb)

Project Structure

maca/
├── main.py # Main training entry point
├── maca_single_agent.py # Single agent hyperparameter tuning and testing
├── model.py # Agent implementation and reward functions
├── debate.py # Multi-agent debate orchestration
├── orchestrator.py # Training coordination and management
├── data.py # Dataset loading and preprocessing
├── parser.py # Answer parsing and grading utilities
├── args.py # Command-line argument definitions
├── utils.py # Utility functions and helpers
├── scheduler.py # Dynamic job scheduling for adapters
├── train_agent_subprocess.py # Subprocess training management
├── analyze_experiment_performance.py # Debate results analysis
├── read_debate_performance.py # Read debate utils
├── data/ # Dataset storage and splits
├── experiments/ # Experiment outputs and results
└── checkpoints/ # Model checkpoints and adapters

Training Methods

Built on Hugging Face TRL, supports multiple paradigms with majority vote variants:

  • MV-SFT: Supervised fine-tuning on consensus examples
  • MV-GRPO: Reinforcement learning with consensus rewards
  • MV-KTO/DPO: Preference optimization methods

See args.py for complete argument documentation.

Citation

This work was developed at Meta AI in collaboration with Meta Superintelligence Labs and the LIINC Lab at Columbia University.

If you use this framework in your research, please cite:

@misc{samanta2024maca,
title={Internalizing Self-Consistency in Language Models: Multi-Agent Consensus Alignment},
author={Ankur Samanta and Akshayaa Magesh and Youliang Yu and Runzhe Wu and Ayush Jain and Daniel Jiang and Boris Vidolov and Paul Sajda and Yonathan Efroni and Kaveh Hassani},
year={2024},
eprint={2509.15172},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://doi.org/10.48550/arXiv.2509.15172}
}

License

MACA is MIT licensed, as found in the LICENSE file.

About

MACA trains language models to be more consistent reasoners through multi-agent debate and consensus-based reinforcement learning.

Resources

Code of conduct

Contributing

Security policy

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

MACA: Multi-Agent Consensus Alignment

Internalizing Self-Consistency in Language Models through Multi-Agent Debate

Policy

📄 Paper: arXiv:2509.15172

Overview

MACA trains language models to be more consistent reasoners through multi-agent debate and consensus-based reinforcement learning.

Key Features:

  • 🤖 Multi-Agent Debate: Orchestrate debates between agents for improved reasoning
  • 🎯 Consensus Training: Post-train on debate outputs using agreement patterns as rewards
  • Distributed Processing: Multi-GPU parallel training with QLoRA adapters
  • 📊 Analysis Tools: Built-in performance tracking and visualization

Quick Start

Installation

conda env create -f env.yml
pip install -e .

Multi-Agent Training

python main.py --model qwen2b --dataset gsm8k --gpus_per_model 1 --max_concurrent_tasks 4 --train_size 1500 --test_size 500 --lora_r 128 --lora_alpha 128 --dpo --epoch_dpo 3 --batch_dpo 6 --lr_dpo 1e-5 --beta_dpo 0.1 --gradient_accumulation_steps_dpo 4 --seed 1 --wandb

Single-Agent Training

python maca_single_agent.py --output_dir q2b_sa_runs --model qwen2b --phase kto --kto --train_datasets math gsm8k mathqa --test_datasets math gsm8k mathqa svamp gpqa csqa --use_full_test --lora_r_range 64 --lora_alpha_range 64 --lr_kto 1e-5 --evaluation_batch_size 24 --wandb

Key Arguments

  • --model: Quantized base model to use (llama1b/3b/8b, phi4b, qwen2b/7b, gemma4b, mistral7b)
  • --dataset: Dataset (gsm8k, math, mathqa, gpqa, svamp, csqa)
  • --agents: Number of agents in the debate (default: 3)
  • --finetune: Enable Majority Vote Supervised Fine-Tuning (SFT)
  • --post_train: Enable Majority Vote Group-Relative Policy Optimization training (GRPO)
  • --dpo: Enable Majority Vote Direct Preference Optimization training
  • --kto: Enable Majority Vote Kahneman-Tversky Optimization training
  • --use_consensus_reward: Enable consensus-based rewards

Wandb logging

  • --wandb: Enable Weights & Biases logging
  • --project_name: W&B project name (default: llm-marl, requires setting --wandb)
  • --entity_name: W&B entity/team name (default: llm-marl, requires setting --wandb)

Project Structure

maca/
├── main.py # Main training entry point
├── maca_single_agent.py # Single agent hyperparameter tuning and testing
├── model.py # Agent implementation and reward functions
├── debate.py # Multi-agent debate orchestration
├── orchestrator.py # Training coordination and management
├── data.py # Dataset loading and preprocessing
├── parser.py # Answer parsing and grading utilities
├── args.py # Command-line argument definitions
├── utils.py # Utility functions and helpers
├── scheduler.py # Dynamic job scheduling for adapters
├── train_agent_subprocess.py # Subprocess training management
├── analyze_experiment_performance.py # Debate results analysis
├── read_debate_performance.py # Read debate utils
├── data/ # Dataset storage and splits
├── experiments/ # Experiment outputs and results
└── checkpoints/ # Model checkpoints and adapters

Training Methods

Built on Hugging Face TRL, supports multiple paradigms with majority vote variants:

  • MV-SFT: Supervised fine-tuning on consensus examples
  • MV-GRPO: Reinforcement learning with consensus rewards
  • MV-KTO/DPO: Preference optimization methods

See args.py for complete argument documentation.

Citation

This work was developed at Meta AI in collaboration with Meta Superintelligence Labs and the LIINC Lab at Columbia University.

If you use this framework in your research, please cite:

@misc{samanta2024maca,
title={Internalizing Self-Consistency in Language Models: Multi-Agent Consensus Alignment},
author={Ankur Samanta and Akshayaa Magesh and Youliang Yu and Runzhe Wu and Ayush Jain and Daniel Jiang and Boris Vidolov and Paul Sajda and Yonathan Efroni and Kaveh Hassani},
year={2024},
eprint={2509.15172},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://doi.org/10.48550/arXiv.2509.15172}
}

License

MACA is MIT licensed, as found in the LICENSE file.

About

MACA trains language models to be more consistent reasoners through multi-agent debate and consensus-based reinforcement learning.

Resources

Code of conduct

Contributing

Security policy

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MACA: Multi-Agent Consensus Alignment

Internalizing Self-Consistency in Language Models through Multi-Agent Debate

Policy

📄 Paper: arXiv:2509.15172

Overview

MACA trains language models to be more consistent reasoners through multi-agent debate and consensus-based reinforcement learning.

Key Features:

  • 🤖 Multi-Agent Debate: Orchestrate debates between agents for improved reasoning
  • 🎯 Consensus Training: Post-train on debate outputs using agreement patterns as rewards
  • Distributed Processing: Multi-GPU parallel training with QLoRA adapters
  • 📊 Analysis Tools: Built-in performance tracking and visualization

Quick Start

Installation

conda env create -f env.yml
pip install -e .

Multi-Agent Training

python main.py --model qwen2b --dataset gsm8k --gpus_per_model 1 --max_concurrent_tasks 4 --train_size 1500 --test_size 500 --lora_r 128 --lora_alpha 128 --dpo --epoch_dpo 3 --batch_dpo 6 --lr_dpo 1e-5 --beta_dpo 0.1 --gradient_accumulation_steps_dpo 4 --seed 1 --wandb

Single-Agent Training

python maca_single_agent.py --output_dir q2b_sa_runs --model qwen2b --phase kto --kto --train_datasets math gsm8k mathqa --test_datasets math gsm8k mathqa svamp gpqa csqa --use_full_test --lora_r_range 64 --lora_alpha_range 64 --lr_kto 1e-5 --evaluation_batch_size 24 --wandb

Key Arguments

  • --model: Quantized base model to use (llama1b/3b/8b, phi4b, qwen2b/7b, gemma4b, mistral7b)
  • --dataset: Dataset (gsm8k, math, mathqa, gpqa, svamp, csqa)
  • --agents: Number of agents in the debate (default: 3)
  • --finetune: Enable Majority Vote Supervised Fine-Tuning (SFT)
  • --post_train: Enable Majority Vote Group-Relative Policy Optimization training (GRPO)
  • --dpo: Enable Majority Vote Direct Preference Optimization training
  • --kto: Enable Majority Vote Kahneman-Tversky Optimization training
  • --use_consensus_reward: Enable consensus-based rewards

Wandb logging

  • --wandb: Enable Weights & Biases logging
  • --project_name: W&B project name (default: llm-marl, requires setting --wandb)
  • --entity_name: W&B entity/team name (default: llm-marl, requires setting --wandb)

Project Structure

maca/
├── main.py # Main training entry point
├── maca_single_agent.py # Single agent hyperparameter tuning and testing
├── model.py # Agent implementation and reward functions
├── debate.py # Multi-agent debate orchestration
├── orchestrator.py # Training coordination and management
├── data.py # Dataset loading and preprocessing
├── parser.py # Answer parsing and grading utilities
├── args.py # Command-line argument definitions
├── utils.py # Utility functions and helpers
├── scheduler.py # Dynamic job scheduling for adapters
├── train_agent_subprocess.py # Subprocess training management
├── analyze_experiment_performance.py # Debate results analysis
├── read_debate_performance.py # Read debate utils
├── data/ # Dataset storage and splits
├── experiments/ # Experiment outputs and results
└── checkpoints/ # Model checkpoints and adapters

Training Methods

Built on Hugging Face TRL, supports multiple paradigms with majority vote variants:

  • MV-SFT: Supervised fine-tuning on consensus examples
  • MV-GRPO: Reinforcement learning with consensus rewards
  • MV-KTO/DPO: Preference optimization methods

See args.py for complete argument documentation.

Citation

This work was developed at Meta AI in collaboration with Meta Superintelligence Labs and the LIINC Lab at Columbia University.

If you use this framework in your research, please cite:

@misc{samanta2024maca,
title={Internalizing Self-Consistency in Language Models: Multi-Agent Consensus Alignment},
author={Ankur Samanta and Akshayaa Magesh and Youliang Yu and Runzhe Wu and Ayush Jain and Daniel Jiang and Boris Vidolov and Paul Sajda and Yonathan Efroni and Kaveh Hassani},
year={2024},
eprint={2509.15172},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://doi.org/10.48550/arXiv.2509.15172}
}

License

MACA is MIT licensed, as found in the LICENSE file.

About

MACA trains language models to be more consistent reasoners through multi-agent debate and consensus-based reinforcement learning.

Resources

Code of conduct

Contributing

Security policy

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MACA: Multi-Agent Consensus Alignment

Internalizing Self-Consistency in Language Models through Multi-Agent Debate

Policy

📄 Paper: arXiv:2509.15172

Overview

MACA trains language models to be more consistent reasoners through multi-agent debate and consensus-based reinforcement learning.

Key Features:

  • 🤖 Multi-Agent Debate: Orchestrate debates between agents for improved reasoning
  • 🎯 Consensus Training: Post-train on debate outputs using agreement patterns as rewards
  • Distributed Processing: Multi-GPU parallel training with QLoRA adapters
  • 📊 Analysis Tools: Built-in performance tracking and visualization

Quick Start

Installation

conda env create -f env.yml
pip install -e .

Multi-Agent Training

python main.py --model qwen2b --dataset gsm8k --gpus_per_model 1 --max_concurrent_tasks 4 --train_size 1500 --test_size 500 --lora_r 128 --lora_alpha 128 --dpo --epoch_dpo 3 --batch_dpo 6 --lr_dpo 1e-5 --beta_dpo 0.1 --gradient_accumulation_steps_dpo 4 --seed 1 --wandb

Single-Agent Training

python maca_single_agent.py --output_dir q2b_sa_runs --model qwen2b --phase kto --kto --train_datasets math gsm8k mathqa --test_datasets math gsm8k mathqa svamp gpqa csqa --use_full_test --lora_r_range 64 --lora_alpha_range 64 --lr_kto 1e-5 --evaluation_batch_size 24 --wandb

Key Arguments

  • --model: Quantized base model to use (llama1b/3b/8b, phi4b, qwen2b/7b, gemma4b, mistral7b)
  • --dataset: Dataset (gsm8k, math, mathqa, gpqa, svamp, csqa)
  • --agents: Number of agents in the debate (default: 3)
  • --finetune: Enable Majority Vote Supervised Fine-Tuning (SFT)
  • --post_train: Enable Majority Vote Group-Relative Policy Optimization training (GRPO)
  • --dpo: Enable Majority Vote Direct Preference Optimization training
  • --kto: Enable Majority Vote Kahneman-Tversky Optimization training
  • --use_consensus_reward: Enable consensus-based rewards

Wandb logging

  • --wandb: Enable Weights & Biases logging
  • --project_name: W&B project name (default: llm-marl, requires setting --wandb)
  • --entity_name: W&B entity/team name (default: llm-marl, requires setting --wandb)

Project Structure

maca/
├── main.py # Main training entry point
├── maca_single_agent.py # Single agent hyperparameter tuning and testing
├── model.py # Agent implementation and reward functions
├── debate.py # Multi-agent debate orchestration
├── orchestrator.py # Training coordination and management
├── data.py # Dataset loading and preprocessing
├── parser.py # Answer parsing and grading utilities
├── args.py # Command-line argument definitions
├── utils.py # Utility functions and helpers
├── scheduler.py # Dynamic job scheduling for adapters
├── train_agent_subprocess.py # Subprocess training management
├── analyze_experiment_performance.py # Debate results analysis
├── read_debate_performance.py # Read debate utils
├── data/ # Dataset storage and splits
├── experiments/ # Experiment outputs and results
└── checkpoints/ # Model checkpoints and adapters

Training Methods

Built on Hugging Face TRL, supports multiple paradigms with majority vote variants:

  • MV-SFT: Supervised fine-tuning on consensus examples
  • MV-GRPO: Reinforcement learning with consensus rewards
  • MV-KTO/DPO: Preference optimization methods

See args.py for complete argument documentation.

Citation

This work was developed at Meta AI in collaboration with Meta Superintelligence Labs and the LIINC Lab at Columbia University.

If you use this framework in your research, please cite:

@misc{samanta2024maca,
title={Internalizing Self-Consistency in Language Models: Multi-Agent Consensus Alignment},
author={Ankur Samanta and Akshayaa Magesh and Youliang Yu and Runzhe Wu and Ayush Jain and Daniel Jiang and Boris Vidolov and Paul Sajda and Yonathan Efroni and Kaveh Hassani},
year={2024},
eprint={2509.15172},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://doi.org/10.48550/arXiv.2509.15172}
}

License

MACA is MIT licensed, as found in the LICENSE file.

About

MACA trains language models to be more consistent reasoners through multi-agent debate and consensus-based reinforcement learning.

Resources

Code of conduct

Contributing

Security policy

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

MACA: Multi-Agent Consensus Alignment

Internalizing Self-Consistency in Language Models through Multi-Agent Debate

Policy

📄 Paper: arXiv:2509.15172

Overview

MACA trains language models to be more consistent reasoners through multi-agent debate and consensus-based reinforcement learning.

Key Features:

  • 🤖 Multi-Agent Debate: Orchestrate debates between agents for improved reasoning
  • 🎯 Consensus Training: Post-train on debate outputs using agreement patterns as rewards
  • Distributed Processing: Multi-GPU parallel training with QLoRA adapters
  • 📊 Analysis Tools: Built-in performance tracking and visualization

Quick Start

Installation

conda env create -f env.yml
pip install -e .

Multi-Agent Training

python main.py --model qwen2b --dataset gsm8k --gpus_per_model 1 --max_concurrent_tasks 4 --train_size 1500 --test_size 500 --lora_r 128 --lora_alpha 128 --dpo --epoch_dpo 3 --batch_dpo 6 --lr_dpo 1e-5 --beta_dpo 0.1 --gradient_accumulation_steps_dpo 4 --seed 1 --wandb

Single-Agent Training

python maca_single_agent.py --output_dir q2b_sa_runs --model qwen2b --phase kto --kto --train_datasets math gsm8k mathqa --test_datasets math gsm8k mathqa svamp gpqa csqa --use_full_test --lora_r_range 64 --lora_alpha_range 64 --lr_kto 1e-5 --evaluation_batch_size 24 --wandb

Key Arguments

  • --model: Quantized base model to use (llama1b/3b/8b, phi4b, qwen2b/7b, gemma4b, mistral7b)
  • --dataset: Dataset (gsm8k, math, mathqa, gpqa, svamp, csqa)
  • --agents: Number of agents in the debate (default: 3)
  • --finetune: Enable Majority Vote Supervised Fine-Tuning (SFT)
  • --post_train: Enable Majority Vote Group-Relative Policy Optimization training (GRPO)
  • --dpo: Enable Majority Vote Direct Preference Optimization training
  • --kto: Enable Majority Vote Kahneman-Tversky Optimization training
  • --use_consensus_reward: Enable consensus-based rewards

Wandb logging

  • --wandb: Enable Weights & Biases logging
  • --project_name: W&B project name (default: llm-marl, requires setting --wandb)
  • --entity_name: W&B entity/team name (default: llm-marl, requires setting --wandb)

Project Structure

maca/
├── main.py # Main training entry point
├── maca_single_agent.py # Single agent hyperparameter tuning and testing
├── model.py # Agent implementation and reward functions
├── debate.py # Multi-agent debate orchestration
├── orchestrator.py # Training coordination and management
├── data.py # Dataset loading and preprocessing
├── parser.py # Answer parsing and grading utilities
├── args.py # Command-line argument definitions
├── utils.py # Utility functions and helpers
├── scheduler.py # Dynamic job scheduling for adapters
├── train_agent_subprocess.py # Subprocess training management
├── analyze_experiment_performance.py # Debate results analysis
├── read_debate_performance.py # Read debate utils
├── data/ # Dataset storage and splits
├── experiments/ # Experiment outputs and results
└── checkpoints/ # Model checkpoints and adapters

Training Methods

Built on Hugging Face TRL, supports multiple paradigms with majority vote variants:

  • MV-SFT: Supervised fine-tuning on consensus examples
  • MV-GRPO: Reinforcement learning with consensus rewards
  • MV-KTO/DPO: Preference optimization methods

See args.py for complete argument documentation.

Citation

This work was developed at Meta AI in collaboration with Meta Superintelligence Labs and the LIINC Lab at Columbia University.

If you use this framework in your research, please cite:

@misc{samanta2024maca,
title={Internalizing Self-Consistency in Language Models: Multi-Agent Consensus Alignment},
author={Ankur Samanta and Akshayaa Magesh and Youliang Yu and Runzhe Wu and Ayush Jain and Daniel Jiang and Boris Vidolov and Paul Sajda and Yonathan Efroni and Kaveh Hassani},
year={2024},
eprint={2509.15172},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://doi.org/10.48550/arXiv.2509.15172}
}

License

MACA is MIT licensed, as found in the LICENSE file.

About

MACA trains language models to be more consistent reasoners through multi-agent debate and consensus-based reinforcement learning.

Resources

Code of conduct

Contributing

Security policy

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MACA: Multi-Agent Consensus Alignment

Internalizing Self-Consistency in Language Models through Multi-Agent Debate

Policy

📄 Paper: arXiv:2509.15172

Overview

MACA trains language models to be more consistent reasoners through multi-agent debate and consensus-based reinforcement learning.

Key Features:

  • 🤖 Multi-Agent Debate: Orchestrate debates between agents for improved reasoning
  • 🎯 Consensus Training: Post-train on debate outputs using agreement patterns as rewards
  • Distributed Processing: Multi-GPU parallel training with QLoRA adapters
  • 📊 Analysis Tools: Built-in performance tracking and visualization

Quick Start

Installation

conda env create -f env.yml
pip install -e .

Multi-Agent Training

python main.py --model qwen2b --dataset gsm8k --gpus_per_model 1 --max_concurrent_tasks 4 --train_size 1500 --test_size 500 --lora_r 128 --lora_alpha 128 --dpo --epoch_dpo 3 --batch_dpo 6 --lr_dpo 1e-5 --beta_dpo 0.1 --gradient_accumulation_steps_dpo 4 --seed 1 --wandb

Single-Agent Training

python maca_single_agent.py --output_dir q2b_sa_runs --model qwen2b --phase kto --kto --train_datasets math gsm8k mathqa --test_datasets math gsm8k mathqa svamp gpqa csqa --use_full_test --lora_r_range 64 --lora_alpha_range 64 --lr_kto 1e-5 --evaluation_batch_size 24 --wandb

Key Arguments

  • --model: Quantized base model to use (llama1b/3b/8b, phi4b, qwen2b/7b, gemma4b, mistral7b)
  • --dataset: Dataset (gsm8k, math, mathqa, gpqa, svamp, csqa)
  • --agents: Number of agents in the debate (default: 3)
  • --finetune: Enable Majority Vote Supervised Fine-Tuning (SFT)
  • --post_train: Enable Majority Vote Group-Relative Policy Optimization training (GRPO)
  • --dpo: Enable Majority Vote Direct Preference Optimization training
  • --kto: Enable Majority Vote Kahneman-Tversky Optimization training
  • --use_consensus_reward: Enable consensus-based rewards

Wandb logging

  • --wandb: Enable Weights & Biases logging
  • --project_name: W&B project name (default: llm-marl, requires setting --wandb)
  • --entity_name: W&B entity/team name (default: llm-marl, requires setting --wandb)

Project Structure

maca/
├── main.py # Main training entry point
├── maca_single_agent.py # Single agent hyperparameter tuning and testing
├── model.py # Agent implementation and reward functions
├── debate.py # Multi-agent debate orchestration
├── orchestrator.py # Training coordination and management
├── data.py # Dataset loading and preprocessing
├── parser.py # Answer parsing and grading utilities
├── args.py # Command-line argument definitions
├── utils.py # Utility functions and helpers
├── scheduler.py # Dynamic job scheduling for adapters
├── train_agent_subprocess.py # Subprocess training management
├── analyze_experiment_performance.py # Debate results analysis
├── read_debate_performance.py # Read debate utils
├── data/ # Dataset storage and splits
├── experiments/ # Experiment outputs and results
└── checkpoints/ # Model checkpoints and adapters

Training Methods

Built on Hugging Face TRL, supports multiple paradigms with majority vote variants:

  • MV-SFT: Supervised fine-tuning on consensus examples
  • MV-GRPO: Reinforcement learning with consensus rewards
  • MV-KTO/DPO: Preference optimization methods

See args.py for complete argument documentation.

Citation

This work was developed at Meta AI in collaboration with Meta Superintelligence Labs and the LIINC Lab at Columbia University.

If you use this framework in your research, please cite:

@misc{samanta2024maca,
title={Internalizing Self-Consistency in Language Models: Multi-Agent Consensus Alignment},
author={Ankur Samanta and Akshayaa Magesh and Youliang Yu and Runzhe Wu and Ayush Jain and Daniel Jiang and Boris Vidolov and Paul Sajda and Yonathan Efroni and Kaveh Hassani},
year={2024},
eprint={2509.15172},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://doi.org/10.48550/arXiv.2509.15172}
}

License

MACA is MIT licensed, as found in the LICENSE file.

About

MACA trains language models to be more consistent reasoners through multi-agent debate and consensus-based reinforcement learning.

Resources

Code of conduct

Contributing

Security policy

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MACA: Multi-Agent Consensus Alignment

Internalizing Self-Consistency in Language Models through Multi-Agent Debate

Policy

📄 Paper: arXiv:2509.15172

Overview

MACA trains language models to be more consistent reasoners through multi-agent debate and consensus-based reinforcement learning.

Key Features:

  • 🤖 Multi-Agent Debate: Orchestrate debates between agents for improved reasoning
  • 🎯 Consensus Training: Post-train on debate outputs using agreement patterns as rewards
  • Distributed Processing: Multi-GPU parallel training with QLoRA adapters
  • 📊 Analysis Tools: Built-in performance tracking and visualization

Quick Start

Installation

conda env create -f env.yml
pip install -e .

Multi-Agent Training

python main.py --model qwen2b --dataset gsm8k --gpus_per_model 1 --max_concurrent_tasks 4 --train_size 1500 --test_size 500 --lora_r 128 --lora_alpha 128 --dpo --epoch_dpo 3 --batch_dpo 6 --lr_dpo 1e-5 --beta_dpo 0.1 --gradient_accumulation_steps_dpo 4 --seed 1 --wandb

Single-Agent Training

python maca_single_agent.py --output_dir q2b_sa_runs --model qwen2b --phase kto --kto --train_datasets math gsm8k mathqa --test_datasets math gsm8k mathqa svamp gpqa csqa --use_full_test --lora_r_range 64 --lora_alpha_range 64 --lr_kto 1e-5 --evaluation_batch_size 24 --wandb

Key Arguments

  • --model: Quantized base model to use (llama1b/3b/8b, phi4b, qwen2b/7b, gemma4b, mistral7b)
  • --dataset: Dataset (gsm8k, math, mathqa, gpqa, svamp, csqa)
  • --agents: Number of agents in the debate (default: 3)
  • --finetune: Enable Majority Vote Supervised Fine-Tuning (SFT)
  • --post_train: Enable Majority Vote Group-Relative Policy Optimization training (GRPO)
  • --dpo: Enable Majority Vote Direct Preference Optimization training
  • --kto: Enable Majority Vote Kahneman-Tversky Optimization training
  • --use_consensus_reward: Enable consensus-based rewards

Wandb logging

  • --wandb: Enable Weights & Biases logging
  • --project_name: W&B project name (default: llm-marl, requires setting --wandb)
  • --entity_name: W&B entity/team name (default: llm-marl, requires setting --wandb)

Project Structure

maca/
├── main.py # Main training entry point
├── maca_single_agent.py # Single agent hyperparameter tuning and testing
├── model.py # Agent implementation and reward functions
├── debate.py # Multi-agent debate orchestration
├── orchestrator.py # Training coordination and management
├── data.py # Dataset loading and preprocessing
├── parser.py # Answer parsing and grading utilities
├── args.py # Command-line argument definitions
├── utils.py # Utility functions and helpers
├── scheduler.py # Dynamic job scheduling for adapters
├── train_agent_subprocess.py # Subprocess training management
├── analyze_experiment_performance.py # Debate results analysis
├── read_debate_performance.py # Read debate utils
├── data/ # Dataset storage and splits
├── experiments/ # Experiment outputs and results
└── checkpoints/ # Model checkpoints and adapters

Training Methods

Built on Hugging Face TRL, supports multiple paradigms with majority vote variants:

  • MV-SFT: Supervised fine-tuning on consensus examples
  • MV-GRPO: Reinforcement learning with consensus rewards
  • MV-KTO/DPO: Preference optimization methods

See args.py for complete argument documentation.

Citation

This work was developed at Meta AI in collaboration with Meta Superintelligence Labs and the LIINC Lab at Columbia University.

If you use this framework in your research, please cite:

@misc{samanta2024maca,
title={Internalizing Self-Consistency in Language Models: Multi-Agent Consensus Alignment},
author={Ankur Samanta and Akshayaa Magesh and Youliang Yu and Runzhe Wu and Ayush Jain and Daniel Jiang and Boris Vidolov and Paul Sajda and Yonathan Efroni and Kaveh Hassani},
year={2024},
eprint={2509.15172},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://doi.org/10.48550/arXiv.2509.15172}
}

License

MACA is MIT licensed, as found in the LICENSE file.

About

MACA trains language models to be more consistent reasoners through multi-agent debate and consensus-based reinforcement learning.

Resources

Code of conduct

Contributing

Security policy

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

MACA: Multi-Agent Consensus Alignment

Internalizing Self-Consistency in Language Models through Multi-Agent Debate

Policy

📄 Paper: arXiv:2509.15172

Overview

MACA trains language models to be more consistent reasoners through multi-agent debate and consensus-based reinforcement learning.

Key Features:

  • 🤖 Multi-Agent Debate: Orchestrate debates between agents for improved reasoning
  • 🎯 Consensus Training: Post-train on debate outputs using agreement patterns as rewards
  • Distributed Processing: Multi-GPU parallel training with QLoRA adapters
  • 📊 Analysis Tools: Built-in performance tracking and visualization

Quick Start

Installation

conda env create -f env.yml
pip install -e .

Multi-Agent Training

python main.py --model qwen2b --dataset gsm8k --gpus_per_model 1 --max_concurrent_tasks 4 --train_size 1500 --test_size 500 --lora_r 128 --lora_alpha 128 --dpo --epoch_dpo 3 --batch_dpo 6 --lr_dpo 1e-5 --beta_dpo 0.1 --gradient_accumulation_steps_dpo 4 --seed 1 --wandb

Single-Agent Training

python maca_single_agent.py --output_dir q2b_sa_runs --model qwen2b --phase kto --kto --train_datasets math gsm8k mathqa --test_datasets math gsm8k mathqa svamp gpqa csqa --use_full_test --lora_r_range 64 --lora_alpha_range 64 --lr_kto 1e-5 --evaluation_batch_size 24 --wandb

Key Arguments

  • --model: Quantized base model to use (llama1b/3b/8b, phi4b, qwen2b/7b, gemma4b, mistral7b)
  • --dataset: Dataset (gsm8k, math, mathqa, gpqa, svamp, csqa)
  • --agents: Number of agents in the debate (default: 3)
  • --finetune: Enable Majority Vote Supervised Fine-Tuning (SFT)
  • --post_train: Enable Majority Vote Group-Relative Policy Optimization training (GRPO)
  • --dpo: Enable Majority Vote Direct Preference Optimization training
  • --kto: Enable Majority Vote Kahneman-Tversky Optimization training
  • --use_consensus_reward: Enable consensus-based rewards

Wandb logging

  • --wandb: Enable Weights & Biases logging
  • --project_name: W&B project name (default: llm-marl, requires setting --wandb)
  • --entity_name: W&B entity/team name (default: llm-marl, requires setting --wandb)

Project Structure

maca/
├── main.py # Main training entry point
├── maca_single_agent.py # Single agent hyperparameter tuning and testing
├── model.py # Agent implementation and reward functions
├── debate.py # Multi-agent debate orchestration
├── orchestrator.py # Training coordination and management
├── data.py # Dataset loading and preprocessing
├── parser.py # Answer parsing and grading utilities
├── args.py # Command-line argument definitions
├── utils.py # Utility functions and helpers
├── scheduler.py # Dynamic job scheduling for adapters
├── train_agent_subprocess.py # Subprocess training management
├── analyze_experiment_performance.py # Debate results analysis
├── read_debate_performance.py # Read debate utils
├── data/ # Dataset storage and splits
├── experiments/ # Experiment outputs and results
└── checkpoints/ # Model checkpoints and adapters

Training Methods

Built on Hugging Face TRL, supports multiple paradigms with majority vote variants:

  • MV-SFT: Supervised fine-tuning on consensus examples
  • MV-GRPO: Reinforcement learning with consensus rewards
  • MV-KTO/DPO: Preference optimization methods

See args.py for complete argument documentation.

Citation

This work was developed at Meta AI in collaboration with Meta Superintelligence Labs and the LIINC Lab at Columbia University.

If you use this framework in your research, please cite:

@misc{samanta2024maca,
title={Internalizing Self-Consistency in Language Models: Multi-Agent Consensus Alignment},
author={Ankur Samanta and Akshayaa Magesh and Youliang Yu and Runzhe Wu and Ayush Jain and Daniel Jiang and Boris Vidolov and Paul Sajda and Yonathan Efroni and Kaveh Hassani},
year={2024},
eprint={2509.15172},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://doi.org/10.48550/arXiv.2509.15172}
}

License

MACA is MIT licensed, as found in the LICENSE file.

About

MACA trains language models to be more consistent reasoners through multi-agent debate and consensus-based reinforcement learning.

Resources

Code of conduct

Contributing

Security policy

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages