Repository files navigation

Code2Math

Can Your Code Agent Effectively Evolve Math Problems Through Exploration?

arXivLicense: MITPython

Code2Math overview

Code2Math is a code-agent framework for synthesizing harder mathematical problems through executable exploration. Starting from seed problems, an evolution agent searches for deeper variants with Python tools, while two verifier agents filter candidates for mathematical solvability and meaningful difficulty increase.

This repository contains the public Code2Math data, prompt templates, and two full code-agent pipeline implementations.

News

Authors

Dadi Guo*, Yuejin Xie*, Qingyu Liu*, Weixian Huang*, Jiayu Liu, Zhiyuan Fan, Qihan Ren, Shuai Shao, Tianyi Zhou, Jianjie Feng, Wenze Su, Yujiu Yang, Dongrui Liu (corresponding author), Yi R. (May) Fung (corresponding author)

*Equal contribution.

Affiliations: Hong Kong University of Science and Technology, Tsinghua University, Zhejiang University, Nanjing Tech University, Shanghai Jiao Tong University, University of Michigan, Independent Researcher.

Repository Layout

Code2Math/
code2math/ # reusable pipeline package
scripts/ # runnable entry points
prompts/ # prompt templates
math_demonstrations/ # few-shot evolution demonstrations
evolved_problems/ # released evolved problem sets
original_problems.json # 100 seed problems

Method

Code2Math decomposes problem evolution into three agents:

  1. Evolution Agent - proposes harder variants of seed problems through thought, empirical inquiry, and Python-backed exploration.
  2. Solvability Verification Agent - rejects ill-formed, inconsistent, or unsolvable candidates.
  3. Difficulty Verification Agent - checks whether the evolved problem increases conceptual discovery burden rather than only adding computation.

The code tools support symbolic computation, graph algorithms, combinatorial enumeration, numerical checks, and constraint-style exploration.

Full Code-Agent Pipelines

We release two complete implementations of the three-stage workflow.

BackendEntry PointWhat It Does
Standard CodeAgentscripts/run_code_agent_pipeline.pyUses the standard smolagents.CodeAgent execution loop.
Interleaved Thinkingscripts/run_interleaved_pipeline.pyPreserves OpenAI-style messages, tool_calls, tool responses, and returned reasoning content in messages.json.

Standard CodeAgent Backend

python scripts/run_code_agent_pipeline.py --max-workers 5

Interleaved-Thinking Backend

python scripts/run_interleaved_pipeline.py --max-workers 5

The interleaved backend is implemented as a small compatibility layer in this repository, so users do not need to install a patched smolagents fork.

Implementation Note

The implementation builds on and follows the design of Hugging Face smolagents, especially its CodeAgent, tool abstraction, Python execution tool, and OpenAI-compatible model wrappers. Code2Math adds the math-evolution prompts, three-agent orchestration, output parsing, released datasets, and an interleaved-thinking backend for preserving tool-call trajectories from thinking-model runs.

Example

Code-driven problem evolution example

Installation

python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env

Edit .env with an OpenAI-compatible endpoint:

URL=https://your-openai-compatible-endpoint/v1KEY=your_api_keyEVOLVE_MODEL=your-evolver-modelVERIFY_MODEL=your-verifier-model

Qwen-style Basic-auth endpoints are also supported through API_AK, API_SK, QWEN_THINKING_URL, and QWEN_INSTRUCT_URL.

Outputs

By default, generated results are written under:

adapted_problems/
codeagent/
interleaved/

Agent logs are written under logs/. These generated outputs are ignored by git.

Each result entry records the pipeline status, failure stage if any, evolved problem, verifier outputs, and final difficulty judgment.

Data Format

Seed problems:

{
"problem_description": "...",
"solution_steps": "...",
"answer": "..."
}

Evolved problem records:

{
"status": "success",
"result_data": {
"status": true,
"new_problem": {
"new_problem": "...",
"new_solution_steps": "...",
"new_answer": "..."
},
"solvability_verifier_output": {},
"difficulty_verifier_output": {}
}
}

Scope

This release focuses on the two full code-agent workflows. The later single-turn no-code pipeline used in internal experiments is intentionally not included.

Please do not commit .env, logs, or generated run outputs.

Citation

If you use Code2Math data, prompts, or code, please cite:

@misc{guo2026code2mathcodeagenteffectively,
title={Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?},
author={Dadi Guo and Yuejin Xie and Qingyu Liu and Weixian Huang and Jiayu Liu and Zhiyuan Fan and Qihan Ren and Shuai Shao and Tianyi Zhou and Jianjie Feng and Wenze Su and Yujiu Yang and Dongrui Liu and Yi R. Fung},
year={2026},
eprint={2603.03202},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2603.03202},
}

License

MIT License. See LICENSE.

About

Evolved math problems and prompt templates from Code2Math project

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Code2Math

Can Your Code Agent Effectively Evolve Math Problems Through Exploration?

arXivLicense: MITPython

Code2Math overview

Code2Math is a code-agent framework for synthesizing harder mathematical problems through executable exploration. Starting from seed problems, an evolution agent searches for deeper variants with Python tools, while two verifier agents filter candidates for mathematical solvability and meaningful difficulty increase.

This repository contains the public Code2Math data, prompt templates, and two full code-agent pipeline implementations.

News

Authors

Dadi Guo*, Yuejin Xie*, Qingyu Liu*, Weixian Huang*, Jiayu Liu, Zhiyuan Fan, Qihan Ren, Shuai Shao, Tianyi Zhou, Jianjie Feng, Wenze Su, Yujiu Yang, Dongrui Liu (corresponding author), Yi R. (May) Fung (corresponding author)

*Equal contribution.

Affiliations: Hong Kong University of Science and Technology, Tsinghua University, Zhejiang University, Nanjing Tech University, Shanghai Jiao Tong University, University of Michigan, Independent Researcher.

Repository Layout

Code2Math/
code2math/ # reusable pipeline package
scripts/ # runnable entry points
prompts/ # prompt templates
math_demonstrations/ # few-shot evolution demonstrations
evolved_problems/ # released evolved problem sets
original_problems.json # 100 seed problems

Method

Code2Math decomposes problem evolution into three agents:

  1. Evolution Agent - proposes harder variants of seed problems through thought, empirical inquiry, and Python-backed exploration.
  2. Solvability Verification Agent - rejects ill-formed, inconsistent, or unsolvable candidates.
  3. Difficulty Verification Agent - checks whether the evolved problem increases conceptual discovery burden rather than only adding computation.

The code tools support symbolic computation, graph algorithms, combinatorial enumeration, numerical checks, and constraint-style exploration.

Full Code-Agent Pipelines

We release two complete implementations of the three-stage workflow.

BackendEntry PointWhat It Does
Standard CodeAgentscripts/run_code_agent_pipeline.pyUses the standard smolagents.CodeAgent execution loop.
Interleaved Thinkingscripts/run_interleaved_pipeline.pyPreserves OpenAI-style messages, tool_calls, tool responses, and returned reasoning content in messages.json.

Standard CodeAgent Backend

python scripts/run_code_agent_pipeline.py --max-workers 5

Interleaved-Thinking Backend

python scripts/run_interleaved_pipeline.py --max-workers 5

The interleaved backend is implemented as a small compatibility layer in this repository, so users do not need to install a patched smolagents fork.

Implementation Note

The implementation builds on and follows the design of Hugging Face smolagents, especially its CodeAgent, tool abstraction, Python execution tool, and OpenAI-compatible model wrappers. Code2Math adds the math-evolution prompts, three-agent orchestration, output parsing, released datasets, and an interleaved-thinking backend for preserving tool-call trajectories from thinking-model runs.

Example

Code-driven problem evolution example

Installation

python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env

Edit .env with an OpenAI-compatible endpoint:

URL=https://your-openai-compatible-endpoint/v1KEY=your_api_keyEVOLVE_MODEL=your-evolver-modelVERIFY_MODEL=your-verifier-model

Qwen-style Basic-auth endpoints are also supported through API_AK, API_SK, QWEN_THINKING_URL, and QWEN_INSTRUCT_URL.

Outputs

By default, generated results are written under:

adapted_problems/
codeagent/
interleaved/

Agent logs are written under logs/. These generated outputs are ignored by git.

Each result entry records the pipeline status, failure stage if any, evolved problem, verifier outputs, and final difficulty judgment.

Data Format

Seed problems:

{
"problem_description": "...",
"solution_steps": "...",
"answer": "..."
}

Evolved problem records:

{
"status": "success",
"result_data": {
"status": true,
"new_problem": {
"new_problem": "...",
"new_solution_steps": "...",
"new_answer": "..."
},
"solvability_verifier_output": {},
"difficulty_verifier_output": {}
}
}

Scope

This release focuses on the two full code-agent workflows. The later single-turn no-code pipeline used in internal experiments is intentionally not included.

Please do not commit .env, logs, or generated run outputs.

Citation

If you use Code2Math data, prompts, or code, please cite:

@misc{guo2026code2mathcodeagenteffectively,
title={Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?},
author={Dadi Guo and Yuejin Xie and Qingyu Liu and Weixian Huang and Jiayu Liu and Zhiyuan Fan and Qihan Ren and Shuai Shao and Tianyi Zhou and Jianjie Feng and Wenze Su and Yujiu Yang and Dongrui Liu and Yi R. Fung},
year={2026},
eprint={2603.03202},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2603.03202},
}

License

MIT License. See LICENSE.

About

Evolved math problems and prompt templates from Code2Math project

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Code2Math

Can Your Code Agent Effectively Evolve Math Problems Through Exploration?

arXivLicense: MITPython

Code2Math overview

Code2Math is a code-agent framework for synthesizing harder mathematical problems through executable exploration. Starting from seed problems, an evolution agent searches for deeper variants with Python tools, while two verifier agents filter candidates for mathematical solvability and meaningful difficulty increase.

This repository contains the public Code2Math data, prompt templates, and two full code-agent pipeline implementations.

News

Authors

Dadi Guo*, Yuejin Xie*, Qingyu Liu*, Weixian Huang*, Jiayu Liu, Zhiyuan Fan, Qihan Ren, Shuai Shao, Tianyi Zhou, Jianjie Feng, Wenze Su, Yujiu Yang, Dongrui Liu (corresponding author), Yi R. (May) Fung (corresponding author)

*Equal contribution.

Affiliations: Hong Kong University of Science and Technology, Tsinghua University, Zhejiang University, Nanjing Tech University, Shanghai Jiao Tong University, University of Michigan, Independent Researcher.

Repository Layout

Code2Math/
code2math/ # reusable pipeline package
scripts/ # runnable entry points
prompts/ # prompt templates
math_demonstrations/ # few-shot evolution demonstrations
evolved_problems/ # released evolved problem sets
original_problems.json # 100 seed problems

Method

Code2Math decomposes problem evolution into three agents:

  1. Evolution Agent - proposes harder variants of seed problems through thought, empirical inquiry, and Python-backed exploration.
  2. Solvability Verification Agent - rejects ill-formed, inconsistent, or unsolvable candidates.
  3. Difficulty Verification Agent - checks whether the evolved problem increases conceptual discovery burden rather than only adding computation.

The code tools support symbolic computation, graph algorithms, combinatorial enumeration, numerical checks, and constraint-style exploration.

Full Code-Agent Pipelines

We release two complete implementations of the three-stage workflow.

BackendEntry PointWhat It Does
Standard CodeAgentscripts/run_code_agent_pipeline.pyUses the standard smolagents.CodeAgent execution loop.
Interleaved Thinkingscripts/run_interleaved_pipeline.pyPreserves OpenAI-style messages, tool_calls, tool responses, and returned reasoning content in messages.json.

Standard CodeAgent Backend

python scripts/run_code_agent_pipeline.py --max-workers 5

Interleaved-Thinking Backend

python scripts/run_interleaved_pipeline.py --max-workers 5

The interleaved backend is implemented as a small compatibility layer in this repository, so users do not need to install a patched smolagents fork.

Implementation Note

The implementation builds on and follows the design of Hugging Face smolagents, especially its CodeAgent, tool abstraction, Python execution tool, and OpenAI-compatible model wrappers. Code2Math adds the math-evolution prompts, three-agent orchestration, output parsing, released datasets, and an interleaved-thinking backend for preserving tool-call trajectories from thinking-model runs.

Example

Code-driven problem evolution example

Installation

python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env

Edit .env with an OpenAI-compatible endpoint:

URL=https://your-openai-compatible-endpoint/v1KEY=your_api_keyEVOLVE_MODEL=your-evolver-modelVERIFY_MODEL=your-verifier-model

Qwen-style Basic-auth endpoints are also supported through API_AK, API_SK, QWEN_THINKING_URL, and QWEN_INSTRUCT_URL.

Outputs

By default, generated results are written under:

adapted_problems/
codeagent/
interleaved/

Agent logs are written under logs/. These generated outputs are ignored by git.

Each result entry records the pipeline status, failure stage if any, evolved problem, verifier outputs, and final difficulty judgment.

Data Format

Seed problems:

{
"problem_description": "...",
"solution_steps": "...",
"answer": "..."
}

Evolved problem records:

{
"status": "success",
"result_data": {
"status": true,
"new_problem": {
"new_problem": "...",
"new_solution_steps": "...",
"new_answer": "..."
},
"solvability_verifier_output": {},
"difficulty_verifier_output": {}
}
}

Scope

This release focuses on the two full code-agent workflows. The later single-turn no-code pipeline used in internal experiments is intentionally not included.

Please do not commit .env, logs, or generated run outputs.

Citation

If you use Code2Math data, prompts, or code, please cite:

@misc{guo2026code2mathcodeagenteffectively,
title={Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?},
author={Dadi Guo and Yuejin Xie and Qingyu Liu and Weixian Huang and Jiayu Liu and Zhiyuan Fan and Qihan Ren and Shuai Shao and Tianyi Zhou and Jianjie Feng and Wenze Su and Yujiu Yang and Dongrui Liu and Yi R. Fung},
year={2026},
eprint={2603.03202},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2603.03202},
}

License

MIT License. See LICENSE.

About

Evolved math problems and prompt templates from Code2Math project

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Code2Math

Can Your Code Agent Effectively Evolve Math Problems Through Exploration?

arXivLicense: MITPython

Code2Math overview

Code2Math is a code-agent framework for synthesizing harder mathematical problems through executable exploration. Starting from seed problems, an evolution agent searches for deeper variants with Python tools, while two verifier agents filter candidates for mathematical solvability and meaningful difficulty increase.

This repository contains the public Code2Math data, prompt templates, and two full code-agent pipeline implementations.

News

Authors

Dadi Guo*, Yuejin Xie*, Qingyu Liu*, Weixian Huang*, Jiayu Liu, Zhiyuan Fan, Qihan Ren, Shuai Shao, Tianyi Zhou, Jianjie Feng, Wenze Su, Yujiu Yang, Dongrui Liu (corresponding author), Yi R. (May) Fung (corresponding author)

*Equal contribution.

Affiliations: Hong Kong University of Science and Technology, Tsinghua University, Zhejiang University, Nanjing Tech University, Shanghai Jiao Tong University, University of Michigan, Independent Researcher.

Repository Layout

Code2Math/
code2math/ # reusable pipeline package
scripts/ # runnable entry points
prompts/ # prompt templates
math_demonstrations/ # few-shot evolution demonstrations
evolved_problems/ # released evolved problem sets
original_problems.json # 100 seed problems

Method

Code2Math decomposes problem evolution into three agents:

  1. Evolution Agent - proposes harder variants of seed problems through thought, empirical inquiry, and Python-backed exploration.
  2. Solvability Verification Agent - rejects ill-formed, inconsistent, or unsolvable candidates.
  3. Difficulty Verification Agent - checks whether the evolved problem increases conceptual discovery burden rather than only adding computation.

The code tools support symbolic computation, graph algorithms, combinatorial enumeration, numerical checks, and constraint-style exploration.

Full Code-Agent Pipelines

We release two complete implementations of the three-stage workflow.

BackendEntry PointWhat It Does
Standard CodeAgentscripts/run_code_agent_pipeline.pyUses the standard smolagents.CodeAgent execution loop.
Interleaved Thinkingscripts/run_interleaved_pipeline.pyPreserves OpenAI-style messages, tool_calls, tool responses, and returned reasoning content in messages.json.

Standard CodeAgent Backend

python scripts/run_code_agent_pipeline.py --max-workers 5

Interleaved-Thinking Backend

python scripts/run_interleaved_pipeline.py --max-workers 5

The interleaved backend is implemented as a small compatibility layer in this repository, so users do not need to install a patched smolagents fork.

Implementation Note

The implementation builds on and follows the design of Hugging Face smolagents, especially its CodeAgent, tool abstraction, Python execution tool, and OpenAI-compatible model wrappers. Code2Math adds the math-evolution prompts, three-agent orchestration, output parsing, released datasets, and an interleaved-thinking backend for preserving tool-call trajectories from thinking-model runs.

Example

Code-driven problem evolution example

Installation

python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env

Edit .env with an OpenAI-compatible endpoint:

URL=https://your-openai-compatible-endpoint/v1KEY=your_api_keyEVOLVE_MODEL=your-evolver-modelVERIFY_MODEL=your-verifier-model

Qwen-style Basic-auth endpoints are also supported through API_AK, API_SK, QWEN_THINKING_URL, and QWEN_INSTRUCT_URL.

Outputs

By default, generated results are written under:

adapted_problems/
codeagent/
interleaved/

Agent logs are written under logs/. These generated outputs are ignored by git.

Each result entry records the pipeline status, failure stage if any, evolved problem, verifier outputs, and final difficulty judgment.

Data Format

Seed problems:

{
"problem_description": "...",
"solution_steps": "...",
"answer": "..."
}

Evolved problem records:

{
"status": "success",
"result_data": {
"status": true,
"new_problem": {
"new_problem": "...",
"new_solution_steps": "...",
"new_answer": "..."
},
"solvability_verifier_output": {},
"difficulty_verifier_output": {}
}
}

Scope

This release focuses on the two full code-agent workflows. The later single-turn no-code pipeline used in internal experiments is intentionally not included.

Please do not commit .env, logs, or generated run outputs.

Citation

If you use Code2Math data, prompts, or code, please cite:

@misc{guo2026code2mathcodeagenteffectively,
title={Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?},
author={Dadi Guo and Yuejin Xie and Qingyu Liu and Weixian Huang and Jiayu Liu and Zhiyuan Fan and Qihan Ren and Shuai Shao and Tianyi Zhou and Jianjie Feng and Wenze Su and Yujiu Yang and Dongrui Liu and Yi R. Fung},
year={2026},
eprint={2603.03202},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2603.03202},
}

License

MIT License. See LICENSE.

About

Evolved math problems and prompt templates from Code2Math project

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Code2Math

Can Your Code Agent Effectively Evolve Math Problems Through Exploration?

arXivLicense: MITPython

Code2Math overview

Code2Math is a code-agent framework for synthesizing harder mathematical problems through executable exploration. Starting from seed problems, an evolution agent searches for deeper variants with Python tools, while two verifier agents filter candidates for mathematical solvability and meaningful difficulty increase.

This repository contains the public Code2Math data, prompt templates, and two full code-agent pipeline implementations.

News

Authors

Dadi Guo*, Yuejin Xie*, Qingyu Liu*, Weixian Huang*, Jiayu Liu, Zhiyuan Fan, Qihan Ren, Shuai Shao, Tianyi Zhou, Jianjie Feng, Wenze Su, Yujiu Yang, Dongrui Liu (corresponding author), Yi R. (May) Fung (corresponding author)

*Equal contribution.

Affiliations: Hong Kong University of Science and Technology, Tsinghua University, Zhejiang University, Nanjing Tech University, Shanghai Jiao Tong University, University of Michigan, Independent Researcher.

Repository Layout

Code2Math/
code2math/ # reusable pipeline package
scripts/ # runnable entry points
prompts/ # prompt templates
math_demonstrations/ # few-shot evolution demonstrations
evolved_problems/ # released evolved problem sets
original_problems.json # 100 seed problems

Method

Code2Math decomposes problem evolution into three agents:

  1. Evolution Agent - proposes harder variants of seed problems through thought, empirical inquiry, and Python-backed exploration.
  2. Solvability Verification Agent - rejects ill-formed, inconsistent, or unsolvable candidates.
  3. Difficulty Verification Agent - checks whether the evolved problem increases conceptual discovery burden rather than only adding computation.

The code tools support symbolic computation, graph algorithms, combinatorial enumeration, numerical checks, and constraint-style exploration.

Full Code-Agent Pipelines

We release two complete implementations of the three-stage workflow.

BackendEntry PointWhat It Does
Standard CodeAgentscripts/run_code_agent_pipeline.pyUses the standard smolagents.CodeAgent execution loop.
Interleaved Thinkingscripts/run_interleaved_pipeline.pyPreserves OpenAI-style messages, tool_calls, tool responses, and returned reasoning content in messages.json.

Standard CodeAgent Backend

python scripts/run_code_agent_pipeline.py --max-workers 5

Interleaved-Thinking Backend

python scripts/run_interleaved_pipeline.py --max-workers 5

The interleaved backend is implemented as a small compatibility layer in this repository, so users do not need to install a patched smolagents fork.

Implementation Note

The implementation builds on and follows the design of Hugging Face smolagents, especially its CodeAgent, tool abstraction, Python execution tool, and OpenAI-compatible model wrappers. Code2Math adds the math-evolution prompts, three-agent orchestration, output parsing, released datasets, and an interleaved-thinking backend for preserving tool-call trajectories from thinking-model runs.

Example

Code-driven problem evolution example

Installation

python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env

Edit .env with an OpenAI-compatible endpoint:

URL=https://your-openai-compatible-endpoint/v1KEY=your_api_keyEVOLVE_MODEL=your-evolver-modelVERIFY_MODEL=your-verifier-model

Qwen-style Basic-auth endpoints are also supported through API_AK, API_SK, QWEN_THINKING_URL, and QWEN_INSTRUCT_URL.

Outputs

By default, generated results are written under:

adapted_problems/
codeagent/
interleaved/

Agent logs are written under logs/. These generated outputs are ignored by git.

Each result entry records the pipeline status, failure stage if any, evolved problem, verifier outputs, and final difficulty judgment.

Data Format

Seed problems:

{
"problem_description": "...",
"solution_steps": "...",
"answer": "..."
}

Evolved problem records:

{
"status": "success",
"result_data": {
"status": true,
"new_problem": {
"new_problem": "...",
"new_solution_steps": "...",
"new_answer": "..."
},
"solvability_verifier_output": {},
"difficulty_verifier_output": {}
}
}

Scope

This release focuses on the two full code-agent workflows. The later single-turn no-code pipeline used in internal experiments is intentionally not included.

Please do not commit .env, logs, or generated run outputs.

Citation

If you use Code2Math data, prompts, or code, please cite:

@misc{guo2026code2mathcodeagenteffectively,
title={Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?},
author={Dadi Guo and Yuejin Xie and Qingyu Liu and Weixian Huang and Jiayu Liu and Zhiyuan Fan and Qihan Ren and Shuai Shao and Tianyi Zhou and Jianjie Feng and Wenze Su and Yujiu Yang and Dongrui Liu and Yi R. Fung},
year={2026},
eprint={2603.03202},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2603.03202},
}

License

MIT License. See LICENSE.

About

Evolved math problems and prompt templates from Code2Math project

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Code2Math

Can Your Code Agent Effectively Evolve Math Problems Through Exploration?

arXivLicense: MITPython

Code2Math overview

Code2Math is a code-agent framework for synthesizing harder mathematical problems through executable exploration. Starting from seed problems, an evolution agent searches for deeper variants with Python tools, while two verifier agents filter candidates for mathematical solvability and meaningful difficulty increase.

This repository contains the public Code2Math data, prompt templates, and two full code-agent pipeline implementations.

News

Authors

Dadi Guo*, Yuejin Xie*, Qingyu Liu*, Weixian Huang*, Jiayu Liu, Zhiyuan Fan, Qihan Ren, Shuai Shao, Tianyi Zhou, Jianjie Feng, Wenze Su, Yujiu Yang, Dongrui Liu (corresponding author), Yi R. (May) Fung (corresponding author)

*Equal contribution.

Affiliations: Hong Kong University of Science and Technology, Tsinghua University, Zhejiang University, Nanjing Tech University, Shanghai Jiao Tong University, University of Michigan, Independent Researcher.

Repository Layout

Code2Math/
code2math/ # reusable pipeline package
scripts/ # runnable entry points
prompts/ # prompt templates
math_demonstrations/ # few-shot evolution demonstrations
evolved_problems/ # released evolved problem sets
original_problems.json # 100 seed problems

Method

Code2Math decomposes problem evolution into three agents:

  1. Evolution Agent - proposes harder variants of seed problems through thought, empirical inquiry, and Python-backed exploration.
  2. Solvability Verification Agent - rejects ill-formed, inconsistent, or unsolvable candidates.
  3. Difficulty Verification Agent - checks whether the evolved problem increases conceptual discovery burden rather than only adding computation.

The code tools support symbolic computation, graph algorithms, combinatorial enumeration, numerical checks, and constraint-style exploration.

Full Code-Agent Pipelines

We release two complete implementations of the three-stage workflow.

BackendEntry PointWhat It Does
Standard CodeAgentscripts/run_code_agent_pipeline.pyUses the standard smolagents.CodeAgent execution loop.
Interleaved Thinkingscripts/run_interleaved_pipeline.pyPreserves OpenAI-style messages, tool_calls, tool responses, and returned reasoning content in messages.json.

Standard CodeAgent Backend

python scripts/run_code_agent_pipeline.py --max-workers 5

Interleaved-Thinking Backend

python scripts/run_interleaved_pipeline.py --max-workers 5

The interleaved backend is implemented as a small compatibility layer in this repository, so users do not need to install a patched smolagents fork.

Implementation Note

The implementation builds on and follows the design of Hugging Face smolagents, especially its CodeAgent, tool abstraction, Python execution tool, and OpenAI-compatible model wrappers. Code2Math adds the math-evolution prompts, three-agent orchestration, output parsing, released datasets, and an interleaved-thinking backend for preserving tool-call trajectories from thinking-model runs.

Example

Code-driven problem evolution example

Installation

python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env

Edit .env with an OpenAI-compatible endpoint:

URL=https://your-openai-compatible-endpoint/v1KEY=your_api_keyEVOLVE_MODEL=your-evolver-modelVERIFY_MODEL=your-verifier-model

Qwen-style Basic-auth endpoints are also supported through API_AK, API_SK, QWEN_THINKING_URL, and QWEN_INSTRUCT_URL.

Outputs

By default, generated results are written under:

adapted_problems/
codeagent/
interleaved/

Agent logs are written under logs/. These generated outputs are ignored by git.

Each result entry records the pipeline status, failure stage if any, evolved problem, verifier outputs, and final difficulty judgment.

Data Format

Seed problems:

{
"problem_description": "...",
"solution_steps": "...",
"answer": "..."
}

Evolved problem records:

{
"status": "success",
"result_data": {
"status": true,
"new_problem": {
"new_problem": "...",
"new_solution_steps": "...",
"new_answer": "..."
},
"solvability_verifier_output": {},
"difficulty_verifier_output": {}
}
}

Scope

This release focuses on the two full code-agent workflows. The later single-turn no-code pipeline used in internal experiments is intentionally not included.

Please do not commit .env, logs, or generated run outputs.

Citation

If you use Code2Math data, prompts, or code, please cite:

@misc{guo2026code2mathcodeagenteffectively,
title={Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?},
author={Dadi Guo and Yuejin Xie and Qingyu Liu and Weixian Huang and Jiayu Liu and Zhiyuan Fan and Qihan Ren and Shuai Shao and Tianyi Zhou and Jianjie Feng and Wenze Su and Yujiu Yang and Dongrui Liu and Yi R. Fung},
year={2026},
eprint={2603.03202},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2603.03202},
}

License

MIT License. See LICENSE.

About

Evolved math problems and prompt templates from Code2Math project

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Code2Math

Can Your Code Agent Effectively Evolve Math Problems Through Exploration?

arXivLicense: MITPython

Code2Math overview

Code2Math is a code-agent framework for synthesizing harder mathematical problems through executable exploration. Starting from seed problems, an evolution agent searches for deeper variants with Python tools, while two verifier agents filter candidates for mathematical solvability and meaningful difficulty increase.

This repository contains the public Code2Math data, prompt templates, and two full code-agent pipeline implementations.

News

Authors

Dadi Guo*, Yuejin Xie*, Qingyu Liu*, Weixian Huang*, Jiayu Liu, Zhiyuan Fan, Qihan Ren, Shuai Shao, Tianyi Zhou, Jianjie Feng, Wenze Su, Yujiu Yang, Dongrui Liu (corresponding author), Yi R. (May) Fung (corresponding author)

*Equal contribution.

Affiliations: Hong Kong University of Science and Technology, Tsinghua University, Zhejiang University, Nanjing Tech University, Shanghai Jiao Tong University, University of Michigan, Independent Researcher.

Repository Layout

Code2Math/
code2math/ # reusable pipeline package
scripts/ # runnable entry points
prompts/ # prompt templates
math_demonstrations/ # few-shot evolution demonstrations
evolved_problems/ # released evolved problem sets
original_problems.json # 100 seed problems

Method

Code2Math decomposes problem evolution into three agents:

  1. Evolution Agent - proposes harder variants of seed problems through thought, empirical inquiry, and Python-backed exploration.
  2. Solvability Verification Agent - rejects ill-formed, inconsistent, or unsolvable candidates.
  3. Difficulty Verification Agent - checks whether the evolved problem increases conceptual discovery burden rather than only adding computation.

The code tools support symbolic computation, graph algorithms, combinatorial enumeration, numerical checks, and constraint-style exploration.

Full Code-Agent Pipelines

We release two complete implementations of the three-stage workflow.

BackendEntry PointWhat It Does
Standard CodeAgentscripts/run_code_agent_pipeline.pyUses the standard smolagents.CodeAgent execution loop.
Interleaved Thinkingscripts/run_interleaved_pipeline.pyPreserves OpenAI-style messages, tool_calls, tool responses, and returned reasoning content in messages.json.

Standard CodeAgent Backend

python scripts/run_code_agent_pipeline.py --max-workers 5

Interleaved-Thinking Backend

python scripts/run_interleaved_pipeline.py --max-workers 5

The interleaved backend is implemented as a small compatibility layer in this repository, so users do not need to install a patched smolagents fork.

Implementation Note

The implementation builds on and follows the design of Hugging Face smolagents, especially its CodeAgent, tool abstraction, Python execution tool, and OpenAI-compatible model wrappers. Code2Math adds the math-evolution prompts, three-agent orchestration, output parsing, released datasets, and an interleaved-thinking backend for preserving tool-call trajectories from thinking-model runs.

Example

Code-driven problem evolution example

Installation

python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env

Edit .env with an OpenAI-compatible endpoint:

URL=https://your-openai-compatible-endpoint/v1KEY=your_api_keyEVOLVE_MODEL=your-evolver-modelVERIFY_MODEL=your-verifier-model

Qwen-style Basic-auth endpoints are also supported through API_AK, API_SK, QWEN_THINKING_URL, and QWEN_INSTRUCT_URL.

Outputs

By default, generated results are written under:

adapted_problems/
codeagent/
interleaved/

Agent logs are written under logs/. These generated outputs are ignored by git.

Each result entry records the pipeline status, failure stage if any, evolved problem, verifier outputs, and final difficulty judgment.

Data Format

Seed problems:

{
"problem_description": "...",
"solution_steps": "...",
"answer": "..."
}

Evolved problem records:

{
"status": "success",
"result_data": {
"status": true,
"new_problem": {
"new_problem": "...",
"new_solution_steps": "...",
"new_answer": "..."
},
"solvability_verifier_output": {},
"difficulty_verifier_output": {}
}
}

Scope

This release focuses on the two full code-agent workflows. The later single-turn no-code pipeline used in internal experiments is intentionally not included.

Please do not commit .env, logs, or generated run outputs.

Citation

If you use Code2Math data, prompts, or code, please cite:

@misc{guo2026code2mathcodeagenteffectively,
title={Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?},
author={Dadi Guo and Yuejin Xie and Qingyu Liu and Weixian Huang and Jiayu Liu and Zhiyuan Fan and Qihan Ren and Shuai Shao and Tianyi Zhou and Jianjie Feng and Wenze Su and Yujiu Yang and Dongrui Liu and Yi R. Fung},
year={2026},
eprint={2603.03202},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2603.03202},
}

License

MIT License. See LICENSE.

About

Evolved math problems and prompt templates from Code2Math project

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Code2Math

Can Your Code Agent Effectively Evolve Math Problems Through Exploration?

arXivLicense: MITPython

Code2Math overview

Code2Math is a code-agent framework for synthesizing harder mathematical problems through executable exploration. Starting from seed problems, an evolution agent searches for deeper variants with Python tools, while two verifier agents filter candidates for mathematical solvability and meaningful difficulty increase.

This repository contains the public Code2Math data, prompt templates, and two full code-agent pipeline implementations.

News

Authors

Dadi Guo*, Yuejin Xie*, Qingyu Liu*, Weixian Huang*, Jiayu Liu, Zhiyuan Fan, Qihan Ren, Shuai Shao, Tianyi Zhou, Jianjie Feng, Wenze Su, Yujiu Yang, Dongrui Liu (corresponding author), Yi R. (May) Fung (corresponding author)

*Equal contribution.

Affiliations: Hong Kong University of Science and Technology, Tsinghua University, Zhejiang University, Nanjing Tech University, Shanghai Jiao Tong University, University of Michigan, Independent Researcher.

Repository Layout

Code2Math/
code2math/ # reusable pipeline package
scripts/ # runnable entry points
prompts/ # prompt templates
math_demonstrations/ # few-shot evolution demonstrations
evolved_problems/ # released evolved problem sets
original_problems.json # 100 seed problems

Method

Code2Math decomposes problem evolution into three agents:

  1. Evolution Agent - proposes harder variants of seed problems through thought, empirical inquiry, and Python-backed exploration.
  2. Solvability Verification Agent - rejects ill-formed, inconsistent, or unsolvable candidates.
  3. Difficulty Verification Agent - checks whether the evolved problem increases conceptual discovery burden rather than only adding computation.

The code tools support symbolic computation, graph algorithms, combinatorial enumeration, numerical checks, and constraint-style exploration.

Full Code-Agent Pipelines

We release two complete implementations of the three-stage workflow.

BackendEntry PointWhat It Does
Standard CodeAgentscripts/run_code_agent_pipeline.pyUses the standard smolagents.CodeAgent execution loop.
Interleaved Thinkingscripts/run_interleaved_pipeline.pyPreserves OpenAI-style messages, tool_calls, tool responses, and returned reasoning content in messages.json.

Standard CodeAgent Backend

python scripts/run_code_agent_pipeline.py --max-workers 5

Interleaved-Thinking Backend

python scripts/run_interleaved_pipeline.py --max-workers 5

The interleaved backend is implemented as a small compatibility layer in this repository, so users do not need to install a patched smolagents fork.

Implementation Note

The implementation builds on and follows the design of Hugging Face smolagents, especially its CodeAgent, tool abstraction, Python execution tool, and OpenAI-compatible model wrappers. Code2Math adds the math-evolution prompts, three-agent orchestration, output parsing, released datasets, and an interleaved-thinking backend for preserving tool-call trajectories from thinking-model runs.

Example

Code-driven problem evolution example

Installation

python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env

Edit .env with an OpenAI-compatible endpoint:

URL=https://your-openai-compatible-endpoint/v1KEY=your_api_keyEVOLVE_MODEL=your-evolver-modelVERIFY_MODEL=your-verifier-model

Qwen-style Basic-auth endpoints are also supported through API_AK, API_SK, QWEN_THINKING_URL, and QWEN_INSTRUCT_URL.

Outputs

By default, generated results are written under:

adapted_problems/
codeagent/
interleaved/

Agent logs are written under logs/. These generated outputs are ignored by git.

Each result entry records the pipeline status, failure stage if any, evolved problem, verifier outputs, and final difficulty judgment.

Data Format

Seed problems:

{
"problem_description": "...",
"solution_steps": "...",
"answer": "..."
}

Evolved problem records:

{
"status": "success",
"result_data": {
"status": true,
"new_problem": {
"new_problem": "...",
"new_solution_steps": "...",
"new_answer": "..."
},
"solvability_verifier_output": {},
"difficulty_verifier_output": {}
}
}

Scope

This release focuses on the two full code-agent workflows. The later single-turn no-code pipeline used in internal experiments is intentionally not included.

Please do not commit .env, logs, or generated run outputs.

Citation

If you use Code2Math data, prompts, or code, please cite:

@misc{guo2026code2mathcodeagenteffectively,
title={Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?},
author={Dadi Guo and Yuejin Xie and Qingyu Liu and Weixian Huang and Jiayu Liu and Zhiyuan Fan and Qihan Ren and Shuai Shao and Tianyi Zhou and Jianjie Feng and Wenze Su and Yujiu Yang and Dongrui Liu and Yi R. Fung},
year={2026},
eprint={2603.03202},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2603.03202},
}

License

MIT License. See LICENSE.

About

Evolved math problems and prompt templates from Code2Math project

Resources

Stars

12 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages