Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

CrowdMath

This repository hosts resources for the paper:

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions

Overview

This dataset captures the collaborative mathematical reasoning that took place on the MIT PRIMES CrowdMath online research program (2016-2025). In CrowdMath, high-school students work together on open research problems in mathematics, posting conjectures, proofs, error corrections, and questions on a shared message board (Art of Problem Solving / AoPS).

Expert annotators read every thread and labeled each post with its role in the mathematical discourse. Those annotations were then grouped into progress chains -- sequences of posts that together advance a single mathematical result from an initial claim to a verified proof.

The dataset file dataset/dataset.json is a JSON array of progress chains.

Example progress chain from the CrowdMath dataset
An example progress chain showing how student posts are annotated with discourse labels.

File format

dataset/dataset.json is a JSON array. Each element is a progress chain object with the fields described below.

Progress chain fields

FieldTypeDescription
result_idstringUnique identifier for the chain, formatted as <topic_id>-<post_number> with an optional letter suffix (e.g. "1228277-14", "1320553-27a").
project_idstringCrowdMath project identifier (e.g. "mitprimes2016", "mitprimes2021"). Each project corresponds to one year of the program; some years have multiple projects (e.g. "mitprimes2024" and "mitprimes2024-2").
open_problem_refstring or nullThe open problem reference as written by the annotator (e.g. "2020-6", "2017-5"). Null when the chain could not be matched to a specific open problem.
open_problem_yearinteger or nullFour-digit year of the matched open problem. Null when unmatched.
open_problem_numberinteger or nullProblem number within that year's problem set. Null when unmatched.
problem_textstring or nullThe mathematical problem statement that this chain addresses. When an open problem was matched, this includes official problem text from the CrowdMath problem set. Otherwise it contains the text of the first "Problem" post in the thread.
problem_resourcesstring or nullSupplementary resources (references, hints, relevant background) associated with the open problem. Non-null for only a small number of chains.
postsarrayOrdered list of post objects that form this chain (see below). The Problem post is excluded from this array because its content is captured in problem_text. Posts are sorted by post_order.

Post fields

Each element of the posts array is an object with these fields:

FieldTypeDescription
topic_idstringAoPS forum thread ID.
post_numberinteger1-based position of this post within its thread.
post_orderintegerGlobal ordering index across the entire dataset. Use this field to sort posts chronologically.
post_idstringAoPS database identifier for this post.
project_idstringCrowdMath project this post belongs to (same format as the chain-level project_id).
titlestringTitle of the forum thread this post appears in.
textstringFull text of the post in BBCode markup (see "Text format" below).
thankedintegerNumber of "thank" reactions this post received from other users.
comment_countintegerNumber of comments on the thread at the time the data was collected.
labelsarrayAnnotation labels assigned to this post (see below). A post may have multiple labels.

Label fields

Each element of the labels array is an object with these fields:

FieldTypeDescription
label_typestringThe annotation category (see "Label types" below).
result_refstringThe progress chain this label contributes to, formatted as <topic_id>-<post_number>. For Published-in-paper labels, this is an arXiv ID and theorem reference instead (e.g. "1704.05211,Theorem-2.3").
prev_refstring or nullA back-reference to a specific earlier post, formatted as <topic_id>-<post_number>. Used by Answer (points to the Question it responds to) and FindError (points to the post whose error is identified). Null for all other label types.

Label types

Labels describe the role a post plays in the mathematical discourse. They fall into two categories.

Discourse labels

These labels indicate how a post contributes to the progress chain:

LabelMeaning
StartIntroduces an initial claim, conjecture, or approach for a result.
ProgressExtends or partially advances an existing line of reasoning.
NewProgressIntroduces a new direction or substantially different approach to an existing result.
ProofProvides a complete proof of the result.
NewProofProvides a complete proof using a substantially different method than prior proofs.
QuestionAsks a mathematical question relevant to the result.
AnswerResponds to a specific Question (see prev_ref).
FindErrorIdentifies a mathematical error in a specific earlier post (see prev_ref).
ErroneousMarks a post whose mathematical content was found to contain an error.
ResultMarks the post that states the final, verified result of the chain. Often co-occurs with Proof.

Metadata labels

LabelMeaning
Published-in-paperIndicates this result was published in a peer-reviewed paper. The result_ref field contains the arXiv ID and theorem number.

Project IDs

Each project_id corresponds to one CrowdMath research project:

Project IDYearChains
mitprimes2016201626
mitprimes2017a201759
mitprimes201820186
mitprimes2019201911
mitprimes2020202021
mitprimes2021202119
mitprimes202220228
mitprimes202320232
mitprimes202420243
mitprimes2024-220242
mitprimes202520255
mitprimes2025-220252

LLM evaluation

The repository also contains the benchmark code from the paper:

Paper taskDirectoryInstancesModel outputs
Task 1 — proof-step classification (4-class)llm_eval/task1_proof_step_classification/275llm_eval/results/task1_solvers.jsonl
Task 2 — next-post prediction (4-way MC)llm_eval/task2_next_post_prediction/81llm_eval/results/task2_solvers.jsonl
  • llm_eval/ — prompts, task instances, the OpenRouter-based generation harness, scoring, and the raw outputs of the six evaluated models. See [llm_eval/README.md](llm_eval/README.md).
pip install -r requirements.txt
python -m llm_eval.score # reproduce Task 1 scores 
python -m llm_eval.score --task task2 # Task 2

License

This dataset is released under the MIT License.

Citation

If you use this dataset, please cite:

@misc{muckatira2026crowdmathdatasetcrowdsourcedmathematical,
title={CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions}, author={Sherin Muckatira and Jesse Geneson and Slava Gerovitch and Pavel Etingof and Mikhail Gronas and Anna Rumshisky},
year={2026},
eprint={2606.06526},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2606.06526}, }

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

CrowdMath

This repository hosts resources for the paper:

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions

Overview

This dataset captures the collaborative mathematical reasoning that took place on the MIT PRIMES CrowdMath online research program (2016-2025). In CrowdMath, high-school students work together on open research problems in mathematics, posting conjectures, proofs, error corrections, and questions on a shared message board (Art of Problem Solving / AoPS).

Expert annotators read every thread and labeled each post with its role in the mathematical discourse. Those annotations were then grouped into progress chains -- sequences of posts that together advance a single mathematical result from an initial claim to a verified proof.

The dataset file dataset/dataset.json is a JSON array of progress chains.

Example progress chain from the CrowdMath dataset
An example progress chain showing how student posts are annotated with discourse labels.

File format

dataset/dataset.json is a JSON array. Each element is a progress chain object with the fields described below.

Progress chain fields

FieldTypeDescription
result_idstringUnique identifier for the chain, formatted as <topic_id>-<post_number> with an optional letter suffix (e.g. "1228277-14", "1320553-27a").
project_idstringCrowdMath project identifier (e.g. "mitprimes2016", "mitprimes2021"). Each project corresponds to one year of the program; some years have multiple projects (e.g. "mitprimes2024" and "mitprimes2024-2").
open_problem_refstring or nullThe open problem reference as written by the annotator (e.g. "2020-6", "2017-5"). Null when the chain could not be matched to a specific open problem.
open_problem_yearinteger or nullFour-digit year of the matched open problem. Null when unmatched.
open_problem_numberinteger or nullProblem number within that year's problem set. Null when unmatched.
problem_textstring or nullThe mathematical problem statement that this chain addresses. When an open problem was matched, this includes official problem text from the CrowdMath problem set. Otherwise it contains the text of the first "Problem" post in the thread.
problem_resourcesstring or nullSupplementary resources (references, hints, relevant background) associated with the open problem. Non-null for only a small number of chains.
postsarrayOrdered list of post objects that form this chain (see below). The Problem post is excluded from this array because its content is captured in problem_text. Posts are sorted by post_order.

Post fields

Each element of the posts array is an object with these fields:

FieldTypeDescription
topic_idstringAoPS forum thread ID.
post_numberinteger1-based position of this post within its thread.
post_orderintegerGlobal ordering index across the entire dataset. Use this field to sort posts chronologically.
post_idstringAoPS database identifier for this post.
project_idstringCrowdMath project this post belongs to (same format as the chain-level project_id).
titlestringTitle of the forum thread this post appears in.
textstringFull text of the post in BBCode markup (see "Text format" below).
thankedintegerNumber of "thank" reactions this post received from other users.
comment_countintegerNumber of comments on the thread at the time the data was collected.
labelsarrayAnnotation labels assigned to this post (see below). A post may have multiple labels.

Label fields

Each element of the labels array is an object with these fields:

FieldTypeDescription
label_typestringThe annotation category (see "Label types" below).
result_refstringThe progress chain this label contributes to, formatted as <topic_id>-<post_number>. For Published-in-paper labels, this is an arXiv ID and theorem reference instead (e.g. "1704.05211,Theorem-2.3").
prev_refstring or nullA back-reference to a specific earlier post, formatted as <topic_id>-<post_number>. Used by Answer (points to the Question it responds to) and FindError (points to the post whose error is identified). Null for all other label types.

Label types

Labels describe the role a post plays in the mathematical discourse. They fall into two categories.

Discourse labels

These labels indicate how a post contributes to the progress chain:

LabelMeaning
StartIntroduces an initial claim, conjecture, or approach for a result.
ProgressExtends or partially advances an existing line of reasoning.
NewProgressIntroduces a new direction or substantially different approach to an existing result.
ProofProvides a complete proof of the result.
NewProofProvides a complete proof using a substantially different method than prior proofs.
QuestionAsks a mathematical question relevant to the result.
AnswerResponds to a specific Question (see prev_ref).
FindErrorIdentifies a mathematical error in a specific earlier post (see prev_ref).
ErroneousMarks a post whose mathematical content was found to contain an error.
ResultMarks the post that states the final, verified result of the chain. Often co-occurs with Proof.

Metadata labels

LabelMeaning
Published-in-paperIndicates this result was published in a peer-reviewed paper. The result_ref field contains the arXiv ID and theorem number.

Project IDs

Each project_id corresponds to one CrowdMath research project:

Project IDYearChains
mitprimes2016201626
mitprimes2017a201759
mitprimes201820186
mitprimes2019201911
mitprimes2020202021
mitprimes2021202119
mitprimes202220228
mitprimes202320232
mitprimes202420243
mitprimes2024-220242
mitprimes202520255
mitprimes2025-220252

LLM evaluation

The repository also contains the benchmark code from the paper:

Paper taskDirectoryInstancesModel outputs
Task 1 — proof-step classification (4-class)llm_eval/task1_proof_step_classification/275llm_eval/results/task1_solvers.jsonl
Task 2 — next-post prediction (4-way MC)llm_eval/task2_next_post_prediction/81llm_eval/results/task2_solvers.jsonl
  • llm_eval/ — prompts, task instances, the OpenRouter-based generation harness, scoring, and the raw outputs of the six evaluated models. See [llm_eval/README.md](llm_eval/README.md).
pip install -r requirements.txt
python -m llm_eval.score # reproduce Task 1 scores 
python -m llm_eval.score --task task2 # Task 2

License

This dataset is released under the MIT License.

Citation

If you use this dataset, please cite:

@misc{muckatira2026crowdmathdatasetcrowdsourcedmathematical,
title={CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions}, author={Sherin Muckatira and Jesse Geneson and Slava Gerovitch and Pavel Etingof and Mikhail Gronas and Anna Rumshisky},
year={2026},
eprint={2606.06526},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2606.06526}, }

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

CrowdMath

This repository hosts resources for the paper:

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions

Overview

This dataset captures the collaborative mathematical reasoning that took place on the MIT PRIMES CrowdMath online research program (2016-2025). In CrowdMath, high-school students work together on open research problems in mathematics, posting conjectures, proofs, error corrections, and questions on a shared message board (Art of Problem Solving / AoPS).

Expert annotators read every thread and labeled each post with its role in the mathematical discourse. Those annotations were then grouped into progress chains -- sequences of posts that together advance a single mathematical result from an initial claim to a verified proof.

The dataset file dataset/dataset.json is a JSON array of progress chains.

Example progress chain from the CrowdMath dataset
An example progress chain showing how student posts are annotated with discourse labels.

File format

dataset/dataset.json is a JSON array. Each element is a progress chain object with the fields described below.

Progress chain fields

FieldTypeDescription
result_idstringUnique identifier for the chain, formatted as <topic_id>-<post_number> with an optional letter suffix (e.g. "1228277-14", "1320553-27a").
project_idstringCrowdMath project identifier (e.g. "mitprimes2016", "mitprimes2021"). Each project corresponds to one year of the program; some years have multiple projects (e.g. "mitprimes2024" and "mitprimes2024-2").
open_problem_refstring or nullThe open problem reference as written by the annotator (e.g. "2020-6", "2017-5"). Null when the chain could not be matched to a specific open problem.
open_problem_yearinteger or nullFour-digit year of the matched open problem. Null when unmatched.
open_problem_numberinteger or nullProblem number within that year's problem set. Null when unmatched.
problem_textstring or nullThe mathematical problem statement that this chain addresses. When an open problem was matched, this includes official problem text from the CrowdMath problem set. Otherwise it contains the text of the first "Problem" post in the thread.
problem_resourcesstring or nullSupplementary resources (references, hints, relevant background) associated with the open problem. Non-null for only a small number of chains.
postsarrayOrdered list of post objects that form this chain (see below). The Problem post is excluded from this array because its content is captured in problem_text. Posts are sorted by post_order.

Post fields

Each element of the posts array is an object with these fields:

FieldTypeDescription
topic_idstringAoPS forum thread ID.
post_numberinteger1-based position of this post within its thread.
post_orderintegerGlobal ordering index across the entire dataset. Use this field to sort posts chronologically.
post_idstringAoPS database identifier for this post.
project_idstringCrowdMath project this post belongs to (same format as the chain-level project_id).
titlestringTitle of the forum thread this post appears in.
textstringFull text of the post in BBCode markup (see "Text format" below).
thankedintegerNumber of "thank" reactions this post received from other users.
comment_countintegerNumber of comments on the thread at the time the data was collected.
labelsarrayAnnotation labels assigned to this post (see below). A post may have multiple labels.

Label fields

Each element of the labels array is an object with these fields:

FieldTypeDescription
label_typestringThe annotation category (see "Label types" below).
result_refstringThe progress chain this label contributes to, formatted as <topic_id>-<post_number>. For Published-in-paper labels, this is an arXiv ID and theorem reference instead (e.g. "1704.05211,Theorem-2.3").
prev_refstring or nullA back-reference to a specific earlier post, formatted as <topic_id>-<post_number>. Used by Answer (points to the Question it responds to) and FindError (points to the post whose error is identified). Null for all other label types.

Label types

Labels describe the role a post plays in the mathematical discourse. They fall into two categories.

Discourse labels

These labels indicate how a post contributes to the progress chain:

LabelMeaning
StartIntroduces an initial claim, conjecture, or approach for a result.
ProgressExtends or partially advances an existing line of reasoning.
NewProgressIntroduces a new direction or substantially different approach to an existing result.
ProofProvides a complete proof of the result.
NewProofProvides a complete proof using a substantially different method than prior proofs.
QuestionAsks a mathematical question relevant to the result.
AnswerResponds to a specific Question (see prev_ref).
FindErrorIdentifies a mathematical error in a specific earlier post (see prev_ref).
ErroneousMarks a post whose mathematical content was found to contain an error.
ResultMarks the post that states the final, verified result of the chain. Often co-occurs with Proof.

Metadata labels

LabelMeaning
Published-in-paperIndicates this result was published in a peer-reviewed paper. The result_ref field contains the arXiv ID and theorem number.

Project IDs

Each project_id corresponds to one CrowdMath research project:

Project IDYearChains
mitprimes2016201626
mitprimes2017a201759
mitprimes201820186
mitprimes2019201911
mitprimes2020202021
mitprimes2021202119
mitprimes202220228
mitprimes202320232
mitprimes202420243
mitprimes2024-220242
mitprimes202520255
mitprimes2025-220252

LLM evaluation

The repository also contains the benchmark code from the paper:

Paper taskDirectoryInstancesModel outputs
Task 1 — proof-step classification (4-class)llm_eval/task1_proof_step_classification/275llm_eval/results/task1_solvers.jsonl
Task 2 — next-post prediction (4-way MC)llm_eval/task2_next_post_prediction/81llm_eval/results/task2_solvers.jsonl
  • llm_eval/ — prompts, task instances, the OpenRouter-based generation harness, scoring, and the raw outputs of the six evaluated models. See [llm_eval/README.md](llm_eval/README.md).
pip install -r requirements.txt
python -m llm_eval.score # reproduce Task 1 scores 
python -m llm_eval.score --task task2 # Task 2

License

This dataset is released under the MIT License.

Citation

If you use this dataset, please cite:

@misc{muckatira2026crowdmathdatasetcrowdsourcedmathematical,
title={CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions}, author={Sherin Muckatira and Jesse Geneson and Slava Gerovitch and Pavel Etingof and Mikhail Gronas and Anna Rumshisky},
year={2026},
eprint={2606.06526},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2606.06526}, }

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

CrowdMath

This repository hosts resources for the paper:

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions

Overview

This dataset captures the collaborative mathematical reasoning that took place on the MIT PRIMES CrowdMath online research program (2016-2025). In CrowdMath, high-school students work together on open research problems in mathematics, posting conjectures, proofs, error corrections, and questions on a shared message board (Art of Problem Solving / AoPS).

Expert annotators read every thread and labeled each post with its role in the mathematical discourse. Those annotations were then grouped into progress chains -- sequences of posts that together advance a single mathematical result from an initial claim to a verified proof.

The dataset file dataset/dataset.json is a JSON array of progress chains.

Example progress chain from the CrowdMath dataset
An example progress chain showing how student posts are annotated with discourse labels.

File format

dataset/dataset.json is a JSON array. Each element is a progress chain object with the fields described below.

Progress chain fields

FieldTypeDescription
result_idstringUnique identifier for the chain, formatted as <topic_id>-<post_number> with an optional letter suffix (e.g. "1228277-14", "1320553-27a").
project_idstringCrowdMath project identifier (e.g. "mitprimes2016", "mitprimes2021"). Each project corresponds to one year of the program; some years have multiple projects (e.g. "mitprimes2024" and "mitprimes2024-2").
open_problem_refstring or nullThe open problem reference as written by the annotator (e.g. "2020-6", "2017-5"). Null when the chain could not be matched to a specific open problem.
open_problem_yearinteger or nullFour-digit year of the matched open problem. Null when unmatched.
open_problem_numberinteger or nullProblem number within that year's problem set. Null when unmatched.
problem_textstring or nullThe mathematical problem statement that this chain addresses. When an open problem was matched, this includes official problem text from the CrowdMath problem set. Otherwise it contains the text of the first "Problem" post in the thread.
problem_resourcesstring or nullSupplementary resources (references, hints, relevant background) associated with the open problem. Non-null for only a small number of chains.
postsarrayOrdered list of post objects that form this chain (see below). The Problem post is excluded from this array because its content is captured in problem_text. Posts are sorted by post_order.

Post fields

Each element of the posts array is an object with these fields:

FieldTypeDescription
topic_idstringAoPS forum thread ID.
post_numberinteger1-based position of this post within its thread.
post_orderintegerGlobal ordering index across the entire dataset. Use this field to sort posts chronologically.
post_idstringAoPS database identifier for this post.
project_idstringCrowdMath project this post belongs to (same format as the chain-level project_id).
titlestringTitle of the forum thread this post appears in.
textstringFull text of the post in BBCode markup (see "Text format" below).
thankedintegerNumber of "thank" reactions this post received from other users.
comment_countintegerNumber of comments on the thread at the time the data was collected.
labelsarrayAnnotation labels assigned to this post (see below). A post may have multiple labels.

Label fields

Each element of the labels array is an object with these fields:

FieldTypeDescription
label_typestringThe annotation category (see "Label types" below).
result_refstringThe progress chain this label contributes to, formatted as <topic_id>-<post_number>. For Published-in-paper labels, this is an arXiv ID and theorem reference instead (e.g. "1704.05211,Theorem-2.3").
prev_refstring or nullA back-reference to a specific earlier post, formatted as <topic_id>-<post_number>. Used by Answer (points to the Question it responds to) and FindError (points to the post whose error is identified). Null for all other label types.

Label types

Labels describe the role a post plays in the mathematical discourse. They fall into two categories.

Discourse labels

These labels indicate how a post contributes to the progress chain:

LabelMeaning
StartIntroduces an initial claim, conjecture, or approach for a result.
ProgressExtends or partially advances an existing line of reasoning.
NewProgressIntroduces a new direction or substantially different approach to an existing result.
ProofProvides a complete proof of the result.
NewProofProvides a complete proof using a substantially different method than prior proofs.
QuestionAsks a mathematical question relevant to the result.
AnswerResponds to a specific Question (see prev_ref).
FindErrorIdentifies a mathematical error in a specific earlier post (see prev_ref).
ErroneousMarks a post whose mathematical content was found to contain an error.
ResultMarks the post that states the final, verified result of the chain. Often co-occurs with Proof.

Metadata labels

LabelMeaning
Published-in-paperIndicates this result was published in a peer-reviewed paper. The result_ref field contains the arXiv ID and theorem number.

Project IDs

Each project_id corresponds to one CrowdMath research project:

Project IDYearChains
mitprimes2016201626
mitprimes2017a201759
mitprimes201820186
mitprimes2019201911
mitprimes2020202021
mitprimes2021202119
mitprimes202220228
mitprimes202320232
mitprimes202420243
mitprimes2024-220242
mitprimes202520255
mitprimes2025-220252

LLM evaluation

The repository also contains the benchmark code from the paper:

Paper taskDirectoryInstancesModel outputs
Task 1 — proof-step classification (4-class)llm_eval/task1_proof_step_classification/275llm_eval/results/task1_solvers.jsonl
Task 2 — next-post prediction (4-way MC)llm_eval/task2_next_post_prediction/81llm_eval/results/task2_solvers.jsonl
  • llm_eval/ — prompts, task instances, the OpenRouter-based generation harness, scoring, and the raw outputs of the six evaluated models. See [llm_eval/README.md](llm_eval/README.md).
pip install -r requirements.txt
python -m llm_eval.score # reproduce Task 1 scores 
python -m llm_eval.score --task task2 # Task 2

License

This dataset is released under the MIT License.

Citation

If you use this dataset, please cite:

@misc{muckatira2026crowdmathdatasetcrowdsourcedmathematical,
title={CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions}, author={Sherin Muckatira and Jesse Geneson and Slava Gerovitch and Pavel Etingof and Mikhail Gronas and Anna Rumshisky},
year={2026},
eprint={2606.06526},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2606.06526}, }

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

CrowdMath

This repository hosts resources for the paper:

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions

Overview

This dataset captures the collaborative mathematical reasoning that took place on the MIT PRIMES CrowdMath online research program (2016-2025). In CrowdMath, high-school students work together on open research problems in mathematics, posting conjectures, proofs, error corrections, and questions on a shared message board (Art of Problem Solving / AoPS).

Expert annotators read every thread and labeled each post with its role in the mathematical discourse. Those annotations were then grouped into progress chains -- sequences of posts that together advance a single mathematical result from an initial claim to a verified proof.

The dataset file dataset/dataset.json is a JSON array of progress chains.

Example progress chain from the CrowdMath dataset
An example progress chain showing how student posts are annotated with discourse labels.

File format

dataset/dataset.json is a JSON array. Each element is a progress chain object with the fields described below.

Progress chain fields

FieldTypeDescription
result_idstringUnique identifier for the chain, formatted as <topic_id>-<post_number> with an optional letter suffix (e.g. "1228277-14", "1320553-27a").
project_idstringCrowdMath project identifier (e.g. "mitprimes2016", "mitprimes2021"). Each project corresponds to one year of the program; some years have multiple projects (e.g. "mitprimes2024" and "mitprimes2024-2").
open_problem_refstring or nullThe open problem reference as written by the annotator (e.g. "2020-6", "2017-5"). Null when the chain could not be matched to a specific open problem.
open_problem_yearinteger or nullFour-digit year of the matched open problem. Null when unmatched.
open_problem_numberinteger or nullProblem number within that year's problem set. Null when unmatched.
problem_textstring or nullThe mathematical problem statement that this chain addresses. When an open problem was matched, this includes official problem text from the CrowdMath problem set. Otherwise it contains the text of the first "Problem" post in the thread.
problem_resourcesstring or nullSupplementary resources (references, hints, relevant background) associated with the open problem. Non-null for only a small number of chains.
postsarrayOrdered list of post objects that form this chain (see below). The Problem post is excluded from this array because its content is captured in problem_text. Posts are sorted by post_order.

Post fields

Each element of the posts array is an object with these fields:

FieldTypeDescription
topic_idstringAoPS forum thread ID.
post_numberinteger1-based position of this post within its thread.
post_orderintegerGlobal ordering index across the entire dataset. Use this field to sort posts chronologically.
post_idstringAoPS database identifier for this post.
project_idstringCrowdMath project this post belongs to (same format as the chain-level project_id).
titlestringTitle of the forum thread this post appears in.
textstringFull text of the post in BBCode markup (see "Text format" below).
thankedintegerNumber of "thank" reactions this post received from other users.
comment_countintegerNumber of comments on the thread at the time the data was collected.
labelsarrayAnnotation labels assigned to this post (see below). A post may have multiple labels.

Label fields

Each element of the labels array is an object with these fields:

FieldTypeDescription
label_typestringThe annotation category (see "Label types" below).
result_refstringThe progress chain this label contributes to, formatted as <topic_id>-<post_number>. For Published-in-paper labels, this is an arXiv ID and theorem reference instead (e.g. "1704.05211,Theorem-2.3").
prev_refstring or nullA back-reference to a specific earlier post, formatted as <topic_id>-<post_number>. Used by Answer (points to the Question it responds to) and FindError (points to the post whose error is identified). Null for all other label types.

Label types

Labels describe the role a post plays in the mathematical discourse. They fall into two categories.

Discourse labels

These labels indicate how a post contributes to the progress chain:

LabelMeaning
StartIntroduces an initial claim, conjecture, or approach for a result.
ProgressExtends or partially advances an existing line of reasoning.
NewProgressIntroduces a new direction or substantially different approach to an existing result.
ProofProvides a complete proof of the result.
NewProofProvides a complete proof using a substantially different method than prior proofs.
QuestionAsks a mathematical question relevant to the result.
AnswerResponds to a specific Question (see prev_ref).
FindErrorIdentifies a mathematical error in a specific earlier post (see prev_ref).
ErroneousMarks a post whose mathematical content was found to contain an error.
ResultMarks the post that states the final, verified result of the chain. Often co-occurs with Proof.

Metadata labels

LabelMeaning
Published-in-paperIndicates this result was published in a peer-reviewed paper. The result_ref field contains the arXiv ID and theorem number.

Project IDs

Each project_id corresponds to one CrowdMath research project:

Project IDYearChains
mitprimes2016201626
mitprimes2017a201759
mitprimes201820186
mitprimes2019201911
mitprimes2020202021
mitprimes2021202119
mitprimes202220228
mitprimes202320232
mitprimes202420243
mitprimes2024-220242
mitprimes202520255
mitprimes2025-220252

LLM evaluation

The repository also contains the benchmark code from the paper:

Paper taskDirectoryInstancesModel outputs
Task 1 — proof-step classification (4-class)llm_eval/task1_proof_step_classification/275llm_eval/results/task1_solvers.jsonl
Task 2 — next-post prediction (4-way MC)llm_eval/task2_next_post_prediction/81llm_eval/results/task2_solvers.jsonl
  • llm_eval/ — prompts, task instances, the OpenRouter-based generation harness, scoring, and the raw outputs of the six evaluated models. See [llm_eval/README.md](llm_eval/README.md).
pip install -r requirements.txt
python -m llm_eval.score # reproduce Task 1 scores 
python -m llm_eval.score --task task2 # Task 2

License

This dataset is released under the MIT License.

Citation

If you use this dataset, please cite:

@misc{muckatira2026crowdmathdatasetcrowdsourcedmathematical,
title={CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions}, author={Sherin Muckatira and Jesse Geneson and Slava Gerovitch and Pavel Etingof and Mikhail Gronas and Anna Rumshisky},
year={2026},
eprint={2606.06526},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2606.06526}, }

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

CrowdMath

This repository hosts resources for the paper:

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions

Overview

This dataset captures the collaborative mathematical reasoning that took place on the MIT PRIMES CrowdMath online research program (2016-2025). In CrowdMath, high-school students work together on open research problems in mathematics, posting conjectures, proofs, error corrections, and questions on a shared message board (Art of Problem Solving / AoPS).

Expert annotators read every thread and labeled each post with its role in the mathematical discourse. Those annotations were then grouped into progress chains -- sequences of posts that together advance a single mathematical result from an initial claim to a verified proof.

The dataset file dataset/dataset.json is a JSON array of progress chains.

Example progress chain from the CrowdMath dataset
An example progress chain showing how student posts are annotated with discourse labels.

File format

dataset/dataset.json is a JSON array. Each element is a progress chain object with the fields described below.

Progress chain fields

FieldTypeDescription
result_idstringUnique identifier for the chain, formatted as <topic_id>-<post_number> with an optional letter suffix (e.g. "1228277-14", "1320553-27a").
project_idstringCrowdMath project identifier (e.g. "mitprimes2016", "mitprimes2021"). Each project corresponds to one year of the program; some years have multiple projects (e.g. "mitprimes2024" and "mitprimes2024-2").
open_problem_refstring or nullThe open problem reference as written by the annotator (e.g. "2020-6", "2017-5"). Null when the chain could not be matched to a specific open problem.
open_problem_yearinteger or nullFour-digit year of the matched open problem. Null when unmatched.
open_problem_numberinteger or nullProblem number within that year's problem set. Null when unmatched.
problem_textstring or nullThe mathematical problem statement that this chain addresses. When an open problem was matched, this includes official problem text from the CrowdMath problem set. Otherwise it contains the text of the first "Problem" post in the thread.
problem_resourcesstring or nullSupplementary resources (references, hints, relevant background) associated with the open problem. Non-null for only a small number of chains.
postsarrayOrdered list of post objects that form this chain (see below). The Problem post is excluded from this array because its content is captured in problem_text. Posts are sorted by post_order.

Post fields

Each element of the posts array is an object with these fields:

FieldTypeDescription
topic_idstringAoPS forum thread ID.
post_numberinteger1-based position of this post within its thread.
post_orderintegerGlobal ordering index across the entire dataset. Use this field to sort posts chronologically.
post_idstringAoPS database identifier for this post.
project_idstringCrowdMath project this post belongs to (same format as the chain-level project_id).
titlestringTitle of the forum thread this post appears in.
textstringFull text of the post in BBCode markup (see "Text format" below).
thankedintegerNumber of "thank" reactions this post received from other users.
comment_countintegerNumber of comments on the thread at the time the data was collected.
labelsarrayAnnotation labels assigned to this post (see below). A post may have multiple labels.

Label fields

Each element of the labels array is an object with these fields:

FieldTypeDescription
label_typestringThe annotation category (see "Label types" below).
result_refstringThe progress chain this label contributes to, formatted as <topic_id>-<post_number>. For Published-in-paper labels, this is an arXiv ID and theorem reference instead (e.g. "1704.05211,Theorem-2.3").
prev_refstring or nullA back-reference to a specific earlier post, formatted as <topic_id>-<post_number>. Used by Answer (points to the Question it responds to) and FindError (points to the post whose error is identified). Null for all other label types.

Label types

Labels describe the role a post plays in the mathematical discourse. They fall into two categories.

Discourse labels

These labels indicate how a post contributes to the progress chain:

LabelMeaning
StartIntroduces an initial claim, conjecture, or approach for a result.
ProgressExtends or partially advances an existing line of reasoning.
NewProgressIntroduces a new direction or substantially different approach to an existing result.
ProofProvides a complete proof of the result.
NewProofProvides a complete proof using a substantially different method than prior proofs.
QuestionAsks a mathematical question relevant to the result.
AnswerResponds to a specific Question (see prev_ref).
FindErrorIdentifies a mathematical error in a specific earlier post (see prev_ref).
ErroneousMarks a post whose mathematical content was found to contain an error.
ResultMarks the post that states the final, verified result of the chain. Often co-occurs with Proof.

Metadata labels

LabelMeaning
Published-in-paperIndicates this result was published in a peer-reviewed paper. The result_ref field contains the arXiv ID and theorem number.

Project IDs

Each project_id corresponds to one CrowdMath research project:

Project IDYearChains
mitprimes2016201626
mitprimes2017a201759
mitprimes201820186
mitprimes2019201911
mitprimes2020202021
mitprimes2021202119
mitprimes202220228
mitprimes202320232
mitprimes202420243
mitprimes2024-220242
mitprimes202520255
mitprimes2025-220252

LLM evaluation

The repository also contains the benchmark code from the paper:

Paper taskDirectoryInstancesModel outputs
Task 1 — proof-step classification (4-class)llm_eval/task1_proof_step_classification/275llm_eval/results/task1_solvers.jsonl
Task 2 — next-post prediction (4-way MC)llm_eval/task2_next_post_prediction/81llm_eval/results/task2_solvers.jsonl
  • llm_eval/ — prompts, task instances, the OpenRouter-based generation harness, scoring, and the raw outputs of the six evaluated models. See [llm_eval/README.md](llm_eval/README.md).
pip install -r requirements.txt
python -m llm_eval.score # reproduce Task 1 scores 
python -m llm_eval.score --task task2 # Task 2

License

This dataset is released under the MIT License.

Citation

If you use this dataset, please cite:

@misc{muckatira2026crowdmathdatasetcrowdsourcedmathematical,
title={CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions}, author={Sherin Muckatira and Jesse Geneson and Slava Gerovitch and Pavel Etingof and Mikhail Gronas and Anna Rumshisky},
year={2026},
eprint={2606.06526},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2606.06526}, }

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

CrowdMath

This repository hosts resources for the paper:

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions

Overview

This dataset captures the collaborative mathematical reasoning that took place on the MIT PRIMES CrowdMath online research program (2016-2025). In CrowdMath, high-school students work together on open research problems in mathematics, posting conjectures, proofs, error corrections, and questions on a shared message board (Art of Problem Solving / AoPS).

Expert annotators read every thread and labeled each post with its role in the mathematical discourse. Those annotations were then grouped into progress chains -- sequences of posts that together advance a single mathematical result from an initial claim to a verified proof.

The dataset file dataset/dataset.json is a JSON array of progress chains.

Example progress chain from the CrowdMath dataset
An example progress chain showing how student posts are annotated with discourse labels.

File format

dataset/dataset.json is a JSON array. Each element is a progress chain object with the fields described below.

Progress chain fields

FieldTypeDescription
result_idstringUnique identifier for the chain, formatted as <topic_id>-<post_number> with an optional letter suffix (e.g. "1228277-14", "1320553-27a").
project_idstringCrowdMath project identifier (e.g. "mitprimes2016", "mitprimes2021"). Each project corresponds to one year of the program; some years have multiple projects (e.g. "mitprimes2024" and "mitprimes2024-2").
open_problem_refstring or nullThe open problem reference as written by the annotator (e.g. "2020-6", "2017-5"). Null when the chain could not be matched to a specific open problem.
open_problem_yearinteger or nullFour-digit year of the matched open problem. Null when unmatched.
open_problem_numberinteger or nullProblem number within that year's problem set. Null when unmatched.
problem_textstring or nullThe mathematical problem statement that this chain addresses. When an open problem was matched, this includes official problem text from the CrowdMath problem set. Otherwise it contains the text of the first "Problem" post in the thread.
problem_resourcesstring or nullSupplementary resources (references, hints, relevant background) associated with the open problem. Non-null for only a small number of chains.
postsarrayOrdered list of post objects that form this chain (see below). The Problem post is excluded from this array because its content is captured in problem_text. Posts are sorted by post_order.

Post fields

Each element of the posts array is an object with these fields:

FieldTypeDescription
topic_idstringAoPS forum thread ID.
post_numberinteger1-based position of this post within its thread.
post_orderintegerGlobal ordering index across the entire dataset. Use this field to sort posts chronologically.
post_idstringAoPS database identifier for this post.
project_idstringCrowdMath project this post belongs to (same format as the chain-level project_id).
titlestringTitle of the forum thread this post appears in.
textstringFull text of the post in BBCode markup (see "Text format" below).
thankedintegerNumber of "thank" reactions this post received from other users.
comment_countintegerNumber of comments on the thread at the time the data was collected.
labelsarrayAnnotation labels assigned to this post (see below). A post may have multiple labels.

Label fields

Each element of the labels array is an object with these fields:

FieldTypeDescription
label_typestringThe annotation category (see "Label types" below).
result_refstringThe progress chain this label contributes to, formatted as <topic_id>-<post_number>. For Published-in-paper labels, this is an arXiv ID and theorem reference instead (e.g. "1704.05211,Theorem-2.3").
prev_refstring or nullA back-reference to a specific earlier post, formatted as <topic_id>-<post_number>. Used by Answer (points to the Question it responds to) and FindError (points to the post whose error is identified). Null for all other label types.

Label types

Labels describe the role a post plays in the mathematical discourse. They fall into two categories.

Discourse labels

These labels indicate how a post contributes to the progress chain:

LabelMeaning
StartIntroduces an initial claim, conjecture, or approach for a result.
ProgressExtends or partially advances an existing line of reasoning.
NewProgressIntroduces a new direction or substantially different approach to an existing result.
ProofProvides a complete proof of the result.
NewProofProvides a complete proof using a substantially different method than prior proofs.
QuestionAsks a mathematical question relevant to the result.
AnswerResponds to a specific Question (see prev_ref).
FindErrorIdentifies a mathematical error in a specific earlier post (see prev_ref).
ErroneousMarks a post whose mathematical content was found to contain an error.
ResultMarks the post that states the final, verified result of the chain. Often co-occurs with Proof.

Metadata labels

LabelMeaning
Published-in-paperIndicates this result was published in a peer-reviewed paper. The result_ref field contains the arXiv ID and theorem number.

Project IDs

Each project_id corresponds to one CrowdMath research project:

Project IDYearChains
mitprimes2016201626
mitprimes2017a201759
mitprimes201820186
mitprimes2019201911
mitprimes2020202021
mitprimes2021202119
mitprimes202220228
mitprimes202320232
mitprimes202420243
mitprimes2024-220242
mitprimes202520255
mitprimes2025-220252

LLM evaluation

The repository also contains the benchmark code from the paper:

Paper taskDirectoryInstancesModel outputs
Task 1 — proof-step classification (4-class)llm_eval/task1_proof_step_classification/275llm_eval/results/task1_solvers.jsonl
Task 2 — next-post prediction (4-way MC)llm_eval/task2_next_post_prediction/81llm_eval/results/task2_solvers.jsonl
  • llm_eval/ — prompts, task instances, the OpenRouter-based generation harness, scoring, and the raw outputs of the six evaluated models. See [llm_eval/README.md](llm_eval/README.md).
pip install -r requirements.txt
python -m llm_eval.score # reproduce Task 1 scores 
python -m llm_eval.score --task task2 # Task 2

License

This dataset is released under the MIT License.

Citation

If you use this dataset, please cite:

@misc{muckatira2026crowdmathdatasetcrowdsourcedmathematical,
title={CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions}, author={Sherin Muckatira and Jesse Geneson and Slava Gerovitch and Pavel Etingof and Mikhail Gronas and Anna Rumshisky},
year={2026},
eprint={2606.06526},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2606.06526}, }

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

CrowdMath

This repository hosts resources for the paper:

CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions

Overview

This dataset captures the collaborative mathematical reasoning that took place on the MIT PRIMES CrowdMath online research program (2016-2025). In CrowdMath, high-school students work together on open research problems in mathematics, posting conjectures, proofs, error corrections, and questions on a shared message board (Art of Problem Solving / AoPS).

Expert annotators read every thread and labeled each post with its role in the mathematical discourse. Those annotations were then grouped into progress chains -- sequences of posts that together advance a single mathematical result from an initial claim to a verified proof.

The dataset file dataset/dataset.json is a JSON array of progress chains.

Example progress chain from the CrowdMath dataset
An example progress chain showing how student posts are annotated with discourse labels.

File format

dataset/dataset.json is a JSON array. Each element is a progress chain object with the fields described below.

Progress chain fields

FieldTypeDescription
result_idstringUnique identifier for the chain, formatted as <topic_id>-<post_number> with an optional letter suffix (e.g. "1228277-14", "1320553-27a").
project_idstringCrowdMath project identifier (e.g. "mitprimes2016", "mitprimes2021"). Each project corresponds to one year of the program; some years have multiple projects (e.g. "mitprimes2024" and "mitprimes2024-2").
open_problem_refstring or nullThe open problem reference as written by the annotator (e.g. "2020-6", "2017-5"). Null when the chain could not be matched to a specific open problem.
open_problem_yearinteger or nullFour-digit year of the matched open problem. Null when unmatched.
open_problem_numberinteger or nullProblem number within that year's problem set. Null when unmatched.
problem_textstring or nullThe mathematical problem statement that this chain addresses. When an open problem was matched, this includes official problem text from the CrowdMath problem set. Otherwise it contains the text of the first "Problem" post in the thread.
problem_resourcesstring or nullSupplementary resources (references, hints, relevant background) associated with the open problem. Non-null for only a small number of chains.
postsarrayOrdered list of post objects that form this chain (see below). The Problem post is excluded from this array because its content is captured in problem_text. Posts are sorted by post_order.

Post fields

Each element of the posts array is an object with these fields:

FieldTypeDescription
topic_idstringAoPS forum thread ID.
post_numberinteger1-based position of this post within its thread.
post_orderintegerGlobal ordering index across the entire dataset. Use this field to sort posts chronologically.
post_idstringAoPS database identifier for this post.
project_idstringCrowdMath project this post belongs to (same format as the chain-level project_id).
titlestringTitle of the forum thread this post appears in.
textstringFull text of the post in BBCode markup (see "Text format" below).
thankedintegerNumber of "thank" reactions this post received from other users.
comment_countintegerNumber of comments on the thread at the time the data was collected.
labelsarrayAnnotation labels assigned to this post (see below). A post may have multiple labels.

Label fields

Each element of the labels array is an object with these fields:

FieldTypeDescription
label_typestringThe annotation category (see "Label types" below).
result_refstringThe progress chain this label contributes to, formatted as <topic_id>-<post_number>. For Published-in-paper labels, this is an arXiv ID and theorem reference instead (e.g. "1704.05211,Theorem-2.3").
prev_refstring or nullA back-reference to a specific earlier post, formatted as <topic_id>-<post_number>. Used by Answer (points to the Question it responds to) and FindError (points to the post whose error is identified). Null for all other label types.

Label types

Labels describe the role a post plays in the mathematical discourse. They fall into two categories.

Discourse labels

These labels indicate how a post contributes to the progress chain:

LabelMeaning
StartIntroduces an initial claim, conjecture, or approach for a result.
ProgressExtends or partially advances an existing line of reasoning.
NewProgressIntroduces a new direction or substantially different approach to an existing result.
ProofProvides a complete proof of the result.
NewProofProvides a complete proof using a substantially different method than prior proofs.
QuestionAsks a mathematical question relevant to the result.
AnswerResponds to a specific Question (see prev_ref).
FindErrorIdentifies a mathematical error in a specific earlier post (see prev_ref).
ErroneousMarks a post whose mathematical content was found to contain an error.
ResultMarks the post that states the final, verified result of the chain. Often co-occurs with Proof.

Metadata labels

LabelMeaning
Published-in-paperIndicates this result was published in a peer-reviewed paper. The result_ref field contains the arXiv ID and theorem number.

Project IDs

Each project_id corresponds to one CrowdMath research project:

Project IDYearChains
mitprimes2016201626
mitprimes2017a201759
mitprimes201820186
mitprimes2019201911
mitprimes2020202021
mitprimes2021202119
mitprimes202220228
mitprimes202320232
mitprimes202420243
mitprimes2024-220242
mitprimes202520255
mitprimes2025-220252

LLM evaluation

The repository also contains the benchmark code from the paper:

Paper taskDirectoryInstancesModel outputs
Task 1 — proof-step classification (4-class)llm_eval/task1_proof_step_classification/275llm_eval/results/task1_solvers.jsonl
Task 2 — next-post prediction (4-way MC)llm_eval/task2_next_post_prediction/81llm_eval/results/task2_solvers.jsonl
  • llm_eval/ — prompts, task instances, the OpenRouter-based generation harness, scoring, and the raw outputs of the six evaluated models. See [llm_eval/README.md](llm_eval/README.md).
pip install -r requirements.txt
python -m llm_eval.score # reproduce Task 1 scores 
python -m llm_eval.score --task task2 # Task 2

License

This dataset is released under the MIT License.

Citation

If you use this dataset, please cite:

@misc{muckatira2026crowdmathdatasetcrowdsourcedmathematical,
title={CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions}, author={Sherin Muckatira and Jesse Geneson and Slava Gerovitch and Pavel Etingof and Mikhail Gronas and Anna Rumshisky},
year={2026},
eprint={2606.06526},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2606.06526}, }

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages