Latest commit

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

📢 News: this work has been accepted at the SIGIR 2026!

SurGE

Welcome to the official GitHub repository for SurGE. SurGE is a benchmark and dataset for end-to-end scientific survey generation in the computer science domain.

SurGE provides a comprehensive resource for evaluating automated survey generation systems through both a large-scale dataset and a fully automated evaluation framework.

Overview

SurGE is designed to push the boundaries of automated survey generation by tackling the complex task of creating coherent, in-depth survey articles from a vast academic literature collection. Unlike traditional IR tasks focused solely on document retrieval, SurGE requires systems to:

  • Retrieve: Identify relevant academic articles from a corpus of over 1 million papers.
  • Organize: Construct a structured and hierarchical survey outline.
  • Synthesize: Generate a coherent narrative with proper citations, reflecting expert-authored surveys.

The benchmark includes 205 carefully curated ground truth surveys, each accompanied by detailed metadata and a corresponding hierarchical structure, along with an extensive literature knowledge base sourced primarily from arXiv.

Data Release

This repository contains all necessary components for working with the SurGE dataset:

Dataset Files & Formats:

  • Ground Truth Surveys: Each survey includes metadata fields such as title, authors, publication year, abstract, hierarchical structure, and citation lists.
  • Literature Knowledge Base: A corpus of 1,086,992 academic papers with key fields (e.g., title, authors, abstract, publication date, and category).
  • Auxiliary Mappings: Topic-to-publication mappings to support systematic survey generation.

The complete dataset can be downloaded from this Google Drive folder.

Then you will get the folder data .

Ground Truth Survey

A ground truth survey contains the full content of a survey and its citation information. However, due to space constraints, we cannot display it in its entirety here.

All ground truth surveys are available in data/surveys.json

A survey consists of the following fields:

FieldDescription
authorsList of contributing researchers.
survey_titleThe title of the survey paper.
yearThe publication year of the survey.
dateThe exact timestamp of publication.
categorySubject classification following the arXiv taxonomy.
abstractThe abstract of the survey paper.
structureHierarchical representation of the survey’s organization.
survey_idA unique identifier for the survey.
all_citesList of document IDs cited in the survey.
Bertopic_CDA diversity measure computed using BERTopic.

Literature Knowledge Base

The corpus containing all literature articles is available in: data/corpus.json

Example : Here, we present how articles are organized in the knowledge base. Overly long abstract has been appropriately shortened.

{
"Title": "Information Geometry of Evolution of Neural Network Parameters While Training",
"Authors": [
"Abhiram Anand Thiruthummal",
"Eun-jin Kim",
"Sergiy Shelyag"
],
"Year": "2024",
"Date": "2024-06-07T23:42:54Z",
"Abstract": "Artificial neural networks (ANNs) are powerful tools capable of approximating any arbitrary mathematical function, but their interpretability remains limited...",
"Category": "cs.LG",
"doc_id": 1086990
}

The following are explanations of each field:

KeyDescription
TitleThe title of the research paper.
AuthorsA list of contributing researchers.
YearThe publication year of the paper.
DateThe exact timestamp of the paper’s release.
AbstractThe abstract of the paper.
CategoryThe subject classification following the arXiv taxonomy.
doc_idA unique identifier assigned for reference and retrieval.

Auxiliary Mappings

The mapping containing all queries and their corresponding articles is available in: data/queries.json

Each query in data/queries.json corresponds to a section or paragraph from the ground truth surveys with high citation extraction quality. The associated articles are the references cited in that part of the survey.

Below is an example, Overly long content has been appropriately shortened.

 {
"original_id": "23870233-7f5b-4ef1-9d38-e6f3adb0fa48",
"query_id": 486,
"date": "2020-07-16T09:23:13Z",
"year": "2020",
"category": "cs.LG",
"content": "}\n{\nMachine learning classifiers can perpetuate and amplify the existing systemic injustices in society . Hence, fairness is becoming another important topic. Traditionally...",
"prefix_titles": [
[
"title",
"Learning from Noisy Labels with Deep Neural Networks: A Survey"
],
[
"section",
"Future Research Directions"
],
[
"subsection",
"{Robust and Fair Training"
]
],
"prefix_titles_query": "What are the future research directions for robust and fair training in the context of learning from noisy labels with deep neural networks?",
"cites": [
7771,
4163,
3899,
8740,
8739
],
"cite_extract_rate": 0.8333333333333334,
"origin_cites_number": 6
}

The following are explanations of each field:

KeyDescription
original_idThe identifier for the section where this query is from.
query_idThe ID associated with the specific query.
contentThe content of the section.
prefix_titlesA hierarchical list of titles of the section/subsection/paragraph
prefix_titles_queryThe question this passage is relevant to. The goal of the question is to retrieve relevant documents.
citesA list of document IDs that are cited within this section.
cite_extract_rateThe ratio of extracted citations to the total number of citations in the original document.
origin_cites_numberThe total number of citations originally present in the section.

Note: Not every section in the ground truth surveys has a corresponding entry in queries.json. Only sections with a high citation extraction rate are included.
These entries can be used to train retrieval models, where prefix_titles_query serves as the query and cites contains the relevant document IDs.
The data in queries.json has not been pre-split into training and development sets—you may divide it manually as needed.

Installation Instructions

Requirements

Before you begin, make sure you have the following packages installed in your environment:

FlagEmbedding
bertopic
safetensors
torch==1.13.1+cu117 rouge-score
sacrebleu
numpy==1.26.4
openai
transformers==4.44.2
gdown
socksio

Setting Up Your Environment

Get started with SurGE on your local machine by following these two steps:

1. Download the SurGE Repository

First, clone the SurGE repository to your computer using Git:

git clone https://github.com/oneal2000/SurGE.git surge
cd surge

2. Create and Configure a Python Environment

Next, create an isolated Python environment and install all necessary dependencies. This keeps your project organized and prevents conflicts with other packages.

conda create -n surge python=3.10 -y
conda activate surge
pip install -r requirements.txt -f https://download.pytorch.org/whl/torch_stable.html

Quick Start

  1. First, follow the environment setup instructions in this repository’s README.

  2. Get Data from the link given in Data Release

    python src/download.py
  3. Prepare your generated survey documents in a folder. Refer to Evaluation part in this repository’s README.

    In each folder named after the survey ID, the corresponding generated output should be stored for evaluation. We have provided the articles generated by three baselines on the test set in the paper. For example, with ID baseline as an illustration, the evaluation can be conducted as follows:

    python src/test_final.py --passage_dir ./baselines/ID/output --save_path ./baselines/ID/output/log.json --device 0 --api_key sk-xxx

​you should fill the sk-xxx with your OpenAI API key.

If some of the evaluations are abnormal, you may need to modify the code that sets client in the init of the SurGEvaluator class in evaluator.py to fit the way your API is called

Evaluation

Each ground truth survey has its own ID. In each folder named after the survey ID, the corresponding survey output should be stored for evaluation. You can refer to the example below to try the evaluation. We have provided the articles generated by three baselines on the test set in the paper. For example, with ID baseline as an illustration, the evaluation can be conducted as follows:

python src/test_final.py --passage_dir ./baselines/ID/output --save_path ./baselines/ID/output/log.json --device 0 --api_key sk-xxx

you should fill the sk-xxx with your OpenAI API key.

If some of the evaluations are abnormal, you may need to modify the code that sets client in the init of the SurGEvaluator class in evaluator.py to fit the way your API is called

Arguments Description

--passage_dir : Directory containing generated survey passages.

--eval_list : List of evaluation metrics (space-separated). Default is ALL.

--survey_path : Path to the surveys JSON file. Default is data/surveys.json.

--corpus_path : Path to the corpus JSON file. Default is data/corpus.json.

--device : Device ID for computation. Default is 0.

--api_key : API key for evaluation services.

--save_path : Path to save evaluation results.

As for eval_list, we have the choices below:

ALL : Evaluate all.

ROUGE-BLEU : Evaluate ROUGE and BLEU

SH-Recall : Evaluate SH-Recall

Structure_Quality : Evaluate Structure_Quality(LLM_as_judge)

Coverage : Evaluate Coverage

Relevance-Paper : Evaluate Relevance-Paper

Relevance-Section : Evaluate Relevance-Section

Relevance-Sentence : Evaluate Relevance-Sentence

Logic : Evaluate Logic

License

This project is licensed under MIT License. Please review the LICENSE file for more details.

About

Code for SurGE, SIGIR 2026 Resource Paper

Resources

Stars

19 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

📢 News: this work has been accepted at the SIGIR 2026!

SurGE

Welcome to the official GitHub repository for SurGE. SurGE is a benchmark and dataset for end-to-end scientific survey generation in the computer science domain.

SurGE provides a comprehensive resource for evaluating automated survey generation systems through both a large-scale dataset and a fully automated evaluation framework.

Overview

SurGE is designed to push the boundaries of automated survey generation by tackling the complex task of creating coherent, in-depth survey articles from a vast academic literature collection. Unlike traditional IR tasks focused solely on document retrieval, SurGE requires systems to:

  • Retrieve: Identify relevant academic articles from a corpus of over 1 million papers.
  • Organize: Construct a structured and hierarchical survey outline.
  • Synthesize: Generate a coherent narrative with proper citations, reflecting expert-authored surveys.

The benchmark includes 205 carefully curated ground truth surveys, each accompanied by detailed metadata and a corresponding hierarchical structure, along with an extensive literature knowledge base sourced primarily from arXiv.

Data Release

This repository contains all necessary components for working with the SurGE dataset:

Dataset Files & Formats:

  • Ground Truth Surveys: Each survey includes metadata fields such as title, authors, publication year, abstract, hierarchical structure, and citation lists.
  • Literature Knowledge Base: A corpus of 1,086,992 academic papers with key fields (e.g., title, authors, abstract, publication date, and category).
  • Auxiliary Mappings: Topic-to-publication mappings to support systematic survey generation.

The complete dataset can be downloaded from this Google Drive folder.

Then you will get the folder data .

Ground Truth Survey

A ground truth survey contains the full content of a survey and its citation information. However, due to space constraints, we cannot display it in its entirety here.

All ground truth surveys are available in data/surveys.json

A survey consists of the following fields:

FieldDescription
authorsList of contributing researchers.
survey_titleThe title of the survey paper.
yearThe publication year of the survey.
dateThe exact timestamp of publication.
categorySubject classification following the arXiv taxonomy.
abstractThe abstract of the survey paper.
structureHierarchical representation of the survey’s organization.
survey_idA unique identifier for the survey.
all_citesList of document IDs cited in the survey.
Bertopic_CDA diversity measure computed using BERTopic.

Literature Knowledge Base

The corpus containing all literature articles is available in: data/corpus.json

Example : Here, we present how articles are organized in the knowledge base. Overly long abstract has been appropriately shortened.

{
"Title": "Information Geometry of Evolution of Neural Network Parameters While Training",
"Authors": [
"Abhiram Anand Thiruthummal",
"Eun-jin Kim",
"Sergiy Shelyag"
],
"Year": "2024",
"Date": "2024-06-07T23:42:54Z",
"Abstract": "Artificial neural networks (ANNs) are powerful tools capable of approximating any arbitrary mathematical function, but their interpretability remains limited...",
"Category": "cs.LG",
"doc_id": 1086990
}

The following are explanations of each field:

KeyDescription
TitleThe title of the research paper.
AuthorsA list of contributing researchers.
YearThe publication year of the paper.
DateThe exact timestamp of the paper’s release.
AbstractThe abstract of the paper.
CategoryThe subject classification following the arXiv taxonomy.
doc_idA unique identifier assigned for reference and retrieval.

Auxiliary Mappings

The mapping containing all queries and their corresponding articles is available in: data/queries.json

Each query in data/queries.json corresponds to a section or paragraph from the ground truth surveys with high citation extraction quality. The associated articles are the references cited in that part of the survey.

Below is an example, Overly long content has been appropriately shortened.

 {
"original_id": "23870233-7f5b-4ef1-9d38-e6f3adb0fa48",
"query_id": 486,
"date": "2020-07-16T09:23:13Z",
"year": "2020",
"category": "cs.LG",
"content": "}\n{\nMachine learning classifiers can perpetuate and amplify the existing systemic injustices in society . Hence, fairness is becoming another important topic. Traditionally...",
"prefix_titles": [
[
"title",
"Learning from Noisy Labels with Deep Neural Networks: A Survey"
],
[
"section",
"Future Research Directions"
],
[
"subsection",
"{Robust and Fair Training"
]
],
"prefix_titles_query": "What are the future research directions for robust and fair training in the context of learning from noisy labels with deep neural networks?",
"cites": [
7771,
4163,
3899,
8740,
8739
],
"cite_extract_rate": 0.8333333333333334,
"origin_cites_number": 6
}

The following are explanations of each field:

KeyDescription
original_idThe identifier for the section where this query is from.
query_idThe ID associated with the specific query.
contentThe content of the section.
prefix_titlesA hierarchical list of titles of the section/subsection/paragraph
prefix_titles_queryThe question this passage is relevant to. The goal of the question is to retrieve relevant documents.
citesA list of document IDs that are cited within this section.
cite_extract_rateThe ratio of extracted citations to the total number of citations in the original document.
origin_cites_numberThe total number of citations originally present in the section.

Note: Not every section in the ground truth surveys has a corresponding entry in queries.json. Only sections with a high citation extraction rate are included.
These entries can be used to train retrieval models, where prefix_titles_query serves as the query and cites contains the relevant document IDs.
The data in queries.json has not been pre-split into training and development sets—you may divide it manually as needed.

Installation Instructions

Requirements

Before you begin, make sure you have the following packages installed in your environment:

FlagEmbedding
bertopic
safetensors
torch==1.13.1+cu117 rouge-score
sacrebleu
numpy==1.26.4
openai
transformers==4.44.2
gdown
socksio

Setting Up Your Environment

Get started with SurGE on your local machine by following these two steps:

1. Download the SurGE Repository

First, clone the SurGE repository to your computer using Git:

git clone https://github.com/oneal2000/SurGE.git surge
cd surge

2. Create and Configure a Python Environment

Next, create an isolated Python environment and install all necessary dependencies. This keeps your project organized and prevents conflicts with other packages.

conda create -n surge python=3.10 -y
conda activate surge
pip install -r requirements.txt -f https://download.pytorch.org/whl/torch_stable.html

Quick Start

  1. First, follow the environment setup instructions in this repository’s README.

  2. Get Data from the link given in Data Release

    python src/download.py
  3. Prepare your generated survey documents in a folder. Refer to Evaluation part in this repository’s README.

    In each folder named after the survey ID, the corresponding generated output should be stored for evaluation. We have provided the articles generated by three baselines on the test set in the paper. For example, with ID baseline as an illustration, the evaluation can be conducted as follows:

    python src/test_final.py --passage_dir ./baselines/ID/output --save_path ./baselines/ID/output/log.json --device 0 --api_key sk-xxx

​you should fill the sk-xxx with your OpenAI API key.

If some of the evaluations are abnormal, you may need to modify the code that sets client in the init of the SurGEvaluator class in evaluator.py to fit the way your API is called

Evaluation

Each ground truth survey has its own ID. In each folder named after the survey ID, the corresponding survey output should be stored for evaluation. You can refer to the example below to try the evaluation. We have provided the articles generated by three baselines on the test set in the paper. For example, with ID baseline as an illustration, the evaluation can be conducted as follows:

python src/test_final.py --passage_dir ./baselines/ID/output --save_path ./baselines/ID/output/log.json --device 0 --api_key sk-xxx

you should fill the sk-xxx with your OpenAI API key.

If some of the evaluations are abnormal, you may need to modify the code that sets client in the init of the SurGEvaluator class in evaluator.py to fit the way your API is called

Arguments Description

--passage_dir : Directory containing generated survey passages.

--eval_list : List of evaluation metrics (space-separated). Default is ALL.

--survey_path : Path to the surveys JSON file. Default is data/surveys.json.

--corpus_path : Path to the corpus JSON file. Default is data/corpus.json.

--device : Device ID for computation. Default is 0.

--api_key : API key for evaluation services.

--save_path : Path to save evaluation results.

As for eval_list, we have the choices below:

ALL : Evaluate all.

ROUGE-BLEU : Evaluate ROUGE and BLEU

SH-Recall : Evaluate SH-Recall

Structure_Quality : Evaluate Structure_Quality(LLM_as_judge)

Coverage : Evaluate Coverage

Relevance-Paper : Evaluate Relevance-Paper

Relevance-Section : Evaluate Relevance-Section

Relevance-Sentence : Evaluate Relevance-Sentence

Logic : Evaluate Logic

License

This project is licensed under MIT License. Please review the LICENSE file for more details.

About

Code for SurGE, SIGIR 2026 Resource Paper

Resources

Stars

19 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

📢 News: this work has been accepted at the SIGIR 2026!

SurGE

Welcome to the official GitHub repository for SurGE. SurGE is a benchmark and dataset for end-to-end scientific survey generation in the computer science domain.

SurGE provides a comprehensive resource for evaluating automated survey generation systems through both a large-scale dataset and a fully automated evaluation framework.

Overview

SurGE is designed to push the boundaries of automated survey generation by tackling the complex task of creating coherent, in-depth survey articles from a vast academic literature collection. Unlike traditional IR tasks focused solely on document retrieval, SurGE requires systems to:

  • Retrieve: Identify relevant academic articles from a corpus of over 1 million papers.
  • Organize: Construct a structured and hierarchical survey outline.
  • Synthesize: Generate a coherent narrative with proper citations, reflecting expert-authored surveys.

The benchmark includes 205 carefully curated ground truth surveys, each accompanied by detailed metadata and a corresponding hierarchical structure, along with an extensive literature knowledge base sourced primarily from arXiv.

Data Release

This repository contains all necessary components for working with the SurGE dataset:

Dataset Files & Formats:

  • Ground Truth Surveys: Each survey includes metadata fields such as title, authors, publication year, abstract, hierarchical structure, and citation lists.
  • Literature Knowledge Base: A corpus of 1,086,992 academic papers with key fields (e.g., title, authors, abstract, publication date, and category).
  • Auxiliary Mappings: Topic-to-publication mappings to support systematic survey generation.

The complete dataset can be downloaded from this Google Drive folder.

Then you will get the folder data .

Ground Truth Survey

A ground truth survey contains the full content of a survey and its citation information. However, due to space constraints, we cannot display it in its entirety here.

All ground truth surveys are available in data/surveys.json

A survey consists of the following fields:

FieldDescription
authorsList of contributing researchers.
survey_titleThe title of the survey paper.
yearThe publication year of the survey.
dateThe exact timestamp of publication.
categorySubject classification following the arXiv taxonomy.
abstractThe abstract of the survey paper.
structureHierarchical representation of the survey’s organization.
survey_idA unique identifier for the survey.
all_citesList of document IDs cited in the survey.
Bertopic_CDA diversity measure computed using BERTopic.

Literature Knowledge Base

The corpus containing all literature articles is available in: data/corpus.json

Example : Here, we present how articles are organized in the knowledge base. Overly long abstract has been appropriately shortened.

{
"Title": "Information Geometry of Evolution of Neural Network Parameters While Training",
"Authors": [
"Abhiram Anand Thiruthummal",
"Eun-jin Kim",
"Sergiy Shelyag"
],
"Year": "2024",
"Date": "2024-06-07T23:42:54Z",
"Abstract": "Artificial neural networks (ANNs) are powerful tools capable of approximating any arbitrary mathematical function, but their interpretability remains limited...",
"Category": "cs.LG",
"doc_id": 1086990
}

The following are explanations of each field:

KeyDescription
TitleThe title of the research paper.
AuthorsA list of contributing researchers.
YearThe publication year of the paper.
DateThe exact timestamp of the paper’s release.
AbstractThe abstract of the paper.
CategoryThe subject classification following the arXiv taxonomy.
doc_idA unique identifier assigned for reference and retrieval.

Auxiliary Mappings

The mapping containing all queries and their corresponding articles is available in: data/queries.json

Each query in data/queries.json corresponds to a section or paragraph from the ground truth surveys with high citation extraction quality. The associated articles are the references cited in that part of the survey.

Below is an example, Overly long content has been appropriately shortened.

 {
"original_id": "23870233-7f5b-4ef1-9d38-e6f3adb0fa48",
"query_id": 486,
"date": "2020-07-16T09:23:13Z",
"year": "2020",
"category": "cs.LG",
"content": "}\n{\nMachine learning classifiers can perpetuate and amplify the existing systemic injustices in society . Hence, fairness is becoming another important topic. Traditionally...",
"prefix_titles": [
[
"title",
"Learning from Noisy Labels with Deep Neural Networks: A Survey"
],
[
"section",
"Future Research Directions"
],
[
"subsection",
"{Robust and Fair Training"
]
],
"prefix_titles_query": "What are the future research directions for robust and fair training in the context of learning from noisy labels with deep neural networks?",
"cites": [
7771,
4163,
3899,
8740,
8739
],
"cite_extract_rate": 0.8333333333333334,
"origin_cites_number": 6
}

The following are explanations of each field:

KeyDescription
original_idThe identifier for the section where this query is from.
query_idThe ID associated with the specific query.
contentThe content of the section.
prefix_titlesA hierarchical list of titles of the section/subsection/paragraph
prefix_titles_queryThe question this passage is relevant to. The goal of the question is to retrieve relevant documents.
citesA list of document IDs that are cited within this section.
cite_extract_rateThe ratio of extracted citations to the total number of citations in the original document.
origin_cites_numberThe total number of citations originally present in the section.

Note: Not every section in the ground truth surveys has a corresponding entry in queries.json. Only sections with a high citation extraction rate are included.
These entries can be used to train retrieval models, where prefix_titles_query serves as the query and cites contains the relevant document IDs.
The data in queries.json has not been pre-split into training and development sets—you may divide it manually as needed.

Installation Instructions

Requirements

Before you begin, make sure you have the following packages installed in your environment:

FlagEmbedding
bertopic
safetensors
torch==1.13.1+cu117 rouge-score
sacrebleu
numpy==1.26.4
openai
transformers==4.44.2
gdown
socksio

Setting Up Your Environment

Get started with SurGE on your local machine by following these two steps:

1. Download the SurGE Repository

First, clone the SurGE repository to your computer using Git:

git clone https://github.com/oneal2000/SurGE.git surge
cd surge

2. Create and Configure a Python Environment

Next, create an isolated Python environment and install all necessary dependencies. This keeps your project organized and prevents conflicts with other packages.

conda create -n surge python=3.10 -y
conda activate surge
pip install -r requirements.txt -f https://download.pytorch.org/whl/torch_stable.html

Quick Start

  1. First, follow the environment setup instructions in this repository’s README.

  2. Get Data from the link given in Data Release

    python src/download.py
  3. Prepare your generated survey documents in a folder. Refer to Evaluation part in this repository’s README.

    In each folder named after the survey ID, the corresponding generated output should be stored for evaluation. We have provided the articles generated by three baselines on the test set in the paper. For example, with ID baseline as an illustration, the evaluation can be conducted as follows:

    python src/test_final.py --passage_dir ./baselines/ID/output --save_path ./baselines/ID/output/log.json --device 0 --api_key sk-xxx

​you should fill the sk-xxx with your OpenAI API key.

If some of the evaluations are abnormal, you may need to modify the code that sets client in the init of the SurGEvaluator class in evaluator.py to fit the way your API is called

Evaluation

Each ground truth survey has its own ID. In each folder named after the survey ID, the corresponding survey output should be stored for evaluation. You can refer to the example below to try the evaluation. We have provided the articles generated by three baselines on the test set in the paper. For example, with ID baseline as an illustration, the evaluation can be conducted as follows:

python src/test_final.py --passage_dir ./baselines/ID/output --save_path ./baselines/ID/output/log.json --device 0 --api_key sk-xxx

you should fill the sk-xxx with your OpenAI API key.

If some of the evaluations are abnormal, you may need to modify the code that sets client in the init of the SurGEvaluator class in evaluator.py to fit the way your API is called

Arguments Description

--passage_dir : Directory containing generated survey passages.

--eval_list : List of evaluation metrics (space-separated). Default is ALL.

--survey_path : Path to the surveys JSON file. Default is data/surveys.json.

--corpus_path : Path to the corpus JSON file. Default is data/corpus.json.

--device : Device ID for computation. Default is 0.

--api_key : API key for evaluation services.

--save_path : Path to save evaluation results.

As for eval_list, we have the choices below:

ALL : Evaluate all.

ROUGE-BLEU : Evaluate ROUGE and BLEU

SH-Recall : Evaluate SH-Recall

Structure_Quality : Evaluate Structure_Quality(LLM_as_judge)

Coverage : Evaluate Coverage

Relevance-Paper : Evaluate Relevance-Paper

Relevance-Section : Evaluate Relevance-Section

Relevance-Sentence : Evaluate Relevance-Sentence

Logic : Evaluate Logic

License

This project is licensed under MIT License. Please review the LICENSE file for more details.

About

Code for SurGE, SIGIR 2026 Resource Paper

Resources

Stars

19 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

📢 News: this work has been accepted at the SIGIR 2026!

SurGE

Welcome to the official GitHub repository for SurGE. SurGE is a benchmark and dataset for end-to-end scientific survey generation in the computer science domain.

SurGE provides a comprehensive resource for evaluating automated survey generation systems through both a large-scale dataset and a fully automated evaluation framework.

Overview

SurGE is designed to push the boundaries of automated survey generation by tackling the complex task of creating coherent, in-depth survey articles from a vast academic literature collection. Unlike traditional IR tasks focused solely on document retrieval, SurGE requires systems to:

  • Retrieve: Identify relevant academic articles from a corpus of over 1 million papers.
  • Organize: Construct a structured and hierarchical survey outline.
  • Synthesize: Generate a coherent narrative with proper citations, reflecting expert-authored surveys.

The benchmark includes 205 carefully curated ground truth surveys, each accompanied by detailed metadata and a corresponding hierarchical structure, along with an extensive literature knowledge base sourced primarily from arXiv.

Data Release

This repository contains all necessary components for working with the SurGE dataset:

Dataset Files & Formats:

  • Ground Truth Surveys: Each survey includes metadata fields such as title, authors, publication year, abstract, hierarchical structure, and citation lists.
  • Literature Knowledge Base: A corpus of 1,086,992 academic papers with key fields (e.g., title, authors, abstract, publication date, and category).
  • Auxiliary Mappings: Topic-to-publication mappings to support systematic survey generation.

The complete dataset can be downloaded from this Google Drive folder.

Then you will get the folder data .

Ground Truth Survey

A ground truth survey contains the full content of a survey and its citation information. However, due to space constraints, we cannot display it in its entirety here.

All ground truth surveys are available in data/surveys.json

A survey consists of the following fields:

FieldDescription
authorsList of contributing researchers.
survey_titleThe title of the survey paper.
yearThe publication year of the survey.
dateThe exact timestamp of publication.
categorySubject classification following the arXiv taxonomy.
abstractThe abstract of the survey paper.
structureHierarchical representation of the survey’s organization.
survey_idA unique identifier for the survey.
all_citesList of document IDs cited in the survey.
Bertopic_CDA diversity measure computed using BERTopic.

Literature Knowledge Base

The corpus containing all literature articles is available in: data/corpus.json

Example : Here, we present how articles are organized in the knowledge base. Overly long abstract has been appropriately shortened.

{
"Title": "Information Geometry of Evolution of Neural Network Parameters While Training",
"Authors": [
"Abhiram Anand Thiruthummal",
"Eun-jin Kim",
"Sergiy Shelyag"
],
"Year": "2024",
"Date": "2024-06-07T23:42:54Z",
"Abstract": "Artificial neural networks (ANNs) are powerful tools capable of approximating any arbitrary mathematical function, but their interpretability remains limited...",
"Category": "cs.LG",
"doc_id": 1086990
}

The following are explanations of each field:

KeyDescription
TitleThe title of the research paper.
AuthorsA list of contributing researchers.
YearThe publication year of the paper.
DateThe exact timestamp of the paper’s release.
AbstractThe abstract of the paper.
CategoryThe subject classification following the arXiv taxonomy.
doc_idA unique identifier assigned for reference and retrieval.

Auxiliary Mappings

The mapping containing all queries and their corresponding articles is available in: data/queries.json

Each query in data/queries.json corresponds to a section or paragraph from the ground truth surveys with high citation extraction quality. The associated articles are the references cited in that part of the survey.

Below is an example, Overly long content has been appropriately shortened.

 {
"original_id": "23870233-7f5b-4ef1-9d38-e6f3adb0fa48",
"query_id": 486,
"date": "2020-07-16T09:23:13Z",
"year": "2020",
"category": "cs.LG",
"content": "}\n{\nMachine learning classifiers can perpetuate and amplify the existing systemic injustices in society . Hence, fairness is becoming another important topic. Traditionally...",
"prefix_titles": [
[
"title",
"Learning from Noisy Labels with Deep Neural Networks: A Survey"
],
[
"section",
"Future Research Directions"
],
[
"subsection",
"{Robust and Fair Training"
]
],
"prefix_titles_query": "What are the future research directions for robust and fair training in the context of learning from noisy labels with deep neural networks?",
"cites": [
7771,
4163,
3899,
8740,
8739
],
"cite_extract_rate": 0.8333333333333334,
"origin_cites_number": 6
}

The following are explanations of each field:

KeyDescription
original_idThe identifier for the section where this query is from.
query_idThe ID associated with the specific query.
contentThe content of the section.
prefix_titlesA hierarchical list of titles of the section/subsection/paragraph
prefix_titles_queryThe question this passage is relevant to. The goal of the question is to retrieve relevant documents.
citesA list of document IDs that are cited within this section.
cite_extract_rateThe ratio of extracted citations to the total number of citations in the original document.
origin_cites_numberThe total number of citations originally present in the section.

Note: Not every section in the ground truth surveys has a corresponding entry in queries.json. Only sections with a high citation extraction rate are included.
These entries can be used to train retrieval models, where prefix_titles_query serves as the query and cites contains the relevant document IDs.
The data in queries.json has not been pre-split into training and development sets—you may divide it manually as needed.

Installation Instructions

Requirements

Before you begin, make sure you have the following packages installed in your environment:

FlagEmbedding
bertopic
safetensors
torch==1.13.1+cu117 rouge-score
sacrebleu
numpy==1.26.4
openai
transformers==4.44.2
gdown
socksio

Setting Up Your Environment

Get started with SurGE on your local machine by following these two steps:

1. Download the SurGE Repository

First, clone the SurGE repository to your computer using Git:

git clone https://github.com/oneal2000/SurGE.git surge
cd surge

2. Create and Configure a Python Environment

Next, create an isolated Python environment and install all necessary dependencies. This keeps your project organized and prevents conflicts with other packages.

conda create -n surge python=3.10 -y
conda activate surge
pip install -r requirements.txt -f https://download.pytorch.org/whl/torch_stable.html

Quick Start

  1. First, follow the environment setup instructions in this repository’s README.

  2. Get Data from the link given in Data Release

    python src/download.py
  3. Prepare your generated survey documents in a folder. Refer to Evaluation part in this repository’s README.

    In each folder named after the survey ID, the corresponding generated output should be stored for evaluation. We have provided the articles generated by three baselines on the test set in the paper. For example, with ID baseline as an illustration, the evaluation can be conducted as follows:

    python src/test_final.py --passage_dir ./baselines/ID/output --save_path ./baselines/ID/output/log.json --device 0 --api_key sk-xxx

​you should fill the sk-xxx with your OpenAI API key.

If some of the evaluations are abnormal, you may need to modify the code that sets client in the init of the SurGEvaluator class in evaluator.py to fit the way your API is called

Evaluation

Each ground truth survey has its own ID. In each folder named after the survey ID, the corresponding survey output should be stored for evaluation. You can refer to the example below to try the evaluation. We have provided the articles generated by three baselines on the test set in the paper. For example, with ID baseline as an illustration, the evaluation can be conducted as follows:

python src/test_final.py --passage_dir ./baselines/ID/output --save_path ./baselines/ID/output/log.json --device 0 --api_key sk-xxx

you should fill the sk-xxx with your OpenAI API key.

If some of the evaluations are abnormal, you may need to modify the code that sets client in the init of the SurGEvaluator class in evaluator.py to fit the way your API is called

Arguments Description

--passage_dir : Directory containing generated survey passages.

--eval_list : List of evaluation metrics (space-separated). Default is ALL.

--survey_path : Path to the surveys JSON file. Default is data/surveys.json.

--corpus_path : Path to the corpus JSON file. Default is data/corpus.json.

--device : Device ID for computation. Default is 0.

--api_key : API key for evaluation services.

--save_path : Path to save evaluation results.

As for eval_list, we have the choices below:

ALL : Evaluate all.

ROUGE-BLEU : Evaluate ROUGE and BLEU

SH-Recall : Evaluate SH-Recall

Structure_Quality : Evaluate Structure_Quality(LLM_as_judge)

Coverage : Evaluate Coverage

Relevance-Paper : Evaluate Relevance-Paper

Relevance-Section : Evaluate Relevance-Section

Relevance-Sentence : Evaluate Relevance-Sentence

Logic : Evaluate Logic

License

This project is licensed under MIT License. Please review the LICENSE file for more details.

About

Code for SurGE, SIGIR 2026 Resource Paper

Resources

Stars

19 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

📢 News: this work has been accepted at the SIGIR 2026!

SurGE

Welcome to the official GitHub repository for SurGE. SurGE is a benchmark and dataset for end-to-end scientific survey generation in the computer science domain.

SurGE provides a comprehensive resource for evaluating automated survey generation systems through both a large-scale dataset and a fully automated evaluation framework.

Overview

SurGE is designed to push the boundaries of automated survey generation by tackling the complex task of creating coherent, in-depth survey articles from a vast academic literature collection. Unlike traditional IR tasks focused solely on document retrieval, SurGE requires systems to:

  • Retrieve: Identify relevant academic articles from a corpus of over 1 million papers.
  • Organize: Construct a structured and hierarchical survey outline.
  • Synthesize: Generate a coherent narrative with proper citations, reflecting expert-authored surveys.

The benchmark includes 205 carefully curated ground truth surveys, each accompanied by detailed metadata and a corresponding hierarchical structure, along with an extensive literature knowledge base sourced primarily from arXiv.

Data Release

This repository contains all necessary components for working with the SurGE dataset:

Dataset Files & Formats:

  • Ground Truth Surveys: Each survey includes metadata fields such as title, authors, publication year, abstract, hierarchical structure, and citation lists.
  • Literature Knowledge Base: A corpus of 1,086,992 academic papers with key fields (e.g., title, authors, abstract, publication date, and category).
  • Auxiliary Mappings: Topic-to-publication mappings to support systematic survey generation.

The complete dataset can be downloaded from this Google Drive folder.

Then you will get the folder data .

Ground Truth Survey

A ground truth survey contains the full content of a survey and its citation information. However, due to space constraints, we cannot display it in its entirety here.

All ground truth surveys are available in data/surveys.json

A survey consists of the following fields:

FieldDescription
authorsList of contributing researchers.
survey_titleThe title of the survey paper.
yearThe publication year of the survey.
dateThe exact timestamp of publication.
categorySubject classification following the arXiv taxonomy.
abstractThe abstract of the survey paper.
structureHierarchical representation of the survey’s organization.
survey_idA unique identifier for the survey.
all_citesList of document IDs cited in the survey.
Bertopic_CDA diversity measure computed using BERTopic.

Literature Knowledge Base

The corpus containing all literature articles is available in: data/corpus.json

Example : Here, we present how articles are organized in the knowledge base. Overly long abstract has been appropriately shortened.

{
"Title": "Information Geometry of Evolution of Neural Network Parameters While Training",
"Authors": [
"Abhiram Anand Thiruthummal",
"Eun-jin Kim",
"Sergiy Shelyag"
],
"Year": "2024",
"Date": "2024-06-07T23:42:54Z",
"Abstract": "Artificial neural networks (ANNs) are powerful tools capable of approximating any arbitrary mathematical function, but their interpretability remains limited...",
"Category": "cs.LG",
"doc_id": 1086990
}

The following are explanations of each field:

KeyDescription
TitleThe title of the research paper.
AuthorsA list of contributing researchers.
YearThe publication year of the paper.
DateThe exact timestamp of the paper’s release.
AbstractThe abstract of the paper.
CategoryThe subject classification following the arXiv taxonomy.
doc_idA unique identifier assigned for reference and retrieval.

Auxiliary Mappings

The mapping containing all queries and their corresponding articles is available in: data/queries.json

Each query in data/queries.json corresponds to a section or paragraph from the ground truth surveys with high citation extraction quality. The associated articles are the references cited in that part of the survey.

Below is an example, Overly long content has been appropriately shortened.

 {
"original_id": "23870233-7f5b-4ef1-9d38-e6f3adb0fa48",
"query_id": 486,
"date": "2020-07-16T09:23:13Z",
"year": "2020",
"category": "cs.LG",
"content": "}\n{\nMachine learning classifiers can perpetuate and amplify the existing systemic injustices in society . Hence, fairness is becoming another important topic. Traditionally...",
"prefix_titles": [
[
"title",
"Learning from Noisy Labels with Deep Neural Networks: A Survey"
],
[
"section",
"Future Research Directions"
],
[
"subsection",
"{Robust and Fair Training"
]
],
"prefix_titles_query": "What are the future research directions for robust and fair training in the context of learning from noisy labels with deep neural networks?",
"cites": [
7771,
4163,
3899,
8740,
8739
],
"cite_extract_rate": 0.8333333333333334,
"origin_cites_number": 6
}

The following are explanations of each field:

KeyDescription
original_idThe identifier for the section where this query is from.
query_idThe ID associated with the specific query.
contentThe content of the section.
prefix_titlesA hierarchical list of titles of the section/subsection/paragraph
prefix_titles_queryThe question this passage is relevant to. The goal of the question is to retrieve relevant documents.
citesA list of document IDs that are cited within this section.
cite_extract_rateThe ratio of extracted citations to the total number of citations in the original document.
origin_cites_numberThe total number of citations originally present in the section.

Note: Not every section in the ground truth surveys has a corresponding entry in queries.json. Only sections with a high citation extraction rate are included.
These entries can be used to train retrieval models, where prefix_titles_query serves as the query and cites contains the relevant document IDs.
The data in queries.json has not been pre-split into training and development sets—you may divide it manually as needed.

Installation Instructions

Requirements

Before you begin, make sure you have the following packages installed in your environment:

FlagEmbedding
bertopic
safetensors
torch==1.13.1+cu117 rouge-score
sacrebleu
numpy==1.26.4
openai
transformers==4.44.2
gdown
socksio

Setting Up Your Environment

Get started with SurGE on your local machine by following these two steps:

1. Download the SurGE Repository

First, clone the SurGE repository to your computer using Git:

git clone https://github.com/oneal2000/SurGE.git surge
cd surge

2. Create and Configure a Python Environment

Next, create an isolated Python environment and install all necessary dependencies. This keeps your project organized and prevents conflicts with other packages.

conda create -n surge python=3.10 -y
conda activate surge
pip install -r requirements.txt -f https://download.pytorch.org/whl/torch_stable.html

Quick Start

  1. First, follow the environment setup instructions in this repository’s README.

  2. Get Data from the link given in Data Release

    python src/download.py
  3. Prepare your generated survey documents in a folder. Refer to Evaluation part in this repository’s README.

    In each folder named after the survey ID, the corresponding generated output should be stored for evaluation. We have provided the articles generated by three baselines on the test set in the paper. For example, with ID baseline as an illustration, the evaluation can be conducted as follows:

    python src/test_final.py --passage_dir ./baselines/ID/output --save_path ./baselines/ID/output/log.json --device 0 --api_key sk-xxx

​you should fill the sk-xxx with your OpenAI API key.

If some of the evaluations are abnormal, you may need to modify the code that sets client in the init of the SurGEvaluator class in evaluator.py to fit the way your API is called

Evaluation

Each ground truth survey has its own ID. In each folder named after the survey ID, the corresponding survey output should be stored for evaluation. You can refer to the example below to try the evaluation. We have provided the articles generated by three baselines on the test set in the paper. For example, with ID baseline as an illustration, the evaluation can be conducted as follows:

python src/test_final.py --passage_dir ./baselines/ID/output --save_path ./baselines/ID/output/log.json --device 0 --api_key sk-xxx

you should fill the sk-xxx with your OpenAI API key.

If some of the evaluations are abnormal, you may need to modify the code that sets client in the init of the SurGEvaluator class in evaluator.py to fit the way your API is called

Arguments Description

--passage_dir : Directory containing generated survey passages.

--eval_list : List of evaluation metrics (space-separated). Default is ALL.

--survey_path : Path to the surveys JSON file. Default is data/surveys.json.

--corpus_path : Path to the corpus JSON file. Default is data/corpus.json.

--device : Device ID for computation. Default is 0.

--api_key : API key for evaluation services.

--save_path : Path to save evaluation results.

As for eval_list, we have the choices below:

ALL : Evaluate all.

ROUGE-BLEU : Evaluate ROUGE and BLEU

SH-Recall : Evaluate SH-Recall

Structure_Quality : Evaluate Structure_Quality(LLM_as_judge)

Coverage : Evaluate Coverage

Relevance-Paper : Evaluate Relevance-Paper

Relevance-Section : Evaluate Relevance-Section

Relevance-Sentence : Evaluate Relevance-Sentence

Logic : Evaluate Logic

License

This project is licensed under MIT License. Please review the LICENSE file for more details.

About

Code for SurGE, SIGIR 2026 Resource Paper

Resources

Stars

19 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

📢 News: this work has been accepted at the SIGIR 2026!

SurGE

Welcome to the official GitHub repository for SurGE. SurGE is a benchmark and dataset for end-to-end scientific survey generation in the computer science domain.

SurGE provides a comprehensive resource for evaluating automated survey generation systems through both a large-scale dataset and a fully automated evaluation framework.

Overview

SurGE is designed to push the boundaries of automated survey generation by tackling the complex task of creating coherent, in-depth survey articles from a vast academic literature collection. Unlike traditional IR tasks focused solely on document retrieval, SurGE requires systems to:

  • Retrieve: Identify relevant academic articles from a corpus of over 1 million papers.
  • Organize: Construct a structured and hierarchical survey outline.
  • Synthesize: Generate a coherent narrative with proper citations, reflecting expert-authored surveys.

The benchmark includes 205 carefully curated ground truth surveys, each accompanied by detailed metadata and a corresponding hierarchical structure, along with an extensive literature knowledge base sourced primarily from arXiv.

Data Release

This repository contains all necessary components for working with the SurGE dataset:

Dataset Files & Formats:

  • Ground Truth Surveys: Each survey includes metadata fields such as title, authors, publication year, abstract, hierarchical structure, and citation lists.
  • Literature Knowledge Base: A corpus of 1,086,992 academic papers with key fields (e.g., title, authors, abstract, publication date, and category).
  • Auxiliary Mappings: Topic-to-publication mappings to support systematic survey generation.

The complete dataset can be downloaded from this Google Drive folder.

Then you will get the folder data .

Ground Truth Survey

A ground truth survey contains the full content of a survey and its citation information. However, due to space constraints, we cannot display it in its entirety here.

All ground truth surveys are available in data/surveys.json

A survey consists of the following fields:

FieldDescription
authorsList of contributing researchers.
survey_titleThe title of the survey paper.
yearThe publication year of the survey.
dateThe exact timestamp of publication.
categorySubject classification following the arXiv taxonomy.
abstractThe abstract of the survey paper.
structureHierarchical representation of the survey’s organization.
survey_idA unique identifier for the survey.
all_citesList of document IDs cited in the survey.
Bertopic_CDA diversity measure computed using BERTopic.

Literature Knowledge Base

The corpus containing all literature articles is available in: data/corpus.json

Example : Here, we present how articles are organized in the knowledge base. Overly long abstract has been appropriately shortened.

{
"Title": "Information Geometry of Evolution of Neural Network Parameters While Training",
"Authors": [
"Abhiram Anand Thiruthummal",
"Eun-jin Kim",
"Sergiy Shelyag"
],
"Year": "2024",
"Date": "2024-06-07T23:42:54Z",
"Abstract": "Artificial neural networks (ANNs) are powerful tools capable of approximating any arbitrary mathematical function, but their interpretability remains limited...",
"Category": "cs.LG",
"doc_id": 1086990
}

The following are explanations of each field:

KeyDescription
TitleThe title of the research paper.
AuthorsA list of contributing researchers.
YearThe publication year of the paper.
DateThe exact timestamp of the paper’s release.
AbstractThe abstract of the paper.
CategoryThe subject classification following the arXiv taxonomy.
doc_idA unique identifier assigned for reference and retrieval.

Auxiliary Mappings

The mapping containing all queries and their corresponding articles is available in: data/queries.json

Each query in data/queries.json corresponds to a section or paragraph from the ground truth surveys with high citation extraction quality. The associated articles are the references cited in that part of the survey.

Below is an example, Overly long content has been appropriately shortened.

 {
"original_id": "23870233-7f5b-4ef1-9d38-e6f3adb0fa48",
"query_id": 486,
"date": "2020-07-16T09:23:13Z",
"year": "2020",
"category": "cs.LG",
"content": "}\n{\nMachine learning classifiers can perpetuate and amplify the existing systemic injustices in society . Hence, fairness is becoming another important topic. Traditionally...",
"prefix_titles": [
[
"title",
"Learning from Noisy Labels with Deep Neural Networks: A Survey"
],
[
"section",
"Future Research Directions"
],
[
"subsection",
"{Robust and Fair Training"
]
],
"prefix_titles_query": "What are the future research directions for robust and fair training in the context of learning from noisy labels with deep neural networks?",
"cites": [
7771,
4163,
3899,
8740,
8739
],
"cite_extract_rate": 0.8333333333333334,
"origin_cites_number": 6
}

The following are explanations of each field:

KeyDescription
original_idThe identifier for the section where this query is from.
query_idThe ID associated with the specific query.
contentThe content of the section.
prefix_titlesA hierarchical list of titles of the section/subsection/paragraph
prefix_titles_queryThe question this passage is relevant to. The goal of the question is to retrieve relevant documents.
citesA list of document IDs that are cited within this section.
cite_extract_rateThe ratio of extracted citations to the total number of citations in the original document.
origin_cites_numberThe total number of citations originally present in the section.

Note: Not every section in the ground truth surveys has a corresponding entry in queries.json. Only sections with a high citation extraction rate are included.
These entries can be used to train retrieval models, where prefix_titles_query serves as the query and cites contains the relevant document IDs.
The data in queries.json has not been pre-split into training and development sets—you may divide it manually as needed.

Installation Instructions

Requirements

Before you begin, make sure you have the following packages installed in your environment:

FlagEmbedding
bertopic
safetensors
torch==1.13.1+cu117 rouge-score
sacrebleu
numpy==1.26.4
openai
transformers==4.44.2
gdown
socksio

Setting Up Your Environment

Get started with SurGE on your local machine by following these two steps:

1. Download the SurGE Repository

First, clone the SurGE repository to your computer using Git:

git clone https://github.com/oneal2000/SurGE.git surge
cd surge

2. Create and Configure a Python Environment

Next, create an isolated Python environment and install all necessary dependencies. This keeps your project organized and prevents conflicts with other packages.

conda create -n surge python=3.10 -y
conda activate surge
pip install -r requirements.txt -f https://download.pytorch.org/whl/torch_stable.html

Quick Start

  1. First, follow the environment setup instructions in this repository’s README.

  2. Get Data from the link given in Data Release

    python src/download.py
  3. Prepare your generated survey documents in a folder. Refer to Evaluation part in this repository’s README.

    In each folder named after the survey ID, the corresponding generated output should be stored for evaluation. We have provided the articles generated by three baselines on the test set in the paper. For example, with ID baseline as an illustration, the evaluation can be conducted as follows:

    python src/test_final.py --passage_dir ./baselines/ID/output --save_path ./baselines/ID/output/log.json --device 0 --api_key sk-xxx

​you should fill the sk-xxx with your OpenAI API key.

If some of the evaluations are abnormal, you may need to modify the code that sets client in the init of the SurGEvaluator class in evaluator.py to fit the way your API is called

Evaluation

Each ground truth survey has its own ID. In each folder named after the survey ID, the corresponding survey output should be stored for evaluation. You can refer to the example below to try the evaluation. We have provided the articles generated by three baselines on the test set in the paper. For example, with ID baseline as an illustration, the evaluation can be conducted as follows:

python src/test_final.py --passage_dir ./baselines/ID/output --save_path ./baselines/ID/output/log.json --device 0 --api_key sk-xxx

you should fill the sk-xxx with your OpenAI API key.

If some of the evaluations are abnormal, you may need to modify the code that sets client in the init of the SurGEvaluator class in evaluator.py to fit the way your API is called

Arguments Description

--passage_dir : Directory containing generated survey passages.

--eval_list : List of evaluation metrics (space-separated). Default is ALL.

--survey_path : Path to the surveys JSON file. Default is data/surveys.json.

--corpus_path : Path to the corpus JSON file. Default is data/corpus.json.

--device : Device ID for computation. Default is 0.

--api_key : API key for evaluation services.

--save_path : Path to save evaluation results.

As for eval_list, we have the choices below:

ALL : Evaluate all.

ROUGE-BLEU : Evaluate ROUGE and BLEU

SH-Recall : Evaluate SH-Recall

Structure_Quality : Evaluate Structure_Quality(LLM_as_judge)

Coverage : Evaluate Coverage

Relevance-Paper : Evaluate Relevance-Paper

Relevance-Section : Evaluate Relevance-Section

Relevance-Sentence : Evaluate Relevance-Sentence

Logic : Evaluate Logic

License

This project is licensed under MIT License. Please review the LICENSE file for more details.

About

Code for SurGE, SIGIR 2026 Resource Paper

Resources

Stars

19 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

📢 News: this work has been accepted at the SIGIR 2026!

SurGE

Welcome to the official GitHub repository for SurGE. SurGE is a benchmark and dataset for end-to-end scientific survey generation in the computer science domain.

SurGE provides a comprehensive resource for evaluating automated survey generation systems through both a large-scale dataset and a fully automated evaluation framework.

Overview

SurGE is designed to push the boundaries of automated survey generation by tackling the complex task of creating coherent, in-depth survey articles from a vast academic literature collection. Unlike traditional IR tasks focused solely on document retrieval, SurGE requires systems to:

  • Retrieve: Identify relevant academic articles from a corpus of over 1 million papers.
  • Organize: Construct a structured and hierarchical survey outline.
  • Synthesize: Generate a coherent narrative with proper citations, reflecting expert-authored surveys.

The benchmark includes 205 carefully curated ground truth surveys, each accompanied by detailed metadata and a corresponding hierarchical structure, along with an extensive literature knowledge base sourced primarily from arXiv.

Data Release

This repository contains all necessary components for working with the SurGE dataset:

Dataset Files & Formats:

  • Ground Truth Surveys: Each survey includes metadata fields such as title, authors, publication year, abstract, hierarchical structure, and citation lists.
  • Literature Knowledge Base: A corpus of 1,086,992 academic papers with key fields (e.g., title, authors, abstract, publication date, and category).
  • Auxiliary Mappings: Topic-to-publication mappings to support systematic survey generation.

The complete dataset can be downloaded from this Google Drive folder.

Then you will get the folder data .

Ground Truth Survey

A ground truth survey contains the full content of a survey and its citation information. However, due to space constraints, we cannot display it in its entirety here.

All ground truth surveys are available in data/surveys.json

A survey consists of the following fields:

FieldDescription
authorsList of contributing researchers.
survey_titleThe title of the survey paper.
yearThe publication year of the survey.
dateThe exact timestamp of publication.
categorySubject classification following the arXiv taxonomy.
abstractThe abstract of the survey paper.
structureHierarchical representation of the survey’s organization.
survey_idA unique identifier for the survey.
all_citesList of document IDs cited in the survey.
Bertopic_CDA diversity measure computed using BERTopic.

Literature Knowledge Base

The corpus containing all literature articles is available in: data/corpus.json

Example : Here, we present how articles are organized in the knowledge base. Overly long abstract has been appropriately shortened.

{
"Title": "Information Geometry of Evolution of Neural Network Parameters While Training",
"Authors": [
"Abhiram Anand Thiruthummal",
"Eun-jin Kim",
"Sergiy Shelyag"
],
"Year": "2024",
"Date": "2024-06-07T23:42:54Z",
"Abstract": "Artificial neural networks (ANNs) are powerful tools capable of approximating any arbitrary mathematical function, but their interpretability remains limited...",
"Category": "cs.LG",
"doc_id": 1086990
}

The following are explanations of each field:

KeyDescription
TitleThe title of the research paper.
AuthorsA list of contributing researchers.
YearThe publication year of the paper.
DateThe exact timestamp of the paper’s release.
AbstractThe abstract of the paper.
CategoryThe subject classification following the arXiv taxonomy.
doc_idA unique identifier assigned for reference and retrieval.

Auxiliary Mappings

The mapping containing all queries and their corresponding articles is available in: data/queries.json

Each query in data/queries.json corresponds to a section or paragraph from the ground truth surveys with high citation extraction quality. The associated articles are the references cited in that part of the survey.

Below is an example, Overly long content has been appropriately shortened.

 {
"original_id": "23870233-7f5b-4ef1-9d38-e6f3adb0fa48",
"query_id": 486,
"date": "2020-07-16T09:23:13Z",
"year": "2020",
"category": "cs.LG",
"content": "}\n{\nMachine learning classifiers can perpetuate and amplify the existing systemic injustices in society . Hence, fairness is becoming another important topic. Traditionally...",
"prefix_titles": [
[
"title",
"Learning from Noisy Labels with Deep Neural Networks: A Survey"
],
[
"section",
"Future Research Directions"
],
[
"subsection",
"{Robust and Fair Training"
]
],
"prefix_titles_query": "What are the future research directions for robust and fair training in the context of learning from noisy labels with deep neural networks?",
"cites": [
7771,
4163,
3899,
8740,
8739
],
"cite_extract_rate": 0.8333333333333334,
"origin_cites_number": 6
}

The following are explanations of each field:

KeyDescription
original_idThe identifier for the section where this query is from.
query_idThe ID associated with the specific query.
contentThe content of the section.
prefix_titlesA hierarchical list of titles of the section/subsection/paragraph
prefix_titles_queryThe question this passage is relevant to. The goal of the question is to retrieve relevant documents.
citesA list of document IDs that are cited within this section.
cite_extract_rateThe ratio of extracted citations to the total number of citations in the original document.
origin_cites_numberThe total number of citations originally present in the section.

Note: Not every section in the ground truth surveys has a corresponding entry in queries.json. Only sections with a high citation extraction rate are included.
These entries can be used to train retrieval models, where prefix_titles_query serves as the query and cites contains the relevant document IDs.
The data in queries.json has not been pre-split into training and development sets—you may divide it manually as needed.

Installation Instructions

Requirements

Before you begin, make sure you have the following packages installed in your environment:

FlagEmbedding
bertopic
safetensors
torch==1.13.1+cu117 rouge-score
sacrebleu
numpy==1.26.4
openai
transformers==4.44.2
gdown
socksio

Setting Up Your Environment

Get started with SurGE on your local machine by following these two steps:

1. Download the SurGE Repository

First, clone the SurGE repository to your computer using Git:

git clone https://github.com/oneal2000/SurGE.git surge
cd surge

2. Create and Configure a Python Environment

Next, create an isolated Python environment and install all necessary dependencies. This keeps your project organized and prevents conflicts with other packages.

conda create -n surge python=3.10 -y
conda activate surge
pip install -r requirements.txt -f https://download.pytorch.org/whl/torch_stable.html

Quick Start

  1. First, follow the environment setup instructions in this repository’s README.

  2. Get Data from the link given in Data Release

    python src/download.py
  3. Prepare your generated survey documents in a folder. Refer to Evaluation part in this repository’s README.

    In each folder named after the survey ID, the corresponding generated output should be stored for evaluation. We have provided the articles generated by three baselines on the test set in the paper. For example, with ID baseline as an illustration, the evaluation can be conducted as follows:

    python src/test_final.py --passage_dir ./baselines/ID/output --save_path ./baselines/ID/output/log.json --device 0 --api_key sk-xxx

​you should fill the sk-xxx with your OpenAI API key.

If some of the evaluations are abnormal, you may need to modify the code that sets client in the init of the SurGEvaluator class in evaluator.py to fit the way your API is called

Evaluation

Each ground truth survey has its own ID. In each folder named after the survey ID, the corresponding survey output should be stored for evaluation. You can refer to the example below to try the evaluation. We have provided the articles generated by three baselines on the test set in the paper. For example, with ID baseline as an illustration, the evaluation can be conducted as follows:

python src/test_final.py --passage_dir ./baselines/ID/output --save_path ./baselines/ID/output/log.json --device 0 --api_key sk-xxx

you should fill the sk-xxx with your OpenAI API key.

If some of the evaluations are abnormal, you may need to modify the code that sets client in the init of the SurGEvaluator class in evaluator.py to fit the way your API is called

Arguments Description

--passage_dir : Directory containing generated survey passages.

--eval_list : List of evaluation metrics (space-separated). Default is ALL.

--survey_path : Path to the surveys JSON file. Default is data/surveys.json.

--corpus_path : Path to the corpus JSON file. Default is data/corpus.json.

--device : Device ID for computation. Default is 0.

--api_key : API key for evaluation services.

--save_path : Path to save evaluation results.

As for eval_list, we have the choices below:

ALL : Evaluate all.

ROUGE-BLEU : Evaluate ROUGE and BLEU

SH-Recall : Evaluate SH-Recall

Structure_Quality : Evaluate Structure_Quality(LLM_as_judge)

Coverage : Evaluate Coverage

Relevance-Paper : Evaluate Relevance-Paper

Relevance-Section : Evaluate Relevance-Section

Relevance-Sentence : Evaluate Relevance-Sentence

Logic : Evaluate Logic

License

This project is licensed under MIT License. Please review the LICENSE file for more details.

About

Code for SurGE, SIGIR 2026 Resource Paper

Resources

Stars

19 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

📢 News: this work has been accepted at the SIGIR 2026!

SurGE

Welcome to the official GitHub repository for SurGE. SurGE is a benchmark and dataset for end-to-end scientific survey generation in the computer science domain.

SurGE provides a comprehensive resource for evaluating automated survey generation systems through both a large-scale dataset and a fully automated evaluation framework.

Overview

SurGE is designed to push the boundaries of automated survey generation by tackling the complex task of creating coherent, in-depth survey articles from a vast academic literature collection. Unlike traditional IR tasks focused solely on document retrieval, SurGE requires systems to:

  • Retrieve: Identify relevant academic articles from a corpus of over 1 million papers.
  • Organize: Construct a structured and hierarchical survey outline.
  • Synthesize: Generate a coherent narrative with proper citations, reflecting expert-authored surveys.

The benchmark includes 205 carefully curated ground truth surveys, each accompanied by detailed metadata and a corresponding hierarchical structure, along with an extensive literature knowledge base sourced primarily from arXiv.

Data Release

This repository contains all necessary components for working with the SurGE dataset:

Dataset Files & Formats:

  • Ground Truth Surveys: Each survey includes metadata fields such as title, authors, publication year, abstract, hierarchical structure, and citation lists.
  • Literature Knowledge Base: A corpus of 1,086,992 academic papers with key fields (e.g., title, authors, abstract, publication date, and category).
  • Auxiliary Mappings: Topic-to-publication mappings to support systematic survey generation.

The complete dataset can be downloaded from this Google Drive folder.

Then you will get the folder data .

Ground Truth Survey

A ground truth survey contains the full content of a survey and its citation information. However, due to space constraints, we cannot display it in its entirety here.

All ground truth surveys are available in data/surveys.json

A survey consists of the following fields:

FieldDescription
authorsList of contributing researchers.
survey_titleThe title of the survey paper.
yearThe publication year of the survey.
dateThe exact timestamp of publication.
categorySubject classification following the arXiv taxonomy.
abstractThe abstract of the survey paper.
structureHierarchical representation of the survey’s organization.
survey_idA unique identifier for the survey.
all_citesList of document IDs cited in the survey.
Bertopic_CDA diversity measure computed using BERTopic.

Literature Knowledge Base

The corpus containing all literature articles is available in: data/corpus.json

Example : Here, we present how articles are organized in the knowledge base. Overly long abstract has been appropriately shortened.

{
"Title": "Information Geometry of Evolution of Neural Network Parameters While Training",
"Authors": [
"Abhiram Anand Thiruthummal",
"Eun-jin Kim",
"Sergiy Shelyag"
],
"Year": "2024",
"Date": "2024-06-07T23:42:54Z",
"Abstract": "Artificial neural networks (ANNs) are powerful tools capable of approximating any arbitrary mathematical function, but their interpretability remains limited...",
"Category": "cs.LG",
"doc_id": 1086990
}

The following are explanations of each field:

KeyDescription
TitleThe title of the research paper.
AuthorsA list of contributing researchers.
YearThe publication year of the paper.
DateThe exact timestamp of the paper’s release.
AbstractThe abstract of the paper.
CategoryThe subject classification following the arXiv taxonomy.
doc_idA unique identifier assigned for reference and retrieval.

Auxiliary Mappings

The mapping containing all queries and their corresponding articles is available in: data/queries.json

Each query in data/queries.json corresponds to a section or paragraph from the ground truth surveys with high citation extraction quality. The associated articles are the references cited in that part of the survey.

Below is an example, Overly long content has been appropriately shortened.

 {
"original_id": "23870233-7f5b-4ef1-9d38-e6f3adb0fa48",
"query_id": 486,
"date": "2020-07-16T09:23:13Z",
"year": "2020",
"category": "cs.LG",
"content": "}\n{\nMachine learning classifiers can perpetuate and amplify the existing systemic injustices in society . Hence, fairness is becoming another important topic. Traditionally...",
"prefix_titles": [
[
"title",
"Learning from Noisy Labels with Deep Neural Networks: A Survey"
],
[
"section",
"Future Research Directions"
],
[
"subsection",
"{Robust and Fair Training"
]
],
"prefix_titles_query": "What are the future research directions for robust and fair training in the context of learning from noisy labels with deep neural networks?",
"cites": [
7771,
4163,
3899,
8740,
8739
],
"cite_extract_rate": 0.8333333333333334,
"origin_cites_number": 6
}

The following are explanations of each field:

KeyDescription
original_idThe identifier for the section where this query is from.
query_idThe ID associated with the specific query.
contentThe content of the section.
prefix_titlesA hierarchical list of titles of the section/subsection/paragraph
prefix_titles_queryThe question this passage is relevant to. The goal of the question is to retrieve relevant documents.
citesA list of document IDs that are cited within this section.
cite_extract_rateThe ratio of extracted citations to the total number of citations in the original document.
origin_cites_numberThe total number of citations originally present in the section.

Note: Not every section in the ground truth surveys has a corresponding entry in queries.json. Only sections with a high citation extraction rate are included.
These entries can be used to train retrieval models, where prefix_titles_query serves as the query and cites contains the relevant document IDs.
The data in queries.json has not been pre-split into training and development sets—you may divide it manually as needed.

Installation Instructions

Requirements

Before you begin, make sure you have the following packages installed in your environment:

FlagEmbedding
bertopic
safetensors
torch==1.13.1+cu117 rouge-score
sacrebleu
numpy==1.26.4
openai
transformers==4.44.2
gdown
socksio

Setting Up Your Environment

Get started with SurGE on your local machine by following these two steps:

1. Download the SurGE Repository

First, clone the SurGE repository to your computer using Git:

git clone https://github.com/oneal2000/SurGE.git surge
cd surge

2. Create and Configure a Python Environment

Next, create an isolated Python environment and install all necessary dependencies. This keeps your project organized and prevents conflicts with other packages.

conda create -n surge python=3.10 -y
conda activate surge
pip install -r requirements.txt -f https://download.pytorch.org/whl/torch_stable.html

Quick Start

  1. First, follow the environment setup instructions in this repository’s README.

  2. Get Data from the link given in Data Release

    python src/download.py
  3. Prepare your generated survey documents in a folder. Refer to Evaluation part in this repository’s README.

    In each folder named after the survey ID, the corresponding generated output should be stored for evaluation. We have provided the articles generated by three baselines on the test set in the paper. For example, with ID baseline as an illustration, the evaluation can be conducted as follows:

    python src/test_final.py --passage_dir ./baselines/ID/output --save_path ./baselines/ID/output/log.json --device 0 --api_key sk-xxx

​you should fill the sk-xxx with your OpenAI API key.

If some of the evaluations are abnormal, you may need to modify the code that sets client in the init of the SurGEvaluator class in evaluator.py to fit the way your API is called

Evaluation

Each ground truth survey has its own ID. In each folder named after the survey ID, the corresponding survey output should be stored for evaluation. You can refer to the example below to try the evaluation. We have provided the articles generated by three baselines on the test set in the paper. For example, with ID baseline as an illustration, the evaluation can be conducted as follows:

python src/test_final.py --passage_dir ./baselines/ID/output --save_path ./baselines/ID/output/log.json --device 0 --api_key sk-xxx

you should fill the sk-xxx with your OpenAI API key.

If some of the evaluations are abnormal, you may need to modify the code that sets client in the init of the SurGEvaluator class in evaluator.py to fit the way your API is called

Arguments Description

--passage_dir : Directory containing generated survey passages.

--eval_list : List of evaluation metrics (space-separated). Default is ALL.

--survey_path : Path to the surveys JSON file. Default is data/surveys.json.

--corpus_path : Path to the corpus JSON file. Default is data/corpus.json.

--device : Device ID for computation. Default is 0.

--api_key : API key for evaluation services.

--save_path : Path to save evaluation results.

As for eval_list, we have the choices below:

ALL : Evaluate all.

ROUGE-BLEU : Evaluate ROUGE and BLEU

SH-Recall : Evaluate SH-Recall

Structure_Quality : Evaluate Structure_Quality(LLM_as_judge)

Coverage : Evaluate Coverage

Relevance-Paper : Evaluate Relevance-Paper

Relevance-Section : Evaluate Relevance-Section

Relevance-Sentence : Evaluate Relevance-Sentence

Logic : Evaluate Logic

License

This project is licensed under MIT License. Please review the LICENSE file for more details.

About

Code for SurGE, SIGIR 2026 Resource Paper

Resources

Stars

19 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages