Repository files navigation

SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science

English | 简体中文

SeedBench is the first multi-task benchmark designed to evaluate large language models (LLMs) in seed science, focusing on seed breeding. This repository includes the dataset, evaluation code, and documentation to support research in this domain. Here is the usage.


🌾 Overview

SeedBench assesses LLMs across three core seed breeding stages:

  • Gene Information Retrieval
  • Gene Function and Regulation Analysis
  • Variety Breeding with Agronomic Trait Optimization

Breeding Workflow
Breeding Expert Workflow Framework

Built with domain experts, SeedBench features 2,264 expert-validated questions across 11 task types and 10 subcategories, initially targeting rice breeding. Future updates will include other crops like maize, soybean, and wheat.

🔎 Dataset Details

  • Corpus: 308,727 publications cleaned to 1.1 billion tokens; 279 segments from 113 documents.

  • Questions: 2,264 across 11 task types, bilingual (English/Chinese), expert-validated.

  • Focus: Rice breeding as a representative case.

    Types and metrics:

    Type IDQuestion TypeMetricCount
    Q&A
    QA-1Multiple ChoiceAccuracy200
    QA-2Multiple AnswerMacro-F1187
    QA-3Fill-in-the-BlankROUGE-L224
    QA-4GenerationROUGE-L242
    Summarization
    SUM-1Simple SummarizationROUGE-L225
    SUM-2Key Information ExtractionROUGE-L225
    Reading Comprehension
    RC-1Multiple ChoiceAccuracy113
    RC-2Multiple AnswerMacro-F1108
    RC-3Fill-in-the-BlankROUGE-L221
    RC-4GenerationROUGE-L240
    RC-5Subcategory ClassificationAccuracy279

    Taxonomy Distribution:

    Taxonomy Distribution

☀️ Key Results

We evaluated 26 LLMs, including proprietary, open-source, and domain-specific models. Highlights:

Performance by Question Type

  • Top Performers: DeepSeek-V3 (68.37), GPT-4 (67.88).

    Proprietary LLM Radar

    Open-Source LLM Radar

Performance by Task Types

ModelQA-1QA-2QA-3QA-4SUM-1SUM-2RC-1RC-2RC-3RC-4RC-5Avg
GPT-460.5073.8721.3536.0758.7362.89100.0096.4487.8662.2986.7467.88
DeepSeek-V372.5079.8429.2940.6348.0654.67100.0097.2287.8955.1986.7468.37
Qwen2-72B59.5075.9819.5531.6231.0863.0999.1294.2472.2051.5889.9662.54

Performance by Subcategory

ModelC1C2C3C4C5C6C7C8C9C10Avg
GPT-459.5960.5576.3261.1656.3459.3563.6764.7460.6567.6662.06
DeepSeek-V3-671B56.0362.4274.8163.1755.2358.8468.2369.0466.4668.4863.30
Qwen2-72B51.1658.1074.0759.7251.5857.7658.8561.6356.6959.1157.62
  • Top Performers: DeepSeek-V3-671B (63.30), GPT-4 (62.06).

🐝 Repository Contents

  • base_model_eval/: Used to test base models without dialogue capabilities, i.e., evaluating performance after pretraining.
  • sft_model_eval/: Used to test SFT (Supervised Fine-Tuning) models, with a total of 2,264 questions covering 10 subcategories (see Fig 2).
    • one-shot/: Organized by 11 task types (see Tab 1).
    • zero-shot/: Organized by 11 task types (see Tab 1).
  • corpus/: 279 high-quality text segments and low-quality questions discarded after expert validation.
  • README.md: This file.

🚀 How to Use SeedBench with OpenCompass

To evaluate models on SeedBench, we utilize OpenCompass. Follow the steps below to set up the environment and run the evaluation.

1. Installation

Clone the OpenCompass repository and install the necessary dependencies (including modelscope for dataset downloading).

git clone https://github.com/open-compass/opencompass opencompass
cd opencompass
pip install -e .
pip install modelscope

2. Evaluation

Set the dataset source environment variable and execute the evaluation script. The example below uses Qwen/Qwen2.5-0.5B-Instruct.

DATASET_SOURCE=ModelScope python run.py --hf-type chat \
--hf-path Qwen/Qwen2.5-0.5B-Instruct \
--datasets seedbench_gen \
--debug

📝 Notes:

  • Dataset Download: The initial run may take a few minutes to automatically download the dataset from ModelScope.
  • Local Models: You can replace Qwen/Qwen2.5-0.5B-Instruct with your absolute local path if necessary.
  • Please see Here for details.

📬 Cite

Open an issue on this repository for questions or contributions.

@inproceedings{ying-etal-2025-seedbench,
title = "{S}eed{B}ench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science",
author = "Ying, Jie and
Chen, Zihong and
Wang, Zhefan and
Jiang, Wanli and
Wang, Chenyang and
Yuan, Zhonghang and
Su, Haoyang and
Kong, Huanjun and
Yang, Fan and
Dong, Nanqing",
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.acl-long.1516/",
pages = "31395--31449",
ISBN = "979-8-89176-251-0",
abstract = "Seed science is essential for modern agriculture, directly influencing crop yields and global food security. However, challenges such as interdisciplinary complexity and high costs with limited returns hinder progress, leading to a shortage of experts and insufficient technological support. While large language models (LLMs) have shown promise across various fields, their application in seed science remains limited due to the scarcity of digital resources, complex gene-trait relationships, and the lack of standardized benchmarks. To address this gap, we introduce SeedBench{---}the first multi-task benchmark specifically designed for seed science. Developed in collaboration with domain experts, SeedBench focuses on seed breeding and simulates key aspects of modern breeding processes. We conduct a comprehensive evaluation of 26 leading LLMs, encompassing proprietary, open-source, and domain-specific fine-tuned models. Our findings not only highlight the substantial gaps between the power of LLMs and the real-world seed science problems, but also make a foundational step for research on LLMs for seed design."
}

About

[ACL 2025] SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science🌾

Topics

Resources

Stars

24 stars

Watchers

4 watching

Forks

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science

English | 简体中文

SeedBench is the first multi-task benchmark designed to evaluate large language models (LLMs) in seed science, focusing on seed breeding. This repository includes the dataset, evaluation code, and documentation to support research in this domain. Here is the usage.


🌾 Overview

SeedBench assesses LLMs across three core seed breeding stages:

  • Gene Information Retrieval
  • Gene Function and Regulation Analysis
  • Variety Breeding with Agronomic Trait Optimization

Breeding Workflow
Breeding Expert Workflow Framework

Built with domain experts, SeedBench features 2,264 expert-validated questions across 11 task types and 10 subcategories, initially targeting rice breeding. Future updates will include other crops like maize, soybean, and wheat.

🔎 Dataset Details

  • Corpus: 308,727 publications cleaned to 1.1 billion tokens; 279 segments from 113 documents.

  • Questions: 2,264 across 11 task types, bilingual (English/Chinese), expert-validated.

  • Focus: Rice breeding as a representative case.

    Types and metrics:

    Type IDQuestion TypeMetricCount
    Q&A
    QA-1Multiple ChoiceAccuracy200
    QA-2Multiple AnswerMacro-F1187
    QA-3Fill-in-the-BlankROUGE-L224
    QA-4GenerationROUGE-L242
    Summarization
    SUM-1Simple SummarizationROUGE-L225
    SUM-2Key Information ExtractionROUGE-L225
    Reading Comprehension
    RC-1Multiple ChoiceAccuracy113
    RC-2Multiple AnswerMacro-F1108
    RC-3Fill-in-the-BlankROUGE-L221
    RC-4GenerationROUGE-L240
    RC-5Subcategory ClassificationAccuracy279

    Taxonomy Distribution:

    Taxonomy Distribution

☀️ Key Results

We evaluated 26 LLMs, including proprietary, open-source, and domain-specific models. Highlights:

Performance by Question Type

  • Top Performers: DeepSeek-V3 (68.37), GPT-4 (67.88).

    Proprietary LLM Radar

    Open-Source LLM Radar

Performance by Task Types

ModelQA-1QA-2QA-3QA-4SUM-1SUM-2RC-1RC-2RC-3RC-4RC-5Avg
GPT-460.5073.8721.3536.0758.7362.89100.0096.4487.8662.2986.7467.88
DeepSeek-V372.5079.8429.2940.6348.0654.67100.0097.2287.8955.1986.7468.37
Qwen2-72B59.5075.9819.5531.6231.0863.0999.1294.2472.2051.5889.9662.54

Performance by Subcategory

ModelC1C2C3C4C5C6C7C8C9C10Avg
GPT-459.5960.5576.3261.1656.3459.3563.6764.7460.6567.6662.06
DeepSeek-V3-671B56.0362.4274.8163.1755.2358.8468.2369.0466.4668.4863.30
Qwen2-72B51.1658.1074.0759.7251.5857.7658.8561.6356.6959.1157.62
  • Top Performers: DeepSeek-V3-671B (63.30), GPT-4 (62.06).

🐝 Repository Contents

  • base_model_eval/: Used to test base models without dialogue capabilities, i.e., evaluating performance after pretraining.
  • sft_model_eval/: Used to test SFT (Supervised Fine-Tuning) models, with a total of 2,264 questions covering 10 subcategories (see Fig 2).
    • one-shot/: Organized by 11 task types (see Tab 1).
    • zero-shot/: Organized by 11 task types (see Tab 1).
  • corpus/: 279 high-quality text segments and low-quality questions discarded after expert validation.
  • README.md: This file.

🚀 How to Use SeedBench with OpenCompass

To evaluate models on SeedBench, we utilize OpenCompass. Follow the steps below to set up the environment and run the evaluation.

1. Installation

Clone the OpenCompass repository and install the necessary dependencies (including modelscope for dataset downloading).

git clone https://github.com/open-compass/opencompass opencompass
cd opencompass
pip install -e .
pip install modelscope

2. Evaluation

Set the dataset source environment variable and execute the evaluation script. The example below uses Qwen/Qwen2.5-0.5B-Instruct.

DATASET_SOURCE=ModelScope python run.py --hf-type chat \
--hf-path Qwen/Qwen2.5-0.5B-Instruct \
--datasets seedbench_gen \
--debug

📝 Notes:

  • Dataset Download: The initial run may take a few minutes to automatically download the dataset from ModelScope.
  • Local Models: You can replace Qwen/Qwen2.5-0.5B-Instruct with your absolute local path if necessary.
  • Please see Here for details.

📬 Cite

Open an issue on this repository for questions or contributions.

@inproceedings{ying-etal-2025-seedbench,
title = "{S}eed{B}ench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science",
author = "Ying, Jie and
Chen, Zihong and
Wang, Zhefan and
Jiang, Wanli and
Wang, Chenyang and
Yuan, Zhonghang and
Su, Haoyang and
Kong, Huanjun and
Yang, Fan and
Dong, Nanqing",
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.acl-long.1516/",
pages = "31395--31449",
ISBN = "979-8-89176-251-0",
abstract = "Seed science is essential for modern agriculture, directly influencing crop yields and global food security. However, challenges such as interdisciplinary complexity and high costs with limited returns hinder progress, leading to a shortage of experts and insufficient technological support. While large language models (LLMs) have shown promise across various fields, their application in seed science remains limited due to the scarcity of digital resources, complex gene-trait relationships, and the lack of standardized benchmarks. To address this gap, we introduce SeedBench{---}the first multi-task benchmark specifically designed for seed science. Developed in collaboration with domain experts, SeedBench focuses on seed breeding and simulates key aspects of modern breeding processes. We conduct a comprehensive evaluation of 26 leading LLMs, encompassing proprietary, open-source, and domain-specific fine-tuned models. Our findings not only highlight the substantial gaps between the power of LLMs and the real-world seed science problems, but also make a foundational step for research on LLMs for seed design."
}

About

[ACL 2025] SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science🌾

Topics

Resources

Stars

24 stars

Watchers

4 watching

Forks

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science

English | 简体中文

SeedBench is the first multi-task benchmark designed to evaluate large language models (LLMs) in seed science, focusing on seed breeding. This repository includes the dataset, evaluation code, and documentation to support research in this domain. Here is the usage.


🌾 Overview

SeedBench assesses LLMs across three core seed breeding stages:

  • Gene Information Retrieval
  • Gene Function and Regulation Analysis
  • Variety Breeding with Agronomic Trait Optimization

Breeding Workflow
Breeding Expert Workflow Framework

Built with domain experts, SeedBench features 2,264 expert-validated questions across 11 task types and 10 subcategories, initially targeting rice breeding. Future updates will include other crops like maize, soybean, and wheat.

🔎 Dataset Details

  • Corpus: 308,727 publications cleaned to 1.1 billion tokens; 279 segments from 113 documents.

  • Questions: 2,264 across 11 task types, bilingual (English/Chinese), expert-validated.

  • Focus: Rice breeding as a representative case.

    Types and metrics:

    Type IDQuestion TypeMetricCount
    Q&A
    QA-1Multiple ChoiceAccuracy200
    QA-2Multiple AnswerMacro-F1187
    QA-3Fill-in-the-BlankROUGE-L224
    QA-4GenerationROUGE-L242
    Summarization
    SUM-1Simple SummarizationROUGE-L225
    SUM-2Key Information ExtractionROUGE-L225
    Reading Comprehension
    RC-1Multiple ChoiceAccuracy113
    RC-2Multiple AnswerMacro-F1108
    RC-3Fill-in-the-BlankROUGE-L221
    RC-4GenerationROUGE-L240
    RC-5Subcategory ClassificationAccuracy279

    Taxonomy Distribution:

    Taxonomy Distribution

☀️ Key Results

We evaluated 26 LLMs, including proprietary, open-source, and domain-specific models. Highlights:

Performance by Question Type

  • Top Performers: DeepSeek-V3 (68.37), GPT-4 (67.88).

    Proprietary LLM Radar

    Open-Source LLM Radar

Performance by Task Types

ModelQA-1QA-2QA-3QA-4SUM-1SUM-2RC-1RC-2RC-3RC-4RC-5Avg
GPT-460.5073.8721.3536.0758.7362.89100.0096.4487.8662.2986.7467.88
DeepSeek-V372.5079.8429.2940.6348.0654.67100.0097.2287.8955.1986.7468.37
Qwen2-72B59.5075.9819.5531.6231.0863.0999.1294.2472.2051.5889.9662.54

Performance by Subcategory

ModelC1C2C3C4C5C6C7C8C9C10Avg
GPT-459.5960.5576.3261.1656.3459.3563.6764.7460.6567.6662.06
DeepSeek-V3-671B56.0362.4274.8163.1755.2358.8468.2369.0466.4668.4863.30
Qwen2-72B51.1658.1074.0759.7251.5857.7658.8561.6356.6959.1157.62
  • Top Performers: DeepSeek-V3-671B (63.30), GPT-4 (62.06).

🐝 Repository Contents

  • base_model_eval/: Used to test base models without dialogue capabilities, i.e., evaluating performance after pretraining.
  • sft_model_eval/: Used to test SFT (Supervised Fine-Tuning) models, with a total of 2,264 questions covering 10 subcategories (see Fig 2).
    • one-shot/: Organized by 11 task types (see Tab 1).
    • zero-shot/: Organized by 11 task types (see Tab 1).
  • corpus/: 279 high-quality text segments and low-quality questions discarded after expert validation.
  • README.md: This file.

🚀 How to Use SeedBench with OpenCompass

To evaluate models on SeedBench, we utilize OpenCompass. Follow the steps below to set up the environment and run the evaluation.

1. Installation

Clone the OpenCompass repository and install the necessary dependencies (including modelscope for dataset downloading).

git clone https://github.com/open-compass/opencompass opencompass
cd opencompass
pip install -e .
pip install modelscope

2. Evaluation

Set the dataset source environment variable and execute the evaluation script. The example below uses Qwen/Qwen2.5-0.5B-Instruct.

DATASET_SOURCE=ModelScope python run.py --hf-type chat \
--hf-path Qwen/Qwen2.5-0.5B-Instruct \
--datasets seedbench_gen \
--debug

📝 Notes:

  • Dataset Download: The initial run may take a few minutes to automatically download the dataset from ModelScope.
  • Local Models: You can replace Qwen/Qwen2.5-0.5B-Instruct with your absolute local path if necessary.
  • Please see Here for details.

📬 Cite

Open an issue on this repository for questions or contributions.

@inproceedings{ying-etal-2025-seedbench,
title = "{S}eed{B}ench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science",
author = "Ying, Jie and
Chen, Zihong and
Wang, Zhefan and
Jiang, Wanli and
Wang, Chenyang and
Yuan, Zhonghang and
Su, Haoyang and
Kong, Huanjun and
Yang, Fan and
Dong, Nanqing",
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.acl-long.1516/",
pages = "31395--31449",
ISBN = "979-8-89176-251-0",
abstract = "Seed science is essential for modern agriculture, directly influencing crop yields and global food security. However, challenges such as interdisciplinary complexity and high costs with limited returns hinder progress, leading to a shortage of experts and insufficient technological support. While large language models (LLMs) have shown promise across various fields, their application in seed science remains limited due to the scarcity of digital resources, complex gene-trait relationships, and the lack of standardized benchmarks. To address this gap, we introduce SeedBench{---}the first multi-task benchmark specifically designed for seed science. Developed in collaboration with domain experts, SeedBench focuses on seed breeding and simulates key aspects of modern breeding processes. We conduct a comprehensive evaluation of 26 leading LLMs, encompassing proprietary, open-source, and domain-specific fine-tuned models. Our findings not only highlight the substantial gaps between the power of LLMs and the real-world seed science problems, but also make a foundational step for research on LLMs for seed design."
}

About

[ACL 2025] SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science🌾

Topics

Resources

Stars

24 stars

Watchers

4 watching

Forks

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science

English | 简体中文

SeedBench is the first multi-task benchmark designed to evaluate large language models (LLMs) in seed science, focusing on seed breeding. This repository includes the dataset, evaluation code, and documentation to support research in this domain. Here is the usage.


🌾 Overview

SeedBench assesses LLMs across three core seed breeding stages:

  • Gene Information Retrieval
  • Gene Function and Regulation Analysis
  • Variety Breeding with Agronomic Trait Optimization

Breeding Workflow
Breeding Expert Workflow Framework

Built with domain experts, SeedBench features 2,264 expert-validated questions across 11 task types and 10 subcategories, initially targeting rice breeding. Future updates will include other crops like maize, soybean, and wheat.

🔎 Dataset Details

  • Corpus: 308,727 publications cleaned to 1.1 billion tokens; 279 segments from 113 documents.

  • Questions: 2,264 across 11 task types, bilingual (English/Chinese), expert-validated.

  • Focus: Rice breeding as a representative case.

    Types and metrics:

    Type IDQuestion TypeMetricCount
    Q&A
    QA-1Multiple ChoiceAccuracy200
    QA-2Multiple AnswerMacro-F1187
    QA-3Fill-in-the-BlankROUGE-L224
    QA-4GenerationROUGE-L242
    Summarization
    SUM-1Simple SummarizationROUGE-L225
    SUM-2Key Information ExtractionROUGE-L225
    Reading Comprehension
    RC-1Multiple ChoiceAccuracy113
    RC-2Multiple AnswerMacro-F1108
    RC-3Fill-in-the-BlankROUGE-L221
    RC-4GenerationROUGE-L240
    RC-5Subcategory ClassificationAccuracy279

    Taxonomy Distribution:

    Taxonomy Distribution

☀️ Key Results

We evaluated 26 LLMs, including proprietary, open-source, and domain-specific models. Highlights:

Performance by Question Type

  • Top Performers: DeepSeek-V3 (68.37), GPT-4 (67.88).

    Proprietary LLM Radar

    Open-Source LLM Radar

Performance by Task Types

ModelQA-1QA-2QA-3QA-4SUM-1SUM-2RC-1RC-2RC-3RC-4RC-5Avg
GPT-460.5073.8721.3536.0758.7362.89100.0096.4487.8662.2986.7467.88
DeepSeek-V372.5079.8429.2940.6348.0654.67100.0097.2287.8955.1986.7468.37
Qwen2-72B59.5075.9819.5531.6231.0863.0999.1294.2472.2051.5889.9662.54

Performance by Subcategory

ModelC1C2C3C4C5C6C7C8C9C10Avg
GPT-459.5960.5576.3261.1656.3459.3563.6764.7460.6567.6662.06
DeepSeek-V3-671B56.0362.4274.8163.1755.2358.8468.2369.0466.4668.4863.30
Qwen2-72B51.1658.1074.0759.7251.5857.7658.8561.6356.6959.1157.62
  • Top Performers: DeepSeek-V3-671B (63.30), GPT-4 (62.06).

🐝 Repository Contents

  • base_model_eval/: Used to test base models without dialogue capabilities, i.e., evaluating performance after pretraining.
  • sft_model_eval/: Used to test SFT (Supervised Fine-Tuning) models, with a total of 2,264 questions covering 10 subcategories (see Fig 2).
    • one-shot/: Organized by 11 task types (see Tab 1).
    • zero-shot/: Organized by 11 task types (see Tab 1).
  • corpus/: 279 high-quality text segments and low-quality questions discarded after expert validation.
  • README.md: This file.

🚀 How to Use SeedBench with OpenCompass

To evaluate models on SeedBench, we utilize OpenCompass. Follow the steps below to set up the environment and run the evaluation.

1. Installation

Clone the OpenCompass repository and install the necessary dependencies (including modelscope for dataset downloading).

git clone https://github.com/open-compass/opencompass opencompass
cd opencompass
pip install -e .
pip install modelscope

2. Evaluation

Set the dataset source environment variable and execute the evaluation script. The example below uses Qwen/Qwen2.5-0.5B-Instruct.

DATASET_SOURCE=ModelScope python run.py --hf-type chat \
--hf-path Qwen/Qwen2.5-0.5B-Instruct \
--datasets seedbench_gen \
--debug

📝 Notes:

  • Dataset Download: The initial run may take a few minutes to automatically download the dataset from ModelScope.
  • Local Models: You can replace Qwen/Qwen2.5-0.5B-Instruct with your absolute local path if necessary.
  • Please see Here for details.

📬 Cite

Open an issue on this repository for questions or contributions.

@inproceedings{ying-etal-2025-seedbench,
title = "{S}eed{B}ench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science",
author = "Ying, Jie and
Chen, Zihong and
Wang, Zhefan and
Jiang, Wanli and
Wang, Chenyang and
Yuan, Zhonghang and
Su, Haoyang and
Kong, Huanjun and
Yang, Fan and
Dong, Nanqing",
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.acl-long.1516/",
pages = "31395--31449",
ISBN = "979-8-89176-251-0",
abstract = "Seed science is essential for modern agriculture, directly influencing crop yields and global food security. However, challenges such as interdisciplinary complexity and high costs with limited returns hinder progress, leading to a shortage of experts and insufficient technological support. While large language models (LLMs) have shown promise across various fields, their application in seed science remains limited due to the scarcity of digital resources, complex gene-trait relationships, and the lack of standardized benchmarks. To address this gap, we introduce SeedBench{---}the first multi-task benchmark specifically designed for seed science. Developed in collaboration with domain experts, SeedBench focuses on seed breeding and simulates key aspects of modern breeding processes. We conduct a comprehensive evaluation of 26 leading LLMs, encompassing proprietary, open-source, and domain-specific fine-tuned models. Our findings not only highlight the substantial gaps between the power of LLMs and the real-world seed science problems, but also make a foundational step for research on LLMs for seed design."
}

About

[ACL 2025] SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science🌾

Topics

Resources

Stars

24 stars

Watchers

4 watching

Forks

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science

English | 简体中文

SeedBench is the first multi-task benchmark designed to evaluate large language models (LLMs) in seed science, focusing on seed breeding. This repository includes the dataset, evaluation code, and documentation to support research in this domain. Here is the usage.


🌾 Overview

SeedBench assesses LLMs across three core seed breeding stages:

  • Gene Information Retrieval
  • Gene Function and Regulation Analysis
  • Variety Breeding with Agronomic Trait Optimization

Breeding Workflow
Breeding Expert Workflow Framework

Built with domain experts, SeedBench features 2,264 expert-validated questions across 11 task types and 10 subcategories, initially targeting rice breeding. Future updates will include other crops like maize, soybean, and wheat.

🔎 Dataset Details

  • Corpus: 308,727 publications cleaned to 1.1 billion tokens; 279 segments from 113 documents.

  • Questions: 2,264 across 11 task types, bilingual (English/Chinese), expert-validated.

  • Focus: Rice breeding as a representative case.

    Types and metrics:

    Type IDQuestion TypeMetricCount
    Q&A
    QA-1Multiple ChoiceAccuracy200
    QA-2Multiple AnswerMacro-F1187
    QA-3Fill-in-the-BlankROUGE-L224
    QA-4GenerationROUGE-L242
    Summarization
    SUM-1Simple SummarizationROUGE-L225
    SUM-2Key Information ExtractionROUGE-L225
    Reading Comprehension
    RC-1Multiple ChoiceAccuracy113
    RC-2Multiple AnswerMacro-F1108
    RC-3Fill-in-the-BlankROUGE-L221
    RC-4GenerationROUGE-L240
    RC-5Subcategory ClassificationAccuracy279

    Taxonomy Distribution:

    Taxonomy Distribution

☀️ Key Results

We evaluated 26 LLMs, including proprietary, open-source, and domain-specific models. Highlights:

Performance by Question Type

  • Top Performers: DeepSeek-V3 (68.37), GPT-4 (67.88).

    Proprietary LLM Radar

    Open-Source LLM Radar

Performance by Task Types

ModelQA-1QA-2QA-3QA-4SUM-1SUM-2RC-1RC-2RC-3RC-4RC-5Avg
GPT-460.5073.8721.3536.0758.7362.89100.0096.4487.8662.2986.7467.88
DeepSeek-V372.5079.8429.2940.6348.0654.67100.0097.2287.8955.1986.7468.37
Qwen2-72B59.5075.9819.5531.6231.0863.0999.1294.2472.2051.5889.9662.54

Performance by Subcategory

ModelC1C2C3C4C5C6C7C8C9C10Avg
GPT-459.5960.5576.3261.1656.3459.3563.6764.7460.6567.6662.06
DeepSeek-V3-671B56.0362.4274.8163.1755.2358.8468.2369.0466.4668.4863.30
Qwen2-72B51.1658.1074.0759.7251.5857.7658.8561.6356.6959.1157.62
  • Top Performers: DeepSeek-V3-671B (63.30), GPT-4 (62.06).

🐝 Repository Contents

  • base_model_eval/: Used to test base models without dialogue capabilities, i.e., evaluating performance after pretraining.
  • sft_model_eval/: Used to test SFT (Supervised Fine-Tuning) models, with a total of 2,264 questions covering 10 subcategories (see Fig 2).
    • one-shot/: Organized by 11 task types (see Tab 1).
    • zero-shot/: Organized by 11 task types (see Tab 1).
  • corpus/: 279 high-quality text segments and low-quality questions discarded after expert validation.
  • README.md: This file.

🚀 How to Use SeedBench with OpenCompass

To evaluate models on SeedBench, we utilize OpenCompass. Follow the steps below to set up the environment and run the evaluation.

1. Installation

Clone the OpenCompass repository and install the necessary dependencies (including modelscope for dataset downloading).

git clone https://github.com/open-compass/opencompass opencompass
cd opencompass
pip install -e .
pip install modelscope

2. Evaluation

Set the dataset source environment variable and execute the evaluation script. The example below uses Qwen/Qwen2.5-0.5B-Instruct.

DATASET_SOURCE=ModelScope python run.py --hf-type chat \
--hf-path Qwen/Qwen2.5-0.5B-Instruct \
--datasets seedbench_gen \
--debug

📝 Notes:

  • Dataset Download: The initial run may take a few minutes to automatically download the dataset from ModelScope.
  • Local Models: You can replace Qwen/Qwen2.5-0.5B-Instruct with your absolute local path if necessary.
  • Please see Here for details.

📬 Cite

Open an issue on this repository for questions or contributions.

@inproceedings{ying-etal-2025-seedbench,
title = "{S}eed{B}ench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science",
author = "Ying, Jie and
Chen, Zihong and
Wang, Zhefan and
Jiang, Wanli and
Wang, Chenyang and
Yuan, Zhonghang and
Su, Haoyang and
Kong, Huanjun and
Yang, Fan and
Dong, Nanqing",
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.acl-long.1516/",
pages = "31395--31449",
ISBN = "979-8-89176-251-0",
abstract = "Seed science is essential for modern agriculture, directly influencing crop yields and global food security. However, challenges such as interdisciplinary complexity and high costs with limited returns hinder progress, leading to a shortage of experts and insufficient technological support. While large language models (LLMs) have shown promise across various fields, their application in seed science remains limited due to the scarcity of digital resources, complex gene-trait relationships, and the lack of standardized benchmarks. To address this gap, we introduce SeedBench{---}the first multi-task benchmark specifically designed for seed science. Developed in collaboration with domain experts, SeedBench focuses on seed breeding and simulates key aspects of modern breeding processes. We conduct a comprehensive evaluation of 26 leading LLMs, encompassing proprietary, open-source, and domain-specific fine-tuned models. Our findings not only highlight the substantial gaps between the power of LLMs and the real-world seed science problems, but also make a foundational step for research on LLMs for seed design."
}

About

[ACL 2025] SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science🌾

Topics

Resources

Stars

24 stars

Watchers

4 watching

Forks

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science

English | 简体中文

SeedBench is the first multi-task benchmark designed to evaluate large language models (LLMs) in seed science, focusing on seed breeding. This repository includes the dataset, evaluation code, and documentation to support research in this domain. Here is the usage.


🌾 Overview

SeedBench assesses LLMs across three core seed breeding stages:

  • Gene Information Retrieval
  • Gene Function and Regulation Analysis
  • Variety Breeding with Agronomic Trait Optimization

Breeding Workflow
Breeding Expert Workflow Framework

Built with domain experts, SeedBench features 2,264 expert-validated questions across 11 task types and 10 subcategories, initially targeting rice breeding. Future updates will include other crops like maize, soybean, and wheat.

🔎 Dataset Details

  • Corpus: 308,727 publications cleaned to 1.1 billion tokens; 279 segments from 113 documents.

  • Questions: 2,264 across 11 task types, bilingual (English/Chinese), expert-validated.

  • Focus: Rice breeding as a representative case.

    Types and metrics:

    Type IDQuestion TypeMetricCount
    Q&A
    QA-1Multiple ChoiceAccuracy200
    QA-2Multiple AnswerMacro-F1187
    QA-3Fill-in-the-BlankROUGE-L224
    QA-4GenerationROUGE-L242
    Summarization
    SUM-1Simple SummarizationROUGE-L225
    SUM-2Key Information ExtractionROUGE-L225
    Reading Comprehension
    RC-1Multiple ChoiceAccuracy113
    RC-2Multiple AnswerMacro-F1108
    RC-3Fill-in-the-BlankROUGE-L221
    RC-4GenerationROUGE-L240
    RC-5Subcategory ClassificationAccuracy279

    Taxonomy Distribution:

    Taxonomy Distribution

☀️ Key Results

We evaluated 26 LLMs, including proprietary, open-source, and domain-specific models. Highlights:

Performance by Question Type

  • Top Performers: DeepSeek-V3 (68.37), GPT-4 (67.88).

    Proprietary LLM Radar

    Open-Source LLM Radar

Performance by Task Types

ModelQA-1QA-2QA-3QA-4SUM-1SUM-2RC-1RC-2RC-3RC-4RC-5Avg
GPT-460.5073.8721.3536.0758.7362.89100.0096.4487.8662.2986.7467.88
DeepSeek-V372.5079.8429.2940.6348.0654.67100.0097.2287.8955.1986.7468.37
Qwen2-72B59.5075.9819.5531.6231.0863.0999.1294.2472.2051.5889.9662.54

Performance by Subcategory

ModelC1C2C3C4C5C6C7C8C9C10Avg
GPT-459.5960.5576.3261.1656.3459.3563.6764.7460.6567.6662.06
DeepSeek-V3-671B56.0362.4274.8163.1755.2358.8468.2369.0466.4668.4863.30
Qwen2-72B51.1658.1074.0759.7251.5857.7658.8561.6356.6959.1157.62
  • Top Performers: DeepSeek-V3-671B (63.30), GPT-4 (62.06).

🐝 Repository Contents

  • base_model_eval/: Used to test base models without dialogue capabilities, i.e., evaluating performance after pretraining.
  • sft_model_eval/: Used to test SFT (Supervised Fine-Tuning) models, with a total of 2,264 questions covering 10 subcategories (see Fig 2).
    • one-shot/: Organized by 11 task types (see Tab 1).
    • zero-shot/: Organized by 11 task types (see Tab 1).
  • corpus/: 279 high-quality text segments and low-quality questions discarded after expert validation.
  • README.md: This file.

🚀 How to Use SeedBench with OpenCompass

To evaluate models on SeedBench, we utilize OpenCompass. Follow the steps below to set up the environment and run the evaluation.

1. Installation

Clone the OpenCompass repository and install the necessary dependencies (including modelscope for dataset downloading).

git clone https://github.com/open-compass/opencompass opencompass
cd opencompass
pip install -e .
pip install modelscope

2. Evaluation

Set the dataset source environment variable and execute the evaluation script. The example below uses Qwen/Qwen2.5-0.5B-Instruct.

DATASET_SOURCE=ModelScope python run.py --hf-type chat \
--hf-path Qwen/Qwen2.5-0.5B-Instruct \
--datasets seedbench_gen \
--debug

📝 Notes:

  • Dataset Download: The initial run may take a few minutes to automatically download the dataset from ModelScope.
  • Local Models: You can replace Qwen/Qwen2.5-0.5B-Instruct with your absolute local path if necessary.
  • Please see Here for details.

📬 Cite

Open an issue on this repository for questions or contributions.

@inproceedings{ying-etal-2025-seedbench,
title = "{S}eed{B}ench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science",
author = "Ying, Jie and
Chen, Zihong and
Wang, Zhefan and
Jiang, Wanli and
Wang, Chenyang and
Yuan, Zhonghang and
Su, Haoyang and
Kong, Huanjun and
Yang, Fan and
Dong, Nanqing",
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.acl-long.1516/",
pages = "31395--31449",
ISBN = "979-8-89176-251-0",
abstract = "Seed science is essential for modern agriculture, directly influencing crop yields and global food security. However, challenges such as interdisciplinary complexity and high costs with limited returns hinder progress, leading to a shortage of experts and insufficient technological support. While large language models (LLMs) have shown promise across various fields, their application in seed science remains limited due to the scarcity of digital resources, complex gene-trait relationships, and the lack of standardized benchmarks. To address this gap, we introduce SeedBench{---}the first multi-task benchmark specifically designed for seed science. Developed in collaboration with domain experts, SeedBench focuses on seed breeding and simulates key aspects of modern breeding processes. We conduct a comprehensive evaluation of 26 leading LLMs, encompassing proprietary, open-source, and domain-specific fine-tuned models. Our findings not only highlight the substantial gaps between the power of LLMs and the real-world seed science problems, but also make a foundational step for research on LLMs for seed design."
}

About

[ACL 2025] SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science🌾

Topics

Resources

Stars

24 stars

Watchers

4 watching

Forks

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science

English | 简体中文

SeedBench is the first multi-task benchmark designed to evaluate large language models (LLMs) in seed science, focusing on seed breeding. This repository includes the dataset, evaluation code, and documentation to support research in this domain. Here is the usage.


🌾 Overview

SeedBench assesses LLMs across three core seed breeding stages:

  • Gene Information Retrieval
  • Gene Function and Regulation Analysis
  • Variety Breeding with Agronomic Trait Optimization

Breeding Workflow
Breeding Expert Workflow Framework

Built with domain experts, SeedBench features 2,264 expert-validated questions across 11 task types and 10 subcategories, initially targeting rice breeding. Future updates will include other crops like maize, soybean, and wheat.

🔎 Dataset Details

  • Corpus: 308,727 publications cleaned to 1.1 billion tokens; 279 segments from 113 documents.

  • Questions: 2,264 across 11 task types, bilingual (English/Chinese), expert-validated.

  • Focus: Rice breeding as a representative case.

    Types and metrics:

    Type IDQuestion TypeMetricCount
    Q&A
    QA-1Multiple ChoiceAccuracy200
    QA-2Multiple AnswerMacro-F1187
    QA-3Fill-in-the-BlankROUGE-L224
    QA-4GenerationROUGE-L242
    Summarization
    SUM-1Simple SummarizationROUGE-L225
    SUM-2Key Information ExtractionROUGE-L225
    Reading Comprehension
    RC-1Multiple ChoiceAccuracy113
    RC-2Multiple AnswerMacro-F1108
    RC-3Fill-in-the-BlankROUGE-L221
    RC-4GenerationROUGE-L240
    RC-5Subcategory ClassificationAccuracy279

    Taxonomy Distribution:

    Taxonomy Distribution

☀️ Key Results

We evaluated 26 LLMs, including proprietary, open-source, and domain-specific models. Highlights:

Performance by Question Type

  • Top Performers: DeepSeek-V3 (68.37), GPT-4 (67.88).

    Proprietary LLM Radar

    Open-Source LLM Radar

Performance by Task Types

ModelQA-1QA-2QA-3QA-4SUM-1SUM-2RC-1RC-2RC-3RC-4RC-5Avg
GPT-460.5073.8721.3536.0758.7362.89100.0096.4487.8662.2986.7467.88
DeepSeek-V372.5079.8429.2940.6348.0654.67100.0097.2287.8955.1986.7468.37
Qwen2-72B59.5075.9819.5531.6231.0863.0999.1294.2472.2051.5889.9662.54

Performance by Subcategory

ModelC1C2C3C4C5C6C7C8C9C10Avg
GPT-459.5960.5576.3261.1656.3459.3563.6764.7460.6567.6662.06
DeepSeek-V3-671B56.0362.4274.8163.1755.2358.8468.2369.0466.4668.4863.30
Qwen2-72B51.1658.1074.0759.7251.5857.7658.8561.6356.6959.1157.62
  • Top Performers: DeepSeek-V3-671B (63.30), GPT-4 (62.06).

🐝 Repository Contents

  • base_model_eval/: Used to test base models without dialogue capabilities, i.e., evaluating performance after pretraining.
  • sft_model_eval/: Used to test SFT (Supervised Fine-Tuning) models, with a total of 2,264 questions covering 10 subcategories (see Fig 2).
    • one-shot/: Organized by 11 task types (see Tab 1).
    • zero-shot/: Organized by 11 task types (see Tab 1).
  • corpus/: 279 high-quality text segments and low-quality questions discarded after expert validation.
  • README.md: This file.

🚀 How to Use SeedBench with OpenCompass

To evaluate models on SeedBench, we utilize OpenCompass. Follow the steps below to set up the environment and run the evaluation.

1. Installation

Clone the OpenCompass repository and install the necessary dependencies (including modelscope for dataset downloading).

git clone https://github.com/open-compass/opencompass opencompass
cd opencompass
pip install -e .
pip install modelscope

2. Evaluation

Set the dataset source environment variable and execute the evaluation script. The example below uses Qwen/Qwen2.5-0.5B-Instruct.

DATASET_SOURCE=ModelScope python run.py --hf-type chat \
--hf-path Qwen/Qwen2.5-0.5B-Instruct \
--datasets seedbench_gen \
--debug

📝 Notes:

  • Dataset Download: The initial run may take a few minutes to automatically download the dataset from ModelScope.
  • Local Models: You can replace Qwen/Qwen2.5-0.5B-Instruct with your absolute local path if necessary.
  • Please see Here for details.

📬 Cite

Open an issue on this repository for questions or contributions.

@inproceedings{ying-etal-2025-seedbench,
title = "{S}eed{B}ench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science",
author = "Ying, Jie and
Chen, Zihong and
Wang, Zhefan and
Jiang, Wanli and
Wang, Chenyang and
Yuan, Zhonghang and
Su, Haoyang and
Kong, Huanjun and
Yang, Fan and
Dong, Nanqing",
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.acl-long.1516/",
pages = "31395--31449",
ISBN = "979-8-89176-251-0",
abstract = "Seed science is essential for modern agriculture, directly influencing crop yields and global food security. However, challenges such as interdisciplinary complexity and high costs with limited returns hinder progress, leading to a shortage of experts and insufficient technological support. While large language models (LLMs) have shown promise across various fields, their application in seed science remains limited due to the scarcity of digital resources, complex gene-trait relationships, and the lack of standardized benchmarks. To address this gap, we introduce SeedBench{---}the first multi-task benchmark specifically designed for seed science. Developed in collaboration with domain experts, SeedBench focuses on seed breeding and simulates key aspects of modern breeding processes. We conduct a comprehensive evaluation of 26 leading LLMs, encompassing proprietary, open-source, and domain-specific fine-tuned models. Our findings not only highlight the substantial gaps between the power of LLMs and the real-world seed science problems, but also make a foundational step for research on LLMs for seed design."
}

About

[ACL 2025] SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science🌾

Topics

Resources

Stars

24 stars

Watchers

4 watching

Forks

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science

English | 简体中文

SeedBench is the first multi-task benchmark designed to evaluate large language models (LLMs) in seed science, focusing on seed breeding. This repository includes the dataset, evaluation code, and documentation to support research in this domain. Here is the usage.


🌾 Overview

SeedBench assesses LLMs across three core seed breeding stages:

  • Gene Information Retrieval
  • Gene Function and Regulation Analysis
  • Variety Breeding with Agronomic Trait Optimization

Breeding Workflow
Breeding Expert Workflow Framework

Built with domain experts, SeedBench features 2,264 expert-validated questions across 11 task types and 10 subcategories, initially targeting rice breeding. Future updates will include other crops like maize, soybean, and wheat.

🔎 Dataset Details

  • Corpus: 308,727 publications cleaned to 1.1 billion tokens; 279 segments from 113 documents.

  • Questions: 2,264 across 11 task types, bilingual (English/Chinese), expert-validated.

  • Focus: Rice breeding as a representative case.

    Types and metrics:

    Type IDQuestion TypeMetricCount
    Q&A
    QA-1Multiple ChoiceAccuracy200
    QA-2Multiple AnswerMacro-F1187
    QA-3Fill-in-the-BlankROUGE-L224
    QA-4GenerationROUGE-L242
    Summarization
    SUM-1Simple SummarizationROUGE-L225
    SUM-2Key Information ExtractionROUGE-L225
    Reading Comprehension
    RC-1Multiple ChoiceAccuracy113
    RC-2Multiple AnswerMacro-F1108
    RC-3Fill-in-the-BlankROUGE-L221
    RC-4GenerationROUGE-L240
    RC-5Subcategory ClassificationAccuracy279

    Taxonomy Distribution:

    Taxonomy Distribution

☀️ Key Results

We evaluated 26 LLMs, including proprietary, open-source, and domain-specific models. Highlights:

Performance by Question Type

  • Top Performers: DeepSeek-V3 (68.37), GPT-4 (67.88).

    Proprietary LLM Radar

    Open-Source LLM Radar

Performance by Task Types

ModelQA-1QA-2QA-3QA-4SUM-1SUM-2RC-1RC-2RC-3RC-4RC-5Avg
GPT-460.5073.8721.3536.0758.7362.89100.0096.4487.8662.2986.7467.88
DeepSeek-V372.5079.8429.2940.6348.0654.67100.0097.2287.8955.1986.7468.37
Qwen2-72B59.5075.9819.5531.6231.0863.0999.1294.2472.2051.5889.9662.54

Performance by Subcategory

ModelC1C2C3C4C5C6C7C8C9C10Avg
GPT-459.5960.5576.3261.1656.3459.3563.6764.7460.6567.6662.06
DeepSeek-V3-671B56.0362.4274.8163.1755.2358.8468.2369.0466.4668.4863.30
Qwen2-72B51.1658.1074.0759.7251.5857.7658.8561.6356.6959.1157.62
  • Top Performers: DeepSeek-V3-671B (63.30), GPT-4 (62.06).

🐝 Repository Contents

  • base_model_eval/: Used to test base models without dialogue capabilities, i.e., evaluating performance after pretraining.
  • sft_model_eval/: Used to test SFT (Supervised Fine-Tuning) models, with a total of 2,264 questions covering 10 subcategories (see Fig 2).
    • one-shot/: Organized by 11 task types (see Tab 1).
    • zero-shot/: Organized by 11 task types (see Tab 1).
  • corpus/: 279 high-quality text segments and low-quality questions discarded after expert validation.
  • README.md: This file.

🚀 How to Use SeedBench with OpenCompass

To evaluate models on SeedBench, we utilize OpenCompass. Follow the steps below to set up the environment and run the evaluation.

1. Installation

Clone the OpenCompass repository and install the necessary dependencies (including modelscope for dataset downloading).

git clone https://github.com/open-compass/opencompass opencompass
cd opencompass
pip install -e .
pip install modelscope

2. Evaluation

Set the dataset source environment variable and execute the evaluation script. The example below uses Qwen/Qwen2.5-0.5B-Instruct.

DATASET_SOURCE=ModelScope python run.py --hf-type chat \
--hf-path Qwen/Qwen2.5-0.5B-Instruct \
--datasets seedbench_gen \
--debug

📝 Notes:

  • Dataset Download: The initial run may take a few minutes to automatically download the dataset from ModelScope.
  • Local Models: You can replace Qwen/Qwen2.5-0.5B-Instruct with your absolute local path if necessary.
  • Please see Here for details.

📬 Cite

Open an issue on this repository for questions or contributions.

@inproceedings{ying-etal-2025-seedbench,
title = "{S}eed{B}ench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science",
author = "Ying, Jie and
Chen, Zihong and
Wang, Zhefan and
Jiang, Wanli and
Wang, Chenyang and
Yuan, Zhonghang and
Su, Haoyang and
Kong, Huanjun and
Yang, Fan and
Dong, Nanqing",
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.acl-long.1516/",
pages = "31395--31449",
ISBN = "979-8-89176-251-0",
abstract = "Seed science is essential for modern agriculture, directly influencing crop yields and global food security. However, challenges such as interdisciplinary complexity and high costs with limited returns hinder progress, leading to a shortage of experts and insufficient technological support. While large language models (LLMs) have shown promise across various fields, their application in seed science remains limited due to the scarcity of digital resources, complex gene-trait relationships, and the lack of standardized benchmarks. To address this gap, we introduce SeedBench{---}the first multi-task benchmark specifically designed for seed science. Developed in collaboration with domain experts, SeedBench focuses on seed breeding and simulates key aspects of modern breeding processes. We conduct a comprehensive evaluation of 26 leading LLMs, encompassing proprietary, open-source, and domain-specific fine-tuned models. Our findings not only highlight the substantial gaps between the power of LLMs and the real-world seed science problems, but also make a foundational step for research on LLMs for seed design."
}

About

[ACL 2025] SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science🌾

Topics

Resources

Stars

24 stars

Watchers

4 watching

Forks

Contributors

Languages