Skip to content

Latest commit

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Beemo

GIF Source.

Dataset Description

Beemo (Benchmark of expert-edited machine-generated outputs) is a benchmark for fine-grained machine-generated text detection, which consists of 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators for various use cases. Furthermore, each machine-generated text is edited by two state-of-the-art LLMs using several diverse editing prompts, which results in 13.1k machine-generated & LLM-edited texts. We make one of the first attempts to address more practical machine-generated text detection scenarios, where the user refines the LLM output or utilizes another LLM to make it more human-like.

Beemo is available in the HuggingFace datasets library and in this GitHub repository. Please refer to our paper for details on our benchmark creation approach, general statistics, and empirical evaluation results.

Our benchmark is named after BMO (abbreviated from "Be MOre", phonetically spelled "Beemo"), one of the main characters of Adventure Time.

  • 📊 Curated by: Toloka, Penn State University, MIT Lincoln Laboratory, and University of Oslo.
  • 🌐 Language(s): English
  • 🗞️ Paper: arxiv.org/abs/2411.04032
  • 🪪 License: MIT

🔥Updates

  • 04.02.2025: Our paper is accepted to NAACL 2025, and we release our human annotation guidelines.
  • 07.11.2024: The release of Beemo, which includes adding machine-generated & LLM-edited texts and a preprint on arXiv.
  • 17.09.2024: The initial release and evaluation of 11 detectors on Beemo.

Benchmark Design

beemo

The Beemo's creation approach involves:

  • (a) 🤖 Machine-generated Text Collection: prompting an instruction-finetuned LLM;
  • (b) 👩🏻‍🔬 Expert-based Editing: editing the LLM's output by an expert annotator;
  • (c) 🦾 LLM-based Editing: editing the LLM's output by two state-of-the-art LLMs.
🤖 Machine-generated Text Collection

The No Robots 🙅‍♂️🤖 dataset is used as the source of prompts and corresponding human-written texts across the following categories: Generation, Rewrite, Summarize, Open QA, and Closed QA. We randomly sample each prompt to generate an output with one of ten open-source instruction-finetuned LLMs using the default 🤗 HuggingFace chat templates and inference hyperparameters.

NameBaseSFT corpusLicensePaper
HuggingFaceH4/zephyr-7b-betaMistral-7B-v0.1UltraChat, UltradFeedbackMITTunstall et al. (2023)
allenai/tulu-2-7bLlama 2 7Bhuman-written and syntheticAI2 ImpACTIvison et al (2023)
allenai/tulu-2-13bLlama 2 13Bhuman-written and syntheticAI2 ImpACTIvison et al. (2023)
google/gemma-2b-itGemma 2Bhuman-written and syntheticGemma licenseGemma Team et al. (2024)
google/gemma-7b-itGemma 7Bhuman-written and syntheticGemma licenseGemma Team et al. (2024)
meta-llama/Llama-2-7b-chat-hfLlama 2 7BMisc.Llama licenseTouvron et al. (2023)
meta-llama/Llama-2-13b-chat-hfLlama 2 13BMisc.Llama licenseTouvron et al. (2023)
meta-llama/Llama-2-70b-chat-hfLlama 2 70BMisc.Llama licenseTouvron et al. (2023)
mistralai/Mistral-7B-Instruct-v0.1Mistral-7B-v0.1Misc.Apache-2.0Jiang et. al (2023)
mistralai/Mixtral-8x7B-Instruct-v0.1Mixtral 8x7BMisc.Apache-2.0Jiang et al. (2024)
meta-llama/Llama-3.1-70B-InstructLlama-3.1Misc.LlamaDubey et al. (2024)
GPT-4oGPT-4Misc.OpenAIOpenAI (2024)
Table 1: Overview of the instruction-finetuned LLMs used to create Beemo. GPT-4o and meta-llama/Llama-3.1-70B-Instruct are used only for LLM-based editing.
👩🏻‍🔬 Expert-based Editing

The machine-generated texts are edited by an in-house team of annotators, who are well experienced in refining content produced by LLMs.

🦾 LLM-based Editing

The machine-generated texts are "humanized" by GPT-4o and meta-llama/Llama-3.1-70B-Instruct using three editing prompts.

  • P1: You are given a prompt and a text generated by AI using this prompt. Your task is to edit the AI-generated text to make it sound human-like and error-free. Ensure your overall edits do not exceed 40% of the generated text and the edited text follows the user request. Output only the edited text and do not explain your edits.\n\nPrompt: {prompt}\n\nAI text: {model_output}
  • P2: You are given a pair containing two components: (1) a user prompt for an AI assistant and (2) the AI assistant’s response. Refine the AI-generated response to make it sound more natural. Vary your editing patterns and the portions of text you choose to modify, and ensure your overall edits are 20-40% of the words in the response.\n\nUser prompt: {prompt}\n\nAI-generated response: {model_output}
  • P3: Modify a machine-generated response to a given prompt to make it appear more like it was written by a native English speaker. Ensure the revised version follows the user's intent. You should just give me the revised version without any other words.\n\nPrompt: {prompt}\n\nMachine-generated response: {model_output}

License

  • The prompts and human-written texts from No Robots 🙅‍♂️🤖 are under the original dataset's license: CC-BY-NC-4.0.
  • The machine-generated texts and their LLM-edited versions are subject to the underlying instruction-finetuned LLMs' licensing terms mentioned in Table 1.
  • The expert-edited machine-generated texts are available under the MIT license, unless otherwise specified in the underlying instruction-finetuned LLMs' licensing terms.

Cite us

@article{artemova2024beemo,
title={Beemo: Benchmark of Expert-edited Machine-generated Outputs},
author={Artemova, Ekaterina and Lucas, Jason and Venkatraman, Saranya and Lee, Jooyoung and Tilga, Sergei and Uchendu, Adaku and Mikhailov, Vladislav},
journal={arXiv preprint arXiv:2411.04032},
year={2024}
}

Contact us

About

Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators.

Resources

Security policy

Stars

11 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - Toloka/beemo: Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators. · GitHub
Skip to content

Latest commit

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Beemo

GIF Source.

Dataset Description

Beemo (Benchmark of expert-edited machine-generated outputs) is a benchmark for fine-grained machine-generated text detection, which consists of 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators for various use cases. Furthermore, each machine-generated text is edited by two state-of-the-art LLMs using several diverse editing prompts, which results in 13.1k machine-generated & LLM-edited texts. We make one of the first attempts to address more practical machine-generated text detection scenarios, where the user refines the LLM output or utilizes another LLM to make it more human-like.

Beemo is available in the HuggingFace datasets library and in this GitHub repository. Please refer to our paper for details on our benchmark creation approach, general statistics, and empirical evaluation results.

Our benchmark is named after BMO (abbreviated from "Be MOre", phonetically spelled "Beemo"), one of the main characters of Adventure Time.

  • 📊 Curated by: Toloka, Penn State University, MIT Lincoln Laboratory, and University of Oslo.
  • 🌐 Language(s): English
  • 🗞️ Paper: arxiv.org/abs/2411.04032
  • 🪪 License: MIT

🔥Updates

  • 04.02.2025: Our paper is accepted to NAACL 2025, and we release our human annotation guidelines.
  • 07.11.2024: The release of Beemo, which includes adding machine-generated & LLM-edited texts and a preprint on arXiv.
  • 17.09.2024: The initial release and evaluation of 11 detectors on Beemo.

Benchmark Design

beemo

The Beemo's creation approach involves:

  • (a) 🤖 Machine-generated Text Collection: prompting an instruction-finetuned LLM;
  • (b) 👩🏻‍🔬 Expert-based Editing: editing the LLM's output by an expert annotator;
  • (c) 🦾 LLM-based Editing: editing the LLM's output by two state-of-the-art LLMs.
🤖 Machine-generated Text Collection

The No Robots 🙅‍♂️🤖 dataset is used as the source of prompts and corresponding human-written texts across the following categories: Generation, Rewrite, Summarize, Open QA, and Closed QA. We randomly sample each prompt to generate an output with one of ten open-source instruction-finetuned LLMs using the default 🤗 HuggingFace chat templates and inference hyperparameters.

NameBaseSFT corpusLicensePaper
HuggingFaceH4/zephyr-7b-betaMistral-7B-v0.1UltraChat, UltradFeedbackMITTunstall et al. (2023)
allenai/tulu-2-7bLlama 2 7Bhuman-written and syntheticAI2 ImpACTIvison et al (2023)
allenai/tulu-2-13bLlama 2 13Bhuman-written and syntheticAI2 ImpACTIvison et al. (2023)
google/gemma-2b-itGemma 2Bhuman-written and syntheticGemma licenseGemma Team et al. (2024)
google/gemma-7b-itGemma 7Bhuman-written and syntheticGemma licenseGemma Team et al. (2024)
meta-llama/Llama-2-7b-chat-hfLlama 2 7BMisc.Llama licenseTouvron et al. (2023)
meta-llama/Llama-2-13b-chat-hfLlama 2 13BMisc.Llama licenseTouvron et al. (2023)
meta-llama/Llama-2-70b-chat-hfLlama 2 70BMisc.Llama licenseTouvron et al. (2023)
mistralai/Mistral-7B-Instruct-v0.1Mistral-7B-v0.1Misc.Apache-2.0Jiang et. al (2023)
mistralai/Mixtral-8x7B-Instruct-v0.1Mixtral 8x7BMisc.Apache-2.0Jiang et al. (2024)
meta-llama/Llama-3.1-70B-InstructLlama-3.1Misc.LlamaDubey et al. (2024)
GPT-4oGPT-4Misc.OpenAIOpenAI (2024)
Table 1: Overview of the instruction-finetuned LLMs used to create Beemo. GPT-4o and meta-llama/Llama-3.1-70B-Instruct are used only for LLM-based editing.
👩🏻‍🔬 Expert-based Editing

The machine-generated texts are edited by an in-house team of annotators, who are well experienced in refining content produced by LLMs.

🦾 LLM-based Editing

The machine-generated texts are "humanized" by GPT-4o and meta-llama/Llama-3.1-70B-Instruct using three editing prompts.

  • P1: You are given a prompt and a text generated by AI using this prompt. Your task is to edit the AI-generated text to make it sound human-like and error-free. Ensure your overall edits do not exceed 40% of the generated text and the edited text follows the user request. Output only the edited text and do not explain your edits.\n\nPrompt: {prompt}\n\nAI text: {model_output}
  • P2: You are given a pair containing two components: (1) a user prompt for an AI assistant and (2) the AI assistant’s response. Refine the AI-generated response to make it sound more natural. Vary your editing patterns and the portions of text you choose to modify, and ensure your overall edits are 20-40% of the words in the response.\n\nUser prompt: {prompt}\n\nAI-generated response: {model_output}
  • P3: Modify a machine-generated response to a given prompt to make it appear more like it was written by a native English speaker. Ensure the revised version follows the user's intent. You should just give me the revised version without any other words.\n\nPrompt: {prompt}\n\nMachine-generated response: {model_output}

License

  • The prompts and human-written texts from No Robots 🙅‍♂️🤖 are under the original dataset's license: CC-BY-NC-4.0.
  • The machine-generated texts and their LLM-edited versions are subject to the underlying instruction-finetuned LLMs' licensing terms mentioned in Table 1.
  • The expert-edited machine-generated texts are available under the MIT license, unless otherwise specified in the underlying instruction-finetuned LLMs' licensing terms.

Cite us

@article{artemova2024beemo,
title={Beemo: Benchmark of Expert-edited Machine-generated Outputs},
author={Artemova, Ekaterina and Lucas, Jason and Venkatraman, Saranya and Lee, Jooyoung and Tilga, Sergei and Uchendu, Adaku and Mikhailov, Vladislav},
journal={arXiv preprint arXiv:2411.04032},
year={2024}
}

Contact us

About

Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators.

Resources

Security policy

Stars

11 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - Toloka/beemo: Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators. · GitHub
Skip to content

Latest commit

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Beemo

GIF Source.

Dataset Description

Beemo (Benchmark of expert-edited machine-generated outputs) is a benchmark for fine-grained machine-generated text detection, which consists of 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators for various use cases. Furthermore, each machine-generated text is edited by two state-of-the-art LLMs using several diverse editing prompts, which results in 13.1k machine-generated & LLM-edited texts. We make one of the first attempts to address more practical machine-generated text detection scenarios, where the user refines the LLM output or utilizes another LLM to make it more human-like.

Beemo is available in the HuggingFace datasets library and in this GitHub repository. Please refer to our paper for details on our benchmark creation approach, general statistics, and empirical evaluation results.

Our benchmark is named after BMO (abbreviated from "Be MOre", phonetically spelled "Beemo"), one of the main characters of Adventure Time.

  • 📊 Curated by: Toloka, Penn State University, MIT Lincoln Laboratory, and University of Oslo.
  • 🌐 Language(s): English
  • 🗞️ Paper: arxiv.org/abs/2411.04032
  • 🪪 License: MIT

🔥Updates

  • 04.02.2025: Our paper is accepted to NAACL 2025, and we release our human annotation guidelines.
  • 07.11.2024: The release of Beemo, which includes adding machine-generated & LLM-edited texts and a preprint on arXiv.
  • 17.09.2024: The initial release and evaluation of 11 detectors on Beemo.

Benchmark Design

beemo

The Beemo's creation approach involves:

  • (a) 🤖 Machine-generated Text Collection: prompting an instruction-finetuned LLM;
  • (b) 👩🏻‍🔬 Expert-based Editing: editing the LLM's output by an expert annotator;
  • (c) 🦾 LLM-based Editing: editing the LLM's output by two state-of-the-art LLMs.
🤖 Machine-generated Text Collection

The No Robots 🙅‍♂️🤖 dataset is used as the source of prompts and corresponding human-written texts across the following categories: Generation, Rewrite, Summarize, Open QA, and Closed QA. We randomly sample each prompt to generate an output with one of ten open-source instruction-finetuned LLMs using the default 🤗 HuggingFace chat templates and inference hyperparameters.

NameBaseSFT corpusLicensePaper
HuggingFaceH4/zephyr-7b-betaMistral-7B-v0.1UltraChat, UltradFeedbackMITTunstall et al. (2023)
allenai/tulu-2-7bLlama 2 7Bhuman-written and syntheticAI2 ImpACTIvison et al (2023)
allenai/tulu-2-13bLlama 2 13Bhuman-written and syntheticAI2 ImpACTIvison et al. (2023)
google/gemma-2b-itGemma 2Bhuman-written and syntheticGemma licenseGemma Team et al. (2024)
google/gemma-7b-itGemma 7Bhuman-written and syntheticGemma licenseGemma Team et al. (2024)
meta-llama/Llama-2-7b-chat-hfLlama 2 7BMisc.Llama licenseTouvron et al. (2023)
meta-llama/Llama-2-13b-chat-hfLlama 2 13BMisc.Llama licenseTouvron et al. (2023)
meta-llama/Llama-2-70b-chat-hfLlama 2 70BMisc.Llama licenseTouvron et al. (2023)
mistralai/Mistral-7B-Instruct-v0.1Mistral-7B-v0.1Misc.Apache-2.0Jiang et. al (2023)
mistralai/Mixtral-8x7B-Instruct-v0.1Mixtral 8x7BMisc.Apache-2.0Jiang et al. (2024)
meta-llama/Llama-3.1-70B-InstructLlama-3.1Misc.LlamaDubey et al. (2024)
GPT-4oGPT-4Misc.OpenAIOpenAI (2024)
Table 1: Overview of the instruction-finetuned LLMs used to create Beemo. GPT-4o and meta-llama/Llama-3.1-70B-Instruct are used only for LLM-based editing.
👩🏻‍🔬 Expert-based Editing

The machine-generated texts are edited by an in-house team of annotators, who are well experienced in refining content produced by LLMs.

🦾 LLM-based Editing

The machine-generated texts are "humanized" by GPT-4o and meta-llama/Llama-3.1-70B-Instruct using three editing prompts.

  • P1: You are given a prompt and a text generated by AI using this prompt. Your task is to edit the AI-generated text to make it sound human-like and error-free. Ensure your overall edits do not exceed 40% of the generated text and the edited text follows the user request. Output only the edited text and do not explain your edits.\n\nPrompt: {prompt}\n\nAI text: {model_output}
  • P2: You are given a pair containing two components: (1) a user prompt for an AI assistant and (2) the AI assistant’s response. Refine the AI-generated response to make it sound more natural. Vary your editing patterns and the portions of text you choose to modify, and ensure your overall edits are 20-40% of the words in the response.\n\nUser prompt: {prompt}\n\nAI-generated response: {model_output}
  • P3: Modify a machine-generated response to a given prompt to make it appear more like it was written by a native English speaker. Ensure the revised version follows the user's intent. You should just give me the revised version without any other words.\n\nPrompt: {prompt}\n\nMachine-generated response: {model_output}

License

  • The prompts and human-written texts from No Robots 🙅‍♂️🤖 are under the original dataset's license: CC-BY-NC-4.0.
  • The machine-generated texts and their LLM-edited versions are subject to the underlying instruction-finetuned LLMs' licensing terms mentioned in Table 1.
  • The expert-edited machine-generated texts are available under the MIT license, unless otherwise specified in the underlying instruction-finetuned LLMs' licensing terms.

Cite us

@article{artemova2024beemo,
title={Beemo: Benchmark of Expert-edited Machine-generated Outputs},
author={Artemova, Ekaterina and Lucas, Jason and Venkatraman, Saranya and Lee, Jooyoung and Tilga, Sergei and Uchendu, Adaku and Mikhailov, Vladislav},
journal={arXiv preprint arXiv:2411.04032},
year={2024}
}

Contact us

About

Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators.

Resources

Security policy

Stars

11 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - Toloka/beemo: Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators. · GitHub
Skip to content

Latest commit

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Beemo

GIF Source.

Dataset Description

Beemo (Benchmark of expert-edited machine-generated outputs) is a benchmark for fine-grained machine-generated text detection, which consists of 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators for various use cases. Furthermore, each machine-generated text is edited by two state-of-the-art LLMs using several diverse editing prompts, which results in 13.1k machine-generated & LLM-edited texts. We make one of the first attempts to address more practical machine-generated text detection scenarios, where the user refines the LLM output or utilizes another LLM to make it more human-like.

Beemo is available in the HuggingFace datasets library and in this GitHub repository. Please refer to our paper for details on our benchmark creation approach, general statistics, and empirical evaluation results.

Our benchmark is named after BMO (abbreviated from "Be MOre", phonetically spelled "Beemo"), one of the main characters of Adventure Time.

  • 📊 Curated by: Toloka, Penn State University, MIT Lincoln Laboratory, and University of Oslo.
  • 🌐 Language(s): English
  • 🗞️ Paper: arxiv.org/abs/2411.04032
  • 🪪 License: MIT

🔥Updates

  • 04.02.2025: Our paper is accepted to NAACL 2025, and we release our human annotation guidelines.
  • 07.11.2024: The release of Beemo, which includes adding machine-generated & LLM-edited texts and a preprint on arXiv.
  • 17.09.2024: The initial release and evaluation of 11 detectors on Beemo.

Benchmark Design

beemo

The Beemo's creation approach involves:

  • (a) 🤖 Machine-generated Text Collection: prompting an instruction-finetuned LLM;
  • (b) 👩🏻‍🔬 Expert-based Editing: editing the LLM's output by an expert annotator;
  • (c) 🦾 LLM-based Editing: editing the LLM's output by two state-of-the-art LLMs.
🤖 Machine-generated Text Collection

The No Robots 🙅‍♂️🤖 dataset is used as the source of prompts and corresponding human-written texts across the following categories: Generation, Rewrite, Summarize, Open QA, and Closed QA. We randomly sample each prompt to generate an output with one of ten open-source instruction-finetuned LLMs using the default 🤗 HuggingFace chat templates and inference hyperparameters.

NameBaseSFT corpusLicensePaper
HuggingFaceH4/zephyr-7b-betaMistral-7B-v0.1UltraChat, UltradFeedbackMITTunstall et al. (2023)
allenai/tulu-2-7bLlama 2 7Bhuman-written and syntheticAI2 ImpACTIvison et al (2023)
allenai/tulu-2-13bLlama 2 13Bhuman-written and syntheticAI2 ImpACTIvison et al. (2023)
google/gemma-2b-itGemma 2Bhuman-written and syntheticGemma licenseGemma Team et al. (2024)
google/gemma-7b-itGemma 7Bhuman-written and syntheticGemma licenseGemma Team et al. (2024)
meta-llama/Llama-2-7b-chat-hfLlama 2 7BMisc.Llama licenseTouvron et al. (2023)
meta-llama/Llama-2-13b-chat-hfLlama 2 13BMisc.Llama licenseTouvron et al. (2023)
meta-llama/Llama-2-70b-chat-hfLlama 2 70BMisc.Llama licenseTouvron et al. (2023)
mistralai/Mistral-7B-Instruct-v0.1Mistral-7B-v0.1Misc.Apache-2.0Jiang et. al (2023)
mistralai/Mixtral-8x7B-Instruct-v0.1Mixtral 8x7BMisc.Apache-2.0Jiang et al. (2024)
meta-llama/Llama-3.1-70B-InstructLlama-3.1Misc.LlamaDubey et al. (2024)
GPT-4oGPT-4Misc.OpenAIOpenAI (2024)
Table 1: Overview of the instruction-finetuned LLMs used to create Beemo. GPT-4o and meta-llama/Llama-3.1-70B-Instruct are used only for LLM-based editing.
👩🏻‍🔬 Expert-based Editing

The machine-generated texts are edited by an in-house team of annotators, who are well experienced in refining content produced by LLMs.

🦾 LLM-based Editing

The machine-generated texts are "humanized" by GPT-4o and meta-llama/Llama-3.1-70B-Instruct using three editing prompts.

  • P1: You are given a prompt and a text generated by AI using this prompt. Your task is to edit the AI-generated text to make it sound human-like and error-free. Ensure your overall edits do not exceed 40% of the generated text and the edited text follows the user request. Output only the edited text and do not explain your edits.\n\nPrompt: {prompt}\n\nAI text: {model_output}
  • P2: You are given a pair containing two components: (1) a user prompt for an AI assistant and (2) the AI assistant’s response. Refine the AI-generated response to make it sound more natural. Vary your editing patterns and the portions of text you choose to modify, and ensure your overall edits are 20-40% of the words in the response.\n\nUser prompt: {prompt}\n\nAI-generated response: {model_output}
  • P3: Modify a machine-generated response to a given prompt to make it appear more like it was written by a native English speaker. Ensure the revised version follows the user's intent. You should just give me the revised version without any other words.\n\nPrompt: {prompt}\n\nMachine-generated response: {model_output}

License

  • The prompts and human-written texts from No Robots 🙅‍♂️🤖 are under the original dataset's license: CC-BY-NC-4.0.
  • The machine-generated texts and their LLM-edited versions are subject to the underlying instruction-finetuned LLMs' licensing terms mentioned in Table 1.
  • The expert-edited machine-generated texts are available under the MIT license, unless otherwise specified in the underlying instruction-finetuned LLMs' licensing terms.

Cite us

@article{artemova2024beemo,
title={Beemo: Benchmark of Expert-edited Machine-generated Outputs},
author={Artemova, Ekaterina and Lucas, Jason and Venkatraman, Saranya and Lee, Jooyoung and Tilga, Sergei and Uchendu, Adaku and Mikhailov, Vladislav},
journal={arXiv preprint arXiv:2411.04032},
year={2024}
}

Contact us

About

Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators.

Resources

Security policy

Stars

11 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - Toloka/beemo: Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators. · GitHub
Skip to content

Latest commit

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Beemo

GIF Source.

Dataset Description

Beemo (Benchmark of expert-edited machine-generated outputs) is a benchmark for fine-grained machine-generated text detection, which consists of 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators for various use cases. Furthermore, each machine-generated text is edited by two state-of-the-art LLMs using several diverse editing prompts, which results in 13.1k machine-generated & LLM-edited texts. We make one of the first attempts to address more practical machine-generated text detection scenarios, where the user refines the LLM output or utilizes another LLM to make it more human-like.

Beemo is available in the HuggingFace datasets library and in this GitHub repository. Please refer to our paper for details on our benchmark creation approach, general statistics, and empirical evaluation results.

Our benchmark is named after BMO (abbreviated from "Be MOre", phonetically spelled "Beemo"), one of the main characters of Adventure Time.

  • 📊 Curated by: Toloka, Penn State University, MIT Lincoln Laboratory, and University of Oslo.
  • 🌐 Language(s): English
  • 🗞️ Paper: arxiv.org/abs/2411.04032
  • 🪪 License: MIT

🔥Updates

  • 04.02.2025: Our paper is accepted to NAACL 2025, and we release our human annotation guidelines.
  • 07.11.2024: The release of Beemo, which includes adding machine-generated & LLM-edited texts and a preprint on arXiv.
  • 17.09.2024: The initial release and evaluation of 11 detectors on Beemo.

Benchmark Design

beemo

The Beemo's creation approach involves:

  • (a) 🤖 Machine-generated Text Collection: prompting an instruction-finetuned LLM;
  • (b) 👩🏻‍🔬 Expert-based Editing: editing the LLM's output by an expert annotator;
  • (c) 🦾 LLM-based Editing: editing the LLM's output by two state-of-the-art LLMs.
🤖 Machine-generated Text Collection

The No Robots 🙅‍♂️🤖 dataset is used as the source of prompts and corresponding human-written texts across the following categories: Generation, Rewrite, Summarize, Open QA, and Closed QA. We randomly sample each prompt to generate an output with one of ten open-source instruction-finetuned LLMs using the default 🤗 HuggingFace chat templates and inference hyperparameters.

NameBaseSFT corpusLicensePaper
HuggingFaceH4/zephyr-7b-betaMistral-7B-v0.1UltraChat, UltradFeedbackMITTunstall et al. (2023)
allenai/tulu-2-7bLlama 2 7Bhuman-written and syntheticAI2 ImpACTIvison et al (2023)
allenai/tulu-2-13bLlama 2 13Bhuman-written and syntheticAI2 ImpACTIvison et al. (2023)
google/gemma-2b-itGemma 2Bhuman-written and syntheticGemma licenseGemma Team et al. (2024)
google/gemma-7b-itGemma 7Bhuman-written and syntheticGemma licenseGemma Team et al. (2024)
meta-llama/Llama-2-7b-chat-hfLlama 2 7BMisc.Llama licenseTouvron et al. (2023)
meta-llama/Llama-2-13b-chat-hfLlama 2 13BMisc.Llama licenseTouvron et al. (2023)
meta-llama/Llama-2-70b-chat-hfLlama 2 70BMisc.Llama licenseTouvron et al. (2023)
mistralai/Mistral-7B-Instruct-v0.1Mistral-7B-v0.1Misc.Apache-2.0Jiang et. al (2023)
mistralai/Mixtral-8x7B-Instruct-v0.1Mixtral 8x7BMisc.Apache-2.0Jiang et al. (2024)
meta-llama/Llama-3.1-70B-InstructLlama-3.1Misc.LlamaDubey et al. (2024)
GPT-4oGPT-4Misc.OpenAIOpenAI (2024)
Table 1: Overview of the instruction-finetuned LLMs used to create Beemo. GPT-4o and meta-llama/Llama-3.1-70B-Instruct are used only for LLM-based editing.
👩🏻‍🔬 Expert-based Editing

The machine-generated texts are edited by an in-house team of annotators, who are well experienced in refining content produced by LLMs.

🦾 LLM-based Editing

The machine-generated texts are "humanized" by GPT-4o and meta-llama/Llama-3.1-70B-Instruct using three editing prompts.

  • P1: You are given a prompt and a text generated by AI using this prompt. Your task is to edit the AI-generated text to make it sound human-like and error-free. Ensure your overall edits do not exceed 40% of the generated text and the edited text follows the user request. Output only the edited text and do not explain your edits.\n\nPrompt: {prompt}\n\nAI text: {model_output}
  • P2: You are given a pair containing two components: (1) a user prompt for an AI assistant and (2) the AI assistant’s response. Refine the AI-generated response to make it sound more natural. Vary your editing patterns and the portions of text you choose to modify, and ensure your overall edits are 20-40% of the words in the response.\n\nUser prompt: {prompt}\n\nAI-generated response: {model_output}
  • P3: Modify a machine-generated response to a given prompt to make it appear more like it was written by a native English speaker. Ensure the revised version follows the user's intent. You should just give me the revised version without any other words.\n\nPrompt: {prompt}\n\nMachine-generated response: {model_output}

License

  • The prompts and human-written texts from No Robots 🙅‍♂️🤖 are under the original dataset's license: CC-BY-NC-4.0.
  • The machine-generated texts and their LLM-edited versions are subject to the underlying instruction-finetuned LLMs' licensing terms mentioned in Table 1.
  • The expert-edited machine-generated texts are available under the MIT license, unless otherwise specified in the underlying instruction-finetuned LLMs' licensing terms.

Cite us

@article{artemova2024beemo,
title={Beemo: Benchmark of Expert-edited Machine-generated Outputs},
author={Artemova, Ekaterina and Lucas, Jason and Venkatraman, Saranya and Lee, Jooyoung and Tilga, Sergei and Uchendu, Adaku and Mikhailov, Vladislav},
journal={arXiv preprint arXiv:2411.04032},
year={2024}
}

Contact us

About

Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators.

Resources

Security policy

Stars

11 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - Toloka/beemo: Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators. · GitHub
Skip to content

Latest commit

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Beemo

GIF Source.

Dataset Description

Beemo (Benchmark of expert-edited machine-generated outputs) is a benchmark for fine-grained machine-generated text detection, which consists of 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators for various use cases. Furthermore, each machine-generated text is edited by two state-of-the-art LLMs using several diverse editing prompts, which results in 13.1k machine-generated & LLM-edited texts. We make one of the first attempts to address more practical machine-generated text detection scenarios, where the user refines the LLM output or utilizes another LLM to make it more human-like.

Beemo is available in the HuggingFace datasets library and in this GitHub repository. Please refer to our paper for details on our benchmark creation approach, general statistics, and empirical evaluation results.

Our benchmark is named after BMO (abbreviated from "Be MOre", phonetically spelled "Beemo"), one of the main characters of Adventure Time.

  • 📊 Curated by: Toloka, Penn State University, MIT Lincoln Laboratory, and University of Oslo.
  • 🌐 Language(s): English
  • 🗞️ Paper: arxiv.org/abs/2411.04032
  • 🪪 License: MIT

🔥Updates

  • 04.02.2025: Our paper is accepted to NAACL 2025, and we release our human annotation guidelines.
  • 07.11.2024: The release of Beemo, which includes adding machine-generated & LLM-edited texts and a preprint on arXiv.
  • 17.09.2024: The initial release and evaluation of 11 detectors on Beemo.

Benchmark Design

beemo

The Beemo's creation approach involves:

  • (a) 🤖 Machine-generated Text Collection: prompting an instruction-finetuned LLM;
  • (b) 👩🏻‍🔬 Expert-based Editing: editing the LLM's output by an expert annotator;
  • (c) 🦾 LLM-based Editing: editing the LLM's output by two state-of-the-art LLMs.
🤖 Machine-generated Text Collection

The No Robots 🙅‍♂️🤖 dataset is used as the source of prompts and corresponding human-written texts across the following categories: Generation, Rewrite, Summarize, Open QA, and Closed QA. We randomly sample each prompt to generate an output with one of ten open-source instruction-finetuned LLMs using the default 🤗 HuggingFace chat templates and inference hyperparameters.

NameBaseSFT corpusLicensePaper
HuggingFaceH4/zephyr-7b-betaMistral-7B-v0.1UltraChat, UltradFeedbackMITTunstall et al. (2023)
allenai/tulu-2-7bLlama 2 7Bhuman-written and syntheticAI2 ImpACTIvison et al (2023)
allenai/tulu-2-13bLlama 2 13Bhuman-written and syntheticAI2 ImpACTIvison et al. (2023)
google/gemma-2b-itGemma 2Bhuman-written and syntheticGemma licenseGemma Team et al. (2024)
google/gemma-7b-itGemma 7Bhuman-written and syntheticGemma licenseGemma Team et al. (2024)
meta-llama/Llama-2-7b-chat-hfLlama 2 7BMisc.Llama licenseTouvron et al. (2023)
meta-llama/Llama-2-13b-chat-hfLlama 2 13BMisc.Llama licenseTouvron et al. (2023)
meta-llama/Llama-2-70b-chat-hfLlama 2 70BMisc.Llama licenseTouvron et al. (2023)
mistralai/Mistral-7B-Instruct-v0.1Mistral-7B-v0.1Misc.Apache-2.0Jiang et. al (2023)
mistralai/Mixtral-8x7B-Instruct-v0.1Mixtral 8x7BMisc.Apache-2.0Jiang et al. (2024)
meta-llama/Llama-3.1-70B-InstructLlama-3.1Misc.LlamaDubey et al. (2024)
GPT-4oGPT-4Misc.OpenAIOpenAI (2024)
Table 1: Overview of the instruction-finetuned LLMs used to create Beemo. GPT-4o and meta-llama/Llama-3.1-70B-Instruct are used only for LLM-based editing.
👩🏻‍🔬 Expert-based Editing

The machine-generated texts are edited by an in-house team of annotators, who are well experienced in refining content produced by LLMs.

🦾 LLM-based Editing

The machine-generated texts are "humanized" by GPT-4o and meta-llama/Llama-3.1-70B-Instruct using three editing prompts.

  • P1: You are given a prompt and a text generated by AI using this prompt. Your task is to edit the AI-generated text to make it sound human-like and error-free. Ensure your overall edits do not exceed 40% of the generated text and the edited text follows the user request. Output only the edited text and do not explain your edits.\n\nPrompt: {prompt}\n\nAI text: {model_output}
  • P2: You are given a pair containing two components: (1) a user prompt for an AI assistant and (2) the AI assistant’s response. Refine the AI-generated response to make it sound more natural. Vary your editing patterns and the portions of text you choose to modify, and ensure your overall edits are 20-40% of the words in the response.\n\nUser prompt: {prompt}\n\nAI-generated response: {model_output}
  • P3: Modify a machine-generated response to a given prompt to make it appear more like it was written by a native English speaker. Ensure the revised version follows the user's intent. You should just give me the revised version without any other words.\n\nPrompt: {prompt}\n\nMachine-generated response: {model_output}

License

  • The prompts and human-written texts from No Robots 🙅‍♂️🤖 are under the original dataset's license: CC-BY-NC-4.0.
  • The machine-generated texts and their LLM-edited versions are subject to the underlying instruction-finetuned LLMs' licensing terms mentioned in Table 1.
  • The expert-edited machine-generated texts are available under the MIT license, unless otherwise specified in the underlying instruction-finetuned LLMs' licensing terms.

Cite us

@article{artemova2024beemo,
title={Beemo: Benchmark of Expert-edited Machine-generated Outputs},
author={Artemova, Ekaterina and Lucas, Jason and Venkatraman, Saranya and Lee, Jooyoung and Tilga, Sergei and Uchendu, Adaku and Mikhailov, Vladislav},
journal={arXiv preprint arXiv:2411.04032},
year={2024}
}

Contact us

About

Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators.

Resources

Security policy

Stars

11 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - Toloka/beemo: Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators. · GitHub
Skip to content

Latest commit

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Beemo

GIF Source.

Dataset Description

Beemo (Benchmark of expert-edited machine-generated outputs) is a benchmark for fine-grained machine-generated text detection, which consists of 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators for various use cases. Furthermore, each machine-generated text is edited by two state-of-the-art LLMs using several diverse editing prompts, which results in 13.1k machine-generated & LLM-edited texts. We make one of the first attempts to address more practical machine-generated text detection scenarios, where the user refines the LLM output or utilizes another LLM to make it more human-like.

Beemo is available in the HuggingFace datasets library and in this GitHub repository. Please refer to our paper for details on our benchmark creation approach, general statistics, and empirical evaluation results.

Our benchmark is named after BMO (abbreviated from "Be MOre", phonetically spelled "Beemo"), one of the main characters of Adventure Time.

  • 📊 Curated by: Toloka, Penn State University, MIT Lincoln Laboratory, and University of Oslo.
  • 🌐 Language(s): English
  • 🗞️ Paper: arxiv.org/abs/2411.04032
  • 🪪 License: MIT

🔥Updates

  • 04.02.2025: Our paper is accepted to NAACL 2025, and we release our human annotation guidelines.
  • 07.11.2024: The release of Beemo, which includes adding machine-generated & LLM-edited texts and a preprint on arXiv.
  • 17.09.2024: The initial release and evaluation of 11 detectors on Beemo.

Benchmark Design

beemo

The Beemo's creation approach involves:

  • (a) 🤖 Machine-generated Text Collection: prompting an instruction-finetuned LLM;
  • (b) 👩🏻‍🔬 Expert-based Editing: editing the LLM's output by an expert annotator;
  • (c) 🦾 LLM-based Editing: editing the LLM's output by two state-of-the-art LLMs.
🤖 Machine-generated Text Collection

The No Robots 🙅‍♂️🤖 dataset is used as the source of prompts and corresponding human-written texts across the following categories: Generation, Rewrite, Summarize, Open QA, and Closed QA. We randomly sample each prompt to generate an output with one of ten open-source instruction-finetuned LLMs using the default 🤗 HuggingFace chat templates and inference hyperparameters.

NameBaseSFT corpusLicensePaper
HuggingFaceH4/zephyr-7b-betaMistral-7B-v0.1UltraChat, UltradFeedbackMITTunstall et al. (2023)
allenai/tulu-2-7bLlama 2 7Bhuman-written and syntheticAI2 ImpACTIvison et al (2023)
allenai/tulu-2-13bLlama 2 13Bhuman-written and syntheticAI2 ImpACTIvison et al. (2023)
google/gemma-2b-itGemma 2Bhuman-written and syntheticGemma licenseGemma Team et al. (2024)
google/gemma-7b-itGemma 7Bhuman-written and syntheticGemma licenseGemma Team et al. (2024)
meta-llama/Llama-2-7b-chat-hfLlama 2 7BMisc.Llama licenseTouvron et al. (2023)
meta-llama/Llama-2-13b-chat-hfLlama 2 13BMisc.Llama licenseTouvron et al. (2023)
meta-llama/Llama-2-70b-chat-hfLlama 2 70BMisc.Llama licenseTouvron et al. (2023)
mistralai/Mistral-7B-Instruct-v0.1Mistral-7B-v0.1Misc.Apache-2.0Jiang et. al (2023)
mistralai/Mixtral-8x7B-Instruct-v0.1Mixtral 8x7BMisc.Apache-2.0Jiang et al. (2024)
meta-llama/Llama-3.1-70B-InstructLlama-3.1Misc.LlamaDubey et al. (2024)
GPT-4oGPT-4Misc.OpenAIOpenAI (2024)
Table 1: Overview of the instruction-finetuned LLMs used to create Beemo. GPT-4o and meta-llama/Llama-3.1-70B-Instruct are used only for LLM-based editing.
👩🏻‍🔬 Expert-based Editing

The machine-generated texts are edited by an in-house team of annotators, who are well experienced in refining content produced by LLMs.

🦾 LLM-based Editing

The machine-generated texts are "humanized" by GPT-4o and meta-llama/Llama-3.1-70B-Instruct using three editing prompts.

  • P1: You are given a prompt and a text generated by AI using this prompt. Your task is to edit the AI-generated text to make it sound human-like and error-free. Ensure your overall edits do not exceed 40% of the generated text and the edited text follows the user request. Output only the edited text and do not explain your edits.\n\nPrompt: {prompt}\n\nAI text: {model_output}
  • P2: You are given a pair containing two components: (1) a user prompt for an AI assistant and (2) the AI assistant’s response. Refine the AI-generated response to make it sound more natural. Vary your editing patterns and the portions of text you choose to modify, and ensure your overall edits are 20-40% of the words in the response.\n\nUser prompt: {prompt}\n\nAI-generated response: {model_output}
  • P3: Modify a machine-generated response to a given prompt to make it appear more like it was written by a native English speaker. Ensure the revised version follows the user's intent. You should just give me the revised version without any other words.\n\nPrompt: {prompt}\n\nMachine-generated response: {model_output}

License

  • The prompts and human-written texts from No Robots 🙅‍♂️🤖 are under the original dataset's license: CC-BY-NC-4.0.
  • The machine-generated texts and their LLM-edited versions are subject to the underlying instruction-finetuned LLMs' licensing terms mentioned in Table 1.
  • The expert-edited machine-generated texts are available under the MIT license, unless otherwise specified in the underlying instruction-finetuned LLMs' licensing terms.

Cite us

@article{artemova2024beemo,
title={Beemo: Benchmark of Expert-edited Machine-generated Outputs},
author={Artemova, Ekaterina and Lucas, Jason and Venkatraman, Saranya and Lee, Jooyoung and Tilga, Sergei and Uchendu, Adaku and Mikhailov, Vladislav},
journal={arXiv preprint arXiv:2411.04032},
year={2024}
}

Contact us

About

Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators.

Resources

Security policy

Stars

11 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - Toloka/beemo: Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators. · GitHub
Skip to content

Latest commit

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Beemo

GIF Source.

Dataset Description

Beemo (Benchmark of expert-edited machine-generated outputs) is a benchmark for fine-grained machine-generated text detection, which consists of 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators for various use cases. Furthermore, each machine-generated text is edited by two state-of-the-art LLMs using several diverse editing prompts, which results in 13.1k machine-generated & LLM-edited texts. We make one of the first attempts to address more practical machine-generated text detection scenarios, where the user refines the LLM output or utilizes another LLM to make it more human-like.

Beemo is available in the HuggingFace datasets library and in this GitHub repository. Please refer to our paper for details on our benchmark creation approach, general statistics, and empirical evaluation results.

Our benchmark is named after BMO (abbreviated from "Be MOre", phonetically spelled "Beemo"), one of the main characters of Adventure Time.

  • 📊 Curated by: Toloka, Penn State University, MIT Lincoln Laboratory, and University of Oslo.
  • 🌐 Language(s): English
  • 🗞️ Paper: arxiv.org/abs/2411.04032
  • 🪪 License: MIT

🔥Updates

  • 04.02.2025: Our paper is accepted to NAACL 2025, and we release our human annotation guidelines.
  • 07.11.2024: The release of Beemo, which includes adding machine-generated & LLM-edited texts and a preprint on arXiv.
  • 17.09.2024: The initial release and evaluation of 11 detectors on Beemo.

Benchmark Design

beemo

The Beemo's creation approach involves:

  • (a) 🤖 Machine-generated Text Collection: prompting an instruction-finetuned LLM;
  • (b) 👩🏻‍🔬 Expert-based Editing: editing the LLM's output by an expert annotator;
  • (c) 🦾 LLM-based Editing: editing the LLM's output by two state-of-the-art LLMs.
🤖 Machine-generated Text Collection

The No Robots 🙅‍♂️🤖 dataset is used as the source of prompts and corresponding human-written texts across the following categories: Generation, Rewrite, Summarize, Open QA, and Closed QA. We randomly sample each prompt to generate an output with one of ten open-source instruction-finetuned LLMs using the default 🤗 HuggingFace chat templates and inference hyperparameters.

NameBaseSFT corpusLicensePaper
HuggingFaceH4/zephyr-7b-betaMistral-7B-v0.1UltraChat, UltradFeedbackMITTunstall et al. (2023)
allenai/tulu-2-7bLlama 2 7Bhuman-written and syntheticAI2 ImpACTIvison et al (2023)
allenai/tulu-2-13bLlama 2 13Bhuman-written and syntheticAI2 ImpACTIvison et al. (2023)
google/gemma-2b-itGemma 2Bhuman-written and syntheticGemma licenseGemma Team et al. (2024)
google/gemma-7b-itGemma 7Bhuman-written and syntheticGemma licenseGemma Team et al. (2024)
meta-llama/Llama-2-7b-chat-hfLlama 2 7BMisc.Llama licenseTouvron et al. (2023)
meta-llama/Llama-2-13b-chat-hfLlama 2 13BMisc.Llama licenseTouvron et al. (2023)
meta-llama/Llama-2-70b-chat-hfLlama 2 70BMisc.Llama licenseTouvron et al. (2023)
mistralai/Mistral-7B-Instruct-v0.1Mistral-7B-v0.1Misc.Apache-2.0Jiang et. al (2023)
mistralai/Mixtral-8x7B-Instruct-v0.1Mixtral 8x7BMisc.Apache-2.0Jiang et al. (2024)
meta-llama/Llama-3.1-70B-InstructLlama-3.1Misc.LlamaDubey et al. (2024)
GPT-4oGPT-4Misc.OpenAIOpenAI (2024)
Table 1: Overview of the instruction-finetuned LLMs used to create Beemo. GPT-4o and meta-llama/Llama-3.1-70B-Instruct are used only for LLM-based editing.
👩🏻‍🔬 Expert-based Editing

The machine-generated texts are edited by an in-house team of annotators, who are well experienced in refining content produced by LLMs.

🦾 LLM-based Editing

The machine-generated texts are "humanized" by GPT-4o and meta-llama/Llama-3.1-70B-Instruct using three editing prompts.

  • P1: You are given a prompt and a text generated by AI using this prompt. Your task is to edit the AI-generated text to make it sound human-like and error-free. Ensure your overall edits do not exceed 40% of the generated text and the edited text follows the user request. Output only the edited text and do not explain your edits.\n\nPrompt: {prompt}\n\nAI text: {model_output}
  • P2: You are given a pair containing two components: (1) a user prompt for an AI assistant and (2) the AI assistant’s response. Refine the AI-generated response to make it sound more natural. Vary your editing patterns and the portions of text you choose to modify, and ensure your overall edits are 20-40% of the words in the response.\n\nUser prompt: {prompt}\n\nAI-generated response: {model_output}
  • P3: Modify a machine-generated response to a given prompt to make it appear more like it was written by a native English speaker. Ensure the revised version follows the user's intent. You should just give me the revised version without any other words.\n\nPrompt: {prompt}\n\nMachine-generated response: {model_output}

License

  • The prompts and human-written texts from No Robots 🙅‍♂️🤖 are under the original dataset's license: CC-BY-NC-4.0.
  • The machine-generated texts and their LLM-edited versions are subject to the underlying instruction-finetuned LLMs' licensing terms mentioned in Table 1.
  • The expert-edited machine-generated texts are available under the MIT license, unless otherwise specified in the underlying instruction-finetuned LLMs' licensing terms.

Cite us

@article{artemova2024beemo,
title={Beemo: Benchmark of Expert-edited Machine-generated Outputs},
author={Artemova, Ekaterina and Lucas, Jason and Venkatraman, Saranya and Lee, Jooyoung and Tilga, Sergei and Uchendu, Adaku and Mikhailov, Vladislav},
journal={arXiv preprint arXiv:2411.04032},
year={2024}
}

Contact us

About

Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators.

Resources

Security policy

Stars

11 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors