Latest commit

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

News

logo

PaperHugging Face DatasetsLicense: CC BY-NC 4.0


Data and evaluation code for the paper MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation).

@inproceedings{tedeschi-navigli-2022-multinerd,
title = "{M}ulti{NERD}: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)",
author = "Tedeschi, Simone and Navigli, Roberto",
booktitle = "Findings of the Association for Computational Linguistics: NAACL 2022",
month = jul,
year = "2022",
address = "Seattle, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.findings-naacl.60",
doi = "10.18653/v1/2022.findings-naacl.60",
pages = "801--812",
abstract = "Named Entity Recognition (NER) is the task of identifying named entities in texts and classifying them through specific semantic categories, a process which is crucial for a wide range of NLP applications. Current datasets for NER focus mainly on coarse-grained entity types, tend to consider a single textual genre and to cover a narrow set of languages, thus limiting the general applicability of NER systems.In this work, we design a new methodology for automatically producing NER annotations, and address the aforementioned limitations by introducing a novel dataset that covers 10 languages, 15 NER categories and 2 textual genres.We also introduce a manually-annotated test set, and extensively evaluate the quality of our novel dataset on both this new test set and standard benchmarks for NER.In addition, in our dataset, we include: i) disambiguation information to enable the development of multilingual entity linking systems, and ii) image URLs to encourage the creation of multimodal systems.We release our dataset at https://github.com/Babelscape/multinerd.",
}

Please consider citing our work if you use data and/or code from this repository.

In a nutshell, MultiNERD is the first language-agnostic methodology for automatically creating multilingual, multi-genre and fine-grained annotations for Named Entity Recognition and Entity Disambiguation. Specifically, it can be seen an extension of the combination of two prior works from our research group that are WikiNEuRal, from which we took inspiration for the state-of-the-art silver-data creation methodology, and NER4EL, from which we took the fine-grained classes and inspiration for the entity linking part.

The produced dataset covers:

  • 10 languages: Chinese, Dutch, English, French, German, Italian, Polish, Portuguese, Russian and Spanish;
  • 15 NER categories: Person (PER), Location (LOC), Organization (ORG}), Animal (ANIM), Biological entity (BIO), Celestial Body (CEL), Disease (DIS), Event (EVE), Food (FOOD), Instrument (INST), Media (MEDIA), Plant (PLANT), Mythological entity (MYTH), Time (TIME) and Vehicle (VEHI);
  • 2 textual genres, i.e. Wikipedia and WikiNews articles;
  • 3 different reference KBs for entity disambiguation, i.e. BabelNet, WikiData and Wikipedia.

Additionally, we included image URLs to encourage the creation of multimodal systems.

Finally, MultiNERD shows consistent improvements of up to against state-of-the-art alternative data production methods on common benchmarks for NER while covering a broader set of NER categories (15 vs. 4):

comparison


Data

Dataset VersionSentencesTokensPERORGLOCANIMBIOCELDISEVEFOODINSTMEDIAMYTHPLANTTIMEVEHIOTHER
MultiNERD EN164.1K3.6M75.8K33.7K78.5K15.5K0.2K2.8K11.2K3.2K11.0K0.4K7.5K0.7K9.5K3.2K0.5K3.1M
MultiNERD ES173.2K4.3M70.9K20.6K90.2K10.5K0.3K2.4K8.6K6.8K7.8K0.6K8.0K1.6K7.6K45.3K0.3K3.8M
MultiNERD NL171.7K3.0M56.9K21.4K78.7K34.4K0.1K2.1K6.1K4.7K5.6K0.2K3.8K1.3K6.3K31.0K0.4K2.7M
MultiNERD DE156.8K2.7M79.2K31.2K72.8K11.5K0.1K1.4K5.2K4.0K3.6K0.1K2.8K0.8K7.8K3.3K0.5K2.4M
MultiNERD RU129.0K2.3M43.4K21.5K75.2K7.3K0.1K1.2K1.9K2.8K3.2K1.1K11.3K0.6K4.8K22.8K0.5K2.0M
MultiNERD IT181.9K4.7M75.3K19.3K98.5K8.8K0.1K5.2K6.5K5.8K5.8K0.8K8.6K1.8K5.1K71.2K0.6K4.2M
MultiNERD FR176.2K4.3M89.6K28.2K90.9K11.4K0.1K2.3K3.1K7.4K3.2K0.7K8.0K2.0K4.4K27.4K0.6K3.8M
MultiNERD PL195.0K3.0M66.5K29.2K100.0K19.7K0.1K3.3K6.5K6.7K3.3K0.6K4.9K1.3K6.6K44.1K0.7K2.5M
MultiNERD PT177.6K3.9M54.0K13.2K124.8K14.7K0.1K4.2K6.8K5.9K5.4K0.6K9.1K1.6K9.2K48.6K0.3K3.4M
MultiNERD ZH195.3K5.8M68.3K20.8K49.6K26.1K0.4K0.8K0.1K5.1K1.9K1.1K55.9K1.8K6.1K0.4K0.3K3.4M

We remark that the datasets are automatically created, and, therefore, they may contain errors. Specifically, the highest-quality classes (in terms of both precision and number of the annotations, according to Table 3 and Figure 1 in the paper are PER, ORG, LOC, CEL, DIS, EVE and MEDIA, while others can be often very noisy due to the ambiguity of their instances.


License

MultiNERD is licensed under the CC BY-SA-NC 4.0 license. The text of the license can be found here.

We underline that the source from which the raw sentences have been extracted are Wikipedia (wikipedia.org) and Wikinews wikinews.org and the NER annotations have been produced by Babelscape.


Acknowledgments

We gratefully acknowledge the support of the ERC Consolidator Grant MOUSSE No. 726487 under the European Union’s Horizon2020 research and innovation programme (http://mousse-project.org/).

About

Repository for the paper "MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)" (NAACL 2022).

Topics

Resources

Stars

43 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

News

logo

PaperHugging Face DatasetsLicense: CC BY-NC 4.0


Data and evaluation code for the paper MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation).

@inproceedings{tedeschi-navigli-2022-multinerd,
title = "{M}ulti{NERD}: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)",
author = "Tedeschi, Simone and Navigli, Roberto",
booktitle = "Findings of the Association for Computational Linguistics: NAACL 2022",
month = jul,
year = "2022",
address = "Seattle, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.findings-naacl.60",
doi = "10.18653/v1/2022.findings-naacl.60",
pages = "801--812",
abstract = "Named Entity Recognition (NER) is the task of identifying named entities in texts and classifying them through specific semantic categories, a process which is crucial for a wide range of NLP applications. Current datasets for NER focus mainly on coarse-grained entity types, tend to consider a single textual genre and to cover a narrow set of languages, thus limiting the general applicability of NER systems.In this work, we design a new methodology for automatically producing NER annotations, and address the aforementioned limitations by introducing a novel dataset that covers 10 languages, 15 NER categories and 2 textual genres.We also introduce a manually-annotated test set, and extensively evaluate the quality of our novel dataset on both this new test set and standard benchmarks for NER.In addition, in our dataset, we include: i) disambiguation information to enable the development of multilingual entity linking systems, and ii) image URLs to encourage the creation of multimodal systems.We release our dataset at https://github.com/Babelscape/multinerd.",
}

Please consider citing our work if you use data and/or code from this repository.

In a nutshell, MultiNERD is the first language-agnostic methodology for automatically creating multilingual, multi-genre and fine-grained annotations for Named Entity Recognition and Entity Disambiguation. Specifically, it can be seen an extension of the combination of two prior works from our research group that are WikiNEuRal, from which we took inspiration for the state-of-the-art silver-data creation methodology, and NER4EL, from which we took the fine-grained classes and inspiration for the entity linking part.

The produced dataset covers:

  • 10 languages: Chinese, Dutch, English, French, German, Italian, Polish, Portuguese, Russian and Spanish;
  • 15 NER categories: Person (PER), Location (LOC), Organization (ORG}), Animal (ANIM), Biological entity (BIO), Celestial Body (CEL), Disease (DIS), Event (EVE), Food (FOOD), Instrument (INST), Media (MEDIA), Plant (PLANT), Mythological entity (MYTH), Time (TIME) and Vehicle (VEHI);
  • 2 textual genres, i.e. Wikipedia and WikiNews articles;
  • 3 different reference KBs for entity disambiguation, i.e. BabelNet, WikiData and Wikipedia.

Additionally, we included image URLs to encourage the creation of multimodal systems.

Finally, MultiNERD shows consistent improvements of up to against state-of-the-art alternative data production methods on common benchmarks for NER while covering a broader set of NER categories (15 vs. 4):

comparison


Data

Dataset VersionSentencesTokensPERORGLOCANIMBIOCELDISEVEFOODINSTMEDIAMYTHPLANTTIMEVEHIOTHER
MultiNERD EN164.1K3.6M75.8K33.7K78.5K15.5K0.2K2.8K11.2K3.2K11.0K0.4K7.5K0.7K9.5K3.2K0.5K3.1M
MultiNERD ES173.2K4.3M70.9K20.6K90.2K10.5K0.3K2.4K8.6K6.8K7.8K0.6K8.0K1.6K7.6K45.3K0.3K3.8M
MultiNERD NL171.7K3.0M56.9K21.4K78.7K34.4K0.1K2.1K6.1K4.7K5.6K0.2K3.8K1.3K6.3K31.0K0.4K2.7M
MultiNERD DE156.8K2.7M79.2K31.2K72.8K11.5K0.1K1.4K5.2K4.0K3.6K0.1K2.8K0.8K7.8K3.3K0.5K2.4M
MultiNERD RU129.0K2.3M43.4K21.5K75.2K7.3K0.1K1.2K1.9K2.8K3.2K1.1K11.3K0.6K4.8K22.8K0.5K2.0M
MultiNERD IT181.9K4.7M75.3K19.3K98.5K8.8K0.1K5.2K6.5K5.8K5.8K0.8K8.6K1.8K5.1K71.2K0.6K4.2M
MultiNERD FR176.2K4.3M89.6K28.2K90.9K11.4K0.1K2.3K3.1K7.4K3.2K0.7K8.0K2.0K4.4K27.4K0.6K3.8M
MultiNERD PL195.0K3.0M66.5K29.2K100.0K19.7K0.1K3.3K6.5K6.7K3.3K0.6K4.9K1.3K6.6K44.1K0.7K2.5M
MultiNERD PT177.6K3.9M54.0K13.2K124.8K14.7K0.1K4.2K6.8K5.9K5.4K0.6K9.1K1.6K9.2K48.6K0.3K3.4M
MultiNERD ZH195.3K5.8M68.3K20.8K49.6K26.1K0.4K0.8K0.1K5.1K1.9K1.1K55.9K1.8K6.1K0.4K0.3K3.4M

We remark that the datasets are automatically created, and, therefore, they may contain errors. Specifically, the highest-quality classes (in terms of both precision and number of the annotations, according to Table 3 and Figure 1 in the paper are PER, ORG, LOC, CEL, DIS, EVE and MEDIA, while others can be often very noisy due to the ambiguity of their instances.


License

MultiNERD is licensed under the CC BY-SA-NC 4.0 license. The text of the license can be found here.

We underline that the source from which the raw sentences have been extracted are Wikipedia (wikipedia.org) and Wikinews wikinews.org and the NER annotations have been produced by Babelscape.


Acknowledgments

We gratefully acknowledge the support of the ERC Consolidator Grant MOUSSE No. 726487 under the European Union’s Horizon2020 research and innovation programme (http://mousse-project.org/).

About

Repository for the paper "MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)" (NAACL 2022).

Topics

Resources

Stars

43 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

News

logo

PaperHugging Face DatasetsLicense: CC BY-NC 4.0


Data and evaluation code for the paper MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation).

@inproceedings{tedeschi-navigli-2022-multinerd,
title = "{M}ulti{NERD}: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)",
author = "Tedeschi, Simone and Navigli, Roberto",
booktitle = "Findings of the Association for Computational Linguistics: NAACL 2022",
month = jul,
year = "2022",
address = "Seattle, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.findings-naacl.60",
doi = "10.18653/v1/2022.findings-naacl.60",
pages = "801--812",
abstract = "Named Entity Recognition (NER) is the task of identifying named entities in texts and classifying them through specific semantic categories, a process which is crucial for a wide range of NLP applications. Current datasets for NER focus mainly on coarse-grained entity types, tend to consider a single textual genre and to cover a narrow set of languages, thus limiting the general applicability of NER systems.In this work, we design a new methodology for automatically producing NER annotations, and address the aforementioned limitations by introducing a novel dataset that covers 10 languages, 15 NER categories and 2 textual genres.We also introduce a manually-annotated test set, and extensively evaluate the quality of our novel dataset on both this new test set and standard benchmarks for NER.In addition, in our dataset, we include: i) disambiguation information to enable the development of multilingual entity linking systems, and ii) image URLs to encourage the creation of multimodal systems.We release our dataset at https://github.com/Babelscape/multinerd.",
}

Please consider citing our work if you use data and/or code from this repository.

In a nutshell, MultiNERD is the first language-agnostic methodology for automatically creating multilingual, multi-genre and fine-grained annotations for Named Entity Recognition and Entity Disambiguation. Specifically, it can be seen an extension of the combination of two prior works from our research group that are WikiNEuRal, from which we took inspiration for the state-of-the-art silver-data creation methodology, and NER4EL, from which we took the fine-grained classes and inspiration for the entity linking part.

The produced dataset covers:

  • 10 languages: Chinese, Dutch, English, French, German, Italian, Polish, Portuguese, Russian and Spanish;
  • 15 NER categories: Person (PER), Location (LOC), Organization (ORG}), Animal (ANIM), Biological entity (BIO), Celestial Body (CEL), Disease (DIS), Event (EVE), Food (FOOD), Instrument (INST), Media (MEDIA), Plant (PLANT), Mythological entity (MYTH), Time (TIME) and Vehicle (VEHI);
  • 2 textual genres, i.e. Wikipedia and WikiNews articles;
  • 3 different reference KBs for entity disambiguation, i.e. BabelNet, WikiData and Wikipedia.

Additionally, we included image URLs to encourage the creation of multimodal systems.

Finally, MultiNERD shows consistent improvements of up to against state-of-the-art alternative data production methods on common benchmarks for NER while covering a broader set of NER categories (15 vs. 4):

comparison


Data

Dataset VersionSentencesTokensPERORGLOCANIMBIOCELDISEVEFOODINSTMEDIAMYTHPLANTTIMEVEHIOTHER
MultiNERD EN164.1K3.6M75.8K33.7K78.5K15.5K0.2K2.8K11.2K3.2K11.0K0.4K7.5K0.7K9.5K3.2K0.5K3.1M
MultiNERD ES173.2K4.3M70.9K20.6K90.2K10.5K0.3K2.4K8.6K6.8K7.8K0.6K8.0K1.6K7.6K45.3K0.3K3.8M
MultiNERD NL171.7K3.0M56.9K21.4K78.7K34.4K0.1K2.1K6.1K4.7K5.6K0.2K3.8K1.3K6.3K31.0K0.4K2.7M
MultiNERD DE156.8K2.7M79.2K31.2K72.8K11.5K0.1K1.4K5.2K4.0K3.6K0.1K2.8K0.8K7.8K3.3K0.5K2.4M
MultiNERD RU129.0K2.3M43.4K21.5K75.2K7.3K0.1K1.2K1.9K2.8K3.2K1.1K11.3K0.6K4.8K22.8K0.5K2.0M
MultiNERD IT181.9K4.7M75.3K19.3K98.5K8.8K0.1K5.2K6.5K5.8K5.8K0.8K8.6K1.8K5.1K71.2K0.6K4.2M
MultiNERD FR176.2K4.3M89.6K28.2K90.9K11.4K0.1K2.3K3.1K7.4K3.2K0.7K8.0K2.0K4.4K27.4K0.6K3.8M
MultiNERD PL195.0K3.0M66.5K29.2K100.0K19.7K0.1K3.3K6.5K6.7K3.3K0.6K4.9K1.3K6.6K44.1K0.7K2.5M
MultiNERD PT177.6K3.9M54.0K13.2K124.8K14.7K0.1K4.2K6.8K5.9K5.4K0.6K9.1K1.6K9.2K48.6K0.3K3.4M
MultiNERD ZH195.3K5.8M68.3K20.8K49.6K26.1K0.4K0.8K0.1K5.1K1.9K1.1K55.9K1.8K6.1K0.4K0.3K3.4M

We remark that the datasets are automatically created, and, therefore, they may contain errors. Specifically, the highest-quality classes (in terms of both precision and number of the annotations, according to Table 3 and Figure 1 in the paper are PER, ORG, LOC, CEL, DIS, EVE and MEDIA, while others can be often very noisy due to the ambiguity of their instances.


License

MultiNERD is licensed under the CC BY-SA-NC 4.0 license. The text of the license can be found here.

We underline that the source from which the raw sentences have been extracted are Wikipedia (wikipedia.org) and Wikinews wikinews.org and the NER annotations have been produced by Babelscape.


Acknowledgments

We gratefully acknowledge the support of the ERC Consolidator Grant MOUSSE No. 726487 under the European Union’s Horizon2020 research and innovation programme (http://mousse-project.org/).

About

Repository for the paper "MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)" (NAACL 2022).

Topics

Resources

Stars

43 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

News

logo

PaperHugging Face DatasetsLicense: CC BY-NC 4.0


Data and evaluation code for the paper MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation).

@inproceedings{tedeschi-navigli-2022-multinerd,
title = "{M}ulti{NERD}: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)",
author = "Tedeschi, Simone and Navigli, Roberto",
booktitle = "Findings of the Association for Computational Linguistics: NAACL 2022",
month = jul,
year = "2022",
address = "Seattle, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.findings-naacl.60",
doi = "10.18653/v1/2022.findings-naacl.60",
pages = "801--812",
abstract = "Named Entity Recognition (NER) is the task of identifying named entities in texts and classifying them through specific semantic categories, a process which is crucial for a wide range of NLP applications. Current datasets for NER focus mainly on coarse-grained entity types, tend to consider a single textual genre and to cover a narrow set of languages, thus limiting the general applicability of NER systems.In this work, we design a new methodology for automatically producing NER annotations, and address the aforementioned limitations by introducing a novel dataset that covers 10 languages, 15 NER categories and 2 textual genres.We also introduce a manually-annotated test set, and extensively evaluate the quality of our novel dataset on both this new test set and standard benchmarks for NER.In addition, in our dataset, we include: i) disambiguation information to enable the development of multilingual entity linking systems, and ii) image URLs to encourage the creation of multimodal systems.We release our dataset at https://github.com/Babelscape/multinerd.",
}

Please consider citing our work if you use data and/or code from this repository.

In a nutshell, MultiNERD is the first language-agnostic methodology for automatically creating multilingual, multi-genre and fine-grained annotations for Named Entity Recognition and Entity Disambiguation. Specifically, it can be seen an extension of the combination of two prior works from our research group that are WikiNEuRal, from which we took inspiration for the state-of-the-art silver-data creation methodology, and NER4EL, from which we took the fine-grained classes and inspiration for the entity linking part.

The produced dataset covers:

  • 10 languages: Chinese, Dutch, English, French, German, Italian, Polish, Portuguese, Russian and Spanish;
  • 15 NER categories: Person (PER), Location (LOC), Organization (ORG}), Animal (ANIM), Biological entity (BIO), Celestial Body (CEL), Disease (DIS), Event (EVE), Food (FOOD), Instrument (INST), Media (MEDIA), Plant (PLANT), Mythological entity (MYTH), Time (TIME) and Vehicle (VEHI);
  • 2 textual genres, i.e. Wikipedia and WikiNews articles;
  • 3 different reference KBs for entity disambiguation, i.e. BabelNet, WikiData and Wikipedia.

Additionally, we included image URLs to encourage the creation of multimodal systems.

Finally, MultiNERD shows consistent improvements of up to against state-of-the-art alternative data production methods on common benchmarks for NER while covering a broader set of NER categories (15 vs. 4):

comparison


Data

Dataset VersionSentencesTokensPERORGLOCANIMBIOCELDISEVEFOODINSTMEDIAMYTHPLANTTIMEVEHIOTHER
MultiNERD EN164.1K3.6M75.8K33.7K78.5K15.5K0.2K2.8K11.2K3.2K11.0K0.4K7.5K0.7K9.5K3.2K0.5K3.1M
MultiNERD ES173.2K4.3M70.9K20.6K90.2K10.5K0.3K2.4K8.6K6.8K7.8K0.6K8.0K1.6K7.6K45.3K0.3K3.8M
MultiNERD NL171.7K3.0M56.9K21.4K78.7K34.4K0.1K2.1K6.1K4.7K5.6K0.2K3.8K1.3K6.3K31.0K0.4K2.7M
MultiNERD DE156.8K2.7M79.2K31.2K72.8K11.5K0.1K1.4K5.2K4.0K3.6K0.1K2.8K0.8K7.8K3.3K0.5K2.4M
MultiNERD RU129.0K2.3M43.4K21.5K75.2K7.3K0.1K1.2K1.9K2.8K3.2K1.1K11.3K0.6K4.8K22.8K0.5K2.0M
MultiNERD IT181.9K4.7M75.3K19.3K98.5K8.8K0.1K5.2K6.5K5.8K5.8K0.8K8.6K1.8K5.1K71.2K0.6K4.2M
MultiNERD FR176.2K4.3M89.6K28.2K90.9K11.4K0.1K2.3K3.1K7.4K3.2K0.7K8.0K2.0K4.4K27.4K0.6K3.8M
MultiNERD PL195.0K3.0M66.5K29.2K100.0K19.7K0.1K3.3K6.5K6.7K3.3K0.6K4.9K1.3K6.6K44.1K0.7K2.5M
MultiNERD PT177.6K3.9M54.0K13.2K124.8K14.7K0.1K4.2K6.8K5.9K5.4K0.6K9.1K1.6K9.2K48.6K0.3K3.4M
MultiNERD ZH195.3K5.8M68.3K20.8K49.6K26.1K0.4K0.8K0.1K5.1K1.9K1.1K55.9K1.8K6.1K0.4K0.3K3.4M

We remark that the datasets are automatically created, and, therefore, they may contain errors. Specifically, the highest-quality classes (in terms of both precision and number of the annotations, according to Table 3 and Figure 1 in the paper are PER, ORG, LOC, CEL, DIS, EVE and MEDIA, while others can be often very noisy due to the ambiguity of their instances.


License

MultiNERD is licensed under the CC BY-SA-NC 4.0 license. The text of the license can be found here.

We underline that the source from which the raw sentences have been extracted are Wikipedia (wikipedia.org) and Wikinews wikinews.org and the NER annotations have been produced by Babelscape.


Acknowledgments

We gratefully acknowledge the support of the ERC Consolidator Grant MOUSSE No. 726487 under the European Union’s Horizon2020 research and innovation programme (http://mousse-project.org/).

About

Repository for the paper "MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)" (NAACL 2022).

Topics

Resources

Stars

43 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

News

logo

PaperHugging Face DatasetsLicense: CC BY-NC 4.0


Data and evaluation code for the paper MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation).

@inproceedings{tedeschi-navigli-2022-multinerd,
title = "{M}ulti{NERD}: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)",
author = "Tedeschi, Simone and Navigli, Roberto",
booktitle = "Findings of the Association for Computational Linguistics: NAACL 2022",
month = jul,
year = "2022",
address = "Seattle, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.findings-naacl.60",
doi = "10.18653/v1/2022.findings-naacl.60",
pages = "801--812",
abstract = "Named Entity Recognition (NER) is the task of identifying named entities in texts and classifying them through specific semantic categories, a process which is crucial for a wide range of NLP applications. Current datasets for NER focus mainly on coarse-grained entity types, tend to consider a single textual genre and to cover a narrow set of languages, thus limiting the general applicability of NER systems.In this work, we design a new methodology for automatically producing NER annotations, and address the aforementioned limitations by introducing a novel dataset that covers 10 languages, 15 NER categories and 2 textual genres.We also introduce a manually-annotated test set, and extensively evaluate the quality of our novel dataset on both this new test set and standard benchmarks for NER.In addition, in our dataset, we include: i) disambiguation information to enable the development of multilingual entity linking systems, and ii) image URLs to encourage the creation of multimodal systems.We release our dataset at https://github.com/Babelscape/multinerd.",
}

Please consider citing our work if you use data and/or code from this repository.

In a nutshell, MultiNERD is the first language-agnostic methodology for automatically creating multilingual, multi-genre and fine-grained annotations for Named Entity Recognition and Entity Disambiguation. Specifically, it can be seen an extension of the combination of two prior works from our research group that are WikiNEuRal, from which we took inspiration for the state-of-the-art silver-data creation methodology, and NER4EL, from which we took the fine-grained classes and inspiration for the entity linking part.

The produced dataset covers:

  • 10 languages: Chinese, Dutch, English, French, German, Italian, Polish, Portuguese, Russian and Spanish;
  • 15 NER categories: Person (PER), Location (LOC), Organization (ORG}), Animal (ANIM), Biological entity (BIO), Celestial Body (CEL), Disease (DIS), Event (EVE), Food (FOOD), Instrument (INST), Media (MEDIA), Plant (PLANT), Mythological entity (MYTH), Time (TIME) and Vehicle (VEHI);
  • 2 textual genres, i.e. Wikipedia and WikiNews articles;
  • 3 different reference KBs for entity disambiguation, i.e. BabelNet, WikiData and Wikipedia.

Additionally, we included image URLs to encourage the creation of multimodal systems.

Finally, MultiNERD shows consistent improvements of up to against state-of-the-art alternative data production methods on common benchmarks for NER while covering a broader set of NER categories (15 vs. 4):

comparison


Data

Dataset VersionSentencesTokensPERORGLOCANIMBIOCELDISEVEFOODINSTMEDIAMYTHPLANTTIMEVEHIOTHER
MultiNERD EN164.1K3.6M75.8K33.7K78.5K15.5K0.2K2.8K11.2K3.2K11.0K0.4K7.5K0.7K9.5K3.2K0.5K3.1M
MultiNERD ES173.2K4.3M70.9K20.6K90.2K10.5K0.3K2.4K8.6K6.8K7.8K0.6K8.0K1.6K7.6K45.3K0.3K3.8M
MultiNERD NL171.7K3.0M56.9K21.4K78.7K34.4K0.1K2.1K6.1K4.7K5.6K0.2K3.8K1.3K6.3K31.0K0.4K2.7M
MultiNERD DE156.8K2.7M79.2K31.2K72.8K11.5K0.1K1.4K5.2K4.0K3.6K0.1K2.8K0.8K7.8K3.3K0.5K2.4M
MultiNERD RU129.0K2.3M43.4K21.5K75.2K7.3K0.1K1.2K1.9K2.8K3.2K1.1K11.3K0.6K4.8K22.8K0.5K2.0M
MultiNERD IT181.9K4.7M75.3K19.3K98.5K8.8K0.1K5.2K6.5K5.8K5.8K0.8K8.6K1.8K5.1K71.2K0.6K4.2M
MultiNERD FR176.2K4.3M89.6K28.2K90.9K11.4K0.1K2.3K3.1K7.4K3.2K0.7K8.0K2.0K4.4K27.4K0.6K3.8M
MultiNERD PL195.0K3.0M66.5K29.2K100.0K19.7K0.1K3.3K6.5K6.7K3.3K0.6K4.9K1.3K6.6K44.1K0.7K2.5M
MultiNERD PT177.6K3.9M54.0K13.2K124.8K14.7K0.1K4.2K6.8K5.9K5.4K0.6K9.1K1.6K9.2K48.6K0.3K3.4M
MultiNERD ZH195.3K5.8M68.3K20.8K49.6K26.1K0.4K0.8K0.1K5.1K1.9K1.1K55.9K1.8K6.1K0.4K0.3K3.4M

We remark that the datasets are automatically created, and, therefore, they may contain errors. Specifically, the highest-quality classes (in terms of both precision and number of the annotations, according to Table 3 and Figure 1 in the paper are PER, ORG, LOC, CEL, DIS, EVE and MEDIA, while others can be often very noisy due to the ambiguity of their instances.


License

MultiNERD is licensed under the CC BY-SA-NC 4.0 license. The text of the license can be found here.

We underline that the source from which the raw sentences have been extracted are Wikipedia (wikipedia.org) and Wikinews wikinews.org and the NER annotations have been produced by Babelscape.


Acknowledgments

We gratefully acknowledge the support of the ERC Consolidator Grant MOUSSE No. 726487 under the European Union’s Horizon2020 research and innovation programme (http://mousse-project.org/).

About

Repository for the paper "MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)" (NAACL 2022).

Topics

Resources

Stars

43 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

News

logo

PaperHugging Face DatasetsLicense: CC BY-NC 4.0


Data and evaluation code for the paper MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation).

@inproceedings{tedeschi-navigli-2022-multinerd,
title = "{M}ulti{NERD}: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)",
author = "Tedeschi, Simone and Navigli, Roberto",
booktitle = "Findings of the Association for Computational Linguistics: NAACL 2022",
month = jul,
year = "2022",
address = "Seattle, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.findings-naacl.60",
doi = "10.18653/v1/2022.findings-naacl.60",
pages = "801--812",
abstract = "Named Entity Recognition (NER) is the task of identifying named entities in texts and classifying them through specific semantic categories, a process which is crucial for a wide range of NLP applications. Current datasets for NER focus mainly on coarse-grained entity types, tend to consider a single textual genre and to cover a narrow set of languages, thus limiting the general applicability of NER systems.In this work, we design a new methodology for automatically producing NER annotations, and address the aforementioned limitations by introducing a novel dataset that covers 10 languages, 15 NER categories and 2 textual genres.We also introduce a manually-annotated test set, and extensively evaluate the quality of our novel dataset on both this new test set and standard benchmarks for NER.In addition, in our dataset, we include: i) disambiguation information to enable the development of multilingual entity linking systems, and ii) image URLs to encourage the creation of multimodal systems.We release our dataset at https://github.com/Babelscape/multinerd.",
}

Please consider citing our work if you use data and/or code from this repository.

In a nutshell, MultiNERD is the first language-agnostic methodology for automatically creating multilingual, multi-genre and fine-grained annotations for Named Entity Recognition and Entity Disambiguation. Specifically, it can be seen an extension of the combination of two prior works from our research group that are WikiNEuRal, from which we took inspiration for the state-of-the-art silver-data creation methodology, and NER4EL, from which we took the fine-grained classes and inspiration for the entity linking part.

The produced dataset covers:

  • 10 languages: Chinese, Dutch, English, French, German, Italian, Polish, Portuguese, Russian and Spanish;
  • 15 NER categories: Person (PER), Location (LOC), Organization (ORG}), Animal (ANIM), Biological entity (BIO), Celestial Body (CEL), Disease (DIS), Event (EVE), Food (FOOD), Instrument (INST), Media (MEDIA), Plant (PLANT), Mythological entity (MYTH), Time (TIME) and Vehicle (VEHI);
  • 2 textual genres, i.e. Wikipedia and WikiNews articles;
  • 3 different reference KBs for entity disambiguation, i.e. BabelNet, WikiData and Wikipedia.

Additionally, we included image URLs to encourage the creation of multimodal systems.

Finally, MultiNERD shows consistent improvements of up to against state-of-the-art alternative data production methods on common benchmarks for NER while covering a broader set of NER categories (15 vs. 4):

comparison


Data

Dataset VersionSentencesTokensPERORGLOCANIMBIOCELDISEVEFOODINSTMEDIAMYTHPLANTTIMEVEHIOTHER
MultiNERD EN164.1K3.6M75.8K33.7K78.5K15.5K0.2K2.8K11.2K3.2K11.0K0.4K7.5K0.7K9.5K3.2K0.5K3.1M
MultiNERD ES173.2K4.3M70.9K20.6K90.2K10.5K0.3K2.4K8.6K6.8K7.8K0.6K8.0K1.6K7.6K45.3K0.3K3.8M
MultiNERD NL171.7K3.0M56.9K21.4K78.7K34.4K0.1K2.1K6.1K4.7K5.6K0.2K3.8K1.3K6.3K31.0K0.4K2.7M
MultiNERD DE156.8K2.7M79.2K31.2K72.8K11.5K0.1K1.4K5.2K4.0K3.6K0.1K2.8K0.8K7.8K3.3K0.5K2.4M
MultiNERD RU129.0K2.3M43.4K21.5K75.2K7.3K0.1K1.2K1.9K2.8K3.2K1.1K11.3K0.6K4.8K22.8K0.5K2.0M
MultiNERD IT181.9K4.7M75.3K19.3K98.5K8.8K0.1K5.2K6.5K5.8K5.8K0.8K8.6K1.8K5.1K71.2K0.6K4.2M
MultiNERD FR176.2K4.3M89.6K28.2K90.9K11.4K0.1K2.3K3.1K7.4K3.2K0.7K8.0K2.0K4.4K27.4K0.6K3.8M
MultiNERD PL195.0K3.0M66.5K29.2K100.0K19.7K0.1K3.3K6.5K6.7K3.3K0.6K4.9K1.3K6.6K44.1K0.7K2.5M
MultiNERD PT177.6K3.9M54.0K13.2K124.8K14.7K0.1K4.2K6.8K5.9K5.4K0.6K9.1K1.6K9.2K48.6K0.3K3.4M
MultiNERD ZH195.3K5.8M68.3K20.8K49.6K26.1K0.4K0.8K0.1K5.1K1.9K1.1K55.9K1.8K6.1K0.4K0.3K3.4M

We remark that the datasets are automatically created, and, therefore, they may contain errors. Specifically, the highest-quality classes (in terms of both precision and number of the annotations, according to Table 3 and Figure 1 in the paper are PER, ORG, LOC, CEL, DIS, EVE and MEDIA, while others can be often very noisy due to the ambiguity of their instances.


License

MultiNERD is licensed under the CC BY-SA-NC 4.0 license. The text of the license can be found here.

We underline that the source from which the raw sentences have been extracted are Wikipedia (wikipedia.org) and Wikinews wikinews.org and the NER annotations have been produced by Babelscape.


Acknowledgments

We gratefully acknowledge the support of the ERC Consolidator Grant MOUSSE No. 726487 under the European Union’s Horizon2020 research and innovation programme (http://mousse-project.org/).

About

Repository for the paper "MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)" (NAACL 2022).

Topics

Resources

Stars

43 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

News

logo

PaperHugging Face DatasetsLicense: CC BY-NC 4.0


Data and evaluation code for the paper MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation).

@inproceedings{tedeschi-navigli-2022-multinerd,
title = "{M}ulti{NERD}: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)",
author = "Tedeschi, Simone and Navigli, Roberto",
booktitle = "Findings of the Association for Computational Linguistics: NAACL 2022",
month = jul,
year = "2022",
address = "Seattle, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.findings-naacl.60",
doi = "10.18653/v1/2022.findings-naacl.60",
pages = "801--812",
abstract = "Named Entity Recognition (NER) is the task of identifying named entities in texts and classifying them through specific semantic categories, a process which is crucial for a wide range of NLP applications. Current datasets for NER focus mainly on coarse-grained entity types, tend to consider a single textual genre and to cover a narrow set of languages, thus limiting the general applicability of NER systems.In this work, we design a new methodology for automatically producing NER annotations, and address the aforementioned limitations by introducing a novel dataset that covers 10 languages, 15 NER categories and 2 textual genres.We also introduce a manually-annotated test set, and extensively evaluate the quality of our novel dataset on both this new test set and standard benchmarks for NER.In addition, in our dataset, we include: i) disambiguation information to enable the development of multilingual entity linking systems, and ii) image URLs to encourage the creation of multimodal systems.We release our dataset at https://github.com/Babelscape/multinerd.",
}

Please consider citing our work if you use data and/or code from this repository.

In a nutshell, MultiNERD is the first language-agnostic methodology for automatically creating multilingual, multi-genre and fine-grained annotations for Named Entity Recognition and Entity Disambiguation. Specifically, it can be seen an extension of the combination of two prior works from our research group that are WikiNEuRal, from which we took inspiration for the state-of-the-art silver-data creation methodology, and NER4EL, from which we took the fine-grained classes and inspiration for the entity linking part.

The produced dataset covers:

  • 10 languages: Chinese, Dutch, English, French, German, Italian, Polish, Portuguese, Russian and Spanish;
  • 15 NER categories: Person (PER), Location (LOC), Organization (ORG}), Animal (ANIM), Biological entity (BIO), Celestial Body (CEL), Disease (DIS), Event (EVE), Food (FOOD), Instrument (INST), Media (MEDIA), Plant (PLANT), Mythological entity (MYTH), Time (TIME) and Vehicle (VEHI);
  • 2 textual genres, i.e. Wikipedia and WikiNews articles;
  • 3 different reference KBs for entity disambiguation, i.e. BabelNet, WikiData and Wikipedia.

Additionally, we included image URLs to encourage the creation of multimodal systems.

Finally, MultiNERD shows consistent improvements of up to against state-of-the-art alternative data production methods on common benchmarks for NER while covering a broader set of NER categories (15 vs. 4):

comparison


Data

Dataset VersionSentencesTokensPERORGLOCANIMBIOCELDISEVEFOODINSTMEDIAMYTHPLANTTIMEVEHIOTHER
MultiNERD EN164.1K3.6M75.8K33.7K78.5K15.5K0.2K2.8K11.2K3.2K11.0K0.4K7.5K0.7K9.5K3.2K0.5K3.1M
MultiNERD ES173.2K4.3M70.9K20.6K90.2K10.5K0.3K2.4K8.6K6.8K7.8K0.6K8.0K1.6K7.6K45.3K0.3K3.8M
MultiNERD NL171.7K3.0M56.9K21.4K78.7K34.4K0.1K2.1K6.1K4.7K5.6K0.2K3.8K1.3K6.3K31.0K0.4K2.7M
MultiNERD DE156.8K2.7M79.2K31.2K72.8K11.5K0.1K1.4K5.2K4.0K3.6K0.1K2.8K0.8K7.8K3.3K0.5K2.4M
MultiNERD RU129.0K2.3M43.4K21.5K75.2K7.3K0.1K1.2K1.9K2.8K3.2K1.1K11.3K0.6K4.8K22.8K0.5K2.0M
MultiNERD IT181.9K4.7M75.3K19.3K98.5K8.8K0.1K5.2K6.5K5.8K5.8K0.8K8.6K1.8K5.1K71.2K0.6K4.2M
MultiNERD FR176.2K4.3M89.6K28.2K90.9K11.4K0.1K2.3K3.1K7.4K3.2K0.7K8.0K2.0K4.4K27.4K0.6K3.8M
MultiNERD PL195.0K3.0M66.5K29.2K100.0K19.7K0.1K3.3K6.5K6.7K3.3K0.6K4.9K1.3K6.6K44.1K0.7K2.5M
MultiNERD PT177.6K3.9M54.0K13.2K124.8K14.7K0.1K4.2K6.8K5.9K5.4K0.6K9.1K1.6K9.2K48.6K0.3K3.4M
MultiNERD ZH195.3K5.8M68.3K20.8K49.6K26.1K0.4K0.8K0.1K5.1K1.9K1.1K55.9K1.8K6.1K0.4K0.3K3.4M

We remark that the datasets are automatically created, and, therefore, they may contain errors. Specifically, the highest-quality classes (in terms of both precision and number of the annotations, according to Table 3 and Figure 1 in the paper are PER, ORG, LOC, CEL, DIS, EVE and MEDIA, while others can be often very noisy due to the ambiguity of their instances.


License

MultiNERD is licensed under the CC BY-SA-NC 4.0 license. The text of the license can be found here.

We underline that the source from which the raw sentences have been extracted are Wikipedia (wikipedia.org) and Wikinews wikinews.org and the NER annotations have been produced by Babelscape.


Acknowledgments

We gratefully acknowledge the support of the ERC Consolidator Grant MOUSSE No. 726487 under the European Union’s Horizon2020 research and innovation programme (http://mousse-project.org/).

About

Repository for the paper "MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)" (NAACL 2022).

Topics

Resources

Stars

43 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

News

logo

PaperHugging Face DatasetsLicense: CC BY-NC 4.0


Data and evaluation code for the paper MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation).

@inproceedings{tedeschi-navigli-2022-multinerd,
title = "{M}ulti{NERD}: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)",
author = "Tedeschi, Simone and Navigli, Roberto",
booktitle = "Findings of the Association for Computational Linguistics: NAACL 2022",
month = jul,
year = "2022",
address = "Seattle, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.findings-naacl.60",
doi = "10.18653/v1/2022.findings-naacl.60",
pages = "801--812",
abstract = "Named Entity Recognition (NER) is the task of identifying named entities in texts and classifying them through specific semantic categories, a process which is crucial for a wide range of NLP applications. Current datasets for NER focus mainly on coarse-grained entity types, tend to consider a single textual genre and to cover a narrow set of languages, thus limiting the general applicability of NER systems.In this work, we design a new methodology for automatically producing NER annotations, and address the aforementioned limitations by introducing a novel dataset that covers 10 languages, 15 NER categories and 2 textual genres.We also introduce a manually-annotated test set, and extensively evaluate the quality of our novel dataset on both this new test set and standard benchmarks for NER.In addition, in our dataset, we include: i) disambiguation information to enable the development of multilingual entity linking systems, and ii) image URLs to encourage the creation of multimodal systems.We release our dataset at https://github.com/Babelscape/multinerd.",
}

Please consider citing our work if you use data and/or code from this repository.

In a nutshell, MultiNERD is the first language-agnostic methodology for automatically creating multilingual, multi-genre and fine-grained annotations for Named Entity Recognition and Entity Disambiguation. Specifically, it can be seen an extension of the combination of two prior works from our research group that are WikiNEuRal, from which we took inspiration for the state-of-the-art silver-data creation methodology, and NER4EL, from which we took the fine-grained classes and inspiration for the entity linking part.

The produced dataset covers:

  • 10 languages: Chinese, Dutch, English, French, German, Italian, Polish, Portuguese, Russian and Spanish;
  • 15 NER categories: Person (PER), Location (LOC), Organization (ORG}), Animal (ANIM), Biological entity (BIO), Celestial Body (CEL), Disease (DIS), Event (EVE), Food (FOOD), Instrument (INST), Media (MEDIA), Plant (PLANT), Mythological entity (MYTH), Time (TIME) and Vehicle (VEHI);
  • 2 textual genres, i.e. Wikipedia and WikiNews articles;
  • 3 different reference KBs for entity disambiguation, i.e. BabelNet, WikiData and Wikipedia.

Additionally, we included image URLs to encourage the creation of multimodal systems.

Finally, MultiNERD shows consistent improvements of up to against state-of-the-art alternative data production methods on common benchmarks for NER while covering a broader set of NER categories (15 vs. 4):

comparison


Data

Dataset VersionSentencesTokensPERORGLOCANIMBIOCELDISEVEFOODINSTMEDIAMYTHPLANTTIMEVEHIOTHER
MultiNERD EN164.1K3.6M75.8K33.7K78.5K15.5K0.2K2.8K11.2K3.2K11.0K0.4K7.5K0.7K9.5K3.2K0.5K3.1M
MultiNERD ES173.2K4.3M70.9K20.6K90.2K10.5K0.3K2.4K8.6K6.8K7.8K0.6K8.0K1.6K7.6K45.3K0.3K3.8M
MultiNERD NL171.7K3.0M56.9K21.4K78.7K34.4K0.1K2.1K6.1K4.7K5.6K0.2K3.8K1.3K6.3K31.0K0.4K2.7M
MultiNERD DE156.8K2.7M79.2K31.2K72.8K11.5K0.1K1.4K5.2K4.0K3.6K0.1K2.8K0.8K7.8K3.3K0.5K2.4M
MultiNERD RU129.0K2.3M43.4K21.5K75.2K7.3K0.1K1.2K1.9K2.8K3.2K1.1K11.3K0.6K4.8K22.8K0.5K2.0M
MultiNERD IT181.9K4.7M75.3K19.3K98.5K8.8K0.1K5.2K6.5K5.8K5.8K0.8K8.6K1.8K5.1K71.2K0.6K4.2M
MultiNERD FR176.2K4.3M89.6K28.2K90.9K11.4K0.1K2.3K3.1K7.4K3.2K0.7K8.0K2.0K4.4K27.4K0.6K3.8M
MultiNERD PL195.0K3.0M66.5K29.2K100.0K19.7K0.1K3.3K6.5K6.7K3.3K0.6K4.9K1.3K6.6K44.1K0.7K2.5M
MultiNERD PT177.6K3.9M54.0K13.2K124.8K14.7K0.1K4.2K6.8K5.9K5.4K0.6K9.1K1.6K9.2K48.6K0.3K3.4M
MultiNERD ZH195.3K5.8M68.3K20.8K49.6K26.1K0.4K0.8K0.1K5.1K1.9K1.1K55.9K1.8K6.1K0.4K0.3K3.4M

We remark that the datasets are automatically created, and, therefore, they may contain errors. Specifically, the highest-quality classes (in terms of both precision and number of the annotations, according to Table 3 and Figure 1 in the paper are PER, ORG, LOC, CEL, DIS, EVE and MEDIA, while others can be often very noisy due to the ambiguity of their instances.


License

MultiNERD is licensed under the CC BY-SA-NC 4.0 license. The text of the license can be found here.

We underline that the source from which the raw sentences have been extracted are Wikipedia (wikipedia.org) and Wikinews wikinews.org and the NER annotations have been produced by Babelscape.


Acknowledgments

We gratefully acknowledge the support of the ERC Consolidator Grant MOUSSE No. 726487 under the European Union’s Horizon2020 research and innovation programme (http://mousse-project.org/).

About

Repository for the paper "MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)" (NAACL 2022).

Topics

Resources

Stars

43 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages