Latest commit

History

137 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SIGMORPHON 2022 Shared Task on Morpheme Segmentation

Morphemes (prefixes, suffixes, root words) are linguistic descriptions, defined as the smallest meaningful unit of words. Our proposed shared task is morpheme segmentation that converts a text into a sequence of morphemes. In order to prepare a dataset for this task, we integrated all basic types of morphological databases (including UniMorph (Kirov et al., 2018b; McCarthy et al., 2020) – inflectional morphology; MorphyNet (Batsuren et al., 2021) – derivational morphology; Universal Dependencies (Nivre et al., 2017) and ten editions of Wiktionary – compound morphology and root words). In the future, we expect the NLP community will benefit a lot by innovating subword-based tokenization with this task. This shared task has two parts:

Please join our Google Group to stay up to date. Click here to register for the task!

Please open the issues if you have any questions.

Two subtasks will be scored separately. Participant teams may submit as many systems as they want to as many subtasks as they want.

Part 1: Word-level Morpheme Segmentation

At the word level, participants will be asked to segment a given word into a sequence of morphemes. Input words contains all types of word forms: root words, derived words, inflected words, and compound words.

Data

Training and development data are UTF-8-encoded tab-separated values files. Each example occupies a single line and consists of input word, the corresponding morpheme sequence, and the corresponding morphological category. The following shows three lines of English data:

inaccuracies in @@accurate @@cy @@s 110
dictionary dictionary 000
screwdriver screw @@drive @@er 011

Note: The third column as the morphological category is an optional feature that can only be used to oversample or undersample training data.

First example is a derived word with prefix (in-) and suffixes (-cy and -s), and second example is a root word. Third example is a compound word. In the test datasets, we will provide only first column of data as input words.

Languages

Development languages are:

  1. ces: Czech
  2. eng: English
  3. fra: French
  4. hun: Hungarian
  5. spa: Spanish
  6. ita: Italian
  7. lat: Latin
  8. rus: Russian

Surprise language is:

  1. mon: Mongolian

Data Statistics

word classEnglishSpanishHungarianFrenchItalianRussianCzechLatinMongolian
100126544502229410662105192253455221760-8319917266
0102031021844924923679834109272970-02201
101137904581011894783171909-035
00010193815843695213619210372921-503381604
0115381821654506140328-00
110106570346862323119126196237104481409-07855
0011699024833201684431259-05
1113059343542791861582658-00
total words5773748845149260983827975537347842143868288232918966

Word category description

For some of the development languages, we are providing the word categories so that participant can deal with imbalanced situation of morphological categories.

word classDescriptionEnglish example (input ==> output)
100Inflection onlyplayed ==> play @@ed
010Derivation onlyplayer ==> play @@er
101Inflection and Compoundwheelbands ==> wheel @@band @@s
000Root wordsprogress ==> progress
011Derivation and Compoundtankbuster ==> tank @@bust @@er
110Inflection and Derivationurbanizes ==> urban @@ize @@s
001Compound onlyhotpot ==> hot @@pot
111Inflection, Derivation, Compoundtrackworkers ==> track @@work @@er @@s

Baseline results

The following table shows the word-level task results of pretrained BertTokenizer on English. This pretrained model was employed from HuggingFace.

word classinflectionderivationcompoundRPF1lev. distance
001nonoyes64.1146.6053.971.42
101yesnoyes50.1251.5750.831.51
011noyesyes38.9336.8537.862.96
111yesyesyes28.3034.8131.223.22
010noyesno33.9025.2928.972.75
110yesyesno26.1724.9225.533.31
100yesnono19.1412.1614.872.73
000nonono5.552.022.962.11
total---28.7920.9924.282.69

Part 2: Sentence-level Morpheme Segmentation

At the sentence level, participating systems are expected to predict a sequence of morphemes for a given sentence. The following shows two lines of English data:

Six weeks of basic training . Six week @@s of base @@ic train @@ing .
Fistfights , please . Fist @@fight @@s , please .

The following shows two lines of Mongolian data:

Гэрт эмээ хоол хийв . Гэр @@т эмээ хоол хийх @@в .
Би өдөр эмээ уусан . Би өдөр эм @@ээ уух @@сан .

In above example, эмээ is a hononym of two different words, first means a grandmother and second is medicine. Depending on the context, the second homonym word is inflectional form of medicine and it is segmentable.

Languages

Development languages are:

  1. ces: Czech
  2. eng: English
  3. mon: Mongolian

Data Statistics

traindevtest
Czech1000500500
English1100717831845
Mongolian1000500600

Baseline results

LanguagePrecisonRecallF-measureLev. distance
Czech36.7630.3533.2521.01
English63.6865.7764.715.50
Mongolian20.0029.9523.9928.86

Evaluation

We will provide python evaluation scripts, reporting the following evaluation measures:

  • Precision - fraction of correctly predicted morphemes on all predicted morphemes
  • Recall - ratio of correctly predicted morphemes on all gold morphemes
  • F-measure - the harmonic mean of the precision and recall
  • Edit distance - average Levenshtein distance between the predicted output and the gold instance.

Submission

Please submit your team's results to khuyagbaatar.b@gmail.com CCing your teammates by May 13th, 2022 (AoE). Each submission should be a .tar.gz or .zip file.

Timeline

Development Phase

Generalization Phase

  • April 8 April 18, 2022: Training and development splits for surprise languages released /data/surprise

Evaluation Phase

  • April 15 April 29, 2022: Test splits for development and surprise languages are released at /data
  • April 29 May 13, 2022: Participants' submissions due.

Write-up Phase

  • May 13 May 27, 2022: Participants' draft system description papers due.
  • May 20 June 3, 2022: Participants' camera-ready system description papers due.

Organizers

  • Khuyagbaatar Batsuren (National University of Mongolia)
  • Gábor Bella (University of Trento)
  • Aryaman Arora (Georgetown University)
  • Viktor Martinović (University of Vienna)
  • Kyle Gorman (Graduate center, City University Of New York)
  • Zdeněk Žabokrtský (Charles University)
  • Amarsanaa Ganbold (National University of Mongolia)
  • Šárka Dohnalová (Charles University)
  • Magda Ševčíková (Charles University)
  • Kateřina Pelegrinová (University of Ostrava)
  • Fausto Giunchiglia (University of Trento)
  • Ryan Cotterell (ETH Zürich)
  • Ekaterina Vylomova (University of Melbourne)

License

The data is released under the Creative Commons Attribution-ShareAlike 3.0 Unported License inherited from Wiktionary itself.

References

Kirov, C., Cotterell, R., Sylak-Glassman, J., Walther, G., Vylomova, E., Xia, P., Faruqui, M., Mielke, S., McCarthy, A., Kübler, S., Yarowsky, D., Eisner, J., and Hulden, M. (2018). UniMorph 2.0: Universal Morphology. Proceedings of LREC 2018.

McCarthy, A.D., Kirov, C., Grella, M., Nidhi, A., Xia, P., Gorman, K., Vylomova, E., Mielke, S.J., Nicolai, G., Silfverberg, M. and Arkhangelskij, T., (2020). UniMorph 3.0: Universal Morphology.. Proceedings of LREC 2020.

Batsuren, K., Bella, G. and Giunchiglia, F., (2021). MorphyNet: a Large Multilingual Database of Derivational and Inflectional Morphology. In Proceedings of SIGMORPHON 2021 (pp. 39-48).

Nivre, J., Agić, Ž., Ahrenberg, L., Antonsen, L., Aranzabe, M.J., Asahara, M., Ateyah, L., Attia, M., Atutxa, A., Augustinus, L. and Badmaeva, E., (2017). Universal Dependencies 2.1.

About

SIGMORPHON 2022 Shared Task on Morpheme Segmentation

Topics

Resources

Stars

36 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

137 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SIGMORPHON 2022 Shared Task on Morpheme Segmentation

Morphemes (prefixes, suffixes, root words) are linguistic descriptions, defined as the smallest meaningful unit of words. Our proposed shared task is morpheme segmentation that converts a text into a sequence of morphemes. In order to prepare a dataset for this task, we integrated all basic types of morphological databases (including UniMorph (Kirov et al., 2018b; McCarthy et al., 2020) – inflectional morphology; MorphyNet (Batsuren et al., 2021) – derivational morphology; Universal Dependencies (Nivre et al., 2017) and ten editions of Wiktionary – compound morphology and root words). In the future, we expect the NLP community will benefit a lot by innovating subword-based tokenization with this task. This shared task has two parts:

Please join our Google Group to stay up to date. Click here to register for the task!

Please open the issues if you have any questions.

Two subtasks will be scored separately. Participant teams may submit as many systems as they want to as many subtasks as they want.

Part 1: Word-level Morpheme Segmentation

At the word level, participants will be asked to segment a given word into a sequence of morphemes. Input words contains all types of word forms: root words, derived words, inflected words, and compound words.

Data

Training and development data are UTF-8-encoded tab-separated values files. Each example occupies a single line and consists of input word, the corresponding morpheme sequence, and the corresponding morphological category. The following shows three lines of English data:

inaccuracies in @@accurate @@cy @@s 110
dictionary dictionary 000
screwdriver screw @@drive @@er 011

Note: The third column as the morphological category is an optional feature that can only be used to oversample or undersample training data.

First example is a derived word with prefix (in-) and suffixes (-cy and -s), and second example is a root word. Third example is a compound word. In the test datasets, we will provide only first column of data as input words.

Languages

Development languages are:

  1. ces: Czech
  2. eng: English
  3. fra: French
  4. hun: Hungarian
  5. spa: Spanish
  6. ita: Italian
  7. lat: Latin
  8. rus: Russian

Surprise language is:

  1. mon: Mongolian

Data Statistics

word classEnglishSpanishHungarianFrenchItalianRussianCzechLatinMongolian
100126544502229410662105192253455221760-8319917266
0102031021844924923679834109272970-02201
101137904581011894783171909-035
00010193815843695213619210372921-503381604
0115381821654506140328-00
110106570346862323119126196237104481409-07855
0011699024833201684431259-05
1113059343542791861582658-00
total words5773748845149260983827975537347842143868288232918966

Word category description

For some of the development languages, we are providing the word categories so that participant can deal with imbalanced situation of morphological categories.

word classDescriptionEnglish example (input ==> output)
100Inflection onlyplayed ==> play @@ed
010Derivation onlyplayer ==> play @@er
101Inflection and Compoundwheelbands ==> wheel @@band @@s
000Root wordsprogress ==> progress
011Derivation and Compoundtankbuster ==> tank @@bust @@er
110Inflection and Derivationurbanizes ==> urban @@ize @@s
001Compound onlyhotpot ==> hot @@pot
111Inflection, Derivation, Compoundtrackworkers ==> track @@work @@er @@s

Baseline results

The following table shows the word-level task results of pretrained BertTokenizer on English. This pretrained model was employed from HuggingFace.

word classinflectionderivationcompoundRPF1lev. distance
001nonoyes64.1146.6053.971.42
101yesnoyes50.1251.5750.831.51
011noyesyes38.9336.8537.862.96
111yesyesyes28.3034.8131.223.22
010noyesno33.9025.2928.972.75
110yesyesno26.1724.9225.533.31
100yesnono19.1412.1614.872.73
000nonono5.552.022.962.11
total---28.7920.9924.282.69

Part 2: Sentence-level Morpheme Segmentation

At the sentence level, participating systems are expected to predict a sequence of morphemes for a given sentence. The following shows two lines of English data:

Six weeks of basic training . Six week @@s of base @@ic train @@ing .
Fistfights , please . Fist @@fight @@s , please .

The following shows two lines of Mongolian data:

Гэрт эмээ хоол хийв . Гэр @@т эмээ хоол хийх @@в .
Би өдөр эмээ уусан . Би өдөр эм @@ээ уух @@сан .

In above example, эмээ is a hononym of two different words, first means a grandmother and second is medicine. Depending on the context, the second homonym word is inflectional form of medicine and it is segmentable.

Languages

Development languages are:

  1. ces: Czech
  2. eng: English
  3. mon: Mongolian

Data Statistics

traindevtest
Czech1000500500
English1100717831845
Mongolian1000500600

Baseline results

LanguagePrecisonRecallF-measureLev. distance
Czech36.7630.3533.2521.01
English63.6865.7764.715.50
Mongolian20.0029.9523.9928.86

Evaluation

We will provide python evaluation scripts, reporting the following evaluation measures:

  • Precision - fraction of correctly predicted morphemes on all predicted morphemes
  • Recall - ratio of correctly predicted morphemes on all gold morphemes
  • F-measure - the harmonic mean of the precision and recall
  • Edit distance - average Levenshtein distance between the predicted output and the gold instance.

Submission

Please submit your team's results to khuyagbaatar.b@gmail.com CCing your teammates by May 13th, 2022 (AoE). Each submission should be a .tar.gz or .zip file.

Timeline

Development Phase

Generalization Phase

  • April 8 April 18, 2022: Training and development splits for surprise languages released /data/surprise

Evaluation Phase

  • April 15 April 29, 2022: Test splits for development and surprise languages are released at /data
  • April 29 May 13, 2022: Participants' submissions due.

Write-up Phase

  • May 13 May 27, 2022: Participants' draft system description papers due.
  • May 20 June 3, 2022: Participants' camera-ready system description papers due.

Organizers

  • Khuyagbaatar Batsuren (National University of Mongolia)
  • Gábor Bella (University of Trento)
  • Aryaman Arora (Georgetown University)
  • Viktor Martinović (University of Vienna)
  • Kyle Gorman (Graduate center, City University Of New York)
  • Zdeněk Žabokrtský (Charles University)
  • Amarsanaa Ganbold (National University of Mongolia)
  • Šárka Dohnalová (Charles University)
  • Magda Ševčíková (Charles University)
  • Kateřina Pelegrinová (University of Ostrava)
  • Fausto Giunchiglia (University of Trento)
  • Ryan Cotterell (ETH Zürich)
  • Ekaterina Vylomova (University of Melbourne)

License

The data is released under the Creative Commons Attribution-ShareAlike 3.0 Unported License inherited from Wiktionary itself.

References

Kirov, C., Cotterell, R., Sylak-Glassman, J., Walther, G., Vylomova, E., Xia, P., Faruqui, M., Mielke, S., McCarthy, A., Kübler, S., Yarowsky, D., Eisner, J., and Hulden, M. (2018). UniMorph 2.0: Universal Morphology. Proceedings of LREC 2018.

McCarthy, A.D., Kirov, C., Grella, M., Nidhi, A., Xia, P., Gorman, K., Vylomova, E., Mielke, S.J., Nicolai, G., Silfverberg, M. and Arkhangelskij, T., (2020). UniMorph 3.0: Universal Morphology.. Proceedings of LREC 2020.

Batsuren, K., Bella, G. and Giunchiglia, F., (2021). MorphyNet: a Large Multilingual Database of Derivational and Inflectional Morphology. In Proceedings of SIGMORPHON 2021 (pp. 39-48).

Nivre, J., Agić, Ž., Ahrenberg, L., Antonsen, L., Aranzabe, M.J., Asahara, M., Ateyah, L., Attia, M., Atutxa, A., Augustinus, L. and Badmaeva, E., (2017). Universal Dependencies 2.1.

About

SIGMORPHON 2022 Shared Task on Morpheme Segmentation

Topics

Resources

Stars

36 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

137 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SIGMORPHON 2022 Shared Task on Morpheme Segmentation

Morphemes (prefixes, suffixes, root words) are linguistic descriptions, defined as the smallest meaningful unit of words. Our proposed shared task is morpheme segmentation that converts a text into a sequence of morphemes. In order to prepare a dataset for this task, we integrated all basic types of morphological databases (including UniMorph (Kirov et al., 2018b; McCarthy et al., 2020) – inflectional morphology; MorphyNet (Batsuren et al., 2021) – derivational morphology; Universal Dependencies (Nivre et al., 2017) and ten editions of Wiktionary – compound morphology and root words). In the future, we expect the NLP community will benefit a lot by innovating subword-based tokenization with this task. This shared task has two parts:

Please join our Google Group to stay up to date. Click here to register for the task!

Please open the issues if you have any questions.

Two subtasks will be scored separately. Participant teams may submit as many systems as they want to as many subtasks as they want.

Part 1: Word-level Morpheme Segmentation

At the word level, participants will be asked to segment a given word into a sequence of morphemes. Input words contains all types of word forms: root words, derived words, inflected words, and compound words.

Data

Training and development data are UTF-8-encoded tab-separated values files. Each example occupies a single line and consists of input word, the corresponding morpheme sequence, and the corresponding morphological category. The following shows three lines of English data:

inaccuracies in @@accurate @@cy @@s 110
dictionary dictionary 000
screwdriver screw @@drive @@er 011

Note: The third column as the morphological category is an optional feature that can only be used to oversample or undersample training data.

First example is a derived word with prefix (in-) and suffixes (-cy and -s), and second example is a root word. Third example is a compound word. In the test datasets, we will provide only first column of data as input words.

Languages

Development languages are:

  1. ces: Czech
  2. eng: English
  3. fra: French
  4. hun: Hungarian
  5. spa: Spanish
  6. ita: Italian
  7. lat: Latin
  8. rus: Russian

Surprise language is:

  1. mon: Mongolian

Data Statistics

word classEnglishSpanishHungarianFrenchItalianRussianCzechLatinMongolian
100126544502229410662105192253455221760-8319917266
0102031021844924923679834109272970-02201
101137904581011894783171909-035
00010193815843695213619210372921-503381604
0115381821654506140328-00
110106570346862323119126196237104481409-07855
0011699024833201684431259-05
1113059343542791861582658-00
total words5773748845149260983827975537347842143868288232918966

Word category description

For some of the development languages, we are providing the word categories so that participant can deal with imbalanced situation of morphological categories.

word classDescriptionEnglish example (input ==> output)
100Inflection onlyplayed ==> play @@ed
010Derivation onlyplayer ==> play @@er
101Inflection and Compoundwheelbands ==> wheel @@band @@s
000Root wordsprogress ==> progress
011Derivation and Compoundtankbuster ==> tank @@bust @@er
110Inflection and Derivationurbanizes ==> urban @@ize @@s
001Compound onlyhotpot ==> hot @@pot
111Inflection, Derivation, Compoundtrackworkers ==> track @@work @@er @@s

Baseline results

The following table shows the word-level task results of pretrained BertTokenizer on English. This pretrained model was employed from HuggingFace.

word classinflectionderivationcompoundRPF1lev. distance
001nonoyes64.1146.6053.971.42
101yesnoyes50.1251.5750.831.51
011noyesyes38.9336.8537.862.96
111yesyesyes28.3034.8131.223.22
010noyesno33.9025.2928.972.75
110yesyesno26.1724.9225.533.31
100yesnono19.1412.1614.872.73
000nonono5.552.022.962.11
total---28.7920.9924.282.69

Part 2: Sentence-level Morpheme Segmentation

At the sentence level, participating systems are expected to predict a sequence of morphemes for a given sentence. The following shows two lines of English data:

Six weeks of basic training . Six week @@s of base @@ic train @@ing .
Fistfights , please . Fist @@fight @@s , please .

The following shows two lines of Mongolian data:

Гэрт эмээ хоол хийв . Гэр @@т эмээ хоол хийх @@в .
Би өдөр эмээ уусан . Би өдөр эм @@ээ уух @@сан .

In above example, эмээ is a hononym of two different words, first means a grandmother and second is medicine. Depending on the context, the second homonym word is inflectional form of medicine and it is segmentable.

Languages

Development languages are:

  1. ces: Czech
  2. eng: English
  3. mon: Mongolian

Data Statistics

traindevtest
Czech1000500500
English1100717831845
Mongolian1000500600

Baseline results

LanguagePrecisonRecallF-measureLev. distance
Czech36.7630.3533.2521.01
English63.6865.7764.715.50
Mongolian20.0029.9523.9928.86

Evaluation

We will provide python evaluation scripts, reporting the following evaluation measures:

  • Precision - fraction of correctly predicted morphemes on all predicted morphemes
  • Recall - ratio of correctly predicted morphemes on all gold morphemes
  • F-measure - the harmonic mean of the precision and recall
  • Edit distance - average Levenshtein distance between the predicted output and the gold instance.

Submission

Please submit your team's results to khuyagbaatar.b@gmail.com CCing your teammates by May 13th, 2022 (AoE). Each submission should be a .tar.gz or .zip file.

Timeline

Development Phase

Generalization Phase

  • April 8 April 18, 2022: Training and development splits for surprise languages released /data/surprise

Evaluation Phase

  • April 15 April 29, 2022: Test splits for development and surprise languages are released at /data
  • April 29 May 13, 2022: Participants' submissions due.

Write-up Phase

  • May 13 May 27, 2022: Participants' draft system description papers due.
  • May 20 June 3, 2022: Participants' camera-ready system description papers due.

Organizers

  • Khuyagbaatar Batsuren (National University of Mongolia)
  • Gábor Bella (University of Trento)
  • Aryaman Arora (Georgetown University)
  • Viktor Martinović (University of Vienna)
  • Kyle Gorman (Graduate center, City University Of New York)
  • Zdeněk Žabokrtský (Charles University)
  • Amarsanaa Ganbold (National University of Mongolia)
  • Šárka Dohnalová (Charles University)
  • Magda Ševčíková (Charles University)
  • Kateřina Pelegrinová (University of Ostrava)
  • Fausto Giunchiglia (University of Trento)
  • Ryan Cotterell (ETH Zürich)
  • Ekaterina Vylomova (University of Melbourne)

License

The data is released under the Creative Commons Attribution-ShareAlike 3.0 Unported License inherited from Wiktionary itself.

References

Kirov, C., Cotterell, R., Sylak-Glassman, J., Walther, G., Vylomova, E., Xia, P., Faruqui, M., Mielke, S., McCarthy, A., Kübler, S., Yarowsky, D., Eisner, J., and Hulden, M. (2018). UniMorph 2.0: Universal Morphology. Proceedings of LREC 2018.

McCarthy, A.D., Kirov, C., Grella, M., Nidhi, A., Xia, P., Gorman, K., Vylomova, E., Mielke, S.J., Nicolai, G., Silfverberg, M. and Arkhangelskij, T., (2020). UniMorph 3.0: Universal Morphology.. Proceedings of LREC 2020.

Batsuren, K., Bella, G. and Giunchiglia, F., (2021). MorphyNet: a Large Multilingual Database of Derivational and Inflectional Morphology. In Proceedings of SIGMORPHON 2021 (pp. 39-48).

Nivre, J., Agić, Ž., Ahrenberg, L., Antonsen, L., Aranzabe, M.J., Asahara, M., Ateyah, L., Attia, M., Atutxa, A., Augustinus, L. and Badmaeva, E., (2017). Universal Dependencies 2.1.

About

SIGMORPHON 2022 Shared Task on Morpheme Segmentation

Topics

Resources

Stars

36 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

137 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SIGMORPHON 2022 Shared Task on Morpheme Segmentation

Morphemes (prefixes, suffixes, root words) are linguistic descriptions, defined as the smallest meaningful unit of words. Our proposed shared task is morpheme segmentation that converts a text into a sequence of morphemes. In order to prepare a dataset for this task, we integrated all basic types of morphological databases (including UniMorph (Kirov et al., 2018b; McCarthy et al., 2020) – inflectional morphology; MorphyNet (Batsuren et al., 2021) – derivational morphology; Universal Dependencies (Nivre et al., 2017) and ten editions of Wiktionary – compound morphology and root words). In the future, we expect the NLP community will benefit a lot by innovating subword-based tokenization with this task. This shared task has two parts:

Please join our Google Group to stay up to date. Click here to register for the task!

Please open the issues if you have any questions.

Two subtasks will be scored separately. Participant teams may submit as many systems as they want to as many subtasks as they want.

Part 1: Word-level Morpheme Segmentation

At the word level, participants will be asked to segment a given word into a sequence of morphemes. Input words contains all types of word forms: root words, derived words, inflected words, and compound words.

Data

Training and development data are UTF-8-encoded tab-separated values files. Each example occupies a single line and consists of input word, the corresponding morpheme sequence, and the corresponding morphological category. The following shows three lines of English data:

inaccuracies in @@accurate @@cy @@s 110
dictionary dictionary 000
screwdriver screw @@drive @@er 011

Note: The third column as the morphological category is an optional feature that can only be used to oversample or undersample training data.

First example is a derived word with prefix (in-) and suffixes (-cy and -s), and second example is a root word. Third example is a compound word. In the test datasets, we will provide only first column of data as input words.

Languages

Development languages are:

  1. ces: Czech
  2. eng: English
  3. fra: French
  4. hun: Hungarian
  5. spa: Spanish
  6. ita: Italian
  7. lat: Latin
  8. rus: Russian

Surprise language is:

  1. mon: Mongolian

Data Statistics

word classEnglishSpanishHungarianFrenchItalianRussianCzechLatinMongolian
100126544502229410662105192253455221760-8319917266
0102031021844924923679834109272970-02201
101137904581011894783171909-035
00010193815843695213619210372921-503381604
0115381821654506140328-00
110106570346862323119126196237104481409-07855
0011699024833201684431259-05
1113059343542791861582658-00
total words5773748845149260983827975537347842143868288232918966

Word category description

For some of the development languages, we are providing the word categories so that participant can deal with imbalanced situation of morphological categories.

word classDescriptionEnglish example (input ==> output)
100Inflection onlyplayed ==> play @@ed
010Derivation onlyplayer ==> play @@er
101Inflection and Compoundwheelbands ==> wheel @@band @@s
000Root wordsprogress ==> progress
011Derivation and Compoundtankbuster ==> tank @@bust @@er
110Inflection and Derivationurbanizes ==> urban @@ize @@s
001Compound onlyhotpot ==> hot @@pot
111Inflection, Derivation, Compoundtrackworkers ==> track @@work @@er @@s

Baseline results

The following table shows the word-level task results of pretrained BertTokenizer on English. This pretrained model was employed from HuggingFace.

word classinflectionderivationcompoundRPF1lev. distance
001nonoyes64.1146.6053.971.42
101yesnoyes50.1251.5750.831.51
011noyesyes38.9336.8537.862.96
111yesyesyes28.3034.8131.223.22
010noyesno33.9025.2928.972.75
110yesyesno26.1724.9225.533.31
100yesnono19.1412.1614.872.73
000nonono5.552.022.962.11
total---28.7920.9924.282.69

Part 2: Sentence-level Morpheme Segmentation

At the sentence level, participating systems are expected to predict a sequence of morphemes for a given sentence. The following shows two lines of English data:

Six weeks of basic training . Six week @@s of base @@ic train @@ing .
Fistfights , please . Fist @@fight @@s , please .

The following shows two lines of Mongolian data:

Гэрт эмээ хоол хийв . Гэр @@т эмээ хоол хийх @@в .
Би өдөр эмээ уусан . Би өдөр эм @@ээ уух @@сан .

In above example, эмээ is a hononym of two different words, first means a grandmother and second is medicine. Depending on the context, the second homonym word is inflectional form of medicine and it is segmentable.

Languages

Development languages are:

  1. ces: Czech
  2. eng: English
  3. mon: Mongolian

Data Statistics

traindevtest
Czech1000500500
English1100717831845
Mongolian1000500600

Baseline results

LanguagePrecisonRecallF-measureLev. distance
Czech36.7630.3533.2521.01
English63.6865.7764.715.50
Mongolian20.0029.9523.9928.86

Evaluation

We will provide python evaluation scripts, reporting the following evaluation measures:

  • Precision - fraction of correctly predicted morphemes on all predicted morphemes
  • Recall - ratio of correctly predicted morphemes on all gold morphemes
  • F-measure - the harmonic mean of the precision and recall
  • Edit distance - average Levenshtein distance between the predicted output and the gold instance.

Submission

Please submit your team's results to khuyagbaatar.b@gmail.com CCing your teammates by May 13th, 2022 (AoE). Each submission should be a .tar.gz or .zip file.

Timeline

Development Phase

Generalization Phase

  • April 8 April 18, 2022: Training and development splits for surprise languages released /data/surprise

Evaluation Phase

  • April 15 April 29, 2022: Test splits for development and surprise languages are released at /data
  • April 29 May 13, 2022: Participants' submissions due.

Write-up Phase

  • May 13 May 27, 2022: Participants' draft system description papers due.
  • May 20 June 3, 2022: Participants' camera-ready system description papers due.

Organizers

  • Khuyagbaatar Batsuren (National University of Mongolia)
  • Gábor Bella (University of Trento)
  • Aryaman Arora (Georgetown University)
  • Viktor Martinović (University of Vienna)
  • Kyle Gorman (Graduate center, City University Of New York)
  • Zdeněk Žabokrtský (Charles University)
  • Amarsanaa Ganbold (National University of Mongolia)
  • Šárka Dohnalová (Charles University)
  • Magda Ševčíková (Charles University)
  • Kateřina Pelegrinová (University of Ostrava)
  • Fausto Giunchiglia (University of Trento)
  • Ryan Cotterell (ETH Zürich)
  • Ekaterina Vylomova (University of Melbourne)

License

The data is released under the Creative Commons Attribution-ShareAlike 3.0 Unported License inherited from Wiktionary itself.

References

Kirov, C., Cotterell, R., Sylak-Glassman, J., Walther, G., Vylomova, E., Xia, P., Faruqui, M., Mielke, S., McCarthy, A., Kübler, S., Yarowsky, D., Eisner, J., and Hulden, M. (2018). UniMorph 2.0: Universal Morphology. Proceedings of LREC 2018.

McCarthy, A.D., Kirov, C., Grella, M., Nidhi, A., Xia, P., Gorman, K., Vylomova, E., Mielke, S.J., Nicolai, G., Silfverberg, M. and Arkhangelskij, T., (2020). UniMorph 3.0: Universal Morphology.. Proceedings of LREC 2020.

Batsuren, K., Bella, G. and Giunchiglia, F., (2021). MorphyNet: a Large Multilingual Database of Derivational and Inflectional Morphology. In Proceedings of SIGMORPHON 2021 (pp. 39-48).

Nivre, J., Agić, Ž., Ahrenberg, L., Antonsen, L., Aranzabe, M.J., Asahara, M., Ateyah, L., Attia, M., Atutxa, A., Augustinus, L. and Badmaeva, E., (2017). Universal Dependencies 2.1.

About

SIGMORPHON 2022 Shared Task on Morpheme Segmentation

Topics

Resources

Stars

36 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

137 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SIGMORPHON 2022 Shared Task on Morpheme Segmentation

Morphemes (prefixes, suffixes, root words) are linguistic descriptions, defined as the smallest meaningful unit of words. Our proposed shared task is morpheme segmentation that converts a text into a sequence of morphemes. In order to prepare a dataset for this task, we integrated all basic types of morphological databases (including UniMorph (Kirov et al., 2018b; McCarthy et al., 2020) – inflectional morphology; MorphyNet (Batsuren et al., 2021) – derivational morphology; Universal Dependencies (Nivre et al., 2017) and ten editions of Wiktionary – compound morphology and root words). In the future, we expect the NLP community will benefit a lot by innovating subword-based tokenization with this task. This shared task has two parts:

Please join our Google Group to stay up to date. Click here to register for the task!

Please open the issues if you have any questions.

Two subtasks will be scored separately. Participant teams may submit as many systems as they want to as many subtasks as they want.

Part 1: Word-level Morpheme Segmentation

At the word level, participants will be asked to segment a given word into a sequence of morphemes. Input words contains all types of word forms: root words, derived words, inflected words, and compound words.

Data

Training and development data are UTF-8-encoded tab-separated values files. Each example occupies a single line and consists of input word, the corresponding morpheme sequence, and the corresponding morphological category. The following shows three lines of English data:

inaccuracies in @@accurate @@cy @@s 110
dictionary dictionary 000
screwdriver screw @@drive @@er 011

Note: The third column as the morphological category is an optional feature that can only be used to oversample or undersample training data.

First example is a derived word with prefix (in-) and suffixes (-cy and -s), and second example is a root word. Third example is a compound word. In the test datasets, we will provide only first column of data as input words.

Languages

Development languages are:

  1. ces: Czech
  2. eng: English
  3. fra: French
  4. hun: Hungarian
  5. spa: Spanish
  6. ita: Italian
  7. lat: Latin
  8. rus: Russian

Surprise language is:

  1. mon: Mongolian

Data Statistics

word classEnglishSpanishHungarianFrenchItalianRussianCzechLatinMongolian
100126544502229410662105192253455221760-8319917266
0102031021844924923679834109272970-02201
101137904581011894783171909-035
00010193815843695213619210372921-503381604
0115381821654506140328-00
110106570346862323119126196237104481409-07855
0011699024833201684431259-05
1113059343542791861582658-00
total words5773748845149260983827975537347842143868288232918966

Word category description

For some of the development languages, we are providing the word categories so that participant can deal with imbalanced situation of morphological categories.

word classDescriptionEnglish example (input ==> output)
100Inflection onlyplayed ==> play @@ed
010Derivation onlyplayer ==> play @@er
101Inflection and Compoundwheelbands ==> wheel @@band @@s
000Root wordsprogress ==> progress
011Derivation and Compoundtankbuster ==> tank @@bust @@er
110Inflection and Derivationurbanizes ==> urban @@ize @@s
001Compound onlyhotpot ==> hot @@pot
111Inflection, Derivation, Compoundtrackworkers ==> track @@work @@er @@s

Baseline results

The following table shows the word-level task results of pretrained BertTokenizer on English. This pretrained model was employed from HuggingFace.

word classinflectionderivationcompoundRPF1lev. distance
001nonoyes64.1146.6053.971.42
101yesnoyes50.1251.5750.831.51
011noyesyes38.9336.8537.862.96
111yesyesyes28.3034.8131.223.22
010noyesno33.9025.2928.972.75
110yesyesno26.1724.9225.533.31
100yesnono19.1412.1614.872.73
000nonono5.552.022.962.11
total---28.7920.9924.282.69

Part 2: Sentence-level Morpheme Segmentation

At the sentence level, participating systems are expected to predict a sequence of morphemes for a given sentence. The following shows two lines of English data:

Six weeks of basic training . Six week @@s of base @@ic train @@ing .
Fistfights , please . Fist @@fight @@s , please .

The following shows two lines of Mongolian data:

Гэрт эмээ хоол хийв . Гэр @@т эмээ хоол хийх @@в .
Би өдөр эмээ уусан . Би өдөр эм @@ээ уух @@сан .

In above example, эмээ is a hononym of two different words, first means a grandmother and second is medicine. Depending on the context, the second homonym word is inflectional form of medicine and it is segmentable.

Languages

Development languages are:

  1. ces: Czech
  2. eng: English
  3. mon: Mongolian

Data Statistics

traindevtest
Czech1000500500
English1100717831845
Mongolian1000500600

Baseline results

LanguagePrecisonRecallF-measureLev. distance
Czech36.7630.3533.2521.01
English63.6865.7764.715.50
Mongolian20.0029.9523.9928.86

Evaluation

We will provide python evaluation scripts, reporting the following evaluation measures:

  • Precision - fraction of correctly predicted morphemes on all predicted morphemes
  • Recall - ratio of correctly predicted morphemes on all gold morphemes
  • F-measure - the harmonic mean of the precision and recall
  • Edit distance - average Levenshtein distance between the predicted output and the gold instance.

Submission

Please submit your team's results to khuyagbaatar.b@gmail.com CCing your teammates by May 13th, 2022 (AoE). Each submission should be a .tar.gz or .zip file.

Timeline

Development Phase

Generalization Phase

  • April 8 April 18, 2022: Training and development splits for surprise languages released /data/surprise

Evaluation Phase

  • April 15 April 29, 2022: Test splits for development and surprise languages are released at /data
  • April 29 May 13, 2022: Participants' submissions due.

Write-up Phase

  • May 13 May 27, 2022: Participants' draft system description papers due.
  • May 20 June 3, 2022: Participants' camera-ready system description papers due.

Organizers

  • Khuyagbaatar Batsuren (National University of Mongolia)
  • Gábor Bella (University of Trento)
  • Aryaman Arora (Georgetown University)
  • Viktor Martinović (University of Vienna)
  • Kyle Gorman (Graduate center, City University Of New York)
  • Zdeněk Žabokrtský (Charles University)
  • Amarsanaa Ganbold (National University of Mongolia)
  • Šárka Dohnalová (Charles University)
  • Magda Ševčíková (Charles University)
  • Kateřina Pelegrinová (University of Ostrava)
  • Fausto Giunchiglia (University of Trento)
  • Ryan Cotterell (ETH Zürich)
  • Ekaterina Vylomova (University of Melbourne)

License

The data is released under the Creative Commons Attribution-ShareAlike 3.0 Unported License inherited from Wiktionary itself.

References

Kirov, C., Cotterell, R., Sylak-Glassman, J., Walther, G., Vylomova, E., Xia, P., Faruqui, M., Mielke, S., McCarthy, A., Kübler, S., Yarowsky, D., Eisner, J., and Hulden, M. (2018). UniMorph 2.0: Universal Morphology. Proceedings of LREC 2018.

McCarthy, A.D., Kirov, C., Grella, M., Nidhi, A., Xia, P., Gorman, K., Vylomova, E., Mielke, S.J., Nicolai, G., Silfverberg, M. and Arkhangelskij, T., (2020). UniMorph 3.0: Universal Morphology.. Proceedings of LREC 2020.

Batsuren, K., Bella, G. and Giunchiglia, F., (2021). MorphyNet: a Large Multilingual Database of Derivational and Inflectional Morphology. In Proceedings of SIGMORPHON 2021 (pp. 39-48).

Nivre, J., Agić, Ž., Ahrenberg, L., Antonsen, L., Aranzabe, M.J., Asahara, M., Ateyah, L., Attia, M., Atutxa, A., Augustinus, L. and Badmaeva, E., (2017). Universal Dependencies 2.1.

About

SIGMORPHON 2022 Shared Task on Morpheme Segmentation

Topics

Resources

Stars

36 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

137 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SIGMORPHON 2022 Shared Task on Morpheme Segmentation

Morphemes (prefixes, suffixes, root words) are linguistic descriptions, defined as the smallest meaningful unit of words. Our proposed shared task is morpheme segmentation that converts a text into a sequence of morphemes. In order to prepare a dataset for this task, we integrated all basic types of morphological databases (including UniMorph (Kirov et al., 2018b; McCarthy et al., 2020) – inflectional morphology; MorphyNet (Batsuren et al., 2021) – derivational morphology; Universal Dependencies (Nivre et al., 2017) and ten editions of Wiktionary – compound morphology and root words). In the future, we expect the NLP community will benefit a lot by innovating subword-based tokenization with this task. This shared task has two parts:

Please join our Google Group to stay up to date. Click here to register for the task!

Please open the issues if you have any questions.

Two subtasks will be scored separately. Participant teams may submit as many systems as they want to as many subtasks as they want.

Part 1: Word-level Morpheme Segmentation

At the word level, participants will be asked to segment a given word into a sequence of morphemes. Input words contains all types of word forms: root words, derived words, inflected words, and compound words.

Data

Training and development data are UTF-8-encoded tab-separated values files. Each example occupies a single line and consists of input word, the corresponding morpheme sequence, and the corresponding morphological category. The following shows three lines of English data:

inaccuracies in @@accurate @@cy @@s 110
dictionary dictionary 000
screwdriver screw @@drive @@er 011

Note: The third column as the morphological category is an optional feature that can only be used to oversample or undersample training data.

First example is a derived word with prefix (in-) and suffixes (-cy and -s), and second example is a root word. Third example is a compound word. In the test datasets, we will provide only first column of data as input words.

Languages

Development languages are:

  1. ces: Czech
  2. eng: English
  3. fra: French
  4. hun: Hungarian
  5. spa: Spanish
  6. ita: Italian
  7. lat: Latin
  8. rus: Russian

Surprise language is:

  1. mon: Mongolian

Data Statistics

word classEnglishSpanishHungarianFrenchItalianRussianCzechLatinMongolian
100126544502229410662105192253455221760-8319917266
0102031021844924923679834109272970-02201
101137904581011894783171909-035
00010193815843695213619210372921-503381604
0115381821654506140328-00
110106570346862323119126196237104481409-07855
0011699024833201684431259-05
1113059343542791861582658-00
total words5773748845149260983827975537347842143868288232918966

Word category description

For some of the development languages, we are providing the word categories so that participant can deal with imbalanced situation of morphological categories.

word classDescriptionEnglish example (input ==> output)
100Inflection onlyplayed ==> play @@ed
010Derivation onlyplayer ==> play @@er
101Inflection and Compoundwheelbands ==> wheel @@band @@s
000Root wordsprogress ==> progress
011Derivation and Compoundtankbuster ==> tank @@bust @@er
110Inflection and Derivationurbanizes ==> urban @@ize @@s
001Compound onlyhotpot ==> hot @@pot
111Inflection, Derivation, Compoundtrackworkers ==> track @@work @@er @@s

Baseline results

The following table shows the word-level task results of pretrained BertTokenizer on English. This pretrained model was employed from HuggingFace.

word classinflectionderivationcompoundRPF1lev. distance
001nonoyes64.1146.6053.971.42
101yesnoyes50.1251.5750.831.51
011noyesyes38.9336.8537.862.96
111yesyesyes28.3034.8131.223.22
010noyesno33.9025.2928.972.75
110yesyesno26.1724.9225.533.31
100yesnono19.1412.1614.872.73
000nonono5.552.022.962.11
total---28.7920.9924.282.69

Part 2: Sentence-level Morpheme Segmentation

At the sentence level, participating systems are expected to predict a sequence of morphemes for a given sentence. The following shows two lines of English data:

Six weeks of basic training . Six week @@s of base @@ic train @@ing .
Fistfights , please . Fist @@fight @@s , please .

The following shows two lines of Mongolian data:

Гэрт эмээ хоол хийв . Гэр @@т эмээ хоол хийх @@в .
Би өдөр эмээ уусан . Би өдөр эм @@ээ уух @@сан .

In above example, эмээ is a hononym of two different words, first means a grandmother and second is medicine. Depending on the context, the second homonym word is inflectional form of medicine and it is segmentable.

Languages

Development languages are:

  1. ces: Czech
  2. eng: English
  3. mon: Mongolian

Data Statistics

traindevtest
Czech1000500500
English1100717831845
Mongolian1000500600

Baseline results

LanguagePrecisonRecallF-measureLev. distance
Czech36.7630.3533.2521.01
English63.6865.7764.715.50
Mongolian20.0029.9523.9928.86

Evaluation

We will provide python evaluation scripts, reporting the following evaluation measures:

  • Precision - fraction of correctly predicted morphemes on all predicted morphemes
  • Recall - ratio of correctly predicted morphemes on all gold morphemes
  • F-measure - the harmonic mean of the precision and recall
  • Edit distance - average Levenshtein distance between the predicted output and the gold instance.

Submission

Please submit your team's results to khuyagbaatar.b@gmail.com CCing your teammates by May 13th, 2022 (AoE). Each submission should be a .tar.gz or .zip file.

Timeline

Development Phase

Generalization Phase

  • April 8 April 18, 2022: Training and development splits for surprise languages released /data/surprise

Evaluation Phase

  • April 15 April 29, 2022: Test splits for development and surprise languages are released at /data
  • April 29 May 13, 2022: Participants' submissions due.

Write-up Phase

  • May 13 May 27, 2022: Participants' draft system description papers due.
  • May 20 June 3, 2022: Participants' camera-ready system description papers due.

Organizers

  • Khuyagbaatar Batsuren (National University of Mongolia)
  • Gábor Bella (University of Trento)
  • Aryaman Arora (Georgetown University)
  • Viktor Martinović (University of Vienna)
  • Kyle Gorman (Graduate center, City University Of New York)
  • Zdeněk Žabokrtský (Charles University)
  • Amarsanaa Ganbold (National University of Mongolia)
  • Šárka Dohnalová (Charles University)
  • Magda Ševčíková (Charles University)
  • Kateřina Pelegrinová (University of Ostrava)
  • Fausto Giunchiglia (University of Trento)
  • Ryan Cotterell (ETH Zürich)
  • Ekaterina Vylomova (University of Melbourne)

License

The data is released under the Creative Commons Attribution-ShareAlike 3.0 Unported License inherited from Wiktionary itself.

References

Kirov, C., Cotterell, R., Sylak-Glassman, J., Walther, G., Vylomova, E., Xia, P., Faruqui, M., Mielke, S., McCarthy, A., Kübler, S., Yarowsky, D., Eisner, J., and Hulden, M. (2018). UniMorph 2.0: Universal Morphology. Proceedings of LREC 2018.

McCarthy, A.D., Kirov, C., Grella, M., Nidhi, A., Xia, P., Gorman, K., Vylomova, E., Mielke, S.J., Nicolai, G., Silfverberg, M. and Arkhangelskij, T., (2020). UniMorph 3.0: Universal Morphology.. Proceedings of LREC 2020.

Batsuren, K., Bella, G. and Giunchiglia, F., (2021). MorphyNet: a Large Multilingual Database of Derivational and Inflectional Morphology. In Proceedings of SIGMORPHON 2021 (pp. 39-48).

Nivre, J., Agić, Ž., Ahrenberg, L., Antonsen, L., Aranzabe, M.J., Asahara, M., Ateyah, L., Attia, M., Atutxa, A., Augustinus, L. and Badmaeva, E., (2017). Universal Dependencies 2.1.

About

SIGMORPHON 2022 Shared Task on Morpheme Segmentation

Topics

Resources

Stars

36 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

137 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SIGMORPHON 2022 Shared Task on Morpheme Segmentation

Morphemes (prefixes, suffixes, root words) are linguistic descriptions, defined as the smallest meaningful unit of words. Our proposed shared task is morpheme segmentation that converts a text into a sequence of morphemes. In order to prepare a dataset for this task, we integrated all basic types of morphological databases (including UniMorph (Kirov et al., 2018b; McCarthy et al., 2020) – inflectional morphology; MorphyNet (Batsuren et al., 2021) – derivational morphology; Universal Dependencies (Nivre et al., 2017) and ten editions of Wiktionary – compound morphology and root words). In the future, we expect the NLP community will benefit a lot by innovating subword-based tokenization with this task. This shared task has two parts:

Please join our Google Group to stay up to date. Click here to register for the task!

Please open the issues if you have any questions.

Two subtasks will be scored separately. Participant teams may submit as many systems as they want to as many subtasks as they want.

Part 1: Word-level Morpheme Segmentation

At the word level, participants will be asked to segment a given word into a sequence of morphemes. Input words contains all types of word forms: root words, derived words, inflected words, and compound words.

Data

Training and development data are UTF-8-encoded tab-separated values files. Each example occupies a single line and consists of input word, the corresponding morpheme sequence, and the corresponding morphological category. The following shows three lines of English data:

inaccuracies in @@accurate @@cy @@s 110
dictionary dictionary 000
screwdriver screw @@drive @@er 011

Note: The third column as the morphological category is an optional feature that can only be used to oversample or undersample training data.

First example is a derived word with prefix (in-) and suffixes (-cy and -s), and second example is a root word. Third example is a compound word. In the test datasets, we will provide only first column of data as input words.

Languages

Development languages are:

  1. ces: Czech
  2. eng: English
  3. fra: French
  4. hun: Hungarian
  5. spa: Spanish
  6. ita: Italian
  7. lat: Latin
  8. rus: Russian

Surprise language is:

  1. mon: Mongolian

Data Statistics

word classEnglishSpanishHungarianFrenchItalianRussianCzechLatinMongolian
100126544502229410662105192253455221760-8319917266
0102031021844924923679834109272970-02201
101137904581011894783171909-035
00010193815843695213619210372921-503381604
0115381821654506140328-00
110106570346862323119126196237104481409-07855
0011699024833201684431259-05
1113059343542791861582658-00
total words5773748845149260983827975537347842143868288232918966

Word category description

For some of the development languages, we are providing the word categories so that participant can deal with imbalanced situation of morphological categories.

word classDescriptionEnglish example (input ==> output)
100Inflection onlyplayed ==> play @@ed
010Derivation onlyplayer ==> play @@er
101Inflection and Compoundwheelbands ==> wheel @@band @@s
000Root wordsprogress ==> progress
011Derivation and Compoundtankbuster ==> tank @@bust @@er
110Inflection and Derivationurbanizes ==> urban @@ize @@s
001Compound onlyhotpot ==> hot @@pot
111Inflection, Derivation, Compoundtrackworkers ==> track @@work @@er @@s

Baseline results

The following table shows the word-level task results of pretrained BertTokenizer on English. This pretrained model was employed from HuggingFace.

word classinflectionderivationcompoundRPF1lev. distance
001nonoyes64.1146.6053.971.42
101yesnoyes50.1251.5750.831.51
011noyesyes38.9336.8537.862.96
111yesyesyes28.3034.8131.223.22
010noyesno33.9025.2928.972.75
110yesyesno26.1724.9225.533.31
100yesnono19.1412.1614.872.73
000nonono5.552.022.962.11
total---28.7920.9924.282.69

Part 2: Sentence-level Morpheme Segmentation

At the sentence level, participating systems are expected to predict a sequence of morphemes for a given sentence. The following shows two lines of English data:

Six weeks of basic training . Six week @@s of base @@ic train @@ing .
Fistfights , please . Fist @@fight @@s , please .

The following shows two lines of Mongolian data:

Гэрт эмээ хоол хийв . Гэр @@т эмээ хоол хийх @@в .
Би өдөр эмээ уусан . Би өдөр эм @@ээ уух @@сан .

In above example, эмээ is a hononym of two different words, first means a grandmother and second is medicine. Depending on the context, the second homonym word is inflectional form of medicine and it is segmentable.

Languages

Development languages are:

  1. ces: Czech
  2. eng: English
  3. mon: Mongolian

Data Statistics

traindevtest
Czech1000500500
English1100717831845
Mongolian1000500600

Baseline results

LanguagePrecisonRecallF-measureLev. distance
Czech36.7630.3533.2521.01
English63.6865.7764.715.50
Mongolian20.0029.9523.9928.86

Evaluation

We will provide python evaluation scripts, reporting the following evaluation measures:

  • Precision - fraction of correctly predicted morphemes on all predicted morphemes
  • Recall - ratio of correctly predicted morphemes on all gold morphemes
  • F-measure - the harmonic mean of the precision and recall
  • Edit distance - average Levenshtein distance between the predicted output and the gold instance.

Submission

Please submit your team's results to khuyagbaatar.b@gmail.com CCing your teammates by May 13th, 2022 (AoE). Each submission should be a .tar.gz or .zip file.

Timeline

Development Phase

Generalization Phase

  • April 8 April 18, 2022: Training and development splits for surprise languages released /data/surprise

Evaluation Phase

  • April 15 April 29, 2022: Test splits for development and surprise languages are released at /data
  • April 29 May 13, 2022: Participants' submissions due.

Write-up Phase

  • May 13 May 27, 2022: Participants' draft system description papers due.
  • May 20 June 3, 2022: Participants' camera-ready system description papers due.

Organizers

  • Khuyagbaatar Batsuren (National University of Mongolia)
  • Gábor Bella (University of Trento)
  • Aryaman Arora (Georgetown University)
  • Viktor Martinović (University of Vienna)
  • Kyle Gorman (Graduate center, City University Of New York)
  • Zdeněk Žabokrtský (Charles University)
  • Amarsanaa Ganbold (National University of Mongolia)
  • Šárka Dohnalová (Charles University)
  • Magda Ševčíková (Charles University)
  • Kateřina Pelegrinová (University of Ostrava)
  • Fausto Giunchiglia (University of Trento)
  • Ryan Cotterell (ETH Zürich)
  • Ekaterina Vylomova (University of Melbourne)

License

The data is released under the Creative Commons Attribution-ShareAlike 3.0 Unported License inherited from Wiktionary itself.

References

Kirov, C., Cotterell, R., Sylak-Glassman, J., Walther, G., Vylomova, E., Xia, P., Faruqui, M., Mielke, S., McCarthy, A., Kübler, S., Yarowsky, D., Eisner, J., and Hulden, M. (2018). UniMorph 2.0: Universal Morphology. Proceedings of LREC 2018.

McCarthy, A.D., Kirov, C., Grella, M., Nidhi, A., Xia, P., Gorman, K., Vylomova, E., Mielke, S.J., Nicolai, G., Silfverberg, M. and Arkhangelskij, T., (2020). UniMorph 3.0: Universal Morphology.. Proceedings of LREC 2020.

Batsuren, K., Bella, G. and Giunchiglia, F., (2021). MorphyNet: a Large Multilingual Database of Derivational and Inflectional Morphology. In Proceedings of SIGMORPHON 2021 (pp. 39-48).

Nivre, J., Agić, Ž., Ahrenberg, L., Antonsen, L., Aranzabe, M.J., Asahara, M., Ateyah, L., Attia, M., Atutxa, A., Augustinus, L. and Badmaeva, E., (2017). Universal Dependencies 2.1.

About

SIGMORPHON 2022 Shared Task on Morpheme Segmentation

Topics

Resources

Stars

36 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

137 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

SIGMORPHON 2022 Shared Task on Morpheme Segmentation

Morphemes (prefixes, suffixes, root words) are linguistic descriptions, defined as the smallest meaningful unit of words. Our proposed shared task is morpheme segmentation that converts a text into a sequence of morphemes. In order to prepare a dataset for this task, we integrated all basic types of morphological databases (including UniMorph (Kirov et al., 2018b; McCarthy et al., 2020) – inflectional morphology; MorphyNet (Batsuren et al., 2021) – derivational morphology; Universal Dependencies (Nivre et al., 2017) and ten editions of Wiktionary – compound morphology and root words). In the future, we expect the NLP community will benefit a lot by innovating subword-based tokenization with this task. This shared task has two parts:

Please join our Google Group to stay up to date. Click here to register for the task!

Please open the issues if you have any questions.

Two subtasks will be scored separately. Participant teams may submit as many systems as they want to as many subtasks as they want.

Part 1: Word-level Morpheme Segmentation

At the word level, participants will be asked to segment a given word into a sequence of morphemes. Input words contains all types of word forms: root words, derived words, inflected words, and compound words.

Data

Training and development data are UTF-8-encoded tab-separated values files. Each example occupies a single line and consists of input word, the corresponding morpheme sequence, and the corresponding morphological category. The following shows three lines of English data:

inaccuracies in @@accurate @@cy @@s 110
dictionary dictionary 000
screwdriver screw @@drive @@er 011

Note: The third column as the morphological category is an optional feature that can only be used to oversample or undersample training data.

First example is a derived word with prefix (in-) and suffixes (-cy and -s), and second example is a root word. Third example is a compound word. In the test datasets, we will provide only first column of data as input words.

Languages

Development languages are:

  1. ces: Czech
  2. eng: English
  3. fra: French
  4. hun: Hungarian
  5. spa: Spanish
  6. ita: Italian
  7. lat: Latin
  8. rus: Russian

Surprise language is:

  1. mon: Mongolian

Data Statistics

word classEnglishSpanishHungarianFrenchItalianRussianCzechLatinMongolian
100126544502229410662105192253455221760-8319917266
0102031021844924923679834109272970-02201
101137904581011894783171909-035
00010193815843695213619210372921-503381604
0115381821654506140328-00
110106570346862323119126196237104481409-07855
0011699024833201684431259-05
1113059343542791861582658-00
total words5773748845149260983827975537347842143868288232918966

Word category description

For some of the development languages, we are providing the word categories so that participant can deal with imbalanced situation of morphological categories.

word classDescriptionEnglish example (input ==> output)
100Inflection onlyplayed ==> play @@ed
010Derivation onlyplayer ==> play @@er
101Inflection and Compoundwheelbands ==> wheel @@band @@s
000Root wordsprogress ==> progress
011Derivation and Compoundtankbuster ==> tank @@bust @@er
110Inflection and Derivationurbanizes ==> urban @@ize @@s
001Compound onlyhotpot ==> hot @@pot
111Inflection, Derivation, Compoundtrackworkers ==> track @@work @@er @@s

Baseline results

The following table shows the word-level task results of pretrained BertTokenizer on English. This pretrained model was employed from HuggingFace.

word classinflectionderivationcompoundRPF1lev. distance
001nonoyes64.1146.6053.971.42
101yesnoyes50.1251.5750.831.51
011noyesyes38.9336.8537.862.96
111yesyesyes28.3034.8131.223.22
010noyesno33.9025.2928.972.75
110yesyesno26.1724.9225.533.31
100yesnono19.1412.1614.872.73
000nonono5.552.022.962.11
total---28.7920.9924.282.69

Part 2: Sentence-level Morpheme Segmentation

At the sentence level, participating systems are expected to predict a sequence of morphemes for a given sentence. The following shows two lines of English data:

Six weeks of basic training . Six week @@s of base @@ic train @@ing .
Fistfights , please . Fist @@fight @@s , please .

The following shows two lines of Mongolian data:

Гэрт эмээ хоол хийв . Гэр @@т эмээ хоол хийх @@в .
Би өдөр эмээ уусан . Би өдөр эм @@ээ уух @@сан .

In above example, эмээ is a hononym of two different words, first means a grandmother and second is medicine. Depending on the context, the second homonym word is inflectional form of medicine and it is segmentable.

Languages

Development languages are:

  1. ces: Czech
  2. eng: English
  3. mon: Mongolian

Data Statistics

traindevtest
Czech1000500500
English1100717831845
Mongolian1000500600

Baseline results

LanguagePrecisonRecallF-measureLev. distance
Czech36.7630.3533.2521.01
English63.6865.7764.715.50
Mongolian20.0029.9523.9928.86

Evaluation

We will provide python evaluation scripts, reporting the following evaluation measures:

  • Precision - fraction of correctly predicted morphemes on all predicted morphemes
  • Recall - ratio of correctly predicted morphemes on all gold morphemes
  • F-measure - the harmonic mean of the precision and recall
  • Edit distance - average Levenshtein distance between the predicted output and the gold instance.

Submission

Please submit your team's results to khuyagbaatar.b@gmail.com CCing your teammates by May 13th, 2022 (AoE). Each submission should be a .tar.gz or .zip file.

Timeline

Development Phase

Generalization Phase

  • April 8 April 18, 2022: Training and development splits for surprise languages released /data/surprise

Evaluation Phase

  • April 15 April 29, 2022: Test splits for development and surprise languages are released at /data
  • April 29 May 13, 2022: Participants' submissions due.

Write-up Phase

  • May 13 May 27, 2022: Participants' draft system description papers due.
  • May 20 June 3, 2022: Participants' camera-ready system description papers due.

Organizers

  • Khuyagbaatar Batsuren (National University of Mongolia)
  • Gábor Bella (University of Trento)
  • Aryaman Arora (Georgetown University)
  • Viktor Martinović (University of Vienna)
  • Kyle Gorman (Graduate center, City University Of New York)
  • Zdeněk Žabokrtský (Charles University)
  • Amarsanaa Ganbold (National University of Mongolia)
  • Šárka Dohnalová (Charles University)
  • Magda Ševčíková (Charles University)
  • Kateřina Pelegrinová (University of Ostrava)
  • Fausto Giunchiglia (University of Trento)
  • Ryan Cotterell (ETH Zürich)
  • Ekaterina Vylomova (University of Melbourne)

License

The data is released under the Creative Commons Attribution-ShareAlike 3.0 Unported License inherited from Wiktionary itself.

References

Kirov, C., Cotterell, R., Sylak-Glassman, J., Walther, G., Vylomova, E., Xia, P., Faruqui, M., Mielke, S., McCarthy, A., Kübler, S., Yarowsky, D., Eisner, J., and Hulden, M. (2018). UniMorph 2.0: Universal Morphology. Proceedings of LREC 2018.

McCarthy, A.D., Kirov, C., Grella, M., Nidhi, A., Xia, P., Gorman, K., Vylomova, E., Mielke, S.J., Nicolai, G., Silfverberg, M. and Arkhangelskij, T., (2020). UniMorph 3.0: Universal Morphology.. Proceedings of LREC 2020.

Batsuren, K., Bella, G. and Giunchiglia, F., (2021). MorphyNet: a Large Multilingual Database of Derivational and Inflectional Morphology. In Proceedings of SIGMORPHON 2021 (pp. 39-48).

Nivre, J., Agić, Ž., Ahrenberg, L., Antonsen, L., Aranzabe, M.J., Asahara, M., Ateyah, L., Attia, M., Atutxa, A., Augustinus, L. and Badmaeva, E., (2017). Universal Dependencies 2.1.

About

SIGMORPHON 2022 Shared Task on Morpheme Segmentation

Topics

Resources

Stars

36 stars

Watchers

7 watching

Forks

Releases

Packages

Contributors

Languages