Repository files navigation

Build StatusDocsWikiLicense: GPL v3

MOTS

MOTS (MOdular Tool for Summarization) is a summarization system, written in Java. It is as modular as possible, and is intended to provide an architecture to implement and test new summarization methods, as well as to ease comparison with already implemented methods, in an unified framework. This system is the first completely modular system for automatic summarization and already allows to summarize using more than a hundred combinations of modules. The need for such a system is important. Indeed, several evaluation campaigns exist in AS field, but summarization algorithms are not easy to compare due to the large variety of pre and post-processings they use.

Getting Started

  • Javadoc available.
  • We provide an example corpus from TAC2009 in /src/main/resources and its associated human summaries.

Prerequisites

  • Maven - Dependency Manager
  • glpk-utils in order to use ILP (sudo apt-get install glpk-utils)
  • At least python3 in order to use WordEmbeddings, make sure to have an updated version of pip3 (sudo pip3 install --upgrade pip)
  • gensim in order to use WordEmbeddings (pip3 install gensim --user)
  • jep in order to use WordEmbeddings (pip3 install jep --user)

Installing

  • You might define $CORPUS_DATA to your DUC/TAC folder.
  • Install ROUGE :
    • Define $ROUGE_HOME to your ROUGE installation folder (or to ./ROUGE-1.5.5/RELEASE-1.5.5).
    • Install XML::DOM module in order to use ROUGE perl script. (sudo cpan install XML::DOM).
    • Run ./rouge_install.sh or :
      • Define $ROUGE_EVAL_HOME to $ROUGE_HOME/data.
      • Recreate database :
       cd data/WordNet-2.0-Exceptions/
      rm WordNet-2.0.exc.db # only if exist
      perl buildExeptionDB.pl . exc WordNet-2.0.exc.db
      cd ..
      rm WordNet-2.0.exc.db # only if exist
      ln -s WordNet-2.0-Exceptions/WordNet-2.0.exc.db WordNet-2.0.exc.db
      
  • Run install.sh script.

Usage

MOTS is a command line tool than can be used like this :

./MOTS mots.X.Y.Z.jar -c <config_file> -m <multicorpus_file> -v <OPTIONAL>

MOTS script encapsulate some environnement variable needed for the execution of WordEmbeddings. If you don't use WordEmbeddings you could launch via :

java -jar mots.X.Y.Z.jar -c <config_file> -m <multicorpus_file> -v <OPTIONAL>

Example config file and multicorpus file are provided in /conf but should be adapted to your setup.

Go Deeper

Each summarization process is defined in a configuration file and the test corpus is defined in a multicorpus configuration file.

Process configuration

Example for LexRank_MMR configuration file :

<CONFIG>
<TASK ID="1">
<LANGUAGE>english</LANGUAGE>
<OUTPUT_PATH>doc/output</OUTPUT_PATH>
<MULTITHREADING>true</MULTITHREADING> <PREPROCESS NAME="GenerateTextModel">
<OPTION NAME="StanfordNLP">true</OPTION>
<OPTION NAME="StopWordListFile">$CORPUS_DATA/stopwords/englishStopWords.txt</OPTION>
</PREPROCESS>
<PROCESS>
<OPTION NAME="CorpusIdToSummarize">all</OPTION>
<OPTION NAME="ReadStopWords">false</OPTION>
<INDEX_BUILDER NAME="TF_IDF.TF_IDF">
</INDEX_BUILDER>
<CARACTERISTIC_BUILDER NAME="vector.TfIdfVectorSentence">
</CARACTERISTIC_BUILDER>
<SCORING_METHOD NAME="graphBased.LexRank">
<OPTION NAME="DampingParameter">0.15</OPTION>
<OPTION NAME="GraphThreshold">0.1</OPTION>
<OPTION NAME="SimilarityMethod">JaccardSimilarity</OPTION>
</SCORING_METHOD>
<SUMMARIZE_METHOD NAME="MMR">
<OPTION NAME="CharLimitBoolean">true</OPTION>
<OPTION NAME="Size">200</OPTION>
<OPTION NAME="SimilarityMethod">JaccardSimilarity</OPTION>
<OPTION NAME="Lambda">0.6</OPTION>
</SUMMARIZE_METHOD>
</PROCESS>
<ROUGE_EVALUATION>
<ROUGE_MEASURE>ROUGE-1	ROUGE-2	ROUGE-SU4</ROUGE_MEASURE>
<MODEL_ROOT>models</MODEL_ROOT>
<PEER_ROOT>systems</PEER_ROOT>
</ROUGE_EVALUATION>
</TASK>
</CONFIG>
  • <CONFIG> is the root node.
    • <TASK> represent a summarization task. You could do multiple in a simple run. At start, stick with one.
      • <LANGUAGE> is the input's document language for preprocessing goal. (english / french for now)
      • <OUTPUT_PATH> is the forlder's output path of the system. It is used to save preprocessed documents, ROUGE xml generated file, old score, ...
      • <MULTITHREADING> (boolean) launch the system in a mutltithreading way or not.
      • <PREPROCESS> is the preprocess step for the system. The preprocess java class to use is pass by the name variable. Here it's GenerateTextModel. It also needs two <OPTION> :
        • <OPTION NAME="StanfordNLP"> (boolean), true if you want to use StanfordNLP pipeline and tool to do the preprocessing.
        • <OPTION NAME="StopWordListPath"> (String), path of the stopwords list you want to use.
      • <PROCESS> is the main step of the system. It should have at least one <SUMMARIZE_METHOD> node and two <OPTION> It often has an <INDEX_BUILDER> node and a <CARACTERISTIC_BUILDER> node :
        • <OPTION NAME="CorpusIdToSummarize"> (String as a list of int separated by \t), the list of CorpusId to summarize from the MultiCorpus configuration file. "all" will do summarization for all corpus.
        • <OPTION NAME="ReadStopWords"> (boolean), state if the system count stopwords as part of the texts or not.
        • <INDEX_BUILDER> is the step where the system generate a computer friendly representation of each text's textual unit. (TF-IDF, Bigram, WordEmbeddings, ...)
        • <CARATERISTIC_BUILDER> is the sentence caracteristic generation step based on the textual unit index building.
        • <SCORING_METHOD> weights each sentences.
        • <SUMMARIZE_BUILDER> generate a summary usually by ranking sentence based on their score.
      • <ROUGE_EVALUATION> is the ROUGE evaluation step. For detail, look at ROUGE readme in /lib/ROUGE folder.
        • <ROUGE_MEASURE> (String as a list of int separated by \t), represent the list of ROUGE measure you want to use.
        • <MODEL_ROOT> is the model's folder name for ROUGE xml input files.
        • <PEER_ROOT> is the peer's folder name for ROUGE xml input files.

The <PROCESS> node is the system's core and you should look for more detail in the javadoc and the source code of the different INDEX_BUILDER, CARACTERISTIC_BUILDER, SCORING_METHOD and SUMMARIZE_METHOD class.

Multicorpus configuration

<?xml version="1.0" encoding="UTF-8"?>
<CONFIG>
<TASK ID="1">
<MULTICORPUS ID="0">
<CORPUS ID="0">
<INPUT_PATH>$CORPUS_DATA/TAC2009/UpdateSumm09_test_docs_files/D0901A/D0901A-A</INPUT_PATH>
<DOCUMENT ID="0">.*</DOCUMENT>
<SUMMARY_PATH>$CORPUS_DATA/TAC2009/UpdateSumm09_eval/ROUGE/models</SUMMARY_PATH>
<SUMMARY ID="0">D0901-A.*</SUMMARY>
</CORPUS>
</MULTICORPUS>
</TASK>
</CONFIG>

For now, all ID are useless and could be avoided.

  • <CONFIG> is the root node.
    • <TASK> represent a summarization task. You could do multiple in a simple run. At start, stick with one.
      • <MULTICORPUS> is a list of <CORPUS>
        • <CORPUS> can be one or more documents. The system will generate one summary per corpus.
          • <INPUT_PATH> is the folder containing the corpus' documents.
          • <DOCUMENT> is the regex for the documents you want to load. You could use multiple <DOCUMENT> node.
          • <SUMMARY_PATH> is the human summaries folder path.
          • <SUMMARY> is the regex for human summary file associating to this corpus. You could use multiple <SUMMARY> node.

Built With

  • Maven - Dependency Management

Authors

License

This project is licensed under the GPL3 License - see the LICENSE.md file for details

Acknowledgments

  • Thanks to Aurélien Bossard, my PhD supervisor.

About

MOTS (MOdular Tool for Summarization) is a summarization system, written in Java. It is as modular as possible, and is intended to provide an architecture to implement and test new summarization methods, as well as to ease comparison with already implemented methods, in an unified framework.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Build StatusDocsWikiLicense: GPL v3

MOTS

MOTS (MOdular Tool for Summarization) is a summarization system, written in Java. It is as modular as possible, and is intended to provide an architecture to implement and test new summarization methods, as well as to ease comparison with already implemented methods, in an unified framework. This system is the first completely modular system for automatic summarization and already allows to summarize using more than a hundred combinations of modules. The need for such a system is important. Indeed, several evaluation campaigns exist in AS field, but summarization algorithms are not easy to compare due to the large variety of pre and post-processings they use.

Getting Started

  • Javadoc available.
  • We provide an example corpus from TAC2009 in /src/main/resources and its associated human summaries.

Prerequisites

  • Maven - Dependency Manager
  • glpk-utils in order to use ILP (sudo apt-get install glpk-utils)
  • At least python3 in order to use WordEmbeddings, make sure to have an updated version of pip3 (sudo pip3 install --upgrade pip)
  • gensim in order to use WordEmbeddings (pip3 install gensim --user)
  • jep in order to use WordEmbeddings (pip3 install jep --user)

Installing

  • You might define $CORPUS_DATA to your DUC/TAC folder.
  • Install ROUGE :
    • Define $ROUGE_HOME to your ROUGE installation folder (or to ./ROUGE-1.5.5/RELEASE-1.5.5).
    • Install XML::DOM module in order to use ROUGE perl script. (sudo cpan install XML::DOM).
    • Run ./rouge_install.sh or :
      • Define $ROUGE_EVAL_HOME to $ROUGE_HOME/data.
      • Recreate database :
       cd data/WordNet-2.0-Exceptions/
      rm WordNet-2.0.exc.db # only if exist
      perl buildExeptionDB.pl . exc WordNet-2.0.exc.db
      cd ..
      rm WordNet-2.0.exc.db # only if exist
      ln -s WordNet-2.0-Exceptions/WordNet-2.0.exc.db WordNet-2.0.exc.db
      
  • Run install.sh script.

Usage

MOTS is a command line tool than can be used like this :

./MOTS mots.X.Y.Z.jar -c <config_file> -m <multicorpus_file> -v <OPTIONAL>

MOTS script encapsulate some environnement variable needed for the execution of WordEmbeddings. If you don't use WordEmbeddings you could launch via :

java -jar mots.X.Y.Z.jar -c <config_file> -m <multicorpus_file> -v <OPTIONAL>

Example config file and multicorpus file are provided in /conf but should be adapted to your setup.

Go Deeper

Each summarization process is defined in a configuration file and the test corpus is defined in a multicorpus configuration file.

Process configuration

Example for LexRank_MMR configuration file :

<CONFIG>
<TASK ID="1">
<LANGUAGE>english</LANGUAGE>
<OUTPUT_PATH>doc/output</OUTPUT_PATH>
<MULTITHREADING>true</MULTITHREADING> <PREPROCESS NAME="GenerateTextModel">
<OPTION NAME="StanfordNLP">true</OPTION>
<OPTION NAME="StopWordListFile">$CORPUS_DATA/stopwords/englishStopWords.txt</OPTION>
</PREPROCESS>
<PROCESS>
<OPTION NAME="CorpusIdToSummarize">all</OPTION>
<OPTION NAME="ReadStopWords">false</OPTION>
<INDEX_BUILDER NAME="TF_IDF.TF_IDF">
</INDEX_BUILDER>
<CARACTERISTIC_BUILDER NAME="vector.TfIdfVectorSentence">
</CARACTERISTIC_BUILDER>
<SCORING_METHOD NAME="graphBased.LexRank">
<OPTION NAME="DampingParameter">0.15</OPTION>
<OPTION NAME="GraphThreshold">0.1</OPTION>
<OPTION NAME="SimilarityMethod">JaccardSimilarity</OPTION>
</SCORING_METHOD>
<SUMMARIZE_METHOD NAME="MMR">
<OPTION NAME="CharLimitBoolean">true</OPTION>
<OPTION NAME="Size">200</OPTION>
<OPTION NAME="SimilarityMethod">JaccardSimilarity</OPTION>
<OPTION NAME="Lambda">0.6</OPTION>
</SUMMARIZE_METHOD>
</PROCESS>
<ROUGE_EVALUATION>
<ROUGE_MEASURE>ROUGE-1	ROUGE-2	ROUGE-SU4</ROUGE_MEASURE>
<MODEL_ROOT>models</MODEL_ROOT>
<PEER_ROOT>systems</PEER_ROOT>
</ROUGE_EVALUATION>
</TASK>
</CONFIG>
  • <CONFIG> is the root node.
    • <TASK> represent a summarization task. You could do multiple in a simple run. At start, stick with one.
      • <LANGUAGE> is the input's document language for preprocessing goal. (english / french for now)
      • <OUTPUT_PATH> is the forlder's output path of the system. It is used to save preprocessed documents, ROUGE xml generated file, old score, ...
      • <MULTITHREADING> (boolean) launch the system in a mutltithreading way or not.
      • <PREPROCESS> is the preprocess step for the system. The preprocess java class to use is pass by the name variable. Here it's GenerateTextModel. It also needs two <OPTION> :
        • <OPTION NAME="StanfordNLP"> (boolean), true if you want to use StanfordNLP pipeline and tool to do the preprocessing.
        • <OPTION NAME="StopWordListPath"> (String), path of the stopwords list you want to use.
      • <PROCESS> is the main step of the system. It should have at least one <SUMMARIZE_METHOD> node and two <OPTION> It often has an <INDEX_BUILDER> node and a <CARACTERISTIC_BUILDER> node :
        • <OPTION NAME="CorpusIdToSummarize"> (String as a list of int separated by \t), the list of CorpusId to summarize from the MultiCorpus configuration file. "all" will do summarization for all corpus.
        • <OPTION NAME="ReadStopWords"> (boolean), state if the system count stopwords as part of the texts or not.
        • <INDEX_BUILDER> is the step where the system generate a computer friendly representation of each text's textual unit. (TF-IDF, Bigram, WordEmbeddings, ...)
        • <CARATERISTIC_BUILDER> is the sentence caracteristic generation step based on the textual unit index building.
        • <SCORING_METHOD> weights each sentences.
        • <SUMMARIZE_BUILDER> generate a summary usually by ranking sentence based on their score.
      • <ROUGE_EVALUATION> is the ROUGE evaluation step. For detail, look at ROUGE readme in /lib/ROUGE folder.
        • <ROUGE_MEASURE> (String as a list of int separated by \t), represent the list of ROUGE measure you want to use.
        • <MODEL_ROOT> is the model's folder name for ROUGE xml input files.
        • <PEER_ROOT> is the peer's folder name for ROUGE xml input files.

The <PROCESS> node is the system's core and you should look for more detail in the javadoc and the source code of the different INDEX_BUILDER, CARACTERISTIC_BUILDER, SCORING_METHOD and SUMMARIZE_METHOD class.

Multicorpus configuration

<?xml version="1.0" encoding="UTF-8"?>
<CONFIG>
<TASK ID="1">
<MULTICORPUS ID="0">
<CORPUS ID="0">
<INPUT_PATH>$CORPUS_DATA/TAC2009/UpdateSumm09_test_docs_files/D0901A/D0901A-A</INPUT_PATH>
<DOCUMENT ID="0">.*</DOCUMENT>
<SUMMARY_PATH>$CORPUS_DATA/TAC2009/UpdateSumm09_eval/ROUGE/models</SUMMARY_PATH>
<SUMMARY ID="0">D0901-A.*</SUMMARY>
</CORPUS>
</MULTICORPUS>
</TASK>
</CONFIG>

For now, all ID are useless and could be avoided.

  • <CONFIG> is the root node.
    • <TASK> represent a summarization task. You could do multiple in a simple run. At start, stick with one.
      • <MULTICORPUS> is a list of <CORPUS>
        • <CORPUS> can be one or more documents. The system will generate one summary per corpus.
          • <INPUT_PATH> is the folder containing the corpus' documents.
          • <DOCUMENT> is the regex for the documents you want to load. You could use multiple <DOCUMENT> node.
          • <SUMMARY_PATH> is the human summaries folder path.
          • <SUMMARY> is the regex for human summary file associating to this corpus. You could use multiple <SUMMARY> node.

Built With

  • Maven - Dependency Management

Authors

License

This project is licensed under the GPL3 License - see the LICENSE.md file for details

Acknowledgments

  • Thanks to Aurélien Bossard, my PhD supervisor.

About

MOTS (MOdular Tool for Summarization) is a summarization system, written in Java. It is as modular as possible, and is intended to provide an architecture to implement and test new summarization methods, as well as to ease comparison with already implemented methods, in an unified framework.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Build StatusDocsWikiLicense: GPL v3

MOTS

MOTS (MOdular Tool for Summarization) is a summarization system, written in Java. It is as modular as possible, and is intended to provide an architecture to implement and test new summarization methods, as well as to ease comparison with already implemented methods, in an unified framework. This system is the first completely modular system for automatic summarization and already allows to summarize using more than a hundred combinations of modules. The need for such a system is important. Indeed, several evaluation campaigns exist in AS field, but summarization algorithms are not easy to compare due to the large variety of pre and post-processings they use.

Getting Started

  • Javadoc available.
  • We provide an example corpus from TAC2009 in /src/main/resources and its associated human summaries.

Prerequisites

  • Maven - Dependency Manager
  • glpk-utils in order to use ILP (sudo apt-get install glpk-utils)
  • At least python3 in order to use WordEmbeddings, make sure to have an updated version of pip3 (sudo pip3 install --upgrade pip)
  • gensim in order to use WordEmbeddings (pip3 install gensim --user)
  • jep in order to use WordEmbeddings (pip3 install jep --user)

Installing

  • You might define $CORPUS_DATA to your DUC/TAC folder.
  • Install ROUGE :
    • Define $ROUGE_HOME to your ROUGE installation folder (or to ./ROUGE-1.5.5/RELEASE-1.5.5).
    • Install XML::DOM module in order to use ROUGE perl script. (sudo cpan install XML::DOM).
    • Run ./rouge_install.sh or :
      • Define $ROUGE_EVAL_HOME to $ROUGE_HOME/data.
      • Recreate database :
       cd data/WordNet-2.0-Exceptions/
      rm WordNet-2.0.exc.db # only if exist
      perl buildExeptionDB.pl . exc WordNet-2.0.exc.db
      cd ..
      rm WordNet-2.0.exc.db # only if exist
      ln -s WordNet-2.0-Exceptions/WordNet-2.0.exc.db WordNet-2.0.exc.db
      
  • Run install.sh script.

Usage

MOTS is a command line tool than can be used like this :

./MOTS mots.X.Y.Z.jar -c <config_file> -m <multicorpus_file> -v <OPTIONAL>

MOTS script encapsulate some environnement variable needed for the execution of WordEmbeddings. If you don't use WordEmbeddings you could launch via :

java -jar mots.X.Y.Z.jar -c <config_file> -m <multicorpus_file> -v <OPTIONAL>

Example config file and multicorpus file are provided in /conf but should be adapted to your setup.

Go Deeper

Each summarization process is defined in a configuration file and the test corpus is defined in a multicorpus configuration file.

Process configuration

Example for LexRank_MMR configuration file :

<CONFIG>
<TASK ID="1">
<LANGUAGE>english</LANGUAGE>
<OUTPUT_PATH>doc/output</OUTPUT_PATH>
<MULTITHREADING>true</MULTITHREADING> <PREPROCESS NAME="GenerateTextModel">
<OPTION NAME="StanfordNLP">true</OPTION>
<OPTION NAME="StopWordListFile">$CORPUS_DATA/stopwords/englishStopWords.txt</OPTION>
</PREPROCESS>
<PROCESS>
<OPTION NAME="CorpusIdToSummarize">all</OPTION>
<OPTION NAME="ReadStopWords">false</OPTION>
<INDEX_BUILDER NAME="TF_IDF.TF_IDF">
</INDEX_BUILDER>
<CARACTERISTIC_BUILDER NAME="vector.TfIdfVectorSentence">
</CARACTERISTIC_BUILDER>
<SCORING_METHOD NAME="graphBased.LexRank">
<OPTION NAME="DampingParameter">0.15</OPTION>
<OPTION NAME="GraphThreshold">0.1</OPTION>
<OPTION NAME="SimilarityMethod">JaccardSimilarity</OPTION>
</SCORING_METHOD>
<SUMMARIZE_METHOD NAME="MMR">
<OPTION NAME="CharLimitBoolean">true</OPTION>
<OPTION NAME="Size">200</OPTION>
<OPTION NAME="SimilarityMethod">JaccardSimilarity</OPTION>
<OPTION NAME="Lambda">0.6</OPTION>
</SUMMARIZE_METHOD>
</PROCESS>
<ROUGE_EVALUATION>
<ROUGE_MEASURE>ROUGE-1	ROUGE-2	ROUGE-SU4</ROUGE_MEASURE>
<MODEL_ROOT>models</MODEL_ROOT>
<PEER_ROOT>systems</PEER_ROOT>
</ROUGE_EVALUATION>
</TASK>
</CONFIG>
  • <CONFIG> is the root node.
    • <TASK> represent a summarization task. You could do multiple in a simple run. At start, stick with one.
      • <LANGUAGE> is the input's document language for preprocessing goal. (english / french for now)
      • <OUTPUT_PATH> is the forlder's output path of the system. It is used to save preprocessed documents, ROUGE xml generated file, old score, ...
      • <MULTITHREADING> (boolean) launch the system in a mutltithreading way or not.
      • <PREPROCESS> is the preprocess step for the system. The preprocess java class to use is pass by the name variable. Here it's GenerateTextModel. It also needs two <OPTION> :
        • <OPTION NAME="StanfordNLP"> (boolean), true if you want to use StanfordNLP pipeline and tool to do the preprocessing.
        • <OPTION NAME="StopWordListPath"> (String), path of the stopwords list you want to use.
      • <PROCESS> is the main step of the system. It should have at least one <SUMMARIZE_METHOD> node and two <OPTION> It often has an <INDEX_BUILDER> node and a <CARACTERISTIC_BUILDER> node :
        • <OPTION NAME="CorpusIdToSummarize"> (String as a list of int separated by \t), the list of CorpusId to summarize from the MultiCorpus configuration file. "all" will do summarization for all corpus.
        • <OPTION NAME="ReadStopWords"> (boolean), state if the system count stopwords as part of the texts or not.
        • <INDEX_BUILDER> is the step where the system generate a computer friendly representation of each text's textual unit. (TF-IDF, Bigram, WordEmbeddings, ...)
        • <CARATERISTIC_BUILDER> is the sentence caracteristic generation step based on the textual unit index building.
        • <SCORING_METHOD> weights each sentences.
        • <SUMMARIZE_BUILDER> generate a summary usually by ranking sentence based on their score.
      • <ROUGE_EVALUATION> is the ROUGE evaluation step. For detail, look at ROUGE readme in /lib/ROUGE folder.
        • <ROUGE_MEASURE> (String as a list of int separated by \t), represent the list of ROUGE measure you want to use.
        • <MODEL_ROOT> is the model's folder name for ROUGE xml input files.
        • <PEER_ROOT> is the peer's folder name for ROUGE xml input files.

The <PROCESS> node is the system's core and you should look for more detail in the javadoc and the source code of the different INDEX_BUILDER, CARACTERISTIC_BUILDER, SCORING_METHOD and SUMMARIZE_METHOD class.

Multicorpus configuration

<?xml version="1.0" encoding="UTF-8"?>
<CONFIG>
<TASK ID="1">
<MULTICORPUS ID="0">
<CORPUS ID="0">
<INPUT_PATH>$CORPUS_DATA/TAC2009/UpdateSumm09_test_docs_files/D0901A/D0901A-A</INPUT_PATH>
<DOCUMENT ID="0">.*</DOCUMENT>
<SUMMARY_PATH>$CORPUS_DATA/TAC2009/UpdateSumm09_eval/ROUGE/models</SUMMARY_PATH>
<SUMMARY ID="0">D0901-A.*</SUMMARY>
</CORPUS>
</MULTICORPUS>
</TASK>
</CONFIG>

For now, all ID are useless and could be avoided.

  • <CONFIG> is the root node.
    • <TASK> represent a summarization task. You could do multiple in a simple run. At start, stick with one.
      • <MULTICORPUS> is a list of <CORPUS>
        • <CORPUS> can be one or more documents. The system will generate one summary per corpus.
          • <INPUT_PATH> is the folder containing the corpus' documents.
          • <DOCUMENT> is the regex for the documents you want to load. You could use multiple <DOCUMENT> node.
          • <SUMMARY_PATH> is the human summaries folder path.
          • <SUMMARY> is the regex for human summary file associating to this corpus. You could use multiple <SUMMARY> node.

Built With

  • Maven - Dependency Management

Authors

License

This project is licensed under the GPL3 License - see the LICENSE.md file for details

Acknowledgments

  • Thanks to Aurélien Bossard, my PhD supervisor.

About

MOTS (MOdular Tool for Summarization) is a summarization system, written in Java. It is as modular as possible, and is intended to provide an architecture to implement and test new summarization methods, as well as to ease comparison with already implemented methods, in an unified framework.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Build StatusDocsWikiLicense: GPL v3

MOTS

MOTS (MOdular Tool for Summarization) is a summarization system, written in Java. It is as modular as possible, and is intended to provide an architecture to implement and test new summarization methods, as well as to ease comparison with already implemented methods, in an unified framework. This system is the first completely modular system for automatic summarization and already allows to summarize using more than a hundred combinations of modules. The need for such a system is important. Indeed, several evaluation campaigns exist in AS field, but summarization algorithms are not easy to compare due to the large variety of pre and post-processings they use.

Getting Started

  • Javadoc available.
  • We provide an example corpus from TAC2009 in /src/main/resources and its associated human summaries.

Prerequisites

  • Maven - Dependency Manager
  • glpk-utils in order to use ILP (sudo apt-get install glpk-utils)
  • At least python3 in order to use WordEmbeddings, make sure to have an updated version of pip3 (sudo pip3 install --upgrade pip)
  • gensim in order to use WordEmbeddings (pip3 install gensim --user)
  • jep in order to use WordEmbeddings (pip3 install jep --user)

Installing

  • You might define $CORPUS_DATA to your DUC/TAC folder.
  • Install ROUGE :
    • Define $ROUGE_HOME to your ROUGE installation folder (or to ./ROUGE-1.5.5/RELEASE-1.5.5).
    • Install XML::DOM module in order to use ROUGE perl script. (sudo cpan install XML::DOM).
    • Run ./rouge_install.sh or :
      • Define $ROUGE_EVAL_HOME to $ROUGE_HOME/data.
      • Recreate database :
       cd data/WordNet-2.0-Exceptions/
      rm WordNet-2.0.exc.db # only if exist
      perl buildExeptionDB.pl . exc WordNet-2.0.exc.db
      cd ..
      rm WordNet-2.0.exc.db # only if exist
      ln -s WordNet-2.0-Exceptions/WordNet-2.0.exc.db WordNet-2.0.exc.db
      
  • Run install.sh script.

Usage

MOTS is a command line tool than can be used like this :

./MOTS mots.X.Y.Z.jar -c <config_file> -m <multicorpus_file> -v <OPTIONAL>

MOTS script encapsulate some environnement variable needed for the execution of WordEmbeddings. If you don't use WordEmbeddings you could launch via :

java -jar mots.X.Y.Z.jar -c <config_file> -m <multicorpus_file> -v <OPTIONAL>

Example config file and multicorpus file are provided in /conf but should be adapted to your setup.

Go Deeper

Each summarization process is defined in a configuration file and the test corpus is defined in a multicorpus configuration file.

Process configuration

Example for LexRank_MMR configuration file :

<CONFIG>
<TASK ID="1">
<LANGUAGE>english</LANGUAGE>
<OUTPUT_PATH>doc/output</OUTPUT_PATH>
<MULTITHREADING>true</MULTITHREADING> <PREPROCESS NAME="GenerateTextModel">
<OPTION NAME="StanfordNLP">true</OPTION>
<OPTION NAME="StopWordListFile">$CORPUS_DATA/stopwords/englishStopWords.txt</OPTION>
</PREPROCESS>
<PROCESS>
<OPTION NAME="CorpusIdToSummarize">all</OPTION>
<OPTION NAME="ReadStopWords">false</OPTION>
<INDEX_BUILDER NAME="TF_IDF.TF_IDF">
</INDEX_BUILDER>
<CARACTERISTIC_BUILDER NAME="vector.TfIdfVectorSentence">
</CARACTERISTIC_BUILDER>
<SCORING_METHOD NAME="graphBased.LexRank">
<OPTION NAME="DampingParameter">0.15</OPTION>
<OPTION NAME="GraphThreshold">0.1</OPTION>
<OPTION NAME="SimilarityMethod">JaccardSimilarity</OPTION>
</SCORING_METHOD>
<SUMMARIZE_METHOD NAME="MMR">
<OPTION NAME="CharLimitBoolean">true</OPTION>
<OPTION NAME="Size">200</OPTION>
<OPTION NAME="SimilarityMethod">JaccardSimilarity</OPTION>
<OPTION NAME="Lambda">0.6</OPTION>
</SUMMARIZE_METHOD>
</PROCESS>
<ROUGE_EVALUATION>
<ROUGE_MEASURE>ROUGE-1	ROUGE-2	ROUGE-SU4</ROUGE_MEASURE>
<MODEL_ROOT>models</MODEL_ROOT>
<PEER_ROOT>systems</PEER_ROOT>
</ROUGE_EVALUATION>
</TASK>
</CONFIG>
  • <CONFIG> is the root node.
    • <TASK> represent a summarization task. You could do multiple in a simple run. At start, stick with one.
      • <LANGUAGE> is the input's document language for preprocessing goal. (english / french for now)
      • <OUTPUT_PATH> is the forlder's output path of the system. It is used to save preprocessed documents, ROUGE xml generated file, old score, ...
      • <MULTITHREADING> (boolean) launch the system in a mutltithreading way or not.
      • <PREPROCESS> is the preprocess step for the system. The preprocess java class to use is pass by the name variable. Here it's GenerateTextModel. It also needs two <OPTION> :
        • <OPTION NAME="StanfordNLP"> (boolean), true if you want to use StanfordNLP pipeline and tool to do the preprocessing.
        • <OPTION NAME="StopWordListPath"> (String), path of the stopwords list you want to use.
      • <PROCESS> is the main step of the system. It should have at least one <SUMMARIZE_METHOD> node and two <OPTION> It often has an <INDEX_BUILDER> node and a <CARACTERISTIC_BUILDER> node :
        • <OPTION NAME="CorpusIdToSummarize"> (String as a list of int separated by \t), the list of CorpusId to summarize from the MultiCorpus configuration file. "all" will do summarization for all corpus.
        • <OPTION NAME="ReadStopWords"> (boolean), state if the system count stopwords as part of the texts or not.
        • <INDEX_BUILDER> is the step where the system generate a computer friendly representation of each text's textual unit. (TF-IDF, Bigram, WordEmbeddings, ...)
        • <CARATERISTIC_BUILDER> is the sentence caracteristic generation step based on the textual unit index building.
        • <SCORING_METHOD> weights each sentences.
        • <SUMMARIZE_BUILDER> generate a summary usually by ranking sentence based on their score.
      • <ROUGE_EVALUATION> is the ROUGE evaluation step. For detail, look at ROUGE readme in /lib/ROUGE folder.
        • <ROUGE_MEASURE> (String as a list of int separated by \t), represent the list of ROUGE measure you want to use.
        • <MODEL_ROOT> is the model's folder name for ROUGE xml input files.
        • <PEER_ROOT> is the peer's folder name for ROUGE xml input files.

The <PROCESS> node is the system's core and you should look for more detail in the javadoc and the source code of the different INDEX_BUILDER, CARACTERISTIC_BUILDER, SCORING_METHOD and SUMMARIZE_METHOD class.

Multicorpus configuration

<?xml version="1.0" encoding="UTF-8"?>
<CONFIG>
<TASK ID="1">
<MULTICORPUS ID="0">
<CORPUS ID="0">
<INPUT_PATH>$CORPUS_DATA/TAC2009/UpdateSumm09_test_docs_files/D0901A/D0901A-A</INPUT_PATH>
<DOCUMENT ID="0">.*</DOCUMENT>
<SUMMARY_PATH>$CORPUS_DATA/TAC2009/UpdateSumm09_eval/ROUGE/models</SUMMARY_PATH>
<SUMMARY ID="0">D0901-A.*</SUMMARY>
</CORPUS>
</MULTICORPUS>
</TASK>
</CONFIG>

For now, all ID are useless and could be avoided.

  • <CONFIG> is the root node.
    • <TASK> represent a summarization task. You could do multiple in a simple run. At start, stick with one.
      • <MULTICORPUS> is a list of <CORPUS>
        • <CORPUS> can be one or more documents. The system will generate one summary per corpus.
          • <INPUT_PATH> is the folder containing the corpus' documents.
          • <DOCUMENT> is the regex for the documents you want to load. You could use multiple <DOCUMENT> node.
          • <SUMMARY_PATH> is the human summaries folder path.
          • <SUMMARY> is the regex for human summary file associating to this corpus. You could use multiple <SUMMARY> node.

Built With

  • Maven - Dependency Management

Authors

License

This project is licensed under the GPL3 License - see the LICENSE.md file for details

Acknowledgments

  • Thanks to Aurélien Bossard, my PhD supervisor.

About

MOTS (MOdular Tool for Summarization) is a summarization system, written in Java. It is as modular as possible, and is intended to provide an architecture to implement and test new summarization methods, as well as to ease comparison with already implemented methods, in an unified framework.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Build StatusDocsWikiLicense: GPL v3

MOTS

MOTS (MOdular Tool for Summarization) is a summarization system, written in Java. It is as modular as possible, and is intended to provide an architecture to implement and test new summarization methods, as well as to ease comparison with already implemented methods, in an unified framework. This system is the first completely modular system for automatic summarization and already allows to summarize using more than a hundred combinations of modules. The need for such a system is important. Indeed, several evaluation campaigns exist in AS field, but summarization algorithms are not easy to compare due to the large variety of pre and post-processings they use.

Getting Started

  • Javadoc available.
  • We provide an example corpus from TAC2009 in /src/main/resources and its associated human summaries.

Prerequisites

  • Maven - Dependency Manager
  • glpk-utils in order to use ILP (sudo apt-get install glpk-utils)
  • At least python3 in order to use WordEmbeddings, make sure to have an updated version of pip3 (sudo pip3 install --upgrade pip)
  • gensim in order to use WordEmbeddings (pip3 install gensim --user)
  • jep in order to use WordEmbeddings (pip3 install jep --user)

Installing

  • You might define $CORPUS_DATA to your DUC/TAC folder.
  • Install ROUGE :
    • Define $ROUGE_HOME to your ROUGE installation folder (or to ./ROUGE-1.5.5/RELEASE-1.5.5).
    • Install XML::DOM module in order to use ROUGE perl script. (sudo cpan install XML::DOM).
    • Run ./rouge_install.sh or :
      • Define $ROUGE_EVAL_HOME to $ROUGE_HOME/data.
      • Recreate database :
       cd data/WordNet-2.0-Exceptions/
      rm WordNet-2.0.exc.db # only if exist
      perl buildExeptionDB.pl . exc WordNet-2.0.exc.db
      cd ..
      rm WordNet-2.0.exc.db # only if exist
      ln -s WordNet-2.0-Exceptions/WordNet-2.0.exc.db WordNet-2.0.exc.db
      
  • Run install.sh script.

Usage

MOTS is a command line tool than can be used like this :

./MOTS mots.X.Y.Z.jar -c <config_file> -m <multicorpus_file> -v <OPTIONAL>

MOTS script encapsulate some environnement variable needed for the execution of WordEmbeddings. If you don't use WordEmbeddings you could launch via :

java -jar mots.X.Y.Z.jar -c <config_file> -m <multicorpus_file> -v <OPTIONAL>

Example config file and multicorpus file are provided in /conf but should be adapted to your setup.

Go Deeper

Each summarization process is defined in a configuration file and the test corpus is defined in a multicorpus configuration file.

Process configuration

Example for LexRank_MMR configuration file :

<CONFIG>
<TASK ID="1">
<LANGUAGE>english</LANGUAGE>
<OUTPUT_PATH>doc/output</OUTPUT_PATH>
<MULTITHREADING>true</MULTITHREADING> <PREPROCESS NAME="GenerateTextModel">
<OPTION NAME="StanfordNLP">true</OPTION>
<OPTION NAME="StopWordListFile">$CORPUS_DATA/stopwords/englishStopWords.txt</OPTION>
</PREPROCESS>
<PROCESS>
<OPTION NAME="CorpusIdToSummarize">all</OPTION>
<OPTION NAME="ReadStopWords">false</OPTION>
<INDEX_BUILDER NAME="TF_IDF.TF_IDF">
</INDEX_BUILDER>
<CARACTERISTIC_BUILDER NAME="vector.TfIdfVectorSentence">
</CARACTERISTIC_BUILDER>
<SCORING_METHOD NAME="graphBased.LexRank">
<OPTION NAME="DampingParameter">0.15</OPTION>
<OPTION NAME="GraphThreshold">0.1</OPTION>
<OPTION NAME="SimilarityMethod">JaccardSimilarity</OPTION>
</SCORING_METHOD>
<SUMMARIZE_METHOD NAME="MMR">
<OPTION NAME="CharLimitBoolean">true</OPTION>
<OPTION NAME="Size">200</OPTION>
<OPTION NAME="SimilarityMethod">JaccardSimilarity</OPTION>
<OPTION NAME="Lambda">0.6</OPTION>
</SUMMARIZE_METHOD>
</PROCESS>
<ROUGE_EVALUATION>
<ROUGE_MEASURE>ROUGE-1	ROUGE-2	ROUGE-SU4</ROUGE_MEASURE>
<MODEL_ROOT>models</MODEL_ROOT>
<PEER_ROOT>systems</PEER_ROOT>
</ROUGE_EVALUATION>
</TASK>
</CONFIG>
  • <CONFIG> is the root node.
    • <TASK> represent a summarization task. You could do multiple in a simple run. At start, stick with one.
      • <LANGUAGE> is the input's document language for preprocessing goal. (english / french for now)
      • <OUTPUT_PATH> is the forlder's output path of the system. It is used to save preprocessed documents, ROUGE xml generated file, old score, ...
      • <MULTITHREADING> (boolean) launch the system in a mutltithreading way or not.
      • <PREPROCESS> is the preprocess step for the system. The preprocess java class to use is pass by the name variable. Here it's GenerateTextModel. It also needs two <OPTION> :
        • <OPTION NAME="StanfordNLP"> (boolean), true if you want to use StanfordNLP pipeline and tool to do the preprocessing.
        • <OPTION NAME="StopWordListPath"> (String), path of the stopwords list you want to use.
      • <PROCESS> is the main step of the system. It should have at least one <SUMMARIZE_METHOD> node and two <OPTION> It often has an <INDEX_BUILDER> node and a <CARACTERISTIC_BUILDER> node :
        • <OPTION NAME="CorpusIdToSummarize"> (String as a list of int separated by \t), the list of CorpusId to summarize from the MultiCorpus configuration file. "all" will do summarization for all corpus.
        • <OPTION NAME="ReadStopWords"> (boolean), state if the system count stopwords as part of the texts or not.
        • <INDEX_BUILDER> is the step where the system generate a computer friendly representation of each text's textual unit. (TF-IDF, Bigram, WordEmbeddings, ...)
        • <CARATERISTIC_BUILDER> is the sentence caracteristic generation step based on the textual unit index building.
        • <SCORING_METHOD> weights each sentences.
        • <SUMMARIZE_BUILDER> generate a summary usually by ranking sentence based on their score.
      • <ROUGE_EVALUATION> is the ROUGE evaluation step. For detail, look at ROUGE readme in /lib/ROUGE folder.
        • <ROUGE_MEASURE> (String as a list of int separated by \t), represent the list of ROUGE measure you want to use.
        • <MODEL_ROOT> is the model's folder name for ROUGE xml input files.
        • <PEER_ROOT> is the peer's folder name for ROUGE xml input files.

The <PROCESS> node is the system's core and you should look for more detail in the javadoc and the source code of the different INDEX_BUILDER, CARACTERISTIC_BUILDER, SCORING_METHOD and SUMMARIZE_METHOD class.

Multicorpus configuration

<?xml version="1.0" encoding="UTF-8"?>
<CONFIG>
<TASK ID="1">
<MULTICORPUS ID="0">
<CORPUS ID="0">
<INPUT_PATH>$CORPUS_DATA/TAC2009/UpdateSumm09_test_docs_files/D0901A/D0901A-A</INPUT_PATH>
<DOCUMENT ID="0">.*</DOCUMENT>
<SUMMARY_PATH>$CORPUS_DATA/TAC2009/UpdateSumm09_eval/ROUGE/models</SUMMARY_PATH>
<SUMMARY ID="0">D0901-A.*</SUMMARY>
</CORPUS>
</MULTICORPUS>
</TASK>
</CONFIG>

For now, all ID are useless and could be avoided.

  • <CONFIG> is the root node.
    • <TASK> represent a summarization task. You could do multiple in a simple run. At start, stick with one.
      • <MULTICORPUS> is a list of <CORPUS>
        • <CORPUS> can be one or more documents. The system will generate one summary per corpus.
          • <INPUT_PATH> is the folder containing the corpus' documents.
          • <DOCUMENT> is the regex for the documents you want to load. You could use multiple <DOCUMENT> node.
          • <SUMMARY_PATH> is the human summaries folder path.
          • <SUMMARY> is the regex for human summary file associating to this corpus. You could use multiple <SUMMARY> node.

Built With

  • Maven - Dependency Management

Authors

License

This project is licensed under the GPL3 License - see the LICENSE.md file for details

Acknowledgments

  • Thanks to Aurélien Bossard, my PhD supervisor.

About

MOTS (MOdular Tool for Summarization) is a summarization system, written in Java. It is as modular as possible, and is intended to provide an architecture to implement and test new summarization methods, as well as to ease comparison with already implemented methods, in an unified framework.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Build StatusDocsWikiLicense: GPL v3

MOTS

MOTS (MOdular Tool for Summarization) is a summarization system, written in Java. It is as modular as possible, and is intended to provide an architecture to implement and test new summarization methods, as well as to ease comparison with already implemented methods, in an unified framework. This system is the first completely modular system for automatic summarization and already allows to summarize using more than a hundred combinations of modules. The need for such a system is important. Indeed, several evaluation campaigns exist in AS field, but summarization algorithms are not easy to compare due to the large variety of pre and post-processings they use.

Getting Started

  • Javadoc available.
  • We provide an example corpus from TAC2009 in /src/main/resources and its associated human summaries.

Prerequisites

  • Maven - Dependency Manager
  • glpk-utils in order to use ILP (sudo apt-get install glpk-utils)
  • At least python3 in order to use WordEmbeddings, make sure to have an updated version of pip3 (sudo pip3 install --upgrade pip)
  • gensim in order to use WordEmbeddings (pip3 install gensim --user)
  • jep in order to use WordEmbeddings (pip3 install jep --user)

Installing

  • You might define $CORPUS_DATA to your DUC/TAC folder.
  • Install ROUGE :
    • Define $ROUGE_HOME to your ROUGE installation folder (or to ./ROUGE-1.5.5/RELEASE-1.5.5).
    • Install XML::DOM module in order to use ROUGE perl script. (sudo cpan install XML::DOM).
    • Run ./rouge_install.sh or :
      • Define $ROUGE_EVAL_HOME to $ROUGE_HOME/data.
      • Recreate database :
       cd data/WordNet-2.0-Exceptions/
      rm WordNet-2.0.exc.db # only if exist
      perl buildExeptionDB.pl . exc WordNet-2.0.exc.db
      cd ..
      rm WordNet-2.0.exc.db # only if exist
      ln -s WordNet-2.0-Exceptions/WordNet-2.0.exc.db WordNet-2.0.exc.db
      
  • Run install.sh script.

Usage

MOTS is a command line tool than can be used like this :

./MOTS mots.X.Y.Z.jar -c <config_file> -m <multicorpus_file> -v <OPTIONAL>

MOTS script encapsulate some environnement variable needed for the execution of WordEmbeddings. If you don't use WordEmbeddings you could launch via :

java -jar mots.X.Y.Z.jar -c <config_file> -m <multicorpus_file> -v <OPTIONAL>

Example config file and multicorpus file are provided in /conf but should be adapted to your setup.

Go Deeper

Each summarization process is defined in a configuration file and the test corpus is defined in a multicorpus configuration file.

Process configuration

Example for LexRank_MMR configuration file :

<CONFIG>
<TASK ID="1">
<LANGUAGE>english</LANGUAGE>
<OUTPUT_PATH>doc/output</OUTPUT_PATH>
<MULTITHREADING>true</MULTITHREADING> <PREPROCESS NAME="GenerateTextModel">
<OPTION NAME="StanfordNLP">true</OPTION>
<OPTION NAME="StopWordListFile">$CORPUS_DATA/stopwords/englishStopWords.txt</OPTION>
</PREPROCESS>
<PROCESS>
<OPTION NAME="CorpusIdToSummarize">all</OPTION>
<OPTION NAME="ReadStopWords">false</OPTION>
<INDEX_BUILDER NAME="TF_IDF.TF_IDF">
</INDEX_BUILDER>
<CARACTERISTIC_BUILDER NAME="vector.TfIdfVectorSentence">
</CARACTERISTIC_BUILDER>
<SCORING_METHOD NAME="graphBased.LexRank">
<OPTION NAME="DampingParameter">0.15</OPTION>
<OPTION NAME="GraphThreshold">0.1</OPTION>
<OPTION NAME="SimilarityMethod">JaccardSimilarity</OPTION>
</SCORING_METHOD>
<SUMMARIZE_METHOD NAME="MMR">
<OPTION NAME="CharLimitBoolean">true</OPTION>
<OPTION NAME="Size">200</OPTION>
<OPTION NAME="SimilarityMethod">JaccardSimilarity</OPTION>
<OPTION NAME="Lambda">0.6</OPTION>
</SUMMARIZE_METHOD>
</PROCESS>
<ROUGE_EVALUATION>
<ROUGE_MEASURE>ROUGE-1	ROUGE-2	ROUGE-SU4</ROUGE_MEASURE>
<MODEL_ROOT>models</MODEL_ROOT>
<PEER_ROOT>systems</PEER_ROOT>
</ROUGE_EVALUATION>
</TASK>
</CONFIG>
  • <CONFIG> is the root node.
    • <TASK> represent a summarization task. You could do multiple in a simple run. At start, stick with one.
      • <LANGUAGE> is the input's document language for preprocessing goal. (english / french for now)
      • <OUTPUT_PATH> is the forlder's output path of the system. It is used to save preprocessed documents, ROUGE xml generated file, old score, ...
      • <MULTITHREADING> (boolean) launch the system in a mutltithreading way or not.
      • <PREPROCESS> is the preprocess step for the system. The preprocess java class to use is pass by the name variable. Here it's GenerateTextModel. It also needs two <OPTION> :
        • <OPTION NAME="StanfordNLP"> (boolean), true if you want to use StanfordNLP pipeline and tool to do the preprocessing.
        • <OPTION NAME="StopWordListPath"> (String), path of the stopwords list you want to use.
      • <PROCESS> is the main step of the system. It should have at least one <SUMMARIZE_METHOD> node and two <OPTION> It often has an <INDEX_BUILDER> node and a <CARACTERISTIC_BUILDER> node :
        • <OPTION NAME="CorpusIdToSummarize"> (String as a list of int separated by \t), the list of CorpusId to summarize from the MultiCorpus configuration file. "all" will do summarization for all corpus.
        • <OPTION NAME="ReadStopWords"> (boolean), state if the system count stopwords as part of the texts or not.
        • <INDEX_BUILDER> is the step where the system generate a computer friendly representation of each text's textual unit. (TF-IDF, Bigram, WordEmbeddings, ...)
        • <CARATERISTIC_BUILDER> is the sentence caracteristic generation step based on the textual unit index building.
        • <SCORING_METHOD> weights each sentences.
        • <SUMMARIZE_BUILDER> generate a summary usually by ranking sentence based on their score.
      • <ROUGE_EVALUATION> is the ROUGE evaluation step. For detail, look at ROUGE readme in /lib/ROUGE folder.
        • <ROUGE_MEASURE> (String as a list of int separated by \t), represent the list of ROUGE measure you want to use.
        • <MODEL_ROOT> is the model's folder name for ROUGE xml input files.
        • <PEER_ROOT> is the peer's folder name for ROUGE xml input files.

The <PROCESS> node is the system's core and you should look for more detail in the javadoc and the source code of the different INDEX_BUILDER, CARACTERISTIC_BUILDER, SCORING_METHOD and SUMMARIZE_METHOD class.

Multicorpus configuration

<?xml version="1.0" encoding="UTF-8"?>
<CONFIG>
<TASK ID="1">
<MULTICORPUS ID="0">
<CORPUS ID="0">
<INPUT_PATH>$CORPUS_DATA/TAC2009/UpdateSumm09_test_docs_files/D0901A/D0901A-A</INPUT_PATH>
<DOCUMENT ID="0">.*</DOCUMENT>
<SUMMARY_PATH>$CORPUS_DATA/TAC2009/UpdateSumm09_eval/ROUGE/models</SUMMARY_PATH>
<SUMMARY ID="0">D0901-A.*</SUMMARY>
</CORPUS>
</MULTICORPUS>
</TASK>
</CONFIG>

For now, all ID are useless and could be avoided.

  • <CONFIG> is the root node.
    • <TASK> represent a summarization task. You could do multiple in a simple run. At start, stick with one.
      • <MULTICORPUS> is a list of <CORPUS>
        • <CORPUS> can be one or more documents. The system will generate one summary per corpus.
          • <INPUT_PATH> is the folder containing the corpus' documents.
          • <DOCUMENT> is the regex for the documents you want to load. You could use multiple <DOCUMENT> node.
          • <SUMMARY_PATH> is the human summaries folder path.
          • <SUMMARY> is the regex for human summary file associating to this corpus. You could use multiple <SUMMARY> node.

Built With

  • Maven - Dependency Management

Authors

License

This project is licensed under the GPL3 License - see the LICENSE.md file for details

Acknowledgments

  • Thanks to Aurélien Bossard, my PhD supervisor.

About

MOTS (MOdular Tool for Summarization) is a summarization system, written in Java. It is as modular as possible, and is intended to provide an architecture to implement and test new summarization methods, as well as to ease comparison with already implemented methods, in an unified framework.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Build StatusDocsWikiLicense: GPL v3

MOTS

MOTS (MOdular Tool for Summarization) is a summarization system, written in Java. It is as modular as possible, and is intended to provide an architecture to implement and test new summarization methods, as well as to ease comparison with already implemented methods, in an unified framework. This system is the first completely modular system for automatic summarization and already allows to summarize using more than a hundred combinations of modules. The need for such a system is important. Indeed, several evaluation campaigns exist in AS field, but summarization algorithms are not easy to compare due to the large variety of pre and post-processings they use.

Getting Started

  • Javadoc available.
  • We provide an example corpus from TAC2009 in /src/main/resources and its associated human summaries.

Prerequisites

  • Maven - Dependency Manager
  • glpk-utils in order to use ILP (sudo apt-get install glpk-utils)
  • At least python3 in order to use WordEmbeddings, make sure to have an updated version of pip3 (sudo pip3 install --upgrade pip)
  • gensim in order to use WordEmbeddings (pip3 install gensim --user)
  • jep in order to use WordEmbeddings (pip3 install jep --user)

Installing

  • You might define $CORPUS_DATA to your DUC/TAC folder.
  • Install ROUGE :
    • Define $ROUGE_HOME to your ROUGE installation folder (or to ./ROUGE-1.5.5/RELEASE-1.5.5).
    • Install XML::DOM module in order to use ROUGE perl script. (sudo cpan install XML::DOM).
    • Run ./rouge_install.sh or :
      • Define $ROUGE_EVAL_HOME to $ROUGE_HOME/data.
      • Recreate database :
       cd data/WordNet-2.0-Exceptions/
      rm WordNet-2.0.exc.db # only if exist
      perl buildExeptionDB.pl . exc WordNet-2.0.exc.db
      cd ..
      rm WordNet-2.0.exc.db # only if exist
      ln -s WordNet-2.0-Exceptions/WordNet-2.0.exc.db WordNet-2.0.exc.db
      
  • Run install.sh script.

Usage

MOTS is a command line tool than can be used like this :

./MOTS mots.X.Y.Z.jar -c <config_file> -m <multicorpus_file> -v <OPTIONAL>

MOTS script encapsulate some environnement variable needed for the execution of WordEmbeddings. If you don't use WordEmbeddings you could launch via :

java -jar mots.X.Y.Z.jar -c <config_file> -m <multicorpus_file> -v <OPTIONAL>

Example config file and multicorpus file are provided in /conf but should be adapted to your setup.

Go Deeper

Each summarization process is defined in a configuration file and the test corpus is defined in a multicorpus configuration file.

Process configuration

Example for LexRank_MMR configuration file :

<CONFIG>
<TASK ID="1">
<LANGUAGE>english</LANGUAGE>
<OUTPUT_PATH>doc/output</OUTPUT_PATH>
<MULTITHREADING>true</MULTITHREADING> <PREPROCESS NAME="GenerateTextModel">
<OPTION NAME="StanfordNLP">true</OPTION>
<OPTION NAME="StopWordListFile">$CORPUS_DATA/stopwords/englishStopWords.txt</OPTION>
</PREPROCESS>
<PROCESS>
<OPTION NAME="CorpusIdToSummarize">all</OPTION>
<OPTION NAME="ReadStopWords">false</OPTION>
<INDEX_BUILDER NAME="TF_IDF.TF_IDF">
</INDEX_BUILDER>
<CARACTERISTIC_BUILDER NAME="vector.TfIdfVectorSentence">
</CARACTERISTIC_BUILDER>
<SCORING_METHOD NAME="graphBased.LexRank">
<OPTION NAME="DampingParameter">0.15</OPTION>
<OPTION NAME="GraphThreshold">0.1</OPTION>
<OPTION NAME="SimilarityMethod">JaccardSimilarity</OPTION>
</SCORING_METHOD>
<SUMMARIZE_METHOD NAME="MMR">
<OPTION NAME="CharLimitBoolean">true</OPTION>
<OPTION NAME="Size">200</OPTION>
<OPTION NAME="SimilarityMethod">JaccardSimilarity</OPTION>
<OPTION NAME="Lambda">0.6</OPTION>
</SUMMARIZE_METHOD>
</PROCESS>
<ROUGE_EVALUATION>
<ROUGE_MEASURE>ROUGE-1	ROUGE-2	ROUGE-SU4</ROUGE_MEASURE>
<MODEL_ROOT>models</MODEL_ROOT>
<PEER_ROOT>systems</PEER_ROOT>
</ROUGE_EVALUATION>
</TASK>
</CONFIG>
  • <CONFIG> is the root node.
    • <TASK> represent a summarization task. You could do multiple in a simple run. At start, stick with one.
      • <LANGUAGE> is the input's document language for preprocessing goal. (english / french for now)
      • <OUTPUT_PATH> is the forlder's output path of the system. It is used to save preprocessed documents, ROUGE xml generated file, old score, ...
      • <MULTITHREADING> (boolean) launch the system in a mutltithreading way or not.
      • <PREPROCESS> is the preprocess step for the system. The preprocess java class to use is pass by the name variable. Here it's GenerateTextModel. It also needs two <OPTION> :
        • <OPTION NAME="StanfordNLP"> (boolean), true if you want to use StanfordNLP pipeline and tool to do the preprocessing.
        • <OPTION NAME="StopWordListPath"> (String), path of the stopwords list you want to use.
      • <PROCESS> is the main step of the system. It should have at least one <SUMMARIZE_METHOD> node and two <OPTION> It often has an <INDEX_BUILDER> node and a <CARACTERISTIC_BUILDER> node :
        • <OPTION NAME="CorpusIdToSummarize"> (String as a list of int separated by \t), the list of CorpusId to summarize from the MultiCorpus configuration file. "all" will do summarization for all corpus.
        • <OPTION NAME="ReadStopWords"> (boolean), state if the system count stopwords as part of the texts or not.
        • <INDEX_BUILDER> is the step where the system generate a computer friendly representation of each text's textual unit. (TF-IDF, Bigram, WordEmbeddings, ...)
        • <CARATERISTIC_BUILDER> is the sentence caracteristic generation step based on the textual unit index building.
        • <SCORING_METHOD> weights each sentences.
        • <SUMMARIZE_BUILDER> generate a summary usually by ranking sentence based on their score.
      • <ROUGE_EVALUATION> is the ROUGE evaluation step. For detail, look at ROUGE readme in /lib/ROUGE folder.
        • <ROUGE_MEASURE> (String as a list of int separated by \t), represent the list of ROUGE measure you want to use.
        • <MODEL_ROOT> is the model's folder name for ROUGE xml input files.
        • <PEER_ROOT> is the peer's folder name for ROUGE xml input files.

The <PROCESS> node is the system's core and you should look for more detail in the javadoc and the source code of the different INDEX_BUILDER, CARACTERISTIC_BUILDER, SCORING_METHOD and SUMMARIZE_METHOD class.

Multicorpus configuration

<?xml version="1.0" encoding="UTF-8"?>
<CONFIG>
<TASK ID="1">
<MULTICORPUS ID="0">
<CORPUS ID="0">
<INPUT_PATH>$CORPUS_DATA/TAC2009/UpdateSumm09_test_docs_files/D0901A/D0901A-A</INPUT_PATH>
<DOCUMENT ID="0">.*</DOCUMENT>
<SUMMARY_PATH>$CORPUS_DATA/TAC2009/UpdateSumm09_eval/ROUGE/models</SUMMARY_PATH>
<SUMMARY ID="0">D0901-A.*</SUMMARY>
</CORPUS>
</MULTICORPUS>
</TASK>
</CONFIG>

For now, all ID are useless and could be avoided.

  • <CONFIG> is the root node.
    • <TASK> represent a summarization task. You could do multiple in a simple run. At start, stick with one.
      • <MULTICORPUS> is a list of <CORPUS>
        • <CORPUS> can be one or more documents. The system will generate one summary per corpus.
          • <INPUT_PATH> is the folder containing the corpus' documents.
          • <DOCUMENT> is the regex for the documents you want to load. You could use multiple <DOCUMENT> node.
          • <SUMMARY_PATH> is the human summaries folder path.
          • <SUMMARY> is the regex for human summary file associating to this corpus. You could use multiple <SUMMARY> node.

Built With

  • Maven - Dependency Management

Authors

License

This project is licensed under the GPL3 License - see the LICENSE.md file for details

Acknowledgments

  • Thanks to Aurélien Bossard, my PhD supervisor.

About

MOTS (MOdular Tool for Summarization) is a summarization system, written in Java. It is as modular as possible, and is intended to provide an architecture to implement and test new summarization methods, as well as to ease comparison with already implemented methods, in an unified framework.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Build StatusDocsWikiLicense: GPL v3

MOTS

MOTS (MOdular Tool for Summarization) is a summarization system, written in Java. It is as modular as possible, and is intended to provide an architecture to implement and test new summarization methods, as well as to ease comparison with already implemented methods, in an unified framework. This system is the first completely modular system for automatic summarization and already allows to summarize using more than a hundred combinations of modules. The need for such a system is important. Indeed, several evaluation campaigns exist in AS field, but summarization algorithms are not easy to compare due to the large variety of pre and post-processings they use.

Getting Started

  • Javadoc available.
  • We provide an example corpus from TAC2009 in /src/main/resources and its associated human summaries.

Prerequisites

  • Maven - Dependency Manager
  • glpk-utils in order to use ILP (sudo apt-get install glpk-utils)
  • At least python3 in order to use WordEmbeddings, make sure to have an updated version of pip3 (sudo pip3 install --upgrade pip)
  • gensim in order to use WordEmbeddings (pip3 install gensim --user)
  • jep in order to use WordEmbeddings (pip3 install jep --user)

Installing

  • You might define $CORPUS_DATA to your DUC/TAC folder.
  • Install ROUGE :
    • Define $ROUGE_HOME to your ROUGE installation folder (or to ./ROUGE-1.5.5/RELEASE-1.5.5).
    • Install XML::DOM module in order to use ROUGE perl script. (sudo cpan install XML::DOM).
    • Run ./rouge_install.sh or :
      • Define $ROUGE_EVAL_HOME to $ROUGE_HOME/data.
      • Recreate database :
       cd data/WordNet-2.0-Exceptions/
      rm WordNet-2.0.exc.db # only if exist
      perl buildExeptionDB.pl . exc WordNet-2.0.exc.db
      cd ..
      rm WordNet-2.0.exc.db # only if exist
      ln -s WordNet-2.0-Exceptions/WordNet-2.0.exc.db WordNet-2.0.exc.db
      
  • Run install.sh script.

Usage

MOTS is a command line tool than can be used like this :

./MOTS mots.X.Y.Z.jar -c <config_file> -m <multicorpus_file> -v <OPTIONAL>

MOTS script encapsulate some environnement variable needed for the execution of WordEmbeddings. If you don't use WordEmbeddings you could launch via :

java -jar mots.X.Y.Z.jar -c <config_file> -m <multicorpus_file> -v <OPTIONAL>

Example config file and multicorpus file are provided in /conf but should be adapted to your setup.

Go Deeper

Each summarization process is defined in a configuration file and the test corpus is defined in a multicorpus configuration file.

Process configuration

Example for LexRank_MMR configuration file :

<CONFIG>
<TASK ID="1">
<LANGUAGE>english</LANGUAGE>
<OUTPUT_PATH>doc/output</OUTPUT_PATH>
<MULTITHREADING>true</MULTITHREADING> <PREPROCESS NAME="GenerateTextModel">
<OPTION NAME="StanfordNLP">true</OPTION>
<OPTION NAME="StopWordListFile">$CORPUS_DATA/stopwords/englishStopWords.txt</OPTION>
</PREPROCESS>
<PROCESS>
<OPTION NAME="CorpusIdToSummarize">all</OPTION>
<OPTION NAME="ReadStopWords">false</OPTION>
<INDEX_BUILDER NAME="TF_IDF.TF_IDF">
</INDEX_BUILDER>
<CARACTERISTIC_BUILDER NAME="vector.TfIdfVectorSentence">
</CARACTERISTIC_BUILDER>
<SCORING_METHOD NAME="graphBased.LexRank">
<OPTION NAME="DampingParameter">0.15</OPTION>
<OPTION NAME="GraphThreshold">0.1</OPTION>
<OPTION NAME="SimilarityMethod">JaccardSimilarity</OPTION>
</SCORING_METHOD>
<SUMMARIZE_METHOD NAME="MMR">
<OPTION NAME="CharLimitBoolean">true</OPTION>
<OPTION NAME="Size">200</OPTION>
<OPTION NAME="SimilarityMethod">JaccardSimilarity</OPTION>
<OPTION NAME="Lambda">0.6</OPTION>
</SUMMARIZE_METHOD>
</PROCESS>
<ROUGE_EVALUATION>
<ROUGE_MEASURE>ROUGE-1	ROUGE-2	ROUGE-SU4</ROUGE_MEASURE>
<MODEL_ROOT>models</MODEL_ROOT>
<PEER_ROOT>systems</PEER_ROOT>
</ROUGE_EVALUATION>
</TASK>
</CONFIG>
  • <CONFIG> is the root node.
    • <TASK> represent a summarization task. You could do multiple in a simple run. At start, stick with one.
      • <LANGUAGE> is the input's document language for preprocessing goal. (english / french for now)
      • <OUTPUT_PATH> is the forlder's output path of the system. It is used to save preprocessed documents, ROUGE xml generated file, old score, ...
      • <MULTITHREADING> (boolean) launch the system in a mutltithreading way or not.
      • <PREPROCESS> is the preprocess step for the system. The preprocess java class to use is pass by the name variable. Here it's GenerateTextModel. It also needs two <OPTION> :
        • <OPTION NAME="StanfordNLP"> (boolean), true if you want to use StanfordNLP pipeline and tool to do the preprocessing.
        • <OPTION NAME="StopWordListPath"> (String), path of the stopwords list you want to use.
      • <PROCESS> is the main step of the system. It should have at least one <SUMMARIZE_METHOD> node and two <OPTION> It often has an <INDEX_BUILDER> node and a <CARACTERISTIC_BUILDER> node :
        • <OPTION NAME="CorpusIdToSummarize"> (String as a list of int separated by \t), the list of CorpusId to summarize from the MultiCorpus configuration file. "all" will do summarization for all corpus.
        • <OPTION NAME="ReadStopWords"> (boolean), state if the system count stopwords as part of the texts or not.
        • <INDEX_BUILDER> is the step where the system generate a computer friendly representation of each text's textual unit. (TF-IDF, Bigram, WordEmbeddings, ...)
        • <CARATERISTIC_BUILDER> is the sentence caracteristic generation step based on the textual unit index building.
        • <SCORING_METHOD> weights each sentences.
        • <SUMMARIZE_BUILDER> generate a summary usually by ranking sentence based on their score.
      • <ROUGE_EVALUATION> is the ROUGE evaluation step. For detail, look at ROUGE readme in /lib/ROUGE folder.
        • <ROUGE_MEASURE> (String as a list of int separated by \t), represent the list of ROUGE measure you want to use.
        • <MODEL_ROOT> is the model's folder name for ROUGE xml input files.
        • <PEER_ROOT> is the peer's folder name for ROUGE xml input files.

The <PROCESS> node is the system's core and you should look for more detail in the javadoc and the source code of the different INDEX_BUILDER, CARACTERISTIC_BUILDER, SCORING_METHOD and SUMMARIZE_METHOD class.

Multicorpus configuration

<?xml version="1.0" encoding="UTF-8"?>
<CONFIG>
<TASK ID="1">
<MULTICORPUS ID="0">
<CORPUS ID="0">
<INPUT_PATH>$CORPUS_DATA/TAC2009/UpdateSumm09_test_docs_files/D0901A/D0901A-A</INPUT_PATH>
<DOCUMENT ID="0">.*</DOCUMENT>
<SUMMARY_PATH>$CORPUS_DATA/TAC2009/UpdateSumm09_eval/ROUGE/models</SUMMARY_PATH>
<SUMMARY ID="0">D0901-A.*</SUMMARY>
</CORPUS>
</MULTICORPUS>
</TASK>
</CONFIG>

For now, all ID are useless and could be avoided.

  • <CONFIG> is the root node.
    • <TASK> represent a summarization task. You could do multiple in a simple run. At start, stick with one.
      • <MULTICORPUS> is a list of <CORPUS>
        • <CORPUS> can be one or more documents. The system will generate one summary per corpus.
          • <INPUT_PATH> is the folder containing the corpus' documents.
          • <DOCUMENT> is the regex for the documents you want to load. You could use multiple <DOCUMENT> node.
          • <SUMMARY_PATH> is the human summaries folder path.
          • <SUMMARY> is the regex for human summary file associating to this corpus. You could use multiple <SUMMARY> node.

Built With

  • Maven - Dependency Management

Authors

License

This project is licensed under the GPL3 License - see the LICENSE.md file for details

Acknowledgments

  • Thanks to Aurélien Bossard, my PhD supervisor.

About

MOTS (MOdular Tool for Summarization) is a summarization system, written in Java. It is as modular as possible, and is intended to provide an architecture to implement and test new summarization methods, as well as to ease comparison with already implemented methods, in an unified framework.

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages