Repository files navigation

RELVM

This repository contains the code accompanying the paper "Learning Informative Representations of Biomedical Relations with Latent Variable Models", Harshil Shah and Julien Fauqueur, EMNLP SustaiNLP 2020, (https://arxiv.org/abs/2011.10285).

Requirements

  • Python 3.7
    • Numpy >= 1.17.2
    • Tensorflow >= 2.0.0

Instructions

Introduction

The code in this repository is for training a latent variable generative model of pairs of entities and the contexts (i.e. sentences) in which the entities occur. The representations from this model can then be used to perform both mention-level and pair-level classification.

Throughout the code, the following conventions are used:

  • x or entities_x will refer to the first entity in a context.
  • y or entities_y will refer to the second entity in a context.
  • c or contexts will refer to the context (i.e. sentence) in which the entities occur.
  • r or labels will refer to the class label when performing either mention-level or pair-level classification.

Data

To avoid out-of-memory issues, the data is stored in memory-mapped Numpy arrays. The metadata is stored in JSON files.

Unsupervised

The unsupervised data directory (e.g. data/unsupervised) should contain the following JSON files (which contain the metadata):

  • vocab.json
    • This is a list of strings. It is the set of possible tokens which can appear in the contexts.
  • entity_types.json
    • This is a list of strings. It is the set of possible values for the entities. If training the model for mention-level classification, this will be a list of entity types. If training the model for pair-level classification, this will be a list of entity identifiers.

It should also contain the following memory-mapped Numpy arrays:

  • entities_x.mmap
    • This is a one-dimensional array. The ith row contains the index to entity_types.json for the first entity in the ith context.
  • entities_y.mmap
    • This is a one-dimensional array. The ith row contains the index to entity_types.json for the second entity in the ith context.
  • contexts.mmap
    • This is a two-dimensional array. The ith row contains the indices to vocab.json for the ith context.

Mention-level classification

The mention-level classification data directory (e.g. data/supervised/mention_level) should contain the following JSON files (which contain the metadata):

  • label_types.json
    • This is a list of strings. It is the set of possible values for the labels.

It should also contain the following memory-mapped Numpy arrays (for the training, validation, and test data respectively):

  • entities_x_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the first entity in the ith context.
  • entities_y_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the second entity in the ith context.
  • contexts_{train,valid,test}.mmap
    • These are two-dimensional arrays. The ith row contains the indices to vocab.json from the unsupervised data directory for the ith context.
  • labels_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to label_types.json for the ith entity pair and context.

Pair-level classification

The pair-level classification data directory (e.g. data/supervised/pair_level) should contain the following JSON files (which contain the metadata):

  • label_types.json
    • This is a list of strings. It is the set of possible values for the labels.
  • pos_entity_pairs.json
    • This is a list of strings containing the entity pairs which have a positive relation. Each element of this list contains the unique entity identifiers for the pair, joined together with a :.

It should also contain the following memory-mapped Numpy arrays (for the training, validation, and test data respectively):

  • entities_x_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the first entity in the ith pair.
  • entities_y_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the second entity in the ith pair.
  • labels_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to label_types.json for the ith entity pair.

Training and evaluating

Unsupervised representation learning

To train the unsupervised representation model, run exp_unsup.py, specifying the directory in which to store the model parameters. For example:

mkdir -p exp_outputs/unsup
python3 exp_unsup.py exp_outputs/unsup

Mention-level classification

Once the unsupervised representation model has been trained, the mention-level classification model can be trained and evaluated by running exp_classification_mention.py. First, the following variables in exp_classification_mention.py must be set:

  • trainer_unsup_pre_trained_dir must be set to the directory with the saved parameters from the unsupervised representation model (e.g. exp_outputs/unsup).
  • unsup_data_dir must be set to the directory with the data used to train the unsupervised representation model (e.g. data/unsupervised).

When running exp_classification_mention.py, the directory in which to store the classification model parameters must be specified. For example:

mkdir -p exp_outputs/classification_mention
python3 exp_classification_mention.py exp_outputs/classification_mention

Pair-level classification

The pair-level classification model can be trained and evaluated by running exp_classification_pair.py in an identical fashion to the mention-level classification model. Again, the following variables in exp_classification_pair.py must be set:

  • trainer_unsup_pre_trained_dir must be set to the directory with the saved parameters from the unsupervised representation model (e.g. exp_outputs/unsup).
  • unsup_data_dir must be set to the directory with the data used to train the unsupervised representation model (e.g. data/unsupervised).

When running exp_classification_pair.py, the directory in which to store the classification model parameters must be specified. For example:

mkdir -p exp_outputs/classification_pair
python3 exp_classification_pair.py exp_outputs/classification_pair

Component testing

From the project root folder, the following 3 scripts should be run in this order and all return OK.

python3 tests/test_unsup.py
python3 tests/test_classification_mention.py
python3 tests/test_classification_pair.py

About

This repository contains the code accompanying the paper "Learning Informative Representations of Biomedical Relations with Latent Variable Models", Harshil Shah and Julien Fauqueur, EMNLP SustaiNLP 2020.

Topics

Resources

Stars

15 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

RELVM

This repository contains the code accompanying the paper "Learning Informative Representations of Biomedical Relations with Latent Variable Models", Harshil Shah and Julien Fauqueur, EMNLP SustaiNLP 2020, (https://arxiv.org/abs/2011.10285).

Requirements

  • Python 3.7
    • Numpy >= 1.17.2
    • Tensorflow >= 2.0.0

Instructions

Introduction

The code in this repository is for training a latent variable generative model of pairs of entities and the contexts (i.e. sentences) in which the entities occur. The representations from this model can then be used to perform both mention-level and pair-level classification.

Throughout the code, the following conventions are used:

  • x or entities_x will refer to the first entity in a context.
  • y or entities_y will refer to the second entity in a context.
  • c or contexts will refer to the context (i.e. sentence) in which the entities occur.
  • r or labels will refer to the class label when performing either mention-level or pair-level classification.

Data

To avoid out-of-memory issues, the data is stored in memory-mapped Numpy arrays. The metadata is stored in JSON files.

Unsupervised

The unsupervised data directory (e.g. data/unsupervised) should contain the following JSON files (which contain the metadata):

  • vocab.json
    • This is a list of strings. It is the set of possible tokens which can appear in the contexts.
  • entity_types.json
    • This is a list of strings. It is the set of possible values for the entities. If training the model for mention-level classification, this will be a list of entity types. If training the model for pair-level classification, this will be a list of entity identifiers.

It should also contain the following memory-mapped Numpy arrays:

  • entities_x.mmap
    • This is a one-dimensional array. The ith row contains the index to entity_types.json for the first entity in the ith context.
  • entities_y.mmap
    • This is a one-dimensional array. The ith row contains the index to entity_types.json for the second entity in the ith context.
  • contexts.mmap
    • This is a two-dimensional array. The ith row contains the indices to vocab.json for the ith context.

Mention-level classification

The mention-level classification data directory (e.g. data/supervised/mention_level) should contain the following JSON files (which contain the metadata):

  • label_types.json
    • This is a list of strings. It is the set of possible values for the labels.

It should also contain the following memory-mapped Numpy arrays (for the training, validation, and test data respectively):

  • entities_x_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the first entity in the ith context.
  • entities_y_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the second entity in the ith context.
  • contexts_{train,valid,test}.mmap
    • These are two-dimensional arrays. The ith row contains the indices to vocab.json from the unsupervised data directory for the ith context.
  • labels_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to label_types.json for the ith entity pair and context.

Pair-level classification

The pair-level classification data directory (e.g. data/supervised/pair_level) should contain the following JSON files (which contain the metadata):

  • label_types.json
    • This is a list of strings. It is the set of possible values for the labels.
  • pos_entity_pairs.json
    • This is a list of strings containing the entity pairs which have a positive relation. Each element of this list contains the unique entity identifiers for the pair, joined together with a :.

It should also contain the following memory-mapped Numpy arrays (for the training, validation, and test data respectively):

  • entities_x_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the first entity in the ith pair.
  • entities_y_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the second entity in the ith pair.
  • labels_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to label_types.json for the ith entity pair.

Training and evaluating

Unsupervised representation learning

To train the unsupervised representation model, run exp_unsup.py, specifying the directory in which to store the model parameters. For example:

mkdir -p exp_outputs/unsup
python3 exp_unsup.py exp_outputs/unsup

Mention-level classification

Once the unsupervised representation model has been trained, the mention-level classification model can be trained and evaluated by running exp_classification_mention.py. First, the following variables in exp_classification_mention.py must be set:

  • trainer_unsup_pre_trained_dir must be set to the directory with the saved parameters from the unsupervised representation model (e.g. exp_outputs/unsup).
  • unsup_data_dir must be set to the directory with the data used to train the unsupervised representation model (e.g. data/unsupervised).

When running exp_classification_mention.py, the directory in which to store the classification model parameters must be specified. For example:

mkdir -p exp_outputs/classification_mention
python3 exp_classification_mention.py exp_outputs/classification_mention

Pair-level classification

The pair-level classification model can be trained and evaluated by running exp_classification_pair.py in an identical fashion to the mention-level classification model. Again, the following variables in exp_classification_pair.py must be set:

  • trainer_unsup_pre_trained_dir must be set to the directory with the saved parameters from the unsupervised representation model (e.g. exp_outputs/unsup).
  • unsup_data_dir must be set to the directory with the data used to train the unsupervised representation model (e.g. data/unsupervised).

When running exp_classification_pair.py, the directory in which to store the classification model parameters must be specified. For example:

mkdir -p exp_outputs/classification_pair
python3 exp_classification_pair.py exp_outputs/classification_pair

Component testing

From the project root folder, the following 3 scripts should be run in this order and all return OK.

python3 tests/test_unsup.py
python3 tests/test_classification_mention.py
python3 tests/test_classification_pair.py

About

This repository contains the code accompanying the paper "Learning Informative Representations of Biomedical Relations with Latent Variable Models", Harshil Shah and Julien Fauqueur, EMNLP SustaiNLP 2020.

Topics

Resources

Stars

15 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

RELVM

This repository contains the code accompanying the paper "Learning Informative Representations of Biomedical Relations with Latent Variable Models", Harshil Shah and Julien Fauqueur, EMNLP SustaiNLP 2020, (https://arxiv.org/abs/2011.10285).

Requirements

  • Python 3.7
    • Numpy >= 1.17.2
    • Tensorflow >= 2.0.0

Instructions

Introduction

The code in this repository is for training a latent variable generative model of pairs of entities and the contexts (i.e. sentences) in which the entities occur. The representations from this model can then be used to perform both mention-level and pair-level classification.

Throughout the code, the following conventions are used:

  • x or entities_x will refer to the first entity in a context.
  • y or entities_y will refer to the second entity in a context.
  • c or contexts will refer to the context (i.e. sentence) in which the entities occur.
  • r or labels will refer to the class label when performing either mention-level or pair-level classification.

Data

To avoid out-of-memory issues, the data is stored in memory-mapped Numpy arrays. The metadata is stored in JSON files.

Unsupervised

The unsupervised data directory (e.g. data/unsupervised) should contain the following JSON files (which contain the metadata):

  • vocab.json
    • This is a list of strings. It is the set of possible tokens which can appear in the contexts.
  • entity_types.json
    • This is a list of strings. It is the set of possible values for the entities. If training the model for mention-level classification, this will be a list of entity types. If training the model for pair-level classification, this will be a list of entity identifiers.

It should also contain the following memory-mapped Numpy arrays:

  • entities_x.mmap
    • This is a one-dimensional array. The ith row contains the index to entity_types.json for the first entity in the ith context.
  • entities_y.mmap
    • This is a one-dimensional array. The ith row contains the index to entity_types.json for the second entity in the ith context.
  • contexts.mmap
    • This is a two-dimensional array. The ith row contains the indices to vocab.json for the ith context.

Mention-level classification

The mention-level classification data directory (e.g. data/supervised/mention_level) should contain the following JSON files (which contain the metadata):

  • label_types.json
    • This is a list of strings. It is the set of possible values for the labels.

It should also contain the following memory-mapped Numpy arrays (for the training, validation, and test data respectively):

  • entities_x_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the first entity in the ith context.
  • entities_y_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the second entity in the ith context.
  • contexts_{train,valid,test}.mmap
    • These are two-dimensional arrays. The ith row contains the indices to vocab.json from the unsupervised data directory for the ith context.
  • labels_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to label_types.json for the ith entity pair and context.

Pair-level classification

The pair-level classification data directory (e.g. data/supervised/pair_level) should contain the following JSON files (which contain the metadata):

  • label_types.json
    • This is a list of strings. It is the set of possible values for the labels.
  • pos_entity_pairs.json
    • This is a list of strings containing the entity pairs which have a positive relation. Each element of this list contains the unique entity identifiers for the pair, joined together with a :.

It should also contain the following memory-mapped Numpy arrays (for the training, validation, and test data respectively):

  • entities_x_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the first entity in the ith pair.
  • entities_y_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the second entity in the ith pair.
  • labels_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to label_types.json for the ith entity pair.

Training and evaluating

Unsupervised representation learning

To train the unsupervised representation model, run exp_unsup.py, specifying the directory in which to store the model parameters. For example:

mkdir -p exp_outputs/unsup
python3 exp_unsup.py exp_outputs/unsup

Mention-level classification

Once the unsupervised representation model has been trained, the mention-level classification model can be trained and evaluated by running exp_classification_mention.py. First, the following variables in exp_classification_mention.py must be set:

  • trainer_unsup_pre_trained_dir must be set to the directory with the saved parameters from the unsupervised representation model (e.g. exp_outputs/unsup).
  • unsup_data_dir must be set to the directory with the data used to train the unsupervised representation model (e.g. data/unsupervised).

When running exp_classification_mention.py, the directory in which to store the classification model parameters must be specified. For example:

mkdir -p exp_outputs/classification_mention
python3 exp_classification_mention.py exp_outputs/classification_mention

Pair-level classification

The pair-level classification model can be trained and evaluated by running exp_classification_pair.py in an identical fashion to the mention-level classification model. Again, the following variables in exp_classification_pair.py must be set:

  • trainer_unsup_pre_trained_dir must be set to the directory with the saved parameters from the unsupervised representation model (e.g. exp_outputs/unsup).
  • unsup_data_dir must be set to the directory with the data used to train the unsupervised representation model (e.g. data/unsupervised).

When running exp_classification_pair.py, the directory in which to store the classification model parameters must be specified. For example:

mkdir -p exp_outputs/classification_pair
python3 exp_classification_pair.py exp_outputs/classification_pair

Component testing

From the project root folder, the following 3 scripts should be run in this order and all return OK.

python3 tests/test_unsup.py
python3 tests/test_classification_mention.py
python3 tests/test_classification_pair.py

About

This repository contains the code accompanying the paper "Learning Informative Representations of Biomedical Relations with Latent Variable Models", Harshil Shah and Julien Fauqueur, EMNLP SustaiNLP 2020.

Topics

Resources

Stars

15 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

RELVM

This repository contains the code accompanying the paper "Learning Informative Representations of Biomedical Relations with Latent Variable Models", Harshil Shah and Julien Fauqueur, EMNLP SustaiNLP 2020, (https://arxiv.org/abs/2011.10285).

Requirements

  • Python 3.7
    • Numpy >= 1.17.2
    • Tensorflow >= 2.0.0

Instructions

Introduction

The code in this repository is for training a latent variable generative model of pairs of entities and the contexts (i.e. sentences) in which the entities occur. The representations from this model can then be used to perform both mention-level and pair-level classification.

Throughout the code, the following conventions are used:

  • x or entities_x will refer to the first entity in a context.
  • y or entities_y will refer to the second entity in a context.
  • c or contexts will refer to the context (i.e. sentence) in which the entities occur.
  • r or labels will refer to the class label when performing either mention-level or pair-level classification.

Data

To avoid out-of-memory issues, the data is stored in memory-mapped Numpy arrays. The metadata is stored in JSON files.

Unsupervised

The unsupervised data directory (e.g. data/unsupervised) should contain the following JSON files (which contain the metadata):

  • vocab.json
    • This is a list of strings. It is the set of possible tokens which can appear in the contexts.
  • entity_types.json
    • This is a list of strings. It is the set of possible values for the entities. If training the model for mention-level classification, this will be a list of entity types. If training the model for pair-level classification, this will be a list of entity identifiers.

It should also contain the following memory-mapped Numpy arrays:

  • entities_x.mmap
    • This is a one-dimensional array. The ith row contains the index to entity_types.json for the first entity in the ith context.
  • entities_y.mmap
    • This is a one-dimensional array. The ith row contains the index to entity_types.json for the second entity in the ith context.
  • contexts.mmap
    • This is a two-dimensional array. The ith row contains the indices to vocab.json for the ith context.

Mention-level classification

The mention-level classification data directory (e.g. data/supervised/mention_level) should contain the following JSON files (which contain the metadata):

  • label_types.json
    • This is a list of strings. It is the set of possible values for the labels.

It should also contain the following memory-mapped Numpy arrays (for the training, validation, and test data respectively):

  • entities_x_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the first entity in the ith context.
  • entities_y_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the second entity in the ith context.
  • contexts_{train,valid,test}.mmap
    • These are two-dimensional arrays. The ith row contains the indices to vocab.json from the unsupervised data directory for the ith context.
  • labels_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to label_types.json for the ith entity pair and context.

Pair-level classification

The pair-level classification data directory (e.g. data/supervised/pair_level) should contain the following JSON files (which contain the metadata):

  • label_types.json
    • This is a list of strings. It is the set of possible values for the labels.
  • pos_entity_pairs.json
    • This is a list of strings containing the entity pairs which have a positive relation. Each element of this list contains the unique entity identifiers for the pair, joined together with a :.

It should also contain the following memory-mapped Numpy arrays (for the training, validation, and test data respectively):

  • entities_x_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the first entity in the ith pair.
  • entities_y_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the second entity in the ith pair.
  • labels_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to label_types.json for the ith entity pair.

Training and evaluating

Unsupervised representation learning

To train the unsupervised representation model, run exp_unsup.py, specifying the directory in which to store the model parameters. For example:

mkdir -p exp_outputs/unsup
python3 exp_unsup.py exp_outputs/unsup

Mention-level classification

Once the unsupervised representation model has been trained, the mention-level classification model can be trained and evaluated by running exp_classification_mention.py. First, the following variables in exp_classification_mention.py must be set:

  • trainer_unsup_pre_trained_dir must be set to the directory with the saved parameters from the unsupervised representation model (e.g. exp_outputs/unsup).
  • unsup_data_dir must be set to the directory with the data used to train the unsupervised representation model (e.g. data/unsupervised).

When running exp_classification_mention.py, the directory in which to store the classification model parameters must be specified. For example:

mkdir -p exp_outputs/classification_mention
python3 exp_classification_mention.py exp_outputs/classification_mention

Pair-level classification

The pair-level classification model can be trained and evaluated by running exp_classification_pair.py in an identical fashion to the mention-level classification model. Again, the following variables in exp_classification_pair.py must be set:

  • trainer_unsup_pre_trained_dir must be set to the directory with the saved parameters from the unsupervised representation model (e.g. exp_outputs/unsup).
  • unsup_data_dir must be set to the directory with the data used to train the unsupervised representation model (e.g. data/unsupervised).

When running exp_classification_pair.py, the directory in which to store the classification model parameters must be specified. For example:

mkdir -p exp_outputs/classification_pair
python3 exp_classification_pair.py exp_outputs/classification_pair

Component testing

From the project root folder, the following 3 scripts should be run in this order and all return OK.

python3 tests/test_unsup.py
python3 tests/test_classification_mention.py
python3 tests/test_classification_pair.py

About

This repository contains the code accompanying the paper "Learning Informative Representations of Biomedical Relations with Latent Variable Models", Harshil Shah and Julien Fauqueur, EMNLP SustaiNLP 2020.

Topics

Resources

Stars

15 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

RELVM

This repository contains the code accompanying the paper "Learning Informative Representations of Biomedical Relations with Latent Variable Models", Harshil Shah and Julien Fauqueur, EMNLP SustaiNLP 2020, (https://arxiv.org/abs/2011.10285).

Requirements

  • Python 3.7
    • Numpy >= 1.17.2
    • Tensorflow >= 2.0.0

Instructions

Introduction

The code in this repository is for training a latent variable generative model of pairs of entities and the contexts (i.e. sentences) in which the entities occur. The representations from this model can then be used to perform both mention-level and pair-level classification.

Throughout the code, the following conventions are used:

  • x or entities_x will refer to the first entity in a context.
  • y or entities_y will refer to the second entity in a context.
  • c or contexts will refer to the context (i.e. sentence) in which the entities occur.
  • r or labels will refer to the class label when performing either mention-level or pair-level classification.

Data

To avoid out-of-memory issues, the data is stored in memory-mapped Numpy arrays. The metadata is stored in JSON files.

Unsupervised

The unsupervised data directory (e.g. data/unsupervised) should contain the following JSON files (which contain the metadata):

  • vocab.json
    • This is a list of strings. It is the set of possible tokens which can appear in the contexts.
  • entity_types.json
    • This is a list of strings. It is the set of possible values for the entities. If training the model for mention-level classification, this will be a list of entity types. If training the model for pair-level classification, this will be a list of entity identifiers.

It should also contain the following memory-mapped Numpy arrays:

  • entities_x.mmap
    • This is a one-dimensional array. The ith row contains the index to entity_types.json for the first entity in the ith context.
  • entities_y.mmap
    • This is a one-dimensional array. The ith row contains the index to entity_types.json for the second entity in the ith context.
  • contexts.mmap
    • This is a two-dimensional array. The ith row contains the indices to vocab.json for the ith context.

Mention-level classification

The mention-level classification data directory (e.g. data/supervised/mention_level) should contain the following JSON files (which contain the metadata):

  • label_types.json
    • This is a list of strings. It is the set of possible values for the labels.

It should also contain the following memory-mapped Numpy arrays (for the training, validation, and test data respectively):

  • entities_x_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the first entity in the ith context.
  • entities_y_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the second entity in the ith context.
  • contexts_{train,valid,test}.mmap
    • These are two-dimensional arrays. The ith row contains the indices to vocab.json from the unsupervised data directory for the ith context.
  • labels_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to label_types.json for the ith entity pair and context.

Pair-level classification

The pair-level classification data directory (e.g. data/supervised/pair_level) should contain the following JSON files (which contain the metadata):

  • label_types.json
    • This is a list of strings. It is the set of possible values for the labels.
  • pos_entity_pairs.json
    • This is a list of strings containing the entity pairs which have a positive relation. Each element of this list contains the unique entity identifiers for the pair, joined together with a :.

It should also contain the following memory-mapped Numpy arrays (for the training, validation, and test data respectively):

  • entities_x_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the first entity in the ith pair.
  • entities_y_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the second entity in the ith pair.
  • labels_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to label_types.json for the ith entity pair.

Training and evaluating

Unsupervised representation learning

To train the unsupervised representation model, run exp_unsup.py, specifying the directory in which to store the model parameters. For example:

mkdir -p exp_outputs/unsup
python3 exp_unsup.py exp_outputs/unsup

Mention-level classification

Once the unsupervised representation model has been trained, the mention-level classification model can be trained and evaluated by running exp_classification_mention.py. First, the following variables in exp_classification_mention.py must be set:

  • trainer_unsup_pre_trained_dir must be set to the directory with the saved parameters from the unsupervised representation model (e.g. exp_outputs/unsup).
  • unsup_data_dir must be set to the directory with the data used to train the unsupervised representation model (e.g. data/unsupervised).

When running exp_classification_mention.py, the directory in which to store the classification model parameters must be specified. For example:

mkdir -p exp_outputs/classification_mention
python3 exp_classification_mention.py exp_outputs/classification_mention

Pair-level classification

The pair-level classification model can be trained and evaluated by running exp_classification_pair.py in an identical fashion to the mention-level classification model. Again, the following variables in exp_classification_pair.py must be set:

  • trainer_unsup_pre_trained_dir must be set to the directory with the saved parameters from the unsupervised representation model (e.g. exp_outputs/unsup).
  • unsup_data_dir must be set to the directory with the data used to train the unsupervised representation model (e.g. data/unsupervised).

When running exp_classification_pair.py, the directory in which to store the classification model parameters must be specified. For example:

mkdir -p exp_outputs/classification_pair
python3 exp_classification_pair.py exp_outputs/classification_pair

Component testing

From the project root folder, the following 3 scripts should be run in this order and all return OK.

python3 tests/test_unsup.py
python3 tests/test_classification_mention.py
python3 tests/test_classification_pair.py

About

This repository contains the code accompanying the paper "Learning Informative Representations of Biomedical Relations with Latent Variable Models", Harshil Shah and Julien Fauqueur, EMNLP SustaiNLP 2020.

Topics

Resources

Stars

15 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

RELVM

This repository contains the code accompanying the paper "Learning Informative Representations of Biomedical Relations with Latent Variable Models", Harshil Shah and Julien Fauqueur, EMNLP SustaiNLP 2020, (https://arxiv.org/abs/2011.10285).

Requirements

  • Python 3.7
    • Numpy >= 1.17.2
    • Tensorflow >= 2.0.0

Instructions

Introduction

The code in this repository is for training a latent variable generative model of pairs of entities and the contexts (i.e. sentences) in which the entities occur. The representations from this model can then be used to perform both mention-level and pair-level classification.

Throughout the code, the following conventions are used:

  • x or entities_x will refer to the first entity in a context.
  • y or entities_y will refer to the second entity in a context.
  • c or contexts will refer to the context (i.e. sentence) in which the entities occur.
  • r or labels will refer to the class label when performing either mention-level or pair-level classification.

Data

To avoid out-of-memory issues, the data is stored in memory-mapped Numpy arrays. The metadata is stored in JSON files.

Unsupervised

The unsupervised data directory (e.g. data/unsupervised) should contain the following JSON files (which contain the metadata):

  • vocab.json
    • This is a list of strings. It is the set of possible tokens which can appear in the contexts.
  • entity_types.json
    • This is a list of strings. It is the set of possible values for the entities. If training the model for mention-level classification, this will be a list of entity types. If training the model for pair-level classification, this will be a list of entity identifiers.

It should also contain the following memory-mapped Numpy arrays:

  • entities_x.mmap
    • This is a one-dimensional array. The ith row contains the index to entity_types.json for the first entity in the ith context.
  • entities_y.mmap
    • This is a one-dimensional array. The ith row contains the index to entity_types.json for the second entity in the ith context.
  • contexts.mmap
    • This is a two-dimensional array. The ith row contains the indices to vocab.json for the ith context.

Mention-level classification

The mention-level classification data directory (e.g. data/supervised/mention_level) should contain the following JSON files (which contain the metadata):

  • label_types.json
    • This is a list of strings. It is the set of possible values for the labels.

It should also contain the following memory-mapped Numpy arrays (for the training, validation, and test data respectively):

  • entities_x_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the first entity in the ith context.
  • entities_y_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the second entity in the ith context.
  • contexts_{train,valid,test}.mmap
    • These are two-dimensional arrays. The ith row contains the indices to vocab.json from the unsupervised data directory for the ith context.
  • labels_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to label_types.json for the ith entity pair and context.

Pair-level classification

The pair-level classification data directory (e.g. data/supervised/pair_level) should contain the following JSON files (which contain the metadata):

  • label_types.json
    • This is a list of strings. It is the set of possible values for the labels.
  • pos_entity_pairs.json
    • This is a list of strings containing the entity pairs which have a positive relation. Each element of this list contains the unique entity identifiers for the pair, joined together with a :.

It should also contain the following memory-mapped Numpy arrays (for the training, validation, and test data respectively):

  • entities_x_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the first entity in the ith pair.
  • entities_y_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the second entity in the ith pair.
  • labels_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to label_types.json for the ith entity pair.

Training and evaluating

Unsupervised representation learning

To train the unsupervised representation model, run exp_unsup.py, specifying the directory in which to store the model parameters. For example:

mkdir -p exp_outputs/unsup
python3 exp_unsup.py exp_outputs/unsup

Mention-level classification

Once the unsupervised representation model has been trained, the mention-level classification model can be trained and evaluated by running exp_classification_mention.py. First, the following variables in exp_classification_mention.py must be set:

  • trainer_unsup_pre_trained_dir must be set to the directory with the saved parameters from the unsupervised representation model (e.g. exp_outputs/unsup).
  • unsup_data_dir must be set to the directory with the data used to train the unsupervised representation model (e.g. data/unsupervised).

When running exp_classification_mention.py, the directory in which to store the classification model parameters must be specified. For example:

mkdir -p exp_outputs/classification_mention
python3 exp_classification_mention.py exp_outputs/classification_mention

Pair-level classification

The pair-level classification model can be trained and evaluated by running exp_classification_pair.py in an identical fashion to the mention-level classification model. Again, the following variables in exp_classification_pair.py must be set:

  • trainer_unsup_pre_trained_dir must be set to the directory with the saved parameters from the unsupervised representation model (e.g. exp_outputs/unsup).
  • unsup_data_dir must be set to the directory with the data used to train the unsupervised representation model (e.g. data/unsupervised).

When running exp_classification_pair.py, the directory in which to store the classification model parameters must be specified. For example:

mkdir -p exp_outputs/classification_pair
python3 exp_classification_pair.py exp_outputs/classification_pair

Component testing

From the project root folder, the following 3 scripts should be run in this order and all return OK.

python3 tests/test_unsup.py
python3 tests/test_classification_mention.py
python3 tests/test_classification_pair.py

About

This repository contains the code accompanying the paper "Learning Informative Representations of Biomedical Relations with Latent Variable Models", Harshil Shah and Julien Fauqueur, EMNLP SustaiNLP 2020.

Topics

Resources

Stars

15 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

RELVM

This repository contains the code accompanying the paper "Learning Informative Representations of Biomedical Relations with Latent Variable Models", Harshil Shah and Julien Fauqueur, EMNLP SustaiNLP 2020, (https://arxiv.org/abs/2011.10285).

Requirements

  • Python 3.7
    • Numpy >= 1.17.2
    • Tensorflow >= 2.0.0

Instructions

Introduction

The code in this repository is for training a latent variable generative model of pairs of entities and the contexts (i.e. sentences) in which the entities occur. The representations from this model can then be used to perform both mention-level and pair-level classification.

Throughout the code, the following conventions are used:

  • x or entities_x will refer to the first entity in a context.
  • y or entities_y will refer to the second entity in a context.
  • c or contexts will refer to the context (i.e. sentence) in which the entities occur.
  • r or labels will refer to the class label when performing either mention-level or pair-level classification.

Data

To avoid out-of-memory issues, the data is stored in memory-mapped Numpy arrays. The metadata is stored in JSON files.

Unsupervised

The unsupervised data directory (e.g. data/unsupervised) should contain the following JSON files (which contain the metadata):

  • vocab.json
    • This is a list of strings. It is the set of possible tokens which can appear in the contexts.
  • entity_types.json
    • This is a list of strings. It is the set of possible values for the entities. If training the model for mention-level classification, this will be a list of entity types. If training the model for pair-level classification, this will be a list of entity identifiers.

It should also contain the following memory-mapped Numpy arrays:

  • entities_x.mmap
    • This is a one-dimensional array. The ith row contains the index to entity_types.json for the first entity in the ith context.
  • entities_y.mmap
    • This is a one-dimensional array. The ith row contains the index to entity_types.json for the second entity in the ith context.
  • contexts.mmap
    • This is a two-dimensional array. The ith row contains the indices to vocab.json for the ith context.

Mention-level classification

The mention-level classification data directory (e.g. data/supervised/mention_level) should contain the following JSON files (which contain the metadata):

  • label_types.json
    • This is a list of strings. It is the set of possible values for the labels.

It should also contain the following memory-mapped Numpy arrays (for the training, validation, and test data respectively):

  • entities_x_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the first entity in the ith context.
  • entities_y_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the second entity in the ith context.
  • contexts_{train,valid,test}.mmap
    • These are two-dimensional arrays. The ith row contains the indices to vocab.json from the unsupervised data directory for the ith context.
  • labels_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to label_types.json for the ith entity pair and context.

Pair-level classification

The pair-level classification data directory (e.g. data/supervised/pair_level) should contain the following JSON files (which contain the metadata):

  • label_types.json
    • This is a list of strings. It is the set of possible values for the labels.
  • pos_entity_pairs.json
    • This is a list of strings containing the entity pairs which have a positive relation. Each element of this list contains the unique entity identifiers for the pair, joined together with a :.

It should also contain the following memory-mapped Numpy arrays (for the training, validation, and test data respectively):

  • entities_x_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the first entity in the ith pair.
  • entities_y_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the second entity in the ith pair.
  • labels_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to label_types.json for the ith entity pair.

Training and evaluating

Unsupervised representation learning

To train the unsupervised representation model, run exp_unsup.py, specifying the directory in which to store the model parameters. For example:

mkdir -p exp_outputs/unsup
python3 exp_unsup.py exp_outputs/unsup

Mention-level classification

Once the unsupervised representation model has been trained, the mention-level classification model can be trained and evaluated by running exp_classification_mention.py. First, the following variables in exp_classification_mention.py must be set:

  • trainer_unsup_pre_trained_dir must be set to the directory with the saved parameters from the unsupervised representation model (e.g. exp_outputs/unsup).
  • unsup_data_dir must be set to the directory with the data used to train the unsupervised representation model (e.g. data/unsupervised).

When running exp_classification_mention.py, the directory in which to store the classification model parameters must be specified. For example:

mkdir -p exp_outputs/classification_mention
python3 exp_classification_mention.py exp_outputs/classification_mention

Pair-level classification

The pair-level classification model can be trained and evaluated by running exp_classification_pair.py in an identical fashion to the mention-level classification model. Again, the following variables in exp_classification_pair.py must be set:

  • trainer_unsup_pre_trained_dir must be set to the directory with the saved parameters from the unsupervised representation model (e.g. exp_outputs/unsup).
  • unsup_data_dir must be set to the directory with the data used to train the unsupervised representation model (e.g. data/unsupervised).

When running exp_classification_pair.py, the directory in which to store the classification model parameters must be specified. For example:

mkdir -p exp_outputs/classification_pair
python3 exp_classification_pair.py exp_outputs/classification_pair

Component testing

From the project root folder, the following 3 scripts should be run in this order and all return OK.

python3 tests/test_unsup.py
python3 tests/test_classification_mention.py
python3 tests/test_classification_pair.py

About

This repository contains the code accompanying the paper "Learning Informative Representations of Biomedical Relations with Latent Variable Models", Harshil Shah and Julien Fauqueur, EMNLP SustaiNLP 2020.

Topics

Resources

Stars

15 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

RELVM

This repository contains the code accompanying the paper "Learning Informative Representations of Biomedical Relations with Latent Variable Models", Harshil Shah and Julien Fauqueur, EMNLP SustaiNLP 2020, (https://arxiv.org/abs/2011.10285).

Requirements

  • Python 3.7
    • Numpy >= 1.17.2
    • Tensorflow >= 2.0.0

Instructions

Introduction

The code in this repository is for training a latent variable generative model of pairs of entities and the contexts (i.e. sentences) in which the entities occur. The representations from this model can then be used to perform both mention-level and pair-level classification.

Throughout the code, the following conventions are used:

  • x or entities_x will refer to the first entity in a context.
  • y or entities_y will refer to the second entity in a context.
  • c or contexts will refer to the context (i.e. sentence) in which the entities occur.
  • r or labels will refer to the class label when performing either mention-level or pair-level classification.

Data

To avoid out-of-memory issues, the data is stored in memory-mapped Numpy arrays. The metadata is stored in JSON files.

Unsupervised

The unsupervised data directory (e.g. data/unsupervised) should contain the following JSON files (which contain the metadata):

  • vocab.json
    • This is a list of strings. It is the set of possible tokens which can appear in the contexts.
  • entity_types.json
    • This is a list of strings. It is the set of possible values for the entities. If training the model for mention-level classification, this will be a list of entity types. If training the model for pair-level classification, this will be a list of entity identifiers.

It should also contain the following memory-mapped Numpy arrays:

  • entities_x.mmap
    • This is a one-dimensional array. The ith row contains the index to entity_types.json for the first entity in the ith context.
  • entities_y.mmap
    • This is a one-dimensional array. The ith row contains the index to entity_types.json for the second entity in the ith context.
  • contexts.mmap
    • This is a two-dimensional array. The ith row contains the indices to vocab.json for the ith context.

Mention-level classification

The mention-level classification data directory (e.g. data/supervised/mention_level) should contain the following JSON files (which contain the metadata):

  • label_types.json
    • This is a list of strings. It is the set of possible values for the labels.

It should also contain the following memory-mapped Numpy arrays (for the training, validation, and test data respectively):

  • entities_x_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the first entity in the ith context.
  • entities_y_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the second entity in the ith context.
  • contexts_{train,valid,test}.mmap
    • These are two-dimensional arrays. The ith row contains the indices to vocab.json from the unsupervised data directory for the ith context.
  • labels_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to label_types.json for the ith entity pair and context.

Pair-level classification

The pair-level classification data directory (e.g. data/supervised/pair_level) should contain the following JSON files (which contain the metadata):

  • label_types.json
    • This is a list of strings. It is the set of possible values for the labels.
  • pos_entity_pairs.json
    • This is a list of strings containing the entity pairs which have a positive relation. Each element of this list contains the unique entity identifiers for the pair, joined together with a :.

It should also contain the following memory-mapped Numpy arrays (for the training, validation, and test data respectively):

  • entities_x_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the first entity in the ith pair.
  • entities_y_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to entity_types.json from the unsupervised data directory for the second entity in the ith pair.
  • labels_{train,valid,test}.mmap
    • These are one-dimensional arrays. The ith row contains the index to label_types.json for the ith entity pair.

Training and evaluating

Unsupervised representation learning

To train the unsupervised representation model, run exp_unsup.py, specifying the directory in which to store the model parameters. For example:

mkdir -p exp_outputs/unsup
python3 exp_unsup.py exp_outputs/unsup

Mention-level classification

Once the unsupervised representation model has been trained, the mention-level classification model can be trained and evaluated by running exp_classification_mention.py. First, the following variables in exp_classification_mention.py must be set:

  • trainer_unsup_pre_trained_dir must be set to the directory with the saved parameters from the unsupervised representation model (e.g. exp_outputs/unsup).
  • unsup_data_dir must be set to the directory with the data used to train the unsupervised representation model (e.g. data/unsupervised).

When running exp_classification_mention.py, the directory in which to store the classification model parameters must be specified. For example:

mkdir -p exp_outputs/classification_mention
python3 exp_classification_mention.py exp_outputs/classification_mention

Pair-level classification

The pair-level classification model can be trained and evaluated by running exp_classification_pair.py in an identical fashion to the mention-level classification model. Again, the following variables in exp_classification_pair.py must be set:

  • trainer_unsup_pre_trained_dir must be set to the directory with the saved parameters from the unsupervised representation model (e.g. exp_outputs/unsup).
  • unsup_data_dir must be set to the directory with the data used to train the unsupervised representation model (e.g. data/unsupervised).

When running exp_classification_pair.py, the directory in which to store the classification model parameters must be specified. For example:

mkdir -p exp_outputs/classification_pair
python3 exp_classification_pair.py exp_outputs/classification_pair

Component testing

From the project root folder, the following 3 scripts should be run in this order and all return OK.

python3 tests/test_unsup.py
python3 tests/test_classification_mention.py
python3 tests/test_classification_pair.py

About

This repository contains the code accompanying the paper "Learning Informative Representations of Biomedical Relations with Latent Variable Models", Harshil Shah and Julien Fauqueur, EMNLP SustaiNLP 2020.

Topics

Resources

Stars

15 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages