Latest commit

History

704 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

NodeNormalization

DOIarXiv

Introduction

Node normalization takes a CURIE, and returns:

  • The preferred CURIE for this entity
  • All other known equivalent identifiers for the entity
  • Semantic types for the entity as defined by the Biolink Model

The data currently served by Node Normalization is created by the prototype project Babel, which attempts to find identifier equivalences, and makes sure that CURIE prefixes are BioLink Model compliant. The NodeNormalization service, however, is independent of Babel and as improved identifier equivalence tools are developed, their results can be easily incorporated.

To determine whether Node Normalization is likely to be useful, check /get_semantic_types, which lists the BioLink semantic types for which normalization has been attempted, and /get_curie_prefixes, which lists the number of times each prefix is used for a semantic type.

For examples of service usage, see the example notebook.

The Node normalization website leverages the R3 (Redis-REST with referencing) Redis data design and configuration.

Users can find the publicly available website at service.

Installation

Create a virtual environment

 python -m venv nodeNormalization-env

Activate the virtual environment

 source nodeNormalization-env/bin/activate

Install requirements

 > pip install -r requirements.txt

Generating equivalence data

The equivalence data can be generated by running Babel. An example of the contents of a compendia file is shown below:

 {"id": {"identifier": "PUBCHEM:50986940"}, "equivalent_identifiers": [{"identifier": "PUBCHEM:50986940"}, {"identifier": "INCHIKEY:CYMOSKLLKPIPCD-UHFFFAOYSA-N"}], "type": ["chemical_substance", "named_thing", "biological_entity", "molecular_entity"]}
{"id": {"identifier": "CHEMBL.COMPOUND:CHEMBL1546789", "label": "CHEMBL1546789"}, "equivalent_identifiers": [{"identifier": "CHEMBL.COMPOUND:CHEMBL1546789", "label": "CHEMBL1546789"}, {"identifier": "PUBCHEM:4879549"}, {"identifier": "INCHIKEY:FUIYIXDZTPMQEH-UHFFFAOYSA-N"}], "type": ["chemical_substance", "named_thing", "biological_entity", "molecular_entity"]}

Creating and loading a Redis container with data

A running instance of Redis is needed to house the node normalization data. a Redis Docker container image can be downloaded from Docker hub. The Redis caonteriner can be started with thie following docker command:

 docker run --name node-norm-redis -p 6379:6379 -d redis redis-server --appendonly yes

Note that the dataset for Node normalization is quite large and 256Gb of memory and disk space should be made available to the Redis instance to insure proper loading of the complete compendia.

Configuration

Insure that the ./config.json file is created and contains the parameters for the node normalization load specific to your environment.

The configuration parameters compendium_directory and data_files specify the location of the compendia files. An example of the files' contents
are listed below:

 {
"compendium_directory": "<path to files>",
"data_files": "anatomy.txt,BiologicalProcess.txt,cell.txt,cellular_component.txt,disease.txt,gene_compendium.txt,gene_family_compendium.txt,MolecularActivity.txt,pathways.txt,phenotypes.txt,taxon_compendium.txt",
"redis_host": "<Redis host server name>",
"redis_port": <Redis connection port>,
"redis_password": "<Redis password",
"test_mode": 1,
"debug_messages": 0
}

Loading of the Redis server with compendia data

The load.py script reads the configuration file for load parameters and the loads the compendia data into the Redis instance.

The redis command line can be used to monitor various aspects of the load.

It is possible to observer the progress of the load opening a command line within the container and issuing Redis commands.

View the number of keys loaded so far.

 redis-cli info keyspace

Once the database has completed loading it is recommended that the Redis database be persisted to disk.

 redis-cli save

Monitor the database to determine if the save has completed.

 redis-cli info persistence

Starting the FASTAPI webserver from the command line

The web server can be started after successful completion of the load.

 cd <Node normalization code root>
pip install -r requirements.txt
uvicorn --host 0.0.0.0 --port 8000 --workers 1 node_normalizer.server:app

Then navigate to http://localhost:8000/docs to run the application. Documentation for the NodeNorm API is available.

Webserver Docker container creation and execution

Much like the Redis Docker container noted above, a Docker container can also be created and executed to run the webserver.

Build the webserver Docker image

 cd <Node normalization code root>
docker build --tag <image_tag> .

Start the container:

Note the Dockerfile specifies port 6380 for the webservice container.

 docker run --name Node-normalization -p 8000:6380 node-norm

Then navigate to: http://localhost:8000/docs to run the application

Kubernetes configurations

Kubernetes configurations and helm charts for this project can be found at:

 https://github.com/helxplatform/translator-devops/helm/r3

Configuration

NodeNorm can be configured by setting environmental variables:

  • SERVER_NAME: The name of this server (defaults to infores:sri-node-normalizer)
  • SERVER_ROOT: The server root (defaults to /)
  • LOG_LEVEL: The log level (defaults to ERROR)
  • TRAPI_VERSION: The TRAPI version this version of NodeNorm supports.
  • MATURITY_VALUE: How mature is this NameRes (defaults to maturity, e.g. development)
  • LOCATION_VALUE: Where is this NameRes setup (defaults to location, e.g. RENCI)
  • EQ_BATCH_SIZE: The size of the get_eqids_and_types() batch size (defaults to 2500)
  • OTEL_ENABLED: Turn on Open TELemetry (default: 'false') -- only 'true' will turn this on.
    • JAEGER_HOST and JAEGER_PORT: Hostname and port for the Jaegar instance to provide telemetry to.
    • JAEGER_SERVICE_NAME: The name of this service (defaults to the value of SERVER_NAME)

About

Service that produces Translator compliant nodes given a curie

Topics

Resources

Stars

17 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

704 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

NodeNormalization

DOIarXiv

Introduction

Node normalization takes a CURIE, and returns:

  • The preferred CURIE for this entity
  • All other known equivalent identifiers for the entity
  • Semantic types for the entity as defined by the Biolink Model

The data currently served by Node Normalization is created by the prototype project Babel, which attempts to find identifier equivalences, and makes sure that CURIE prefixes are BioLink Model compliant. The NodeNormalization service, however, is independent of Babel and as improved identifier equivalence tools are developed, their results can be easily incorporated.

To determine whether Node Normalization is likely to be useful, check /get_semantic_types, which lists the BioLink semantic types for which normalization has been attempted, and /get_curie_prefixes, which lists the number of times each prefix is used for a semantic type.

For examples of service usage, see the example notebook.

The Node normalization website leverages the R3 (Redis-REST with referencing) Redis data design and configuration.

Users can find the publicly available website at service.

Installation

Create a virtual environment

 python -m venv nodeNormalization-env

Activate the virtual environment

 source nodeNormalization-env/bin/activate

Install requirements

 > pip install -r requirements.txt

Generating equivalence data

The equivalence data can be generated by running Babel. An example of the contents of a compendia file is shown below:

 {"id": {"identifier": "PUBCHEM:50986940"}, "equivalent_identifiers": [{"identifier": "PUBCHEM:50986940"}, {"identifier": "INCHIKEY:CYMOSKLLKPIPCD-UHFFFAOYSA-N"}], "type": ["chemical_substance", "named_thing", "biological_entity", "molecular_entity"]}
{"id": {"identifier": "CHEMBL.COMPOUND:CHEMBL1546789", "label": "CHEMBL1546789"}, "equivalent_identifiers": [{"identifier": "CHEMBL.COMPOUND:CHEMBL1546789", "label": "CHEMBL1546789"}, {"identifier": "PUBCHEM:4879549"}, {"identifier": "INCHIKEY:FUIYIXDZTPMQEH-UHFFFAOYSA-N"}], "type": ["chemical_substance", "named_thing", "biological_entity", "molecular_entity"]}

Creating and loading a Redis container with data

A running instance of Redis is needed to house the node normalization data. a Redis Docker container image can be downloaded from Docker hub. The Redis caonteriner can be started with thie following docker command:

 docker run --name node-norm-redis -p 6379:6379 -d redis redis-server --appendonly yes

Note that the dataset for Node normalization is quite large and 256Gb of memory and disk space should be made available to the Redis instance to insure proper loading of the complete compendia.

Configuration

Insure that the ./config.json file is created and contains the parameters for the node normalization load specific to your environment.

The configuration parameters compendium_directory and data_files specify the location of the compendia files. An example of the files' contents
are listed below:

 {
"compendium_directory": "<path to files>",
"data_files": "anatomy.txt,BiologicalProcess.txt,cell.txt,cellular_component.txt,disease.txt,gene_compendium.txt,gene_family_compendium.txt,MolecularActivity.txt,pathways.txt,phenotypes.txt,taxon_compendium.txt",
"redis_host": "<Redis host server name>",
"redis_port": <Redis connection port>,
"redis_password": "<Redis password",
"test_mode": 1,
"debug_messages": 0
}

Loading of the Redis server with compendia data

The load.py script reads the configuration file for load parameters and the loads the compendia data into the Redis instance.

The redis command line can be used to monitor various aspects of the load.

It is possible to observer the progress of the load opening a command line within the container and issuing Redis commands.

View the number of keys loaded so far.

 redis-cli info keyspace

Once the database has completed loading it is recommended that the Redis database be persisted to disk.

 redis-cli save

Monitor the database to determine if the save has completed.

 redis-cli info persistence

Starting the FASTAPI webserver from the command line

The web server can be started after successful completion of the load.

 cd <Node normalization code root>
pip install -r requirements.txt
uvicorn --host 0.0.0.0 --port 8000 --workers 1 node_normalizer.server:app

Then navigate to http://localhost:8000/docs to run the application. Documentation for the NodeNorm API is available.

Webserver Docker container creation and execution

Much like the Redis Docker container noted above, a Docker container can also be created and executed to run the webserver.

Build the webserver Docker image

 cd <Node normalization code root>
docker build --tag <image_tag> .

Start the container:

Note the Dockerfile specifies port 6380 for the webservice container.

 docker run --name Node-normalization -p 8000:6380 node-norm

Then navigate to: http://localhost:8000/docs to run the application

Kubernetes configurations

Kubernetes configurations and helm charts for this project can be found at:

 https://github.com/helxplatform/translator-devops/helm/r3

Configuration

NodeNorm can be configured by setting environmental variables:

  • SERVER_NAME: The name of this server (defaults to infores:sri-node-normalizer)
  • SERVER_ROOT: The server root (defaults to /)
  • LOG_LEVEL: The log level (defaults to ERROR)
  • TRAPI_VERSION: The TRAPI version this version of NodeNorm supports.
  • MATURITY_VALUE: How mature is this NameRes (defaults to maturity, e.g. development)
  • LOCATION_VALUE: Where is this NameRes setup (defaults to location, e.g. RENCI)
  • EQ_BATCH_SIZE: The size of the get_eqids_and_types() batch size (defaults to 2500)
  • OTEL_ENABLED: Turn on Open TELemetry (default: 'false') -- only 'true' will turn this on.
    • JAEGER_HOST and JAEGER_PORT: Hostname and port for the Jaegar instance to provide telemetry to.
    • JAEGER_SERVICE_NAME: The name of this service (defaults to the value of SERVER_NAME)

About

Service that produces Translator compliant nodes given a curie

Topics

Resources

Stars

17 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

704 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

NodeNormalization

DOIarXiv

Introduction

Node normalization takes a CURIE, and returns:

  • The preferred CURIE for this entity
  • All other known equivalent identifiers for the entity
  • Semantic types for the entity as defined by the Biolink Model

The data currently served by Node Normalization is created by the prototype project Babel, which attempts to find identifier equivalences, and makes sure that CURIE prefixes are BioLink Model compliant. The NodeNormalization service, however, is independent of Babel and as improved identifier equivalence tools are developed, their results can be easily incorporated.

To determine whether Node Normalization is likely to be useful, check /get_semantic_types, which lists the BioLink semantic types for which normalization has been attempted, and /get_curie_prefixes, which lists the number of times each prefix is used for a semantic type.

For examples of service usage, see the example notebook.

The Node normalization website leverages the R3 (Redis-REST with referencing) Redis data design and configuration.

Users can find the publicly available website at service.

Installation

Create a virtual environment

 python -m venv nodeNormalization-env

Activate the virtual environment

 source nodeNormalization-env/bin/activate

Install requirements

 > pip install -r requirements.txt

Generating equivalence data

The equivalence data can be generated by running Babel. An example of the contents of a compendia file is shown below:

 {"id": {"identifier": "PUBCHEM:50986940"}, "equivalent_identifiers": [{"identifier": "PUBCHEM:50986940"}, {"identifier": "INCHIKEY:CYMOSKLLKPIPCD-UHFFFAOYSA-N"}], "type": ["chemical_substance", "named_thing", "biological_entity", "molecular_entity"]}
{"id": {"identifier": "CHEMBL.COMPOUND:CHEMBL1546789", "label": "CHEMBL1546789"}, "equivalent_identifiers": [{"identifier": "CHEMBL.COMPOUND:CHEMBL1546789", "label": "CHEMBL1546789"}, {"identifier": "PUBCHEM:4879549"}, {"identifier": "INCHIKEY:FUIYIXDZTPMQEH-UHFFFAOYSA-N"}], "type": ["chemical_substance", "named_thing", "biological_entity", "molecular_entity"]}

Creating and loading a Redis container with data

A running instance of Redis is needed to house the node normalization data. a Redis Docker container image can be downloaded from Docker hub. The Redis caonteriner can be started with thie following docker command:

 docker run --name node-norm-redis -p 6379:6379 -d redis redis-server --appendonly yes

Note that the dataset for Node normalization is quite large and 256Gb of memory and disk space should be made available to the Redis instance to insure proper loading of the complete compendia.

Configuration

Insure that the ./config.json file is created and contains the parameters for the node normalization load specific to your environment.

The configuration parameters compendium_directory and data_files specify the location of the compendia files. An example of the files' contents
are listed below:

 {
"compendium_directory": "<path to files>",
"data_files": "anatomy.txt,BiologicalProcess.txt,cell.txt,cellular_component.txt,disease.txt,gene_compendium.txt,gene_family_compendium.txt,MolecularActivity.txt,pathways.txt,phenotypes.txt,taxon_compendium.txt",
"redis_host": "<Redis host server name>",
"redis_port": <Redis connection port>,
"redis_password": "<Redis password",
"test_mode": 1,
"debug_messages": 0
}

Loading of the Redis server with compendia data

The load.py script reads the configuration file for load parameters and the loads the compendia data into the Redis instance.

The redis command line can be used to monitor various aspects of the load.

It is possible to observer the progress of the load opening a command line within the container and issuing Redis commands.

View the number of keys loaded so far.

 redis-cli info keyspace

Once the database has completed loading it is recommended that the Redis database be persisted to disk.

 redis-cli save

Monitor the database to determine if the save has completed.

 redis-cli info persistence

Starting the FASTAPI webserver from the command line

The web server can be started after successful completion of the load.

 cd <Node normalization code root>
pip install -r requirements.txt
uvicorn --host 0.0.0.0 --port 8000 --workers 1 node_normalizer.server:app

Then navigate to http://localhost:8000/docs to run the application. Documentation for the NodeNorm API is available.

Webserver Docker container creation and execution

Much like the Redis Docker container noted above, a Docker container can also be created and executed to run the webserver.

Build the webserver Docker image

 cd <Node normalization code root>
docker build --tag <image_tag> .

Start the container:

Note the Dockerfile specifies port 6380 for the webservice container.

 docker run --name Node-normalization -p 8000:6380 node-norm

Then navigate to: http://localhost:8000/docs to run the application

Kubernetes configurations

Kubernetes configurations and helm charts for this project can be found at:

 https://github.com/helxplatform/translator-devops/helm/r3

Configuration

NodeNorm can be configured by setting environmental variables:

  • SERVER_NAME: The name of this server (defaults to infores:sri-node-normalizer)
  • SERVER_ROOT: The server root (defaults to /)
  • LOG_LEVEL: The log level (defaults to ERROR)
  • TRAPI_VERSION: The TRAPI version this version of NodeNorm supports.
  • MATURITY_VALUE: How mature is this NameRes (defaults to maturity, e.g. development)
  • LOCATION_VALUE: Where is this NameRes setup (defaults to location, e.g. RENCI)
  • EQ_BATCH_SIZE: The size of the get_eqids_and_types() batch size (defaults to 2500)
  • OTEL_ENABLED: Turn on Open TELemetry (default: 'false') -- only 'true' will turn this on.
    • JAEGER_HOST and JAEGER_PORT: Hostname and port for the Jaegar instance to provide telemetry to.
    • JAEGER_SERVICE_NAME: The name of this service (defaults to the value of SERVER_NAME)

About

Service that produces Translator compliant nodes given a curie

Topics

Resources

Stars

17 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

704 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

NodeNormalization

DOIarXiv

Introduction

Node normalization takes a CURIE, and returns:

  • The preferred CURIE for this entity
  • All other known equivalent identifiers for the entity
  • Semantic types for the entity as defined by the Biolink Model

The data currently served by Node Normalization is created by the prototype project Babel, which attempts to find identifier equivalences, and makes sure that CURIE prefixes are BioLink Model compliant. The NodeNormalization service, however, is independent of Babel and as improved identifier equivalence tools are developed, their results can be easily incorporated.

To determine whether Node Normalization is likely to be useful, check /get_semantic_types, which lists the BioLink semantic types for which normalization has been attempted, and /get_curie_prefixes, which lists the number of times each prefix is used for a semantic type.

For examples of service usage, see the example notebook.

The Node normalization website leverages the R3 (Redis-REST with referencing) Redis data design and configuration.

Users can find the publicly available website at service.

Installation

Create a virtual environment

 python -m venv nodeNormalization-env

Activate the virtual environment

 source nodeNormalization-env/bin/activate

Install requirements

 > pip install -r requirements.txt

Generating equivalence data

The equivalence data can be generated by running Babel. An example of the contents of a compendia file is shown below:

 {"id": {"identifier": "PUBCHEM:50986940"}, "equivalent_identifiers": [{"identifier": "PUBCHEM:50986940"}, {"identifier": "INCHIKEY:CYMOSKLLKPIPCD-UHFFFAOYSA-N"}], "type": ["chemical_substance", "named_thing", "biological_entity", "molecular_entity"]}
{"id": {"identifier": "CHEMBL.COMPOUND:CHEMBL1546789", "label": "CHEMBL1546789"}, "equivalent_identifiers": [{"identifier": "CHEMBL.COMPOUND:CHEMBL1546789", "label": "CHEMBL1546789"}, {"identifier": "PUBCHEM:4879549"}, {"identifier": "INCHIKEY:FUIYIXDZTPMQEH-UHFFFAOYSA-N"}], "type": ["chemical_substance", "named_thing", "biological_entity", "molecular_entity"]}

Creating and loading a Redis container with data

A running instance of Redis is needed to house the node normalization data. a Redis Docker container image can be downloaded from Docker hub. The Redis caonteriner can be started with thie following docker command:

 docker run --name node-norm-redis -p 6379:6379 -d redis redis-server --appendonly yes

Note that the dataset for Node normalization is quite large and 256Gb of memory and disk space should be made available to the Redis instance to insure proper loading of the complete compendia.

Configuration

Insure that the ./config.json file is created and contains the parameters for the node normalization load specific to your environment.

The configuration parameters compendium_directory and data_files specify the location of the compendia files. An example of the files' contents
are listed below:

 {
"compendium_directory": "<path to files>",
"data_files": "anatomy.txt,BiologicalProcess.txt,cell.txt,cellular_component.txt,disease.txt,gene_compendium.txt,gene_family_compendium.txt,MolecularActivity.txt,pathways.txt,phenotypes.txt,taxon_compendium.txt",
"redis_host": "<Redis host server name>",
"redis_port": <Redis connection port>,
"redis_password": "<Redis password",
"test_mode": 1,
"debug_messages": 0
}

Loading of the Redis server with compendia data

The load.py script reads the configuration file for load parameters and the loads the compendia data into the Redis instance.

The redis command line can be used to monitor various aspects of the load.

It is possible to observer the progress of the load opening a command line within the container and issuing Redis commands.

View the number of keys loaded so far.

 redis-cli info keyspace

Once the database has completed loading it is recommended that the Redis database be persisted to disk.

 redis-cli save

Monitor the database to determine if the save has completed.

 redis-cli info persistence

Starting the FASTAPI webserver from the command line

The web server can be started after successful completion of the load.

 cd <Node normalization code root>
pip install -r requirements.txt
uvicorn --host 0.0.0.0 --port 8000 --workers 1 node_normalizer.server:app

Then navigate to http://localhost:8000/docs to run the application. Documentation for the NodeNorm API is available.

Webserver Docker container creation and execution

Much like the Redis Docker container noted above, a Docker container can also be created and executed to run the webserver.

Build the webserver Docker image

 cd <Node normalization code root>
docker build --tag <image_tag> .

Start the container:

Note the Dockerfile specifies port 6380 for the webservice container.

 docker run --name Node-normalization -p 8000:6380 node-norm

Then navigate to: http://localhost:8000/docs to run the application

Kubernetes configurations

Kubernetes configurations and helm charts for this project can be found at:

 https://github.com/helxplatform/translator-devops/helm/r3

Configuration

NodeNorm can be configured by setting environmental variables:

  • SERVER_NAME: The name of this server (defaults to infores:sri-node-normalizer)
  • SERVER_ROOT: The server root (defaults to /)
  • LOG_LEVEL: The log level (defaults to ERROR)
  • TRAPI_VERSION: The TRAPI version this version of NodeNorm supports.
  • MATURITY_VALUE: How mature is this NameRes (defaults to maturity, e.g. development)
  • LOCATION_VALUE: Where is this NameRes setup (defaults to location, e.g. RENCI)
  • EQ_BATCH_SIZE: The size of the get_eqids_and_types() batch size (defaults to 2500)
  • OTEL_ENABLED: Turn on Open TELemetry (default: 'false') -- only 'true' will turn this on.
    • JAEGER_HOST and JAEGER_PORT: Hostname and port for the Jaegar instance to provide telemetry to.
    • JAEGER_SERVICE_NAME: The name of this service (defaults to the value of SERVER_NAME)

About

Service that produces Translator compliant nodes given a curie

Topics

Resources

Stars

17 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

704 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

NodeNormalization

DOIarXiv

Introduction

Node normalization takes a CURIE, and returns:

  • The preferred CURIE for this entity
  • All other known equivalent identifiers for the entity
  • Semantic types for the entity as defined by the Biolink Model

The data currently served by Node Normalization is created by the prototype project Babel, which attempts to find identifier equivalences, and makes sure that CURIE prefixes are BioLink Model compliant. The NodeNormalization service, however, is independent of Babel and as improved identifier equivalence tools are developed, their results can be easily incorporated.

To determine whether Node Normalization is likely to be useful, check /get_semantic_types, which lists the BioLink semantic types for which normalization has been attempted, and /get_curie_prefixes, which lists the number of times each prefix is used for a semantic type.

For examples of service usage, see the example notebook.

The Node normalization website leverages the R3 (Redis-REST with referencing) Redis data design and configuration.

Users can find the publicly available website at service.

Installation

Create a virtual environment

 python -m venv nodeNormalization-env

Activate the virtual environment

 source nodeNormalization-env/bin/activate

Install requirements

 > pip install -r requirements.txt

Generating equivalence data

The equivalence data can be generated by running Babel. An example of the contents of a compendia file is shown below:

 {"id": {"identifier": "PUBCHEM:50986940"}, "equivalent_identifiers": [{"identifier": "PUBCHEM:50986940"}, {"identifier": "INCHIKEY:CYMOSKLLKPIPCD-UHFFFAOYSA-N"}], "type": ["chemical_substance", "named_thing", "biological_entity", "molecular_entity"]}
{"id": {"identifier": "CHEMBL.COMPOUND:CHEMBL1546789", "label": "CHEMBL1546789"}, "equivalent_identifiers": [{"identifier": "CHEMBL.COMPOUND:CHEMBL1546789", "label": "CHEMBL1546789"}, {"identifier": "PUBCHEM:4879549"}, {"identifier": "INCHIKEY:FUIYIXDZTPMQEH-UHFFFAOYSA-N"}], "type": ["chemical_substance", "named_thing", "biological_entity", "molecular_entity"]}

Creating and loading a Redis container with data

A running instance of Redis is needed to house the node normalization data. a Redis Docker container image can be downloaded from Docker hub. The Redis caonteriner can be started with thie following docker command:

 docker run --name node-norm-redis -p 6379:6379 -d redis redis-server --appendonly yes

Note that the dataset for Node normalization is quite large and 256Gb of memory and disk space should be made available to the Redis instance to insure proper loading of the complete compendia.

Configuration

Insure that the ./config.json file is created and contains the parameters for the node normalization load specific to your environment.

The configuration parameters compendium_directory and data_files specify the location of the compendia files. An example of the files' contents
are listed below:

 {
"compendium_directory": "<path to files>",
"data_files": "anatomy.txt,BiologicalProcess.txt,cell.txt,cellular_component.txt,disease.txt,gene_compendium.txt,gene_family_compendium.txt,MolecularActivity.txt,pathways.txt,phenotypes.txt,taxon_compendium.txt",
"redis_host": "<Redis host server name>",
"redis_port": <Redis connection port>,
"redis_password": "<Redis password",
"test_mode": 1,
"debug_messages": 0
}

Loading of the Redis server with compendia data

The load.py script reads the configuration file for load parameters and the loads the compendia data into the Redis instance.

The redis command line can be used to monitor various aspects of the load.

It is possible to observer the progress of the load opening a command line within the container and issuing Redis commands.

View the number of keys loaded so far.

 redis-cli info keyspace

Once the database has completed loading it is recommended that the Redis database be persisted to disk.

 redis-cli save

Monitor the database to determine if the save has completed.

 redis-cli info persistence

Starting the FASTAPI webserver from the command line

The web server can be started after successful completion of the load.

 cd <Node normalization code root>
pip install -r requirements.txt
uvicorn --host 0.0.0.0 --port 8000 --workers 1 node_normalizer.server:app

Then navigate to http://localhost:8000/docs to run the application. Documentation for the NodeNorm API is available.

Webserver Docker container creation and execution

Much like the Redis Docker container noted above, a Docker container can also be created and executed to run the webserver.

Build the webserver Docker image

 cd <Node normalization code root>
docker build --tag <image_tag> .

Start the container:

Note the Dockerfile specifies port 6380 for the webservice container.

 docker run --name Node-normalization -p 8000:6380 node-norm

Then navigate to: http://localhost:8000/docs to run the application

Kubernetes configurations

Kubernetes configurations and helm charts for this project can be found at:

 https://github.com/helxplatform/translator-devops/helm/r3

Configuration

NodeNorm can be configured by setting environmental variables:

  • SERVER_NAME: The name of this server (defaults to infores:sri-node-normalizer)
  • SERVER_ROOT: The server root (defaults to /)
  • LOG_LEVEL: The log level (defaults to ERROR)
  • TRAPI_VERSION: The TRAPI version this version of NodeNorm supports.
  • MATURITY_VALUE: How mature is this NameRes (defaults to maturity, e.g. development)
  • LOCATION_VALUE: Where is this NameRes setup (defaults to location, e.g. RENCI)
  • EQ_BATCH_SIZE: The size of the get_eqids_and_types() batch size (defaults to 2500)
  • OTEL_ENABLED: Turn on Open TELemetry (default: 'false') -- only 'true' will turn this on.
    • JAEGER_HOST and JAEGER_PORT: Hostname and port for the Jaegar instance to provide telemetry to.
    • JAEGER_SERVICE_NAME: The name of this service (defaults to the value of SERVER_NAME)

About

Service that produces Translator compliant nodes given a curie

Topics

Resources

Stars

17 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

704 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

NodeNormalization

DOIarXiv

Introduction

Node normalization takes a CURIE, and returns:

  • The preferred CURIE for this entity
  • All other known equivalent identifiers for the entity
  • Semantic types for the entity as defined by the Biolink Model

The data currently served by Node Normalization is created by the prototype project Babel, which attempts to find identifier equivalences, and makes sure that CURIE prefixes are BioLink Model compliant. The NodeNormalization service, however, is independent of Babel and as improved identifier equivalence tools are developed, their results can be easily incorporated.

To determine whether Node Normalization is likely to be useful, check /get_semantic_types, which lists the BioLink semantic types for which normalization has been attempted, and /get_curie_prefixes, which lists the number of times each prefix is used for a semantic type.

For examples of service usage, see the example notebook.

The Node normalization website leverages the R3 (Redis-REST with referencing) Redis data design and configuration.

Users can find the publicly available website at service.

Installation

Create a virtual environment

 python -m venv nodeNormalization-env

Activate the virtual environment

 source nodeNormalization-env/bin/activate

Install requirements

 > pip install -r requirements.txt

Generating equivalence data

The equivalence data can be generated by running Babel. An example of the contents of a compendia file is shown below:

 {"id": {"identifier": "PUBCHEM:50986940"}, "equivalent_identifiers": [{"identifier": "PUBCHEM:50986940"}, {"identifier": "INCHIKEY:CYMOSKLLKPIPCD-UHFFFAOYSA-N"}], "type": ["chemical_substance", "named_thing", "biological_entity", "molecular_entity"]}
{"id": {"identifier": "CHEMBL.COMPOUND:CHEMBL1546789", "label": "CHEMBL1546789"}, "equivalent_identifiers": [{"identifier": "CHEMBL.COMPOUND:CHEMBL1546789", "label": "CHEMBL1546789"}, {"identifier": "PUBCHEM:4879549"}, {"identifier": "INCHIKEY:FUIYIXDZTPMQEH-UHFFFAOYSA-N"}], "type": ["chemical_substance", "named_thing", "biological_entity", "molecular_entity"]}

Creating and loading a Redis container with data

A running instance of Redis is needed to house the node normalization data. a Redis Docker container image can be downloaded from Docker hub. The Redis caonteriner can be started with thie following docker command:

 docker run --name node-norm-redis -p 6379:6379 -d redis redis-server --appendonly yes

Note that the dataset for Node normalization is quite large and 256Gb of memory and disk space should be made available to the Redis instance to insure proper loading of the complete compendia.

Configuration

Insure that the ./config.json file is created and contains the parameters for the node normalization load specific to your environment.

The configuration parameters compendium_directory and data_files specify the location of the compendia files. An example of the files' contents
are listed below:

 {
"compendium_directory": "<path to files>",
"data_files": "anatomy.txt,BiologicalProcess.txt,cell.txt,cellular_component.txt,disease.txt,gene_compendium.txt,gene_family_compendium.txt,MolecularActivity.txt,pathways.txt,phenotypes.txt,taxon_compendium.txt",
"redis_host": "<Redis host server name>",
"redis_port": <Redis connection port>,
"redis_password": "<Redis password",
"test_mode": 1,
"debug_messages": 0
}

Loading of the Redis server with compendia data

The load.py script reads the configuration file for load parameters and the loads the compendia data into the Redis instance.

The redis command line can be used to monitor various aspects of the load.

It is possible to observer the progress of the load opening a command line within the container and issuing Redis commands.

View the number of keys loaded so far.

 redis-cli info keyspace

Once the database has completed loading it is recommended that the Redis database be persisted to disk.

 redis-cli save

Monitor the database to determine if the save has completed.

 redis-cli info persistence

Starting the FASTAPI webserver from the command line

The web server can be started after successful completion of the load.

 cd <Node normalization code root>
pip install -r requirements.txt
uvicorn --host 0.0.0.0 --port 8000 --workers 1 node_normalizer.server:app

Then navigate to http://localhost:8000/docs to run the application. Documentation for the NodeNorm API is available.

Webserver Docker container creation and execution

Much like the Redis Docker container noted above, a Docker container can also be created and executed to run the webserver.

Build the webserver Docker image

 cd <Node normalization code root>
docker build --tag <image_tag> .

Start the container:

Note the Dockerfile specifies port 6380 for the webservice container.

 docker run --name Node-normalization -p 8000:6380 node-norm

Then navigate to: http://localhost:8000/docs to run the application

Kubernetes configurations

Kubernetes configurations and helm charts for this project can be found at:

 https://github.com/helxplatform/translator-devops/helm/r3

Configuration

NodeNorm can be configured by setting environmental variables:

  • SERVER_NAME: The name of this server (defaults to infores:sri-node-normalizer)
  • SERVER_ROOT: The server root (defaults to /)
  • LOG_LEVEL: The log level (defaults to ERROR)
  • TRAPI_VERSION: The TRAPI version this version of NodeNorm supports.
  • MATURITY_VALUE: How mature is this NameRes (defaults to maturity, e.g. development)
  • LOCATION_VALUE: Where is this NameRes setup (defaults to location, e.g. RENCI)
  • EQ_BATCH_SIZE: The size of the get_eqids_and_types() batch size (defaults to 2500)
  • OTEL_ENABLED: Turn on Open TELemetry (default: 'false') -- only 'true' will turn this on.
    • JAEGER_HOST and JAEGER_PORT: Hostname and port for the Jaegar instance to provide telemetry to.
    • JAEGER_SERVICE_NAME: The name of this service (defaults to the value of SERVER_NAME)

About

Service that produces Translator compliant nodes given a curie

Topics

Resources

Stars

17 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

704 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

NodeNormalization

DOIarXiv

Introduction

Node normalization takes a CURIE, and returns:

  • The preferred CURIE for this entity
  • All other known equivalent identifiers for the entity
  • Semantic types for the entity as defined by the Biolink Model

The data currently served by Node Normalization is created by the prototype project Babel, which attempts to find identifier equivalences, and makes sure that CURIE prefixes are BioLink Model compliant. The NodeNormalization service, however, is independent of Babel and as improved identifier equivalence tools are developed, their results can be easily incorporated.

To determine whether Node Normalization is likely to be useful, check /get_semantic_types, which lists the BioLink semantic types for which normalization has been attempted, and /get_curie_prefixes, which lists the number of times each prefix is used for a semantic type.

For examples of service usage, see the example notebook.

The Node normalization website leverages the R3 (Redis-REST with referencing) Redis data design and configuration.

Users can find the publicly available website at service.

Installation

Create a virtual environment

 python -m venv nodeNormalization-env

Activate the virtual environment

 source nodeNormalization-env/bin/activate

Install requirements

 > pip install -r requirements.txt

Generating equivalence data

The equivalence data can be generated by running Babel. An example of the contents of a compendia file is shown below:

 {"id": {"identifier": "PUBCHEM:50986940"}, "equivalent_identifiers": [{"identifier": "PUBCHEM:50986940"}, {"identifier": "INCHIKEY:CYMOSKLLKPIPCD-UHFFFAOYSA-N"}], "type": ["chemical_substance", "named_thing", "biological_entity", "molecular_entity"]}
{"id": {"identifier": "CHEMBL.COMPOUND:CHEMBL1546789", "label": "CHEMBL1546789"}, "equivalent_identifiers": [{"identifier": "CHEMBL.COMPOUND:CHEMBL1546789", "label": "CHEMBL1546789"}, {"identifier": "PUBCHEM:4879549"}, {"identifier": "INCHIKEY:FUIYIXDZTPMQEH-UHFFFAOYSA-N"}], "type": ["chemical_substance", "named_thing", "biological_entity", "molecular_entity"]}

Creating and loading a Redis container with data

A running instance of Redis is needed to house the node normalization data. a Redis Docker container image can be downloaded from Docker hub. The Redis caonteriner can be started with thie following docker command:

 docker run --name node-norm-redis -p 6379:6379 -d redis redis-server --appendonly yes

Note that the dataset for Node normalization is quite large and 256Gb of memory and disk space should be made available to the Redis instance to insure proper loading of the complete compendia.

Configuration

Insure that the ./config.json file is created and contains the parameters for the node normalization load specific to your environment.

The configuration parameters compendium_directory and data_files specify the location of the compendia files. An example of the files' contents
are listed below:

 {
"compendium_directory": "<path to files>",
"data_files": "anatomy.txt,BiologicalProcess.txt,cell.txt,cellular_component.txt,disease.txt,gene_compendium.txt,gene_family_compendium.txt,MolecularActivity.txt,pathways.txt,phenotypes.txt,taxon_compendium.txt",
"redis_host": "<Redis host server name>",
"redis_port": <Redis connection port>,
"redis_password": "<Redis password",
"test_mode": 1,
"debug_messages": 0
}

Loading of the Redis server with compendia data

The load.py script reads the configuration file for load parameters and the loads the compendia data into the Redis instance.

The redis command line can be used to monitor various aspects of the load.

It is possible to observer the progress of the load opening a command line within the container and issuing Redis commands.

View the number of keys loaded so far.

 redis-cli info keyspace

Once the database has completed loading it is recommended that the Redis database be persisted to disk.

 redis-cli save

Monitor the database to determine if the save has completed.

 redis-cli info persistence

Starting the FASTAPI webserver from the command line

The web server can be started after successful completion of the load.

 cd <Node normalization code root>
pip install -r requirements.txt
uvicorn --host 0.0.0.0 --port 8000 --workers 1 node_normalizer.server:app

Then navigate to http://localhost:8000/docs to run the application. Documentation for the NodeNorm API is available.

Webserver Docker container creation and execution

Much like the Redis Docker container noted above, a Docker container can also be created and executed to run the webserver.

Build the webserver Docker image

 cd <Node normalization code root>
docker build --tag <image_tag> .

Start the container:

Note the Dockerfile specifies port 6380 for the webservice container.

 docker run --name Node-normalization -p 8000:6380 node-norm

Then navigate to: http://localhost:8000/docs to run the application

Kubernetes configurations

Kubernetes configurations and helm charts for this project can be found at:

 https://github.com/helxplatform/translator-devops/helm/r3

Configuration

NodeNorm can be configured by setting environmental variables:

  • SERVER_NAME: The name of this server (defaults to infores:sri-node-normalizer)
  • SERVER_ROOT: The server root (defaults to /)
  • LOG_LEVEL: The log level (defaults to ERROR)
  • TRAPI_VERSION: The TRAPI version this version of NodeNorm supports.
  • MATURITY_VALUE: How mature is this NameRes (defaults to maturity, e.g. development)
  • LOCATION_VALUE: Where is this NameRes setup (defaults to location, e.g. RENCI)
  • EQ_BATCH_SIZE: The size of the get_eqids_and_types() batch size (defaults to 2500)
  • OTEL_ENABLED: Turn on Open TELemetry (default: 'false') -- only 'true' will turn this on.
    • JAEGER_HOST and JAEGER_PORT: Hostname and port for the Jaegar instance to provide telemetry to.
    • JAEGER_SERVICE_NAME: The name of this service (defaults to the value of SERVER_NAME)

About

Service that produces Translator compliant nodes given a curie

Topics

Resources

Stars

17 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

704 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

NodeNormalization

DOIarXiv

Introduction

Node normalization takes a CURIE, and returns:

  • The preferred CURIE for this entity
  • All other known equivalent identifiers for the entity
  • Semantic types for the entity as defined by the Biolink Model

The data currently served by Node Normalization is created by the prototype project Babel, which attempts to find identifier equivalences, and makes sure that CURIE prefixes are BioLink Model compliant. The NodeNormalization service, however, is independent of Babel and as improved identifier equivalence tools are developed, their results can be easily incorporated.

To determine whether Node Normalization is likely to be useful, check /get_semantic_types, which lists the BioLink semantic types for which normalization has been attempted, and /get_curie_prefixes, which lists the number of times each prefix is used for a semantic type.

For examples of service usage, see the example notebook.

The Node normalization website leverages the R3 (Redis-REST with referencing) Redis data design and configuration.

Users can find the publicly available website at service.

Installation

Create a virtual environment

 python -m venv nodeNormalization-env

Activate the virtual environment

 source nodeNormalization-env/bin/activate

Install requirements

 > pip install -r requirements.txt

Generating equivalence data

The equivalence data can be generated by running Babel. An example of the contents of a compendia file is shown below:

 {"id": {"identifier": "PUBCHEM:50986940"}, "equivalent_identifiers": [{"identifier": "PUBCHEM:50986940"}, {"identifier": "INCHIKEY:CYMOSKLLKPIPCD-UHFFFAOYSA-N"}], "type": ["chemical_substance", "named_thing", "biological_entity", "molecular_entity"]}
{"id": {"identifier": "CHEMBL.COMPOUND:CHEMBL1546789", "label": "CHEMBL1546789"}, "equivalent_identifiers": [{"identifier": "CHEMBL.COMPOUND:CHEMBL1546789", "label": "CHEMBL1546789"}, {"identifier": "PUBCHEM:4879549"}, {"identifier": "INCHIKEY:FUIYIXDZTPMQEH-UHFFFAOYSA-N"}], "type": ["chemical_substance", "named_thing", "biological_entity", "molecular_entity"]}

Creating and loading a Redis container with data

A running instance of Redis is needed to house the node normalization data. a Redis Docker container image can be downloaded from Docker hub. The Redis caonteriner can be started with thie following docker command:

 docker run --name node-norm-redis -p 6379:6379 -d redis redis-server --appendonly yes

Note that the dataset for Node normalization is quite large and 256Gb of memory and disk space should be made available to the Redis instance to insure proper loading of the complete compendia.

Configuration

Insure that the ./config.json file is created and contains the parameters for the node normalization load specific to your environment.

The configuration parameters compendium_directory and data_files specify the location of the compendia files. An example of the files' contents
are listed below:

 {
"compendium_directory": "<path to files>",
"data_files": "anatomy.txt,BiologicalProcess.txt,cell.txt,cellular_component.txt,disease.txt,gene_compendium.txt,gene_family_compendium.txt,MolecularActivity.txt,pathways.txt,phenotypes.txt,taxon_compendium.txt",
"redis_host": "<Redis host server name>",
"redis_port": <Redis connection port>,
"redis_password": "<Redis password",
"test_mode": 1,
"debug_messages": 0
}

Loading of the Redis server with compendia data

The load.py script reads the configuration file for load parameters and the loads the compendia data into the Redis instance.

The redis command line can be used to monitor various aspects of the load.

It is possible to observer the progress of the load opening a command line within the container and issuing Redis commands.

View the number of keys loaded so far.

 redis-cli info keyspace

Once the database has completed loading it is recommended that the Redis database be persisted to disk.

 redis-cli save

Monitor the database to determine if the save has completed.

 redis-cli info persistence

Starting the FASTAPI webserver from the command line

The web server can be started after successful completion of the load.

 cd <Node normalization code root>
pip install -r requirements.txt
uvicorn --host 0.0.0.0 --port 8000 --workers 1 node_normalizer.server:app

Then navigate to http://localhost:8000/docs to run the application. Documentation for the NodeNorm API is available.

Webserver Docker container creation and execution

Much like the Redis Docker container noted above, a Docker container can also be created and executed to run the webserver.

Build the webserver Docker image

 cd <Node normalization code root>
docker build --tag <image_tag> .

Start the container:

Note the Dockerfile specifies port 6380 for the webservice container.

 docker run --name Node-normalization -p 8000:6380 node-norm

Then navigate to: http://localhost:8000/docs to run the application

Kubernetes configurations

Kubernetes configurations and helm charts for this project can be found at:

 https://github.com/helxplatform/translator-devops/helm/r3

Configuration

NodeNorm can be configured by setting environmental variables:

  • SERVER_NAME: The name of this server (defaults to infores:sri-node-normalizer)
  • SERVER_ROOT: The server root (defaults to /)
  • LOG_LEVEL: The log level (defaults to ERROR)
  • TRAPI_VERSION: The TRAPI version this version of NodeNorm supports.
  • MATURITY_VALUE: How mature is this NameRes (defaults to maturity, e.g. development)
  • LOCATION_VALUE: Where is this NameRes setup (defaults to location, e.g. RENCI)
  • EQ_BATCH_SIZE: The size of the get_eqids_and_types() batch size (defaults to 2500)
  • OTEL_ENABLED: Turn on Open TELemetry (default: 'false') -- only 'true' will turn this on.
    • JAEGER_HOST and JAEGER_PORT: Hostname and port for the Jaegar instance to provide telemetry to.
    • JAEGER_SERVICE_NAME: The name of this service (defaults to the value of SERVER_NAME)

About

Service that produces Translator compliant nodes given a curie

Topics

Resources

Stars

17 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages