Skip to content

Repository files navigation

VarSAn - Variant Set Annotator

Xiaoman Xie [xiaoman2@illinois.edu] and Saurabh Sinha

University of Illinois Urbana-Champaign

Table of Contents

  1. Introduction
  2. Installation
  3. Tutorial
  4. Acknowledgments and credits

Introduction

Genotype-to-phenotype studies continue to identify sets of genomic variants associated with diseases, e.g., through GWAS, cataloging of somatic mutations in cancer, or identification of de novo mutations from family-based studies. There is a pressing need to interpret these variants mechanistically, e.g., through characterization of molecular pathways impacted by them. Such insights are especially useful in studies of complex diseases where no single variant explains etiology. Here we present a computational tool called ‘VarSAn’ (Variant Set Analysis) that uses a powerful graph mining algorithm to identify pathways relevant to a given set of variants. VarSAn aggregates diverse information about the user-provided collection of variants, along with prior knowledge about genes and pathways, to provide systems-level insights into those variants.

VarSAn uses a pre-built and configurable heterogeneous network whose nodes represent SNPs, genes and pathways, and edges represent relationships among these entities. The network includes: (1) SNP-gene relationships based on expression quantitative trait loci (eQTL) studies, protein activity impact prediction and genomic proximity, (2) gene-gene relationships based on known interactions among encoded proteins, and (3) gene-pathway membership. VarSAn then takes a “query set” of variants and uses the Random Walk with Restarts algorithm on the network to rank pathway nodes for relevance to the query set, reporting p-values for pathway relevance.

Method Overview

Return to TOC

Installation

Please first clone this repository from Github:

git clone https://github.com/UIUCSinhaLab/VarSAn.git

Please also download the edges in the network from: data.zip The data directory should be unzipped and placed in the downloaded repository.

Tutorial

This section of the README is meant to walk a user through the process of using VarSAn to find pathways that relate to a SNP set of interest. In these examples, we will examine a query set of 798 SNPs.

Creating SNP Query Set

The first step is to create a file that includes the SNPs in the query set. This query set file format should list one SNP on each row with SNPs formatted as 'chromosome_location'. For example, the sample query set file contains:

chr10_121589810
chr2_217098337
chr2_217056046
chr12_28021884

Please put the query file in the gene_sets folder as the sample query set file.

Run VarSAn

VarSAn should be called using the command:

./run_all.sh <snplist> <all> <pathway database> <gene-gene network> <weight> <number of random query set>
ArgumentDescriptionAccepted Value
snplistThe filename of the file with query set variants. File extension should be excluded. For example, "SampleQuery" instead of "SampleQuery.txt" should be passed.e.g. "Sample Query"
allSpecify the tissue for tissue-specific mode or use "all" to run VarSAn in the pan-tissue mode.all, tissue from the list
pathway databasePathway Database. Only Reactome pathways are supported in the Github version."Reactome"
weightIf variants in the query set should be weighted according to the number of connections to control for linkage disequilibrium among non-coding SNPs."1" or "0"
number of random query setThe number of random query sets used to provide a background distribution for the empirical p-value calculation. A higher empirical p-value resolution can be achieved by increasing the number of random query sets, but which would also require longer running time.Any integer, we recommend to use numbers greater than 20.

A sample command could be:

./run_all.sh SampleQuery Breast_Mammary_Tissue Reactome yes 1 10

Acknowledgments and Credits

Edges in the downloadable network include SNP-gene edges.

Edges in the downloadable network include gene-pathway edges. These are based on pathway membership information according to the REACTOME database https://reactome.org/.

Edges in the downloadable network include gene-gene edges. These are based on protein-protein interactions according to the HumanNet database (https://www.inetbio.org/humannet/) and were obtained from https://knoweng.org/kn-overview/

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - UIUCSinhaLab/VarSAn · GitHub
Skip to content

Repository files navigation

VarSAn - Variant Set Annotator

Xiaoman Xie [xiaoman2@illinois.edu] and Saurabh Sinha

University of Illinois Urbana-Champaign

Table of Contents

  1. Introduction
  2. Installation
  3. Tutorial
  4. Acknowledgments and credits

Introduction

Genotype-to-phenotype studies continue to identify sets of genomic variants associated with diseases, e.g., through GWAS, cataloging of somatic mutations in cancer, or identification of de novo mutations from family-based studies. There is a pressing need to interpret these variants mechanistically, e.g., through characterization of molecular pathways impacted by them. Such insights are especially useful in studies of complex diseases where no single variant explains etiology. Here we present a computational tool called ‘VarSAn’ (Variant Set Analysis) that uses a powerful graph mining algorithm to identify pathways relevant to a given set of variants. VarSAn aggregates diverse information about the user-provided collection of variants, along with prior knowledge about genes and pathways, to provide systems-level insights into those variants.

VarSAn uses a pre-built and configurable heterogeneous network whose nodes represent SNPs, genes and pathways, and edges represent relationships among these entities. The network includes: (1) SNP-gene relationships based on expression quantitative trait loci (eQTL) studies, protein activity impact prediction and genomic proximity, (2) gene-gene relationships based on known interactions among encoded proteins, and (3) gene-pathway membership. VarSAn then takes a “query set” of variants and uses the Random Walk with Restarts algorithm on the network to rank pathway nodes for relevance to the query set, reporting p-values for pathway relevance.

Method Overview

Return to TOC

Installation

Please first clone this repository from Github:

git clone https://github.com/UIUCSinhaLab/VarSAn.git

Please also download the edges in the network from: data.zip The data directory should be unzipped and placed in the downloaded repository.

Tutorial

This section of the README is meant to walk a user through the process of using VarSAn to find pathways that relate to a SNP set of interest. In these examples, we will examine a query set of 798 SNPs.

Creating SNP Query Set

The first step is to create a file that includes the SNPs in the query set. This query set file format should list one SNP on each row with SNPs formatted as 'chromosome_location'. For example, the sample query set file contains:

chr10_121589810
chr2_217098337
chr2_217056046
chr12_28021884

Please put the query file in the gene_sets folder as the sample query set file.

Run VarSAn

VarSAn should be called using the command:

./run_all.sh <snplist> <all> <pathway database> <gene-gene network> <weight> <number of random query set>
ArgumentDescriptionAccepted Value
snplistThe filename of the file with query set variants. File extension should be excluded. For example, "SampleQuery" instead of "SampleQuery.txt" should be passed.e.g. "Sample Query"
allSpecify the tissue for tissue-specific mode or use "all" to run VarSAn in the pan-tissue mode.all, tissue from the list
pathway databasePathway Database. Only Reactome pathways are supported in the Github version."Reactome"
weightIf variants in the query set should be weighted according to the number of connections to control for linkage disequilibrium among non-coding SNPs."1" or "0"
number of random query setThe number of random query sets used to provide a background distribution for the empirical p-value calculation. A higher empirical p-value resolution can be achieved by increasing the number of random query sets, but which would also require longer running time.Any integer, we recommend to use numbers greater than 20.

A sample command could be:

./run_all.sh SampleQuery Breast_Mammary_Tissue Reactome yes 1 10

Acknowledgments and Credits

Edges in the downloadable network include SNP-gene edges.

Edges in the downloadable network include gene-pathway edges. These are based on pathway membership information according to the REACTOME database https://reactome.org/.

Edges in the downloadable network include gene-gene edges. These are based on protein-protein interactions according to the HumanNet database (https://www.inetbio.org/humannet/) and were obtained from https://knoweng.org/kn-overview/

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - UIUCSinhaLab/VarSAn · GitHub
Skip to content

Repository files navigation

VarSAn - Variant Set Annotator

Xiaoman Xie [xiaoman2@illinois.edu] and Saurabh Sinha

University of Illinois Urbana-Champaign

Table of Contents

  1. Introduction
  2. Installation
  3. Tutorial
  4. Acknowledgments and credits

Introduction

Genotype-to-phenotype studies continue to identify sets of genomic variants associated with diseases, e.g., through GWAS, cataloging of somatic mutations in cancer, or identification of de novo mutations from family-based studies. There is a pressing need to interpret these variants mechanistically, e.g., through characterization of molecular pathways impacted by them. Such insights are especially useful in studies of complex diseases where no single variant explains etiology. Here we present a computational tool called ‘VarSAn’ (Variant Set Analysis) that uses a powerful graph mining algorithm to identify pathways relevant to a given set of variants. VarSAn aggregates diverse information about the user-provided collection of variants, along with prior knowledge about genes and pathways, to provide systems-level insights into those variants.

VarSAn uses a pre-built and configurable heterogeneous network whose nodes represent SNPs, genes and pathways, and edges represent relationships among these entities. The network includes: (1) SNP-gene relationships based on expression quantitative trait loci (eQTL) studies, protein activity impact prediction and genomic proximity, (2) gene-gene relationships based on known interactions among encoded proteins, and (3) gene-pathway membership. VarSAn then takes a “query set” of variants and uses the Random Walk with Restarts algorithm on the network to rank pathway nodes for relevance to the query set, reporting p-values for pathway relevance.

Method Overview

Return to TOC

Installation

Please first clone this repository from Github:

git clone https://github.com/UIUCSinhaLab/VarSAn.git

Please also download the edges in the network from: data.zip The data directory should be unzipped and placed in the downloaded repository.

Tutorial

This section of the README is meant to walk a user through the process of using VarSAn to find pathways that relate to a SNP set of interest. In these examples, we will examine a query set of 798 SNPs.

Creating SNP Query Set

The first step is to create a file that includes the SNPs in the query set. This query set file format should list one SNP on each row with SNPs formatted as 'chromosome_location'. For example, the sample query set file contains:

chr10_121589810
chr2_217098337
chr2_217056046
chr12_28021884

Please put the query file in the gene_sets folder as the sample query set file.

Run VarSAn

VarSAn should be called using the command:

./run_all.sh <snplist> <all> <pathway database> <gene-gene network> <weight> <number of random query set>
ArgumentDescriptionAccepted Value
snplistThe filename of the file with query set variants. File extension should be excluded. For example, "SampleQuery" instead of "SampleQuery.txt" should be passed.e.g. "Sample Query"
allSpecify the tissue for tissue-specific mode or use "all" to run VarSAn in the pan-tissue mode.all, tissue from the list
pathway databasePathway Database. Only Reactome pathways are supported in the Github version."Reactome"
weightIf variants in the query set should be weighted according to the number of connections to control for linkage disequilibrium among non-coding SNPs."1" or "0"
number of random query setThe number of random query sets used to provide a background distribution for the empirical p-value calculation. A higher empirical p-value resolution can be achieved by increasing the number of random query sets, but which would also require longer running time.Any integer, we recommend to use numbers greater than 20.

A sample command could be:

./run_all.sh SampleQuery Breast_Mammary_Tissue Reactome yes 1 10

Acknowledgments and Credits

Edges in the downloadable network include SNP-gene edges.

Edges in the downloadable network include gene-pathway edges. These are based on pathway membership information according to the REACTOME database https://reactome.org/.

Edges in the downloadable network include gene-gene edges. These are based on protein-protein interactions according to the HumanNet database (https://www.inetbio.org/humannet/) and were obtained from https://knoweng.org/kn-overview/

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - UIUCSinhaLab/VarSAn · GitHub
Skip to content

Repository files navigation

VarSAn - Variant Set Annotator

Xiaoman Xie [xiaoman2@illinois.edu] and Saurabh Sinha

University of Illinois Urbana-Champaign

Table of Contents

  1. Introduction
  2. Installation
  3. Tutorial
  4. Acknowledgments and credits

Introduction

Genotype-to-phenotype studies continue to identify sets of genomic variants associated with diseases, e.g., through GWAS, cataloging of somatic mutations in cancer, or identification of de novo mutations from family-based studies. There is a pressing need to interpret these variants mechanistically, e.g., through characterization of molecular pathways impacted by them. Such insights are especially useful in studies of complex diseases where no single variant explains etiology. Here we present a computational tool called ‘VarSAn’ (Variant Set Analysis) that uses a powerful graph mining algorithm to identify pathways relevant to a given set of variants. VarSAn aggregates diverse information about the user-provided collection of variants, along with prior knowledge about genes and pathways, to provide systems-level insights into those variants.

VarSAn uses a pre-built and configurable heterogeneous network whose nodes represent SNPs, genes and pathways, and edges represent relationships among these entities. The network includes: (1) SNP-gene relationships based on expression quantitative trait loci (eQTL) studies, protein activity impact prediction and genomic proximity, (2) gene-gene relationships based on known interactions among encoded proteins, and (3) gene-pathway membership. VarSAn then takes a “query set” of variants and uses the Random Walk with Restarts algorithm on the network to rank pathway nodes for relevance to the query set, reporting p-values for pathway relevance.

Method Overview

Return to TOC

Installation

Please first clone this repository from Github:

git clone https://github.com/UIUCSinhaLab/VarSAn.git

Please also download the edges in the network from: data.zip The data directory should be unzipped and placed in the downloaded repository.

Tutorial

This section of the README is meant to walk a user through the process of using VarSAn to find pathways that relate to a SNP set of interest. In these examples, we will examine a query set of 798 SNPs.

Creating SNP Query Set

The first step is to create a file that includes the SNPs in the query set. This query set file format should list one SNP on each row with SNPs formatted as 'chromosome_location'. For example, the sample query set file contains:

chr10_121589810
chr2_217098337
chr2_217056046
chr12_28021884

Please put the query file in the gene_sets folder as the sample query set file.

Run VarSAn

VarSAn should be called using the command:

./run_all.sh <snplist> <all> <pathway database> <gene-gene network> <weight> <number of random query set>
ArgumentDescriptionAccepted Value
snplistThe filename of the file with query set variants. File extension should be excluded. For example, "SampleQuery" instead of "SampleQuery.txt" should be passed.e.g. "Sample Query"
allSpecify the tissue for tissue-specific mode or use "all" to run VarSAn in the pan-tissue mode.all, tissue from the list
pathway databasePathway Database. Only Reactome pathways are supported in the Github version."Reactome"
weightIf variants in the query set should be weighted according to the number of connections to control for linkage disequilibrium among non-coding SNPs."1" or "0"
number of random query setThe number of random query sets used to provide a background distribution for the empirical p-value calculation. A higher empirical p-value resolution can be achieved by increasing the number of random query sets, but which would also require longer running time.Any integer, we recommend to use numbers greater than 20.

A sample command could be:

./run_all.sh SampleQuery Breast_Mammary_Tissue Reactome yes 1 10

Acknowledgments and Credits

Edges in the downloadable network include SNP-gene edges.

Edges in the downloadable network include gene-pathway edges. These are based on pathway membership information according to the REACTOME database https://reactome.org/.

Edges in the downloadable network include gene-gene edges. These are based on protein-protein interactions according to the HumanNet database (https://www.inetbio.org/humannet/) and were obtained from https://knoweng.org/kn-overview/

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - UIUCSinhaLab/VarSAn · GitHub
Skip to content

Repository files navigation

VarSAn - Variant Set Annotator

Xiaoman Xie [xiaoman2@illinois.edu] and Saurabh Sinha

University of Illinois Urbana-Champaign

Table of Contents

  1. Introduction
  2. Installation
  3. Tutorial
  4. Acknowledgments and credits

Introduction

Genotype-to-phenotype studies continue to identify sets of genomic variants associated with diseases, e.g., through GWAS, cataloging of somatic mutations in cancer, or identification of de novo mutations from family-based studies. There is a pressing need to interpret these variants mechanistically, e.g., through characterization of molecular pathways impacted by them. Such insights are especially useful in studies of complex diseases where no single variant explains etiology. Here we present a computational tool called ‘VarSAn’ (Variant Set Analysis) that uses a powerful graph mining algorithm to identify pathways relevant to a given set of variants. VarSAn aggregates diverse information about the user-provided collection of variants, along with prior knowledge about genes and pathways, to provide systems-level insights into those variants.

VarSAn uses a pre-built and configurable heterogeneous network whose nodes represent SNPs, genes and pathways, and edges represent relationships among these entities. The network includes: (1) SNP-gene relationships based on expression quantitative trait loci (eQTL) studies, protein activity impact prediction and genomic proximity, (2) gene-gene relationships based on known interactions among encoded proteins, and (3) gene-pathway membership. VarSAn then takes a “query set” of variants and uses the Random Walk with Restarts algorithm on the network to rank pathway nodes for relevance to the query set, reporting p-values for pathway relevance.

Method Overview

Return to TOC

Installation

Please first clone this repository from Github:

git clone https://github.com/UIUCSinhaLab/VarSAn.git

Please also download the edges in the network from: data.zip The data directory should be unzipped and placed in the downloaded repository.

Tutorial

This section of the README is meant to walk a user through the process of using VarSAn to find pathways that relate to a SNP set of interest. In these examples, we will examine a query set of 798 SNPs.

Creating SNP Query Set

The first step is to create a file that includes the SNPs in the query set. This query set file format should list one SNP on each row with SNPs formatted as 'chromosome_location'. For example, the sample query set file contains:

chr10_121589810
chr2_217098337
chr2_217056046
chr12_28021884

Please put the query file in the gene_sets folder as the sample query set file.

Run VarSAn

VarSAn should be called using the command:

./run_all.sh <snplist> <all> <pathway database> <gene-gene network> <weight> <number of random query set>
ArgumentDescriptionAccepted Value
snplistThe filename of the file with query set variants. File extension should be excluded. For example, "SampleQuery" instead of "SampleQuery.txt" should be passed.e.g. "Sample Query"
allSpecify the tissue for tissue-specific mode or use "all" to run VarSAn in the pan-tissue mode.all, tissue from the list
pathway databasePathway Database. Only Reactome pathways are supported in the Github version."Reactome"
weightIf variants in the query set should be weighted according to the number of connections to control for linkage disequilibrium among non-coding SNPs."1" or "0"
number of random query setThe number of random query sets used to provide a background distribution for the empirical p-value calculation. A higher empirical p-value resolution can be achieved by increasing the number of random query sets, but which would also require longer running time.Any integer, we recommend to use numbers greater than 20.

A sample command could be:

./run_all.sh SampleQuery Breast_Mammary_Tissue Reactome yes 1 10

Acknowledgments and Credits

Edges in the downloadable network include SNP-gene edges.

Edges in the downloadable network include gene-pathway edges. These are based on pathway membership information according to the REACTOME database https://reactome.org/.

Edges in the downloadable network include gene-gene edges. These are based on protein-protein interactions according to the HumanNet database (https://www.inetbio.org/humannet/) and were obtained from https://knoweng.org/kn-overview/

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - UIUCSinhaLab/VarSAn · GitHub
Skip to content

Repository files navigation

VarSAn - Variant Set Annotator

Xiaoman Xie [xiaoman2@illinois.edu] and Saurabh Sinha

University of Illinois Urbana-Champaign

Table of Contents

  1. Introduction
  2. Installation
  3. Tutorial
  4. Acknowledgments and credits

Introduction

Genotype-to-phenotype studies continue to identify sets of genomic variants associated with diseases, e.g., through GWAS, cataloging of somatic mutations in cancer, or identification of de novo mutations from family-based studies. There is a pressing need to interpret these variants mechanistically, e.g., through characterization of molecular pathways impacted by them. Such insights are especially useful in studies of complex diseases where no single variant explains etiology. Here we present a computational tool called ‘VarSAn’ (Variant Set Analysis) that uses a powerful graph mining algorithm to identify pathways relevant to a given set of variants. VarSAn aggregates diverse information about the user-provided collection of variants, along with prior knowledge about genes and pathways, to provide systems-level insights into those variants.

VarSAn uses a pre-built and configurable heterogeneous network whose nodes represent SNPs, genes and pathways, and edges represent relationships among these entities. The network includes: (1) SNP-gene relationships based on expression quantitative trait loci (eQTL) studies, protein activity impact prediction and genomic proximity, (2) gene-gene relationships based on known interactions among encoded proteins, and (3) gene-pathway membership. VarSAn then takes a “query set” of variants and uses the Random Walk with Restarts algorithm on the network to rank pathway nodes for relevance to the query set, reporting p-values for pathway relevance.

Method Overview

Return to TOC

Installation

Please first clone this repository from Github:

git clone https://github.com/UIUCSinhaLab/VarSAn.git

Please also download the edges in the network from: data.zip The data directory should be unzipped and placed in the downloaded repository.

Tutorial

This section of the README is meant to walk a user through the process of using VarSAn to find pathways that relate to a SNP set of interest. In these examples, we will examine a query set of 798 SNPs.

Creating SNP Query Set

The first step is to create a file that includes the SNPs in the query set. This query set file format should list one SNP on each row with SNPs formatted as 'chromosome_location'. For example, the sample query set file contains:

chr10_121589810
chr2_217098337
chr2_217056046
chr12_28021884

Please put the query file in the gene_sets folder as the sample query set file.

Run VarSAn

VarSAn should be called using the command:

./run_all.sh <snplist> <all> <pathway database> <gene-gene network> <weight> <number of random query set>
ArgumentDescriptionAccepted Value
snplistThe filename of the file with query set variants. File extension should be excluded. For example, "SampleQuery" instead of "SampleQuery.txt" should be passed.e.g. "Sample Query"
allSpecify the tissue for tissue-specific mode or use "all" to run VarSAn in the pan-tissue mode.all, tissue from the list
pathway databasePathway Database. Only Reactome pathways are supported in the Github version."Reactome"
weightIf variants in the query set should be weighted according to the number of connections to control for linkage disequilibrium among non-coding SNPs."1" or "0"
number of random query setThe number of random query sets used to provide a background distribution for the empirical p-value calculation. A higher empirical p-value resolution can be achieved by increasing the number of random query sets, but which would also require longer running time.Any integer, we recommend to use numbers greater than 20.

A sample command could be:

./run_all.sh SampleQuery Breast_Mammary_Tissue Reactome yes 1 10

Acknowledgments and Credits

Edges in the downloadable network include SNP-gene edges.

Edges in the downloadable network include gene-pathway edges. These are based on pathway membership information according to the REACTOME database https://reactome.org/.

Edges in the downloadable network include gene-gene edges. These are based on protein-protein interactions according to the HumanNet database (https://www.inetbio.org/humannet/) and were obtained from https://knoweng.org/kn-overview/

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - UIUCSinhaLab/VarSAn · GitHub
Skip to content

Repository files navigation

VarSAn - Variant Set Annotator

Xiaoman Xie [xiaoman2@illinois.edu] and Saurabh Sinha

University of Illinois Urbana-Champaign

Table of Contents

  1. Introduction
  2. Installation
  3. Tutorial
  4. Acknowledgments and credits

Introduction

Genotype-to-phenotype studies continue to identify sets of genomic variants associated with diseases, e.g., through GWAS, cataloging of somatic mutations in cancer, or identification of de novo mutations from family-based studies. There is a pressing need to interpret these variants mechanistically, e.g., through characterization of molecular pathways impacted by them. Such insights are especially useful in studies of complex diseases where no single variant explains etiology. Here we present a computational tool called ‘VarSAn’ (Variant Set Analysis) that uses a powerful graph mining algorithm to identify pathways relevant to a given set of variants. VarSAn aggregates diverse information about the user-provided collection of variants, along with prior knowledge about genes and pathways, to provide systems-level insights into those variants.

VarSAn uses a pre-built and configurable heterogeneous network whose nodes represent SNPs, genes and pathways, and edges represent relationships among these entities. The network includes: (1) SNP-gene relationships based on expression quantitative trait loci (eQTL) studies, protein activity impact prediction and genomic proximity, (2) gene-gene relationships based on known interactions among encoded proteins, and (3) gene-pathway membership. VarSAn then takes a “query set” of variants and uses the Random Walk with Restarts algorithm on the network to rank pathway nodes for relevance to the query set, reporting p-values for pathway relevance.

Method Overview

Return to TOC

Installation

Please first clone this repository from Github:

git clone https://github.com/UIUCSinhaLab/VarSAn.git

Please also download the edges in the network from: data.zip The data directory should be unzipped and placed in the downloaded repository.

Tutorial

This section of the README is meant to walk a user through the process of using VarSAn to find pathways that relate to a SNP set of interest. In these examples, we will examine a query set of 798 SNPs.

Creating SNP Query Set

The first step is to create a file that includes the SNPs in the query set. This query set file format should list one SNP on each row with SNPs formatted as 'chromosome_location'. For example, the sample query set file contains:

chr10_121589810
chr2_217098337
chr2_217056046
chr12_28021884

Please put the query file in the gene_sets folder as the sample query set file.

Run VarSAn

VarSAn should be called using the command:

./run_all.sh <snplist> <all> <pathway database> <gene-gene network> <weight> <number of random query set>
ArgumentDescriptionAccepted Value
snplistThe filename of the file with query set variants. File extension should be excluded. For example, "SampleQuery" instead of "SampleQuery.txt" should be passed.e.g. "Sample Query"
allSpecify the tissue for tissue-specific mode or use "all" to run VarSAn in the pan-tissue mode.all, tissue from the list
pathway databasePathway Database. Only Reactome pathways are supported in the Github version."Reactome"
weightIf variants in the query set should be weighted according to the number of connections to control for linkage disequilibrium among non-coding SNPs."1" or "0"
number of random query setThe number of random query sets used to provide a background distribution for the empirical p-value calculation. A higher empirical p-value resolution can be achieved by increasing the number of random query sets, but which would also require longer running time.Any integer, we recommend to use numbers greater than 20.

A sample command could be:

./run_all.sh SampleQuery Breast_Mammary_Tissue Reactome yes 1 10

Acknowledgments and Credits

Edges in the downloadable network include SNP-gene edges.

Edges in the downloadable network include gene-pathway edges. These are based on pathway membership information according to the REACTOME database https://reactome.org/.

Edges in the downloadable network include gene-gene edges. These are based on protein-protein interactions according to the HumanNet database (https://www.inetbio.org/humannet/) and were obtained from https://knoweng.org/kn-overview/

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - UIUCSinhaLab/VarSAn · GitHub
Skip to content

Repository files navigation

VarSAn - Variant Set Annotator

Xiaoman Xie [xiaoman2@illinois.edu] and Saurabh Sinha

University of Illinois Urbana-Champaign

Table of Contents

  1. Introduction
  2. Installation
  3. Tutorial
  4. Acknowledgments and credits

Introduction

Genotype-to-phenotype studies continue to identify sets of genomic variants associated with diseases, e.g., through GWAS, cataloging of somatic mutations in cancer, or identification of de novo mutations from family-based studies. There is a pressing need to interpret these variants mechanistically, e.g., through characterization of molecular pathways impacted by them. Such insights are especially useful in studies of complex diseases where no single variant explains etiology. Here we present a computational tool called ‘VarSAn’ (Variant Set Analysis) that uses a powerful graph mining algorithm to identify pathways relevant to a given set of variants. VarSAn aggregates diverse information about the user-provided collection of variants, along with prior knowledge about genes and pathways, to provide systems-level insights into those variants.

VarSAn uses a pre-built and configurable heterogeneous network whose nodes represent SNPs, genes and pathways, and edges represent relationships among these entities. The network includes: (1) SNP-gene relationships based on expression quantitative trait loci (eQTL) studies, protein activity impact prediction and genomic proximity, (2) gene-gene relationships based on known interactions among encoded proteins, and (3) gene-pathway membership. VarSAn then takes a “query set” of variants and uses the Random Walk with Restarts algorithm on the network to rank pathway nodes for relevance to the query set, reporting p-values for pathway relevance.

Method Overview

Return to TOC

Installation

Please first clone this repository from Github:

git clone https://github.com/UIUCSinhaLab/VarSAn.git

Please also download the edges in the network from: data.zip The data directory should be unzipped and placed in the downloaded repository.

Tutorial

This section of the README is meant to walk a user through the process of using VarSAn to find pathways that relate to a SNP set of interest. In these examples, we will examine a query set of 798 SNPs.

Creating SNP Query Set

The first step is to create a file that includes the SNPs in the query set. This query set file format should list one SNP on each row with SNPs formatted as 'chromosome_location'. For example, the sample query set file contains:

chr10_121589810
chr2_217098337
chr2_217056046
chr12_28021884

Please put the query file in the gene_sets folder as the sample query set file.

Run VarSAn

VarSAn should be called using the command:

./run_all.sh <snplist> <all> <pathway database> <gene-gene network> <weight> <number of random query set>
ArgumentDescriptionAccepted Value
snplistThe filename of the file with query set variants. File extension should be excluded. For example, "SampleQuery" instead of "SampleQuery.txt" should be passed.e.g. "Sample Query"
allSpecify the tissue for tissue-specific mode or use "all" to run VarSAn in the pan-tissue mode.all, tissue from the list
pathway databasePathway Database. Only Reactome pathways are supported in the Github version."Reactome"
weightIf variants in the query set should be weighted according to the number of connections to control for linkage disequilibrium among non-coding SNPs."1" or "0"
number of random query setThe number of random query sets used to provide a background distribution for the empirical p-value calculation. A higher empirical p-value resolution can be achieved by increasing the number of random query sets, but which would also require longer running time.Any integer, we recommend to use numbers greater than 20.

A sample command could be:

./run_all.sh SampleQuery Breast_Mammary_Tissue Reactome yes 1 10

Acknowledgments and Credits

Edges in the downloadable network include SNP-gene edges.

Edges in the downloadable network include gene-pathway edges. These are based on pathway membership information according to the REACTOME database https://reactome.org/.

Edges in the downloadable network include gene-gene edges. These are based on protein-protein interactions according to the HumanNet database (https://www.inetbio.org/humannet/) and were obtained from https://knoweng.org/kn-overview/

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages