Latest commit

History

50 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

cWords is a software tool that measure correlations of short oligonucleotide sequence (word) occurrences and i.e. expression changes in a two condition experiment. In summary, it produces a statistic that quantifies how over-represented a word is in a ranked list of sequences.

1. REQUIREMENTS

All software components require Ruby (2.2 or newer, http://www.ruby-lang.org/) and JRuby (9.0.5.0 or newer, http://jruby.org/). The software has only been tested on a Unix platform (both Linux and Mac OS X Capitan).

2. INSTALL

  • Install Ruby (www.ruby-lang.org, check if it is already installed: 'ruby -v')
  • Install Java (www.java.com, check if it is already installed: 'java -version'),
  • Install JRruby (www.jruby.org, check if it is already installed: 'jruby -v'), make sure that you have the 'jruby' command in your path.

3. USAGE

The full list of options for each script is available by running the program with -h flag. Here the main options are described.

Options:

Usage: cwords [options]
-w, --wordsize ARG word length(s) you wish to search in (default 6,7)
-b, --bg ARG Order of Markov background nucleotide model (default 0)
-t, --threads ARG use multiple threads to parallelize computations (default 1)
-C, --custom_IDs Use your own sequences with matching IDs in rank and sequence files
-A, --anders_ids Use integer IDs
-x, --rank_split_mean Split ranked list at mean
-r, --rankfile ARG Rank file with IDs and one or more columns (will calculate mean across columns) of a metric of expression changes (tab or space delimiter) or just one column with ordered IDs most down-regulated to most up-regulated
-s, --seqfile ARG Sequence file - Rank and sequence IDs should be one of the compatible IDs for the species you use
--gen_plot ARG Generate plot data and plots for the top k words, takes 2 passes.
--mkplot ARG Make Word Cluster Plot - highlight (with black border) the 8mer seed site (ex for miR-1: ACATTCCA) and its corresponding 7mer and 6mer seed sites. To highlight nothing write a word not in the [acgt] alphabet.
-N, --noAnn No miRNA-annotation on the Word Cluster Plot
--annotFile ARG Supply you own annotations, for Word Cluster Plot and word ranking.
--species ARG Different ID systems are used for different species, what's the species of your data? Currently we support Human Ensembl sequences as default (write human), Mouse (write mouse), Fruit Fly (write fruitfly) and C. elegans (write roundworm)

Example runs:

cWords is built to utilize many cores in parallel on a large computer. The following are standard analyses to run with cWords.

tri-nucleotide background model, word sizes 6, 7 and 8, using 40 processors
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 2 -p 40
IDs in you sequences and rank-file match
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 2 -C
Mononucleotide background model, word cluster plot (highlighting ACATTCCA, ACATTCC, CATTCCA, ACATTC, CATTCC), Enrichment profile graphs for the 20 most significant words
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 0 -mkplot ACATTCCA -gen_plot 20

4. INTERPRETATION

Results mainly compose of three elements. A ranking of most strongly correlated words, a word cluster plot and enrichment profile plots.

Positive and negative set

The analysis is divided into 2 passes; one where words that are over-represented in up-regulated genes (ie. positively correlated words) and one where negatively correlated words are evaluated. If all genes are included in the analysis in each pass, all words of the length in question will be divided into the negative set and positive set except the words that occur 5 or less times across all sequences. These sets can be divided in different ways (see options -h), and you can consider only most regulated genes in the two passes. If you consider the most down-regulated words in the negative pass and the most up-regulated in the positive pass, the same word will be present in both the negative and the positive set.

The output

The final output of this analysis produces a summary of the top correlating words (a list for each end of the ranked list), i.e. words over-represented in beginning of list:

Top 10 words
rank word z-score p-value fdr ledge annotation 1 cactgcc 22.54 1.00e-10 6.58e-09 1651 hsa-miR-34a-5p,hsa-miR-34c-5p,hsa-miR-449a,hsa-miR-449b-5p
2 actgcca 20.62 1.00e-10 6.98e-09 1115 hsa-miR-34a-5p,hsa-miR-34c-5p,hsa-miR-449a,hsa-miR-449b-5p,hsa-miR-548au-3p
3 cctgccc 20.54 1.00e-10 6.03e-09 2702 hsa-miR-6721-5p
4 ccctgcc 19.82 1.00e-10 6.17e-09 3252 hsa-miR-1207-5p,hsa-miR-4763-3p
5 ccctggg 18.74 1.00e-10 6.22e-09 3500 hsa-miR-1915-3p
6 ctgcccc 17.16 1.00e-10 5.73e-09 2546 hsa-miR-486-3p 7 ctgccct 16.97 1.00e-10 5.69e-09 3248 hsa-miR-4632-5p,hsa-miR-4436b-3p
8 ccagccc 16.54 1.00e-10 6.47e-09 2690 hsa-miR-762,hsa-miR-4492,hsa-miR-4498,hsa-miR-5001-5p
9 ccccagc 16.53 1.00e-10 6.32e-09 2706 hsa-miR-4731-5p
10 ctgggcc 16.42 1.00e-10 5.90e-09 3524 ...
  • 'z-score' is a correlation statistic for the given word after correction for correlations obtained from random gene list orderings.
  • 'fdr' (false discovery rate) is the estimated proportion of false discoveries for the given z-score threshold.
  • 'ledge' is the leading-edge which denotes the position in the gene list where the maximum imbalance was measured; genes before this threshold are relatively enriched for the word compared to genes after the threshold.

Invalid IDs

If you do not use -C the IDs in the rank-file will be mapped to the sequences, when it was not possible to map the rank-file ID to a sequence we report the ID as invalid and. Problems with many invalid IDs typically occur when one uses ID associated with different versions or even assemblies (like sequences hg19 and microArray probes from hg18). Invalid IDs are written to a file facilitating further investigation (file name: invalid_ids.txt).

Computer resources

Memory consumption and running time can vary a lot. The word length (-w) has a significant impact on this, if you want to run for word lengths > 8 works best with 4 GB memory or more (depending on number of genes in your analysis). Word lengths > 9 works best wih 10 GB or more.

Generally running time grows exponentially with word length and linearly with the number of bases in the sequences that need to be analysed.

5. How to cite

Rasmussen S, Jacobsen A and Krogh A;cWords - systematic microRNA regulatory motif discovery from mRNA expression data; Silence (2013)

6. LICENSE

Copyright (c) 2011, Simon H. Rasmussen. The software is open source and released under the MIT license.

About

Regulatory motif discovery tool.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

50 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

cWords is a software tool that measure correlations of short oligonucleotide sequence (word) occurrences and i.e. expression changes in a two condition experiment. In summary, it produces a statistic that quantifies how over-represented a word is in a ranked list of sequences.

1. REQUIREMENTS

All software components require Ruby (2.2 or newer, http://www.ruby-lang.org/) and JRuby (9.0.5.0 or newer, http://jruby.org/). The software has only been tested on a Unix platform (both Linux and Mac OS X Capitan).

2. INSTALL

  • Install Ruby (www.ruby-lang.org, check if it is already installed: 'ruby -v')
  • Install Java (www.java.com, check if it is already installed: 'java -version'),
  • Install JRruby (www.jruby.org, check if it is already installed: 'jruby -v'), make sure that you have the 'jruby' command in your path.

3. USAGE

The full list of options for each script is available by running the program with -h flag. Here the main options are described.

Options:

Usage: cwords [options]
-w, --wordsize ARG word length(s) you wish to search in (default 6,7)
-b, --bg ARG Order of Markov background nucleotide model (default 0)
-t, --threads ARG use multiple threads to parallelize computations (default 1)
-C, --custom_IDs Use your own sequences with matching IDs in rank and sequence files
-A, --anders_ids Use integer IDs
-x, --rank_split_mean Split ranked list at mean
-r, --rankfile ARG Rank file with IDs and one or more columns (will calculate mean across columns) of a metric of expression changes (tab or space delimiter) or just one column with ordered IDs most down-regulated to most up-regulated
-s, --seqfile ARG Sequence file - Rank and sequence IDs should be one of the compatible IDs for the species you use
--gen_plot ARG Generate plot data and plots for the top k words, takes 2 passes.
--mkplot ARG Make Word Cluster Plot - highlight (with black border) the 8mer seed site (ex for miR-1: ACATTCCA) and its corresponding 7mer and 6mer seed sites. To highlight nothing write a word not in the [acgt] alphabet.
-N, --noAnn No miRNA-annotation on the Word Cluster Plot
--annotFile ARG Supply you own annotations, for Word Cluster Plot and word ranking.
--species ARG Different ID systems are used for different species, what's the species of your data? Currently we support Human Ensembl sequences as default (write human), Mouse (write mouse), Fruit Fly (write fruitfly) and C. elegans (write roundworm)

Example runs:

cWords is built to utilize many cores in parallel on a large computer. The following are standard analyses to run with cWords.

tri-nucleotide background model, word sizes 6, 7 and 8, using 40 processors
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 2 -p 40
IDs in you sequences and rank-file match
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 2 -C
Mononucleotide background model, word cluster plot (highlighting ACATTCCA, ACATTCC, CATTCCA, ACATTC, CATTCC), Enrichment profile graphs for the 20 most significant words
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 0 -mkplot ACATTCCA -gen_plot 20

4. INTERPRETATION

Results mainly compose of three elements. A ranking of most strongly correlated words, a word cluster plot and enrichment profile plots.

Positive and negative set

The analysis is divided into 2 passes; one where words that are over-represented in up-regulated genes (ie. positively correlated words) and one where negatively correlated words are evaluated. If all genes are included in the analysis in each pass, all words of the length in question will be divided into the negative set and positive set except the words that occur 5 or less times across all sequences. These sets can be divided in different ways (see options -h), and you can consider only most regulated genes in the two passes. If you consider the most down-regulated words in the negative pass and the most up-regulated in the positive pass, the same word will be present in both the negative and the positive set.

The output

The final output of this analysis produces a summary of the top correlating words (a list for each end of the ranked list), i.e. words over-represented in beginning of list:

Top 10 words
rank word z-score p-value fdr ledge annotation 1 cactgcc 22.54 1.00e-10 6.58e-09 1651 hsa-miR-34a-5p,hsa-miR-34c-5p,hsa-miR-449a,hsa-miR-449b-5p
2 actgcca 20.62 1.00e-10 6.98e-09 1115 hsa-miR-34a-5p,hsa-miR-34c-5p,hsa-miR-449a,hsa-miR-449b-5p,hsa-miR-548au-3p
3 cctgccc 20.54 1.00e-10 6.03e-09 2702 hsa-miR-6721-5p
4 ccctgcc 19.82 1.00e-10 6.17e-09 3252 hsa-miR-1207-5p,hsa-miR-4763-3p
5 ccctggg 18.74 1.00e-10 6.22e-09 3500 hsa-miR-1915-3p
6 ctgcccc 17.16 1.00e-10 5.73e-09 2546 hsa-miR-486-3p 7 ctgccct 16.97 1.00e-10 5.69e-09 3248 hsa-miR-4632-5p,hsa-miR-4436b-3p
8 ccagccc 16.54 1.00e-10 6.47e-09 2690 hsa-miR-762,hsa-miR-4492,hsa-miR-4498,hsa-miR-5001-5p
9 ccccagc 16.53 1.00e-10 6.32e-09 2706 hsa-miR-4731-5p
10 ctgggcc 16.42 1.00e-10 5.90e-09 3524 ...
  • 'z-score' is a correlation statistic for the given word after correction for correlations obtained from random gene list orderings.
  • 'fdr' (false discovery rate) is the estimated proportion of false discoveries for the given z-score threshold.
  • 'ledge' is the leading-edge which denotes the position in the gene list where the maximum imbalance was measured; genes before this threshold are relatively enriched for the word compared to genes after the threshold.

Invalid IDs

If you do not use -C the IDs in the rank-file will be mapped to the sequences, when it was not possible to map the rank-file ID to a sequence we report the ID as invalid and. Problems with many invalid IDs typically occur when one uses ID associated with different versions or even assemblies (like sequences hg19 and microArray probes from hg18). Invalid IDs are written to a file facilitating further investigation (file name: invalid_ids.txt).

Computer resources

Memory consumption and running time can vary a lot. The word length (-w) has a significant impact on this, if you want to run for word lengths > 8 works best with 4 GB memory or more (depending on number of genes in your analysis). Word lengths > 9 works best wih 10 GB or more.

Generally running time grows exponentially with word length and linearly with the number of bases in the sequences that need to be analysed.

5. How to cite

Rasmussen S, Jacobsen A and Krogh A;cWords - systematic microRNA regulatory motif discovery from mRNA expression data; Silence (2013)

6. LICENSE

Copyright (c) 2011, Simon H. Rasmussen. The software is open source and released under the MIT license.

About

Regulatory motif discovery tool.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

50 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

cWords is a software tool that measure correlations of short oligonucleotide sequence (word) occurrences and i.e. expression changes in a two condition experiment. In summary, it produces a statistic that quantifies how over-represented a word is in a ranked list of sequences.

1. REQUIREMENTS

All software components require Ruby (2.2 or newer, http://www.ruby-lang.org/) and JRuby (9.0.5.0 or newer, http://jruby.org/). The software has only been tested on a Unix platform (both Linux and Mac OS X Capitan).

2. INSTALL

  • Install Ruby (www.ruby-lang.org, check if it is already installed: 'ruby -v')
  • Install Java (www.java.com, check if it is already installed: 'java -version'),
  • Install JRruby (www.jruby.org, check if it is already installed: 'jruby -v'), make sure that you have the 'jruby' command in your path.

3. USAGE

The full list of options for each script is available by running the program with -h flag. Here the main options are described.

Options:

Usage: cwords [options]
-w, --wordsize ARG word length(s) you wish to search in (default 6,7)
-b, --bg ARG Order of Markov background nucleotide model (default 0)
-t, --threads ARG use multiple threads to parallelize computations (default 1)
-C, --custom_IDs Use your own sequences with matching IDs in rank and sequence files
-A, --anders_ids Use integer IDs
-x, --rank_split_mean Split ranked list at mean
-r, --rankfile ARG Rank file with IDs and one or more columns (will calculate mean across columns) of a metric of expression changes (tab or space delimiter) or just one column with ordered IDs most down-regulated to most up-regulated
-s, --seqfile ARG Sequence file - Rank and sequence IDs should be one of the compatible IDs for the species you use
--gen_plot ARG Generate plot data and plots for the top k words, takes 2 passes.
--mkplot ARG Make Word Cluster Plot - highlight (with black border) the 8mer seed site (ex for miR-1: ACATTCCA) and its corresponding 7mer and 6mer seed sites. To highlight nothing write a word not in the [acgt] alphabet.
-N, --noAnn No miRNA-annotation on the Word Cluster Plot
--annotFile ARG Supply you own annotations, for Word Cluster Plot and word ranking.
--species ARG Different ID systems are used for different species, what's the species of your data? Currently we support Human Ensembl sequences as default (write human), Mouse (write mouse), Fruit Fly (write fruitfly) and C. elegans (write roundworm)

Example runs:

cWords is built to utilize many cores in parallel on a large computer. The following are standard analyses to run with cWords.

tri-nucleotide background model, word sizes 6, 7 and 8, using 40 processors
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 2 -p 40
IDs in you sequences and rank-file match
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 2 -C
Mononucleotide background model, word cluster plot (highlighting ACATTCCA, ACATTCC, CATTCCA, ACATTC, CATTCC), Enrichment profile graphs for the 20 most significant words
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 0 -mkplot ACATTCCA -gen_plot 20

4. INTERPRETATION

Results mainly compose of three elements. A ranking of most strongly correlated words, a word cluster plot and enrichment profile plots.

Positive and negative set

The analysis is divided into 2 passes; one where words that are over-represented in up-regulated genes (ie. positively correlated words) and one where negatively correlated words are evaluated. If all genes are included in the analysis in each pass, all words of the length in question will be divided into the negative set and positive set except the words that occur 5 or less times across all sequences. These sets can be divided in different ways (see options -h), and you can consider only most regulated genes in the two passes. If you consider the most down-regulated words in the negative pass and the most up-regulated in the positive pass, the same word will be present in both the negative and the positive set.

The output

The final output of this analysis produces a summary of the top correlating words (a list for each end of the ranked list), i.e. words over-represented in beginning of list:

Top 10 words
rank word z-score p-value fdr ledge annotation 1 cactgcc 22.54 1.00e-10 6.58e-09 1651 hsa-miR-34a-5p,hsa-miR-34c-5p,hsa-miR-449a,hsa-miR-449b-5p
2 actgcca 20.62 1.00e-10 6.98e-09 1115 hsa-miR-34a-5p,hsa-miR-34c-5p,hsa-miR-449a,hsa-miR-449b-5p,hsa-miR-548au-3p
3 cctgccc 20.54 1.00e-10 6.03e-09 2702 hsa-miR-6721-5p
4 ccctgcc 19.82 1.00e-10 6.17e-09 3252 hsa-miR-1207-5p,hsa-miR-4763-3p
5 ccctggg 18.74 1.00e-10 6.22e-09 3500 hsa-miR-1915-3p
6 ctgcccc 17.16 1.00e-10 5.73e-09 2546 hsa-miR-486-3p 7 ctgccct 16.97 1.00e-10 5.69e-09 3248 hsa-miR-4632-5p,hsa-miR-4436b-3p
8 ccagccc 16.54 1.00e-10 6.47e-09 2690 hsa-miR-762,hsa-miR-4492,hsa-miR-4498,hsa-miR-5001-5p
9 ccccagc 16.53 1.00e-10 6.32e-09 2706 hsa-miR-4731-5p
10 ctgggcc 16.42 1.00e-10 5.90e-09 3524 ...
  • 'z-score' is a correlation statistic for the given word after correction for correlations obtained from random gene list orderings.
  • 'fdr' (false discovery rate) is the estimated proportion of false discoveries for the given z-score threshold.
  • 'ledge' is the leading-edge which denotes the position in the gene list where the maximum imbalance was measured; genes before this threshold are relatively enriched for the word compared to genes after the threshold.

Invalid IDs

If you do not use -C the IDs in the rank-file will be mapped to the sequences, when it was not possible to map the rank-file ID to a sequence we report the ID as invalid and. Problems with many invalid IDs typically occur when one uses ID associated with different versions or even assemblies (like sequences hg19 and microArray probes from hg18). Invalid IDs are written to a file facilitating further investigation (file name: invalid_ids.txt).

Computer resources

Memory consumption and running time can vary a lot. The word length (-w) has a significant impact on this, if you want to run for word lengths > 8 works best with 4 GB memory or more (depending on number of genes in your analysis). Word lengths > 9 works best wih 10 GB or more.

Generally running time grows exponentially with word length and linearly with the number of bases in the sequences that need to be analysed.

5. How to cite

Rasmussen S, Jacobsen A and Krogh A;cWords - systematic microRNA regulatory motif discovery from mRNA expression data; Silence (2013)

6. LICENSE

Copyright (c) 2011, Simon H. Rasmussen. The software is open source and released under the MIT license.

About

Regulatory motif discovery tool.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

50 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

cWords is a software tool that measure correlations of short oligonucleotide sequence (word) occurrences and i.e. expression changes in a two condition experiment. In summary, it produces a statistic that quantifies how over-represented a word is in a ranked list of sequences.

1. REQUIREMENTS

All software components require Ruby (2.2 or newer, http://www.ruby-lang.org/) and JRuby (9.0.5.0 or newer, http://jruby.org/). The software has only been tested on a Unix platform (both Linux and Mac OS X Capitan).

2. INSTALL

  • Install Ruby (www.ruby-lang.org, check if it is already installed: 'ruby -v')
  • Install Java (www.java.com, check if it is already installed: 'java -version'),
  • Install JRruby (www.jruby.org, check if it is already installed: 'jruby -v'), make sure that you have the 'jruby' command in your path.

3. USAGE

The full list of options for each script is available by running the program with -h flag. Here the main options are described.

Options:

Usage: cwords [options]
-w, --wordsize ARG word length(s) you wish to search in (default 6,7)
-b, --bg ARG Order of Markov background nucleotide model (default 0)
-t, --threads ARG use multiple threads to parallelize computations (default 1)
-C, --custom_IDs Use your own sequences with matching IDs in rank and sequence files
-A, --anders_ids Use integer IDs
-x, --rank_split_mean Split ranked list at mean
-r, --rankfile ARG Rank file with IDs and one or more columns (will calculate mean across columns) of a metric of expression changes (tab or space delimiter) or just one column with ordered IDs most down-regulated to most up-regulated
-s, --seqfile ARG Sequence file - Rank and sequence IDs should be one of the compatible IDs for the species you use
--gen_plot ARG Generate plot data and plots for the top k words, takes 2 passes.
--mkplot ARG Make Word Cluster Plot - highlight (with black border) the 8mer seed site (ex for miR-1: ACATTCCA) and its corresponding 7mer and 6mer seed sites. To highlight nothing write a word not in the [acgt] alphabet.
-N, --noAnn No miRNA-annotation on the Word Cluster Plot
--annotFile ARG Supply you own annotations, for Word Cluster Plot and word ranking.
--species ARG Different ID systems are used for different species, what's the species of your data? Currently we support Human Ensembl sequences as default (write human), Mouse (write mouse), Fruit Fly (write fruitfly) and C. elegans (write roundworm)

Example runs:

cWords is built to utilize many cores in parallel on a large computer. The following are standard analyses to run with cWords.

tri-nucleotide background model, word sizes 6, 7 and 8, using 40 processors
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 2 -p 40
IDs in you sequences and rank-file match
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 2 -C
Mononucleotide background model, word cluster plot (highlighting ACATTCCA, ACATTCC, CATTCCA, ACATTC, CATTCC), Enrichment profile graphs for the 20 most significant words
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 0 -mkplot ACATTCCA -gen_plot 20

4. INTERPRETATION

Results mainly compose of three elements. A ranking of most strongly correlated words, a word cluster plot and enrichment profile plots.

Positive and negative set

The analysis is divided into 2 passes; one where words that are over-represented in up-regulated genes (ie. positively correlated words) and one where negatively correlated words are evaluated. If all genes are included in the analysis in each pass, all words of the length in question will be divided into the negative set and positive set except the words that occur 5 or less times across all sequences. These sets can be divided in different ways (see options -h), and you can consider only most regulated genes in the two passes. If you consider the most down-regulated words in the negative pass and the most up-regulated in the positive pass, the same word will be present in both the negative and the positive set.

The output

The final output of this analysis produces a summary of the top correlating words (a list for each end of the ranked list), i.e. words over-represented in beginning of list:

Top 10 words
rank word z-score p-value fdr ledge annotation 1 cactgcc 22.54 1.00e-10 6.58e-09 1651 hsa-miR-34a-5p,hsa-miR-34c-5p,hsa-miR-449a,hsa-miR-449b-5p
2 actgcca 20.62 1.00e-10 6.98e-09 1115 hsa-miR-34a-5p,hsa-miR-34c-5p,hsa-miR-449a,hsa-miR-449b-5p,hsa-miR-548au-3p
3 cctgccc 20.54 1.00e-10 6.03e-09 2702 hsa-miR-6721-5p
4 ccctgcc 19.82 1.00e-10 6.17e-09 3252 hsa-miR-1207-5p,hsa-miR-4763-3p
5 ccctggg 18.74 1.00e-10 6.22e-09 3500 hsa-miR-1915-3p
6 ctgcccc 17.16 1.00e-10 5.73e-09 2546 hsa-miR-486-3p 7 ctgccct 16.97 1.00e-10 5.69e-09 3248 hsa-miR-4632-5p,hsa-miR-4436b-3p
8 ccagccc 16.54 1.00e-10 6.47e-09 2690 hsa-miR-762,hsa-miR-4492,hsa-miR-4498,hsa-miR-5001-5p
9 ccccagc 16.53 1.00e-10 6.32e-09 2706 hsa-miR-4731-5p
10 ctgggcc 16.42 1.00e-10 5.90e-09 3524 ...
  • 'z-score' is a correlation statistic for the given word after correction for correlations obtained from random gene list orderings.
  • 'fdr' (false discovery rate) is the estimated proportion of false discoveries for the given z-score threshold.
  • 'ledge' is the leading-edge which denotes the position in the gene list where the maximum imbalance was measured; genes before this threshold are relatively enriched for the word compared to genes after the threshold.

Invalid IDs

If you do not use -C the IDs in the rank-file will be mapped to the sequences, when it was not possible to map the rank-file ID to a sequence we report the ID as invalid and. Problems with many invalid IDs typically occur when one uses ID associated with different versions or even assemblies (like sequences hg19 and microArray probes from hg18). Invalid IDs are written to a file facilitating further investigation (file name: invalid_ids.txt).

Computer resources

Memory consumption and running time can vary a lot. The word length (-w) has a significant impact on this, if you want to run for word lengths > 8 works best with 4 GB memory or more (depending on number of genes in your analysis). Word lengths > 9 works best wih 10 GB or more.

Generally running time grows exponentially with word length and linearly with the number of bases in the sequences that need to be analysed.

5. How to cite

Rasmussen S, Jacobsen A and Krogh A;cWords - systematic microRNA regulatory motif discovery from mRNA expression data; Silence (2013)

6. LICENSE

Copyright (c) 2011, Simon H. Rasmussen. The software is open source and released under the MIT license.

About

Regulatory motif discovery tool.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

50 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

cWords is a software tool that measure correlations of short oligonucleotide sequence (word) occurrences and i.e. expression changes in a two condition experiment. In summary, it produces a statistic that quantifies how over-represented a word is in a ranked list of sequences.

1. REQUIREMENTS

All software components require Ruby (2.2 or newer, http://www.ruby-lang.org/) and JRuby (9.0.5.0 or newer, http://jruby.org/). The software has only been tested on a Unix platform (both Linux and Mac OS X Capitan).

2. INSTALL

  • Install Ruby (www.ruby-lang.org, check if it is already installed: 'ruby -v')
  • Install Java (www.java.com, check if it is already installed: 'java -version'),
  • Install JRruby (www.jruby.org, check if it is already installed: 'jruby -v'), make sure that you have the 'jruby' command in your path.

3. USAGE

The full list of options for each script is available by running the program with -h flag. Here the main options are described.

Options:

Usage: cwords [options]
-w, --wordsize ARG word length(s) you wish to search in (default 6,7)
-b, --bg ARG Order of Markov background nucleotide model (default 0)
-t, --threads ARG use multiple threads to parallelize computations (default 1)
-C, --custom_IDs Use your own sequences with matching IDs in rank and sequence files
-A, --anders_ids Use integer IDs
-x, --rank_split_mean Split ranked list at mean
-r, --rankfile ARG Rank file with IDs and one or more columns (will calculate mean across columns) of a metric of expression changes (tab or space delimiter) or just one column with ordered IDs most down-regulated to most up-regulated
-s, --seqfile ARG Sequence file - Rank and sequence IDs should be one of the compatible IDs for the species you use
--gen_plot ARG Generate plot data and plots for the top k words, takes 2 passes.
--mkplot ARG Make Word Cluster Plot - highlight (with black border) the 8mer seed site (ex for miR-1: ACATTCCA) and its corresponding 7mer and 6mer seed sites. To highlight nothing write a word not in the [acgt] alphabet.
-N, --noAnn No miRNA-annotation on the Word Cluster Plot
--annotFile ARG Supply you own annotations, for Word Cluster Plot and word ranking.
--species ARG Different ID systems are used for different species, what's the species of your data? Currently we support Human Ensembl sequences as default (write human), Mouse (write mouse), Fruit Fly (write fruitfly) and C. elegans (write roundworm)

Example runs:

cWords is built to utilize many cores in parallel on a large computer. The following are standard analyses to run with cWords.

tri-nucleotide background model, word sizes 6, 7 and 8, using 40 processors
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 2 -p 40
IDs in you sequences and rank-file match
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 2 -C
Mononucleotide background model, word cluster plot (highlighting ACATTCCA, ACATTCC, CATTCCA, ACATTC, CATTCC), Enrichment profile graphs for the 20 most significant words
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 0 -mkplot ACATTCCA -gen_plot 20

4. INTERPRETATION

Results mainly compose of three elements. A ranking of most strongly correlated words, a word cluster plot and enrichment profile plots.

Positive and negative set

The analysis is divided into 2 passes; one where words that are over-represented in up-regulated genes (ie. positively correlated words) and one where negatively correlated words are evaluated. If all genes are included in the analysis in each pass, all words of the length in question will be divided into the negative set and positive set except the words that occur 5 or less times across all sequences. These sets can be divided in different ways (see options -h), and you can consider only most regulated genes in the two passes. If you consider the most down-regulated words in the negative pass and the most up-regulated in the positive pass, the same word will be present in both the negative and the positive set.

The output

The final output of this analysis produces a summary of the top correlating words (a list for each end of the ranked list), i.e. words over-represented in beginning of list:

Top 10 words
rank word z-score p-value fdr ledge annotation 1 cactgcc 22.54 1.00e-10 6.58e-09 1651 hsa-miR-34a-5p,hsa-miR-34c-5p,hsa-miR-449a,hsa-miR-449b-5p
2 actgcca 20.62 1.00e-10 6.98e-09 1115 hsa-miR-34a-5p,hsa-miR-34c-5p,hsa-miR-449a,hsa-miR-449b-5p,hsa-miR-548au-3p
3 cctgccc 20.54 1.00e-10 6.03e-09 2702 hsa-miR-6721-5p
4 ccctgcc 19.82 1.00e-10 6.17e-09 3252 hsa-miR-1207-5p,hsa-miR-4763-3p
5 ccctggg 18.74 1.00e-10 6.22e-09 3500 hsa-miR-1915-3p
6 ctgcccc 17.16 1.00e-10 5.73e-09 2546 hsa-miR-486-3p 7 ctgccct 16.97 1.00e-10 5.69e-09 3248 hsa-miR-4632-5p,hsa-miR-4436b-3p
8 ccagccc 16.54 1.00e-10 6.47e-09 2690 hsa-miR-762,hsa-miR-4492,hsa-miR-4498,hsa-miR-5001-5p
9 ccccagc 16.53 1.00e-10 6.32e-09 2706 hsa-miR-4731-5p
10 ctgggcc 16.42 1.00e-10 5.90e-09 3524 ...
  • 'z-score' is a correlation statistic for the given word after correction for correlations obtained from random gene list orderings.
  • 'fdr' (false discovery rate) is the estimated proportion of false discoveries for the given z-score threshold.
  • 'ledge' is the leading-edge which denotes the position in the gene list where the maximum imbalance was measured; genes before this threshold are relatively enriched for the word compared to genes after the threshold.

Invalid IDs

If you do not use -C the IDs in the rank-file will be mapped to the sequences, when it was not possible to map the rank-file ID to a sequence we report the ID as invalid and. Problems with many invalid IDs typically occur when one uses ID associated with different versions or even assemblies (like sequences hg19 and microArray probes from hg18). Invalid IDs are written to a file facilitating further investigation (file name: invalid_ids.txt).

Computer resources

Memory consumption and running time can vary a lot. The word length (-w) has a significant impact on this, if you want to run for word lengths > 8 works best with 4 GB memory or more (depending on number of genes in your analysis). Word lengths > 9 works best wih 10 GB or more.

Generally running time grows exponentially with word length and linearly with the number of bases in the sequences that need to be analysed.

5. How to cite

Rasmussen S, Jacobsen A and Krogh A;cWords - systematic microRNA regulatory motif discovery from mRNA expression data; Silence (2013)

6. LICENSE

Copyright (c) 2011, Simon H. Rasmussen. The software is open source and released under the MIT license.

About

Regulatory motif discovery tool.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

50 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

cWords is a software tool that measure correlations of short oligonucleotide sequence (word) occurrences and i.e. expression changes in a two condition experiment. In summary, it produces a statistic that quantifies how over-represented a word is in a ranked list of sequences.

1. REQUIREMENTS

All software components require Ruby (2.2 or newer, http://www.ruby-lang.org/) and JRuby (9.0.5.0 or newer, http://jruby.org/). The software has only been tested on a Unix platform (both Linux and Mac OS X Capitan).

2. INSTALL

  • Install Ruby (www.ruby-lang.org, check if it is already installed: 'ruby -v')
  • Install Java (www.java.com, check if it is already installed: 'java -version'),
  • Install JRruby (www.jruby.org, check if it is already installed: 'jruby -v'), make sure that you have the 'jruby' command in your path.

3. USAGE

The full list of options for each script is available by running the program with -h flag. Here the main options are described.

Options:

Usage: cwords [options]
-w, --wordsize ARG word length(s) you wish to search in (default 6,7)
-b, --bg ARG Order of Markov background nucleotide model (default 0)
-t, --threads ARG use multiple threads to parallelize computations (default 1)
-C, --custom_IDs Use your own sequences with matching IDs in rank and sequence files
-A, --anders_ids Use integer IDs
-x, --rank_split_mean Split ranked list at mean
-r, --rankfile ARG Rank file with IDs and one or more columns (will calculate mean across columns) of a metric of expression changes (tab or space delimiter) or just one column with ordered IDs most down-regulated to most up-regulated
-s, --seqfile ARG Sequence file - Rank and sequence IDs should be one of the compatible IDs for the species you use
--gen_plot ARG Generate plot data and plots for the top k words, takes 2 passes.
--mkplot ARG Make Word Cluster Plot - highlight (with black border) the 8mer seed site (ex for miR-1: ACATTCCA) and its corresponding 7mer and 6mer seed sites. To highlight nothing write a word not in the [acgt] alphabet.
-N, --noAnn No miRNA-annotation on the Word Cluster Plot
--annotFile ARG Supply you own annotations, for Word Cluster Plot and word ranking.
--species ARG Different ID systems are used for different species, what's the species of your data? Currently we support Human Ensembl sequences as default (write human), Mouse (write mouse), Fruit Fly (write fruitfly) and C. elegans (write roundworm)

Example runs:

cWords is built to utilize many cores in parallel on a large computer. The following are standard analyses to run with cWords.

tri-nucleotide background model, word sizes 6, 7 and 8, using 40 processors
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 2 -p 40
IDs in you sequences and rank-file match
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 2 -C
Mononucleotide background model, word cluster plot (highlighting ACATTCCA, ACATTCC, CATTCCA, ACATTC, CATTCC), Enrichment profile graphs for the 20 most significant words
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 0 -mkplot ACATTCCA -gen_plot 20

4. INTERPRETATION

Results mainly compose of three elements. A ranking of most strongly correlated words, a word cluster plot and enrichment profile plots.

Positive and negative set

The analysis is divided into 2 passes; one where words that are over-represented in up-regulated genes (ie. positively correlated words) and one where negatively correlated words are evaluated. If all genes are included in the analysis in each pass, all words of the length in question will be divided into the negative set and positive set except the words that occur 5 or less times across all sequences. These sets can be divided in different ways (see options -h), and you can consider only most regulated genes in the two passes. If you consider the most down-regulated words in the negative pass and the most up-regulated in the positive pass, the same word will be present in both the negative and the positive set.

The output

The final output of this analysis produces a summary of the top correlating words (a list for each end of the ranked list), i.e. words over-represented in beginning of list:

Top 10 words
rank word z-score p-value fdr ledge annotation 1 cactgcc 22.54 1.00e-10 6.58e-09 1651 hsa-miR-34a-5p,hsa-miR-34c-5p,hsa-miR-449a,hsa-miR-449b-5p
2 actgcca 20.62 1.00e-10 6.98e-09 1115 hsa-miR-34a-5p,hsa-miR-34c-5p,hsa-miR-449a,hsa-miR-449b-5p,hsa-miR-548au-3p
3 cctgccc 20.54 1.00e-10 6.03e-09 2702 hsa-miR-6721-5p
4 ccctgcc 19.82 1.00e-10 6.17e-09 3252 hsa-miR-1207-5p,hsa-miR-4763-3p
5 ccctggg 18.74 1.00e-10 6.22e-09 3500 hsa-miR-1915-3p
6 ctgcccc 17.16 1.00e-10 5.73e-09 2546 hsa-miR-486-3p 7 ctgccct 16.97 1.00e-10 5.69e-09 3248 hsa-miR-4632-5p,hsa-miR-4436b-3p
8 ccagccc 16.54 1.00e-10 6.47e-09 2690 hsa-miR-762,hsa-miR-4492,hsa-miR-4498,hsa-miR-5001-5p
9 ccccagc 16.53 1.00e-10 6.32e-09 2706 hsa-miR-4731-5p
10 ctgggcc 16.42 1.00e-10 5.90e-09 3524 ...
  • 'z-score' is a correlation statistic for the given word after correction for correlations obtained from random gene list orderings.
  • 'fdr' (false discovery rate) is the estimated proportion of false discoveries for the given z-score threshold.
  • 'ledge' is the leading-edge which denotes the position in the gene list where the maximum imbalance was measured; genes before this threshold are relatively enriched for the word compared to genes after the threshold.

Invalid IDs

If you do not use -C the IDs in the rank-file will be mapped to the sequences, when it was not possible to map the rank-file ID to a sequence we report the ID as invalid and. Problems with many invalid IDs typically occur when one uses ID associated with different versions or even assemblies (like sequences hg19 and microArray probes from hg18). Invalid IDs are written to a file facilitating further investigation (file name: invalid_ids.txt).

Computer resources

Memory consumption and running time can vary a lot. The word length (-w) has a significant impact on this, if you want to run for word lengths > 8 works best with 4 GB memory or more (depending on number of genes in your analysis). Word lengths > 9 works best wih 10 GB or more.

Generally running time grows exponentially with word length and linearly with the number of bases in the sequences that need to be analysed.

5. How to cite

Rasmussen S, Jacobsen A and Krogh A;cWords - systematic microRNA regulatory motif discovery from mRNA expression data; Silence (2013)

6. LICENSE

Copyright (c) 2011, Simon H. Rasmussen. The software is open source and released under the MIT license.

About

Regulatory motif discovery tool.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

50 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

cWords is a software tool that measure correlations of short oligonucleotide sequence (word) occurrences and i.e. expression changes in a two condition experiment. In summary, it produces a statistic that quantifies how over-represented a word is in a ranked list of sequences.

1. REQUIREMENTS

All software components require Ruby (2.2 or newer, http://www.ruby-lang.org/) and JRuby (9.0.5.0 or newer, http://jruby.org/). The software has only been tested on a Unix platform (both Linux and Mac OS X Capitan).

2. INSTALL

  • Install Ruby (www.ruby-lang.org, check if it is already installed: 'ruby -v')
  • Install Java (www.java.com, check if it is already installed: 'java -version'),
  • Install JRruby (www.jruby.org, check if it is already installed: 'jruby -v'), make sure that you have the 'jruby' command in your path.

3. USAGE

The full list of options for each script is available by running the program with -h flag. Here the main options are described.

Options:

Usage: cwords [options]
-w, --wordsize ARG word length(s) you wish to search in (default 6,7)
-b, --bg ARG Order of Markov background nucleotide model (default 0)
-t, --threads ARG use multiple threads to parallelize computations (default 1)
-C, --custom_IDs Use your own sequences with matching IDs in rank and sequence files
-A, --anders_ids Use integer IDs
-x, --rank_split_mean Split ranked list at mean
-r, --rankfile ARG Rank file with IDs and one or more columns (will calculate mean across columns) of a metric of expression changes (tab or space delimiter) or just one column with ordered IDs most down-regulated to most up-regulated
-s, --seqfile ARG Sequence file - Rank and sequence IDs should be one of the compatible IDs for the species you use
--gen_plot ARG Generate plot data and plots for the top k words, takes 2 passes.
--mkplot ARG Make Word Cluster Plot - highlight (with black border) the 8mer seed site (ex for miR-1: ACATTCCA) and its corresponding 7mer and 6mer seed sites. To highlight nothing write a word not in the [acgt] alphabet.
-N, --noAnn No miRNA-annotation on the Word Cluster Plot
--annotFile ARG Supply you own annotations, for Word Cluster Plot and word ranking.
--species ARG Different ID systems are used for different species, what's the species of your data? Currently we support Human Ensembl sequences as default (write human), Mouse (write mouse), Fruit Fly (write fruitfly) and C. elegans (write roundworm)

Example runs:

cWords is built to utilize many cores in parallel on a large computer. The following are standard analyses to run with cWords.

tri-nucleotide background model, word sizes 6, 7 and 8, using 40 processors
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 2 -p 40
IDs in you sequences and rank-file match
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 2 -C
Mononucleotide background model, word cluster plot (highlighting ACATTCCA, ACATTCC, CATTCCA, ACATTC, CATTCC), Enrichment profile graphs for the 20 most significant words
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 0 -mkplot ACATTCCA -gen_plot 20

4. INTERPRETATION

Results mainly compose of three elements. A ranking of most strongly correlated words, a word cluster plot and enrichment profile plots.

Positive and negative set

The analysis is divided into 2 passes; one where words that are over-represented in up-regulated genes (ie. positively correlated words) and one where negatively correlated words are evaluated. If all genes are included in the analysis in each pass, all words of the length in question will be divided into the negative set and positive set except the words that occur 5 or less times across all sequences. These sets can be divided in different ways (see options -h), and you can consider only most regulated genes in the two passes. If you consider the most down-regulated words in the negative pass and the most up-regulated in the positive pass, the same word will be present in both the negative and the positive set.

The output

The final output of this analysis produces a summary of the top correlating words (a list for each end of the ranked list), i.e. words over-represented in beginning of list:

Top 10 words
rank word z-score p-value fdr ledge annotation 1 cactgcc 22.54 1.00e-10 6.58e-09 1651 hsa-miR-34a-5p,hsa-miR-34c-5p,hsa-miR-449a,hsa-miR-449b-5p
2 actgcca 20.62 1.00e-10 6.98e-09 1115 hsa-miR-34a-5p,hsa-miR-34c-5p,hsa-miR-449a,hsa-miR-449b-5p,hsa-miR-548au-3p
3 cctgccc 20.54 1.00e-10 6.03e-09 2702 hsa-miR-6721-5p
4 ccctgcc 19.82 1.00e-10 6.17e-09 3252 hsa-miR-1207-5p,hsa-miR-4763-3p
5 ccctggg 18.74 1.00e-10 6.22e-09 3500 hsa-miR-1915-3p
6 ctgcccc 17.16 1.00e-10 5.73e-09 2546 hsa-miR-486-3p 7 ctgccct 16.97 1.00e-10 5.69e-09 3248 hsa-miR-4632-5p,hsa-miR-4436b-3p
8 ccagccc 16.54 1.00e-10 6.47e-09 2690 hsa-miR-762,hsa-miR-4492,hsa-miR-4498,hsa-miR-5001-5p
9 ccccagc 16.53 1.00e-10 6.32e-09 2706 hsa-miR-4731-5p
10 ctgggcc 16.42 1.00e-10 5.90e-09 3524 ...
  • 'z-score' is a correlation statistic for the given word after correction for correlations obtained from random gene list orderings.
  • 'fdr' (false discovery rate) is the estimated proportion of false discoveries for the given z-score threshold.
  • 'ledge' is the leading-edge which denotes the position in the gene list where the maximum imbalance was measured; genes before this threshold are relatively enriched for the word compared to genes after the threshold.

Invalid IDs

If you do not use -C the IDs in the rank-file will be mapped to the sequences, when it was not possible to map the rank-file ID to a sequence we report the ID as invalid and. Problems with many invalid IDs typically occur when one uses ID associated with different versions or even assemblies (like sequences hg19 and microArray probes from hg18). Invalid IDs are written to a file facilitating further investigation (file name: invalid_ids.txt).

Computer resources

Memory consumption and running time can vary a lot. The word length (-w) has a significant impact on this, if you want to run for word lengths > 8 works best with 4 GB memory or more (depending on number of genes in your analysis). Word lengths > 9 works best wih 10 GB or more.

Generally running time grows exponentially with word length and linearly with the number of bases in the sequences that need to be analysed.

5. How to cite

Rasmussen S, Jacobsen A and Krogh A;cWords - systematic microRNA regulatory motif discovery from mRNA expression data; Silence (2013)

6. LICENSE

Copyright (c) 2011, Simon H. Rasmussen. The software is open source and released under the MIT license.

About

Regulatory motif discovery tool.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

50 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

cWords is a software tool that measure correlations of short oligonucleotide sequence (word) occurrences and i.e. expression changes in a two condition experiment. In summary, it produces a statistic that quantifies how over-represented a word is in a ranked list of sequences.

1. REQUIREMENTS

All software components require Ruby (2.2 or newer, http://www.ruby-lang.org/) and JRuby (9.0.5.0 or newer, http://jruby.org/). The software has only been tested on a Unix platform (both Linux and Mac OS X Capitan).

2. INSTALL

  • Install Ruby (www.ruby-lang.org, check if it is already installed: 'ruby -v')
  • Install Java (www.java.com, check if it is already installed: 'java -version'),
  • Install JRruby (www.jruby.org, check if it is already installed: 'jruby -v'), make sure that you have the 'jruby' command in your path.

3. USAGE

The full list of options for each script is available by running the program with -h flag. Here the main options are described.

Options:

Usage: cwords [options]
-w, --wordsize ARG word length(s) you wish to search in (default 6,7)
-b, --bg ARG Order of Markov background nucleotide model (default 0)
-t, --threads ARG use multiple threads to parallelize computations (default 1)
-C, --custom_IDs Use your own sequences with matching IDs in rank and sequence files
-A, --anders_ids Use integer IDs
-x, --rank_split_mean Split ranked list at mean
-r, --rankfile ARG Rank file with IDs and one or more columns (will calculate mean across columns) of a metric of expression changes (tab or space delimiter) or just one column with ordered IDs most down-regulated to most up-regulated
-s, --seqfile ARG Sequence file - Rank and sequence IDs should be one of the compatible IDs for the species you use
--gen_plot ARG Generate plot data and plots for the top k words, takes 2 passes.
--mkplot ARG Make Word Cluster Plot - highlight (with black border) the 8mer seed site (ex for miR-1: ACATTCCA) and its corresponding 7mer and 6mer seed sites. To highlight nothing write a word not in the [acgt] alphabet.
-N, --noAnn No miRNA-annotation on the Word Cluster Plot
--annotFile ARG Supply you own annotations, for Word Cluster Plot and word ranking.
--species ARG Different ID systems are used for different species, what's the species of your data? Currently we support Human Ensembl sequences as default (write human), Mouse (write mouse), Fruit Fly (write fruitfly) and C. elegans (write roundworm)

Example runs:

cWords is built to utilize many cores in parallel on a large computer. The following are standard analyses to run with cWords.

tri-nucleotide background model, word sizes 6, 7 and 8, using 40 processors
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 2 -p 40
IDs in you sequences and rank-file match
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 2 -C
Mononucleotide background model, word cluster plot (highlighting ACATTCCA, ACATTCC, CATTCCA, ACATTC, CATTCC), Enrichment profile graphs for the 20 most significant words
jruby -J-Xmx4g scripts/cwords_Mult_MEM.rb -s <fasta-file> -r <rank-file> -w 6,7,8 -b 0 -mkplot ACATTCCA -gen_plot 20

4. INTERPRETATION

Results mainly compose of three elements. A ranking of most strongly correlated words, a word cluster plot and enrichment profile plots.

Positive and negative set

The analysis is divided into 2 passes; one where words that are over-represented in up-regulated genes (ie. positively correlated words) and one where negatively correlated words are evaluated. If all genes are included in the analysis in each pass, all words of the length in question will be divided into the negative set and positive set except the words that occur 5 or less times across all sequences. These sets can be divided in different ways (see options -h), and you can consider only most regulated genes in the two passes. If you consider the most down-regulated words in the negative pass and the most up-regulated in the positive pass, the same word will be present in both the negative and the positive set.

The output

The final output of this analysis produces a summary of the top correlating words (a list for each end of the ranked list), i.e. words over-represented in beginning of list:

Top 10 words
rank word z-score p-value fdr ledge annotation 1 cactgcc 22.54 1.00e-10 6.58e-09 1651 hsa-miR-34a-5p,hsa-miR-34c-5p,hsa-miR-449a,hsa-miR-449b-5p
2 actgcca 20.62 1.00e-10 6.98e-09 1115 hsa-miR-34a-5p,hsa-miR-34c-5p,hsa-miR-449a,hsa-miR-449b-5p,hsa-miR-548au-3p
3 cctgccc 20.54 1.00e-10 6.03e-09 2702 hsa-miR-6721-5p
4 ccctgcc 19.82 1.00e-10 6.17e-09 3252 hsa-miR-1207-5p,hsa-miR-4763-3p
5 ccctggg 18.74 1.00e-10 6.22e-09 3500 hsa-miR-1915-3p
6 ctgcccc 17.16 1.00e-10 5.73e-09 2546 hsa-miR-486-3p 7 ctgccct 16.97 1.00e-10 5.69e-09 3248 hsa-miR-4632-5p,hsa-miR-4436b-3p
8 ccagccc 16.54 1.00e-10 6.47e-09 2690 hsa-miR-762,hsa-miR-4492,hsa-miR-4498,hsa-miR-5001-5p
9 ccccagc 16.53 1.00e-10 6.32e-09 2706 hsa-miR-4731-5p
10 ctgggcc 16.42 1.00e-10 5.90e-09 3524 ...
  • 'z-score' is a correlation statistic for the given word after correction for correlations obtained from random gene list orderings.
  • 'fdr' (false discovery rate) is the estimated proportion of false discoveries for the given z-score threshold.
  • 'ledge' is the leading-edge which denotes the position in the gene list where the maximum imbalance was measured; genes before this threshold are relatively enriched for the word compared to genes after the threshold.

Invalid IDs

If you do not use -C the IDs in the rank-file will be mapped to the sequences, when it was not possible to map the rank-file ID to a sequence we report the ID as invalid and. Problems with many invalid IDs typically occur when one uses ID associated with different versions or even assemblies (like sequences hg19 and microArray probes from hg18). Invalid IDs are written to a file facilitating further investigation (file name: invalid_ids.txt).

Computer resources

Memory consumption and running time can vary a lot. The word length (-w) has a significant impact on this, if you want to run for word lengths > 8 works best with 4 GB memory or more (depending on number of genes in your analysis). Word lengths > 9 works best wih 10 GB or more.

Generally running time grows exponentially with word length and linearly with the number of bases in the sequences that need to be analysed.

5. How to cite

Rasmussen S, Jacobsen A and Krogh A;cWords - systematic microRNA regulatory motif discovery from mRNA expression data; Silence (2013)

6. LICENSE

Copyright (c) 2011, Simon H. Rasmussen. The software is open source and released under the MIT license.

About

Regulatory motif discovery tool.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages