Repository files navigation

How to get cloneHD and filterHD?

The current stable release, as well as pre-compiled executable binaries for Mac OS X and GNU Linux (64bit), can be found here. Watch/Star this repo to receive updates.

Build Status

Compilation

For Mac OSX and GNU Linux (64bit), pre-compiled binaries are available here. To compile cloneHD yourself, you need the GNU scientific library (GSL) v1.15 or later. Change the paths in the Makefile to point to your local GSL installation (if non-standard). Then type

$ make

in the src directory. The executables will be in build. For debugging with gdb, use make -f Makefile.debug.

If you use Mac OSX, cloneHD is also available with Homebrew. This will install cloneHD plus all dependencies:

$ brew tap homebrew/science
$ brew install clonehd

The executables will be automatically added to your path.

Run a test with simulated data

After downloading cloneHD from the release site, you can test both filterHD and cloneHD by running

$ sh run-example.sh

where you can see a typical workflow of analysing read depth and BAF data with a matched normal. All command line arguments are explained below.

Report bugs

To report bugs, use the issue interface of GitHub.

Full documentation

The full documentation can be found in the /docs/ subfolder. Click below.

What are cloneHD and filterHD for?

cloneHD is a software for reconstructing the subclonal structure of a population from short-read sequencing data. Read depth data, B-allele count data and somatic nucleotide variant (SNV) data can be used for the inference. cloneHD can estimate the number of subclonal populations, their fractions in the sample, their individual total copy number profiles, their B-allele status and all the SNV genotypes with high resolution.

filterHD is a general purpose probabilistic filtering algorithm for one-dimensional discrete data, similar in spirit to a Kalman filter. It is a continuous state space Hidden Markov model with Poisson or Binomial emissions and a jump-diffusion propagator. It can be used for scale-free smoothing, fuzzy data segmentation and data filtering.

cna gofbaf gofcna postcna realbaf postbaf realsnv gof

Visualization of the cloneHD output for the simulated data set. From top to bottom: (i) The bias corrected read depth data and the cloneHD prediction (red). (ii) The BAF (B-allele frequency), reflected at 0.5 and the cloneHD prediction (red). (iii) The total copy number posterior. (iv) The real total copy number profile. (v) The minor copy number posterior. (vi) The real minor copy number profile. (vii) The observed SNV frequencies, corrected for local ploidy, and per genotype (SNVs are assigned ramdomly according to the cloneHD SNV posterior). (All plots are created with Wolfram Mathematica.)

Tips and tricks

  • The read depth input files for filterHD and cloneHD can be generated from a bam-file with samtools. Given a bed file of non-overlapping 1kb windows, e.g.

     human-genome.1kb-grid.bed :
    1	9000	10000
    1	10000	11000
    ...
    

    $ samtools bedcov human-genome.1kb-grid.bed sample.bam > read-depth.sample.txt

     read-depth.sample.txt :
    1	9000	10000	12009
    1	10000	11000	213557
    ...
    

    $ awk '{print $1,$3,int(0.5+$4/1000.0),1}' read-depth.sample.txt > read-depth.sample.cloneHD.txt

     read-depth.sample.cloneHD.txt :
    1 10000 12 1
    1 11000 214 1
    ...
    

    For 10 kb windows, you would use $ awk '{print $1,$3,int(0.5+$4/10000.0),10} etc. For windows smaller than 1kb, the observations might not be approximately independent.

  • All input files are assumed to be sorted by chromosome and genomic coordinate. With Unix, this can be achieved with sort -k1n,1 -k2n,2 file.txt > sorted-file.txt.

  • Pre-filtering of data is very important. If filterHD predicts many more jumps than you would expect, it might be necessary to pre-filter the data, removing variable regions, outliers or very short segments (use programs pre-filter and filterHD).

  • Make sure that the bias field for the tumor CNA data is meaningful. If a matched normal sample was sequenced with the same pipeline, its read depth profile, as predicted by filterHD, can be used as a bias field for the tumor CNA data. Follow the logic of the example data given here.

  • If the matched-normal sample was sequenced at lower coverage than the tumor sample, it might be necessary to run filterHD with a higher-than-optimal diffusion constant (set with --sigma [double]) to obtain a more faithful bias field. Otherwise, the filterHD solution is too stiff and you loose bias detail.

  • filterHD can sometimes run into local optima. In this case, it might be useful to set initial values for the parameters via --jumpi [double] etc.

  • By default, cloneHD runs with mass-gauging enabled. This seems wasteful, but is actually quite useful because you can see some alternative explanations during the course of the analysis.

  • Don't put too much weight on the BIC criterion. It was calibrated using simulated data. For real data, it should be supplemented with common sense and biological knowledge. Use --force [int] to use a fixed number of subclones and --max-tcn [int] to set the maximum possible total copy number.

  • If high copy numbers are expected only in a few chromosomes, you can increase performance by using the --max-tcn [file] option to specify per-chromosome upper limits.

  • For exome sequencing data, the read depth bias can be enormous. The filterHD estimate of the bias field might not be very useful, especially in segmenting the tumor data. Use rather, if available, the jumps seen in the BAF data for both CNA and BAF data (give the BAF jumps file to both --cna-jumps and --baf-jumps).

Who is using cloneHD?

Google Scholar lists some of the references where cloneHD has been used by us and other researchers. We’d like to highlight:

  • T McKerrell et al. Development and validation of a comprehensive genomic diagnostic tool for myeloid malignancies. Blood 128:e1-e9 (2016) DOI: 10.1182/blood-2015-11-683334.

  • I Vázquez-García et al. Clonal heterogeneity influences the fate of new adaptive mutations. Cell Reports 21 (3), 732-744 (2017) DOI: 10.1016/j.celrep.2017.09.046.

How to cite

The cloneHD and filterHD software is free under the GNU General Public License v3. If you use this software in your work, please cite the accompanying publication:

Andrej Fischer, Ignacio Vázquez-García, Christopher J.R. Illingworth and Ville Mustonen. High-definition reconstruction of clonal composition in cancer. Cell Reports 7 (5), 1740-1752 (2014). DOI: 10.1016/j.celrep.2014.04.055.

About

High-definition reconstruction of clonal composition from high-throughput sequencing data

Resources

Stars

1 star

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

How to get cloneHD and filterHD?

The current stable release, as well as pre-compiled executable binaries for Mac OS X and GNU Linux (64bit), can be found here. Watch/Star this repo to receive updates.

Build Status

Compilation

For Mac OSX and GNU Linux (64bit), pre-compiled binaries are available here. To compile cloneHD yourself, you need the GNU scientific library (GSL) v1.15 or later. Change the paths in the Makefile to point to your local GSL installation (if non-standard). Then type

$ make

in the src directory. The executables will be in build. For debugging with gdb, use make -f Makefile.debug.

If you use Mac OSX, cloneHD is also available with Homebrew. This will install cloneHD plus all dependencies:

$ brew tap homebrew/science
$ brew install clonehd

The executables will be automatically added to your path.

Run a test with simulated data

After downloading cloneHD from the release site, you can test both filterHD and cloneHD by running

$ sh run-example.sh

where you can see a typical workflow of analysing read depth and BAF data with a matched normal. All command line arguments are explained below.

Report bugs

To report bugs, use the issue interface of GitHub.

Full documentation

The full documentation can be found in the /docs/ subfolder. Click below.

What are cloneHD and filterHD for?

cloneHD is a software for reconstructing the subclonal structure of a population from short-read sequencing data. Read depth data, B-allele count data and somatic nucleotide variant (SNV) data can be used for the inference. cloneHD can estimate the number of subclonal populations, their fractions in the sample, their individual total copy number profiles, their B-allele status and all the SNV genotypes with high resolution.

filterHD is a general purpose probabilistic filtering algorithm for one-dimensional discrete data, similar in spirit to a Kalman filter. It is a continuous state space Hidden Markov model with Poisson or Binomial emissions and a jump-diffusion propagator. It can be used for scale-free smoothing, fuzzy data segmentation and data filtering.

cna gofbaf gofcna postcna realbaf postbaf realsnv gof

Visualization of the cloneHD output for the simulated data set. From top to bottom: (i) The bias corrected read depth data and the cloneHD prediction (red). (ii) The BAF (B-allele frequency), reflected at 0.5 and the cloneHD prediction (red). (iii) The total copy number posterior. (iv) The real total copy number profile. (v) The minor copy number posterior. (vi) The real minor copy number profile. (vii) The observed SNV frequencies, corrected for local ploidy, and per genotype (SNVs are assigned ramdomly according to the cloneHD SNV posterior). (All plots are created with Wolfram Mathematica.)

Tips and tricks

  • The read depth input files for filterHD and cloneHD can be generated from a bam-file with samtools. Given a bed file of non-overlapping 1kb windows, e.g.

     human-genome.1kb-grid.bed :
    1	9000	10000
    1	10000	11000
    ...
    

    $ samtools bedcov human-genome.1kb-grid.bed sample.bam > read-depth.sample.txt

     read-depth.sample.txt :
    1	9000	10000	12009
    1	10000	11000	213557
    ...
    

    $ awk '{print $1,$3,int(0.5+$4/1000.0),1}' read-depth.sample.txt > read-depth.sample.cloneHD.txt

     read-depth.sample.cloneHD.txt :
    1 10000 12 1
    1 11000 214 1
    ...
    

    For 10 kb windows, you would use $ awk '{print $1,$3,int(0.5+$4/10000.0),10} etc. For windows smaller than 1kb, the observations might not be approximately independent.

  • All input files are assumed to be sorted by chromosome and genomic coordinate. With Unix, this can be achieved with sort -k1n,1 -k2n,2 file.txt > sorted-file.txt.

  • Pre-filtering of data is very important. If filterHD predicts many more jumps than you would expect, it might be necessary to pre-filter the data, removing variable regions, outliers or very short segments (use programs pre-filter and filterHD).

  • Make sure that the bias field for the tumor CNA data is meaningful. If a matched normal sample was sequenced with the same pipeline, its read depth profile, as predicted by filterHD, can be used as a bias field for the tumor CNA data. Follow the logic of the example data given here.

  • If the matched-normal sample was sequenced at lower coverage than the tumor sample, it might be necessary to run filterHD with a higher-than-optimal diffusion constant (set with --sigma [double]) to obtain a more faithful bias field. Otherwise, the filterHD solution is too stiff and you loose bias detail.

  • filterHD can sometimes run into local optima. In this case, it might be useful to set initial values for the parameters via --jumpi [double] etc.

  • By default, cloneHD runs with mass-gauging enabled. This seems wasteful, but is actually quite useful because you can see some alternative explanations during the course of the analysis.

  • Don't put too much weight on the BIC criterion. It was calibrated using simulated data. For real data, it should be supplemented with common sense and biological knowledge. Use --force [int] to use a fixed number of subclones and --max-tcn [int] to set the maximum possible total copy number.

  • If high copy numbers are expected only in a few chromosomes, you can increase performance by using the --max-tcn [file] option to specify per-chromosome upper limits.

  • For exome sequencing data, the read depth bias can be enormous. The filterHD estimate of the bias field might not be very useful, especially in segmenting the tumor data. Use rather, if available, the jumps seen in the BAF data for both CNA and BAF data (give the BAF jumps file to both --cna-jumps and --baf-jumps).

Who is using cloneHD?

Google Scholar lists some of the references where cloneHD has been used by us and other researchers. We’d like to highlight:

  • T McKerrell et al. Development and validation of a comprehensive genomic diagnostic tool for myeloid malignancies. Blood 128:e1-e9 (2016) DOI: 10.1182/blood-2015-11-683334.

  • I Vázquez-García et al. Clonal heterogeneity influences the fate of new adaptive mutations. Cell Reports 21 (3), 732-744 (2017) DOI: 10.1016/j.celrep.2017.09.046.

How to cite

The cloneHD and filterHD software is free under the GNU General Public License v3. If you use this software in your work, please cite the accompanying publication:

Andrej Fischer, Ignacio Vázquez-García, Christopher J.R. Illingworth and Ville Mustonen. High-definition reconstruction of clonal composition in cancer. Cell Reports 7 (5), 1740-1752 (2014). DOI: 10.1016/j.celrep.2014.04.055.

About

High-definition reconstruction of clonal composition from high-throughput sequencing data

Resources

Stars

1 star

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

How to get cloneHD and filterHD?

The current stable release, as well as pre-compiled executable binaries for Mac OS X and GNU Linux (64bit), can be found here. Watch/Star this repo to receive updates.

Build Status

Compilation

For Mac OSX and GNU Linux (64bit), pre-compiled binaries are available here. To compile cloneHD yourself, you need the GNU scientific library (GSL) v1.15 or later. Change the paths in the Makefile to point to your local GSL installation (if non-standard). Then type

$ make

in the src directory. The executables will be in build. For debugging with gdb, use make -f Makefile.debug.

If you use Mac OSX, cloneHD is also available with Homebrew. This will install cloneHD plus all dependencies:

$ brew tap homebrew/science
$ brew install clonehd

The executables will be automatically added to your path.

Run a test with simulated data

After downloading cloneHD from the release site, you can test both filterHD and cloneHD by running

$ sh run-example.sh

where you can see a typical workflow of analysing read depth and BAF data with a matched normal. All command line arguments are explained below.

Report bugs

To report bugs, use the issue interface of GitHub.

Full documentation

The full documentation can be found in the /docs/ subfolder. Click below.

What are cloneHD and filterHD for?

cloneHD is a software for reconstructing the subclonal structure of a population from short-read sequencing data. Read depth data, B-allele count data and somatic nucleotide variant (SNV) data can be used for the inference. cloneHD can estimate the number of subclonal populations, their fractions in the sample, their individual total copy number profiles, their B-allele status and all the SNV genotypes with high resolution.

filterHD is a general purpose probabilistic filtering algorithm for one-dimensional discrete data, similar in spirit to a Kalman filter. It is a continuous state space Hidden Markov model with Poisson or Binomial emissions and a jump-diffusion propagator. It can be used for scale-free smoothing, fuzzy data segmentation and data filtering.

cna gofbaf gofcna postcna realbaf postbaf realsnv gof

Visualization of the cloneHD output for the simulated data set. From top to bottom: (i) The bias corrected read depth data and the cloneHD prediction (red). (ii) The BAF (B-allele frequency), reflected at 0.5 and the cloneHD prediction (red). (iii) The total copy number posterior. (iv) The real total copy number profile. (v) The minor copy number posterior. (vi) The real minor copy number profile. (vii) The observed SNV frequencies, corrected for local ploidy, and per genotype (SNVs are assigned ramdomly according to the cloneHD SNV posterior). (All plots are created with Wolfram Mathematica.)

Tips and tricks

  • The read depth input files for filterHD and cloneHD can be generated from a bam-file with samtools. Given a bed file of non-overlapping 1kb windows, e.g.

     human-genome.1kb-grid.bed :
    1	9000	10000
    1	10000	11000
    ...
    

    $ samtools bedcov human-genome.1kb-grid.bed sample.bam > read-depth.sample.txt

     read-depth.sample.txt :
    1	9000	10000	12009
    1	10000	11000	213557
    ...
    

    $ awk '{print $1,$3,int(0.5+$4/1000.0),1}' read-depth.sample.txt > read-depth.sample.cloneHD.txt

     read-depth.sample.cloneHD.txt :
    1 10000 12 1
    1 11000 214 1
    ...
    

    For 10 kb windows, you would use $ awk '{print $1,$3,int(0.5+$4/10000.0),10} etc. For windows smaller than 1kb, the observations might not be approximately independent.

  • All input files are assumed to be sorted by chromosome and genomic coordinate. With Unix, this can be achieved with sort -k1n,1 -k2n,2 file.txt > sorted-file.txt.

  • Pre-filtering of data is very important. If filterHD predicts many more jumps than you would expect, it might be necessary to pre-filter the data, removing variable regions, outliers or very short segments (use programs pre-filter and filterHD).

  • Make sure that the bias field for the tumor CNA data is meaningful. If a matched normal sample was sequenced with the same pipeline, its read depth profile, as predicted by filterHD, can be used as a bias field for the tumor CNA data. Follow the logic of the example data given here.

  • If the matched-normal sample was sequenced at lower coverage than the tumor sample, it might be necessary to run filterHD with a higher-than-optimal diffusion constant (set with --sigma [double]) to obtain a more faithful bias field. Otherwise, the filterHD solution is too stiff and you loose bias detail.

  • filterHD can sometimes run into local optima. In this case, it might be useful to set initial values for the parameters via --jumpi [double] etc.

  • By default, cloneHD runs with mass-gauging enabled. This seems wasteful, but is actually quite useful because you can see some alternative explanations during the course of the analysis.

  • Don't put too much weight on the BIC criterion. It was calibrated using simulated data. For real data, it should be supplemented with common sense and biological knowledge. Use --force [int] to use a fixed number of subclones and --max-tcn [int] to set the maximum possible total copy number.

  • If high copy numbers are expected only in a few chromosomes, you can increase performance by using the --max-tcn [file] option to specify per-chromosome upper limits.

  • For exome sequencing data, the read depth bias can be enormous. The filterHD estimate of the bias field might not be very useful, especially in segmenting the tumor data. Use rather, if available, the jumps seen in the BAF data for both CNA and BAF data (give the BAF jumps file to both --cna-jumps and --baf-jumps).

Who is using cloneHD?

Google Scholar lists some of the references where cloneHD has been used by us and other researchers. We’d like to highlight:

  • T McKerrell et al. Development and validation of a comprehensive genomic diagnostic tool for myeloid malignancies. Blood 128:e1-e9 (2016) DOI: 10.1182/blood-2015-11-683334.

  • I Vázquez-García et al. Clonal heterogeneity influences the fate of new adaptive mutations. Cell Reports 21 (3), 732-744 (2017) DOI: 10.1016/j.celrep.2017.09.046.

How to cite

The cloneHD and filterHD software is free under the GNU General Public License v3. If you use this software in your work, please cite the accompanying publication:

Andrej Fischer, Ignacio Vázquez-García, Christopher J.R. Illingworth and Ville Mustonen. High-definition reconstruction of clonal composition in cancer. Cell Reports 7 (5), 1740-1752 (2014). DOI: 10.1016/j.celrep.2014.04.055.

About

High-definition reconstruction of clonal composition from high-throughput sequencing data

Resources

Stars

1 star

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

How to get cloneHD and filterHD?

The current stable release, as well as pre-compiled executable binaries for Mac OS X and GNU Linux (64bit), can be found here. Watch/Star this repo to receive updates.

Build Status

Compilation

For Mac OSX and GNU Linux (64bit), pre-compiled binaries are available here. To compile cloneHD yourself, you need the GNU scientific library (GSL) v1.15 or later. Change the paths in the Makefile to point to your local GSL installation (if non-standard). Then type

$ make

in the src directory. The executables will be in build. For debugging with gdb, use make -f Makefile.debug.

If you use Mac OSX, cloneHD is also available with Homebrew. This will install cloneHD plus all dependencies:

$ brew tap homebrew/science
$ brew install clonehd

The executables will be automatically added to your path.

Run a test with simulated data

After downloading cloneHD from the release site, you can test both filterHD and cloneHD by running

$ sh run-example.sh

where you can see a typical workflow of analysing read depth and BAF data with a matched normal. All command line arguments are explained below.

Report bugs

To report bugs, use the issue interface of GitHub.

Full documentation

The full documentation can be found in the /docs/ subfolder. Click below.

What are cloneHD and filterHD for?

cloneHD is a software for reconstructing the subclonal structure of a population from short-read sequencing data. Read depth data, B-allele count data and somatic nucleotide variant (SNV) data can be used for the inference. cloneHD can estimate the number of subclonal populations, their fractions in the sample, their individual total copy number profiles, their B-allele status and all the SNV genotypes with high resolution.

filterHD is a general purpose probabilistic filtering algorithm for one-dimensional discrete data, similar in spirit to a Kalman filter. It is a continuous state space Hidden Markov model with Poisson or Binomial emissions and a jump-diffusion propagator. It can be used for scale-free smoothing, fuzzy data segmentation and data filtering.

cna gofbaf gofcna postcna realbaf postbaf realsnv gof

Visualization of the cloneHD output for the simulated data set. From top to bottom: (i) The bias corrected read depth data and the cloneHD prediction (red). (ii) The BAF (B-allele frequency), reflected at 0.5 and the cloneHD prediction (red). (iii) The total copy number posterior. (iv) The real total copy number profile. (v) The minor copy number posterior. (vi) The real minor copy number profile. (vii) The observed SNV frequencies, corrected for local ploidy, and per genotype (SNVs are assigned ramdomly according to the cloneHD SNV posterior). (All plots are created with Wolfram Mathematica.)

Tips and tricks

  • The read depth input files for filterHD and cloneHD can be generated from a bam-file with samtools. Given a bed file of non-overlapping 1kb windows, e.g.

     human-genome.1kb-grid.bed :
    1	9000	10000
    1	10000	11000
    ...
    

    $ samtools bedcov human-genome.1kb-grid.bed sample.bam > read-depth.sample.txt

     read-depth.sample.txt :
    1	9000	10000	12009
    1	10000	11000	213557
    ...
    

    $ awk '{print $1,$3,int(0.5+$4/1000.0),1}' read-depth.sample.txt > read-depth.sample.cloneHD.txt

     read-depth.sample.cloneHD.txt :
    1 10000 12 1
    1 11000 214 1
    ...
    

    For 10 kb windows, you would use $ awk '{print $1,$3,int(0.5+$4/10000.0),10} etc. For windows smaller than 1kb, the observations might not be approximately independent.

  • All input files are assumed to be sorted by chromosome and genomic coordinate. With Unix, this can be achieved with sort -k1n,1 -k2n,2 file.txt > sorted-file.txt.

  • Pre-filtering of data is very important. If filterHD predicts many more jumps than you would expect, it might be necessary to pre-filter the data, removing variable regions, outliers or very short segments (use programs pre-filter and filterHD).

  • Make sure that the bias field for the tumor CNA data is meaningful. If a matched normal sample was sequenced with the same pipeline, its read depth profile, as predicted by filterHD, can be used as a bias field for the tumor CNA data. Follow the logic of the example data given here.

  • If the matched-normal sample was sequenced at lower coverage than the tumor sample, it might be necessary to run filterHD with a higher-than-optimal diffusion constant (set with --sigma [double]) to obtain a more faithful bias field. Otherwise, the filterHD solution is too stiff and you loose bias detail.

  • filterHD can sometimes run into local optima. In this case, it might be useful to set initial values for the parameters via --jumpi [double] etc.

  • By default, cloneHD runs with mass-gauging enabled. This seems wasteful, but is actually quite useful because you can see some alternative explanations during the course of the analysis.

  • Don't put too much weight on the BIC criterion. It was calibrated using simulated data. For real data, it should be supplemented with common sense and biological knowledge. Use --force [int] to use a fixed number of subclones and --max-tcn [int] to set the maximum possible total copy number.

  • If high copy numbers are expected only in a few chromosomes, you can increase performance by using the --max-tcn [file] option to specify per-chromosome upper limits.

  • For exome sequencing data, the read depth bias can be enormous. The filterHD estimate of the bias field might not be very useful, especially in segmenting the tumor data. Use rather, if available, the jumps seen in the BAF data for both CNA and BAF data (give the BAF jumps file to both --cna-jumps and --baf-jumps).

Who is using cloneHD?

Google Scholar lists some of the references where cloneHD has been used by us and other researchers. We’d like to highlight:

  • T McKerrell et al. Development and validation of a comprehensive genomic diagnostic tool for myeloid malignancies. Blood 128:e1-e9 (2016) DOI: 10.1182/blood-2015-11-683334.

  • I Vázquez-García et al. Clonal heterogeneity influences the fate of new adaptive mutations. Cell Reports 21 (3), 732-744 (2017) DOI: 10.1016/j.celrep.2017.09.046.

How to cite

The cloneHD and filterHD software is free under the GNU General Public License v3. If you use this software in your work, please cite the accompanying publication:

Andrej Fischer, Ignacio Vázquez-García, Christopher J.R. Illingworth and Ville Mustonen. High-definition reconstruction of clonal composition in cancer. Cell Reports 7 (5), 1740-1752 (2014). DOI: 10.1016/j.celrep.2014.04.055.

About

High-definition reconstruction of clonal composition from high-throughput sequencing data

Resources

Stars

1 star

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

How to get cloneHD and filterHD?

The current stable release, as well as pre-compiled executable binaries for Mac OS X and GNU Linux (64bit), can be found here. Watch/Star this repo to receive updates.

Build Status

Compilation

For Mac OSX and GNU Linux (64bit), pre-compiled binaries are available here. To compile cloneHD yourself, you need the GNU scientific library (GSL) v1.15 or later. Change the paths in the Makefile to point to your local GSL installation (if non-standard). Then type

$ make

in the src directory. The executables will be in build. For debugging with gdb, use make -f Makefile.debug.

If you use Mac OSX, cloneHD is also available with Homebrew. This will install cloneHD plus all dependencies:

$ brew tap homebrew/science
$ brew install clonehd

The executables will be automatically added to your path.

Run a test with simulated data

After downloading cloneHD from the release site, you can test both filterHD and cloneHD by running

$ sh run-example.sh

where you can see a typical workflow of analysing read depth and BAF data with a matched normal. All command line arguments are explained below.

Report bugs

To report bugs, use the issue interface of GitHub.

Full documentation

The full documentation can be found in the /docs/ subfolder. Click below.

What are cloneHD and filterHD for?

cloneHD is a software for reconstructing the subclonal structure of a population from short-read sequencing data. Read depth data, B-allele count data and somatic nucleotide variant (SNV) data can be used for the inference. cloneHD can estimate the number of subclonal populations, their fractions in the sample, their individual total copy number profiles, their B-allele status and all the SNV genotypes with high resolution.

filterHD is a general purpose probabilistic filtering algorithm for one-dimensional discrete data, similar in spirit to a Kalman filter. It is a continuous state space Hidden Markov model with Poisson or Binomial emissions and a jump-diffusion propagator. It can be used for scale-free smoothing, fuzzy data segmentation and data filtering.

cna gofbaf gofcna postcna realbaf postbaf realsnv gof

Visualization of the cloneHD output for the simulated data set. From top to bottom: (i) The bias corrected read depth data and the cloneHD prediction (red). (ii) The BAF (B-allele frequency), reflected at 0.5 and the cloneHD prediction (red). (iii) The total copy number posterior. (iv) The real total copy number profile. (v) The minor copy number posterior. (vi) The real minor copy number profile. (vii) The observed SNV frequencies, corrected for local ploidy, and per genotype (SNVs are assigned ramdomly according to the cloneHD SNV posterior). (All plots are created with Wolfram Mathematica.)

Tips and tricks

  • The read depth input files for filterHD and cloneHD can be generated from a bam-file with samtools. Given a bed file of non-overlapping 1kb windows, e.g.

     human-genome.1kb-grid.bed :
    1	9000	10000
    1	10000	11000
    ...
    

    $ samtools bedcov human-genome.1kb-grid.bed sample.bam > read-depth.sample.txt

     read-depth.sample.txt :
    1	9000	10000	12009
    1	10000	11000	213557
    ...
    

    $ awk '{print $1,$3,int(0.5+$4/1000.0),1}' read-depth.sample.txt > read-depth.sample.cloneHD.txt

     read-depth.sample.cloneHD.txt :
    1 10000 12 1
    1 11000 214 1
    ...
    

    For 10 kb windows, you would use $ awk '{print $1,$3,int(0.5+$4/10000.0),10} etc. For windows smaller than 1kb, the observations might not be approximately independent.

  • All input files are assumed to be sorted by chromosome and genomic coordinate. With Unix, this can be achieved with sort -k1n,1 -k2n,2 file.txt > sorted-file.txt.

  • Pre-filtering of data is very important. If filterHD predicts many more jumps than you would expect, it might be necessary to pre-filter the data, removing variable regions, outliers or very short segments (use programs pre-filter and filterHD).

  • Make sure that the bias field for the tumor CNA data is meaningful. If a matched normal sample was sequenced with the same pipeline, its read depth profile, as predicted by filterHD, can be used as a bias field for the tumor CNA data. Follow the logic of the example data given here.

  • If the matched-normal sample was sequenced at lower coverage than the tumor sample, it might be necessary to run filterHD with a higher-than-optimal diffusion constant (set with --sigma [double]) to obtain a more faithful bias field. Otherwise, the filterHD solution is too stiff and you loose bias detail.

  • filterHD can sometimes run into local optima. In this case, it might be useful to set initial values for the parameters via --jumpi [double] etc.

  • By default, cloneHD runs with mass-gauging enabled. This seems wasteful, but is actually quite useful because you can see some alternative explanations during the course of the analysis.

  • Don't put too much weight on the BIC criterion. It was calibrated using simulated data. For real data, it should be supplemented with common sense and biological knowledge. Use --force [int] to use a fixed number of subclones and --max-tcn [int] to set the maximum possible total copy number.

  • If high copy numbers are expected only in a few chromosomes, you can increase performance by using the --max-tcn [file] option to specify per-chromosome upper limits.

  • For exome sequencing data, the read depth bias can be enormous. The filterHD estimate of the bias field might not be very useful, especially in segmenting the tumor data. Use rather, if available, the jumps seen in the BAF data for both CNA and BAF data (give the BAF jumps file to both --cna-jumps and --baf-jumps).

Who is using cloneHD?

Google Scholar lists some of the references where cloneHD has been used by us and other researchers. We’d like to highlight:

  • T McKerrell et al. Development and validation of a comprehensive genomic diagnostic tool for myeloid malignancies. Blood 128:e1-e9 (2016) DOI: 10.1182/blood-2015-11-683334.

  • I Vázquez-García et al. Clonal heterogeneity influences the fate of new adaptive mutations. Cell Reports 21 (3), 732-744 (2017) DOI: 10.1016/j.celrep.2017.09.046.

How to cite

The cloneHD and filterHD software is free under the GNU General Public License v3. If you use this software in your work, please cite the accompanying publication:

Andrej Fischer, Ignacio Vázquez-García, Christopher J.R. Illingworth and Ville Mustonen. High-definition reconstruction of clonal composition in cancer. Cell Reports 7 (5), 1740-1752 (2014). DOI: 10.1016/j.celrep.2014.04.055.

About

High-definition reconstruction of clonal composition from high-throughput sequencing data

Resources

Stars

1 star

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

How to get cloneHD and filterHD?

The current stable release, as well as pre-compiled executable binaries for Mac OS X and GNU Linux (64bit), can be found here. Watch/Star this repo to receive updates.

Build Status

Compilation

For Mac OSX and GNU Linux (64bit), pre-compiled binaries are available here. To compile cloneHD yourself, you need the GNU scientific library (GSL) v1.15 or later. Change the paths in the Makefile to point to your local GSL installation (if non-standard). Then type

$ make

in the src directory. The executables will be in build. For debugging with gdb, use make -f Makefile.debug.

If you use Mac OSX, cloneHD is also available with Homebrew. This will install cloneHD plus all dependencies:

$ brew tap homebrew/science
$ brew install clonehd

The executables will be automatically added to your path.

Run a test with simulated data

After downloading cloneHD from the release site, you can test both filterHD and cloneHD by running

$ sh run-example.sh

where you can see a typical workflow of analysing read depth and BAF data with a matched normal. All command line arguments are explained below.

Report bugs

To report bugs, use the issue interface of GitHub.

Full documentation

The full documentation can be found in the /docs/ subfolder. Click below.

What are cloneHD and filterHD for?

cloneHD is a software for reconstructing the subclonal structure of a population from short-read sequencing data. Read depth data, B-allele count data and somatic nucleotide variant (SNV) data can be used for the inference. cloneHD can estimate the number of subclonal populations, their fractions in the sample, their individual total copy number profiles, their B-allele status and all the SNV genotypes with high resolution.

filterHD is a general purpose probabilistic filtering algorithm for one-dimensional discrete data, similar in spirit to a Kalman filter. It is a continuous state space Hidden Markov model with Poisson or Binomial emissions and a jump-diffusion propagator. It can be used for scale-free smoothing, fuzzy data segmentation and data filtering.

cna gofbaf gofcna postcna realbaf postbaf realsnv gof

Visualization of the cloneHD output for the simulated data set. From top to bottom: (i) The bias corrected read depth data and the cloneHD prediction (red). (ii) The BAF (B-allele frequency), reflected at 0.5 and the cloneHD prediction (red). (iii) The total copy number posterior. (iv) The real total copy number profile. (v) The minor copy number posterior. (vi) The real minor copy number profile. (vii) The observed SNV frequencies, corrected for local ploidy, and per genotype (SNVs are assigned ramdomly according to the cloneHD SNV posterior). (All plots are created with Wolfram Mathematica.)

Tips and tricks

  • The read depth input files for filterHD and cloneHD can be generated from a bam-file with samtools. Given a bed file of non-overlapping 1kb windows, e.g.

     human-genome.1kb-grid.bed :
    1	9000	10000
    1	10000	11000
    ...
    

    $ samtools bedcov human-genome.1kb-grid.bed sample.bam > read-depth.sample.txt

     read-depth.sample.txt :
    1	9000	10000	12009
    1	10000	11000	213557
    ...
    

    $ awk '{print $1,$3,int(0.5+$4/1000.0),1}' read-depth.sample.txt > read-depth.sample.cloneHD.txt

     read-depth.sample.cloneHD.txt :
    1 10000 12 1
    1 11000 214 1
    ...
    

    For 10 kb windows, you would use $ awk '{print $1,$3,int(0.5+$4/10000.0),10} etc. For windows smaller than 1kb, the observations might not be approximately independent.

  • All input files are assumed to be sorted by chromosome and genomic coordinate. With Unix, this can be achieved with sort -k1n,1 -k2n,2 file.txt > sorted-file.txt.

  • Pre-filtering of data is very important. If filterHD predicts many more jumps than you would expect, it might be necessary to pre-filter the data, removing variable regions, outliers or very short segments (use programs pre-filter and filterHD).

  • Make sure that the bias field for the tumor CNA data is meaningful. If a matched normal sample was sequenced with the same pipeline, its read depth profile, as predicted by filterHD, can be used as a bias field for the tumor CNA data. Follow the logic of the example data given here.

  • If the matched-normal sample was sequenced at lower coverage than the tumor sample, it might be necessary to run filterHD with a higher-than-optimal diffusion constant (set with --sigma [double]) to obtain a more faithful bias field. Otherwise, the filterHD solution is too stiff and you loose bias detail.

  • filterHD can sometimes run into local optima. In this case, it might be useful to set initial values for the parameters via --jumpi [double] etc.

  • By default, cloneHD runs with mass-gauging enabled. This seems wasteful, but is actually quite useful because you can see some alternative explanations during the course of the analysis.

  • Don't put too much weight on the BIC criterion. It was calibrated using simulated data. For real data, it should be supplemented with common sense and biological knowledge. Use --force [int] to use a fixed number of subclones and --max-tcn [int] to set the maximum possible total copy number.

  • If high copy numbers are expected only in a few chromosomes, you can increase performance by using the --max-tcn [file] option to specify per-chromosome upper limits.

  • For exome sequencing data, the read depth bias can be enormous. The filterHD estimate of the bias field might not be very useful, especially in segmenting the tumor data. Use rather, if available, the jumps seen in the BAF data for both CNA and BAF data (give the BAF jumps file to both --cna-jumps and --baf-jumps).

Who is using cloneHD?

Google Scholar lists some of the references where cloneHD has been used by us and other researchers. We’d like to highlight:

  • T McKerrell et al. Development and validation of a comprehensive genomic diagnostic tool for myeloid malignancies. Blood 128:e1-e9 (2016) DOI: 10.1182/blood-2015-11-683334.

  • I Vázquez-García et al. Clonal heterogeneity influences the fate of new adaptive mutations. Cell Reports 21 (3), 732-744 (2017) DOI: 10.1016/j.celrep.2017.09.046.

How to cite

The cloneHD and filterHD software is free under the GNU General Public License v3. If you use this software in your work, please cite the accompanying publication:

Andrej Fischer, Ignacio Vázquez-García, Christopher J.R. Illingworth and Ville Mustonen. High-definition reconstruction of clonal composition in cancer. Cell Reports 7 (5), 1740-1752 (2014). DOI: 10.1016/j.celrep.2014.04.055.

About

High-definition reconstruction of clonal composition from high-throughput sequencing data

Resources

Stars

1 star

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

How to get cloneHD and filterHD?

The current stable release, as well as pre-compiled executable binaries for Mac OS X and GNU Linux (64bit), can be found here. Watch/Star this repo to receive updates.

Build Status

Compilation

For Mac OSX and GNU Linux (64bit), pre-compiled binaries are available here. To compile cloneHD yourself, you need the GNU scientific library (GSL) v1.15 or later. Change the paths in the Makefile to point to your local GSL installation (if non-standard). Then type

$ make

in the src directory. The executables will be in build. For debugging with gdb, use make -f Makefile.debug.

If you use Mac OSX, cloneHD is also available with Homebrew. This will install cloneHD plus all dependencies:

$ brew tap homebrew/science
$ brew install clonehd

The executables will be automatically added to your path.

Run a test with simulated data

After downloading cloneHD from the release site, you can test both filterHD and cloneHD by running

$ sh run-example.sh

where you can see a typical workflow of analysing read depth and BAF data with a matched normal. All command line arguments are explained below.

Report bugs

To report bugs, use the issue interface of GitHub.

Full documentation

The full documentation can be found in the /docs/ subfolder. Click below.

What are cloneHD and filterHD for?

cloneHD is a software for reconstructing the subclonal structure of a population from short-read sequencing data. Read depth data, B-allele count data and somatic nucleotide variant (SNV) data can be used for the inference. cloneHD can estimate the number of subclonal populations, their fractions in the sample, their individual total copy number profiles, their B-allele status and all the SNV genotypes with high resolution.

filterHD is a general purpose probabilistic filtering algorithm for one-dimensional discrete data, similar in spirit to a Kalman filter. It is a continuous state space Hidden Markov model with Poisson or Binomial emissions and a jump-diffusion propagator. It can be used for scale-free smoothing, fuzzy data segmentation and data filtering.

cna gofbaf gofcna postcna realbaf postbaf realsnv gof

Visualization of the cloneHD output for the simulated data set. From top to bottom: (i) The bias corrected read depth data and the cloneHD prediction (red). (ii) The BAF (B-allele frequency), reflected at 0.5 and the cloneHD prediction (red). (iii) The total copy number posterior. (iv) The real total copy number profile. (v) The minor copy number posterior. (vi) The real minor copy number profile. (vii) The observed SNV frequencies, corrected for local ploidy, and per genotype (SNVs are assigned ramdomly according to the cloneHD SNV posterior). (All plots are created with Wolfram Mathematica.)

Tips and tricks

  • The read depth input files for filterHD and cloneHD can be generated from a bam-file with samtools. Given a bed file of non-overlapping 1kb windows, e.g.

     human-genome.1kb-grid.bed :
    1	9000	10000
    1	10000	11000
    ...
    

    $ samtools bedcov human-genome.1kb-grid.bed sample.bam > read-depth.sample.txt

     read-depth.sample.txt :
    1	9000	10000	12009
    1	10000	11000	213557
    ...
    

    $ awk '{print $1,$3,int(0.5+$4/1000.0),1}' read-depth.sample.txt > read-depth.sample.cloneHD.txt

     read-depth.sample.cloneHD.txt :
    1 10000 12 1
    1 11000 214 1
    ...
    

    For 10 kb windows, you would use $ awk '{print $1,$3,int(0.5+$4/10000.0),10} etc. For windows smaller than 1kb, the observations might not be approximately independent.

  • All input files are assumed to be sorted by chromosome and genomic coordinate. With Unix, this can be achieved with sort -k1n,1 -k2n,2 file.txt > sorted-file.txt.

  • Pre-filtering of data is very important. If filterHD predicts many more jumps than you would expect, it might be necessary to pre-filter the data, removing variable regions, outliers or very short segments (use programs pre-filter and filterHD).

  • Make sure that the bias field for the tumor CNA data is meaningful. If a matched normal sample was sequenced with the same pipeline, its read depth profile, as predicted by filterHD, can be used as a bias field for the tumor CNA data. Follow the logic of the example data given here.

  • If the matched-normal sample was sequenced at lower coverage than the tumor sample, it might be necessary to run filterHD with a higher-than-optimal diffusion constant (set with --sigma [double]) to obtain a more faithful bias field. Otherwise, the filterHD solution is too stiff and you loose bias detail.

  • filterHD can sometimes run into local optima. In this case, it might be useful to set initial values for the parameters via --jumpi [double] etc.

  • By default, cloneHD runs with mass-gauging enabled. This seems wasteful, but is actually quite useful because you can see some alternative explanations during the course of the analysis.

  • Don't put too much weight on the BIC criterion. It was calibrated using simulated data. For real data, it should be supplemented with common sense and biological knowledge. Use --force [int] to use a fixed number of subclones and --max-tcn [int] to set the maximum possible total copy number.

  • If high copy numbers are expected only in a few chromosomes, you can increase performance by using the --max-tcn [file] option to specify per-chromosome upper limits.

  • For exome sequencing data, the read depth bias can be enormous. The filterHD estimate of the bias field might not be very useful, especially in segmenting the tumor data. Use rather, if available, the jumps seen in the BAF data for both CNA and BAF data (give the BAF jumps file to both --cna-jumps and --baf-jumps).

Who is using cloneHD?

Google Scholar lists some of the references where cloneHD has been used by us and other researchers. We’d like to highlight:

  • T McKerrell et al. Development and validation of a comprehensive genomic diagnostic tool for myeloid malignancies. Blood 128:e1-e9 (2016) DOI: 10.1182/blood-2015-11-683334.

  • I Vázquez-García et al. Clonal heterogeneity influences the fate of new adaptive mutations. Cell Reports 21 (3), 732-744 (2017) DOI: 10.1016/j.celrep.2017.09.046.

How to cite

The cloneHD and filterHD software is free under the GNU General Public License v3. If you use this software in your work, please cite the accompanying publication:

Andrej Fischer, Ignacio Vázquez-García, Christopher J.R. Illingworth and Ville Mustonen. High-definition reconstruction of clonal composition in cancer. Cell Reports 7 (5), 1740-1752 (2014). DOI: 10.1016/j.celrep.2014.04.055.

About

High-definition reconstruction of clonal composition from high-throughput sequencing data

Resources

Stars

1 star

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

How to get cloneHD and filterHD?

The current stable release, as well as pre-compiled executable binaries for Mac OS X and GNU Linux (64bit), can be found here. Watch/Star this repo to receive updates.

Build Status

Compilation

For Mac OSX and GNU Linux (64bit), pre-compiled binaries are available here. To compile cloneHD yourself, you need the GNU scientific library (GSL) v1.15 or later. Change the paths in the Makefile to point to your local GSL installation (if non-standard). Then type

$ make

in the src directory. The executables will be in build. For debugging with gdb, use make -f Makefile.debug.

If you use Mac OSX, cloneHD is also available with Homebrew. This will install cloneHD plus all dependencies:

$ brew tap homebrew/science
$ brew install clonehd

The executables will be automatically added to your path.

Run a test with simulated data

After downloading cloneHD from the release site, you can test both filterHD and cloneHD by running

$ sh run-example.sh

where you can see a typical workflow of analysing read depth and BAF data with a matched normal. All command line arguments are explained below.

Report bugs

To report bugs, use the issue interface of GitHub.

Full documentation

The full documentation can be found in the /docs/ subfolder. Click below.

What are cloneHD and filterHD for?

cloneHD is a software for reconstructing the subclonal structure of a population from short-read sequencing data. Read depth data, B-allele count data and somatic nucleotide variant (SNV) data can be used for the inference. cloneHD can estimate the number of subclonal populations, their fractions in the sample, their individual total copy number profiles, their B-allele status and all the SNV genotypes with high resolution.

filterHD is a general purpose probabilistic filtering algorithm for one-dimensional discrete data, similar in spirit to a Kalman filter. It is a continuous state space Hidden Markov model with Poisson or Binomial emissions and a jump-diffusion propagator. It can be used for scale-free smoothing, fuzzy data segmentation and data filtering.

cna gofbaf gofcna postcna realbaf postbaf realsnv gof

Visualization of the cloneHD output for the simulated data set. From top to bottom: (i) The bias corrected read depth data and the cloneHD prediction (red). (ii) The BAF (B-allele frequency), reflected at 0.5 and the cloneHD prediction (red). (iii) The total copy number posterior. (iv) The real total copy number profile. (v) The minor copy number posterior. (vi) The real minor copy number profile. (vii) The observed SNV frequencies, corrected for local ploidy, and per genotype (SNVs are assigned ramdomly according to the cloneHD SNV posterior). (All plots are created with Wolfram Mathematica.)

Tips and tricks

  • The read depth input files for filterHD and cloneHD can be generated from a bam-file with samtools. Given a bed file of non-overlapping 1kb windows, e.g.

     human-genome.1kb-grid.bed :
    1	9000	10000
    1	10000	11000
    ...
    

    $ samtools bedcov human-genome.1kb-grid.bed sample.bam > read-depth.sample.txt

     read-depth.sample.txt :
    1	9000	10000	12009
    1	10000	11000	213557
    ...
    

    $ awk '{print $1,$3,int(0.5+$4/1000.0),1}' read-depth.sample.txt > read-depth.sample.cloneHD.txt

     read-depth.sample.cloneHD.txt :
    1 10000 12 1
    1 11000 214 1
    ...
    

    For 10 kb windows, you would use $ awk '{print $1,$3,int(0.5+$4/10000.0),10} etc. For windows smaller than 1kb, the observations might not be approximately independent.

  • All input files are assumed to be sorted by chromosome and genomic coordinate. With Unix, this can be achieved with sort -k1n,1 -k2n,2 file.txt > sorted-file.txt.

  • Pre-filtering of data is very important. If filterHD predicts many more jumps than you would expect, it might be necessary to pre-filter the data, removing variable regions, outliers or very short segments (use programs pre-filter and filterHD).

  • Make sure that the bias field for the tumor CNA data is meaningful. If a matched normal sample was sequenced with the same pipeline, its read depth profile, as predicted by filterHD, can be used as a bias field for the tumor CNA data. Follow the logic of the example data given here.

  • If the matched-normal sample was sequenced at lower coverage than the tumor sample, it might be necessary to run filterHD with a higher-than-optimal diffusion constant (set with --sigma [double]) to obtain a more faithful bias field. Otherwise, the filterHD solution is too stiff and you loose bias detail.

  • filterHD can sometimes run into local optima. In this case, it might be useful to set initial values for the parameters via --jumpi [double] etc.

  • By default, cloneHD runs with mass-gauging enabled. This seems wasteful, but is actually quite useful because you can see some alternative explanations during the course of the analysis.

  • Don't put too much weight on the BIC criterion. It was calibrated using simulated data. For real data, it should be supplemented with common sense and biological knowledge. Use --force [int] to use a fixed number of subclones and --max-tcn [int] to set the maximum possible total copy number.

  • If high copy numbers are expected only in a few chromosomes, you can increase performance by using the --max-tcn [file] option to specify per-chromosome upper limits.

  • For exome sequencing data, the read depth bias can be enormous. The filterHD estimate of the bias field might not be very useful, especially in segmenting the tumor data. Use rather, if available, the jumps seen in the BAF data for both CNA and BAF data (give the BAF jumps file to both --cna-jumps and --baf-jumps).

Who is using cloneHD?

Google Scholar lists some of the references where cloneHD has been used by us and other researchers. We’d like to highlight:

  • T McKerrell et al. Development and validation of a comprehensive genomic diagnostic tool for myeloid malignancies. Blood 128:e1-e9 (2016) DOI: 10.1182/blood-2015-11-683334.

  • I Vázquez-García et al. Clonal heterogeneity influences the fate of new adaptive mutations. Cell Reports 21 (3), 732-744 (2017) DOI: 10.1016/j.celrep.2017.09.046.

How to cite

The cloneHD and filterHD software is free under the GNU General Public License v3. If you use this software in your work, please cite the accompanying publication:

Andrej Fischer, Ignacio Vázquez-García, Christopher J.R. Illingworth and Ville Mustonen. High-definition reconstruction of clonal composition in cancer. Cell Reports 7 (5), 1740-1752 (2014). DOI: 10.1016/j.celrep.2014.04.055.

About

High-definition reconstruction of clonal composition from high-throughput sequencing data

Resources

Stars

1 star

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages