Latest commit

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Tapioca

Tapioca is a pipeline for Illumina Casava 1.8 genome analyzer/hiseq data. Main features:

  • contaminant filtering
  • fastq statistical summary
  • collating/binning of casava chunks

(why tapioca? "In Brazil, the plant (cassava) is named "mandioca", while its starch is called "tapioca" https://en.wikipedia.org/wiki/Tapioca )

Depends on

Setup

In addition to the software dependencies, you'll need

  • A directory containing your Illumina sequencing instrument output.
  • Two bowtie libraries for contaminant filtering. We created one called phix and one called 'other' for adapters and primers.

Walkthrough

  • Make a new directory for the Casava & Tapioca output. Dont work in the instrument's output directory.
mkdir tap-work
cd tap-work
  • Create file samplesheet.csv. Either using Illumina's experiment manager software, or by a script to pull data from your internal LIMS. The samplesheet.csv format is described in Illumina's documentation.
  • Run Casava 1.8 to generate an Unaligned/ directory and makefile. Example:
configureBclToFastq.pl \
--input-dir /your/instrument/output/run_flowcell/Data/Intensities/BaseCalls/ \
--output-dir ./Unaligned \
--sample-sheet samplesheet.csv \
--with-failed-reads

note: It is recommended to use option --with-failed-reads, then tapioca will later separate failed chastity reads into a separate file. See Casava user's guide for other options, e.g. --use-bases-mask etc.

  • Start Casava by cd into Unaligned and running make.
cd Unaligned
make # or make -j [cores]
  • After Casava make finishes then configure Tapioca by running tap_configure_postprocessing. The last parameter is the Unaligned directory created by Casava 1.8. Like Casava, Tapioca uses Make for dependency tracking and job parallelism, so a makefile is the output of the configuration script.
cd ..
export PATH=/your/tapioca/bin:$PATH
tap_configure_postprocessing \
--contam-phix-index /your/contam_libs/tapioca_phix_contam \
--contam-phix-pct 80 \
--contam-other-index /your/contam_libs/tapioca_other_contam \
--contam-other-pct 20 \
--deployed /your/deployed/dir \
./Unaligned
  • Now the makefile was created. First run the precheck target; it does some sanity checking on the casava run and will output some warnings if it notices anything wrong off the bat.
make precheck
  • Making the 'all' target will perform the contaminant filtering and summary reporting. Technically it is not 'all' because the deploy step is a separate target.
make all
# or make -j [cores] all

make -j 16 will use 16 cores. Alternately, the qmake script could be submitted to a SGE cluster if more parallelism is required. Qmake job submission has not been tested.

  • Now check results as necessary, in the various ./Project directories. Run make deploy when ready
make deploy

The deploy target collates all the chunks of data from the casava output into the --deployed directory. You could add more targets to the makefile to perform additional processing after the deploy is finished.

Output

Look in the Deployed directory. It should be pretty self explanatory how things are organized by subdirectory. Sorry this is not better documented.

Cleanup

There is no 'make clean' target, and please be aware the intermediate Project directories created by tapioca have uncompressed fastq files in them, and so should not be left on disk long term. Delete the directories yourself.

Authors

John Crow https://github.com/crowja , Alex Rice (agr@ncgr.org)

License

# Tapioca
# Copyright (C) 2013 National Center for Genome Resources - http://ncgr.org
#
# This program is free software; you can redistribute it and/or modify
# it under the terms of the GNU General Public License as published by
# the Free Software Foundation; either version 2 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU General Public License for more details.
#
# You should have received a copy of the GNU General Public License along
# with this program; if not, write to the Free Software Foundation, Inc.,
# 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA.

About

Tapioca is a pipeline for Illumina Casava 1.8 genome analyzer/hiseq data.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Latest commit

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Tapioca

Tapioca is a pipeline for Illumina Casava 1.8 genome analyzer/hiseq data. Main features:

  • contaminant filtering
  • fastq statistical summary
  • collating/binning of casava chunks

(why tapioca? "In Brazil, the plant (cassava) is named "mandioca", while its starch is called "tapioca" https://en.wikipedia.org/wiki/Tapioca )

Depends on

Setup

In addition to the software dependencies, you'll need

  • A directory containing your Illumina sequencing instrument output.
  • Two bowtie libraries for contaminant filtering. We created one called phix and one called 'other' for adapters and primers.

Walkthrough

  • Make a new directory for the Casava & Tapioca output. Dont work in the instrument's output directory.
mkdir tap-work
cd tap-work
  • Create file samplesheet.csv. Either using Illumina's experiment manager software, or by a script to pull data from your internal LIMS. The samplesheet.csv format is described in Illumina's documentation.
  • Run Casava 1.8 to generate an Unaligned/ directory and makefile. Example:
configureBclToFastq.pl \
--input-dir /your/instrument/output/run_flowcell/Data/Intensities/BaseCalls/ \
--output-dir ./Unaligned \
--sample-sheet samplesheet.csv \
--with-failed-reads

note: It is recommended to use option --with-failed-reads, then tapioca will later separate failed chastity reads into a separate file. See Casava user's guide for other options, e.g. --use-bases-mask etc.

  • Start Casava by cd into Unaligned and running make.
cd Unaligned
make # or make -j [cores]
  • After Casava make finishes then configure Tapioca by running tap_configure_postprocessing. The last parameter is the Unaligned directory created by Casava 1.8. Like Casava, Tapioca uses Make for dependency tracking and job parallelism, so a makefile is the output of the configuration script.
cd ..
export PATH=/your/tapioca/bin:$PATH
tap_configure_postprocessing \
--contam-phix-index /your/contam_libs/tapioca_phix_contam \
--contam-phix-pct 80 \
--contam-other-index /your/contam_libs/tapioca_other_contam \
--contam-other-pct 20 \
--deployed /your/deployed/dir \
./Unaligned
  • Now the makefile was created. First run the precheck target; it does some sanity checking on the casava run and will output some warnings if it notices anything wrong off the bat.
make precheck
  • Making the 'all' target will perform the contaminant filtering and summary reporting. Technically it is not 'all' because the deploy step is a separate target.
make all
# or make -j [cores] all

make -j 16 will use 16 cores. Alternately, the qmake script could be submitted to a SGE cluster if more parallelism is required. Qmake job submission has not been tested.

  • Now check results as necessary, in the various ./Project directories. Run make deploy when ready
make deploy

The deploy target collates all the chunks of data from the casava output into the --deployed directory. You could add more targets to the makefile to perform additional processing after the deploy is finished.

Output

Look in the Deployed directory. It should be pretty self explanatory how things are organized by subdirectory. Sorry this is not better documented.

Cleanup

There is no 'make clean' target, and please be aware the intermediate Project directories created by tapioca have uncompressed fastq files in them, and so should not be left on disk long term. Delete the directories yourself.

Authors

John Crow https://github.com/crowja , Alex Rice (agr@ncgr.org)

License

# Tapioca
# Copyright (C) 2013 National Center for Genome Resources - http://ncgr.org
#
# This program is free software; you can redistribute it and/or modify
# it under the terms of the GNU General Public License as published by
# the Free Software Foundation; either version 2 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU General Public License for more details.
#
# You should have received a copy of the GNU General Public License along
# with this program; if not, write to the Free Software Foundation, Inc.,
# 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA.

About

Tapioca is a pipeline for Illumina Casava 1.8 genome analyzer/hiseq data.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Tapioca

Tapioca is a pipeline for Illumina Casava 1.8 genome analyzer/hiseq data. Main features:

  • contaminant filtering
  • fastq statistical summary
  • collating/binning of casava chunks

(why tapioca? "In Brazil, the plant (cassava) is named "mandioca", while its starch is called "tapioca" https://en.wikipedia.org/wiki/Tapioca )

Depends on

Setup

In addition to the software dependencies, you'll need

  • A directory containing your Illumina sequencing instrument output.
  • Two bowtie libraries for contaminant filtering. We created one called phix and one called 'other' for adapters and primers.

Walkthrough

  • Make a new directory for the Casava & Tapioca output. Dont work in the instrument's output directory.
mkdir tap-work
cd tap-work
  • Create file samplesheet.csv. Either using Illumina's experiment manager software, or by a script to pull data from your internal LIMS. The samplesheet.csv format is described in Illumina's documentation.
  • Run Casava 1.8 to generate an Unaligned/ directory and makefile. Example:
configureBclToFastq.pl \
--input-dir /your/instrument/output/run_flowcell/Data/Intensities/BaseCalls/ \
--output-dir ./Unaligned \
--sample-sheet samplesheet.csv \
--with-failed-reads

note: It is recommended to use option --with-failed-reads, then tapioca will later separate failed chastity reads into a separate file. See Casava user's guide for other options, e.g. --use-bases-mask etc.

  • Start Casava by cd into Unaligned and running make.
cd Unaligned
make # or make -j [cores]
  • After Casava make finishes then configure Tapioca by running tap_configure_postprocessing. The last parameter is the Unaligned directory created by Casava 1.8. Like Casava, Tapioca uses Make for dependency tracking and job parallelism, so a makefile is the output of the configuration script.
cd ..
export PATH=/your/tapioca/bin:$PATH
tap_configure_postprocessing \
--contam-phix-index /your/contam_libs/tapioca_phix_contam \
--contam-phix-pct 80 \
--contam-other-index /your/contam_libs/tapioca_other_contam \
--contam-other-pct 20 \
--deployed /your/deployed/dir \
./Unaligned
  • Now the makefile was created. First run the precheck target; it does some sanity checking on the casava run and will output some warnings if it notices anything wrong off the bat.
make precheck
  • Making the 'all' target will perform the contaminant filtering and summary reporting. Technically it is not 'all' because the deploy step is a separate target.
make all
# or make -j [cores] all

make -j 16 will use 16 cores. Alternately, the qmake script could be submitted to a SGE cluster if more parallelism is required. Qmake job submission has not been tested.

  • Now check results as necessary, in the various ./Project directories. Run make deploy when ready
make deploy

The deploy target collates all the chunks of data from the casava output into the --deployed directory. You could add more targets to the makefile to perform additional processing after the deploy is finished.

Output

Look in the Deployed directory. It should be pretty self explanatory how things are organized by subdirectory. Sorry this is not better documented.

Cleanup

There is no 'make clean' target, and please be aware the intermediate Project directories created by tapioca have uncompressed fastq files in them, and so should not be left on disk long term. Delete the directories yourself.

Authors

John Crow https://github.com/crowja , Alex Rice (agr@ncgr.org)

License

# Tapioca
# Copyright (C) 2013 National Center for Genome Resources - http://ncgr.org
#
# This program is free software; you can redistribute it and/or modify
# it under the terms of the GNU General Public License as published by
# the Free Software Foundation; either version 2 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU General Public License for more details.
#
# You should have received a copy of the GNU General Public License along
# with this program; if not, write to the Free Software Foundation, Inc.,
# 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA.

About

Tapioca is a pipeline for Illumina Casava 1.8 genome analyzer/hiseq data.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Tapioca

Tapioca is a pipeline for Illumina Casava 1.8 genome analyzer/hiseq data. Main features:

  • contaminant filtering
  • fastq statistical summary
  • collating/binning of casava chunks

(why tapioca? "In Brazil, the plant (cassava) is named "mandioca", while its starch is called "tapioca" https://en.wikipedia.org/wiki/Tapioca )

Depends on

Setup

In addition to the software dependencies, you'll need

  • A directory containing your Illumina sequencing instrument output.
  • Two bowtie libraries for contaminant filtering. We created one called phix and one called 'other' for adapters and primers.

Walkthrough

  • Make a new directory for the Casava & Tapioca output. Dont work in the instrument's output directory.
mkdir tap-work
cd tap-work
  • Create file samplesheet.csv. Either using Illumina's experiment manager software, or by a script to pull data from your internal LIMS. The samplesheet.csv format is described in Illumina's documentation.
  • Run Casava 1.8 to generate an Unaligned/ directory and makefile. Example:
configureBclToFastq.pl \
--input-dir /your/instrument/output/run_flowcell/Data/Intensities/BaseCalls/ \
--output-dir ./Unaligned \
--sample-sheet samplesheet.csv \
--with-failed-reads

note: It is recommended to use option --with-failed-reads, then tapioca will later separate failed chastity reads into a separate file. See Casava user's guide for other options, e.g. --use-bases-mask etc.

  • Start Casava by cd into Unaligned and running make.
cd Unaligned
make # or make -j [cores]
  • After Casava make finishes then configure Tapioca by running tap_configure_postprocessing. The last parameter is the Unaligned directory created by Casava 1.8. Like Casava, Tapioca uses Make for dependency tracking and job parallelism, so a makefile is the output of the configuration script.
cd ..
export PATH=/your/tapioca/bin:$PATH
tap_configure_postprocessing \
--contam-phix-index /your/contam_libs/tapioca_phix_contam \
--contam-phix-pct 80 \
--contam-other-index /your/contam_libs/tapioca_other_contam \
--contam-other-pct 20 \
--deployed /your/deployed/dir \
./Unaligned
  • Now the makefile was created. First run the precheck target; it does some sanity checking on the casava run and will output some warnings if it notices anything wrong off the bat.
make precheck
  • Making the 'all' target will perform the contaminant filtering and summary reporting. Technically it is not 'all' because the deploy step is a separate target.
make all
# or make -j [cores] all

make -j 16 will use 16 cores. Alternately, the qmake script could be submitted to a SGE cluster if more parallelism is required. Qmake job submission has not been tested.

  • Now check results as necessary, in the various ./Project directories. Run make deploy when ready
make deploy

The deploy target collates all the chunks of data from the casava output into the --deployed directory. You could add more targets to the makefile to perform additional processing after the deploy is finished.

Output

Look in the Deployed directory. It should be pretty self explanatory how things are organized by subdirectory. Sorry this is not better documented.

Cleanup

There is no 'make clean' target, and please be aware the intermediate Project directories created by tapioca have uncompressed fastq files in them, and so should not be left on disk long term. Delete the directories yourself.

Authors

John Crow https://github.com/crowja , Alex Rice (agr@ncgr.org)

License

# Tapioca
# Copyright (C) 2013 National Center for Genome Resources - http://ncgr.org
#
# This program is free software; you can redistribute it and/or modify
# it under the terms of the GNU General Public License as published by
# the Free Software Foundation; either version 2 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU General Public License for more details.
#
# You should have received a copy of the GNU General Public License along
# with this program; if not, write to the Free Software Foundation, Inc.,
# 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA.

About

Tapioca is a pipeline for Illumina Casava 1.8 genome analyzer/hiseq data.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Latest commit

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Tapioca

Tapioca is a pipeline for Illumina Casava 1.8 genome analyzer/hiseq data. Main features:

  • contaminant filtering
  • fastq statistical summary
  • collating/binning of casava chunks

(why tapioca? "In Brazil, the plant (cassava) is named "mandioca", while its starch is called "tapioca" https://en.wikipedia.org/wiki/Tapioca )

Depends on

Setup

In addition to the software dependencies, you'll need

  • A directory containing your Illumina sequencing instrument output.
  • Two bowtie libraries for contaminant filtering. We created one called phix and one called 'other' for adapters and primers.

Walkthrough

  • Make a new directory for the Casava & Tapioca output. Dont work in the instrument's output directory.
mkdir tap-work
cd tap-work
  • Create file samplesheet.csv. Either using Illumina's experiment manager software, or by a script to pull data from your internal LIMS. The samplesheet.csv format is described in Illumina's documentation.
  • Run Casava 1.8 to generate an Unaligned/ directory and makefile. Example:
configureBclToFastq.pl \
--input-dir /your/instrument/output/run_flowcell/Data/Intensities/BaseCalls/ \
--output-dir ./Unaligned \
--sample-sheet samplesheet.csv \
--with-failed-reads

note: It is recommended to use option --with-failed-reads, then tapioca will later separate failed chastity reads into a separate file. See Casava user's guide for other options, e.g. --use-bases-mask etc.

  • Start Casava by cd into Unaligned and running make.
cd Unaligned
make # or make -j [cores]
  • After Casava make finishes then configure Tapioca by running tap_configure_postprocessing. The last parameter is the Unaligned directory created by Casava 1.8. Like Casava, Tapioca uses Make for dependency tracking and job parallelism, so a makefile is the output of the configuration script.
cd ..
export PATH=/your/tapioca/bin:$PATH
tap_configure_postprocessing \
--contam-phix-index /your/contam_libs/tapioca_phix_contam \
--contam-phix-pct 80 \
--contam-other-index /your/contam_libs/tapioca_other_contam \
--contam-other-pct 20 \
--deployed /your/deployed/dir \
./Unaligned
  • Now the makefile was created. First run the precheck target; it does some sanity checking on the casava run and will output some warnings if it notices anything wrong off the bat.
make precheck
  • Making the 'all' target will perform the contaminant filtering and summary reporting. Technically it is not 'all' because the deploy step is a separate target.
make all
# or make -j [cores] all

make -j 16 will use 16 cores. Alternately, the qmake script could be submitted to a SGE cluster if more parallelism is required. Qmake job submission has not been tested.

  • Now check results as necessary, in the various ./Project directories. Run make deploy when ready
make deploy

The deploy target collates all the chunks of data from the casava output into the --deployed directory. You could add more targets to the makefile to perform additional processing after the deploy is finished.

Output

Look in the Deployed directory. It should be pretty self explanatory how things are organized by subdirectory. Sorry this is not better documented.

Cleanup

There is no 'make clean' target, and please be aware the intermediate Project directories created by tapioca have uncompressed fastq files in them, and so should not be left on disk long term. Delete the directories yourself.

Authors

John Crow https://github.com/crowja , Alex Rice (agr@ncgr.org)

License

# Tapioca
# Copyright (C) 2013 National Center for Genome Resources - http://ncgr.org
#
# This program is free software; you can redistribute it and/or modify
# it under the terms of the GNU General Public License as published by
# the Free Software Foundation; either version 2 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU General Public License for more details.
#
# You should have received a copy of the GNU General Public License along
# with this program; if not, write to the Free Software Foundation, Inc.,
# 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA.

About

Tapioca is a pipeline for Illumina Casava 1.8 genome analyzer/hiseq data.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Tapioca

Tapioca is a pipeline for Illumina Casava 1.8 genome analyzer/hiseq data. Main features:

  • contaminant filtering
  • fastq statistical summary
  • collating/binning of casava chunks

(why tapioca? "In Brazil, the plant (cassava) is named "mandioca", while its starch is called "tapioca" https://en.wikipedia.org/wiki/Tapioca )

Depends on

Setup

In addition to the software dependencies, you'll need

  • A directory containing your Illumina sequencing instrument output.
  • Two bowtie libraries for contaminant filtering. We created one called phix and one called 'other' for adapters and primers.

Walkthrough

  • Make a new directory for the Casava & Tapioca output. Dont work in the instrument's output directory.
mkdir tap-work
cd tap-work
  • Create file samplesheet.csv. Either using Illumina's experiment manager software, or by a script to pull data from your internal LIMS. The samplesheet.csv format is described in Illumina's documentation.
  • Run Casava 1.8 to generate an Unaligned/ directory and makefile. Example:
configureBclToFastq.pl \
--input-dir /your/instrument/output/run_flowcell/Data/Intensities/BaseCalls/ \
--output-dir ./Unaligned \
--sample-sheet samplesheet.csv \
--with-failed-reads

note: It is recommended to use option --with-failed-reads, then tapioca will later separate failed chastity reads into a separate file. See Casava user's guide for other options, e.g. --use-bases-mask etc.

  • Start Casava by cd into Unaligned and running make.
cd Unaligned
make # or make -j [cores]
  • After Casava make finishes then configure Tapioca by running tap_configure_postprocessing. The last parameter is the Unaligned directory created by Casava 1.8. Like Casava, Tapioca uses Make for dependency tracking and job parallelism, so a makefile is the output of the configuration script.
cd ..
export PATH=/your/tapioca/bin:$PATH
tap_configure_postprocessing \
--contam-phix-index /your/contam_libs/tapioca_phix_contam \
--contam-phix-pct 80 \
--contam-other-index /your/contam_libs/tapioca_other_contam \
--contam-other-pct 20 \
--deployed /your/deployed/dir \
./Unaligned
  • Now the makefile was created. First run the precheck target; it does some sanity checking on the casava run and will output some warnings if it notices anything wrong off the bat.
make precheck
  • Making the 'all' target will perform the contaminant filtering and summary reporting. Technically it is not 'all' because the deploy step is a separate target.
make all
# or make -j [cores] all

make -j 16 will use 16 cores. Alternately, the qmake script could be submitted to a SGE cluster if more parallelism is required. Qmake job submission has not been tested.

  • Now check results as necessary, in the various ./Project directories. Run make deploy when ready
make deploy

The deploy target collates all the chunks of data from the casava output into the --deployed directory. You could add more targets to the makefile to perform additional processing after the deploy is finished.

Output

Look in the Deployed directory. It should be pretty self explanatory how things are organized by subdirectory. Sorry this is not better documented.

Cleanup

There is no 'make clean' target, and please be aware the intermediate Project directories created by tapioca have uncompressed fastq files in them, and so should not be left on disk long term. Delete the directories yourself.

Authors

John Crow https://github.com/crowja , Alex Rice (agr@ncgr.org)

License

# Tapioca
# Copyright (C) 2013 National Center for Genome Resources - http://ncgr.org
#
# This program is free software; you can redistribute it and/or modify
# it under the terms of the GNU General Public License as published by
# the Free Software Foundation; either version 2 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU General Public License for more details.
#
# You should have received a copy of the GNU General Public License along
# with this program; if not, write to the Free Software Foundation, Inc.,
# 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA.

About

Tapioca is a pipeline for Illumina Casava 1.8 genome analyzer/hiseq data.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Latest commit

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Tapioca

Tapioca is a pipeline for Illumina Casava 1.8 genome analyzer/hiseq data. Main features:

  • contaminant filtering
  • fastq statistical summary
  • collating/binning of casava chunks

(why tapioca? "In Brazil, the plant (cassava) is named "mandioca", while its starch is called "tapioca" https://en.wikipedia.org/wiki/Tapioca )

Depends on

Setup

In addition to the software dependencies, you'll need

  • A directory containing your Illumina sequencing instrument output.
  • Two bowtie libraries for contaminant filtering. We created one called phix and one called 'other' for adapters and primers.

Walkthrough

  • Make a new directory for the Casava & Tapioca output. Dont work in the instrument's output directory.
mkdir tap-work
cd tap-work
  • Create file samplesheet.csv. Either using Illumina's experiment manager software, or by a script to pull data from your internal LIMS. The samplesheet.csv format is described in Illumina's documentation.
  • Run Casava 1.8 to generate an Unaligned/ directory and makefile. Example:
configureBclToFastq.pl \
--input-dir /your/instrument/output/run_flowcell/Data/Intensities/BaseCalls/ \
--output-dir ./Unaligned \
--sample-sheet samplesheet.csv \
--with-failed-reads

note: It is recommended to use option --with-failed-reads, then tapioca will later separate failed chastity reads into a separate file. See Casava user's guide for other options, e.g. --use-bases-mask etc.

  • Start Casava by cd into Unaligned and running make.
cd Unaligned
make # or make -j [cores]
  • After Casava make finishes then configure Tapioca by running tap_configure_postprocessing. The last parameter is the Unaligned directory created by Casava 1.8. Like Casava, Tapioca uses Make for dependency tracking and job parallelism, so a makefile is the output of the configuration script.
cd ..
export PATH=/your/tapioca/bin:$PATH
tap_configure_postprocessing \
--contam-phix-index /your/contam_libs/tapioca_phix_contam \
--contam-phix-pct 80 \
--contam-other-index /your/contam_libs/tapioca_other_contam \
--contam-other-pct 20 \
--deployed /your/deployed/dir \
./Unaligned
  • Now the makefile was created. First run the precheck target; it does some sanity checking on the casava run and will output some warnings if it notices anything wrong off the bat.
make precheck
  • Making the 'all' target will perform the contaminant filtering and summary reporting. Technically it is not 'all' because the deploy step is a separate target.
make all
# or make -j [cores] all

make -j 16 will use 16 cores. Alternately, the qmake script could be submitted to a SGE cluster if more parallelism is required. Qmake job submission has not been tested.

  • Now check results as necessary, in the various ./Project directories. Run make deploy when ready
make deploy

The deploy target collates all the chunks of data from the casava output into the --deployed directory. You could add more targets to the makefile to perform additional processing after the deploy is finished.

Output

Look in the Deployed directory. It should be pretty self explanatory how things are organized by subdirectory. Sorry this is not better documented.

Cleanup

There is no 'make clean' target, and please be aware the intermediate Project directories created by tapioca have uncompressed fastq files in them, and so should not be left on disk long term. Delete the directories yourself.

Authors

John Crow https://github.com/crowja , Alex Rice (agr@ncgr.org)

License

# Tapioca
# Copyright (C) 2013 National Center for Genome Resources - http://ncgr.org
#
# This program is free software; you can redistribute it and/or modify
# it under the terms of the GNU General Public License as published by
# the Free Software Foundation; either version 2 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU General Public License for more details.
#
# You should have received a copy of the GNU General Public License along
# with this program; if not, write to the Free Software Foundation, Inc.,
# 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA.

About

Tapioca is a pipeline for Illumina Casava 1.8 genome analyzer/hiseq data.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Latest commit

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date

Repository files navigation

Tapioca

Tapioca is a pipeline for Illumina Casava 1.8 genome analyzer/hiseq data. Main features:

  • contaminant filtering
  • fastq statistical summary
  • collating/binning of casava chunks

(why tapioca? "In Brazil, the plant (cassava) is named "mandioca", while its starch is called "tapioca" https://en.wikipedia.org/wiki/Tapioca )

Depends on

Setup

In addition to the software dependencies, you'll need

  • A directory containing your Illumina sequencing instrument output.
  • Two bowtie libraries for contaminant filtering. We created one called phix and one called 'other' for adapters and primers.

Walkthrough

  • Make a new directory for the Casava & Tapioca output. Dont work in the instrument's output directory.
mkdir tap-work
cd tap-work
  • Create file samplesheet.csv. Either using Illumina's experiment manager software, or by a script to pull data from your internal LIMS. The samplesheet.csv format is described in Illumina's documentation.
  • Run Casava 1.8 to generate an Unaligned/ directory and makefile. Example:
configureBclToFastq.pl \
--input-dir /your/instrument/output/run_flowcell/Data/Intensities/BaseCalls/ \
--output-dir ./Unaligned \
--sample-sheet samplesheet.csv \
--with-failed-reads

note: It is recommended to use option --with-failed-reads, then tapioca will later separate failed chastity reads into a separate file. See Casava user's guide for other options, e.g. --use-bases-mask etc.

  • Start Casava by cd into Unaligned and running make.
cd Unaligned
make # or make -j [cores]
  • After Casava make finishes then configure Tapioca by running tap_configure_postprocessing. The last parameter is the Unaligned directory created by Casava 1.8. Like Casava, Tapioca uses Make for dependency tracking and job parallelism, so a makefile is the output of the configuration script.
cd ..
export PATH=/your/tapioca/bin:$PATH
tap_configure_postprocessing \
--contam-phix-index /your/contam_libs/tapioca_phix_contam \
--contam-phix-pct 80 \
--contam-other-index /your/contam_libs/tapioca_other_contam \
--contam-other-pct 20 \
--deployed /your/deployed/dir \
./Unaligned
  • Now the makefile was created. First run the precheck target; it does some sanity checking on the casava run and will output some warnings if it notices anything wrong off the bat.
make precheck
  • Making the 'all' target will perform the contaminant filtering and summary reporting. Technically it is not 'all' because the deploy step is a separate target.
make all
# or make -j [cores] all

make -j 16 will use 16 cores. Alternately, the qmake script could be submitted to a SGE cluster if more parallelism is required. Qmake job submission has not been tested.

  • Now check results as necessary, in the various ./Project directories. Run make deploy when ready
make deploy

The deploy target collates all the chunks of data from the casava output into the --deployed directory. You could add more targets to the makefile to perform additional processing after the deploy is finished.

Output

Look in the Deployed directory. It should be pretty self explanatory how things are organized by subdirectory. Sorry this is not better documented.

Cleanup

There is no 'make clean' target, and please be aware the intermediate Project directories created by tapioca have uncompressed fastq files in them, and so should not be left on disk long term. Delete the directories yourself.

Authors

John Crow https://github.com/crowja , Alex Rice (agr@ncgr.org)

License

# Tapioca
# Copyright (C) 2013 National Center for Genome Resources - http://ncgr.org
#
# This program is free software; you can redistribute it and/or modify
# it under the terms of the GNU General Public License as published by
# the Free Software Foundation; either version 2 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU General Public License for more details.
#
# You should have received a copy of the GNU General Public License along
# with this program; if not, write to the Free Software Foundation, Inc.,
# 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA.

About

Tapioca is a pipeline for Illumina Casava 1.8 genome analyzer/hiseq data.

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors