Repository files navigation

CLIP_mapping

This is the pipeline I use to quickly map CLIP data. Modified from rnaroids pipeline (Neel Mukherjee) with tips from Philipp Boss (author of omniCLIP).

This pipeline is desinged to work with SUN Grid engine. (Or can be run locally on a machine with at least 80G memory).

This pipeline maps single end CLIP data to hg19. If the sequencing is paired-end (as in eCLIP), it will only take read1.

This pipeline uses snakemake (Köster 2012, https://academic.oup.com/bioinformatics/article/28/19/2520/290322). Snakefile is the main pipeline file. Config file config_clippipe.yaml lists variables that Snakefile needs (genome, annotation, adapter sequences, number of threads). Config file cluster_config.json regulates memory and other requirements when submitting jobs to the cluster. The bash script snakecharmer.sh is used to submit snakemake job to the cluster.

Preparation

Get the genome and annotation from Gencode and unzip it
wget ftp://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/GRCh37.p13.genome.fa.gz
wget ftp://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/gencode.v19.annotation.gtf.gz
Install miniconda and install the environment

Download the respective file from miniconda depending on your system (https://docs.conda.io/en/latest/miniconda.html). Follow installation instructions.

For example, for Linux, do:

wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
bash Miniconda3-latest-Linux-x86_64.sh
eval "$($HOME/miniconda3/bin/conda shell.bash hook)"
## create environment for CLIP mapping
conda create --name CLIP_mapping --file CLIP_mapping_conda_env.txt
conda activate CLIP_mapping
Get CLIP data

I use a small fastq fetching script for files from NCBI SRA, especially handy if there are many CLIP files from different RBPs. Make a sample sheet that looks like:

RBP1	SRR100{1..3}
RBP2	SRR100{4..5}

Run the script:

bash fetchFastq.sh example_sample_sheet

It should create directories "RBP1", "RBP2" etc. and put fastq files inside. In the example case, there should be 1 directory "ELAVL1" with 3 fastq files.

Note: I use --split-files by default to be able to always use the same script whether it is single end, paired end or has barcode read. Snakemake will always take read 1, and search for input files which look like *_1.fastq.gz. To use own fastq files, please add _1 to the file base name.

Run the pipeline

To check that everything OK, try snakemake dry run:

snakemake -nrp

It should give you wildcards for sample names, all rules that will be executed and output files. Then run the pipeline:

  • on a cluster qsub snakecharmer.sh or bash snakecharmer.sh
  • without cluster bash snakecharmer.sh or snakemake

If you are on the cluster, you might change "account" in cluster_config.json to your email address: you will get an email when the job is done, or aborted.

If you are not on the cluster, keep in mind that it may need quite some memory. Segemehl index for hg19 requires at least 64G and alignment usually takes ~75G.

Changing parameters

For small RNAs, mapping parameters could be a source of debate :-) Depending on library quality, one might want to make them more strict or more relaxed. This pipeline is trying to be very universal and find a compromise between many variants of CLIP data. Several things can be adjusted:

  • cutadapt -m 18 : this removes all reads shorter than 18nt, because it is progressively more difficult to map shorter reads. If one still might want to try and include those, change this number to 16, for example.
  • I try to include all possible small RNA adapters used. Of course, if you know exact adapter sequences for the sample (not always easy to find out in GEO!) you might want to add them to cutadapt rule, if they are not listed yet in the config file.
  • Deduplication: since I assume an old CLIP without UMIs, I do simple read collapsing. If one has UMIs, it is better to use UMI-tools(https://github.com/CGATOxford/UMI-tools) to deduplicate reads.
  • Mapping: I use segemehl (Hoffmann 2009, PMID: 19750212) because it especially well captures deletions which can be used as diagnostic events (Kassuhn 2016, PMID:26776207). segemehl parameters -M -D and -S
    • -S regulates split alignments. It can be removed if you presume the RBP doesn't bind pre-mRNAs since it makes bam files much larger in size. -S option generated three additional segemehl output files (which for now cannot be redirected to output directory): trns.txt, sngl.bed, mult.bed
    • -D 2 : allows mismatches in seeds, especially helpful with 4SU and 6SG based CLIP if we expect >1 T>C or G>A conversions next to each other
    • -M 1 : controls the number of multiple hits. I only take unique hits, but up to -M 3 may give you additional usable reads
Calling peaks (not included in the pipeline)

This pipeline is optimized to be followed by the omniCLIP peak caller (Drewe-Boss 2018, https://github.com/philippdre/omniCLIP) which needs at least 2 replicates of CLIP and an input/background sample to call peaks against. I prefer to use RNA-seq data in the same cell line as background (1 replicate is enough).

About

a pipeline to map CLIP-seq data

Topics

Resources

Stars

2 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

CLIP_mapping

This is the pipeline I use to quickly map CLIP data. Modified from rnaroids pipeline (Neel Mukherjee) with tips from Philipp Boss (author of omniCLIP).

This pipeline is desinged to work with SUN Grid engine. (Or can be run locally on a machine with at least 80G memory).

This pipeline maps single end CLIP data to hg19. If the sequencing is paired-end (as in eCLIP), it will only take read1.

This pipeline uses snakemake (Köster 2012, https://academic.oup.com/bioinformatics/article/28/19/2520/290322). Snakefile is the main pipeline file. Config file config_clippipe.yaml lists variables that Snakefile needs (genome, annotation, adapter sequences, number of threads). Config file cluster_config.json regulates memory and other requirements when submitting jobs to the cluster. The bash script snakecharmer.sh is used to submit snakemake job to the cluster.

Preparation

Get the genome and annotation from Gencode and unzip it
wget ftp://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/GRCh37.p13.genome.fa.gz
wget ftp://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/gencode.v19.annotation.gtf.gz
Install miniconda and install the environment

Download the respective file from miniconda depending on your system (https://docs.conda.io/en/latest/miniconda.html). Follow installation instructions.

For example, for Linux, do:

wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
bash Miniconda3-latest-Linux-x86_64.sh
eval "$($HOME/miniconda3/bin/conda shell.bash hook)"
## create environment for CLIP mapping
conda create --name CLIP_mapping --file CLIP_mapping_conda_env.txt
conda activate CLIP_mapping
Get CLIP data

I use a small fastq fetching script for files from NCBI SRA, especially handy if there are many CLIP files from different RBPs. Make a sample sheet that looks like:

RBP1	SRR100{1..3}
RBP2	SRR100{4..5}

Run the script:

bash fetchFastq.sh example_sample_sheet

It should create directories "RBP1", "RBP2" etc. and put fastq files inside. In the example case, there should be 1 directory "ELAVL1" with 3 fastq files.

Note: I use --split-files by default to be able to always use the same script whether it is single end, paired end or has barcode read. Snakemake will always take read 1, and search for input files which look like *_1.fastq.gz. To use own fastq files, please add _1 to the file base name.

Run the pipeline

To check that everything OK, try snakemake dry run:

snakemake -nrp

It should give you wildcards for sample names, all rules that will be executed and output files. Then run the pipeline:

  • on a cluster qsub snakecharmer.sh or bash snakecharmer.sh
  • without cluster bash snakecharmer.sh or snakemake

If you are on the cluster, you might change "account" in cluster_config.json to your email address: you will get an email when the job is done, or aborted.

If you are not on the cluster, keep in mind that it may need quite some memory. Segemehl index for hg19 requires at least 64G and alignment usually takes ~75G.

Changing parameters

For small RNAs, mapping parameters could be a source of debate :-) Depending on library quality, one might want to make them more strict or more relaxed. This pipeline is trying to be very universal and find a compromise between many variants of CLIP data. Several things can be adjusted:

  • cutadapt -m 18 : this removes all reads shorter than 18nt, because it is progressively more difficult to map shorter reads. If one still might want to try and include those, change this number to 16, for example.
  • I try to include all possible small RNA adapters used. Of course, if you know exact adapter sequences for the sample (not always easy to find out in GEO!) you might want to add them to cutadapt rule, if they are not listed yet in the config file.
  • Deduplication: since I assume an old CLIP without UMIs, I do simple read collapsing. If one has UMIs, it is better to use UMI-tools(https://github.com/CGATOxford/UMI-tools) to deduplicate reads.
  • Mapping: I use segemehl (Hoffmann 2009, PMID: 19750212) because it especially well captures deletions which can be used as diagnostic events (Kassuhn 2016, PMID:26776207). segemehl parameters -M -D and -S
    • -S regulates split alignments. It can be removed if you presume the RBP doesn't bind pre-mRNAs since it makes bam files much larger in size. -S option generated three additional segemehl output files (which for now cannot be redirected to output directory): trns.txt, sngl.bed, mult.bed
    • -D 2 : allows mismatches in seeds, especially helpful with 4SU and 6SG based CLIP if we expect >1 T>C or G>A conversions next to each other
    • -M 1 : controls the number of multiple hits. I only take unique hits, but up to -M 3 may give you additional usable reads
Calling peaks (not included in the pipeline)

This pipeline is optimized to be followed by the omniCLIP peak caller (Drewe-Boss 2018, https://github.com/philippdre/omniCLIP) which needs at least 2 replicates of CLIP and an input/background sample to call peaks against. I prefer to use RNA-seq data in the same cell line as background (1 replicate is enough).

About

a pipeline to map CLIP-seq data

Topics

Resources

Stars

2 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CLIP_mapping

This is the pipeline I use to quickly map CLIP data. Modified from rnaroids pipeline (Neel Mukherjee) with tips from Philipp Boss (author of omniCLIP).

This pipeline is desinged to work with SUN Grid engine. (Or can be run locally on a machine with at least 80G memory).

This pipeline maps single end CLIP data to hg19. If the sequencing is paired-end (as in eCLIP), it will only take read1.

This pipeline uses snakemake (Köster 2012, https://academic.oup.com/bioinformatics/article/28/19/2520/290322). Snakefile is the main pipeline file. Config file config_clippipe.yaml lists variables that Snakefile needs (genome, annotation, adapter sequences, number of threads). Config file cluster_config.json regulates memory and other requirements when submitting jobs to the cluster. The bash script snakecharmer.sh is used to submit snakemake job to the cluster.

Preparation

Get the genome and annotation from Gencode and unzip it
wget ftp://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/GRCh37.p13.genome.fa.gz
wget ftp://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/gencode.v19.annotation.gtf.gz
Install miniconda and install the environment

Download the respective file from miniconda depending on your system (https://docs.conda.io/en/latest/miniconda.html). Follow installation instructions.

For example, for Linux, do:

wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
bash Miniconda3-latest-Linux-x86_64.sh
eval "$($HOME/miniconda3/bin/conda shell.bash hook)"
## create environment for CLIP mapping
conda create --name CLIP_mapping --file CLIP_mapping_conda_env.txt
conda activate CLIP_mapping
Get CLIP data

I use a small fastq fetching script for files from NCBI SRA, especially handy if there are many CLIP files from different RBPs. Make a sample sheet that looks like:

RBP1	SRR100{1..3}
RBP2	SRR100{4..5}

Run the script:

bash fetchFastq.sh example_sample_sheet

It should create directories "RBP1", "RBP2" etc. and put fastq files inside. In the example case, there should be 1 directory "ELAVL1" with 3 fastq files.

Note: I use --split-files by default to be able to always use the same script whether it is single end, paired end or has barcode read. Snakemake will always take read 1, and search for input files which look like *_1.fastq.gz. To use own fastq files, please add _1 to the file base name.

Run the pipeline

To check that everything OK, try snakemake dry run:

snakemake -nrp

It should give you wildcards for sample names, all rules that will be executed and output files. Then run the pipeline:

  • on a cluster qsub snakecharmer.sh or bash snakecharmer.sh
  • without cluster bash snakecharmer.sh or snakemake

If you are on the cluster, you might change "account" in cluster_config.json to your email address: you will get an email when the job is done, or aborted.

If you are not on the cluster, keep in mind that it may need quite some memory. Segemehl index for hg19 requires at least 64G and alignment usually takes ~75G.

Changing parameters

For small RNAs, mapping parameters could be a source of debate :-) Depending on library quality, one might want to make them more strict or more relaxed. This pipeline is trying to be very universal and find a compromise between many variants of CLIP data. Several things can be adjusted:

  • cutadapt -m 18 : this removes all reads shorter than 18nt, because it is progressively more difficult to map shorter reads. If one still might want to try and include those, change this number to 16, for example.
  • I try to include all possible small RNA adapters used. Of course, if you know exact adapter sequences for the sample (not always easy to find out in GEO!) you might want to add them to cutadapt rule, if they are not listed yet in the config file.
  • Deduplication: since I assume an old CLIP without UMIs, I do simple read collapsing. If one has UMIs, it is better to use UMI-tools(https://github.com/CGATOxford/UMI-tools) to deduplicate reads.
  • Mapping: I use segemehl (Hoffmann 2009, PMID: 19750212) because it especially well captures deletions which can be used as diagnostic events (Kassuhn 2016, PMID:26776207). segemehl parameters -M -D and -S
    • -S regulates split alignments. It can be removed if you presume the RBP doesn't bind pre-mRNAs since it makes bam files much larger in size. -S option generated three additional segemehl output files (which for now cannot be redirected to output directory): trns.txt, sngl.bed, mult.bed
    • -D 2 : allows mismatches in seeds, especially helpful with 4SU and 6SG based CLIP if we expect >1 T>C or G>A conversions next to each other
    • -M 1 : controls the number of multiple hits. I only take unique hits, but up to -M 3 may give you additional usable reads
Calling peaks (not included in the pipeline)

This pipeline is optimized to be followed by the omniCLIP peak caller (Drewe-Boss 2018, https://github.com/philippdre/omniCLIP) which needs at least 2 replicates of CLIP and an input/background sample to call peaks against. I prefer to use RNA-seq data in the same cell line as background (1 replicate is enough).

About

a pipeline to map CLIP-seq data

Topics

Resources

Stars

2 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CLIP_mapping

This is the pipeline I use to quickly map CLIP data. Modified from rnaroids pipeline (Neel Mukherjee) with tips from Philipp Boss (author of omniCLIP).

This pipeline is desinged to work with SUN Grid engine. (Or can be run locally on a machine with at least 80G memory).

This pipeline maps single end CLIP data to hg19. If the sequencing is paired-end (as in eCLIP), it will only take read1.

This pipeline uses snakemake (Köster 2012, https://academic.oup.com/bioinformatics/article/28/19/2520/290322). Snakefile is the main pipeline file. Config file config_clippipe.yaml lists variables that Snakefile needs (genome, annotation, adapter sequences, number of threads). Config file cluster_config.json regulates memory and other requirements when submitting jobs to the cluster. The bash script snakecharmer.sh is used to submit snakemake job to the cluster.

Preparation

Get the genome and annotation from Gencode and unzip it
wget ftp://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/GRCh37.p13.genome.fa.gz
wget ftp://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/gencode.v19.annotation.gtf.gz
Install miniconda and install the environment

Download the respective file from miniconda depending on your system (https://docs.conda.io/en/latest/miniconda.html). Follow installation instructions.

For example, for Linux, do:

wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
bash Miniconda3-latest-Linux-x86_64.sh
eval "$($HOME/miniconda3/bin/conda shell.bash hook)"
## create environment for CLIP mapping
conda create --name CLIP_mapping --file CLIP_mapping_conda_env.txt
conda activate CLIP_mapping
Get CLIP data

I use a small fastq fetching script for files from NCBI SRA, especially handy if there are many CLIP files from different RBPs. Make a sample sheet that looks like:

RBP1	SRR100{1..3}
RBP2	SRR100{4..5}

Run the script:

bash fetchFastq.sh example_sample_sheet

It should create directories "RBP1", "RBP2" etc. and put fastq files inside. In the example case, there should be 1 directory "ELAVL1" with 3 fastq files.

Note: I use --split-files by default to be able to always use the same script whether it is single end, paired end or has barcode read. Snakemake will always take read 1, and search for input files which look like *_1.fastq.gz. To use own fastq files, please add _1 to the file base name.

Run the pipeline

To check that everything OK, try snakemake dry run:

snakemake -nrp

It should give you wildcards for sample names, all rules that will be executed and output files. Then run the pipeline:

  • on a cluster qsub snakecharmer.sh or bash snakecharmer.sh
  • without cluster bash snakecharmer.sh or snakemake

If you are on the cluster, you might change "account" in cluster_config.json to your email address: you will get an email when the job is done, or aborted.

If you are not on the cluster, keep in mind that it may need quite some memory. Segemehl index for hg19 requires at least 64G and alignment usually takes ~75G.

Changing parameters

For small RNAs, mapping parameters could be a source of debate :-) Depending on library quality, one might want to make them more strict or more relaxed. This pipeline is trying to be very universal and find a compromise between many variants of CLIP data. Several things can be adjusted:

  • cutadapt -m 18 : this removes all reads shorter than 18nt, because it is progressively more difficult to map shorter reads. If one still might want to try and include those, change this number to 16, for example.
  • I try to include all possible small RNA adapters used. Of course, if you know exact adapter sequences for the sample (not always easy to find out in GEO!) you might want to add them to cutadapt rule, if they are not listed yet in the config file.
  • Deduplication: since I assume an old CLIP without UMIs, I do simple read collapsing. If one has UMIs, it is better to use UMI-tools(https://github.com/CGATOxford/UMI-tools) to deduplicate reads.
  • Mapping: I use segemehl (Hoffmann 2009, PMID: 19750212) because it especially well captures deletions which can be used as diagnostic events (Kassuhn 2016, PMID:26776207). segemehl parameters -M -D and -S
    • -S regulates split alignments. It can be removed if you presume the RBP doesn't bind pre-mRNAs since it makes bam files much larger in size. -S option generated three additional segemehl output files (which for now cannot be redirected to output directory): trns.txt, sngl.bed, mult.bed
    • -D 2 : allows mismatches in seeds, especially helpful with 4SU and 6SG based CLIP if we expect >1 T>C or G>A conversions next to each other
    • -M 1 : controls the number of multiple hits. I only take unique hits, but up to -M 3 may give you additional usable reads
Calling peaks (not included in the pipeline)

This pipeline is optimized to be followed by the omniCLIP peak caller (Drewe-Boss 2018, https://github.com/philippdre/omniCLIP) which needs at least 2 replicates of CLIP and an input/background sample to call peaks against. I prefer to use RNA-seq data in the same cell line as background (1 replicate is enough).

About

a pipeline to map CLIP-seq data

Topics

Resources

Stars

2 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

CLIP_mapping

This is the pipeline I use to quickly map CLIP data. Modified from rnaroids pipeline (Neel Mukherjee) with tips from Philipp Boss (author of omniCLIP).

This pipeline is desinged to work with SUN Grid engine. (Or can be run locally on a machine with at least 80G memory).

This pipeline maps single end CLIP data to hg19. If the sequencing is paired-end (as in eCLIP), it will only take read1.

This pipeline uses snakemake (Köster 2012, https://academic.oup.com/bioinformatics/article/28/19/2520/290322). Snakefile is the main pipeline file. Config file config_clippipe.yaml lists variables that Snakefile needs (genome, annotation, adapter sequences, number of threads). Config file cluster_config.json regulates memory and other requirements when submitting jobs to the cluster. The bash script snakecharmer.sh is used to submit snakemake job to the cluster.

Preparation

Get the genome and annotation from Gencode and unzip it
wget ftp://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/GRCh37.p13.genome.fa.gz
wget ftp://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/gencode.v19.annotation.gtf.gz
Install miniconda and install the environment

Download the respective file from miniconda depending on your system (https://docs.conda.io/en/latest/miniconda.html). Follow installation instructions.

For example, for Linux, do:

wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
bash Miniconda3-latest-Linux-x86_64.sh
eval "$($HOME/miniconda3/bin/conda shell.bash hook)"
## create environment for CLIP mapping
conda create --name CLIP_mapping --file CLIP_mapping_conda_env.txt
conda activate CLIP_mapping
Get CLIP data

I use a small fastq fetching script for files from NCBI SRA, especially handy if there are many CLIP files from different RBPs. Make a sample sheet that looks like:

RBP1	SRR100{1..3}
RBP2	SRR100{4..5}

Run the script:

bash fetchFastq.sh example_sample_sheet

It should create directories "RBP1", "RBP2" etc. and put fastq files inside. In the example case, there should be 1 directory "ELAVL1" with 3 fastq files.

Note: I use --split-files by default to be able to always use the same script whether it is single end, paired end or has barcode read. Snakemake will always take read 1, and search for input files which look like *_1.fastq.gz. To use own fastq files, please add _1 to the file base name.

Run the pipeline

To check that everything OK, try snakemake dry run:

snakemake -nrp

It should give you wildcards for sample names, all rules that will be executed and output files. Then run the pipeline:

  • on a cluster qsub snakecharmer.sh or bash snakecharmer.sh
  • without cluster bash snakecharmer.sh or snakemake

If you are on the cluster, you might change "account" in cluster_config.json to your email address: you will get an email when the job is done, or aborted.

If you are not on the cluster, keep in mind that it may need quite some memory. Segemehl index for hg19 requires at least 64G and alignment usually takes ~75G.

Changing parameters

For small RNAs, mapping parameters could be a source of debate :-) Depending on library quality, one might want to make them more strict or more relaxed. This pipeline is trying to be very universal and find a compromise between many variants of CLIP data. Several things can be adjusted:

  • cutadapt -m 18 : this removes all reads shorter than 18nt, because it is progressively more difficult to map shorter reads. If one still might want to try and include those, change this number to 16, for example.
  • I try to include all possible small RNA adapters used. Of course, if you know exact adapter sequences for the sample (not always easy to find out in GEO!) you might want to add them to cutadapt rule, if they are not listed yet in the config file.
  • Deduplication: since I assume an old CLIP without UMIs, I do simple read collapsing. If one has UMIs, it is better to use UMI-tools(https://github.com/CGATOxford/UMI-tools) to deduplicate reads.
  • Mapping: I use segemehl (Hoffmann 2009, PMID: 19750212) because it especially well captures deletions which can be used as diagnostic events (Kassuhn 2016, PMID:26776207). segemehl parameters -M -D and -S
    • -S regulates split alignments. It can be removed if you presume the RBP doesn't bind pre-mRNAs since it makes bam files much larger in size. -S option generated three additional segemehl output files (which for now cannot be redirected to output directory): trns.txt, sngl.bed, mult.bed
    • -D 2 : allows mismatches in seeds, especially helpful with 4SU and 6SG based CLIP if we expect >1 T>C or G>A conversions next to each other
    • -M 1 : controls the number of multiple hits. I only take unique hits, but up to -M 3 may give you additional usable reads
Calling peaks (not included in the pipeline)

This pipeline is optimized to be followed by the omniCLIP peak caller (Drewe-Boss 2018, https://github.com/philippdre/omniCLIP) which needs at least 2 replicates of CLIP and an input/background sample to call peaks against. I prefer to use RNA-seq data in the same cell line as background (1 replicate is enough).

About

a pipeline to map CLIP-seq data

Topics

Resources

Stars

2 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CLIP_mapping

This is the pipeline I use to quickly map CLIP data. Modified from rnaroids pipeline (Neel Mukherjee) with tips from Philipp Boss (author of omniCLIP).

This pipeline is desinged to work with SUN Grid engine. (Or can be run locally on a machine with at least 80G memory).

This pipeline maps single end CLIP data to hg19. If the sequencing is paired-end (as in eCLIP), it will only take read1.

This pipeline uses snakemake (Köster 2012, https://academic.oup.com/bioinformatics/article/28/19/2520/290322). Snakefile is the main pipeline file. Config file config_clippipe.yaml lists variables that Snakefile needs (genome, annotation, adapter sequences, number of threads). Config file cluster_config.json regulates memory and other requirements when submitting jobs to the cluster. The bash script snakecharmer.sh is used to submit snakemake job to the cluster.

Preparation

Get the genome and annotation from Gencode and unzip it
wget ftp://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/GRCh37.p13.genome.fa.gz
wget ftp://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/gencode.v19.annotation.gtf.gz
Install miniconda and install the environment

Download the respective file from miniconda depending on your system (https://docs.conda.io/en/latest/miniconda.html). Follow installation instructions.

For example, for Linux, do:

wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
bash Miniconda3-latest-Linux-x86_64.sh
eval "$($HOME/miniconda3/bin/conda shell.bash hook)"
## create environment for CLIP mapping
conda create --name CLIP_mapping --file CLIP_mapping_conda_env.txt
conda activate CLIP_mapping
Get CLIP data

I use a small fastq fetching script for files from NCBI SRA, especially handy if there are many CLIP files from different RBPs. Make a sample sheet that looks like:

RBP1	SRR100{1..3}
RBP2	SRR100{4..5}

Run the script:

bash fetchFastq.sh example_sample_sheet

It should create directories "RBP1", "RBP2" etc. and put fastq files inside. In the example case, there should be 1 directory "ELAVL1" with 3 fastq files.

Note: I use --split-files by default to be able to always use the same script whether it is single end, paired end or has barcode read. Snakemake will always take read 1, and search for input files which look like *_1.fastq.gz. To use own fastq files, please add _1 to the file base name.

Run the pipeline

To check that everything OK, try snakemake dry run:

snakemake -nrp

It should give you wildcards for sample names, all rules that will be executed and output files. Then run the pipeline:

  • on a cluster qsub snakecharmer.sh or bash snakecharmer.sh
  • without cluster bash snakecharmer.sh or snakemake

If you are on the cluster, you might change "account" in cluster_config.json to your email address: you will get an email when the job is done, or aborted.

If you are not on the cluster, keep in mind that it may need quite some memory. Segemehl index for hg19 requires at least 64G and alignment usually takes ~75G.

Changing parameters

For small RNAs, mapping parameters could be a source of debate :-) Depending on library quality, one might want to make them more strict or more relaxed. This pipeline is trying to be very universal and find a compromise between many variants of CLIP data. Several things can be adjusted:

  • cutadapt -m 18 : this removes all reads shorter than 18nt, because it is progressively more difficult to map shorter reads. If one still might want to try and include those, change this number to 16, for example.
  • I try to include all possible small RNA adapters used. Of course, if you know exact adapter sequences for the sample (not always easy to find out in GEO!) you might want to add them to cutadapt rule, if they are not listed yet in the config file.
  • Deduplication: since I assume an old CLIP without UMIs, I do simple read collapsing. If one has UMIs, it is better to use UMI-tools(https://github.com/CGATOxford/UMI-tools) to deduplicate reads.
  • Mapping: I use segemehl (Hoffmann 2009, PMID: 19750212) because it especially well captures deletions which can be used as diagnostic events (Kassuhn 2016, PMID:26776207). segemehl parameters -M -D and -S
    • -S regulates split alignments. It can be removed if you presume the RBP doesn't bind pre-mRNAs since it makes bam files much larger in size. -S option generated three additional segemehl output files (which for now cannot be redirected to output directory): trns.txt, sngl.bed, mult.bed
    • -D 2 : allows mismatches in seeds, especially helpful with 4SU and 6SG based CLIP if we expect >1 T>C or G>A conversions next to each other
    • -M 1 : controls the number of multiple hits. I only take unique hits, but up to -M 3 may give you additional usable reads
Calling peaks (not included in the pipeline)

This pipeline is optimized to be followed by the omniCLIP peak caller (Drewe-Boss 2018, https://github.com/philippdre/omniCLIP) which needs at least 2 replicates of CLIP and an input/background sample to call peaks against. I prefer to use RNA-seq data in the same cell line as background (1 replicate is enough).

About

a pipeline to map CLIP-seq data

Topics

Resources

Stars

2 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

CLIP_mapping

This is the pipeline I use to quickly map CLIP data. Modified from rnaroids pipeline (Neel Mukherjee) with tips from Philipp Boss (author of omniCLIP).

This pipeline is desinged to work with SUN Grid engine. (Or can be run locally on a machine with at least 80G memory).

This pipeline maps single end CLIP data to hg19. If the sequencing is paired-end (as in eCLIP), it will only take read1.

This pipeline uses snakemake (Köster 2012, https://academic.oup.com/bioinformatics/article/28/19/2520/290322). Snakefile is the main pipeline file. Config file config_clippipe.yaml lists variables that Snakefile needs (genome, annotation, adapter sequences, number of threads). Config file cluster_config.json regulates memory and other requirements when submitting jobs to the cluster. The bash script snakecharmer.sh is used to submit snakemake job to the cluster.

Preparation

Get the genome and annotation from Gencode and unzip it
wget ftp://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/GRCh37.p13.genome.fa.gz
wget ftp://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/gencode.v19.annotation.gtf.gz
Install miniconda and install the environment

Download the respective file from miniconda depending on your system (https://docs.conda.io/en/latest/miniconda.html). Follow installation instructions.

For example, for Linux, do:

wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
bash Miniconda3-latest-Linux-x86_64.sh
eval "$($HOME/miniconda3/bin/conda shell.bash hook)"
## create environment for CLIP mapping
conda create --name CLIP_mapping --file CLIP_mapping_conda_env.txt
conda activate CLIP_mapping
Get CLIP data

I use a small fastq fetching script for files from NCBI SRA, especially handy if there are many CLIP files from different RBPs. Make a sample sheet that looks like:

RBP1	SRR100{1..3}
RBP2	SRR100{4..5}

Run the script:

bash fetchFastq.sh example_sample_sheet

It should create directories "RBP1", "RBP2" etc. and put fastq files inside. In the example case, there should be 1 directory "ELAVL1" with 3 fastq files.

Note: I use --split-files by default to be able to always use the same script whether it is single end, paired end or has barcode read. Snakemake will always take read 1, and search for input files which look like *_1.fastq.gz. To use own fastq files, please add _1 to the file base name.

Run the pipeline

To check that everything OK, try snakemake dry run:

snakemake -nrp

It should give you wildcards for sample names, all rules that will be executed and output files. Then run the pipeline:

  • on a cluster qsub snakecharmer.sh or bash snakecharmer.sh
  • without cluster bash snakecharmer.sh or snakemake

If you are on the cluster, you might change "account" in cluster_config.json to your email address: you will get an email when the job is done, or aborted.

If you are not on the cluster, keep in mind that it may need quite some memory. Segemehl index for hg19 requires at least 64G and alignment usually takes ~75G.

Changing parameters

For small RNAs, mapping parameters could be a source of debate :-) Depending on library quality, one might want to make them more strict or more relaxed. This pipeline is trying to be very universal and find a compromise between many variants of CLIP data. Several things can be adjusted:

  • cutadapt -m 18 : this removes all reads shorter than 18nt, because it is progressively more difficult to map shorter reads. If one still might want to try and include those, change this number to 16, for example.
  • I try to include all possible small RNA adapters used. Of course, if you know exact adapter sequences for the sample (not always easy to find out in GEO!) you might want to add them to cutadapt rule, if they are not listed yet in the config file.
  • Deduplication: since I assume an old CLIP without UMIs, I do simple read collapsing. If one has UMIs, it is better to use UMI-tools(https://github.com/CGATOxford/UMI-tools) to deduplicate reads.
  • Mapping: I use segemehl (Hoffmann 2009, PMID: 19750212) because it especially well captures deletions which can be used as diagnostic events (Kassuhn 2016, PMID:26776207). segemehl parameters -M -D and -S
    • -S regulates split alignments. It can be removed if you presume the RBP doesn't bind pre-mRNAs since it makes bam files much larger in size. -S option generated three additional segemehl output files (which for now cannot be redirected to output directory): trns.txt, sngl.bed, mult.bed
    • -D 2 : allows mismatches in seeds, especially helpful with 4SU and 6SG based CLIP if we expect >1 T>C or G>A conversions next to each other
    • -M 1 : controls the number of multiple hits. I only take unique hits, but up to -M 3 may give you additional usable reads
Calling peaks (not included in the pipeline)

This pipeline is optimized to be followed by the omniCLIP peak caller (Drewe-Boss 2018, https://github.com/philippdre/omniCLIP) which needs at least 2 replicates of CLIP and an input/background sample to call peaks against. I prefer to use RNA-seq data in the same cell line as background (1 replicate is enough).

About

a pipeline to map CLIP-seq data

Topics

Resources

Stars

2 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

CLIP_mapping

This is the pipeline I use to quickly map CLIP data. Modified from rnaroids pipeline (Neel Mukherjee) with tips from Philipp Boss (author of omniCLIP).

This pipeline is desinged to work with SUN Grid engine. (Or can be run locally on a machine with at least 80G memory).

This pipeline maps single end CLIP data to hg19. If the sequencing is paired-end (as in eCLIP), it will only take read1.

This pipeline uses snakemake (Köster 2012, https://academic.oup.com/bioinformatics/article/28/19/2520/290322). Snakefile is the main pipeline file. Config file config_clippipe.yaml lists variables that Snakefile needs (genome, annotation, adapter sequences, number of threads). Config file cluster_config.json regulates memory and other requirements when submitting jobs to the cluster. The bash script snakecharmer.sh is used to submit snakemake job to the cluster.

Preparation

Get the genome and annotation from Gencode and unzip it
wget ftp://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/GRCh37.p13.genome.fa.gz
wget ftp://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/gencode.v19.annotation.gtf.gz
Install miniconda and install the environment

Download the respective file from miniconda depending on your system (https://docs.conda.io/en/latest/miniconda.html). Follow installation instructions.

For example, for Linux, do:

wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh
bash Miniconda3-latest-Linux-x86_64.sh
eval "$($HOME/miniconda3/bin/conda shell.bash hook)"
## create environment for CLIP mapping
conda create --name CLIP_mapping --file CLIP_mapping_conda_env.txt
conda activate CLIP_mapping
Get CLIP data

I use a small fastq fetching script for files from NCBI SRA, especially handy if there are many CLIP files from different RBPs. Make a sample sheet that looks like:

RBP1	SRR100{1..3}
RBP2	SRR100{4..5}

Run the script:

bash fetchFastq.sh example_sample_sheet

It should create directories "RBP1", "RBP2" etc. and put fastq files inside. In the example case, there should be 1 directory "ELAVL1" with 3 fastq files.

Note: I use --split-files by default to be able to always use the same script whether it is single end, paired end or has barcode read. Snakemake will always take read 1, and search for input files which look like *_1.fastq.gz. To use own fastq files, please add _1 to the file base name.

Run the pipeline

To check that everything OK, try snakemake dry run:

snakemake -nrp

It should give you wildcards for sample names, all rules that will be executed and output files. Then run the pipeline:

  • on a cluster qsub snakecharmer.sh or bash snakecharmer.sh
  • without cluster bash snakecharmer.sh or snakemake

If you are on the cluster, you might change "account" in cluster_config.json to your email address: you will get an email when the job is done, or aborted.

If you are not on the cluster, keep in mind that it may need quite some memory. Segemehl index for hg19 requires at least 64G and alignment usually takes ~75G.

Changing parameters

For small RNAs, mapping parameters could be a source of debate :-) Depending on library quality, one might want to make them more strict or more relaxed. This pipeline is trying to be very universal and find a compromise between many variants of CLIP data. Several things can be adjusted:

  • cutadapt -m 18 : this removes all reads shorter than 18nt, because it is progressively more difficult to map shorter reads. If one still might want to try and include those, change this number to 16, for example.
  • I try to include all possible small RNA adapters used. Of course, if you know exact adapter sequences for the sample (not always easy to find out in GEO!) you might want to add them to cutadapt rule, if they are not listed yet in the config file.
  • Deduplication: since I assume an old CLIP without UMIs, I do simple read collapsing. If one has UMIs, it is better to use UMI-tools(https://github.com/CGATOxford/UMI-tools) to deduplicate reads.
  • Mapping: I use segemehl (Hoffmann 2009, PMID: 19750212) because it especially well captures deletions which can be used as diagnostic events (Kassuhn 2016, PMID:26776207). segemehl parameters -M -D and -S
    • -S regulates split alignments. It can be removed if you presume the RBP doesn't bind pre-mRNAs since it makes bam files much larger in size. -S option generated three additional segemehl output files (which for now cannot be redirected to output directory): trns.txt, sngl.bed, mult.bed
    • -D 2 : allows mismatches in seeds, especially helpful with 4SU and 6SG based CLIP if we expect >1 T>C or G>A conversions next to each other
    • -M 1 : controls the number of multiple hits. I only take unique hits, but up to -M 3 may give you additional usable reads
Calling peaks (not included in the pipeline)

This pipeline is optimized to be followed by the omniCLIP peak caller (Drewe-Boss 2018, https://github.com/philippdre/omniCLIP) which needs at least 2 replicates of CLIP and an input/background sample to call peaks against. I prefer to use RNA-seq data in the same cell line as background (1 replicate is enough).

About

a pipeline to map CLIP-seq data

Topics

Resources

Stars

2 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages