Repository files navigation

funkea

Perform functional enrichment analysis at scale.

Install

funkea is available on PyPI, and can be installed using pip:

pip install funkea
# OR
pip install 'funkea[jax]'# optional JAX dependency -- enables LDSC on GPU

If you want to build from source, you can clone the repository and install using:

git clone https://github.com/BenevolentAI/funkea
cd funkea
pip install .

Note that funkea requires Python 3.10 or higher, and that it uses Scala for some of its user-defined functions. The above command will fail if you do not have a Scala compiler installed, or if it is not on your path. If you do not have a Scala compiler installed, you can install one from here (version 2.12.x). Note that the Scala installer can sometimes have issues adding the compiler to your path, so you may need to do this manually.

Quickstart

funkea was built with composability in mind, but also ships with 5 popular enrichment methods out of the box. The simplest and fastest method is a Fisher's exact test on the overlapped annotations:

fromfunkea.implementationsimportFisherfrompyspark.sqlimportDataFramesumstats: DataFrame= ... # your GWAS sumstats# perform tissue enrichment using GTEx datamodel=Fisher.default()
enrichment=model.transform(sumstats)

This assumes that the default filepaths are set in the file-registry. If you have not set the file-registry, you can pass the filepaths directly to the data components:

fromfunkea.implementationsimportFisherfromfunkea.coreimportdatafrompyspark.sqlimportDataFramesumstats: DataFrame= ... # your GWAS sumstats# define your annotation componentannotation=data.AnnotationComponent(
# the column names in annotation dataframecolumns=data.AnnotationColumns(
annotation_id="gene_id",
partition_id="tissue_id",
),
# the path to the GTEx dataset (needs to be in parquet format)dataset="path/to/gtex_dataset.parquet",
# the type of partitioning used in the annotation dataset# (either data.PartitionType.SOFT or data.PartitionType.HARD)# see docs for more informationpartition_type=data.PartitionType.SOFT,
)
# perform tissue enrichment using GTEx datamodel=Fisher.default(annotation=annotation)

The dataset can also be passed as a DataFrame directly, if you have already loaded it into your Spark session.

fromfunkea.coreimportdatafrompyspark.sqlimportDataFrame, SparkSession# load the GTEx datasetspark: SparkSession= ... # your Spark sessiongtex: DataFrame=spark.read.parquet("path/to/gtex_dataset.parquet")
# define your annotation componentannotation=data.AnnotationComponent(
columns=data.AnnotationColumns(
annotation_id="gene_id",
partition_id="tissue_id",
),
# the GTEx datasetdataset=gtex, # <-- pass the DataFrame directlypartition_type=data.PartitionType.SOFT,
)

It is generally recommended to set the file-paths using the file-registry, as it is easy to forget to pass the filepaths to all the components. The above example is simple, where there is only one required data source (barring the GWAS sumstats), but more complex methods may require multiple data sources. For example, the GARFIELD method requires linkage disequilibrium (LD) estimates, variant-level controlling covariates, and the annotation component.

Introduction

funkea is a Python library for large-scale functional enrichment analysis. It provides 5 popular enrichment methods, and also allows for experimentation by composing different components. It is written in Spark, and allows users to run an arbitrary number of GWAS studies concurrently, given the resources are available. It also provides a CLI for managing data sources.

funkea has a few concepts used for abstraction, such that all methods could be unified. A view of the schematic is outlined below

schematic

i.e. each workflow consists of (1) a data pipeline; and (2) an enrichment method. The former filters down the sumstats (variant_selection), creates loci from the remaining variants (locus_definition) and then finally associates these loci with annotations (annotation). The latter then takes the loci (including their annotations) and computes the study-wide enrichments for each annotation partition, and its respective significance.

The variant selection and locus definitions are composed by the user, but each of the enrichment methods provided by funkea provide default configurations. The user can also define their own annotation component, which is required for all enrichment methods.

Setting default filepaths

For ease of use, funkea uses a file-registry for its source of truth of various data sources. These need to be set by a user, which can be set easily by using the funkea-cli. For example, if the user has put all the data into a single directory, the registry can be set like so:

funkea-cli init --from-stem <PATH_TO_DATASET>

This will set the default filepaths for all the data sources, where the filenames will be the key in JSON object. For example, if the user has put all the data into /path/to/data, the registry will look like:

{
"gtex": "/path/to/data/gtex",
"ld_reference_data": "/path/to/data/ld_reference_data",
"depict_null_loci": "/path/to/data/depict_null_loci",
"chromosomes": "/path/to/data/chromosomes",
"snpsea_background_variants": "/path/to/data/snpsea_background_variants",
"garfield_control_covariates": "/path/to/data/garfield_control_covariates",
"ldsc_controlling_ld_scores": "/path/to/data/ldsc_controlling_ld_scores",
"ldsc_weighting_ld_scores": "/path/to/data/ldsc_weighting_ld_scores"
}

The filepaths can also be set individually using the funkea-cli:

funkea-cli init

This will start an interactive prompt, where the user can set the filepaths individually. The registry can also be edited manually at funkea/core/resources/file_registry.json, but it is recommended to use the CLI, as it can be difficult to find the registry file.

Alternatively, the file paths can be specified directly in the components (see quickstart). It is generally recommended to use the CLI, as it is easy to forget to pass the filepaths to all the components, especially for enrichment methods which require multiple data sources (LD data, controlling covariates etc.).

Testing

funkea uses pytest for testing. To run the tests, run the following command from the root directory:

make test

This will install most dependencies if they are not present, except the Java runtime, as this depends heavily on the user's system. Make sure to install this yourself first before running the tests.

Note: some tests will be skipped on Linux systems with Aarch64 architecture, as JAX does not support this architecture. This is a known issue, and will (hopefully) be fixed in the future (see here).

Documentation

The documentation is hosted on readthedocs. It is automatically built from the docs directory in the main branch. To build the documentation locally, run the following command from the root directory:

make docs

About

Perform functional enrichment analysis at scale.

Resources

Stars

5 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

funkea

Perform functional enrichment analysis at scale.

Install

funkea is available on PyPI, and can be installed using pip:

pip install funkea
# OR
pip install 'funkea[jax]'# optional JAX dependency -- enables LDSC on GPU

If you want to build from source, you can clone the repository and install using:

git clone https://github.com/BenevolentAI/funkea
cd funkea
pip install .

Note that funkea requires Python 3.10 or higher, and that it uses Scala for some of its user-defined functions. The above command will fail if you do not have a Scala compiler installed, or if it is not on your path. If you do not have a Scala compiler installed, you can install one from here (version 2.12.x). Note that the Scala installer can sometimes have issues adding the compiler to your path, so you may need to do this manually.

Quickstart

funkea was built with composability in mind, but also ships with 5 popular enrichment methods out of the box. The simplest and fastest method is a Fisher's exact test on the overlapped annotations:

fromfunkea.implementationsimportFisherfrompyspark.sqlimportDataFramesumstats: DataFrame= ... # your GWAS sumstats# perform tissue enrichment using GTEx datamodel=Fisher.default()
enrichment=model.transform(sumstats)

This assumes that the default filepaths are set in the file-registry. If you have not set the file-registry, you can pass the filepaths directly to the data components:

fromfunkea.implementationsimportFisherfromfunkea.coreimportdatafrompyspark.sqlimportDataFramesumstats: DataFrame= ... # your GWAS sumstats# define your annotation componentannotation=data.AnnotationComponent(
# the column names in annotation dataframecolumns=data.AnnotationColumns(
annotation_id="gene_id",
partition_id="tissue_id",
),
# the path to the GTEx dataset (needs to be in parquet format)dataset="path/to/gtex_dataset.parquet",
# the type of partitioning used in the annotation dataset# (either data.PartitionType.SOFT or data.PartitionType.HARD)# see docs for more informationpartition_type=data.PartitionType.SOFT,
)
# perform tissue enrichment using GTEx datamodel=Fisher.default(annotation=annotation)

The dataset can also be passed as a DataFrame directly, if you have already loaded it into your Spark session.

fromfunkea.coreimportdatafrompyspark.sqlimportDataFrame, SparkSession# load the GTEx datasetspark: SparkSession= ... # your Spark sessiongtex: DataFrame=spark.read.parquet("path/to/gtex_dataset.parquet")
# define your annotation componentannotation=data.AnnotationComponent(
columns=data.AnnotationColumns(
annotation_id="gene_id",
partition_id="tissue_id",
),
# the GTEx datasetdataset=gtex, # <-- pass the DataFrame directlypartition_type=data.PartitionType.SOFT,
)

It is generally recommended to set the file-paths using the file-registry, as it is easy to forget to pass the filepaths to all the components. The above example is simple, where there is only one required data source (barring the GWAS sumstats), but more complex methods may require multiple data sources. For example, the GARFIELD method requires linkage disequilibrium (LD) estimates, variant-level controlling covariates, and the annotation component.

Introduction

funkea is a Python library for large-scale functional enrichment analysis. It provides 5 popular enrichment methods, and also allows for experimentation by composing different components. It is written in Spark, and allows users to run an arbitrary number of GWAS studies concurrently, given the resources are available. It also provides a CLI for managing data sources.

funkea has a few concepts used for abstraction, such that all methods could be unified. A view of the schematic is outlined below

schematic

i.e. each workflow consists of (1) a data pipeline; and (2) an enrichment method. The former filters down the sumstats (variant_selection), creates loci from the remaining variants (locus_definition) and then finally associates these loci with annotations (annotation). The latter then takes the loci (including their annotations) and computes the study-wide enrichments for each annotation partition, and its respective significance.

The variant selection and locus definitions are composed by the user, but each of the enrichment methods provided by funkea provide default configurations. The user can also define their own annotation component, which is required for all enrichment methods.

Setting default filepaths

For ease of use, funkea uses a file-registry for its source of truth of various data sources. These need to be set by a user, which can be set easily by using the funkea-cli. For example, if the user has put all the data into a single directory, the registry can be set like so:

funkea-cli init --from-stem <PATH_TO_DATASET>

This will set the default filepaths for all the data sources, where the filenames will be the key in JSON object. For example, if the user has put all the data into /path/to/data, the registry will look like:

{
"gtex": "/path/to/data/gtex",
"ld_reference_data": "/path/to/data/ld_reference_data",
"depict_null_loci": "/path/to/data/depict_null_loci",
"chromosomes": "/path/to/data/chromosomes",
"snpsea_background_variants": "/path/to/data/snpsea_background_variants",
"garfield_control_covariates": "/path/to/data/garfield_control_covariates",
"ldsc_controlling_ld_scores": "/path/to/data/ldsc_controlling_ld_scores",
"ldsc_weighting_ld_scores": "/path/to/data/ldsc_weighting_ld_scores"
}

The filepaths can also be set individually using the funkea-cli:

funkea-cli init

This will start an interactive prompt, where the user can set the filepaths individually. The registry can also be edited manually at funkea/core/resources/file_registry.json, but it is recommended to use the CLI, as it can be difficult to find the registry file.

Alternatively, the file paths can be specified directly in the components (see quickstart). It is generally recommended to use the CLI, as it is easy to forget to pass the filepaths to all the components, especially for enrichment methods which require multiple data sources (LD data, controlling covariates etc.).

Testing

funkea uses pytest for testing. To run the tests, run the following command from the root directory:

make test

This will install most dependencies if they are not present, except the Java runtime, as this depends heavily on the user's system. Make sure to install this yourself first before running the tests.

Note: some tests will be skipped on Linux systems with Aarch64 architecture, as JAX does not support this architecture. This is a known issue, and will (hopefully) be fixed in the future (see here).

Documentation

The documentation is hosted on readthedocs. It is automatically built from the docs directory in the main branch. To build the documentation locally, run the following command from the root directory:

make docs

About

Perform functional enrichment analysis at scale.

Resources

Stars

5 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

funkea

Perform functional enrichment analysis at scale.

Install

funkea is available on PyPI, and can be installed using pip:

pip install funkea
# OR
pip install 'funkea[jax]'# optional JAX dependency -- enables LDSC on GPU

If you want to build from source, you can clone the repository and install using:

git clone https://github.com/BenevolentAI/funkea
cd funkea
pip install .

Note that funkea requires Python 3.10 or higher, and that it uses Scala for some of its user-defined functions. The above command will fail if you do not have a Scala compiler installed, or if it is not on your path. If you do not have a Scala compiler installed, you can install one from here (version 2.12.x). Note that the Scala installer can sometimes have issues adding the compiler to your path, so you may need to do this manually.

Quickstart

funkea was built with composability in mind, but also ships with 5 popular enrichment methods out of the box. The simplest and fastest method is a Fisher's exact test on the overlapped annotations:

fromfunkea.implementationsimportFisherfrompyspark.sqlimportDataFramesumstats: DataFrame= ... # your GWAS sumstats# perform tissue enrichment using GTEx datamodel=Fisher.default()
enrichment=model.transform(sumstats)

This assumes that the default filepaths are set in the file-registry. If you have not set the file-registry, you can pass the filepaths directly to the data components:

fromfunkea.implementationsimportFisherfromfunkea.coreimportdatafrompyspark.sqlimportDataFramesumstats: DataFrame= ... # your GWAS sumstats# define your annotation componentannotation=data.AnnotationComponent(
# the column names in annotation dataframecolumns=data.AnnotationColumns(
annotation_id="gene_id",
partition_id="tissue_id",
),
# the path to the GTEx dataset (needs to be in parquet format)dataset="path/to/gtex_dataset.parquet",
# the type of partitioning used in the annotation dataset# (either data.PartitionType.SOFT or data.PartitionType.HARD)# see docs for more informationpartition_type=data.PartitionType.SOFT,
)
# perform tissue enrichment using GTEx datamodel=Fisher.default(annotation=annotation)

The dataset can also be passed as a DataFrame directly, if you have already loaded it into your Spark session.

fromfunkea.coreimportdatafrompyspark.sqlimportDataFrame, SparkSession# load the GTEx datasetspark: SparkSession= ... # your Spark sessiongtex: DataFrame=spark.read.parquet("path/to/gtex_dataset.parquet")
# define your annotation componentannotation=data.AnnotationComponent(
columns=data.AnnotationColumns(
annotation_id="gene_id",
partition_id="tissue_id",
),
# the GTEx datasetdataset=gtex, # <-- pass the DataFrame directlypartition_type=data.PartitionType.SOFT,
)

It is generally recommended to set the file-paths using the file-registry, as it is easy to forget to pass the filepaths to all the components. The above example is simple, where there is only one required data source (barring the GWAS sumstats), but more complex methods may require multiple data sources. For example, the GARFIELD method requires linkage disequilibrium (LD) estimates, variant-level controlling covariates, and the annotation component.

Introduction

funkea is a Python library for large-scale functional enrichment analysis. It provides 5 popular enrichment methods, and also allows for experimentation by composing different components. It is written in Spark, and allows users to run an arbitrary number of GWAS studies concurrently, given the resources are available. It also provides a CLI for managing data sources.

funkea has a few concepts used for abstraction, such that all methods could be unified. A view of the schematic is outlined below

schematic

i.e. each workflow consists of (1) a data pipeline; and (2) an enrichment method. The former filters down the sumstats (variant_selection), creates loci from the remaining variants (locus_definition) and then finally associates these loci with annotations (annotation). The latter then takes the loci (including their annotations) and computes the study-wide enrichments for each annotation partition, and its respective significance.

The variant selection and locus definitions are composed by the user, but each of the enrichment methods provided by funkea provide default configurations. The user can also define their own annotation component, which is required for all enrichment methods.

Setting default filepaths

For ease of use, funkea uses a file-registry for its source of truth of various data sources. These need to be set by a user, which can be set easily by using the funkea-cli. For example, if the user has put all the data into a single directory, the registry can be set like so:

funkea-cli init --from-stem <PATH_TO_DATASET>

This will set the default filepaths for all the data sources, where the filenames will be the key in JSON object. For example, if the user has put all the data into /path/to/data, the registry will look like:

{
"gtex": "/path/to/data/gtex",
"ld_reference_data": "/path/to/data/ld_reference_data",
"depict_null_loci": "/path/to/data/depict_null_loci",
"chromosomes": "/path/to/data/chromosomes",
"snpsea_background_variants": "/path/to/data/snpsea_background_variants",
"garfield_control_covariates": "/path/to/data/garfield_control_covariates",
"ldsc_controlling_ld_scores": "/path/to/data/ldsc_controlling_ld_scores",
"ldsc_weighting_ld_scores": "/path/to/data/ldsc_weighting_ld_scores"
}

The filepaths can also be set individually using the funkea-cli:

funkea-cli init

This will start an interactive prompt, where the user can set the filepaths individually. The registry can also be edited manually at funkea/core/resources/file_registry.json, but it is recommended to use the CLI, as it can be difficult to find the registry file.

Alternatively, the file paths can be specified directly in the components (see quickstart). It is generally recommended to use the CLI, as it is easy to forget to pass the filepaths to all the components, especially for enrichment methods which require multiple data sources (LD data, controlling covariates etc.).

Testing

funkea uses pytest for testing. To run the tests, run the following command from the root directory:

make test

This will install most dependencies if they are not present, except the Java runtime, as this depends heavily on the user's system. Make sure to install this yourself first before running the tests.

Note: some tests will be skipped on Linux systems with Aarch64 architecture, as JAX does not support this architecture. This is a known issue, and will (hopefully) be fixed in the future (see here).

Documentation

The documentation is hosted on readthedocs. It is automatically built from the docs directory in the main branch. To build the documentation locally, run the following command from the root directory:

make docs

About

Perform functional enrichment analysis at scale.

Resources

Stars

5 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

funkea

Perform functional enrichment analysis at scale.

Install

funkea is available on PyPI, and can be installed using pip:

pip install funkea
# OR
pip install 'funkea[jax]'# optional JAX dependency -- enables LDSC on GPU

If you want to build from source, you can clone the repository and install using:

git clone https://github.com/BenevolentAI/funkea
cd funkea
pip install .

Note that funkea requires Python 3.10 or higher, and that it uses Scala for some of its user-defined functions. The above command will fail if you do not have a Scala compiler installed, or if it is not on your path. If you do not have a Scala compiler installed, you can install one from here (version 2.12.x). Note that the Scala installer can sometimes have issues adding the compiler to your path, so you may need to do this manually.

Quickstart

funkea was built with composability in mind, but also ships with 5 popular enrichment methods out of the box. The simplest and fastest method is a Fisher's exact test on the overlapped annotations:

fromfunkea.implementationsimportFisherfrompyspark.sqlimportDataFramesumstats: DataFrame= ... # your GWAS sumstats# perform tissue enrichment using GTEx datamodel=Fisher.default()
enrichment=model.transform(sumstats)

This assumes that the default filepaths are set in the file-registry. If you have not set the file-registry, you can pass the filepaths directly to the data components:

fromfunkea.implementationsimportFisherfromfunkea.coreimportdatafrompyspark.sqlimportDataFramesumstats: DataFrame= ... # your GWAS sumstats# define your annotation componentannotation=data.AnnotationComponent(
# the column names in annotation dataframecolumns=data.AnnotationColumns(
annotation_id="gene_id",
partition_id="tissue_id",
),
# the path to the GTEx dataset (needs to be in parquet format)dataset="path/to/gtex_dataset.parquet",
# the type of partitioning used in the annotation dataset# (either data.PartitionType.SOFT or data.PartitionType.HARD)# see docs for more informationpartition_type=data.PartitionType.SOFT,
)
# perform tissue enrichment using GTEx datamodel=Fisher.default(annotation=annotation)

The dataset can also be passed as a DataFrame directly, if you have already loaded it into your Spark session.

fromfunkea.coreimportdatafrompyspark.sqlimportDataFrame, SparkSession# load the GTEx datasetspark: SparkSession= ... # your Spark sessiongtex: DataFrame=spark.read.parquet("path/to/gtex_dataset.parquet")
# define your annotation componentannotation=data.AnnotationComponent(
columns=data.AnnotationColumns(
annotation_id="gene_id",
partition_id="tissue_id",
),
# the GTEx datasetdataset=gtex, # <-- pass the DataFrame directlypartition_type=data.PartitionType.SOFT,
)

It is generally recommended to set the file-paths using the file-registry, as it is easy to forget to pass the filepaths to all the components. The above example is simple, where there is only one required data source (barring the GWAS sumstats), but more complex methods may require multiple data sources. For example, the GARFIELD method requires linkage disequilibrium (LD) estimates, variant-level controlling covariates, and the annotation component.

Introduction

funkea is a Python library for large-scale functional enrichment analysis. It provides 5 popular enrichment methods, and also allows for experimentation by composing different components. It is written in Spark, and allows users to run an arbitrary number of GWAS studies concurrently, given the resources are available. It also provides a CLI for managing data sources.

funkea has a few concepts used for abstraction, such that all methods could be unified. A view of the schematic is outlined below

schematic

i.e. each workflow consists of (1) a data pipeline; and (2) an enrichment method. The former filters down the sumstats (variant_selection), creates loci from the remaining variants (locus_definition) and then finally associates these loci with annotations (annotation). The latter then takes the loci (including their annotations) and computes the study-wide enrichments for each annotation partition, and its respective significance.

The variant selection and locus definitions are composed by the user, but each of the enrichment methods provided by funkea provide default configurations. The user can also define their own annotation component, which is required for all enrichment methods.

Setting default filepaths

For ease of use, funkea uses a file-registry for its source of truth of various data sources. These need to be set by a user, which can be set easily by using the funkea-cli. For example, if the user has put all the data into a single directory, the registry can be set like so:

funkea-cli init --from-stem <PATH_TO_DATASET>

This will set the default filepaths for all the data sources, where the filenames will be the key in JSON object. For example, if the user has put all the data into /path/to/data, the registry will look like:

{
"gtex": "/path/to/data/gtex",
"ld_reference_data": "/path/to/data/ld_reference_data",
"depict_null_loci": "/path/to/data/depict_null_loci",
"chromosomes": "/path/to/data/chromosomes",
"snpsea_background_variants": "/path/to/data/snpsea_background_variants",
"garfield_control_covariates": "/path/to/data/garfield_control_covariates",
"ldsc_controlling_ld_scores": "/path/to/data/ldsc_controlling_ld_scores",
"ldsc_weighting_ld_scores": "/path/to/data/ldsc_weighting_ld_scores"
}

The filepaths can also be set individually using the funkea-cli:

funkea-cli init

This will start an interactive prompt, where the user can set the filepaths individually. The registry can also be edited manually at funkea/core/resources/file_registry.json, but it is recommended to use the CLI, as it can be difficult to find the registry file.

Alternatively, the file paths can be specified directly in the components (see quickstart). It is generally recommended to use the CLI, as it is easy to forget to pass the filepaths to all the components, especially for enrichment methods which require multiple data sources (LD data, controlling covariates etc.).

Testing

funkea uses pytest for testing. To run the tests, run the following command from the root directory:

make test

This will install most dependencies if they are not present, except the Java runtime, as this depends heavily on the user's system. Make sure to install this yourself first before running the tests.

Note: some tests will be skipped on Linux systems with Aarch64 architecture, as JAX does not support this architecture. This is a known issue, and will (hopefully) be fixed in the future (see here).

Documentation

The documentation is hosted on readthedocs. It is automatically built from the docs directory in the main branch. To build the documentation locally, run the following command from the root directory:

make docs

About

Perform functional enrichment analysis at scale.

Resources

Stars

5 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

funkea

Perform functional enrichment analysis at scale.

Install

funkea is available on PyPI, and can be installed using pip:

pip install funkea
# OR
pip install 'funkea[jax]'# optional JAX dependency -- enables LDSC on GPU

If you want to build from source, you can clone the repository and install using:

git clone https://github.com/BenevolentAI/funkea
cd funkea
pip install .

Note that funkea requires Python 3.10 or higher, and that it uses Scala for some of its user-defined functions. The above command will fail if you do not have a Scala compiler installed, or if it is not on your path. If you do not have a Scala compiler installed, you can install one from here (version 2.12.x). Note that the Scala installer can sometimes have issues adding the compiler to your path, so you may need to do this manually.

Quickstart

funkea was built with composability in mind, but also ships with 5 popular enrichment methods out of the box. The simplest and fastest method is a Fisher's exact test on the overlapped annotations:

fromfunkea.implementationsimportFisherfrompyspark.sqlimportDataFramesumstats: DataFrame= ... # your GWAS sumstats# perform tissue enrichment using GTEx datamodel=Fisher.default()
enrichment=model.transform(sumstats)

This assumes that the default filepaths are set in the file-registry. If you have not set the file-registry, you can pass the filepaths directly to the data components:

fromfunkea.implementationsimportFisherfromfunkea.coreimportdatafrompyspark.sqlimportDataFramesumstats: DataFrame= ... # your GWAS sumstats# define your annotation componentannotation=data.AnnotationComponent(
# the column names in annotation dataframecolumns=data.AnnotationColumns(
annotation_id="gene_id",
partition_id="tissue_id",
),
# the path to the GTEx dataset (needs to be in parquet format)dataset="path/to/gtex_dataset.parquet",
# the type of partitioning used in the annotation dataset# (either data.PartitionType.SOFT or data.PartitionType.HARD)# see docs for more informationpartition_type=data.PartitionType.SOFT,
)
# perform tissue enrichment using GTEx datamodel=Fisher.default(annotation=annotation)

The dataset can also be passed as a DataFrame directly, if you have already loaded it into your Spark session.

fromfunkea.coreimportdatafrompyspark.sqlimportDataFrame, SparkSession# load the GTEx datasetspark: SparkSession= ... # your Spark sessiongtex: DataFrame=spark.read.parquet("path/to/gtex_dataset.parquet")
# define your annotation componentannotation=data.AnnotationComponent(
columns=data.AnnotationColumns(
annotation_id="gene_id",
partition_id="tissue_id",
),
# the GTEx datasetdataset=gtex, # <-- pass the DataFrame directlypartition_type=data.PartitionType.SOFT,
)

It is generally recommended to set the file-paths using the file-registry, as it is easy to forget to pass the filepaths to all the components. The above example is simple, where there is only one required data source (barring the GWAS sumstats), but more complex methods may require multiple data sources. For example, the GARFIELD method requires linkage disequilibrium (LD) estimates, variant-level controlling covariates, and the annotation component.

Introduction

funkea is a Python library for large-scale functional enrichment analysis. It provides 5 popular enrichment methods, and also allows for experimentation by composing different components. It is written in Spark, and allows users to run an arbitrary number of GWAS studies concurrently, given the resources are available. It also provides a CLI for managing data sources.

funkea has a few concepts used for abstraction, such that all methods could be unified. A view of the schematic is outlined below

schematic

i.e. each workflow consists of (1) a data pipeline; and (2) an enrichment method. The former filters down the sumstats (variant_selection), creates loci from the remaining variants (locus_definition) and then finally associates these loci with annotations (annotation). The latter then takes the loci (including their annotations) and computes the study-wide enrichments for each annotation partition, and its respective significance.

The variant selection and locus definitions are composed by the user, but each of the enrichment methods provided by funkea provide default configurations. The user can also define their own annotation component, which is required for all enrichment methods.

Setting default filepaths

For ease of use, funkea uses a file-registry for its source of truth of various data sources. These need to be set by a user, which can be set easily by using the funkea-cli. For example, if the user has put all the data into a single directory, the registry can be set like so:

funkea-cli init --from-stem <PATH_TO_DATASET>

This will set the default filepaths for all the data sources, where the filenames will be the key in JSON object. For example, if the user has put all the data into /path/to/data, the registry will look like:

{
"gtex": "/path/to/data/gtex",
"ld_reference_data": "/path/to/data/ld_reference_data",
"depict_null_loci": "/path/to/data/depict_null_loci",
"chromosomes": "/path/to/data/chromosomes",
"snpsea_background_variants": "/path/to/data/snpsea_background_variants",
"garfield_control_covariates": "/path/to/data/garfield_control_covariates",
"ldsc_controlling_ld_scores": "/path/to/data/ldsc_controlling_ld_scores",
"ldsc_weighting_ld_scores": "/path/to/data/ldsc_weighting_ld_scores"
}

The filepaths can also be set individually using the funkea-cli:

funkea-cli init

This will start an interactive prompt, where the user can set the filepaths individually. The registry can also be edited manually at funkea/core/resources/file_registry.json, but it is recommended to use the CLI, as it can be difficult to find the registry file.

Alternatively, the file paths can be specified directly in the components (see quickstart). It is generally recommended to use the CLI, as it is easy to forget to pass the filepaths to all the components, especially for enrichment methods which require multiple data sources (LD data, controlling covariates etc.).

Testing

funkea uses pytest for testing. To run the tests, run the following command from the root directory:

make test

This will install most dependencies if they are not present, except the Java runtime, as this depends heavily on the user's system. Make sure to install this yourself first before running the tests.

Note: some tests will be skipped on Linux systems with Aarch64 architecture, as JAX does not support this architecture. This is a known issue, and will (hopefully) be fixed in the future (see here).

Documentation

The documentation is hosted on readthedocs. It is automatically built from the docs directory in the main branch. To build the documentation locally, run the following command from the root directory:

make docs

About

Perform functional enrichment analysis at scale.

Resources

Stars

5 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

funkea

Perform functional enrichment analysis at scale.

Install

funkea is available on PyPI, and can be installed using pip:

pip install funkea
# OR
pip install 'funkea[jax]'# optional JAX dependency -- enables LDSC on GPU

If you want to build from source, you can clone the repository and install using:

git clone https://github.com/BenevolentAI/funkea
cd funkea
pip install .

Note that funkea requires Python 3.10 or higher, and that it uses Scala for some of its user-defined functions. The above command will fail if you do not have a Scala compiler installed, or if it is not on your path. If you do not have a Scala compiler installed, you can install one from here (version 2.12.x). Note that the Scala installer can sometimes have issues adding the compiler to your path, so you may need to do this manually.

Quickstart

funkea was built with composability in mind, but also ships with 5 popular enrichment methods out of the box. The simplest and fastest method is a Fisher's exact test on the overlapped annotations:

fromfunkea.implementationsimportFisherfrompyspark.sqlimportDataFramesumstats: DataFrame= ... # your GWAS sumstats# perform tissue enrichment using GTEx datamodel=Fisher.default()
enrichment=model.transform(sumstats)

This assumes that the default filepaths are set in the file-registry. If you have not set the file-registry, you can pass the filepaths directly to the data components:

fromfunkea.implementationsimportFisherfromfunkea.coreimportdatafrompyspark.sqlimportDataFramesumstats: DataFrame= ... # your GWAS sumstats# define your annotation componentannotation=data.AnnotationComponent(
# the column names in annotation dataframecolumns=data.AnnotationColumns(
annotation_id="gene_id",
partition_id="tissue_id",
),
# the path to the GTEx dataset (needs to be in parquet format)dataset="path/to/gtex_dataset.parquet",
# the type of partitioning used in the annotation dataset# (either data.PartitionType.SOFT or data.PartitionType.HARD)# see docs for more informationpartition_type=data.PartitionType.SOFT,
)
# perform tissue enrichment using GTEx datamodel=Fisher.default(annotation=annotation)

The dataset can also be passed as a DataFrame directly, if you have already loaded it into your Spark session.

fromfunkea.coreimportdatafrompyspark.sqlimportDataFrame, SparkSession# load the GTEx datasetspark: SparkSession= ... # your Spark sessiongtex: DataFrame=spark.read.parquet("path/to/gtex_dataset.parquet")
# define your annotation componentannotation=data.AnnotationComponent(
columns=data.AnnotationColumns(
annotation_id="gene_id",
partition_id="tissue_id",
),
# the GTEx datasetdataset=gtex, # <-- pass the DataFrame directlypartition_type=data.PartitionType.SOFT,
)

It is generally recommended to set the file-paths using the file-registry, as it is easy to forget to pass the filepaths to all the components. The above example is simple, where there is only one required data source (barring the GWAS sumstats), but more complex methods may require multiple data sources. For example, the GARFIELD method requires linkage disequilibrium (LD) estimates, variant-level controlling covariates, and the annotation component.

Introduction

funkea is a Python library for large-scale functional enrichment analysis. It provides 5 popular enrichment methods, and also allows for experimentation by composing different components. It is written in Spark, and allows users to run an arbitrary number of GWAS studies concurrently, given the resources are available. It also provides a CLI for managing data sources.

funkea has a few concepts used for abstraction, such that all methods could be unified. A view of the schematic is outlined below

schematic

i.e. each workflow consists of (1) a data pipeline; and (2) an enrichment method. The former filters down the sumstats (variant_selection), creates loci from the remaining variants (locus_definition) and then finally associates these loci with annotations (annotation). The latter then takes the loci (including their annotations) and computes the study-wide enrichments for each annotation partition, and its respective significance.

The variant selection and locus definitions are composed by the user, but each of the enrichment methods provided by funkea provide default configurations. The user can also define their own annotation component, which is required for all enrichment methods.

Setting default filepaths

For ease of use, funkea uses a file-registry for its source of truth of various data sources. These need to be set by a user, which can be set easily by using the funkea-cli. For example, if the user has put all the data into a single directory, the registry can be set like so:

funkea-cli init --from-stem <PATH_TO_DATASET>

This will set the default filepaths for all the data sources, where the filenames will be the key in JSON object. For example, if the user has put all the data into /path/to/data, the registry will look like:

{
"gtex": "/path/to/data/gtex",
"ld_reference_data": "/path/to/data/ld_reference_data",
"depict_null_loci": "/path/to/data/depict_null_loci",
"chromosomes": "/path/to/data/chromosomes",
"snpsea_background_variants": "/path/to/data/snpsea_background_variants",
"garfield_control_covariates": "/path/to/data/garfield_control_covariates",
"ldsc_controlling_ld_scores": "/path/to/data/ldsc_controlling_ld_scores",
"ldsc_weighting_ld_scores": "/path/to/data/ldsc_weighting_ld_scores"
}

The filepaths can also be set individually using the funkea-cli:

funkea-cli init

This will start an interactive prompt, where the user can set the filepaths individually. The registry can also be edited manually at funkea/core/resources/file_registry.json, but it is recommended to use the CLI, as it can be difficult to find the registry file.

Alternatively, the file paths can be specified directly in the components (see quickstart). It is generally recommended to use the CLI, as it is easy to forget to pass the filepaths to all the components, especially for enrichment methods which require multiple data sources (LD data, controlling covariates etc.).

Testing

funkea uses pytest for testing. To run the tests, run the following command from the root directory:

make test

This will install most dependencies if they are not present, except the Java runtime, as this depends heavily on the user's system. Make sure to install this yourself first before running the tests.

Note: some tests will be skipped on Linux systems with Aarch64 architecture, as JAX does not support this architecture. This is a known issue, and will (hopefully) be fixed in the future (see here).

Documentation

The documentation is hosted on readthedocs. It is automatically built from the docs directory in the main branch. To build the documentation locally, run the following command from the root directory:

make docs

About

Perform functional enrichment analysis at scale.

Resources

Stars

5 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

funkea

Perform functional enrichment analysis at scale.

Install

funkea is available on PyPI, and can be installed using pip:

pip install funkea
# OR
pip install 'funkea[jax]'# optional JAX dependency -- enables LDSC on GPU

If you want to build from source, you can clone the repository and install using:

git clone https://github.com/BenevolentAI/funkea
cd funkea
pip install .

Note that funkea requires Python 3.10 or higher, and that it uses Scala for some of its user-defined functions. The above command will fail if you do not have a Scala compiler installed, or if it is not on your path. If you do not have a Scala compiler installed, you can install one from here (version 2.12.x). Note that the Scala installer can sometimes have issues adding the compiler to your path, so you may need to do this manually.

Quickstart

funkea was built with composability in mind, but also ships with 5 popular enrichment methods out of the box. The simplest and fastest method is a Fisher's exact test on the overlapped annotations:

fromfunkea.implementationsimportFisherfrompyspark.sqlimportDataFramesumstats: DataFrame= ... # your GWAS sumstats# perform tissue enrichment using GTEx datamodel=Fisher.default()
enrichment=model.transform(sumstats)

This assumes that the default filepaths are set in the file-registry. If you have not set the file-registry, you can pass the filepaths directly to the data components:

fromfunkea.implementationsimportFisherfromfunkea.coreimportdatafrompyspark.sqlimportDataFramesumstats: DataFrame= ... # your GWAS sumstats# define your annotation componentannotation=data.AnnotationComponent(
# the column names in annotation dataframecolumns=data.AnnotationColumns(
annotation_id="gene_id",
partition_id="tissue_id",
),
# the path to the GTEx dataset (needs to be in parquet format)dataset="path/to/gtex_dataset.parquet",
# the type of partitioning used in the annotation dataset# (either data.PartitionType.SOFT or data.PartitionType.HARD)# see docs for more informationpartition_type=data.PartitionType.SOFT,
)
# perform tissue enrichment using GTEx datamodel=Fisher.default(annotation=annotation)

The dataset can also be passed as a DataFrame directly, if you have already loaded it into your Spark session.

fromfunkea.coreimportdatafrompyspark.sqlimportDataFrame, SparkSession# load the GTEx datasetspark: SparkSession= ... # your Spark sessiongtex: DataFrame=spark.read.parquet("path/to/gtex_dataset.parquet")
# define your annotation componentannotation=data.AnnotationComponent(
columns=data.AnnotationColumns(
annotation_id="gene_id",
partition_id="tissue_id",
),
# the GTEx datasetdataset=gtex, # <-- pass the DataFrame directlypartition_type=data.PartitionType.SOFT,
)

It is generally recommended to set the file-paths using the file-registry, as it is easy to forget to pass the filepaths to all the components. The above example is simple, where there is only one required data source (barring the GWAS sumstats), but more complex methods may require multiple data sources. For example, the GARFIELD method requires linkage disequilibrium (LD) estimates, variant-level controlling covariates, and the annotation component.

Introduction

funkea is a Python library for large-scale functional enrichment analysis. It provides 5 popular enrichment methods, and also allows for experimentation by composing different components. It is written in Spark, and allows users to run an arbitrary number of GWAS studies concurrently, given the resources are available. It also provides a CLI for managing data sources.

funkea has a few concepts used for abstraction, such that all methods could be unified. A view of the schematic is outlined below

schematic

i.e. each workflow consists of (1) a data pipeline; and (2) an enrichment method. The former filters down the sumstats (variant_selection), creates loci from the remaining variants (locus_definition) and then finally associates these loci with annotations (annotation). The latter then takes the loci (including their annotations) and computes the study-wide enrichments for each annotation partition, and its respective significance.

The variant selection and locus definitions are composed by the user, but each of the enrichment methods provided by funkea provide default configurations. The user can also define their own annotation component, which is required for all enrichment methods.

Setting default filepaths

For ease of use, funkea uses a file-registry for its source of truth of various data sources. These need to be set by a user, which can be set easily by using the funkea-cli. For example, if the user has put all the data into a single directory, the registry can be set like so:

funkea-cli init --from-stem <PATH_TO_DATASET>

This will set the default filepaths for all the data sources, where the filenames will be the key in JSON object. For example, if the user has put all the data into /path/to/data, the registry will look like:

{
"gtex": "/path/to/data/gtex",
"ld_reference_data": "/path/to/data/ld_reference_data",
"depict_null_loci": "/path/to/data/depict_null_loci",
"chromosomes": "/path/to/data/chromosomes",
"snpsea_background_variants": "/path/to/data/snpsea_background_variants",
"garfield_control_covariates": "/path/to/data/garfield_control_covariates",
"ldsc_controlling_ld_scores": "/path/to/data/ldsc_controlling_ld_scores",
"ldsc_weighting_ld_scores": "/path/to/data/ldsc_weighting_ld_scores"
}

The filepaths can also be set individually using the funkea-cli:

funkea-cli init

This will start an interactive prompt, where the user can set the filepaths individually. The registry can also be edited manually at funkea/core/resources/file_registry.json, but it is recommended to use the CLI, as it can be difficult to find the registry file.

Alternatively, the file paths can be specified directly in the components (see quickstart). It is generally recommended to use the CLI, as it is easy to forget to pass the filepaths to all the components, especially for enrichment methods which require multiple data sources (LD data, controlling covariates etc.).

Testing

funkea uses pytest for testing. To run the tests, run the following command from the root directory:

make test

This will install most dependencies if they are not present, except the Java runtime, as this depends heavily on the user's system. Make sure to install this yourself first before running the tests.

Note: some tests will be skipped on Linux systems with Aarch64 architecture, as JAX does not support this architecture. This is a known issue, and will (hopefully) be fixed in the future (see here).

Documentation

The documentation is hosted on readthedocs. It is automatically built from the docs directory in the main branch. To build the documentation locally, run the following command from the root directory:

make docs

About

Perform functional enrichment analysis at scale.

Resources

Stars

5 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

funkea

Perform functional enrichment analysis at scale.

Install

funkea is available on PyPI, and can be installed using pip:

pip install funkea
# OR
pip install 'funkea[jax]'# optional JAX dependency -- enables LDSC on GPU

If you want to build from source, you can clone the repository and install using:

git clone https://github.com/BenevolentAI/funkea
cd funkea
pip install .

Note that funkea requires Python 3.10 or higher, and that it uses Scala for some of its user-defined functions. The above command will fail if you do not have a Scala compiler installed, or if it is not on your path. If you do not have a Scala compiler installed, you can install one from here (version 2.12.x). Note that the Scala installer can sometimes have issues adding the compiler to your path, so you may need to do this manually.

Quickstart

funkea was built with composability in mind, but also ships with 5 popular enrichment methods out of the box. The simplest and fastest method is a Fisher's exact test on the overlapped annotations:

fromfunkea.implementationsimportFisherfrompyspark.sqlimportDataFramesumstats: DataFrame= ... # your GWAS sumstats# perform tissue enrichment using GTEx datamodel=Fisher.default()
enrichment=model.transform(sumstats)

This assumes that the default filepaths are set in the file-registry. If you have not set the file-registry, you can pass the filepaths directly to the data components:

fromfunkea.implementationsimportFisherfromfunkea.coreimportdatafrompyspark.sqlimportDataFramesumstats: DataFrame= ... # your GWAS sumstats# define your annotation componentannotation=data.AnnotationComponent(
# the column names in annotation dataframecolumns=data.AnnotationColumns(
annotation_id="gene_id",
partition_id="tissue_id",
),
# the path to the GTEx dataset (needs to be in parquet format)dataset="path/to/gtex_dataset.parquet",
# the type of partitioning used in the annotation dataset# (either data.PartitionType.SOFT or data.PartitionType.HARD)# see docs for more informationpartition_type=data.PartitionType.SOFT,
)
# perform tissue enrichment using GTEx datamodel=Fisher.default(annotation=annotation)

The dataset can also be passed as a DataFrame directly, if you have already loaded it into your Spark session.

fromfunkea.coreimportdatafrompyspark.sqlimportDataFrame, SparkSession# load the GTEx datasetspark: SparkSession= ... # your Spark sessiongtex: DataFrame=spark.read.parquet("path/to/gtex_dataset.parquet")
# define your annotation componentannotation=data.AnnotationComponent(
columns=data.AnnotationColumns(
annotation_id="gene_id",
partition_id="tissue_id",
),
# the GTEx datasetdataset=gtex, # <-- pass the DataFrame directlypartition_type=data.PartitionType.SOFT,
)

It is generally recommended to set the file-paths using the file-registry, as it is easy to forget to pass the filepaths to all the components. The above example is simple, where there is only one required data source (barring the GWAS sumstats), but more complex methods may require multiple data sources. For example, the GARFIELD method requires linkage disequilibrium (LD) estimates, variant-level controlling covariates, and the annotation component.

Introduction

funkea is a Python library for large-scale functional enrichment analysis. It provides 5 popular enrichment methods, and also allows for experimentation by composing different components. It is written in Spark, and allows users to run an arbitrary number of GWAS studies concurrently, given the resources are available. It also provides a CLI for managing data sources.

funkea has a few concepts used for abstraction, such that all methods could be unified. A view of the schematic is outlined below

schematic

i.e. each workflow consists of (1) a data pipeline; and (2) an enrichment method. The former filters down the sumstats (variant_selection), creates loci from the remaining variants (locus_definition) and then finally associates these loci with annotations (annotation). The latter then takes the loci (including their annotations) and computes the study-wide enrichments for each annotation partition, and its respective significance.

The variant selection and locus definitions are composed by the user, but each of the enrichment methods provided by funkea provide default configurations. The user can also define their own annotation component, which is required for all enrichment methods.

Setting default filepaths

For ease of use, funkea uses a file-registry for its source of truth of various data sources. These need to be set by a user, which can be set easily by using the funkea-cli. For example, if the user has put all the data into a single directory, the registry can be set like so:

funkea-cli init --from-stem <PATH_TO_DATASET>

This will set the default filepaths for all the data sources, where the filenames will be the key in JSON object. For example, if the user has put all the data into /path/to/data, the registry will look like:

{
"gtex": "/path/to/data/gtex",
"ld_reference_data": "/path/to/data/ld_reference_data",
"depict_null_loci": "/path/to/data/depict_null_loci",
"chromosomes": "/path/to/data/chromosomes",
"snpsea_background_variants": "/path/to/data/snpsea_background_variants",
"garfield_control_covariates": "/path/to/data/garfield_control_covariates",
"ldsc_controlling_ld_scores": "/path/to/data/ldsc_controlling_ld_scores",
"ldsc_weighting_ld_scores": "/path/to/data/ldsc_weighting_ld_scores"
}

The filepaths can also be set individually using the funkea-cli:

funkea-cli init

This will start an interactive prompt, where the user can set the filepaths individually. The registry can also be edited manually at funkea/core/resources/file_registry.json, but it is recommended to use the CLI, as it can be difficult to find the registry file.

Alternatively, the file paths can be specified directly in the components (see quickstart). It is generally recommended to use the CLI, as it is easy to forget to pass the filepaths to all the components, especially for enrichment methods which require multiple data sources (LD data, controlling covariates etc.).

Testing

funkea uses pytest for testing. To run the tests, run the following command from the root directory:

make test

This will install most dependencies if they are not present, except the Java runtime, as this depends heavily on the user's system. Make sure to install this yourself first before running the tests.

Note: some tests will be skipped on Linux systems with Aarch64 architecture, as JAX does not support this architecture. This is a known issue, and will (hopefully) be fixed in the future (see here).

Documentation

The documentation is hosted on readthedocs. It is automatically built from the docs directory in the main branch. To build the documentation locally, run the following command from the root directory:

make docs

About

Perform functional enrichment analysis at scale.

Resources

Stars

5 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages