Repository files navigation

Spatial Segmentation

A one sentence summary of purpose and methodology. Used for creating an overview tables.

Repository: openproblems-bio/task_spatial_segmentation

Description

Provide a clear and concise description of your task, detailing the specific problem it aims to solve. Outline the input data types, the expected output, and any assumptions or constraints. Be sure to explain any terminology or concepts that are essential for understanding the task.

Explain the motivation behind your proposed task. Describe the biological or computational problem you aim to address and why it’s important. Discuss the current state of research in this area and any gaps or challenges that your task could help address. This section should convince readers of the significance and relevance of your task.

Authors & contributors

nameroles
John Doeauthor, maintainer

API

flowchart TB
file_common_ist("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-ist-dataset'>Common iST Dataset</a>")
comp_data_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-data-processor'>Data processor</a>"/]
file_scrnaseq_reference("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-scrna-seq-reference'>scRNA-seq Reference</a>")
file_spatial_dataset("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-raw-ist-dataset'>Raw iST Dataset</a>")
comp_control_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-control-method'>Control Method</a>"/]
comp_metric[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-metric'>Metric</a>"/]
comp_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-method'>Method</a>"/]
file_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-predicted-data'>Predicted data</a>")
file_score("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-score'>Score</a>")
file_common_scrnaseq("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-sc-dataset'>Common SC Dataset</a>")
file_common_ist---comp_data_processor
comp_data_processor-->file_scrnaseq_reference
comp_data_processor-->file_spatial_dataset
file_scrnaseq_reference---comp_control_method
file_scrnaseq_reference---comp_metric
file_spatial_dataset---comp_control_method
file_spatial_dataset---comp_method
comp_control_method-->file_prediction
comp_metric-->file_score
comp_method-->file_prediction
file_prediction---comp_metric
file_common_scrnaseq---comp_data_processor
Loading

File format: Common iST Dataset

An unprocessed spatial imaging dataset stored as a zarr file.

Example file: resources_test/common/2023_10x_mouse_brain_xenium_rep1/dataset.zarr

Description:

This dataset contains raw images, labels, points, shapes, and tables as output by a dataset loader.

Format:

Data structure:

Component type: Data processor

A data processor.

Arguments:

NameTypeDescription
--input_spfileAn unprocessed spatial imaging dataset stored as a zarr file.
--input_scfileAn unprocessed dataset as output by a dataset loader.
--output_spatial_datasetfile(Output) A spatial transcriptomics dataset, preprocessed for this benchmark.
--output_scrnaseqfile(Output) A single-cell reference dataset, preprocessed for this benchmark.

File format: scRNA-seq Reference

A single-cell reference dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/scrnaseq_reference.h5ad

Description:

This dataset contains preprocessed counts and metadata for single-cell RNA-seq data.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

File format: Raw iST Dataset

A spatial transcriptomics dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_dataset.zarr

Description:

This dataset contains preprocessed images, labels, points, shapes, and tables for spatial transcriptomics data.

Format:

Data structure:

Component type: Control Method

Quality control methods for verifying the pipeline.

Arguments:

NameTypeDescription
--inputfileA spatial transcriptomics dataset, preprocessed for this benchmark.
--input_scrnaseq_referencefileA single-cell reference dataset, preprocessed for this benchmark.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Metric

A task template metric.

Arguments:

NameTypeDescription
--input_predictionfileA predicted dataset as output by a method.
--input_scrnaseq_referencefileA single-cell reference dataset, preprocessed for this benchmark.
--outputfile(Output) File indicating the score of a metric.

Component type: Method

A method.

Arguments:

NameTypeDescription
--inputfileA spatial transcriptomics dataset, preprocessed for this benchmark.
--outputfile(Output) A predicted dataset as output by a method.

File format: Predicted data

A predicted dataset as output by a method.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/prediction.h5ad

Format:

Data structure:

File format: Score

File indicating the score of a metric.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/score.h5ad

Format:

AnnData object
uns: 'dataset_id', 'normalization_id', 'method_id', 'metric_ids', 'metric_values'

Data structure:

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["normalization_id"]stringWhich normalization was used.
uns["method_id"]stringA unique identifier for the method.
uns["metric_ids"]stringOne or more unique metric identifiers.
uns["metric_values"]doubleThe metric values obtained for the given prediction. Must be of same length as ‘metric_ids’.

File format: Common SC Dataset

An unprocessed dataset as output by a dataset loader.

Example file: resources_test/common/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

Description:

This dataset contains raw counts and metadata as output by a dataset loader.

The format of this file is mainly derived from the CELLxGENE schema v4.0.0.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

About

No description, website, or topics provided.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Spatial Segmentation

A one sentence summary of purpose and methodology. Used for creating an overview tables.

Repository: openproblems-bio/task_spatial_segmentation

Description

Provide a clear and concise description of your task, detailing the specific problem it aims to solve. Outline the input data types, the expected output, and any assumptions or constraints. Be sure to explain any terminology or concepts that are essential for understanding the task.

Explain the motivation behind your proposed task. Describe the biological or computational problem you aim to address and why it’s important. Discuss the current state of research in this area and any gaps or challenges that your task could help address. This section should convince readers of the significance and relevance of your task.

Authors & contributors

nameroles
John Doeauthor, maintainer

API

flowchart TB
file_common_ist("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-ist-dataset'>Common iST Dataset</a>")
comp_data_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-data-processor'>Data processor</a>"/]
file_scrnaseq_reference("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-scrna-seq-reference'>scRNA-seq Reference</a>")
file_spatial_dataset("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-raw-ist-dataset'>Raw iST Dataset</a>")
comp_control_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-control-method'>Control Method</a>"/]
comp_metric[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-metric'>Metric</a>"/]
comp_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-method'>Method</a>"/]
file_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-predicted-data'>Predicted data</a>")
file_score("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-score'>Score</a>")
file_common_scrnaseq("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-sc-dataset'>Common SC Dataset</a>")
file_common_ist---comp_data_processor
comp_data_processor-->file_scrnaseq_reference
comp_data_processor-->file_spatial_dataset
file_scrnaseq_reference---comp_control_method
file_scrnaseq_reference---comp_metric
file_spatial_dataset---comp_control_method
file_spatial_dataset---comp_method
comp_control_method-->file_prediction
comp_metric-->file_score
comp_method-->file_prediction
file_prediction---comp_metric
file_common_scrnaseq---comp_data_processor
Loading

File format: Common iST Dataset

An unprocessed spatial imaging dataset stored as a zarr file.

Example file: resources_test/common/2023_10x_mouse_brain_xenium_rep1/dataset.zarr

Description:

This dataset contains raw images, labels, points, shapes, and tables as output by a dataset loader.

Format:

Data structure:

Component type: Data processor

A data processor.

Arguments:

NameTypeDescription
--input_spfileAn unprocessed spatial imaging dataset stored as a zarr file.
--input_scfileAn unprocessed dataset as output by a dataset loader.
--output_spatial_datasetfile(Output) A spatial transcriptomics dataset, preprocessed for this benchmark.
--output_scrnaseqfile(Output) A single-cell reference dataset, preprocessed for this benchmark.

File format: scRNA-seq Reference

A single-cell reference dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/scrnaseq_reference.h5ad

Description:

This dataset contains preprocessed counts and metadata for single-cell RNA-seq data.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

File format: Raw iST Dataset

A spatial transcriptomics dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_dataset.zarr

Description:

This dataset contains preprocessed images, labels, points, shapes, and tables for spatial transcriptomics data.

Format:

Data structure:

Component type: Control Method

Quality control methods for verifying the pipeline.

Arguments:

NameTypeDescription
--inputfileA spatial transcriptomics dataset, preprocessed for this benchmark.
--input_scrnaseq_referencefileA single-cell reference dataset, preprocessed for this benchmark.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Metric

A task template metric.

Arguments:

NameTypeDescription
--input_predictionfileA predicted dataset as output by a method.
--input_scrnaseq_referencefileA single-cell reference dataset, preprocessed for this benchmark.
--outputfile(Output) File indicating the score of a metric.

Component type: Method

A method.

Arguments:

NameTypeDescription
--inputfileA spatial transcriptomics dataset, preprocessed for this benchmark.
--outputfile(Output) A predicted dataset as output by a method.

File format: Predicted data

A predicted dataset as output by a method.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/prediction.h5ad

Format:

Data structure:

File format: Score

File indicating the score of a metric.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/score.h5ad

Format:

AnnData object
uns: 'dataset_id', 'normalization_id', 'method_id', 'metric_ids', 'metric_values'

Data structure:

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["normalization_id"]stringWhich normalization was used.
uns["method_id"]stringA unique identifier for the method.
uns["metric_ids"]stringOne or more unique metric identifiers.
uns["metric_values"]doubleThe metric values obtained for the given prediction. Must be of same length as ‘metric_ids’.

File format: Common SC Dataset

An unprocessed dataset as output by a dataset loader.

Example file: resources_test/common/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

Description:

This dataset contains raw counts and metadata as output by a dataset loader.

The format of this file is mainly derived from the CELLxGENE schema v4.0.0.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

About

No description, website, or topics provided.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Spatial Segmentation

A one sentence summary of purpose and methodology. Used for creating an overview tables.

Repository: openproblems-bio/task_spatial_segmentation

Description

Provide a clear and concise description of your task, detailing the specific problem it aims to solve. Outline the input data types, the expected output, and any assumptions or constraints. Be sure to explain any terminology or concepts that are essential for understanding the task.

Explain the motivation behind your proposed task. Describe the biological or computational problem you aim to address and why it’s important. Discuss the current state of research in this area and any gaps or challenges that your task could help address. This section should convince readers of the significance and relevance of your task.

Authors & contributors

nameroles
John Doeauthor, maintainer

API

flowchart TB
file_common_ist("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-ist-dataset'>Common iST Dataset</a>")
comp_data_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-data-processor'>Data processor</a>"/]
file_scrnaseq_reference("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-scrna-seq-reference'>scRNA-seq Reference</a>")
file_spatial_dataset("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-raw-ist-dataset'>Raw iST Dataset</a>")
comp_control_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-control-method'>Control Method</a>"/]
comp_metric[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-metric'>Metric</a>"/]
comp_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-method'>Method</a>"/]
file_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-predicted-data'>Predicted data</a>")
file_score("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-score'>Score</a>")
file_common_scrnaseq("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-sc-dataset'>Common SC Dataset</a>")
file_common_ist---comp_data_processor
comp_data_processor-->file_scrnaseq_reference
comp_data_processor-->file_spatial_dataset
file_scrnaseq_reference---comp_control_method
file_scrnaseq_reference---comp_metric
file_spatial_dataset---comp_control_method
file_spatial_dataset---comp_method
comp_control_method-->file_prediction
comp_metric-->file_score
comp_method-->file_prediction
file_prediction---comp_metric
file_common_scrnaseq---comp_data_processor
Loading

File format: Common iST Dataset

An unprocessed spatial imaging dataset stored as a zarr file.

Example file: resources_test/common/2023_10x_mouse_brain_xenium_rep1/dataset.zarr

Description:

This dataset contains raw images, labels, points, shapes, and tables as output by a dataset loader.

Format:

Data structure:

Component type: Data processor

A data processor.

Arguments:

NameTypeDescription
--input_spfileAn unprocessed spatial imaging dataset stored as a zarr file.
--input_scfileAn unprocessed dataset as output by a dataset loader.
--output_spatial_datasetfile(Output) A spatial transcriptomics dataset, preprocessed for this benchmark.
--output_scrnaseqfile(Output) A single-cell reference dataset, preprocessed for this benchmark.

File format: scRNA-seq Reference

A single-cell reference dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/scrnaseq_reference.h5ad

Description:

This dataset contains preprocessed counts and metadata for single-cell RNA-seq data.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

File format: Raw iST Dataset

A spatial transcriptomics dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_dataset.zarr

Description:

This dataset contains preprocessed images, labels, points, shapes, and tables for spatial transcriptomics data.

Format:

Data structure:

Component type: Control Method

Quality control methods for verifying the pipeline.

Arguments:

NameTypeDescription
--inputfileA spatial transcriptomics dataset, preprocessed for this benchmark.
--input_scrnaseq_referencefileA single-cell reference dataset, preprocessed for this benchmark.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Metric

A task template metric.

Arguments:

NameTypeDescription
--input_predictionfileA predicted dataset as output by a method.
--input_scrnaseq_referencefileA single-cell reference dataset, preprocessed for this benchmark.
--outputfile(Output) File indicating the score of a metric.

Component type: Method

A method.

Arguments:

NameTypeDescription
--inputfileA spatial transcriptomics dataset, preprocessed for this benchmark.
--outputfile(Output) A predicted dataset as output by a method.

File format: Predicted data

A predicted dataset as output by a method.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/prediction.h5ad

Format:

Data structure:

File format: Score

File indicating the score of a metric.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/score.h5ad

Format:

AnnData object
uns: 'dataset_id', 'normalization_id', 'method_id', 'metric_ids', 'metric_values'

Data structure:

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["normalization_id"]stringWhich normalization was used.
uns["method_id"]stringA unique identifier for the method.
uns["metric_ids"]stringOne or more unique metric identifiers.
uns["metric_values"]doubleThe metric values obtained for the given prediction. Must be of same length as ‘metric_ids’.

File format: Common SC Dataset

An unprocessed dataset as output by a dataset loader.

Example file: resources_test/common/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

Description:

This dataset contains raw counts and metadata as output by a dataset loader.

The format of this file is mainly derived from the CELLxGENE schema v4.0.0.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

About

No description, website, or topics provided.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Spatial Segmentation

A one sentence summary of purpose and methodology. Used for creating an overview tables.

Repository: openproblems-bio/task_spatial_segmentation

Description

Provide a clear and concise description of your task, detailing the specific problem it aims to solve. Outline the input data types, the expected output, and any assumptions or constraints. Be sure to explain any terminology or concepts that are essential for understanding the task.

Explain the motivation behind your proposed task. Describe the biological or computational problem you aim to address and why it’s important. Discuss the current state of research in this area and any gaps or challenges that your task could help address. This section should convince readers of the significance and relevance of your task.

Authors & contributors

nameroles
John Doeauthor, maintainer

API

flowchart TB
file_common_ist("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-ist-dataset'>Common iST Dataset</a>")
comp_data_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-data-processor'>Data processor</a>"/]
file_scrnaseq_reference("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-scrna-seq-reference'>scRNA-seq Reference</a>")
file_spatial_dataset("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-raw-ist-dataset'>Raw iST Dataset</a>")
comp_control_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-control-method'>Control Method</a>"/]
comp_metric[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-metric'>Metric</a>"/]
comp_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-method'>Method</a>"/]
file_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-predicted-data'>Predicted data</a>")
file_score("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-score'>Score</a>")
file_common_scrnaseq("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-sc-dataset'>Common SC Dataset</a>")
file_common_ist---comp_data_processor
comp_data_processor-->file_scrnaseq_reference
comp_data_processor-->file_spatial_dataset
file_scrnaseq_reference---comp_control_method
file_scrnaseq_reference---comp_metric
file_spatial_dataset---comp_control_method
file_spatial_dataset---comp_method
comp_control_method-->file_prediction
comp_metric-->file_score
comp_method-->file_prediction
file_prediction---comp_metric
file_common_scrnaseq---comp_data_processor
Loading

File format: Common iST Dataset

An unprocessed spatial imaging dataset stored as a zarr file.

Example file: resources_test/common/2023_10x_mouse_brain_xenium_rep1/dataset.zarr

Description:

This dataset contains raw images, labels, points, shapes, and tables as output by a dataset loader.

Format:

Data structure:

Component type: Data processor

A data processor.

Arguments:

NameTypeDescription
--input_spfileAn unprocessed spatial imaging dataset stored as a zarr file.
--input_scfileAn unprocessed dataset as output by a dataset loader.
--output_spatial_datasetfile(Output) A spatial transcriptomics dataset, preprocessed for this benchmark.
--output_scrnaseqfile(Output) A single-cell reference dataset, preprocessed for this benchmark.

File format: scRNA-seq Reference

A single-cell reference dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/scrnaseq_reference.h5ad

Description:

This dataset contains preprocessed counts and metadata for single-cell RNA-seq data.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

File format: Raw iST Dataset

A spatial transcriptomics dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_dataset.zarr

Description:

This dataset contains preprocessed images, labels, points, shapes, and tables for spatial transcriptomics data.

Format:

Data structure:

Component type: Control Method

Quality control methods for verifying the pipeline.

Arguments:

NameTypeDescription
--inputfileA spatial transcriptomics dataset, preprocessed for this benchmark.
--input_scrnaseq_referencefileA single-cell reference dataset, preprocessed for this benchmark.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Metric

A task template metric.

Arguments:

NameTypeDescription
--input_predictionfileA predicted dataset as output by a method.
--input_scrnaseq_referencefileA single-cell reference dataset, preprocessed for this benchmark.
--outputfile(Output) File indicating the score of a metric.

Component type: Method

A method.

Arguments:

NameTypeDescription
--inputfileA spatial transcriptomics dataset, preprocessed for this benchmark.
--outputfile(Output) A predicted dataset as output by a method.

File format: Predicted data

A predicted dataset as output by a method.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/prediction.h5ad

Format:

Data structure:

File format: Score

File indicating the score of a metric.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/score.h5ad

Format:

AnnData object
uns: 'dataset_id', 'normalization_id', 'method_id', 'metric_ids', 'metric_values'

Data structure:

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["normalization_id"]stringWhich normalization was used.
uns["method_id"]stringA unique identifier for the method.
uns["metric_ids"]stringOne or more unique metric identifiers.
uns["metric_values"]doubleThe metric values obtained for the given prediction. Must be of same length as ‘metric_ids’.

File format: Common SC Dataset

An unprocessed dataset as output by a dataset loader.

Example file: resources_test/common/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

Description:

This dataset contains raw counts and metadata as output by a dataset loader.

The format of this file is mainly derived from the CELLxGENE schema v4.0.0.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

About

No description, website, or topics provided.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Spatial Segmentation

A one sentence summary of purpose and methodology. Used for creating an overview tables.

Repository: openproblems-bio/task_spatial_segmentation

Description

Provide a clear and concise description of your task, detailing the specific problem it aims to solve. Outline the input data types, the expected output, and any assumptions or constraints. Be sure to explain any terminology or concepts that are essential for understanding the task.

Explain the motivation behind your proposed task. Describe the biological or computational problem you aim to address and why it’s important. Discuss the current state of research in this area and any gaps or challenges that your task could help address. This section should convince readers of the significance and relevance of your task.

Authors & contributors

nameroles
John Doeauthor, maintainer

API

flowchart TB
file_common_ist("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-ist-dataset'>Common iST Dataset</a>")
comp_data_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-data-processor'>Data processor</a>"/]
file_scrnaseq_reference("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-scrna-seq-reference'>scRNA-seq Reference</a>")
file_spatial_dataset("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-raw-ist-dataset'>Raw iST Dataset</a>")
comp_control_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-control-method'>Control Method</a>"/]
comp_metric[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-metric'>Metric</a>"/]
comp_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-method'>Method</a>"/]
file_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-predicted-data'>Predicted data</a>")
file_score("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-score'>Score</a>")
file_common_scrnaseq("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-sc-dataset'>Common SC Dataset</a>")
file_common_ist---comp_data_processor
comp_data_processor-->file_scrnaseq_reference
comp_data_processor-->file_spatial_dataset
file_scrnaseq_reference---comp_control_method
file_scrnaseq_reference---comp_metric
file_spatial_dataset---comp_control_method
file_spatial_dataset---comp_method
comp_control_method-->file_prediction
comp_metric-->file_score
comp_method-->file_prediction
file_prediction---comp_metric
file_common_scrnaseq---comp_data_processor
Loading

File format: Common iST Dataset

An unprocessed spatial imaging dataset stored as a zarr file.

Example file: resources_test/common/2023_10x_mouse_brain_xenium_rep1/dataset.zarr

Description:

This dataset contains raw images, labels, points, shapes, and tables as output by a dataset loader.

Format:

Data structure:

Component type: Data processor

A data processor.

Arguments:

NameTypeDescription
--input_spfileAn unprocessed spatial imaging dataset stored as a zarr file.
--input_scfileAn unprocessed dataset as output by a dataset loader.
--output_spatial_datasetfile(Output) A spatial transcriptomics dataset, preprocessed for this benchmark.
--output_scrnaseqfile(Output) A single-cell reference dataset, preprocessed for this benchmark.

File format: scRNA-seq Reference

A single-cell reference dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/scrnaseq_reference.h5ad

Description:

This dataset contains preprocessed counts and metadata for single-cell RNA-seq data.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

File format: Raw iST Dataset

A spatial transcriptomics dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_dataset.zarr

Description:

This dataset contains preprocessed images, labels, points, shapes, and tables for spatial transcriptomics data.

Format:

Data structure:

Component type: Control Method

Quality control methods for verifying the pipeline.

Arguments:

NameTypeDescription
--inputfileA spatial transcriptomics dataset, preprocessed for this benchmark.
--input_scrnaseq_referencefileA single-cell reference dataset, preprocessed for this benchmark.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Metric

A task template metric.

Arguments:

NameTypeDescription
--input_predictionfileA predicted dataset as output by a method.
--input_scrnaseq_referencefileA single-cell reference dataset, preprocessed for this benchmark.
--outputfile(Output) File indicating the score of a metric.

Component type: Method

A method.

Arguments:

NameTypeDescription
--inputfileA spatial transcriptomics dataset, preprocessed for this benchmark.
--outputfile(Output) A predicted dataset as output by a method.

File format: Predicted data

A predicted dataset as output by a method.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/prediction.h5ad

Format:

Data structure:

File format: Score

File indicating the score of a metric.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/score.h5ad

Format:

AnnData object
uns: 'dataset_id', 'normalization_id', 'method_id', 'metric_ids', 'metric_values'

Data structure:

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["normalization_id"]stringWhich normalization was used.
uns["method_id"]stringA unique identifier for the method.
uns["metric_ids"]stringOne or more unique metric identifiers.
uns["metric_values"]doubleThe metric values obtained for the given prediction. Must be of same length as ‘metric_ids’.

File format: Common SC Dataset

An unprocessed dataset as output by a dataset loader.

Example file: resources_test/common/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

Description:

This dataset contains raw counts and metadata as output by a dataset loader.

The format of this file is mainly derived from the CELLxGENE schema v4.0.0.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

About

No description, website, or topics provided.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Spatial Segmentation

A one sentence summary of purpose and methodology. Used for creating an overview tables.

Repository: openproblems-bio/task_spatial_segmentation

Description

Provide a clear and concise description of your task, detailing the specific problem it aims to solve. Outline the input data types, the expected output, and any assumptions or constraints. Be sure to explain any terminology or concepts that are essential for understanding the task.

Explain the motivation behind your proposed task. Describe the biological or computational problem you aim to address and why it’s important. Discuss the current state of research in this area and any gaps or challenges that your task could help address. This section should convince readers of the significance and relevance of your task.

Authors & contributors

nameroles
John Doeauthor, maintainer

API

flowchart TB
file_common_ist("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-ist-dataset'>Common iST Dataset</a>")
comp_data_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-data-processor'>Data processor</a>"/]
file_scrnaseq_reference("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-scrna-seq-reference'>scRNA-seq Reference</a>")
file_spatial_dataset("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-raw-ist-dataset'>Raw iST Dataset</a>")
comp_control_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-control-method'>Control Method</a>"/]
comp_metric[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-metric'>Metric</a>"/]
comp_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-method'>Method</a>"/]
file_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-predicted-data'>Predicted data</a>")
file_score("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-score'>Score</a>")
file_common_scrnaseq("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-sc-dataset'>Common SC Dataset</a>")
file_common_ist---comp_data_processor
comp_data_processor-->file_scrnaseq_reference
comp_data_processor-->file_spatial_dataset
file_scrnaseq_reference---comp_control_method
file_scrnaseq_reference---comp_metric
file_spatial_dataset---comp_control_method
file_spatial_dataset---comp_method
comp_control_method-->file_prediction
comp_metric-->file_score
comp_method-->file_prediction
file_prediction---comp_metric
file_common_scrnaseq---comp_data_processor
Loading

File format: Common iST Dataset

An unprocessed spatial imaging dataset stored as a zarr file.

Example file: resources_test/common/2023_10x_mouse_brain_xenium_rep1/dataset.zarr

Description:

This dataset contains raw images, labels, points, shapes, and tables as output by a dataset loader.

Format:

Data structure:

Component type: Data processor

A data processor.

Arguments:

NameTypeDescription
--input_spfileAn unprocessed spatial imaging dataset stored as a zarr file.
--input_scfileAn unprocessed dataset as output by a dataset loader.
--output_spatial_datasetfile(Output) A spatial transcriptomics dataset, preprocessed for this benchmark.
--output_scrnaseqfile(Output) A single-cell reference dataset, preprocessed for this benchmark.

File format: scRNA-seq Reference

A single-cell reference dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/scrnaseq_reference.h5ad

Description:

This dataset contains preprocessed counts and metadata for single-cell RNA-seq data.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

File format: Raw iST Dataset

A spatial transcriptomics dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_dataset.zarr

Description:

This dataset contains preprocessed images, labels, points, shapes, and tables for spatial transcriptomics data.

Format:

Data structure:

Component type: Control Method

Quality control methods for verifying the pipeline.

Arguments:

NameTypeDescription
--inputfileA spatial transcriptomics dataset, preprocessed for this benchmark.
--input_scrnaseq_referencefileA single-cell reference dataset, preprocessed for this benchmark.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Metric

A task template metric.

Arguments:

NameTypeDescription
--input_predictionfileA predicted dataset as output by a method.
--input_scrnaseq_referencefileA single-cell reference dataset, preprocessed for this benchmark.
--outputfile(Output) File indicating the score of a metric.

Component type: Method

A method.

Arguments:

NameTypeDescription
--inputfileA spatial transcriptomics dataset, preprocessed for this benchmark.
--outputfile(Output) A predicted dataset as output by a method.

File format: Predicted data

A predicted dataset as output by a method.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/prediction.h5ad

Format:

Data structure:

File format: Score

File indicating the score of a metric.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/score.h5ad

Format:

AnnData object
uns: 'dataset_id', 'normalization_id', 'method_id', 'metric_ids', 'metric_values'

Data structure:

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["normalization_id"]stringWhich normalization was used.
uns["method_id"]stringA unique identifier for the method.
uns["metric_ids"]stringOne or more unique metric identifiers.
uns["metric_values"]doubleThe metric values obtained for the given prediction. Must be of same length as ‘metric_ids’.

File format: Common SC Dataset

An unprocessed dataset as output by a dataset loader.

Example file: resources_test/common/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

Description:

This dataset contains raw counts and metadata as output by a dataset loader.

The format of this file is mainly derived from the CELLxGENE schema v4.0.0.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

About

No description, website, or topics provided.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Spatial Segmentation

A one sentence summary of purpose and methodology. Used for creating an overview tables.

Repository: openproblems-bio/task_spatial_segmentation

Description

Provide a clear and concise description of your task, detailing the specific problem it aims to solve. Outline the input data types, the expected output, and any assumptions or constraints. Be sure to explain any terminology or concepts that are essential for understanding the task.

Explain the motivation behind your proposed task. Describe the biological or computational problem you aim to address and why it’s important. Discuss the current state of research in this area and any gaps or challenges that your task could help address. This section should convince readers of the significance and relevance of your task.

Authors & contributors

nameroles
John Doeauthor, maintainer

API

flowchart TB
file_common_ist("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-ist-dataset'>Common iST Dataset</a>")
comp_data_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-data-processor'>Data processor</a>"/]
file_scrnaseq_reference("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-scrna-seq-reference'>scRNA-seq Reference</a>")
file_spatial_dataset("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-raw-ist-dataset'>Raw iST Dataset</a>")
comp_control_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-control-method'>Control Method</a>"/]
comp_metric[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-metric'>Metric</a>"/]
comp_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-method'>Method</a>"/]
file_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-predicted-data'>Predicted data</a>")
file_score("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-score'>Score</a>")
file_common_scrnaseq("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-sc-dataset'>Common SC Dataset</a>")
file_common_ist---comp_data_processor
comp_data_processor-->file_scrnaseq_reference
comp_data_processor-->file_spatial_dataset
file_scrnaseq_reference---comp_control_method
file_scrnaseq_reference---comp_metric
file_spatial_dataset---comp_control_method
file_spatial_dataset---comp_method
comp_control_method-->file_prediction
comp_metric-->file_score
comp_method-->file_prediction
file_prediction---comp_metric
file_common_scrnaseq---comp_data_processor
Loading

File format: Common iST Dataset

An unprocessed spatial imaging dataset stored as a zarr file.

Example file: resources_test/common/2023_10x_mouse_brain_xenium_rep1/dataset.zarr

Description:

This dataset contains raw images, labels, points, shapes, and tables as output by a dataset loader.

Format:

Data structure:

Component type: Data processor

A data processor.

Arguments:

NameTypeDescription
--input_spfileAn unprocessed spatial imaging dataset stored as a zarr file.
--input_scfileAn unprocessed dataset as output by a dataset loader.
--output_spatial_datasetfile(Output) A spatial transcriptomics dataset, preprocessed for this benchmark.
--output_scrnaseqfile(Output) A single-cell reference dataset, preprocessed for this benchmark.

File format: scRNA-seq Reference

A single-cell reference dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/scrnaseq_reference.h5ad

Description:

This dataset contains preprocessed counts and metadata for single-cell RNA-seq data.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

File format: Raw iST Dataset

A spatial transcriptomics dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_dataset.zarr

Description:

This dataset contains preprocessed images, labels, points, shapes, and tables for spatial transcriptomics data.

Format:

Data structure:

Component type: Control Method

Quality control methods for verifying the pipeline.

Arguments:

NameTypeDescription
--inputfileA spatial transcriptomics dataset, preprocessed for this benchmark.
--input_scrnaseq_referencefileA single-cell reference dataset, preprocessed for this benchmark.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Metric

A task template metric.

Arguments:

NameTypeDescription
--input_predictionfileA predicted dataset as output by a method.
--input_scrnaseq_referencefileA single-cell reference dataset, preprocessed for this benchmark.
--outputfile(Output) File indicating the score of a metric.

Component type: Method

A method.

Arguments:

NameTypeDescription
--inputfileA spatial transcriptomics dataset, preprocessed for this benchmark.
--outputfile(Output) A predicted dataset as output by a method.

File format: Predicted data

A predicted dataset as output by a method.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/prediction.h5ad

Format:

Data structure:

File format: Score

File indicating the score of a metric.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/score.h5ad

Format:

AnnData object
uns: 'dataset_id', 'normalization_id', 'method_id', 'metric_ids', 'metric_values'

Data structure:

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["normalization_id"]stringWhich normalization was used.
uns["method_id"]stringA unique identifier for the method.
uns["metric_ids"]stringOne or more unique metric identifiers.
uns["metric_values"]doubleThe metric values obtained for the given prediction. Must be of same length as ‘metric_ids’.

File format: Common SC Dataset

An unprocessed dataset as output by a dataset loader.

Example file: resources_test/common/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

Description:

This dataset contains raw counts and metadata as output by a dataset loader.

The format of this file is mainly derived from the CELLxGENE schema v4.0.0.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

About

No description, website, or topics provided.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Spatial Segmentation

A one sentence summary of purpose and methodology. Used for creating an overview tables.

Repository: openproblems-bio/task_spatial_segmentation

Description

Provide a clear and concise description of your task, detailing the specific problem it aims to solve. Outline the input data types, the expected output, and any assumptions or constraints. Be sure to explain any terminology or concepts that are essential for understanding the task.

Explain the motivation behind your proposed task. Describe the biological or computational problem you aim to address and why it’s important. Discuss the current state of research in this area and any gaps or challenges that your task could help address. This section should convince readers of the significance and relevance of your task.

Authors & contributors

nameroles
John Doeauthor, maintainer

API

flowchart TB
file_common_ist("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-ist-dataset'>Common iST Dataset</a>")
comp_data_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-data-processor'>Data processor</a>"/]
file_scrnaseq_reference("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-scrna-seq-reference'>scRNA-seq Reference</a>")
file_spatial_dataset("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-raw-ist-dataset'>Raw iST Dataset</a>")
comp_control_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-control-method'>Control Method</a>"/]
comp_metric[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-metric'>Metric</a>"/]
comp_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-method'>Method</a>"/]
file_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-predicted-data'>Predicted data</a>")
file_score("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-score'>Score</a>")
file_common_scrnaseq("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-sc-dataset'>Common SC Dataset</a>")
file_common_ist---comp_data_processor
comp_data_processor-->file_scrnaseq_reference
comp_data_processor-->file_spatial_dataset
file_scrnaseq_reference---comp_control_method
file_scrnaseq_reference---comp_metric
file_spatial_dataset---comp_control_method
file_spatial_dataset---comp_method
comp_control_method-->file_prediction
comp_metric-->file_score
comp_method-->file_prediction
file_prediction---comp_metric
file_common_scrnaseq---comp_data_processor
Loading

File format: Common iST Dataset

An unprocessed spatial imaging dataset stored as a zarr file.

Example file: resources_test/common/2023_10x_mouse_brain_xenium_rep1/dataset.zarr

Description:

This dataset contains raw images, labels, points, shapes, and tables as output by a dataset loader.

Format:

Data structure:

Component type: Data processor

A data processor.

Arguments:

NameTypeDescription
--input_spfileAn unprocessed spatial imaging dataset stored as a zarr file.
--input_scfileAn unprocessed dataset as output by a dataset loader.
--output_spatial_datasetfile(Output) A spatial transcriptomics dataset, preprocessed for this benchmark.
--output_scrnaseqfile(Output) A single-cell reference dataset, preprocessed for this benchmark.

File format: scRNA-seq Reference

A single-cell reference dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/scrnaseq_reference.h5ad

Description:

This dataset contains preprocessed counts and metadata for single-cell RNA-seq data.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

File format: Raw iST Dataset

A spatial transcriptomics dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_dataset.zarr

Description:

This dataset contains preprocessed images, labels, points, shapes, and tables for spatial transcriptomics data.

Format:

Data structure:

Component type: Control Method

Quality control methods for verifying the pipeline.

Arguments:

NameTypeDescription
--inputfileA spatial transcriptomics dataset, preprocessed for this benchmark.
--input_scrnaseq_referencefileA single-cell reference dataset, preprocessed for this benchmark.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Metric

A task template metric.

Arguments:

NameTypeDescription
--input_predictionfileA predicted dataset as output by a method.
--input_scrnaseq_referencefileA single-cell reference dataset, preprocessed for this benchmark.
--outputfile(Output) File indicating the score of a metric.

Component type: Method

A method.

Arguments:

NameTypeDescription
--inputfileA spatial transcriptomics dataset, preprocessed for this benchmark.
--outputfile(Output) A predicted dataset as output by a method.

File format: Predicted data

A predicted dataset as output by a method.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/prediction.h5ad

Format:

Data structure:

File format: Score

File indicating the score of a metric.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/score.h5ad

Format:

AnnData object
uns: 'dataset_id', 'normalization_id', 'method_id', 'metric_ids', 'metric_values'

Data structure:

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["normalization_id"]stringWhich normalization was used.
uns["method_id"]stringA unique identifier for the method.
uns["metric_ids"]stringOne or more unique metric identifiers.
uns["metric_values"]doubleThe metric values obtained for the given prediction. Must be of same length as ‘metric_ids’.

File format: Common SC Dataset

An unprocessed dataset as output by a dataset loader.

Example file: resources_test/common/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

Description:

This dataset contains raw counts and metadata as output by a dataset loader.

The format of this file is mainly derived from the CELLxGENE schema v4.0.0.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

About

No description, website, or topics provided.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages