Repository files navigation

Spatial Segmentation

A one sentence summary of purpose and methodology. Used for creating an overview tables.

Repository: openproblems-bio/task_spatial_segmentation

Description

Provide a clear and concise description of your task, detailing the specific problem it aims to solve. Outline the input data types, the expected output, and any assumptions or constraints. Be sure to explain any terminology or concepts that are essential for understanding the task.

Explain the motivation behind your proposed task. Describe the biological or computational problem you aim to address and why it’s important. Discuss the current state of research in this area and any gaps or challenges that your task could help address. This section should convince readers of the significance and relevance of your task.

Authors & contributors

NameRolesOrcidGithub
Daria Romanovskaiamaintainer, author0000-0003-2831-0919dariarom94
Florian Heylmaintainer, author0000-0002-3651-5685heylf
Robrecht Cannoodtauthor0000-0003-3641-729Xrcannood

API

flowchart TB
file_common_ist("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-ist-dataset'>Common iST Dataset</a>")
comp_data_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-data-processor'>Data processor</a>"/]
file_spatial_unlabelled("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-unlabelled'>Unlabelled</a>")
file_spatial_solution("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-solution'>Solution</a>")
file_scrnaseq_reference("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-scrna-seq-reference'>scRNA-seq Reference</a>")
comp_control_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-control-method'>Control Method</a>"/]
comp_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-method'>Method</a>"/]
comp_output_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-output-processor'>Output processor</a>"/]
comp_metric[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-metric'>Metric</a>"/]
file_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-predicted-data'>Predicted data</a>")
file_processed_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-processed-prediction'>Processed prediction</a>")
file_score("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-score'>Score</a>")
file_common_scrnaseq("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-sc-dataset'>Common SC Dataset</a>")
file_common_ist---comp_data_processor
comp_data_processor-->file_spatial_unlabelled
comp_data_processor-->file_spatial_solution
comp_data_processor-->file_scrnaseq_reference
file_spatial_unlabelled---comp_control_method
file_spatial_unlabelled---comp_method
file_spatial_unlabelled---comp_output_processor
file_spatial_solution---comp_control_method
file_spatial_solution---comp_metric
comp_control_method-->file_prediction
comp_method-->file_prediction
comp_output_processor-->file_processed_prediction
comp_metric-->file_score
file_prediction---comp_output_processor
file_processed_prediction---comp_metric
file_common_scrnaseq---comp_data_processor
Loading

File format: Common iST Dataset

An unprocessed spatial imaging dataset stored as a zarr file.

Example file: resources_test/common/2023_10x_mouse_brain_xenium_rep1/dataset.zarr

Description:

This dataset contains raw images, labels, points, shapes, and tables as output by a dataset loader.

Format:

SpatialData object
images: 'image', 'image_3D', 'he_image'
labels: 'cell_labels', 'nucleus_labels', 'groundtruth_cell_labels'
points: 'transcripts'
shapes: 'cell_boundaries', 'nucleus_boundaries'
tables: 'table'
coordinate_systems: 'global'

Data structure:

images

NameDescription
imageThe raw image data.
image_3D(Optional) The raw 3D image data.
he_image(Optional) H&E image data.

labels

NameDescription
cell_labelsVendor-provided cell segmentation labels.
nucleus_labelsVendor-provided nucleus segmentation labels.
groundtruth_cell_labels(Optional) Manually annotated cell segmentation labels used as ground truth for evaluation.

points

transcripts: Point cloud data of transcripts.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
cell_idinteger(Optional) Unique identifier of the cell.
nucleus_idinteger(Optional) Unique identifier of the nucleus.
cell_typestring(Optional) Cell type of the cell.
qvfloat(Optional) Quality value of the point.
transcript_idlongUnique identifier of the transcript.
overlaps_nucleusboolean(Optional) Whether the point overlaps with a nucleus.

shapes

cell_boundaries: Cell boundaries.

ColumnTypeDescription
geometryobjectGeometry of the cell boundary.

nucleus_boundaries: Nucleus boundaries.

ColumnTypeDescription
geometryobjectGeometry of the nucleus boundary.

tables

table: Metadata of spatial dataset.

SlotTypeDescription
obs["cell_id"]stringA unique identifier for the cell.
obs["groundtruth_cell_type"]string(Optional) Manually curated cell type annotations which serves as ground truth for evaluations.
var["gene_ids"]stringUnique identifier for the gene.
var["feature_types"]stringType of the feature.
obsm["spatial"]doubleSpatial coordinates of the cell.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]stringLink to the original source of the dataset.
uns["dataset_reference"]stringBibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]stringThe organism of the sample in the dataset.
uns["segmentation_id"]stringA unique identifier for the segmentation.

coordinate_systems

NameDescription
globalCoordinate system of the replicate.

Component type: Data processor

A data processor.

Arguments:

NameTypeDescription
--input_spfileAn unprocessed spatial imaging dataset stored as a zarr file.
--input_scfileAn unprocessed dataset as output by a dataset loader.
--output_spatial_unlabelledfile(Output) Preprocessed spatial transcriptomics data without segmentation labels for method input.
--output_spatial_solutionfile(Output) Ground truth segmentation labels and cell assignments for method evaluation.
--output_scrnaseq_referencefile(Output) A single-cell reference dataset, preprocessed for this benchmark.

File format: Unlabelled

Preprocessed spatial transcriptomics data without segmentation labels for method input.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_unlabelled.zarr

Description:

This dataset contains preprocessed images and transcript point clouds for spatial transcriptomics data. Ground truth segmentation labels are intentionally excluded to prevent methods from cheating.

Format:

SpatialData object
images: 'image'
labels: 'cell_labels', 'nucleus_labels'
points: 'transcripts'
tables: 'table'
coordinate_systems: 'global'

Data structure:

images

NameDescription
imageThe raw image data.

labels

NameDescription
cell_labels(Optional) Vendor-provided cell segmentation labels, exposed as a segmentation prior.
nucleus_labels(Optional) Vendor-provided nucleus segmentation labels, exposed as a segmentation prior.

points

transcripts: Point cloud data of transcripts.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
qvfloat(Optional) Quality value of the point.
transcript_idlongUnique identifier of the transcript.
overlaps_nucleusboolean(Optional) Whether the point overlaps with the nucleus (derived from morphology).
cell_idinteger(Optional) Vendor-provided cell assignment from the raw data, exposed as a segmentation prior. This is NOT the ground truth used for evaluation (which is held out in spatial_solution); methods may freely condition on it.

tables

table: Metadata of spatial dataset.

SlotTypeDescription
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]stringLink to the original source of the dataset.
uns["dataset_reference"]stringBibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]stringThe organism of the sample in the dataset.
uns["orig_dataset_id"]stringThe identifier of the original dataset from which this dataset was derived (if applicable).

coordinate_systems

NameDescription
globalCoordinate system of the replicate.

File format: Solution

Ground truth segmentation labels and cell assignments for method evaluation.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_solution.zarr

Description:

This dataset contains the ground truth cell and nucleus segmentation labels, cell boundaries, and a reference table matching each cell to its label region.

Format:

SpatialData object
labels: 'cell_labels', 'nucleus_labels', 'groundtruth_cell_labels'
points: 'transcripts'
shapes: 'cell_boundaries', 'nucleus_boundaries'
tables: 'table'

Data structure:

labels

NameDescription
cell_labelsVendor-provided cell segmentation labels.
nucleus_labelsVendor-provided nucleus segmentation labels.
groundtruth_cell_labels(Optional) Manually annotated cell segmentation labels used as ground truth for evaluation.

points

transcripts: Point cloud data of transcripts with ground truth cell assignments.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
cell_idintegerGround truth cell assignment (0 = background).
transcript_idlongUnique identifier of the transcript.

shapes

cell_boundaries: Ground truth cell boundary shapes.

ColumnTypeDescription
geometryobjectGeometry of the cell boundary.

nucleus_boundaries: Ground truth nucleus boundary shapes.

ColumnTypeDescription
geometryobjectGeometry of the nucleus boundary.

tables

table: Reference cell metadata table.

SlotTypeDescription
obs["cell_id"]integerUnique cell identifier, matching instance IDs in the label images.
obs["region"]stringName of the label image this cell belongs to (e.g. ‘cell_labels’).
obs["cell_area"]double(Optional) Area of the cell in pixels.
obs["transcript_counts"]integer(Optional) Total number of transcripts assigned to this cell.
obs["groundtruth_cell_type"]string(Optional) Manually curated cell type annotations which serves as ground truth for evaluations.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["orig_dataset_id"]stringThe identifier of the original dataset from which this dataset was derived (if applicable).

File format: scRNA-seq Reference

A single-cell reference dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/scrnaseq_reference.h5ad

Description:

This dataset contains preprocessed counts and metadata for single-cell RNA-seq data.

Format:

AnnData object
obs: 'cell_type'
var: 'feature_id', 'feature_name', 'hvg'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized', 'normalized_log', 'normalized_log_scaled'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]doubleNormalized expression values.
layers["normalized_log"]doubleLog1p normalized expression values.
layers["normalized_log_scaled"]doubleLog1p normalized expression values scaled to unit variance and zero mean.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

Component type: Control Method

Quality control methods for verifying the pipeline.

Arguments:

NameTypeDescription
--inputfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--input_solutionfileGround truth segmentation labels and cell assignments for method evaluation.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Method

A method.

Arguments:

NameTypeDescription
--inputfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Output processor

An output processor for the prediction.

Arguments:

NameTypeDescription
--input_predictionfileA predicted dataset as output by a method.
--input_spatial_unlabelledfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--outputfile(Output) A processed predicted dataset, ready to be used as input for the evaluation.

Component type: Metric

A task template metric.

Arguments:

NameTypeDescription
--input_predictionfileA processed predicted dataset, ready to be used as input for the evaluation.
--input_solutionfileGround truth segmentation labels and cell assignments for method evaluation.
--outputfile(Output) File indicating the score of a metric.

File format: Predicted data

A predicted dataset as output by a method.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/prediction.zarr

Format:

SpatialData object
labels: 'segmentation'
tables: 'table'

Data structure:

labels

NameDescription
segmentationSegmentation of the data.

tables

table: AnnData table.

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["method_id"]stringA unique identifier for the method.

File format: Processed prediction

A processed predicted dataset, ready to be used as input for the evaluation.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/processed_prediction.zarr

Format:

SpatialData object
labels: 'segmentation'
tables: 'table'

Data structure:

labels

NameDescription
segmentationSegmentation of the data.

tables

table: AnnData table.

SlotTypeDescription
obs["cell_id"]stringCell ID.
obs["region"]stringRegion.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
layers["counts"]integerRaw counts.
layers["normalized"]doubleNormalized expression values.
layers["normalized_log"]doubleLog1p normalized expression values.
layers["normalized_log_scaled"]doubleLog1p normalized expression values scaled to unit variance and zero mean.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["method_id"]stringA unique identifier for the method.

File format: Score

File indicating the score of a metric.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/score.h5ad

Format:

AnnData object
uns: 'dataset_id', 'normalization_id', 'method_id', 'metric_ids', 'metric_values'

Data structure:

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["normalization_id"]stringWhich normalization was used.
uns["method_id"]stringA unique identifier for the method.
uns["metric_ids"]stringOne or more unique metric identifiers.
uns["metric_values"]doubleThe metric values obtained for the given prediction. Must be of same length as ‘metric_ids’.

File format: Common SC Dataset

An unprocessed dataset as output by a dataset loader.

Example file: resources_test/common/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

Description:

This dataset contains raw counts and metadata as output by a dataset loader.

The format of this file is mainly derived from the CELLxGENE schema v4.0.0.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

About

No description, website, or topics provided.

Resources

Contributing

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Spatial Segmentation

A one sentence summary of purpose and methodology. Used for creating an overview tables.

Repository: openproblems-bio/task_spatial_segmentation

Description

Provide a clear and concise description of your task, detailing the specific problem it aims to solve. Outline the input data types, the expected output, and any assumptions or constraints. Be sure to explain any terminology or concepts that are essential for understanding the task.

Explain the motivation behind your proposed task. Describe the biological or computational problem you aim to address and why it’s important. Discuss the current state of research in this area and any gaps or challenges that your task could help address. This section should convince readers of the significance and relevance of your task.

Authors & contributors

NameRolesOrcidGithub
Daria Romanovskaiamaintainer, author0000-0003-2831-0919dariarom94
Florian Heylmaintainer, author0000-0002-3651-5685heylf
Robrecht Cannoodtauthor0000-0003-3641-729Xrcannood

API

flowchart TB
file_common_ist("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-ist-dataset'>Common iST Dataset</a>")
comp_data_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-data-processor'>Data processor</a>"/]
file_spatial_unlabelled("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-unlabelled'>Unlabelled</a>")
file_spatial_solution("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-solution'>Solution</a>")
file_scrnaseq_reference("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-scrna-seq-reference'>scRNA-seq Reference</a>")
comp_control_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-control-method'>Control Method</a>"/]
comp_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-method'>Method</a>"/]
comp_output_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-output-processor'>Output processor</a>"/]
comp_metric[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-metric'>Metric</a>"/]
file_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-predicted-data'>Predicted data</a>")
file_processed_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-processed-prediction'>Processed prediction</a>")
file_score("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-score'>Score</a>")
file_common_scrnaseq("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-sc-dataset'>Common SC Dataset</a>")
file_common_ist---comp_data_processor
comp_data_processor-->file_spatial_unlabelled
comp_data_processor-->file_spatial_solution
comp_data_processor-->file_scrnaseq_reference
file_spatial_unlabelled---comp_control_method
file_spatial_unlabelled---comp_method
file_spatial_unlabelled---comp_output_processor
file_spatial_solution---comp_control_method
file_spatial_solution---comp_metric
comp_control_method-->file_prediction
comp_method-->file_prediction
comp_output_processor-->file_processed_prediction
comp_metric-->file_score
file_prediction---comp_output_processor
file_processed_prediction---comp_metric
file_common_scrnaseq---comp_data_processor
Loading

File format: Common iST Dataset

An unprocessed spatial imaging dataset stored as a zarr file.

Example file: resources_test/common/2023_10x_mouse_brain_xenium_rep1/dataset.zarr

Description:

This dataset contains raw images, labels, points, shapes, and tables as output by a dataset loader.

Format:

SpatialData object
images: 'image', 'image_3D', 'he_image'
labels: 'cell_labels', 'nucleus_labels', 'groundtruth_cell_labels'
points: 'transcripts'
shapes: 'cell_boundaries', 'nucleus_boundaries'
tables: 'table'
coordinate_systems: 'global'

Data structure:

images

NameDescription
imageThe raw image data.
image_3D(Optional) The raw 3D image data.
he_image(Optional) H&E image data.

labels

NameDescription
cell_labelsVendor-provided cell segmentation labels.
nucleus_labelsVendor-provided nucleus segmentation labels.
groundtruth_cell_labels(Optional) Manually annotated cell segmentation labels used as ground truth for evaluation.

points

transcripts: Point cloud data of transcripts.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
cell_idinteger(Optional) Unique identifier of the cell.
nucleus_idinteger(Optional) Unique identifier of the nucleus.
cell_typestring(Optional) Cell type of the cell.
qvfloat(Optional) Quality value of the point.
transcript_idlongUnique identifier of the transcript.
overlaps_nucleusboolean(Optional) Whether the point overlaps with a nucleus.

shapes

cell_boundaries: Cell boundaries.

ColumnTypeDescription
geometryobjectGeometry of the cell boundary.

nucleus_boundaries: Nucleus boundaries.

ColumnTypeDescription
geometryobjectGeometry of the nucleus boundary.

tables

table: Metadata of spatial dataset.

SlotTypeDescription
obs["cell_id"]stringA unique identifier for the cell.
obs["groundtruth_cell_type"]string(Optional) Manually curated cell type annotations which serves as ground truth for evaluations.
var["gene_ids"]stringUnique identifier for the gene.
var["feature_types"]stringType of the feature.
obsm["spatial"]doubleSpatial coordinates of the cell.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]stringLink to the original source of the dataset.
uns["dataset_reference"]stringBibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]stringThe organism of the sample in the dataset.
uns["segmentation_id"]stringA unique identifier for the segmentation.

coordinate_systems

NameDescription
globalCoordinate system of the replicate.

Component type: Data processor

A data processor.

Arguments:

NameTypeDescription
--input_spfileAn unprocessed spatial imaging dataset stored as a zarr file.
--input_scfileAn unprocessed dataset as output by a dataset loader.
--output_spatial_unlabelledfile(Output) Preprocessed spatial transcriptomics data without segmentation labels for method input.
--output_spatial_solutionfile(Output) Ground truth segmentation labels and cell assignments for method evaluation.
--output_scrnaseq_referencefile(Output) A single-cell reference dataset, preprocessed for this benchmark.

File format: Unlabelled

Preprocessed spatial transcriptomics data without segmentation labels for method input.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_unlabelled.zarr

Description:

This dataset contains preprocessed images and transcript point clouds for spatial transcriptomics data. Ground truth segmentation labels are intentionally excluded to prevent methods from cheating.

Format:

SpatialData object
images: 'image'
labels: 'cell_labels', 'nucleus_labels'
points: 'transcripts'
tables: 'table'
coordinate_systems: 'global'

Data structure:

images

NameDescription
imageThe raw image data.

labels

NameDescription
cell_labels(Optional) Vendor-provided cell segmentation labels, exposed as a segmentation prior.
nucleus_labels(Optional) Vendor-provided nucleus segmentation labels, exposed as a segmentation prior.

points

transcripts: Point cloud data of transcripts.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
qvfloat(Optional) Quality value of the point.
transcript_idlongUnique identifier of the transcript.
overlaps_nucleusboolean(Optional) Whether the point overlaps with the nucleus (derived from morphology).
cell_idinteger(Optional) Vendor-provided cell assignment from the raw data, exposed as a segmentation prior. This is NOT the ground truth used for evaluation (which is held out in spatial_solution); methods may freely condition on it.

tables

table: Metadata of spatial dataset.

SlotTypeDescription
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]stringLink to the original source of the dataset.
uns["dataset_reference"]stringBibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]stringThe organism of the sample in the dataset.
uns["orig_dataset_id"]stringThe identifier of the original dataset from which this dataset was derived (if applicable).

coordinate_systems

NameDescription
globalCoordinate system of the replicate.

File format: Solution

Ground truth segmentation labels and cell assignments for method evaluation.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_solution.zarr

Description:

This dataset contains the ground truth cell and nucleus segmentation labels, cell boundaries, and a reference table matching each cell to its label region.

Format:

SpatialData object
labels: 'cell_labels', 'nucleus_labels', 'groundtruth_cell_labels'
points: 'transcripts'
shapes: 'cell_boundaries', 'nucleus_boundaries'
tables: 'table'

Data structure:

labels

NameDescription
cell_labelsVendor-provided cell segmentation labels.
nucleus_labelsVendor-provided nucleus segmentation labels.
groundtruth_cell_labels(Optional) Manually annotated cell segmentation labels used as ground truth for evaluation.

points

transcripts: Point cloud data of transcripts with ground truth cell assignments.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
cell_idintegerGround truth cell assignment (0 = background).
transcript_idlongUnique identifier of the transcript.

shapes

cell_boundaries: Ground truth cell boundary shapes.

ColumnTypeDescription
geometryobjectGeometry of the cell boundary.

nucleus_boundaries: Ground truth nucleus boundary shapes.

ColumnTypeDescription
geometryobjectGeometry of the nucleus boundary.

tables

table: Reference cell metadata table.

SlotTypeDescription
obs["cell_id"]integerUnique cell identifier, matching instance IDs in the label images.
obs["region"]stringName of the label image this cell belongs to (e.g. ‘cell_labels’).
obs["cell_area"]double(Optional) Area of the cell in pixels.
obs["transcript_counts"]integer(Optional) Total number of transcripts assigned to this cell.
obs["groundtruth_cell_type"]string(Optional) Manually curated cell type annotations which serves as ground truth for evaluations.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["orig_dataset_id"]stringThe identifier of the original dataset from which this dataset was derived (if applicable).

File format: scRNA-seq Reference

A single-cell reference dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/scrnaseq_reference.h5ad

Description:

This dataset contains preprocessed counts and metadata for single-cell RNA-seq data.

Format:

AnnData object
obs: 'cell_type'
var: 'feature_id', 'feature_name', 'hvg'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized', 'normalized_log', 'normalized_log_scaled'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]doubleNormalized expression values.
layers["normalized_log"]doubleLog1p normalized expression values.
layers["normalized_log_scaled"]doubleLog1p normalized expression values scaled to unit variance and zero mean.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

Component type: Control Method

Quality control methods for verifying the pipeline.

Arguments:

NameTypeDescription
--inputfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--input_solutionfileGround truth segmentation labels and cell assignments for method evaluation.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Method

A method.

Arguments:

NameTypeDescription
--inputfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Output processor

An output processor for the prediction.

Arguments:

NameTypeDescription
--input_predictionfileA predicted dataset as output by a method.
--input_spatial_unlabelledfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--outputfile(Output) A processed predicted dataset, ready to be used as input for the evaluation.

Component type: Metric

A task template metric.

Arguments:

NameTypeDescription
--input_predictionfileA processed predicted dataset, ready to be used as input for the evaluation.
--input_solutionfileGround truth segmentation labels and cell assignments for method evaluation.
--outputfile(Output) File indicating the score of a metric.

File format: Predicted data

A predicted dataset as output by a method.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/prediction.zarr

Format:

SpatialData object
labels: 'segmentation'
tables: 'table'

Data structure:

labels

NameDescription
segmentationSegmentation of the data.

tables

table: AnnData table.

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["method_id"]stringA unique identifier for the method.

File format: Processed prediction

A processed predicted dataset, ready to be used as input for the evaluation.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/processed_prediction.zarr

Format:

SpatialData object
labels: 'segmentation'
tables: 'table'

Data structure:

labels

NameDescription
segmentationSegmentation of the data.

tables

table: AnnData table.

SlotTypeDescription
obs["cell_id"]stringCell ID.
obs["region"]stringRegion.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
layers["counts"]integerRaw counts.
layers["normalized"]doubleNormalized expression values.
layers["normalized_log"]doubleLog1p normalized expression values.
layers["normalized_log_scaled"]doubleLog1p normalized expression values scaled to unit variance and zero mean.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["method_id"]stringA unique identifier for the method.

File format: Score

File indicating the score of a metric.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/score.h5ad

Format:

AnnData object
uns: 'dataset_id', 'normalization_id', 'method_id', 'metric_ids', 'metric_values'

Data structure:

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["normalization_id"]stringWhich normalization was used.
uns["method_id"]stringA unique identifier for the method.
uns["metric_ids"]stringOne or more unique metric identifiers.
uns["metric_values"]doubleThe metric values obtained for the given prediction. Must be of same length as ‘metric_ids’.

File format: Common SC Dataset

An unprocessed dataset as output by a dataset loader.

Example file: resources_test/common/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

Description:

This dataset contains raw counts and metadata as output by a dataset loader.

The format of this file is mainly derived from the CELLxGENE schema v4.0.0.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

About

No description, website, or topics provided.

Resources

Contributing

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Spatial Segmentation

A one sentence summary of purpose and methodology. Used for creating an overview tables.

Repository: openproblems-bio/task_spatial_segmentation

Description

Provide a clear and concise description of your task, detailing the specific problem it aims to solve. Outline the input data types, the expected output, and any assumptions or constraints. Be sure to explain any terminology or concepts that are essential for understanding the task.

Explain the motivation behind your proposed task. Describe the biological or computational problem you aim to address and why it’s important. Discuss the current state of research in this area and any gaps or challenges that your task could help address. This section should convince readers of the significance and relevance of your task.

Authors & contributors

NameRolesOrcidGithub
Daria Romanovskaiamaintainer, author0000-0003-2831-0919dariarom94
Florian Heylmaintainer, author0000-0002-3651-5685heylf
Robrecht Cannoodtauthor0000-0003-3641-729Xrcannood

API

flowchart TB
file_common_ist("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-ist-dataset'>Common iST Dataset</a>")
comp_data_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-data-processor'>Data processor</a>"/]
file_spatial_unlabelled("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-unlabelled'>Unlabelled</a>")
file_spatial_solution("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-solution'>Solution</a>")
file_scrnaseq_reference("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-scrna-seq-reference'>scRNA-seq Reference</a>")
comp_control_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-control-method'>Control Method</a>"/]
comp_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-method'>Method</a>"/]
comp_output_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-output-processor'>Output processor</a>"/]
comp_metric[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-metric'>Metric</a>"/]
file_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-predicted-data'>Predicted data</a>")
file_processed_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-processed-prediction'>Processed prediction</a>")
file_score("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-score'>Score</a>")
file_common_scrnaseq("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-sc-dataset'>Common SC Dataset</a>")
file_common_ist---comp_data_processor
comp_data_processor-->file_spatial_unlabelled
comp_data_processor-->file_spatial_solution
comp_data_processor-->file_scrnaseq_reference
file_spatial_unlabelled---comp_control_method
file_spatial_unlabelled---comp_method
file_spatial_unlabelled---comp_output_processor
file_spatial_solution---comp_control_method
file_spatial_solution---comp_metric
comp_control_method-->file_prediction
comp_method-->file_prediction
comp_output_processor-->file_processed_prediction
comp_metric-->file_score
file_prediction---comp_output_processor
file_processed_prediction---comp_metric
file_common_scrnaseq---comp_data_processor
Loading

File format: Common iST Dataset

An unprocessed spatial imaging dataset stored as a zarr file.

Example file: resources_test/common/2023_10x_mouse_brain_xenium_rep1/dataset.zarr

Description:

This dataset contains raw images, labels, points, shapes, and tables as output by a dataset loader.

Format:

SpatialData object
images: 'image', 'image_3D', 'he_image'
labels: 'cell_labels', 'nucleus_labels', 'groundtruth_cell_labels'
points: 'transcripts'
shapes: 'cell_boundaries', 'nucleus_boundaries'
tables: 'table'
coordinate_systems: 'global'

Data structure:

images

NameDescription
imageThe raw image data.
image_3D(Optional) The raw 3D image data.
he_image(Optional) H&E image data.

labels

NameDescription
cell_labelsVendor-provided cell segmentation labels.
nucleus_labelsVendor-provided nucleus segmentation labels.
groundtruth_cell_labels(Optional) Manually annotated cell segmentation labels used as ground truth for evaluation.

points

transcripts: Point cloud data of transcripts.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
cell_idinteger(Optional) Unique identifier of the cell.
nucleus_idinteger(Optional) Unique identifier of the nucleus.
cell_typestring(Optional) Cell type of the cell.
qvfloat(Optional) Quality value of the point.
transcript_idlongUnique identifier of the transcript.
overlaps_nucleusboolean(Optional) Whether the point overlaps with a nucleus.

shapes

cell_boundaries: Cell boundaries.

ColumnTypeDescription
geometryobjectGeometry of the cell boundary.

nucleus_boundaries: Nucleus boundaries.

ColumnTypeDescription
geometryobjectGeometry of the nucleus boundary.

tables

table: Metadata of spatial dataset.

SlotTypeDescription
obs["cell_id"]stringA unique identifier for the cell.
obs["groundtruth_cell_type"]string(Optional) Manually curated cell type annotations which serves as ground truth for evaluations.
var["gene_ids"]stringUnique identifier for the gene.
var["feature_types"]stringType of the feature.
obsm["spatial"]doubleSpatial coordinates of the cell.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]stringLink to the original source of the dataset.
uns["dataset_reference"]stringBibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]stringThe organism of the sample in the dataset.
uns["segmentation_id"]stringA unique identifier for the segmentation.

coordinate_systems

NameDescription
globalCoordinate system of the replicate.

Component type: Data processor

A data processor.

Arguments:

NameTypeDescription
--input_spfileAn unprocessed spatial imaging dataset stored as a zarr file.
--input_scfileAn unprocessed dataset as output by a dataset loader.
--output_spatial_unlabelledfile(Output) Preprocessed spatial transcriptomics data without segmentation labels for method input.
--output_spatial_solutionfile(Output) Ground truth segmentation labels and cell assignments for method evaluation.
--output_scrnaseq_referencefile(Output) A single-cell reference dataset, preprocessed for this benchmark.

File format: Unlabelled

Preprocessed spatial transcriptomics data without segmentation labels for method input.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_unlabelled.zarr

Description:

This dataset contains preprocessed images and transcript point clouds for spatial transcriptomics data. Ground truth segmentation labels are intentionally excluded to prevent methods from cheating.

Format:

SpatialData object
images: 'image'
labels: 'cell_labels', 'nucleus_labels'
points: 'transcripts'
tables: 'table'
coordinate_systems: 'global'

Data structure:

images

NameDescription
imageThe raw image data.

labels

NameDescription
cell_labels(Optional) Vendor-provided cell segmentation labels, exposed as a segmentation prior.
nucleus_labels(Optional) Vendor-provided nucleus segmentation labels, exposed as a segmentation prior.

points

transcripts: Point cloud data of transcripts.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
qvfloat(Optional) Quality value of the point.
transcript_idlongUnique identifier of the transcript.
overlaps_nucleusboolean(Optional) Whether the point overlaps with the nucleus (derived from morphology).
cell_idinteger(Optional) Vendor-provided cell assignment from the raw data, exposed as a segmentation prior. This is NOT the ground truth used for evaluation (which is held out in spatial_solution); methods may freely condition on it.

tables

table: Metadata of spatial dataset.

SlotTypeDescription
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]stringLink to the original source of the dataset.
uns["dataset_reference"]stringBibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]stringThe organism of the sample in the dataset.
uns["orig_dataset_id"]stringThe identifier of the original dataset from which this dataset was derived (if applicable).

coordinate_systems

NameDescription
globalCoordinate system of the replicate.

File format: Solution

Ground truth segmentation labels and cell assignments for method evaluation.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_solution.zarr

Description:

This dataset contains the ground truth cell and nucleus segmentation labels, cell boundaries, and a reference table matching each cell to its label region.

Format:

SpatialData object
labels: 'cell_labels', 'nucleus_labels', 'groundtruth_cell_labels'
points: 'transcripts'
shapes: 'cell_boundaries', 'nucleus_boundaries'
tables: 'table'

Data structure:

labels

NameDescription
cell_labelsVendor-provided cell segmentation labels.
nucleus_labelsVendor-provided nucleus segmentation labels.
groundtruth_cell_labels(Optional) Manually annotated cell segmentation labels used as ground truth for evaluation.

points

transcripts: Point cloud data of transcripts with ground truth cell assignments.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
cell_idintegerGround truth cell assignment (0 = background).
transcript_idlongUnique identifier of the transcript.

shapes

cell_boundaries: Ground truth cell boundary shapes.

ColumnTypeDescription
geometryobjectGeometry of the cell boundary.

nucleus_boundaries: Ground truth nucleus boundary shapes.

ColumnTypeDescription
geometryobjectGeometry of the nucleus boundary.

tables

table: Reference cell metadata table.

SlotTypeDescription
obs["cell_id"]integerUnique cell identifier, matching instance IDs in the label images.
obs["region"]stringName of the label image this cell belongs to (e.g. ‘cell_labels’).
obs["cell_area"]double(Optional) Area of the cell in pixels.
obs["transcript_counts"]integer(Optional) Total number of transcripts assigned to this cell.
obs["groundtruth_cell_type"]string(Optional) Manually curated cell type annotations which serves as ground truth for evaluations.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["orig_dataset_id"]stringThe identifier of the original dataset from which this dataset was derived (if applicable).

File format: scRNA-seq Reference

A single-cell reference dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/scrnaseq_reference.h5ad

Description:

This dataset contains preprocessed counts and metadata for single-cell RNA-seq data.

Format:

AnnData object
obs: 'cell_type'
var: 'feature_id', 'feature_name', 'hvg'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized', 'normalized_log', 'normalized_log_scaled'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]doubleNormalized expression values.
layers["normalized_log"]doubleLog1p normalized expression values.
layers["normalized_log_scaled"]doubleLog1p normalized expression values scaled to unit variance and zero mean.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

Component type: Control Method

Quality control methods for verifying the pipeline.

Arguments:

NameTypeDescription
--inputfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--input_solutionfileGround truth segmentation labels and cell assignments for method evaluation.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Method

A method.

Arguments:

NameTypeDescription
--inputfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Output processor

An output processor for the prediction.

Arguments:

NameTypeDescription
--input_predictionfileA predicted dataset as output by a method.
--input_spatial_unlabelledfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--outputfile(Output) A processed predicted dataset, ready to be used as input for the evaluation.

Component type: Metric

A task template metric.

Arguments:

NameTypeDescription
--input_predictionfileA processed predicted dataset, ready to be used as input for the evaluation.
--input_solutionfileGround truth segmentation labels and cell assignments for method evaluation.
--outputfile(Output) File indicating the score of a metric.

File format: Predicted data

A predicted dataset as output by a method.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/prediction.zarr

Format:

SpatialData object
labels: 'segmentation'
tables: 'table'

Data structure:

labels

NameDescription
segmentationSegmentation of the data.

tables

table: AnnData table.

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["method_id"]stringA unique identifier for the method.

File format: Processed prediction

A processed predicted dataset, ready to be used as input for the evaluation.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/processed_prediction.zarr

Format:

SpatialData object
labels: 'segmentation'
tables: 'table'

Data structure:

labels

NameDescription
segmentationSegmentation of the data.

tables

table: AnnData table.

SlotTypeDescription
obs["cell_id"]stringCell ID.
obs["region"]stringRegion.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
layers["counts"]integerRaw counts.
layers["normalized"]doubleNormalized expression values.
layers["normalized_log"]doubleLog1p normalized expression values.
layers["normalized_log_scaled"]doubleLog1p normalized expression values scaled to unit variance and zero mean.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["method_id"]stringA unique identifier for the method.

File format: Score

File indicating the score of a metric.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/score.h5ad

Format:

AnnData object
uns: 'dataset_id', 'normalization_id', 'method_id', 'metric_ids', 'metric_values'

Data structure:

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["normalization_id"]stringWhich normalization was used.
uns["method_id"]stringA unique identifier for the method.
uns["metric_ids"]stringOne or more unique metric identifiers.
uns["metric_values"]doubleThe metric values obtained for the given prediction. Must be of same length as ‘metric_ids’.

File format: Common SC Dataset

An unprocessed dataset as output by a dataset loader.

Example file: resources_test/common/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

Description:

This dataset contains raw counts and metadata as output by a dataset loader.

The format of this file is mainly derived from the CELLxGENE schema v4.0.0.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

About

No description, website, or topics provided.

Resources

Contributing

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Spatial Segmentation

A one sentence summary of purpose and methodology. Used for creating an overview tables.

Repository: openproblems-bio/task_spatial_segmentation

Description

Provide a clear and concise description of your task, detailing the specific problem it aims to solve. Outline the input data types, the expected output, and any assumptions or constraints. Be sure to explain any terminology or concepts that are essential for understanding the task.

Explain the motivation behind your proposed task. Describe the biological or computational problem you aim to address and why it’s important. Discuss the current state of research in this area and any gaps or challenges that your task could help address. This section should convince readers of the significance and relevance of your task.

Authors & contributors

NameRolesOrcidGithub
Daria Romanovskaiamaintainer, author0000-0003-2831-0919dariarom94
Florian Heylmaintainer, author0000-0002-3651-5685heylf
Robrecht Cannoodtauthor0000-0003-3641-729Xrcannood

API

flowchart TB
file_common_ist("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-ist-dataset'>Common iST Dataset</a>")
comp_data_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-data-processor'>Data processor</a>"/]
file_spatial_unlabelled("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-unlabelled'>Unlabelled</a>")
file_spatial_solution("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-solution'>Solution</a>")
file_scrnaseq_reference("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-scrna-seq-reference'>scRNA-seq Reference</a>")
comp_control_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-control-method'>Control Method</a>"/]
comp_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-method'>Method</a>"/]
comp_output_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-output-processor'>Output processor</a>"/]
comp_metric[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-metric'>Metric</a>"/]
file_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-predicted-data'>Predicted data</a>")
file_processed_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-processed-prediction'>Processed prediction</a>")
file_score("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-score'>Score</a>")
file_common_scrnaseq("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-sc-dataset'>Common SC Dataset</a>")
file_common_ist---comp_data_processor
comp_data_processor-->file_spatial_unlabelled
comp_data_processor-->file_spatial_solution
comp_data_processor-->file_scrnaseq_reference
file_spatial_unlabelled---comp_control_method
file_spatial_unlabelled---comp_method
file_spatial_unlabelled---comp_output_processor
file_spatial_solution---comp_control_method
file_spatial_solution---comp_metric
comp_control_method-->file_prediction
comp_method-->file_prediction
comp_output_processor-->file_processed_prediction
comp_metric-->file_score
file_prediction---comp_output_processor
file_processed_prediction---comp_metric
file_common_scrnaseq---comp_data_processor
Loading

File format: Common iST Dataset

An unprocessed spatial imaging dataset stored as a zarr file.

Example file: resources_test/common/2023_10x_mouse_brain_xenium_rep1/dataset.zarr

Description:

This dataset contains raw images, labels, points, shapes, and tables as output by a dataset loader.

Format:

SpatialData object
images: 'image', 'image_3D', 'he_image'
labels: 'cell_labels', 'nucleus_labels', 'groundtruth_cell_labels'
points: 'transcripts'
shapes: 'cell_boundaries', 'nucleus_boundaries'
tables: 'table'
coordinate_systems: 'global'

Data structure:

images

NameDescription
imageThe raw image data.
image_3D(Optional) The raw 3D image data.
he_image(Optional) H&E image data.

labels

NameDescription
cell_labelsVendor-provided cell segmentation labels.
nucleus_labelsVendor-provided nucleus segmentation labels.
groundtruth_cell_labels(Optional) Manually annotated cell segmentation labels used as ground truth for evaluation.

points

transcripts: Point cloud data of transcripts.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
cell_idinteger(Optional) Unique identifier of the cell.
nucleus_idinteger(Optional) Unique identifier of the nucleus.
cell_typestring(Optional) Cell type of the cell.
qvfloat(Optional) Quality value of the point.
transcript_idlongUnique identifier of the transcript.
overlaps_nucleusboolean(Optional) Whether the point overlaps with a nucleus.

shapes

cell_boundaries: Cell boundaries.

ColumnTypeDescription
geometryobjectGeometry of the cell boundary.

nucleus_boundaries: Nucleus boundaries.

ColumnTypeDescription
geometryobjectGeometry of the nucleus boundary.

tables

table: Metadata of spatial dataset.

SlotTypeDescription
obs["cell_id"]stringA unique identifier for the cell.
obs["groundtruth_cell_type"]string(Optional) Manually curated cell type annotations which serves as ground truth for evaluations.
var["gene_ids"]stringUnique identifier for the gene.
var["feature_types"]stringType of the feature.
obsm["spatial"]doubleSpatial coordinates of the cell.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]stringLink to the original source of the dataset.
uns["dataset_reference"]stringBibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]stringThe organism of the sample in the dataset.
uns["segmentation_id"]stringA unique identifier for the segmentation.

coordinate_systems

NameDescription
globalCoordinate system of the replicate.

Component type: Data processor

A data processor.

Arguments:

NameTypeDescription
--input_spfileAn unprocessed spatial imaging dataset stored as a zarr file.
--input_scfileAn unprocessed dataset as output by a dataset loader.
--output_spatial_unlabelledfile(Output) Preprocessed spatial transcriptomics data without segmentation labels for method input.
--output_spatial_solutionfile(Output) Ground truth segmentation labels and cell assignments for method evaluation.
--output_scrnaseq_referencefile(Output) A single-cell reference dataset, preprocessed for this benchmark.

File format: Unlabelled

Preprocessed spatial transcriptomics data without segmentation labels for method input.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_unlabelled.zarr

Description:

This dataset contains preprocessed images and transcript point clouds for spatial transcriptomics data. Ground truth segmentation labels are intentionally excluded to prevent methods from cheating.

Format:

SpatialData object
images: 'image'
labels: 'cell_labels', 'nucleus_labels'
points: 'transcripts'
tables: 'table'
coordinate_systems: 'global'

Data structure:

images

NameDescription
imageThe raw image data.

labels

NameDescription
cell_labels(Optional) Vendor-provided cell segmentation labels, exposed as a segmentation prior.
nucleus_labels(Optional) Vendor-provided nucleus segmentation labels, exposed as a segmentation prior.

points

transcripts: Point cloud data of transcripts.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
qvfloat(Optional) Quality value of the point.
transcript_idlongUnique identifier of the transcript.
overlaps_nucleusboolean(Optional) Whether the point overlaps with the nucleus (derived from morphology).
cell_idinteger(Optional) Vendor-provided cell assignment from the raw data, exposed as a segmentation prior. This is NOT the ground truth used for evaluation (which is held out in spatial_solution); methods may freely condition on it.

tables

table: Metadata of spatial dataset.

SlotTypeDescription
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]stringLink to the original source of the dataset.
uns["dataset_reference"]stringBibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]stringThe organism of the sample in the dataset.
uns["orig_dataset_id"]stringThe identifier of the original dataset from which this dataset was derived (if applicable).

coordinate_systems

NameDescription
globalCoordinate system of the replicate.

File format: Solution

Ground truth segmentation labels and cell assignments for method evaluation.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_solution.zarr

Description:

This dataset contains the ground truth cell and nucleus segmentation labels, cell boundaries, and a reference table matching each cell to its label region.

Format:

SpatialData object
labels: 'cell_labels', 'nucleus_labels', 'groundtruth_cell_labels'
points: 'transcripts'
shapes: 'cell_boundaries', 'nucleus_boundaries'
tables: 'table'

Data structure:

labels

NameDescription
cell_labelsVendor-provided cell segmentation labels.
nucleus_labelsVendor-provided nucleus segmentation labels.
groundtruth_cell_labels(Optional) Manually annotated cell segmentation labels used as ground truth for evaluation.

points

transcripts: Point cloud data of transcripts with ground truth cell assignments.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
cell_idintegerGround truth cell assignment (0 = background).
transcript_idlongUnique identifier of the transcript.

shapes

cell_boundaries: Ground truth cell boundary shapes.

ColumnTypeDescription
geometryobjectGeometry of the cell boundary.

nucleus_boundaries: Ground truth nucleus boundary shapes.

ColumnTypeDescription
geometryobjectGeometry of the nucleus boundary.

tables

table: Reference cell metadata table.

SlotTypeDescription
obs["cell_id"]integerUnique cell identifier, matching instance IDs in the label images.
obs["region"]stringName of the label image this cell belongs to (e.g. ‘cell_labels’).
obs["cell_area"]double(Optional) Area of the cell in pixels.
obs["transcript_counts"]integer(Optional) Total number of transcripts assigned to this cell.
obs["groundtruth_cell_type"]string(Optional) Manually curated cell type annotations which serves as ground truth for evaluations.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["orig_dataset_id"]stringThe identifier of the original dataset from which this dataset was derived (if applicable).

File format: scRNA-seq Reference

A single-cell reference dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/scrnaseq_reference.h5ad

Description:

This dataset contains preprocessed counts and metadata for single-cell RNA-seq data.

Format:

AnnData object
obs: 'cell_type'
var: 'feature_id', 'feature_name', 'hvg'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized', 'normalized_log', 'normalized_log_scaled'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]doubleNormalized expression values.
layers["normalized_log"]doubleLog1p normalized expression values.
layers["normalized_log_scaled"]doubleLog1p normalized expression values scaled to unit variance and zero mean.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

Component type: Control Method

Quality control methods for verifying the pipeline.

Arguments:

NameTypeDescription
--inputfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--input_solutionfileGround truth segmentation labels and cell assignments for method evaluation.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Method

A method.

Arguments:

NameTypeDescription
--inputfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Output processor

An output processor for the prediction.

Arguments:

NameTypeDescription
--input_predictionfileA predicted dataset as output by a method.
--input_spatial_unlabelledfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--outputfile(Output) A processed predicted dataset, ready to be used as input for the evaluation.

Component type: Metric

A task template metric.

Arguments:

NameTypeDescription
--input_predictionfileA processed predicted dataset, ready to be used as input for the evaluation.
--input_solutionfileGround truth segmentation labels and cell assignments for method evaluation.
--outputfile(Output) File indicating the score of a metric.

File format: Predicted data

A predicted dataset as output by a method.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/prediction.zarr

Format:

SpatialData object
labels: 'segmentation'
tables: 'table'

Data structure:

labels

NameDescription
segmentationSegmentation of the data.

tables

table: AnnData table.

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["method_id"]stringA unique identifier for the method.

File format: Processed prediction

A processed predicted dataset, ready to be used as input for the evaluation.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/processed_prediction.zarr

Format:

SpatialData object
labels: 'segmentation'
tables: 'table'

Data structure:

labels

NameDescription
segmentationSegmentation of the data.

tables

table: AnnData table.

SlotTypeDescription
obs["cell_id"]stringCell ID.
obs["region"]stringRegion.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
layers["counts"]integerRaw counts.
layers["normalized"]doubleNormalized expression values.
layers["normalized_log"]doubleLog1p normalized expression values.
layers["normalized_log_scaled"]doubleLog1p normalized expression values scaled to unit variance and zero mean.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["method_id"]stringA unique identifier for the method.

File format: Score

File indicating the score of a metric.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/score.h5ad

Format:

AnnData object
uns: 'dataset_id', 'normalization_id', 'method_id', 'metric_ids', 'metric_values'

Data structure:

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["normalization_id"]stringWhich normalization was used.
uns["method_id"]stringA unique identifier for the method.
uns["metric_ids"]stringOne or more unique metric identifiers.
uns["metric_values"]doubleThe metric values obtained for the given prediction. Must be of same length as ‘metric_ids’.

File format: Common SC Dataset

An unprocessed dataset as output by a dataset loader.

Example file: resources_test/common/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

Description:

This dataset contains raw counts and metadata as output by a dataset loader.

The format of this file is mainly derived from the CELLxGENE schema v4.0.0.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

About

No description, website, or topics provided.

Resources

Contributing

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Spatial Segmentation

A one sentence summary of purpose and methodology. Used for creating an overview tables.

Repository: openproblems-bio/task_spatial_segmentation

Description

Provide a clear and concise description of your task, detailing the specific problem it aims to solve. Outline the input data types, the expected output, and any assumptions or constraints. Be sure to explain any terminology or concepts that are essential for understanding the task.

Explain the motivation behind your proposed task. Describe the biological or computational problem you aim to address and why it’s important. Discuss the current state of research in this area and any gaps or challenges that your task could help address. This section should convince readers of the significance and relevance of your task.

Authors & contributors

NameRolesOrcidGithub
Daria Romanovskaiamaintainer, author0000-0003-2831-0919dariarom94
Florian Heylmaintainer, author0000-0002-3651-5685heylf
Robrecht Cannoodtauthor0000-0003-3641-729Xrcannood

API

flowchart TB
file_common_ist("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-ist-dataset'>Common iST Dataset</a>")
comp_data_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-data-processor'>Data processor</a>"/]
file_spatial_unlabelled("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-unlabelled'>Unlabelled</a>")
file_spatial_solution("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-solution'>Solution</a>")
file_scrnaseq_reference("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-scrna-seq-reference'>scRNA-seq Reference</a>")
comp_control_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-control-method'>Control Method</a>"/]
comp_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-method'>Method</a>"/]
comp_output_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-output-processor'>Output processor</a>"/]
comp_metric[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-metric'>Metric</a>"/]
file_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-predicted-data'>Predicted data</a>")
file_processed_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-processed-prediction'>Processed prediction</a>")
file_score("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-score'>Score</a>")
file_common_scrnaseq("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-sc-dataset'>Common SC Dataset</a>")
file_common_ist---comp_data_processor
comp_data_processor-->file_spatial_unlabelled
comp_data_processor-->file_spatial_solution
comp_data_processor-->file_scrnaseq_reference
file_spatial_unlabelled---comp_control_method
file_spatial_unlabelled---comp_method
file_spatial_unlabelled---comp_output_processor
file_spatial_solution---comp_control_method
file_spatial_solution---comp_metric
comp_control_method-->file_prediction
comp_method-->file_prediction
comp_output_processor-->file_processed_prediction
comp_metric-->file_score
file_prediction---comp_output_processor
file_processed_prediction---comp_metric
file_common_scrnaseq---comp_data_processor
Loading

File format: Common iST Dataset

An unprocessed spatial imaging dataset stored as a zarr file.

Example file: resources_test/common/2023_10x_mouse_brain_xenium_rep1/dataset.zarr

Description:

This dataset contains raw images, labels, points, shapes, and tables as output by a dataset loader.

Format:

SpatialData object
images: 'image', 'image_3D', 'he_image'
labels: 'cell_labels', 'nucleus_labels', 'groundtruth_cell_labels'
points: 'transcripts'
shapes: 'cell_boundaries', 'nucleus_boundaries'
tables: 'table'
coordinate_systems: 'global'

Data structure:

images

NameDescription
imageThe raw image data.
image_3D(Optional) The raw 3D image data.
he_image(Optional) H&E image data.

labels

NameDescription
cell_labelsVendor-provided cell segmentation labels.
nucleus_labelsVendor-provided nucleus segmentation labels.
groundtruth_cell_labels(Optional) Manually annotated cell segmentation labels used as ground truth for evaluation.

points

transcripts: Point cloud data of transcripts.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
cell_idinteger(Optional) Unique identifier of the cell.
nucleus_idinteger(Optional) Unique identifier of the nucleus.
cell_typestring(Optional) Cell type of the cell.
qvfloat(Optional) Quality value of the point.
transcript_idlongUnique identifier of the transcript.
overlaps_nucleusboolean(Optional) Whether the point overlaps with a nucleus.

shapes

cell_boundaries: Cell boundaries.

ColumnTypeDescription
geometryobjectGeometry of the cell boundary.

nucleus_boundaries: Nucleus boundaries.

ColumnTypeDescription
geometryobjectGeometry of the nucleus boundary.

tables

table: Metadata of spatial dataset.

SlotTypeDescription
obs["cell_id"]stringA unique identifier for the cell.
obs["groundtruth_cell_type"]string(Optional) Manually curated cell type annotations which serves as ground truth for evaluations.
var["gene_ids"]stringUnique identifier for the gene.
var["feature_types"]stringType of the feature.
obsm["spatial"]doubleSpatial coordinates of the cell.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]stringLink to the original source of the dataset.
uns["dataset_reference"]stringBibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]stringThe organism of the sample in the dataset.
uns["segmentation_id"]stringA unique identifier for the segmentation.

coordinate_systems

NameDescription
globalCoordinate system of the replicate.

Component type: Data processor

A data processor.

Arguments:

NameTypeDescription
--input_spfileAn unprocessed spatial imaging dataset stored as a zarr file.
--input_scfileAn unprocessed dataset as output by a dataset loader.
--output_spatial_unlabelledfile(Output) Preprocessed spatial transcriptomics data without segmentation labels for method input.
--output_spatial_solutionfile(Output) Ground truth segmentation labels and cell assignments for method evaluation.
--output_scrnaseq_referencefile(Output) A single-cell reference dataset, preprocessed for this benchmark.

File format: Unlabelled

Preprocessed spatial transcriptomics data without segmentation labels for method input.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_unlabelled.zarr

Description:

This dataset contains preprocessed images and transcript point clouds for spatial transcriptomics data. Ground truth segmentation labels are intentionally excluded to prevent methods from cheating.

Format:

SpatialData object
images: 'image'
labels: 'cell_labels', 'nucleus_labels'
points: 'transcripts'
tables: 'table'
coordinate_systems: 'global'

Data structure:

images

NameDescription
imageThe raw image data.

labels

NameDescription
cell_labels(Optional) Vendor-provided cell segmentation labels, exposed as a segmentation prior.
nucleus_labels(Optional) Vendor-provided nucleus segmentation labels, exposed as a segmentation prior.

points

transcripts: Point cloud data of transcripts.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
qvfloat(Optional) Quality value of the point.
transcript_idlongUnique identifier of the transcript.
overlaps_nucleusboolean(Optional) Whether the point overlaps with the nucleus (derived from morphology).
cell_idinteger(Optional) Vendor-provided cell assignment from the raw data, exposed as a segmentation prior. This is NOT the ground truth used for evaluation (which is held out in spatial_solution); methods may freely condition on it.

tables

table: Metadata of spatial dataset.

SlotTypeDescription
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]stringLink to the original source of the dataset.
uns["dataset_reference"]stringBibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]stringThe organism of the sample in the dataset.
uns["orig_dataset_id"]stringThe identifier of the original dataset from which this dataset was derived (if applicable).

coordinate_systems

NameDescription
globalCoordinate system of the replicate.

File format: Solution

Ground truth segmentation labels and cell assignments for method evaluation.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_solution.zarr

Description:

This dataset contains the ground truth cell and nucleus segmentation labels, cell boundaries, and a reference table matching each cell to its label region.

Format:

SpatialData object
labels: 'cell_labels', 'nucleus_labels', 'groundtruth_cell_labels'
points: 'transcripts'
shapes: 'cell_boundaries', 'nucleus_boundaries'
tables: 'table'

Data structure:

labels

NameDescription
cell_labelsVendor-provided cell segmentation labels.
nucleus_labelsVendor-provided nucleus segmentation labels.
groundtruth_cell_labels(Optional) Manually annotated cell segmentation labels used as ground truth for evaluation.

points

transcripts: Point cloud data of transcripts with ground truth cell assignments.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
cell_idintegerGround truth cell assignment (0 = background).
transcript_idlongUnique identifier of the transcript.

shapes

cell_boundaries: Ground truth cell boundary shapes.

ColumnTypeDescription
geometryobjectGeometry of the cell boundary.

nucleus_boundaries: Ground truth nucleus boundary shapes.

ColumnTypeDescription
geometryobjectGeometry of the nucleus boundary.

tables

table: Reference cell metadata table.

SlotTypeDescription
obs["cell_id"]integerUnique cell identifier, matching instance IDs in the label images.
obs["region"]stringName of the label image this cell belongs to (e.g. ‘cell_labels’).
obs["cell_area"]double(Optional) Area of the cell in pixels.
obs["transcript_counts"]integer(Optional) Total number of transcripts assigned to this cell.
obs["groundtruth_cell_type"]string(Optional) Manually curated cell type annotations which serves as ground truth for evaluations.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["orig_dataset_id"]stringThe identifier of the original dataset from which this dataset was derived (if applicable).

File format: scRNA-seq Reference

A single-cell reference dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/scrnaseq_reference.h5ad

Description:

This dataset contains preprocessed counts and metadata for single-cell RNA-seq data.

Format:

AnnData object
obs: 'cell_type'
var: 'feature_id', 'feature_name', 'hvg'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized', 'normalized_log', 'normalized_log_scaled'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]doubleNormalized expression values.
layers["normalized_log"]doubleLog1p normalized expression values.
layers["normalized_log_scaled"]doubleLog1p normalized expression values scaled to unit variance and zero mean.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

Component type: Control Method

Quality control methods for verifying the pipeline.

Arguments:

NameTypeDescription
--inputfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--input_solutionfileGround truth segmentation labels and cell assignments for method evaluation.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Method

A method.

Arguments:

NameTypeDescription
--inputfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Output processor

An output processor for the prediction.

Arguments:

NameTypeDescription
--input_predictionfileA predicted dataset as output by a method.
--input_spatial_unlabelledfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--outputfile(Output) A processed predicted dataset, ready to be used as input for the evaluation.

Component type: Metric

A task template metric.

Arguments:

NameTypeDescription
--input_predictionfileA processed predicted dataset, ready to be used as input for the evaluation.
--input_solutionfileGround truth segmentation labels and cell assignments for method evaluation.
--outputfile(Output) File indicating the score of a metric.

File format: Predicted data

A predicted dataset as output by a method.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/prediction.zarr

Format:

SpatialData object
labels: 'segmentation'
tables: 'table'

Data structure:

labels

NameDescription
segmentationSegmentation of the data.

tables

table: AnnData table.

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["method_id"]stringA unique identifier for the method.

File format: Processed prediction

A processed predicted dataset, ready to be used as input for the evaluation.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/processed_prediction.zarr

Format:

SpatialData object
labels: 'segmentation'
tables: 'table'

Data structure:

labels

NameDescription
segmentationSegmentation of the data.

tables

table: AnnData table.

SlotTypeDescription
obs["cell_id"]stringCell ID.
obs["region"]stringRegion.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
layers["counts"]integerRaw counts.
layers["normalized"]doubleNormalized expression values.
layers["normalized_log"]doubleLog1p normalized expression values.
layers["normalized_log_scaled"]doubleLog1p normalized expression values scaled to unit variance and zero mean.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["method_id"]stringA unique identifier for the method.

File format: Score

File indicating the score of a metric.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/score.h5ad

Format:

AnnData object
uns: 'dataset_id', 'normalization_id', 'method_id', 'metric_ids', 'metric_values'

Data structure:

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["normalization_id"]stringWhich normalization was used.
uns["method_id"]stringA unique identifier for the method.
uns["metric_ids"]stringOne or more unique metric identifiers.
uns["metric_values"]doubleThe metric values obtained for the given prediction. Must be of same length as ‘metric_ids’.

File format: Common SC Dataset

An unprocessed dataset as output by a dataset loader.

Example file: resources_test/common/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

Description:

This dataset contains raw counts and metadata as output by a dataset loader.

The format of this file is mainly derived from the CELLxGENE schema v4.0.0.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

About

No description, website, or topics provided.

Resources

Contributing

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Spatial Segmentation

A one sentence summary of purpose and methodology. Used for creating an overview tables.

Repository: openproblems-bio/task_spatial_segmentation

Description

Provide a clear and concise description of your task, detailing the specific problem it aims to solve. Outline the input data types, the expected output, and any assumptions or constraints. Be sure to explain any terminology or concepts that are essential for understanding the task.

Explain the motivation behind your proposed task. Describe the biological or computational problem you aim to address and why it’s important. Discuss the current state of research in this area and any gaps or challenges that your task could help address. This section should convince readers of the significance and relevance of your task.

Authors & contributors

NameRolesOrcidGithub
Daria Romanovskaiamaintainer, author0000-0003-2831-0919dariarom94
Florian Heylmaintainer, author0000-0002-3651-5685heylf
Robrecht Cannoodtauthor0000-0003-3641-729Xrcannood

API

flowchart TB
file_common_ist("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-ist-dataset'>Common iST Dataset</a>")
comp_data_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-data-processor'>Data processor</a>"/]
file_spatial_unlabelled("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-unlabelled'>Unlabelled</a>")
file_spatial_solution("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-solution'>Solution</a>")
file_scrnaseq_reference("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-scrna-seq-reference'>scRNA-seq Reference</a>")
comp_control_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-control-method'>Control Method</a>"/]
comp_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-method'>Method</a>"/]
comp_output_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-output-processor'>Output processor</a>"/]
comp_metric[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-metric'>Metric</a>"/]
file_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-predicted-data'>Predicted data</a>")
file_processed_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-processed-prediction'>Processed prediction</a>")
file_score("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-score'>Score</a>")
file_common_scrnaseq("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-sc-dataset'>Common SC Dataset</a>")
file_common_ist---comp_data_processor
comp_data_processor-->file_spatial_unlabelled
comp_data_processor-->file_spatial_solution
comp_data_processor-->file_scrnaseq_reference
file_spatial_unlabelled---comp_control_method
file_spatial_unlabelled---comp_method
file_spatial_unlabelled---comp_output_processor
file_spatial_solution---comp_control_method
file_spatial_solution---comp_metric
comp_control_method-->file_prediction
comp_method-->file_prediction
comp_output_processor-->file_processed_prediction
comp_metric-->file_score
file_prediction---comp_output_processor
file_processed_prediction---comp_metric
file_common_scrnaseq---comp_data_processor
Loading

File format: Common iST Dataset

An unprocessed spatial imaging dataset stored as a zarr file.

Example file: resources_test/common/2023_10x_mouse_brain_xenium_rep1/dataset.zarr

Description:

This dataset contains raw images, labels, points, shapes, and tables as output by a dataset loader.

Format:

SpatialData object
images: 'image', 'image_3D', 'he_image'
labels: 'cell_labels', 'nucleus_labels', 'groundtruth_cell_labels'
points: 'transcripts'
shapes: 'cell_boundaries', 'nucleus_boundaries'
tables: 'table'
coordinate_systems: 'global'

Data structure:

images

NameDescription
imageThe raw image data.
image_3D(Optional) The raw 3D image data.
he_image(Optional) H&E image data.

labels

NameDescription
cell_labelsVendor-provided cell segmentation labels.
nucleus_labelsVendor-provided nucleus segmentation labels.
groundtruth_cell_labels(Optional) Manually annotated cell segmentation labels used as ground truth for evaluation.

points

transcripts: Point cloud data of transcripts.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
cell_idinteger(Optional) Unique identifier of the cell.
nucleus_idinteger(Optional) Unique identifier of the nucleus.
cell_typestring(Optional) Cell type of the cell.
qvfloat(Optional) Quality value of the point.
transcript_idlongUnique identifier of the transcript.
overlaps_nucleusboolean(Optional) Whether the point overlaps with a nucleus.

shapes

cell_boundaries: Cell boundaries.

ColumnTypeDescription
geometryobjectGeometry of the cell boundary.

nucleus_boundaries: Nucleus boundaries.

ColumnTypeDescription
geometryobjectGeometry of the nucleus boundary.

tables

table: Metadata of spatial dataset.

SlotTypeDescription
obs["cell_id"]stringA unique identifier for the cell.
obs["groundtruth_cell_type"]string(Optional) Manually curated cell type annotations which serves as ground truth for evaluations.
var["gene_ids"]stringUnique identifier for the gene.
var["feature_types"]stringType of the feature.
obsm["spatial"]doubleSpatial coordinates of the cell.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]stringLink to the original source of the dataset.
uns["dataset_reference"]stringBibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]stringThe organism of the sample in the dataset.
uns["segmentation_id"]stringA unique identifier for the segmentation.

coordinate_systems

NameDescription
globalCoordinate system of the replicate.

Component type: Data processor

A data processor.

Arguments:

NameTypeDescription
--input_spfileAn unprocessed spatial imaging dataset stored as a zarr file.
--input_scfileAn unprocessed dataset as output by a dataset loader.
--output_spatial_unlabelledfile(Output) Preprocessed spatial transcriptomics data without segmentation labels for method input.
--output_spatial_solutionfile(Output) Ground truth segmentation labels and cell assignments for method evaluation.
--output_scrnaseq_referencefile(Output) A single-cell reference dataset, preprocessed for this benchmark.

File format: Unlabelled

Preprocessed spatial transcriptomics data without segmentation labels for method input.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_unlabelled.zarr

Description:

This dataset contains preprocessed images and transcript point clouds for spatial transcriptomics data. Ground truth segmentation labels are intentionally excluded to prevent methods from cheating.

Format:

SpatialData object
images: 'image'
labels: 'cell_labels', 'nucleus_labels'
points: 'transcripts'
tables: 'table'
coordinate_systems: 'global'

Data structure:

images

NameDescription
imageThe raw image data.

labels

NameDescription
cell_labels(Optional) Vendor-provided cell segmentation labels, exposed as a segmentation prior.
nucleus_labels(Optional) Vendor-provided nucleus segmentation labels, exposed as a segmentation prior.

points

transcripts: Point cloud data of transcripts.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
qvfloat(Optional) Quality value of the point.
transcript_idlongUnique identifier of the transcript.
overlaps_nucleusboolean(Optional) Whether the point overlaps with the nucleus (derived from morphology).
cell_idinteger(Optional) Vendor-provided cell assignment from the raw data, exposed as a segmentation prior. This is NOT the ground truth used for evaluation (which is held out in spatial_solution); methods may freely condition on it.

tables

table: Metadata of spatial dataset.

SlotTypeDescription
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]stringLink to the original source of the dataset.
uns["dataset_reference"]stringBibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]stringThe organism of the sample in the dataset.
uns["orig_dataset_id"]stringThe identifier of the original dataset from which this dataset was derived (if applicable).

coordinate_systems

NameDescription
globalCoordinate system of the replicate.

File format: Solution

Ground truth segmentation labels and cell assignments for method evaluation.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_solution.zarr

Description:

This dataset contains the ground truth cell and nucleus segmentation labels, cell boundaries, and a reference table matching each cell to its label region.

Format:

SpatialData object
labels: 'cell_labels', 'nucleus_labels', 'groundtruth_cell_labels'
points: 'transcripts'
shapes: 'cell_boundaries', 'nucleus_boundaries'
tables: 'table'

Data structure:

labels

NameDescription
cell_labelsVendor-provided cell segmentation labels.
nucleus_labelsVendor-provided nucleus segmentation labels.
groundtruth_cell_labels(Optional) Manually annotated cell segmentation labels used as ground truth for evaluation.

points

transcripts: Point cloud data of transcripts with ground truth cell assignments.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
cell_idintegerGround truth cell assignment (0 = background).
transcript_idlongUnique identifier of the transcript.

shapes

cell_boundaries: Ground truth cell boundary shapes.

ColumnTypeDescription
geometryobjectGeometry of the cell boundary.

nucleus_boundaries: Ground truth nucleus boundary shapes.

ColumnTypeDescription
geometryobjectGeometry of the nucleus boundary.

tables

table: Reference cell metadata table.

SlotTypeDescription
obs["cell_id"]integerUnique cell identifier, matching instance IDs in the label images.
obs["region"]stringName of the label image this cell belongs to (e.g. ‘cell_labels’).
obs["cell_area"]double(Optional) Area of the cell in pixels.
obs["transcript_counts"]integer(Optional) Total number of transcripts assigned to this cell.
obs["groundtruth_cell_type"]string(Optional) Manually curated cell type annotations which serves as ground truth for evaluations.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["orig_dataset_id"]stringThe identifier of the original dataset from which this dataset was derived (if applicable).

File format: scRNA-seq Reference

A single-cell reference dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/scrnaseq_reference.h5ad

Description:

This dataset contains preprocessed counts and metadata for single-cell RNA-seq data.

Format:

AnnData object
obs: 'cell_type'
var: 'feature_id', 'feature_name', 'hvg'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized', 'normalized_log', 'normalized_log_scaled'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]doubleNormalized expression values.
layers["normalized_log"]doubleLog1p normalized expression values.
layers["normalized_log_scaled"]doubleLog1p normalized expression values scaled to unit variance and zero mean.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

Component type: Control Method

Quality control methods for verifying the pipeline.

Arguments:

NameTypeDescription
--inputfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--input_solutionfileGround truth segmentation labels and cell assignments for method evaluation.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Method

A method.

Arguments:

NameTypeDescription
--inputfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Output processor

An output processor for the prediction.

Arguments:

NameTypeDescription
--input_predictionfileA predicted dataset as output by a method.
--input_spatial_unlabelledfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--outputfile(Output) A processed predicted dataset, ready to be used as input for the evaluation.

Component type: Metric

A task template metric.

Arguments:

NameTypeDescription
--input_predictionfileA processed predicted dataset, ready to be used as input for the evaluation.
--input_solutionfileGround truth segmentation labels and cell assignments for method evaluation.
--outputfile(Output) File indicating the score of a metric.

File format: Predicted data

A predicted dataset as output by a method.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/prediction.zarr

Format:

SpatialData object
labels: 'segmentation'
tables: 'table'

Data structure:

labels

NameDescription
segmentationSegmentation of the data.

tables

table: AnnData table.

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["method_id"]stringA unique identifier for the method.

File format: Processed prediction

A processed predicted dataset, ready to be used as input for the evaluation.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/processed_prediction.zarr

Format:

SpatialData object
labels: 'segmentation'
tables: 'table'

Data structure:

labels

NameDescription
segmentationSegmentation of the data.

tables

table: AnnData table.

SlotTypeDescription
obs["cell_id"]stringCell ID.
obs["region"]stringRegion.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
layers["counts"]integerRaw counts.
layers["normalized"]doubleNormalized expression values.
layers["normalized_log"]doubleLog1p normalized expression values.
layers["normalized_log_scaled"]doubleLog1p normalized expression values scaled to unit variance and zero mean.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["method_id"]stringA unique identifier for the method.

File format: Score

File indicating the score of a metric.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/score.h5ad

Format:

AnnData object
uns: 'dataset_id', 'normalization_id', 'method_id', 'metric_ids', 'metric_values'

Data structure:

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["normalization_id"]stringWhich normalization was used.
uns["method_id"]stringA unique identifier for the method.
uns["metric_ids"]stringOne or more unique metric identifiers.
uns["metric_values"]doubleThe metric values obtained for the given prediction. Must be of same length as ‘metric_ids’.

File format: Common SC Dataset

An unprocessed dataset as output by a dataset loader.

Example file: resources_test/common/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

Description:

This dataset contains raw counts and metadata as output by a dataset loader.

The format of this file is mainly derived from the CELLxGENE schema v4.0.0.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

About

No description, website, or topics provided.

Resources

Contributing

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Spatial Segmentation

A one sentence summary of purpose and methodology. Used for creating an overview tables.

Repository: openproblems-bio/task_spatial_segmentation

Description

Provide a clear and concise description of your task, detailing the specific problem it aims to solve. Outline the input data types, the expected output, and any assumptions or constraints. Be sure to explain any terminology or concepts that are essential for understanding the task.

Explain the motivation behind your proposed task. Describe the biological or computational problem you aim to address and why it’s important. Discuss the current state of research in this area and any gaps or challenges that your task could help address. This section should convince readers of the significance and relevance of your task.

Authors & contributors

NameRolesOrcidGithub
Daria Romanovskaiamaintainer, author0000-0003-2831-0919dariarom94
Florian Heylmaintainer, author0000-0002-3651-5685heylf
Robrecht Cannoodtauthor0000-0003-3641-729Xrcannood

API

flowchart TB
file_common_ist("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-ist-dataset'>Common iST Dataset</a>")
comp_data_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-data-processor'>Data processor</a>"/]
file_spatial_unlabelled("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-unlabelled'>Unlabelled</a>")
file_spatial_solution("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-solution'>Solution</a>")
file_scrnaseq_reference("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-scrna-seq-reference'>scRNA-seq Reference</a>")
comp_control_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-control-method'>Control Method</a>"/]
comp_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-method'>Method</a>"/]
comp_output_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-output-processor'>Output processor</a>"/]
comp_metric[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-metric'>Metric</a>"/]
file_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-predicted-data'>Predicted data</a>")
file_processed_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-processed-prediction'>Processed prediction</a>")
file_score("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-score'>Score</a>")
file_common_scrnaseq("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-sc-dataset'>Common SC Dataset</a>")
file_common_ist---comp_data_processor
comp_data_processor-->file_spatial_unlabelled
comp_data_processor-->file_spatial_solution
comp_data_processor-->file_scrnaseq_reference
file_spatial_unlabelled---comp_control_method
file_spatial_unlabelled---comp_method
file_spatial_unlabelled---comp_output_processor
file_spatial_solution---comp_control_method
file_spatial_solution---comp_metric
comp_control_method-->file_prediction
comp_method-->file_prediction
comp_output_processor-->file_processed_prediction
comp_metric-->file_score
file_prediction---comp_output_processor
file_processed_prediction---comp_metric
file_common_scrnaseq---comp_data_processor
Loading

File format: Common iST Dataset

An unprocessed spatial imaging dataset stored as a zarr file.

Example file: resources_test/common/2023_10x_mouse_brain_xenium_rep1/dataset.zarr

Description:

This dataset contains raw images, labels, points, shapes, and tables as output by a dataset loader.

Format:

SpatialData object
images: 'image', 'image_3D', 'he_image'
labels: 'cell_labels', 'nucleus_labels', 'groundtruth_cell_labels'
points: 'transcripts'
shapes: 'cell_boundaries', 'nucleus_boundaries'
tables: 'table'
coordinate_systems: 'global'

Data structure:

images

NameDescription
imageThe raw image data.
image_3D(Optional) The raw 3D image data.
he_image(Optional) H&E image data.

labels

NameDescription
cell_labelsVendor-provided cell segmentation labels.
nucleus_labelsVendor-provided nucleus segmentation labels.
groundtruth_cell_labels(Optional) Manually annotated cell segmentation labels used as ground truth for evaluation.

points

transcripts: Point cloud data of transcripts.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
cell_idinteger(Optional) Unique identifier of the cell.
nucleus_idinteger(Optional) Unique identifier of the nucleus.
cell_typestring(Optional) Cell type of the cell.
qvfloat(Optional) Quality value of the point.
transcript_idlongUnique identifier of the transcript.
overlaps_nucleusboolean(Optional) Whether the point overlaps with a nucleus.

shapes

cell_boundaries: Cell boundaries.

ColumnTypeDescription
geometryobjectGeometry of the cell boundary.

nucleus_boundaries: Nucleus boundaries.

ColumnTypeDescription
geometryobjectGeometry of the nucleus boundary.

tables

table: Metadata of spatial dataset.

SlotTypeDescription
obs["cell_id"]stringA unique identifier for the cell.
obs["groundtruth_cell_type"]string(Optional) Manually curated cell type annotations which serves as ground truth for evaluations.
var["gene_ids"]stringUnique identifier for the gene.
var["feature_types"]stringType of the feature.
obsm["spatial"]doubleSpatial coordinates of the cell.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]stringLink to the original source of the dataset.
uns["dataset_reference"]stringBibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]stringThe organism of the sample in the dataset.
uns["segmentation_id"]stringA unique identifier for the segmentation.

coordinate_systems

NameDescription
globalCoordinate system of the replicate.

Component type: Data processor

A data processor.

Arguments:

NameTypeDescription
--input_spfileAn unprocessed spatial imaging dataset stored as a zarr file.
--input_scfileAn unprocessed dataset as output by a dataset loader.
--output_spatial_unlabelledfile(Output) Preprocessed spatial transcriptomics data without segmentation labels for method input.
--output_spatial_solutionfile(Output) Ground truth segmentation labels and cell assignments for method evaluation.
--output_scrnaseq_referencefile(Output) A single-cell reference dataset, preprocessed for this benchmark.

File format: Unlabelled

Preprocessed spatial transcriptomics data without segmentation labels for method input.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_unlabelled.zarr

Description:

This dataset contains preprocessed images and transcript point clouds for spatial transcriptomics data. Ground truth segmentation labels are intentionally excluded to prevent methods from cheating.

Format:

SpatialData object
images: 'image'
labels: 'cell_labels', 'nucleus_labels'
points: 'transcripts'
tables: 'table'
coordinate_systems: 'global'

Data structure:

images

NameDescription
imageThe raw image data.

labels

NameDescription
cell_labels(Optional) Vendor-provided cell segmentation labels, exposed as a segmentation prior.
nucleus_labels(Optional) Vendor-provided nucleus segmentation labels, exposed as a segmentation prior.

points

transcripts: Point cloud data of transcripts.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
qvfloat(Optional) Quality value of the point.
transcript_idlongUnique identifier of the transcript.
overlaps_nucleusboolean(Optional) Whether the point overlaps with the nucleus (derived from morphology).
cell_idinteger(Optional) Vendor-provided cell assignment from the raw data, exposed as a segmentation prior. This is NOT the ground truth used for evaluation (which is held out in spatial_solution); methods may freely condition on it.

tables

table: Metadata of spatial dataset.

SlotTypeDescription
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]stringLink to the original source of the dataset.
uns["dataset_reference"]stringBibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]stringThe organism of the sample in the dataset.
uns["orig_dataset_id"]stringThe identifier of the original dataset from which this dataset was derived (if applicable).

coordinate_systems

NameDescription
globalCoordinate system of the replicate.

File format: Solution

Ground truth segmentation labels and cell assignments for method evaluation.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_solution.zarr

Description:

This dataset contains the ground truth cell and nucleus segmentation labels, cell boundaries, and a reference table matching each cell to its label region.

Format:

SpatialData object
labels: 'cell_labels', 'nucleus_labels', 'groundtruth_cell_labels'
points: 'transcripts'
shapes: 'cell_boundaries', 'nucleus_boundaries'
tables: 'table'

Data structure:

labels

NameDescription
cell_labelsVendor-provided cell segmentation labels.
nucleus_labelsVendor-provided nucleus segmentation labels.
groundtruth_cell_labels(Optional) Manually annotated cell segmentation labels used as ground truth for evaluation.

points

transcripts: Point cloud data of transcripts with ground truth cell assignments.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
cell_idintegerGround truth cell assignment (0 = background).
transcript_idlongUnique identifier of the transcript.

shapes

cell_boundaries: Ground truth cell boundary shapes.

ColumnTypeDescription
geometryobjectGeometry of the cell boundary.

nucleus_boundaries: Ground truth nucleus boundary shapes.

ColumnTypeDescription
geometryobjectGeometry of the nucleus boundary.

tables

table: Reference cell metadata table.

SlotTypeDescription
obs["cell_id"]integerUnique cell identifier, matching instance IDs in the label images.
obs["region"]stringName of the label image this cell belongs to (e.g. ‘cell_labels’).
obs["cell_area"]double(Optional) Area of the cell in pixels.
obs["transcript_counts"]integer(Optional) Total number of transcripts assigned to this cell.
obs["groundtruth_cell_type"]string(Optional) Manually curated cell type annotations which serves as ground truth for evaluations.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["orig_dataset_id"]stringThe identifier of the original dataset from which this dataset was derived (if applicable).

File format: scRNA-seq Reference

A single-cell reference dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/scrnaseq_reference.h5ad

Description:

This dataset contains preprocessed counts and metadata for single-cell RNA-seq data.

Format:

AnnData object
obs: 'cell_type'
var: 'feature_id', 'feature_name', 'hvg'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized', 'normalized_log', 'normalized_log_scaled'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]doubleNormalized expression values.
layers["normalized_log"]doubleLog1p normalized expression values.
layers["normalized_log_scaled"]doubleLog1p normalized expression values scaled to unit variance and zero mean.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

Component type: Control Method

Quality control methods for verifying the pipeline.

Arguments:

NameTypeDescription
--inputfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--input_solutionfileGround truth segmentation labels and cell assignments for method evaluation.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Method

A method.

Arguments:

NameTypeDescription
--inputfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Output processor

An output processor for the prediction.

Arguments:

NameTypeDescription
--input_predictionfileA predicted dataset as output by a method.
--input_spatial_unlabelledfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--outputfile(Output) A processed predicted dataset, ready to be used as input for the evaluation.

Component type: Metric

A task template metric.

Arguments:

NameTypeDescription
--input_predictionfileA processed predicted dataset, ready to be used as input for the evaluation.
--input_solutionfileGround truth segmentation labels and cell assignments for method evaluation.
--outputfile(Output) File indicating the score of a metric.

File format: Predicted data

A predicted dataset as output by a method.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/prediction.zarr

Format:

SpatialData object
labels: 'segmentation'
tables: 'table'

Data structure:

labels

NameDescription
segmentationSegmentation of the data.

tables

table: AnnData table.

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["method_id"]stringA unique identifier for the method.

File format: Processed prediction

A processed predicted dataset, ready to be used as input for the evaluation.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/processed_prediction.zarr

Format:

SpatialData object
labels: 'segmentation'
tables: 'table'

Data structure:

labels

NameDescription
segmentationSegmentation of the data.

tables

table: AnnData table.

SlotTypeDescription
obs["cell_id"]stringCell ID.
obs["region"]stringRegion.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
layers["counts"]integerRaw counts.
layers["normalized"]doubleNormalized expression values.
layers["normalized_log"]doubleLog1p normalized expression values.
layers["normalized_log_scaled"]doubleLog1p normalized expression values scaled to unit variance and zero mean.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["method_id"]stringA unique identifier for the method.

File format: Score

File indicating the score of a metric.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/score.h5ad

Format:

AnnData object
uns: 'dataset_id', 'normalization_id', 'method_id', 'metric_ids', 'metric_values'

Data structure:

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["normalization_id"]stringWhich normalization was used.
uns["method_id"]stringA unique identifier for the method.
uns["metric_ids"]stringOne or more unique metric identifiers.
uns["metric_values"]doubleThe metric values obtained for the given prediction. Must be of same length as ‘metric_ids’.

File format: Common SC Dataset

An unprocessed dataset as output by a dataset loader.

Example file: resources_test/common/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

Description:

This dataset contains raw counts and metadata as output by a dataset loader.

The format of this file is mainly derived from the CELLxGENE schema v4.0.0.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

About

No description, website, or topics provided.

Resources

Contributing

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Spatial Segmentation

A one sentence summary of purpose and methodology. Used for creating an overview tables.

Repository: openproblems-bio/task_spatial_segmentation

Description

Provide a clear and concise description of your task, detailing the specific problem it aims to solve. Outline the input data types, the expected output, and any assumptions or constraints. Be sure to explain any terminology or concepts that are essential for understanding the task.

Explain the motivation behind your proposed task. Describe the biological or computational problem you aim to address and why it’s important. Discuss the current state of research in this area and any gaps or challenges that your task could help address. This section should convince readers of the significance and relevance of your task.

Authors & contributors

NameRolesOrcidGithub
Daria Romanovskaiamaintainer, author0000-0003-2831-0919dariarom94
Florian Heylmaintainer, author0000-0002-3651-5685heylf
Robrecht Cannoodtauthor0000-0003-3641-729Xrcannood

API

flowchart TB
file_common_ist("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-ist-dataset'>Common iST Dataset</a>")
comp_data_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-data-processor'>Data processor</a>"/]
file_spatial_unlabelled("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-unlabelled'>Unlabelled</a>")
file_spatial_solution("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-solution'>Solution</a>")
file_scrnaseq_reference("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-scrna-seq-reference'>scRNA-seq Reference</a>")
comp_control_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-control-method'>Control Method</a>"/]
comp_method[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-method'>Method</a>"/]
comp_output_processor[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-output-processor'>Output processor</a>"/]
comp_metric[/"<a href='https://github.com/openproblems-bio/task_spatial_segmentation#component-type-metric'>Metric</a>"/]
file_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-predicted-data'>Predicted data</a>")
file_processed_prediction("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-processed-prediction'>Processed prediction</a>")
file_score("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-score'>Score</a>")
file_common_scrnaseq("<a href='https://github.com/openproblems-bio/task_spatial_segmentation#file-format-common-sc-dataset'>Common SC Dataset</a>")
file_common_ist---comp_data_processor
comp_data_processor-->file_spatial_unlabelled
comp_data_processor-->file_spatial_solution
comp_data_processor-->file_scrnaseq_reference
file_spatial_unlabelled---comp_control_method
file_spatial_unlabelled---comp_method
file_spatial_unlabelled---comp_output_processor
file_spatial_solution---comp_control_method
file_spatial_solution---comp_metric
comp_control_method-->file_prediction
comp_method-->file_prediction
comp_output_processor-->file_processed_prediction
comp_metric-->file_score
file_prediction---comp_output_processor
file_processed_prediction---comp_metric
file_common_scrnaseq---comp_data_processor
Loading

File format: Common iST Dataset

An unprocessed spatial imaging dataset stored as a zarr file.

Example file: resources_test/common/2023_10x_mouse_brain_xenium_rep1/dataset.zarr

Description:

This dataset contains raw images, labels, points, shapes, and tables as output by a dataset loader.

Format:

SpatialData object
images: 'image', 'image_3D', 'he_image'
labels: 'cell_labels', 'nucleus_labels', 'groundtruth_cell_labels'
points: 'transcripts'
shapes: 'cell_boundaries', 'nucleus_boundaries'
tables: 'table'
coordinate_systems: 'global'

Data structure:

images

NameDescription
imageThe raw image data.
image_3D(Optional) The raw 3D image data.
he_image(Optional) H&E image data.

labels

NameDescription
cell_labelsVendor-provided cell segmentation labels.
nucleus_labelsVendor-provided nucleus segmentation labels.
groundtruth_cell_labels(Optional) Manually annotated cell segmentation labels used as ground truth for evaluation.

points

transcripts: Point cloud data of transcripts.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
cell_idinteger(Optional) Unique identifier of the cell.
nucleus_idinteger(Optional) Unique identifier of the nucleus.
cell_typestring(Optional) Cell type of the cell.
qvfloat(Optional) Quality value of the point.
transcript_idlongUnique identifier of the transcript.
overlaps_nucleusboolean(Optional) Whether the point overlaps with a nucleus.

shapes

cell_boundaries: Cell boundaries.

ColumnTypeDescription
geometryobjectGeometry of the cell boundary.

nucleus_boundaries: Nucleus boundaries.

ColumnTypeDescription
geometryobjectGeometry of the nucleus boundary.

tables

table: Metadata of spatial dataset.

SlotTypeDescription
obs["cell_id"]stringA unique identifier for the cell.
obs["groundtruth_cell_type"]string(Optional) Manually curated cell type annotations which serves as ground truth for evaluations.
var["gene_ids"]stringUnique identifier for the gene.
var["feature_types"]stringType of the feature.
obsm["spatial"]doubleSpatial coordinates of the cell.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]stringLink to the original source of the dataset.
uns["dataset_reference"]stringBibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]stringThe organism of the sample in the dataset.
uns["segmentation_id"]stringA unique identifier for the segmentation.

coordinate_systems

NameDescription
globalCoordinate system of the replicate.

Component type: Data processor

A data processor.

Arguments:

NameTypeDescription
--input_spfileAn unprocessed spatial imaging dataset stored as a zarr file.
--input_scfileAn unprocessed dataset as output by a dataset loader.
--output_spatial_unlabelledfile(Output) Preprocessed spatial transcriptomics data without segmentation labels for method input.
--output_spatial_solutionfile(Output) Ground truth segmentation labels and cell assignments for method evaluation.
--output_scrnaseq_referencefile(Output) A single-cell reference dataset, preprocessed for this benchmark.

File format: Unlabelled

Preprocessed spatial transcriptomics data without segmentation labels for method input.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_unlabelled.zarr

Description:

This dataset contains preprocessed images and transcript point clouds for spatial transcriptomics data. Ground truth segmentation labels are intentionally excluded to prevent methods from cheating.

Format:

SpatialData object
images: 'image'
labels: 'cell_labels', 'nucleus_labels'
points: 'transcripts'
tables: 'table'
coordinate_systems: 'global'

Data structure:

images

NameDescription
imageThe raw image data.

labels

NameDescription
cell_labels(Optional) Vendor-provided cell segmentation labels, exposed as a segmentation prior.
nucleus_labels(Optional) Vendor-provided nucleus segmentation labels, exposed as a segmentation prior.

points

transcripts: Point cloud data of transcripts.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
qvfloat(Optional) Quality value of the point.
transcript_idlongUnique identifier of the transcript.
overlaps_nucleusboolean(Optional) Whether the point overlaps with the nucleus (derived from morphology).
cell_idinteger(Optional) Vendor-provided cell assignment from the raw data, exposed as a segmentation prior. This is NOT the ground truth used for evaluation (which is held out in spatial_solution); methods may freely condition on it.

tables

table: Metadata of spatial dataset.

SlotTypeDescription
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]stringLink to the original source of the dataset.
uns["dataset_reference"]stringBibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]stringThe organism of the sample in the dataset.
uns["orig_dataset_id"]stringThe identifier of the original dataset from which this dataset was derived (if applicable).

coordinate_systems

NameDescription
globalCoordinate system of the replicate.

File format: Solution

Ground truth segmentation labels and cell assignments for method evaluation.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/spatial_solution.zarr

Description:

This dataset contains the ground truth cell and nucleus segmentation labels, cell boundaries, and a reference table matching each cell to its label region.

Format:

SpatialData object
labels: 'cell_labels', 'nucleus_labels', 'groundtruth_cell_labels'
points: 'transcripts'
shapes: 'cell_boundaries', 'nucleus_boundaries'
tables: 'table'

Data structure:

labels

NameDescription
cell_labelsVendor-provided cell segmentation labels.
nucleus_labelsVendor-provided nucleus segmentation labels.
groundtruth_cell_labels(Optional) Manually annotated cell segmentation labels used as ground truth for evaluation.

points

transcripts: Point cloud data of transcripts with ground truth cell assignments.

ColumnTypeDescription
xfloatx-coordinate of the point.
yfloaty-coordinate of the point.
zfloat(Optional) z-coordinate of the point.
feature_namecategoricalName of the feature.
cell_idintegerGround truth cell assignment (0 = background).
transcript_idlongUnique identifier of the transcript.

shapes

cell_boundaries: Ground truth cell boundary shapes.

ColumnTypeDescription
geometryobjectGeometry of the cell boundary.

nucleus_boundaries: Ground truth nucleus boundary shapes.

ColumnTypeDescription
geometryobjectGeometry of the nucleus boundary.

tables

table: Reference cell metadata table.

SlotTypeDescription
obs["cell_id"]integerUnique cell identifier, matching instance IDs in the label images.
obs["region"]stringName of the label image this cell belongs to (e.g. ‘cell_labels’).
obs["cell_area"]double(Optional) Area of the cell in pixels.
obs["transcript_counts"]integer(Optional) Total number of transcripts assigned to this cell.
obs["groundtruth_cell_type"]string(Optional) Manually curated cell type annotations which serves as ground truth for evaluations.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["orig_dataset_id"]stringThe identifier of the original dataset from which this dataset was derived (if applicable).

File format: scRNA-seq Reference

A single-cell reference dataset, preprocessed for this benchmark.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/scrnaseq_reference.h5ad

Description:

This dataset contains preprocessed counts and metadata for single-cell RNA-seq data.

Format:

AnnData object
obs: 'cell_type'
var: 'feature_id', 'feature_name', 'hvg'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized', 'normalized_log', 'normalized_log_scaled'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]doubleNormalized expression values.
layers["normalized_log"]doubleLog1p normalized expression values.
layers["normalized_log_scaled"]doubleLog1p normalized expression values scaled to unit variance and zero mean.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

Component type: Control Method

Quality control methods for verifying the pipeline.

Arguments:

NameTypeDescription
--inputfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--input_solutionfileGround truth segmentation labels and cell assignments for method evaluation.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Method

A method.

Arguments:

NameTypeDescription
--inputfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--outputfile(Output) A predicted dataset as output by a method.

Component type: Output processor

An output processor for the prediction.

Arguments:

NameTypeDescription
--input_predictionfileA predicted dataset as output by a method.
--input_spatial_unlabelledfilePreprocessed spatial transcriptomics data without segmentation labels for method input.
--outputfile(Output) A processed predicted dataset, ready to be used as input for the evaluation.

Component type: Metric

A task template metric.

Arguments:

NameTypeDescription
--input_predictionfileA processed predicted dataset, ready to be used as input for the evaluation.
--input_solutionfileGround truth segmentation labels and cell assignments for method evaluation.
--outputfile(Output) File indicating the score of a metric.

File format: Predicted data

A predicted dataset as output by a method.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/prediction.zarr

Format:

SpatialData object
labels: 'segmentation'
tables: 'table'

Data structure:

labels

NameDescription
segmentationSegmentation of the data.

tables

table: AnnData table.

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["method_id"]stringA unique identifier for the method.

File format: Processed prediction

A processed predicted dataset, ready to be used as input for the evaluation.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/processed_prediction.zarr

Format:

SpatialData object
labels: 'segmentation'
tables: 'table'

Data structure:

labels

NameDescription
segmentationSegmentation of the data.

tables

table: AnnData table.

SlotTypeDescription
obs["cell_id"]stringCell ID.
obs["region"]stringRegion.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
layers["counts"]integerRaw counts.
layers["normalized"]doubleNormalized expression values.
layers["normalized_log"]doubleLog1p normalized expression values.
layers["normalized_log_scaled"]doubleLog1p normalized expression values scaled to unit variance and zero mean.
uns["dataset_id"]stringA unique identifier for the dataset.
uns["method_id"]stringA unique identifier for the method.

File format: Score

File indicating the score of a metric.

Example file: resources_test/task_spatial_segmentation/mouse_brain_combined/score.h5ad

Format:

AnnData object
uns: 'dataset_id', 'normalization_id', 'method_id', 'metric_ids', 'metric_values'

Data structure:

SlotTypeDescription
uns["dataset_id"]stringA unique identifier for the dataset.
uns["normalization_id"]stringWhich normalization was used.
uns["method_id"]stringA unique identifier for the method.
uns["metric_ids"]stringOne or more unique metric identifiers.
uns["metric_values"]doubleThe metric values obtained for the given prediction. Must be of same length as ‘metric_ids’.

File format: Common SC Dataset

An unprocessed dataset as output by a dataset loader.

Example file: resources_test/common/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

Description:

This dataset contains raw counts and metadata as output by a dataset loader.

The format of this file is mainly derived from the CELLxGENE schema v4.0.0.

Format:

AnnData object
obs: 'cell_type', 'cell_type_level2', 'cell_type_level3', 'cell_type_level4', 'dataset_id', 'assay', 'assay_ontology_term_id', 'cell_type_ontology_term_id', 'development_stage', 'development_stage_ontology_term_id', 'disease', 'disease_ontology_term_id', 'donor_id', 'is_primary_data', 'organism', 'organism_ontology_term_id', 'self_reported_ethnicity', 'self_reported_ethnicity_ontology_term_id', 'sex', 'sex_ontology_term_id', 'suspension_type', 'tissue', 'tissue_ontology_term_id', 'tissue_general', 'tissue_general_ontology_term_id', 'batch', 'soma_joinid'
var: 'feature_id', 'feature_name', 'soma_joinid', 'hvg', 'hvg_score'
obsm: 'X_pca'
obsp: 'knn_distances', 'knn_connectivities'
varm: 'pca_loadings'
layers: 'counts', 'normalized'
uns: 'dataset_id', 'dataset_name', 'dataset_url', 'dataset_reference', 'dataset_summary', 'dataset_description', 'dataset_organism'

Data structure:

SlotTypeDescription
obs["cell_type"]stringClassification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level2"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level3"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["cell_type_level4"]string(Optional) Classification of the cell type based on its characteristics and function within the tissue or organism.
obs["dataset_id"]string(Optional) Identifier for the dataset from which the cell data is derived, useful for tracking and referencing purposes.
obs["assay"]string(Optional) Type of assay used to generate the cell data, indicating the methodology or technique employed.
obs["assay_ontology_term_id"]string(Optional) Experimental Factor Ontology (EFO:) term identifier for the assay, providing a standardized reference to the assay type.
obs["cell_type_ontology_term_id"]string(Optional) Cell Ontology (CL:) term identifier for the cell type, offering a standardized reference to the specific cell classification.
obs["development_stage"]string(Optional) Stage of development of the organism or tissue from which the cell is derived, indicating its maturity or developmental phase.
obs["development_stage_ontology_term_id"]string(Optional) Ontology term identifier for the developmental stage, providing a standardized reference to the organism’s developmental phase. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Developmental Stages (HsapDv:) ontology is used. If the organism is mouse (organism_ontology_term_id == 'NCBITaxon:10090'), then the Mouse Developmental Stages (MmusDv:) ontology is used. Otherwise, the Uberon (UBERON:) ontology is used.
obs["disease"]string(Optional) Information on any disease or pathological condition associated with the cell or donor.
obs["disease_ontology_term_id"]string(Optional) Ontology term identifier for the disease, enabling standardized disease classification and referencing. Must be a term from the Mondo Disease Ontology (MONDO:) ontology term, or PATO:0000461 from the Phenotype And Trait Ontology (PATO:).
obs["donor_id"]string(Optional) Identifier for the donor from whom the cell sample is obtained.
obs["is_primary_data"]boolean(Optional) Indicates whether the data is primary (directly obtained from experiments) or has been computationally derived from other primary data.
obs["organism"]string(Optional) Organism from which the cell sample is obtained.
obs["organism_ontology_term_id"]string(Optional) Ontology term identifier for the organism, providing a standardized reference for the organism. Must be a term from the NCBI Taxonomy Ontology (NCBITaxon:) which is a child of NCBITaxon:33208.
obs["self_reported_ethnicity"]string(Optional) Ethnicity of the donor as self-reported, relevant for studies considering genetic diversity and population-specific traits.
obs["self_reported_ethnicity_ontology_term_id"]string(Optional) Ontology term identifier for the self-reported ethnicity, providing a standardized reference for ethnic classifications. If the organism is human (organism_ontology_term_id == 'NCBITaxon:9606'), then the Human Ancestry Ontology (HANCESTRO:) is used.
obs["sex"]string(Optional) Biological sex of the donor or source organism, crucial for studies involving sex-specific traits or conditions.
obs["sex_ontology_term_id"]string(Optional) Ontology term identifier for the biological sex, ensuring standardized classification of sex. Only PATO:0000383, PATO:0000384 and PATO:0001340 are allowed.
obs["suspension_type"]string(Optional) Type of suspension or medium in which the cells were stored or processed, important for understanding cell handling and conditions.
obs["tissue"]string(Optional) Specific tissue from which the cells were derived, key for context and specificity in cell studies.
obs["tissue_ontology_term_id"]string(Optional) Ontology term identifier for the tissue, providing a standardized reference for the tissue type. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["tissue_general"]string(Optional) General category or classification of the tissue, useful for broader grouping and comparison of cell data.
obs["tissue_general_ontology_term_id"]string(Optional) Ontology term identifier for the general tissue category, aiding in standardizing and grouping tissue types. For organoid or tissue samples, the Uber-anatomy ontology (UBERON:) is used. The term ids must be a child term of UBERON:0001062 (anatomical entity). For cell cultures, the Cell Ontology (CL:) is used. The term ids cannot be CL:0000255, CL:0000257 or CL:0000548.
obs["batch"]string(Optional) A batch identifier. This label is very context-dependent and may be a combination of the tissue, assay, donor, etc.
obs["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the cell.
var["feature_id"]string(Optional) Unique identifier for the feature, usually a ENSEMBL gene id.
var["feature_name"]stringA human-readable name for the feature, usually a gene symbol.
var["soma_joinid"]integer(Optional) If the dataset was retrieved from CELLxGENE census, this is a unique identifier for the feature.
var["hvg"]booleanWhether or not the feature is considered to be a ‘highly variable gene’.
var["hvg_score"]doubleA score for the feature indicating how highly variable it is.
obsm["X_pca"]doubleThe resulting PCA embedding.
obsp["knn_distances"]doubleK nearest neighbors distance matrix.
obsp["knn_connectivities"]doubleK nearest neighbors connectivities matrix.
varm["pca_loadings"]doubleThe PCA loadings matrix.
layers["counts"]integerRaw counts.
layers["normalized"]integerNormalized expression values.
uns["dataset_id"]stringA unique identifier for the dataset. This is different from the obs.dataset_id field, which is the identifier for the dataset from which the cell data is derived.
uns["dataset_name"]stringA human-readable name for the dataset.
uns["dataset_url"]string(Optional) Link to the original source of the dataset.
uns["dataset_reference"]string(Optional) Bibtex reference of the paper in which the dataset was published.
uns["dataset_summary"]stringShort description of the dataset.
uns["dataset_description"]stringLong description of the dataset.
uns["dataset_organism"]string(Optional) The organism of the sample in the dataset.

About

No description, website, or topics provided.

Resources

Contributing

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages