Skip to content

Add method to calculate embeddings for variable by distance aggregation - #807

Open
LLehner wants to merge 39 commits into
mainfrom
var_by_distance_clustering
Open

Add method to calculate embeddings for variable by distance aggregation#807
LLehner wants to merge 39 commits into
mainfrom
var_by_distance_clustering

Conversation

@LLehner

@LLehnerLLehner commented Mar 4, 2024

Copy link
Copy Markdown
Member

Description

Adds a method in tools to calculate embeddings of variables by their counts aggregated by distance.

Example usage

import squidpy as sq

load example data set
adata = sq.datasets.seqfish()

Calculate distances of each observation to a specified anchor point (e.g. cell type or tissue location). Here we use cell type "Endothelium" in the annotation column "celltype_mapped_refined":
sq.tl.var_by_distance(adata, groups="Endothelium", cluster_key="celltype_mapped_refined")

The resulting distances are stored in adata.obsm["design_matrix"]. Now we can calculate the embeddings, which are returned as a new anndata object:
adata_new = sq.tl.var_embeddings(adata, group="Endothelium", design_matrix_key="design_matrix")

Note that by default the bin of distance 0, meaning the counts that belong to the anchor point, are excluded. This can be changed by setting include_anchor=True in sq.tl.var_embeddings().

adata_new.X contains the aggregated var x distance_bin count matrix.
adata_new.obs contains the variables as a categorical matrix, which is required to highlight them in plots.

TODO

  • Add a plotting function so this doesn't need to be done manually.
  • Allow flexible embedding calculations

@LLehner
LLehner requested a review from timtreisMarch 4, 2024 22:56
@codecov-commenter

codecov-commenter commented Mar 4, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 33.33333% with 24 lines in your changes are missing coverage. Please review.

Project coverage is 69.75%. Comparing base (df8e042) to head (8ee07ba).

Additional details and impacted files
@@ Coverage Diff @@## main #807 +/- ##
==========================================
- Coverage 69.99% 69.75% -0.24% 
==========================================
Files 39 40 +1 Lines 5525 5561 +36 Branches 1029 1037 +8 ==========================================
+ Hits 3867 3879 +12 - Misses 1363 1387 +24 
Partials 295 295 
FilesCoverage Δ
src/squidpy/tl/_var_embeddings.py33.33% <33.33%> (ø)

@giovp

Copy link
Copy Markdown
Member

hi @LLehner , thank you for this, would you mind elaborating a bit when this would be used? also, what if the embedding are pre-calculated, or the user would like to use something other than the UMAP, should that be an option? finally, I think a test would be required before we get this in, thanks!

@timtreis

Copy link
Copy Markdown
Member

Hey @giovp, this feature was coming out of a discussion with @maiiashulman. We ran into a situation in which the "literature-curated" signature for hypoxia was either 20 or 4000 genes, the latter obviously being useless. So we wondered which other genes maybe show the same spatially variable pattern as a function of distance to a certain cell-type (e.g. epithelial). This is essentially a graphical method to see if a given set of genes (f.e. the 20 gene signature) even varies in a similar pattern.

But I agree with your points; if we see that it's actually doing something useful, we should make it a bit more flexible.

@LLehner
LLehner marked this pull request as draft April 22, 2024 22:05
@timtreis
timtreis marked this pull request as ready for review July 9, 2024 21:02
@timtreistimtreis added squidpy2.0 Everything releated to a Squidpy 2.0 release feature PR introduces a new feature labels Jul 9, 2024
@LLehner
LLehner marked this pull request as draft August 8, 2024 10:22
@LLehner

LLehner commented Aug 8, 2024

Copy link
Copy Markdown
MemberAuthor

@timtreis this function now returns an anndata object, which is i think simplifies further processing, compared to storing the new count matrix somewhere in .varm or .uns. Because if we want to make us of already implemented dimreduction and clustering methods from scanpy, then the count matrix needs to be in .X and for visualization we need the variable names stored as categories in .obs. Doing all of this in the same anndata will just make things cluttered.

Additionally the question is whether a spatialdata object should be required as input instead of an anndataone, because then a new table could be added directly instead of having multiple disconnected tables.

The function call would change from:
adata_new = sq.tl.var_embeddings(adata, group="Endothelium", design_matrix_key="design_matrix")
to
sq.tl.var_embeddings(sdata, group="Endothelium", design_matrix_key="design_matrix")

@LLehner
LLehner marked this pull request as ready for review October 10, 2024 13:06
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

featurePR introduces a new featuresquidpy2.0Everything releated to a Squidpy 2.0 release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@LLehner@codecov-commenter@giovp@timtreis
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Add method to calculate embeddings for variable by distance aggregation by LLehner · Pull Request #807 · scverse/squidpy · GitHub
Skip to content

Add method to calculate embeddings for variable by distance aggregation - #807

Open
LLehner wants to merge 39 commits into
mainfrom
var_by_distance_clustering
Open

Add method to calculate embeddings for variable by distance aggregation#807
LLehner wants to merge 39 commits into
mainfrom
var_by_distance_clustering

Conversation

@LLehner

@LLehnerLLehner commented Mar 4, 2024

Copy link
Copy Markdown
Member

Description

Adds a method in tools to calculate embeddings of variables by their counts aggregated by distance.

Example usage

import squidpy as sq

load example data set
adata = sq.datasets.seqfish()

Calculate distances of each observation to a specified anchor point (e.g. cell type or tissue location). Here we use cell type "Endothelium" in the annotation column "celltype_mapped_refined":
sq.tl.var_by_distance(adata, groups="Endothelium", cluster_key="celltype_mapped_refined")

The resulting distances are stored in adata.obsm["design_matrix"]. Now we can calculate the embeddings, which are returned as a new anndata object:
adata_new = sq.tl.var_embeddings(adata, group="Endothelium", design_matrix_key="design_matrix")

Note that by default the bin of distance 0, meaning the counts that belong to the anchor point, are excluded. This can be changed by setting include_anchor=True in sq.tl.var_embeddings().

adata_new.X contains the aggregated var x distance_bin count matrix.
adata_new.obs contains the variables as a categorical matrix, which is required to highlight them in plots.

TODO

  • Add a plotting function so this doesn't need to be done manually.
  • Allow flexible embedding calculations

@LLehner
LLehner requested a review from timtreisMarch 4, 2024 22:56
@codecov-commenter

codecov-commenter commented Mar 4, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 33.33333% with 24 lines in your changes are missing coverage. Please review.

Project coverage is 69.75%. Comparing base (df8e042) to head (8ee07ba).

Additional details and impacted files
@@ Coverage Diff @@## main #807 +/- ##
==========================================
- Coverage 69.99% 69.75% -0.24% 
==========================================
Files 39 40 +1 Lines 5525 5561 +36 Branches 1029 1037 +8 ==========================================
+ Hits 3867 3879 +12 - Misses 1363 1387 +24 
Partials 295 295 
FilesCoverage Δ
src/squidpy/tl/_var_embeddings.py33.33% <33.33%> (ø)

@giovp

Copy link
Copy Markdown
Member

hi @LLehner , thank you for this, would you mind elaborating a bit when this would be used? also, what if the embedding are pre-calculated, or the user would like to use something other than the UMAP, should that be an option? finally, I think a test would be required before we get this in, thanks!

@timtreis

Copy link
Copy Markdown
Member

Hey @giovp, this feature was coming out of a discussion with @maiiashulman. We ran into a situation in which the "literature-curated" signature for hypoxia was either 20 or 4000 genes, the latter obviously being useless. So we wondered which other genes maybe show the same spatially variable pattern as a function of distance to a certain cell-type (e.g. epithelial). This is essentially a graphical method to see if a given set of genes (f.e. the 20 gene signature) even varies in a similar pattern.

But I agree with your points; if we see that it's actually doing something useful, we should make it a bit more flexible.

@LLehner
LLehner marked this pull request as draft April 22, 2024 22:05
@timtreis
timtreis marked this pull request as ready for review July 9, 2024 21:02
@timtreistimtreis added squidpy2.0 Everything releated to a Squidpy 2.0 release feature PR introduces a new feature labels Jul 9, 2024
@LLehner
LLehner marked this pull request as draft August 8, 2024 10:22
@LLehner

LLehner commented Aug 8, 2024

Copy link
Copy Markdown
MemberAuthor

@timtreis this function now returns an anndata object, which is i think simplifies further processing, compared to storing the new count matrix somewhere in .varm or .uns. Because if we want to make us of already implemented dimreduction and clustering methods from scanpy, then the count matrix needs to be in .X and for visualization we need the variable names stored as categories in .obs. Doing all of this in the same anndata will just make things cluttered.

Additionally the question is whether a spatialdata object should be required as input instead of an anndataone, because then a new table could be added directly instead of having multiple disconnected tables.

The function call would change from:
adata_new = sq.tl.var_embeddings(adata, group="Endothelium", design_matrix_key="design_matrix")
to
sq.tl.var_embeddings(sdata, group="Endothelium", design_matrix_key="design_matrix")

@LLehner
LLehner marked this pull request as ready for review October 10, 2024 13:06
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

featurePR introduces a new featuresquidpy2.0Everything releated to a Squidpy 2.0 release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@LLehner@codecov-commenter@giovp@timtreis
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add method to calculate embeddings for variable by distance aggregation by LLehner · Pull Request #807 · scverse/squidpy · GitHub
Skip to content

Add method to calculate embeddings for variable by distance aggregation - #807

Open
LLehner wants to merge 39 commits into
mainfrom
var_by_distance_clustering
Open

Add method to calculate embeddings for variable by distance aggregation#807
LLehner wants to merge 39 commits into
mainfrom
var_by_distance_clustering

Conversation

@LLehner

@LLehnerLLehner commented Mar 4, 2024

Copy link
Copy Markdown
Member

Description

Adds a method in tools to calculate embeddings of variables by their counts aggregated by distance.

Example usage

import squidpy as sq

load example data set
adata = sq.datasets.seqfish()

Calculate distances of each observation to a specified anchor point (e.g. cell type or tissue location). Here we use cell type "Endothelium" in the annotation column "celltype_mapped_refined":
sq.tl.var_by_distance(adata, groups="Endothelium", cluster_key="celltype_mapped_refined")

The resulting distances are stored in adata.obsm["design_matrix"]. Now we can calculate the embeddings, which are returned as a new anndata object:
adata_new = sq.tl.var_embeddings(adata, group="Endothelium", design_matrix_key="design_matrix")

Note that by default the bin of distance 0, meaning the counts that belong to the anchor point, are excluded. This can be changed by setting include_anchor=True in sq.tl.var_embeddings().

adata_new.X contains the aggregated var x distance_bin count matrix.
adata_new.obs contains the variables as a categorical matrix, which is required to highlight them in plots.

TODO

  • Add a plotting function so this doesn't need to be done manually.
  • Allow flexible embedding calculations

@LLehner
LLehner requested a review from timtreisMarch 4, 2024 22:56
@codecov-commenter

codecov-commenter commented Mar 4, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 33.33333% with 24 lines in your changes are missing coverage. Please review.

Project coverage is 69.75%. Comparing base (df8e042) to head (8ee07ba).

Additional details and impacted files
@@ Coverage Diff @@## main #807 +/- ##
==========================================
- Coverage 69.99% 69.75% -0.24% 
==========================================
Files 39 40 +1 Lines 5525 5561 +36 Branches 1029 1037 +8 ==========================================
+ Hits 3867 3879 +12 - Misses 1363 1387 +24 
Partials 295 295 
FilesCoverage Δ
src/squidpy/tl/_var_embeddings.py33.33% <33.33%> (ø)

@giovp

Copy link
Copy Markdown
Member

hi @LLehner , thank you for this, would you mind elaborating a bit when this would be used? also, what if the embedding are pre-calculated, or the user would like to use something other than the UMAP, should that be an option? finally, I think a test would be required before we get this in, thanks!

@timtreis

Copy link
Copy Markdown
Member

Hey @giovp, this feature was coming out of a discussion with @maiiashulman. We ran into a situation in which the "literature-curated" signature for hypoxia was either 20 or 4000 genes, the latter obviously being useless. So we wondered which other genes maybe show the same spatially variable pattern as a function of distance to a certain cell-type (e.g. epithelial). This is essentially a graphical method to see if a given set of genes (f.e. the 20 gene signature) even varies in a similar pattern.

But I agree with your points; if we see that it's actually doing something useful, we should make it a bit more flexible.

@LLehner
LLehner marked this pull request as draft April 22, 2024 22:05
@timtreis
timtreis marked this pull request as ready for review July 9, 2024 21:02
@timtreistimtreis added squidpy2.0 Everything releated to a Squidpy 2.0 release feature PR introduces a new feature labels Jul 9, 2024
@LLehner
LLehner marked this pull request as draft August 8, 2024 10:22
@LLehner

LLehner commented Aug 8, 2024

Copy link
Copy Markdown
MemberAuthor

@timtreis this function now returns an anndata object, which is i think simplifies further processing, compared to storing the new count matrix somewhere in .varm or .uns. Because if we want to make us of already implemented dimreduction and clustering methods from scanpy, then the count matrix needs to be in .X and for visualization we need the variable names stored as categories in .obs. Doing all of this in the same anndata will just make things cluttered.

Additionally the question is whether a spatialdata object should be required as input instead of an anndataone, because then a new table could be added directly instead of having multiple disconnected tables.

The function call would change from:
adata_new = sq.tl.var_embeddings(adata, group="Endothelium", design_matrix_key="design_matrix")
to
sq.tl.var_embeddings(sdata, group="Endothelium", design_matrix_key="design_matrix")

@LLehner
LLehner marked this pull request as ready for review October 10, 2024 13:06
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

featurePR introduces a new featuresquidpy2.0Everything releated to a Squidpy 2.0 release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@LLehner@codecov-commenter@giovp@timtreis
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add method to calculate embeddings for variable by distance aggregation by LLehner · Pull Request #807 · scverse/squidpy · GitHub
Skip to content

Add method to calculate embeddings for variable by distance aggregation - #807

Open
LLehner wants to merge 39 commits into
mainfrom
var_by_distance_clustering
Open

Add method to calculate embeddings for variable by distance aggregation#807
LLehner wants to merge 39 commits into
mainfrom
var_by_distance_clustering

Conversation

@LLehner

@LLehnerLLehner commented Mar 4, 2024

Copy link
Copy Markdown
Member

Description

Adds a method in tools to calculate embeddings of variables by their counts aggregated by distance.

Example usage

import squidpy as sq

load example data set
adata = sq.datasets.seqfish()

Calculate distances of each observation to a specified anchor point (e.g. cell type or tissue location). Here we use cell type "Endothelium" in the annotation column "celltype_mapped_refined":
sq.tl.var_by_distance(adata, groups="Endothelium", cluster_key="celltype_mapped_refined")

The resulting distances are stored in adata.obsm["design_matrix"]. Now we can calculate the embeddings, which are returned as a new anndata object:
adata_new = sq.tl.var_embeddings(adata, group="Endothelium", design_matrix_key="design_matrix")

Note that by default the bin of distance 0, meaning the counts that belong to the anchor point, are excluded. This can be changed by setting include_anchor=True in sq.tl.var_embeddings().

adata_new.X contains the aggregated var x distance_bin count matrix.
adata_new.obs contains the variables as a categorical matrix, which is required to highlight them in plots.

TODO

  • Add a plotting function so this doesn't need to be done manually.
  • Allow flexible embedding calculations

@LLehner
LLehner requested a review from timtreisMarch 4, 2024 22:56
@codecov-commenter

codecov-commenter commented Mar 4, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 33.33333% with 24 lines in your changes are missing coverage. Please review.

Project coverage is 69.75%. Comparing base (df8e042) to head (8ee07ba).

Additional details and impacted files
@@ Coverage Diff @@## main #807 +/- ##
==========================================
- Coverage 69.99% 69.75% -0.24% 
==========================================
Files 39 40 +1 Lines 5525 5561 +36 Branches 1029 1037 +8 ==========================================
+ Hits 3867 3879 +12 - Misses 1363 1387 +24 
Partials 295 295 
FilesCoverage Δ
src/squidpy/tl/_var_embeddings.py33.33% <33.33%> (ø)

@giovp

Copy link
Copy Markdown
Member

hi @LLehner , thank you for this, would you mind elaborating a bit when this would be used? also, what if the embedding are pre-calculated, or the user would like to use something other than the UMAP, should that be an option? finally, I think a test would be required before we get this in, thanks!

@timtreis

Copy link
Copy Markdown
Member

Hey @giovp, this feature was coming out of a discussion with @maiiashulman. We ran into a situation in which the "literature-curated" signature for hypoxia was either 20 or 4000 genes, the latter obviously being useless. So we wondered which other genes maybe show the same spatially variable pattern as a function of distance to a certain cell-type (e.g. epithelial). This is essentially a graphical method to see if a given set of genes (f.e. the 20 gene signature) even varies in a similar pattern.

But I agree with your points; if we see that it's actually doing something useful, we should make it a bit more flexible.

@LLehner
LLehner marked this pull request as draft April 22, 2024 22:05
@timtreis
timtreis marked this pull request as ready for review July 9, 2024 21:02
@timtreistimtreis added squidpy2.0 Everything releated to a Squidpy 2.0 release feature PR introduces a new feature labels Jul 9, 2024
@LLehner
LLehner marked this pull request as draft August 8, 2024 10:22
@LLehner

LLehner commented Aug 8, 2024

Copy link
Copy Markdown
MemberAuthor

@timtreis this function now returns an anndata object, which is i think simplifies further processing, compared to storing the new count matrix somewhere in .varm or .uns. Because if we want to make us of already implemented dimreduction and clustering methods from scanpy, then the count matrix needs to be in .X and for visualization we need the variable names stored as categories in .obs. Doing all of this in the same anndata will just make things cluttered.

Additionally the question is whether a spatialdata object should be required as input instead of an anndataone, because then a new table could be added directly instead of having multiple disconnected tables.

The function call would change from:
adata_new = sq.tl.var_embeddings(adata, group="Endothelium", design_matrix_key="design_matrix")
to
sq.tl.var_embeddings(sdata, group="Endothelium", design_matrix_key="design_matrix")

@LLehner
LLehner marked this pull request as ready for review October 10, 2024 13:06
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

featurePR introduces a new featuresquidpy2.0Everything releated to a Squidpy 2.0 release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@LLehner@codecov-commenter@giovp@timtreis
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Add method to calculate embeddings for variable by distance aggregation by LLehner · Pull Request #807 · scverse/squidpy · GitHub
Skip to content

Add method to calculate embeddings for variable by distance aggregation - #807

Open
LLehner wants to merge 39 commits into
mainfrom
var_by_distance_clustering
Open

Add method to calculate embeddings for variable by distance aggregation#807
LLehner wants to merge 39 commits into
mainfrom
var_by_distance_clustering

Conversation

@LLehner

@LLehnerLLehner commented Mar 4, 2024

Copy link
Copy Markdown
Member

Description

Adds a method in tools to calculate embeddings of variables by their counts aggregated by distance.

Example usage

import squidpy as sq

load example data set
adata = sq.datasets.seqfish()

Calculate distances of each observation to a specified anchor point (e.g. cell type or tissue location). Here we use cell type "Endothelium" in the annotation column "celltype_mapped_refined":
sq.tl.var_by_distance(adata, groups="Endothelium", cluster_key="celltype_mapped_refined")

The resulting distances are stored in adata.obsm["design_matrix"]. Now we can calculate the embeddings, which are returned as a new anndata object:
adata_new = sq.tl.var_embeddings(adata, group="Endothelium", design_matrix_key="design_matrix")

Note that by default the bin of distance 0, meaning the counts that belong to the anchor point, are excluded. This can be changed by setting include_anchor=True in sq.tl.var_embeddings().

adata_new.X contains the aggregated var x distance_bin count matrix.
adata_new.obs contains the variables as a categorical matrix, which is required to highlight them in plots.

TODO

  • Add a plotting function so this doesn't need to be done manually.
  • Allow flexible embedding calculations

@LLehner
LLehner requested a review from timtreisMarch 4, 2024 22:56
@codecov-commenter

codecov-commenter commented Mar 4, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 33.33333% with 24 lines in your changes are missing coverage. Please review.

Project coverage is 69.75%. Comparing base (df8e042) to head (8ee07ba).

Additional details and impacted files
@@ Coverage Diff @@## main #807 +/- ##
==========================================
- Coverage 69.99% 69.75% -0.24% 
==========================================
Files 39 40 +1 Lines 5525 5561 +36 Branches 1029 1037 +8 ==========================================
+ Hits 3867 3879 +12 - Misses 1363 1387 +24 
Partials 295 295 
FilesCoverage Δ
src/squidpy/tl/_var_embeddings.py33.33% <33.33%> (ø)

@giovp

Copy link
Copy Markdown
Member

hi @LLehner , thank you for this, would you mind elaborating a bit when this would be used? also, what if the embedding are pre-calculated, or the user would like to use something other than the UMAP, should that be an option? finally, I think a test would be required before we get this in, thanks!

@timtreis

Copy link
Copy Markdown
Member

Hey @giovp, this feature was coming out of a discussion with @maiiashulman. We ran into a situation in which the "literature-curated" signature for hypoxia was either 20 or 4000 genes, the latter obviously being useless. So we wondered which other genes maybe show the same spatially variable pattern as a function of distance to a certain cell-type (e.g. epithelial). This is essentially a graphical method to see if a given set of genes (f.e. the 20 gene signature) even varies in a similar pattern.

But I agree with your points; if we see that it's actually doing something useful, we should make it a bit more flexible.

@LLehner
LLehner marked this pull request as draft April 22, 2024 22:05
@timtreis
timtreis marked this pull request as ready for review July 9, 2024 21:02
@timtreistimtreis added squidpy2.0 Everything releated to a Squidpy 2.0 release feature PR introduces a new feature labels Jul 9, 2024
@LLehner
LLehner marked this pull request as draft August 8, 2024 10:22
@LLehner

LLehner commented Aug 8, 2024

Copy link
Copy Markdown
MemberAuthor

@timtreis this function now returns an anndata object, which is i think simplifies further processing, compared to storing the new count matrix somewhere in .varm or .uns. Because if we want to make us of already implemented dimreduction and clustering methods from scanpy, then the count matrix needs to be in .X and for visualization we need the variable names stored as categories in .obs. Doing all of this in the same anndata will just make things cluttered.

Additionally the question is whether a spatialdata object should be required as input instead of an anndataone, because then a new table could be added directly instead of having multiple disconnected tables.

The function call would change from:
adata_new = sq.tl.var_embeddings(adata, group="Endothelium", design_matrix_key="design_matrix")
to
sq.tl.var_embeddings(sdata, group="Endothelium", design_matrix_key="design_matrix")

@LLehner
LLehner marked this pull request as ready for review October 10, 2024 13:06
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

featurePR introduces a new featuresquidpy2.0Everything releated to a Squidpy 2.0 release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@LLehner@codecov-commenter@giovp@timtreis
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add method to calculate embeddings for variable by distance aggregation by LLehner · Pull Request #807 · scverse/squidpy · GitHub
Skip to content

Add method to calculate embeddings for variable by distance aggregation - #807

Open
LLehner wants to merge 39 commits into
mainfrom
var_by_distance_clustering
Open

Add method to calculate embeddings for variable by distance aggregation#807
LLehner wants to merge 39 commits into
mainfrom
var_by_distance_clustering

Conversation

@LLehner

@LLehnerLLehner commented Mar 4, 2024

Copy link
Copy Markdown
Member

Description

Adds a method in tools to calculate embeddings of variables by their counts aggregated by distance.

Example usage

import squidpy as sq

load example data set
adata = sq.datasets.seqfish()

Calculate distances of each observation to a specified anchor point (e.g. cell type or tissue location). Here we use cell type "Endothelium" in the annotation column "celltype_mapped_refined":
sq.tl.var_by_distance(adata, groups="Endothelium", cluster_key="celltype_mapped_refined")

The resulting distances are stored in adata.obsm["design_matrix"]. Now we can calculate the embeddings, which are returned as a new anndata object:
adata_new = sq.tl.var_embeddings(adata, group="Endothelium", design_matrix_key="design_matrix")

Note that by default the bin of distance 0, meaning the counts that belong to the anchor point, are excluded. This can be changed by setting include_anchor=True in sq.tl.var_embeddings().

adata_new.X contains the aggregated var x distance_bin count matrix.
adata_new.obs contains the variables as a categorical matrix, which is required to highlight them in plots.

TODO

  • Add a plotting function so this doesn't need to be done manually.
  • Allow flexible embedding calculations

@LLehner
LLehner requested a review from timtreisMarch 4, 2024 22:56
@codecov-commenter

codecov-commenter commented Mar 4, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 33.33333% with 24 lines in your changes are missing coverage. Please review.

Project coverage is 69.75%. Comparing base (df8e042) to head (8ee07ba).

Additional details and impacted files
@@ Coverage Diff @@## main #807 +/- ##
==========================================
- Coverage 69.99% 69.75% -0.24% 
==========================================
Files 39 40 +1 Lines 5525 5561 +36 Branches 1029 1037 +8 ==========================================
+ Hits 3867 3879 +12 - Misses 1363 1387 +24 
Partials 295 295 
FilesCoverage Δ
src/squidpy/tl/_var_embeddings.py33.33% <33.33%> (ø)

@giovp

Copy link
Copy Markdown
Member

hi @LLehner , thank you for this, would you mind elaborating a bit when this would be used? also, what if the embedding are pre-calculated, or the user would like to use something other than the UMAP, should that be an option? finally, I think a test would be required before we get this in, thanks!

@timtreis

Copy link
Copy Markdown
Member

Hey @giovp, this feature was coming out of a discussion with @maiiashulman. We ran into a situation in which the "literature-curated" signature for hypoxia was either 20 or 4000 genes, the latter obviously being useless. So we wondered which other genes maybe show the same spatially variable pattern as a function of distance to a certain cell-type (e.g. epithelial). This is essentially a graphical method to see if a given set of genes (f.e. the 20 gene signature) even varies in a similar pattern.

But I agree with your points; if we see that it's actually doing something useful, we should make it a bit more flexible.

@LLehner
LLehner marked this pull request as draft April 22, 2024 22:05
@timtreis
timtreis marked this pull request as ready for review July 9, 2024 21:02
@timtreistimtreis added squidpy2.0 Everything releated to a Squidpy 2.0 release feature PR introduces a new feature labels Jul 9, 2024
@LLehner
LLehner marked this pull request as draft August 8, 2024 10:22
@LLehner

LLehner commented Aug 8, 2024

Copy link
Copy Markdown
MemberAuthor

@timtreis this function now returns an anndata object, which is i think simplifies further processing, compared to storing the new count matrix somewhere in .varm or .uns. Because if we want to make us of already implemented dimreduction and clustering methods from scanpy, then the count matrix needs to be in .X and for visualization we need the variable names stored as categories in .obs. Doing all of this in the same anndata will just make things cluttered.

Additionally the question is whether a spatialdata object should be required as input instead of an anndataone, because then a new table could be added directly instead of having multiple disconnected tables.

The function call would change from:
adata_new = sq.tl.var_embeddings(adata, group="Endothelium", design_matrix_key="design_matrix")
to
sq.tl.var_embeddings(sdata, group="Endothelium", design_matrix_key="design_matrix")

@LLehner
LLehner marked this pull request as ready for review October 10, 2024 13:06
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

featurePR introduces a new featuresquidpy2.0Everything releated to a Squidpy 2.0 release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@LLehner@codecov-commenter@giovp@timtreis
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add method to calculate embeddings for variable by distance aggregation by LLehner · Pull Request #807 · scverse/squidpy · GitHub
Skip to content

Add method to calculate embeddings for variable by distance aggregation - #807

Open
LLehner wants to merge 39 commits into
mainfrom
var_by_distance_clustering
Open

Add method to calculate embeddings for variable by distance aggregation#807
LLehner wants to merge 39 commits into
mainfrom
var_by_distance_clustering

Conversation

@LLehner

@LLehnerLLehner commented Mar 4, 2024

Copy link
Copy Markdown
Member

Description

Adds a method in tools to calculate embeddings of variables by their counts aggregated by distance.

Example usage

import squidpy as sq

load example data set
adata = sq.datasets.seqfish()

Calculate distances of each observation to a specified anchor point (e.g. cell type or tissue location). Here we use cell type "Endothelium" in the annotation column "celltype_mapped_refined":
sq.tl.var_by_distance(adata, groups="Endothelium", cluster_key="celltype_mapped_refined")

The resulting distances are stored in adata.obsm["design_matrix"]. Now we can calculate the embeddings, which are returned as a new anndata object:
adata_new = sq.tl.var_embeddings(adata, group="Endothelium", design_matrix_key="design_matrix")

Note that by default the bin of distance 0, meaning the counts that belong to the anchor point, are excluded. This can be changed by setting include_anchor=True in sq.tl.var_embeddings().

adata_new.X contains the aggregated var x distance_bin count matrix.
adata_new.obs contains the variables as a categorical matrix, which is required to highlight them in plots.

TODO

  • Add a plotting function so this doesn't need to be done manually.
  • Allow flexible embedding calculations

@LLehner
LLehner requested a review from timtreisMarch 4, 2024 22:56
@codecov-commenter

codecov-commenter commented Mar 4, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 33.33333% with 24 lines in your changes are missing coverage. Please review.

Project coverage is 69.75%. Comparing base (df8e042) to head (8ee07ba).

Additional details and impacted files
@@ Coverage Diff @@## main #807 +/- ##
==========================================
- Coverage 69.99% 69.75% -0.24% 
==========================================
Files 39 40 +1 Lines 5525 5561 +36 Branches 1029 1037 +8 ==========================================
+ Hits 3867 3879 +12 - Misses 1363 1387 +24 
Partials 295 295 
FilesCoverage Δ
src/squidpy/tl/_var_embeddings.py33.33% <33.33%> (ø)

@giovp

Copy link
Copy Markdown
Member

hi @LLehner , thank you for this, would you mind elaborating a bit when this would be used? also, what if the embedding are pre-calculated, or the user would like to use something other than the UMAP, should that be an option? finally, I think a test would be required before we get this in, thanks!

@timtreis

Copy link
Copy Markdown
Member

Hey @giovp, this feature was coming out of a discussion with @maiiashulman. We ran into a situation in which the "literature-curated" signature for hypoxia was either 20 or 4000 genes, the latter obviously being useless. So we wondered which other genes maybe show the same spatially variable pattern as a function of distance to a certain cell-type (e.g. epithelial). This is essentially a graphical method to see if a given set of genes (f.e. the 20 gene signature) even varies in a similar pattern.

But I agree with your points; if we see that it's actually doing something useful, we should make it a bit more flexible.

@LLehner
LLehner marked this pull request as draft April 22, 2024 22:05
@timtreis
timtreis marked this pull request as ready for review July 9, 2024 21:02
@timtreistimtreis added squidpy2.0 Everything releated to a Squidpy 2.0 release feature PR introduces a new feature labels Jul 9, 2024
@LLehner
LLehner marked this pull request as draft August 8, 2024 10:22
@LLehner

LLehner commented Aug 8, 2024

Copy link
Copy Markdown
MemberAuthor

@timtreis this function now returns an anndata object, which is i think simplifies further processing, compared to storing the new count matrix somewhere in .varm or .uns. Because if we want to make us of already implemented dimreduction and clustering methods from scanpy, then the count matrix needs to be in .X and for visualization we need the variable names stored as categories in .obs. Doing all of this in the same anndata will just make things cluttered.

Additionally the question is whether a spatialdata object should be required as input instead of an anndataone, because then a new table could be added directly instead of having multiple disconnected tables.

The function call would change from:
adata_new = sq.tl.var_embeddings(adata, group="Endothelium", design_matrix_key="design_matrix")
to
sq.tl.var_embeddings(sdata, group="Endothelium", design_matrix_key="design_matrix")

@LLehner
LLehner marked this pull request as ready for review October 10, 2024 13:06
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

featurePR introduces a new featuresquidpy2.0Everything releated to a Squidpy 2.0 release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@LLehner@codecov-commenter@giovp@timtreis
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Add method to calculate embeddings for variable by distance aggregation by LLehner · Pull Request #807 · scverse/squidpy · GitHub
Skip to content

Add method to calculate embeddings for variable by distance aggregation - #807

Open
LLehner wants to merge 39 commits into
mainfrom
var_by_distance_clustering
Open

Add method to calculate embeddings for variable by distance aggregation#807
LLehner wants to merge 39 commits into
mainfrom
var_by_distance_clustering

Conversation

@LLehner

@LLehnerLLehner commented Mar 4, 2024

Copy link
Copy Markdown
Member

Description

Adds a method in tools to calculate embeddings of variables by their counts aggregated by distance.

Example usage

import squidpy as sq

load example data set
adata = sq.datasets.seqfish()

Calculate distances of each observation to a specified anchor point (e.g. cell type or tissue location). Here we use cell type "Endothelium" in the annotation column "celltype_mapped_refined":
sq.tl.var_by_distance(adata, groups="Endothelium", cluster_key="celltype_mapped_refined")

The resulting distances are stored in adata.obsm["design_matrix"]. Now we can calculate the embeddings, which are returned as a new anndata object:
adata_new = sq.tl.var_embeddings(adata, group="Endothelium", design_matrix_key="design_matrix")

Note that by default the bin of distance 0, meaning the counts that belong to the anchor point, are excluded. This can be changed by setting include_anchor=True in sq.tl.var_embeddings().

adata_new.X contains the aggregated var x distance_bin count matrix.
adata_new.obs contains the variables as a categorical matrix, which is required to highlight them in plots.

TODO

  • Add a plotting function so this doesn't need to be done manually.
  • Allow flexible embedding calculations

@LLehner
LLehner requested a review from timtreisMarch 4, 2024 22:56
@codecov-commenter

codecov-commenter commented Mar 4, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 33.33333% with 24 lines in your changes are missing coverage. Please review.

Project coverage is 69.75%. Comparing base (df8e042) to head (8ee07ba).

Additional details and impacted files
@@ Coverage Diff @@## main #807 +/- ##
==========================================
- Coverage 69.99% 69.75% -0.24% 
==========================================
Files 39 40 +1 Lines 5525 5561 +36 Branches 1029 1037 +8 ==========================================
+ Hits 3867 3879 +12 - Misses 1363 1387 +24 
Partials 295 295 
FilesCoverage Δ
src/squidpy/tl/_var_embeddings.py33.33% <33.33%> (ø)

@giovp

Copy link
Copy Markdown
Member

hi @LLehner , thank you for this, would you mind elaborating a bit when this would be used? also, what if the embedding are pre-calculated, or the user would like to use something other than the UMAP, should that be an option? finally, I think a test would be required before we get this in, thanks!

@timtreis

Copy link
Copy Markdown
Member

Hey @giovp, this feature was coming out of a discussion with @maiiashulman. We ran into a situation in which the "literature-curated" signature for hypoxia was either 20 or 4000 genes, the latter obviously being useless. So we wondered which other genes maybe show the same spatially variable pattern as a function of distance to a certain cell-type (e.g. epithelial). This is essentially a graphical method to see if a given set of genes (f.e. the 20 gene signature) even varies in a similar pattern.

But I agree with your points; if we see that it's actually doing something useful, we should make it a bit more flexible.

@LLehner
LLehner marked this pull request as draft April 22, 2024 22:05
@timtreis
timtreis marked this pull request as ready for review July 9, 2024 21:02
@timtreistimtreis added squidpy2.0 Everything releated to a Squidpy 2.0 release feature PR introduces a new feature labels Jul 9, 2024
@LLehner
LLehner marked this pull request as draft August 8, 2024 10:22
@LLehner

LLehner commented Aug 8, 2024

Copy link
Copy Markdown
MemberAuthor

@timtreis this function now returns an anndata object, which is i think simplifies further processing, compared to storing the new count matrix somewhere in .varm or .uns. Because if we want to make us of already implemented dimreduction and clustering methods from scanpy, then the count matrix needs to be in .X and for visualization we need the variable names stored as categories in .obs. Doing all of this in the same anndata will just make things cluttered.

Additionally the question is whether a spatialdata object should be required as input instead of an anndataone, because then a new table could be added directly instead of having multiple disconnected tables.

The function call would change from:
adata_new = sq.tl.var_embeddings(adata, group="Endothelium", design_matrix_key="design_matrix")
to
sq.tl.var_embeddings(sdata, group="Endothelium", design_matrix_key="design_matrix")

@LLehner
LLehner marked this pull request as ready for review October 10, 2024 13:06
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

featurePR introduces a new featuresquidpy2.0Everything releated to a Squidpy 2.0 release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@LLehner@codecov-commenter@giovp@timtreis