Skip to content

Partial reading - #765

Merged
LucaMarconato merged 11 commits into
scverse:mainfrom
aeisenbarth:partial-reading
Jan 31, 2025
Merged

Partial reading#765
LucaMarconato merged 11 commits into
scverse:mainfrom
aeisenbarth:partial-reading

Conversation

@aeisenbarth

@aeisenbarthaeisenbarth commented Nov 6, 2024

Copy link
Copy Markdown
Contributor

This PR implements modes for read_zarr that allow to continue reading SpatialData elements even if some elements caused read errors, for example due to corrupted files (crash during write, a race condition or copy errors). This is useful, because with growing datasets (by number of elements), a corruption of a single file can make all other elements inaccessible. The ability to read at least the unaffected data helps for data recovery.

This is implemented by wrapping element-specific reader calls into a context managerhandle_read_errors. If enabled with on_bad_files="warn", it converts exceptions into warnings, and continues execution. Since execution is not aborted, all subsequent code within an iteration must be wrapped in the context manager to avoid that it is executed after a handled error. In order to identify which element failed to be read, the warning message is prepended with its location, e.g.:

shapes/blobs_polygons: JSONDecodeError: Expecting value: line 1 column 1 (char 0)

At the end, a SpatialData object is always returned, containing only the successfully read elements.

The current implementation is at the granularity level of whole elements, that means if anything in an element causes an error, the whole element is skipped. In principle it could be made more fine-granular, e.g. returning an element with just invalid attributes skipped, or returning a table with unreadable columns skipped. Especially for large single tables with annotations for all elements, this would be very beneficial (although they are column-based, so that a corrupted file would destroy a column's values for all annotations). However, the relevant code is not contained in this repository, but requires changes in anndata.read_zarr. To avoid losing annotations due to file corruption, I would recommend not to use a single table, but one table per SpatialData element.

For tests, see test_read_zarr_with_error and test_read_zarr_with_warnings.

Closes#457

Release notes

- Added option `on_bad_files="warn"` to `read_zarr` for partially reading corrupted datasets #457

@codecov

codecovBot commented Nov 6, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 89.89899% with 10 lines in your changes missing coverage. Please review.

Project coverage is 92.00%. Comparing base (a7e9e99) to head (9b1c7a3).
Report is 25 commits behind head on main.

Files with missing linesPatch %Lines
src/spatialdata/_io/io_zarr.py87.71%7 Missing ⚠️
src/spatialdata/_io/io_table.py86.36%3 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #765 +/- ##
==========================================
+ Coverage 91.95% 92.00% +0.05% 
==========================================
Files 47 47 Lines 7293 7331 +38 ==========================================
+ Hits 6706 6745 +39 + Misses 587 586 -1 
Files with missing linesCoverage Δ
src/spatialdata/_io/_utils.py88.44% <100.00%> (+1.21%)⬆️
src/spatialdata/_io/io_table.py94.44% <86.36%> (+0.56%)⬆️
src/spatialdata/_io/io_zarr.py88.46% <87.71%> (+0.68%)⬆️

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@LucaMarconato

Copy link
Copy Markdown
Member

Thanks @aeisenbarth, great PR and very well structured tests! I made some very minor adjustments here: I made a test more strict and replaced shutil.rmtree() with another approach; with the other devs we opted to minimize/avoid calls to rmtree(), for possible safety concerns.

@LucaMarconato
LucaMarconato enabled auto-merge (squash) January 31, 2025 15:00
@LucaMarconato
LucaMarconato merged commit ce88f15 into scverse:mainJan 31, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Tolerance when reading corrupted data

2 participants

@aeisenbarth@LucaMarconato
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Partial reading by aeisenbarth · Pull Request #765 · scverse/spatialdata · GitHub
Skip to content

Partial reading - #765

Merged
LucaMarconato merged 11 commits into
scverse:mainfrom
aeisenbarth:partial-reading
Jan 31, 2025
Merged

Partial reading#765
LucaMarconato merged 11 commits into
scverse:mainfrom
aeisenbarth:partial-reading

Conversation

@aeisenbarth

@aeisenbarthaeisenbarth commented Nov 6, 2024

Copy link
Copy Markdown
Contributor

This PR implements modes for read_zarr that allow to continue reading SpatialData elements even if some elements caused read errors, for example due to corrupted files (crash during write, a race condition or copy errors). This is useful, because with growing datasets (by number of elements), a corruption of a single file can make all other elements inaccessible. The ability to read at least the unaffected data helps for data recovery.

This is implemented by wrapping element-specific reader calls into a context managerhandle_read_errors. If enabled with on_bad_files="warn", it converts exceptions into warnings, and continues execution. Since execution is not aborted, all subsequent code within an iteration must be wrapped in the context manager to avoid that it is executed after a handled error. In order to identify which element failed to be read, the warning message is prepended with its location, e.g.:

shapes/blobs_polygons: JSONDecodeError: Expecting value: line 1 column 1 (char 0)

At the end, a SpatialData object is always returned, containing only the successfully read elements.

The current implementation is at the granularity level of whole elements, that means if anything in an element causes an error, the whole element is skipped. In principle it could be made more fine-granular, e.g. returning an element with just invalid attributes skipped, or returning a table with unreadable columns skipped. Especially for large single tables with annotations for all elements, this would be very beneficial (although they are column-based, so that a corrupted file would destroy a column's values for all annotations). However, the relevant code is not contained in this repository, but requires changes in anndata.read_zarr. To avoid losing annotations due to file corruption, I would recommend not to use a single table, but one table per SpatialData element.

For tests, see test_read_zarr_with_error and test_read_zarr_with_warnings.

Closes#457

Release notes

- Added option `on_bad_files="warn"` to `read_zarr` for partially reading corrupted datasets #457

@codecov

codecovBot commented Nov 6, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 89.89899% with 10 lines in your changes missing coverage. Please review.

Project coverage is 92.00%. Comparing base (a7e9e99) to head (9b1c7a3).
Report is 25 commits behind head on main.

Files with missing linesPatch %Lines
src/spatialdata/_io/io_zarr.py87.71%7 Missing ⚠️
src/spatialdata/_io/io_table.py86.36%3 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #765 +/- ##
==========================================
+ Coverage 91.95% 92.00% +0.05% 
==========================================
Files 47 47 Lines 7293 7331 +38 ==========================================
+ Hits 6706 6745 +39 + Misses 587 586 -1 
Files with missing linesCoverage Δ
src/spatialdata/_io/_utils.py88.44% <100.00%> (+1.21%)⬆️
src/spatialdata/_io/io_table.py94.44% <86.36%> (+0.56%)⬆️
src/spatialdata/_io/io_zarr.py88.46% <87.71%> (+0.68%)⬆️

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@LucaMarconato

Copy link
Copy Markdown
Member

Thanks @aeisenbarth, great PR and very well structured tests! I made some very minor adjustments here: I made a test more strict and replaced shutil.rmtree() with another approach; with the other devs we opted to minimize/avoid calls to rmtree(), for possible safety concerns.

@LucaMarconato
LucaMarconato enabled auto-merge (squash) January 31, 2025 15:00
@LucaMarconato
LucaMarconato merged commit ce88f15 into scverse:mainJan 31, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Tolerance when reading corrupted data

2 participants

@aeisenbarth@LucaMarconato
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Partial reading by aeisenbarth · Pull Request #765 · scverse/spatialdata · GitHub
Skip to content

Partial reading - #765

Merged
LucaMarconato merged 11 commits into
scverse:mainfrom
aeisenbarth:partial-reading
Jan 31, 2025
Merged

Partial reading#765
LucaMarconato merged 11 commits into
scverse:mainfrom
aeisenbarth:partial-reading

Conversation

@aeisenbarth

@aeisenbarthaeisenbarth commented Nov 6, 2024

Copy link
Copy Markdown
Contributor

This PR implements modes for read_zarr that allow to continue reading SpatialData elements even if some elements caused read errors, for example due to corrupted files (crash during write, a race condition or copy errors). This is useful, because with growing datasets (by number of elements), a corruption of a single file can make all other elements inaccessible. The ability to read at least the unaffected data helps for data recovery.

This is implemented by wrapping element-specific reader calls into a context managerhandle_read_errors. If enabled with on_bad_files="warn", it converts exceptions into warnings, and continues execution. Since execution is not aborted, all subsequent code within an iteration must be wrapped in the context manager to avoid that it is executed after a handled error. In order to identify which element failed to be read, the warning message is prepended with its location, e.g.:

shapes/blobs_polygons: JSONDecodeError: Expecting value: line 1 column 1 (char 0)

At the end, a SpatialData object is always returned, containing only the successfully read elements.

The current implementation is at the granularity level of whole elements, that means if anything in an element causes an error, the whole element is skipped. In principle it could be made more fine-granular, e.g. returning an element with just invalid attributes skipped, or returning a table with unreadable columns skipped. Especially for large single tables with annotations for all elements, this would be very beneficial (although they are column-based, so that a corrupted file would destroy a column's values for all annotations). However, the relevant code is not contained in this repository, but requires changes in anndata.read_zarr. To avoid losing annotations due to file corruption, I would recommend not to use a single table, but one table per SpatialData element.

For tests, see test_read_zarr_with_error and test_read_zarr_with_warnings.

Closes#457

Release notes

- Added option `on_bad_files="warn"` to `read_zarr` for partially reading corrupted datasets #457

@codecov

codecovBot commented Nov 6, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 89.89899% with 10 lines in your changes missing coverage. Please review.

Project coverage is 92.00%. Comparing base (a7e9e99) to head (9b1c7a3).
Report is 25 commits behind head on main.

Files with missing linesPatch %Lines
src/spatialdata/_io/io_zarr.py87.71%7 Missing ⚠️
src/spatialdata/_io/io_table.py86.36%3 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #765 +/- ##
==========================================
+ Coverage 91.95% 92.00% +0.05% 
==========================================
Files 47 47 Lines 7293 7331 +38 ==========================================
+ Hits 6706 6745 +39 + Misses 587 586 -1 
Files with missing linesCoverage Δ
src/spatialdata/_io/_utils.py88.44% <100.00%> (+1.21%)⬆️
src/spatialdata/_io/io_table.py94.44% <86.36%> (+0.56%)⬆️
src/spatialdata/_io/io_zarr.py88.46% <87.71%> (+0.68%)⬆️

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@LucaMarconato

Copy link
Copy Markdown
Member

Thanks @aeisenbarth, great PR and very well structured tests! I made some very minor adjustments here: I made a test more strict and replaced shutil.rmtree() with another approach; with the other devs we opted to minimize/avoid calls to rmtree(), for possible safety concerns.

@LucaMarconato
LucaMarconato enabled auto-merge (squash) January 31, 2025 15:00
@LucaMarconato
LucaMarconato merged commit ce88f15 into scverse:mainJan 31, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Tolerance when reading corrupted data

2 participants

@aeisenbarth@LucaMarconato
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Partial reading by aeisenbarth · Pull Request #765 · scverse/spatialdata · GitHub
Skip to content

Partial reading - #765

Merged
LucaMarconato merged 11 commits into
scverse:mainfrom
aeisenbarth:partial-reading
Jan 31, 2025
Merged

Partial reading#765
LucaMarconato merged 11 commits into
scverse:mainfrom
aeisenbarth:partial-reading

Conversation

@aeisenbarth

@aeisenbarthaeisenbarth commented Nov 6, 2024

Copy link
Copy Markdown
Contributor

This PR implements modes for read_zarr that allow to continue reading SpatialData elements even if some elements caused read errors, for example due to corrupted files (crash during write, a race condition or copy errors). This is useful, because with growing datasets (by number of elements), a corruption of a single file can make all other elements inaccessible. The ability to read at least the unaffected data helps for data recovery.

This is implemented by wrapping element-specific reader calls into a context managerhandle_read_errors. If enabled with on_bad_files="warn", it converts exceptions into warnings, and continues execution. Since execution is not aborted, all subsequent code within an iteration must be wrapped in the context manager to avoid that it is executed after a handled error. In order to identify which element failed to be read, the warning message is prepended with its location, e.g.:

shapes/blobs_polygons: JSONDecodeError: Expecting value: line 1 column 1 (char 0)

At the end, a SpatialData object is always returned, containing only the successfully read elements.

The current implementation is at the granularity level of whole elements, that means if anything in an element causes an error, the whole element is skipped. In principle it could be made more fine-granular, e.g. returning an element with just invalid attributes skipped, or returning a table with unreadable columns skipped. Especially for large single tables with annotations for all elements, this would be very beneficial (although they are column-based, so that a corrupted file would destroy a column's values for all annotations). However, the relevant code is not contained in this repository, but requires changes in anndata.read_zarr. To avoid losing annotations due to file corruption, I would recommend not to use a single table, but one table per SpatialData element.

For tests, see test_read_zarr_with_error and test_read_zarr_with_warnings.

Closes#457

Release notes

- Added option `on_bad_files="warn"` to `read_zarr` for partially reading corrupted datasets #457

@codecov

codecovBot commented Nov 6, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 89.89899% with 10 lines in your changes missing coverage. Please review.

Project coverage is 92.00%. Comparing base (a7e9e99) to head (9b1c7a3).
Report is 25 commits behind head on main.

Files with missing linesPatch %Lines
src/spatialdata/_io/io_zarr.py87.71%7 Missing ⚠️
src/spatialdata/_io/io_table.py86.36%3 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #765 +/- ##
==========================================
+ Coverage 91.95% 92.00% +0.05% 
==========================================
Files 47 47 Lines 7293 7331 +38 ==========================================
+ Hits 6706 6745 +39 + Misses 587 586 -1 
Files with missing linesCoverage Δ
src/spatialdata/_io/_utils.py88.44% <100.00%> (+1.21%)⬆️
src/spatialdata/_io/io_table.py94.44% <86.36%> (+0.56%)⬆️
src/spatialdata/_io/io_zarr.py88.46% <87.71%> (+0.68%)⬆️

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@LucaMarconato

Copy link
Copy Markdown
Member

Thanks @aeisenbarth, great PR and very well structured tests! I made some very minor adjustments here: I made a test more strict and replaced shutil.rmtree() with another approach; with the other devs we opted to minimize/avoid calls to rmtree(), for possible safety concerns.

@LucaMarconato
LucaMarconato enabled auto-merge (squash) January 31, 2025 15:00
@LucaMarconato
LucaMarconato merged commit ce88f15 into scverse:mainJan 31, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Tolerance when reading corrupted data

2 participants

@aeisenbarth@LucaMarconato
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Partial reading by aeisenbarth · Pull Request #765 · scverse/spatialdata · GitHub
Skip to content

Partial reading - #765

Merged
LucaMarconato merged 11 commits into
scverse:mainfrom
aeisenbarth:partial-reading
Jan 31, 2025
Merged

Partial reading#765
LucaMarconato merged 11 commits into
scverse:mainfrom
aeisenbarth:partial-reading

Conversation

@aeisenbarth

@aeisenbarthaeisenbarth commented Nov 6, 2024

Copy link
Copy Markdown
Contributor

This PR implements modes for read_zarr that allow to continue reading SpatialData elements even if some elements caused read errors, for example due to corrupted files (crash during write, a race condition or copy errors). This is useful, because with growing datasets (by number of elements), a corruption of a single file can make all other elements inaccessible. The ability to read at least the unaffected data helps for data recovery.

This is implemented by wrapping element-specific reader calls into a context managerhandle_read_errors. If enabled with on_bad_files="warn", it converts exceptions into warnings, and continues execution. Since execution is not aborted, all subsequent code within an iteration must be wrapped in the context manager to avoid that it is executed after a handled error. In order to identify which element failed to be read, the warning message is prepended with its location, e.g.:

shapes/blobs_polygons: JSONDecodeError: Expecting value: line 1 column 1 (char 0)

At the end, a SpatialData object is always returned, containing only the successfully read elements.

The current implementation is at the granularity level of whole elements, that means if anything in an element causes an error, the whole element is skipped. In principle it could be made more fine-granular, e.g. returning an element with just invalid attributes skipped, or returning a table with unreadable columns skipped. Especially for large single tables with annotations for all elements, this would be very beneficial (although they are column-based, so that a corrupted file would destroy a column's values for all annotations). However, the relevant code is not contained in this repository, but requires changes in anndata.read_zarr. To avoid losing annotations due to file corruption, I would recommend not to use a single table, but one table per SpatialData element.

For tests, see test_read_zarr_with_error and test_read_zarr_with_warnings.

Closes#457

Release notes

- Added option `on_bad_files="warn"` to `read_zarr` for partially reading corrupted datasets #457

@codecov

codecovBot commented Nov 6, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 89.89899% with 10 lines in your changes missing coverage. Please review.

Project coverage is 92.00%. Comparing base (a7e9e99) to head (9b1c7a3).
Report is 25 commits behind head on main.

Files with missing linesPatch %Lines
src/spatialdata/_io/io_zarr.py87.71%7 Missing ⚠️
src/spatialdata/_io/io_table.py86.36%3 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #765 +/- ##
==========================================
+ Coverage 91.95% 92.00% +0.05% 
==========================================
Files 47 47 Lines 7293 7331 +38 ==========================================
+ Hits 6706 6745 +39 + Misses 587 586 -1 
Files with missing linesCoverage Δ
src/spatialdata/_io/_utils.py88.44% <100.00%> (+1.21%)⬆️
src/spatialdata/_io/io_table.py94.44% <86.36%> (+0.56%)⬆️
src/spatialdata/_io/io_zarr.py88.46% <87.71%> (+0.68%)⬆️

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@LucaMarconato

Copy link
Copy Markdown
Member

Thanks @aeisenbarth, great PR and very well structured tests! I made some very minor adjustments here: I made a test more strict and replaced shutil.rmtree() with another approach; with the other devs we opted to minimize/avoid calls to rmtree(), for possible safety concerns.

@LucaMarconato
LucaMarconato enabled auto-merge (squash) January 31, 2025 15:00
@LucaMarconato
LucaMarconato merged commit ce88f15 into scverse:mainJan 31, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Tolerance when reading corrupted data

2 participants

@aeisenbarth@LucaMarconato
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Partial reading by aeisenbarth · Pull Request #765 · scverse/spatialdata · GitHub
Skip to content

Partial reading - #765

Merged
LucaMarconato merged 11 commits into
scverse:mainfrom
aeisenbarth:partial-reading
Jan 31, 2025
Merged

Partial reading#765
LucaMarconato merged 11 commits into
scverse:mainfrom
aeisenbarth:partial-reading

Conversation

@aeisenbarth

@aeisenbarthaeisenbarth commented Nov 6, 2024

Copy link
Copy Markdown
Contributor

This PR implements modes for read_zarr that allow to continue reading SpatialData elements even if some elements caused read errors, for example due to corrupted files (crash during write, a race condition or copy errors). This is useful, because with growing datasets (by number of elements), a corruption of a single file can make all other elements inaccessible. The ability to read at least the unaffected data helps for data recovery.

This is implemented by wrapping element-specific reader calls into a context managerhandle_read_errors. If enabled with on_bad_files="warn", it converts exceptions into warnings, and continues execution. Since execution is not aborted, all subsequent code within an iteration must be wrapped in the context manager to avoid that it is executed after a handled error. In order to identify which element failed to be read, the warning message is prepended with its location, e.g.:

shapes/blobs_polygons: JSONDecodeError: Expecting value: line 1 column 1 (char 0)

At the end, a SpatialData object is always returned, containing only the successfully read elements.

The current implementation is at the granularity level of whole elements, that means if anything in an element causes an error, the whole element is skipped. In principle it could be made more fine-granular, e.g. returning an element with just invalid attributes skipped, or returning a table with unreadable columns skipped. Especially for large single tables with annotations for all elements, this would be very beneficial (although they are column-based, so that a corrupted file would destroy a column's values for all annotations). However, the relevant code is not contained in this repository, but requires changes in anndata.read_zarr. To avoid losing annotations due to file corruption, I would recommend not to use a single table, but one table per SpatialData element.

For tests, see test_read_zarr_with_error and test_read_zarr_with_warnings.

Closes#457

Release notes

- Added option `on_bad_files="warn"` to `read_zarr` for partially reading corrupted datasets #457

@codecov

codecovBot commented Nov 6, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 89.89899% with 10 lines in your changes missing coverage. Please review.

Project coverage is 92.00%. Comparing base (a7e9e99) to head (9b1c7a3).
Report is 25 commits behind head on main.

Files with missing linesPatch %Lines
src/spatialdata/_io/io_zarr.py87.71%7 Missing ⚠️
src/spatialdata/_io/io_table.py86.36%3 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #765 +/- ##
==========================================
+ Coverage 91.95% 92.00% +0.05% 
==========================================
Files 47 47 Lines 7293 7331 +38 ==========================================
+ Hits 6706 6745 +39 + Misses 587 586 -1 
Files with missing linesCoverage Δ
src/spatialdata/_io/_utils.py88.44% <100.00%> (+1.21%)⬆️
src/spatialdata/_io/io_table.py94.44% <86.36%> (+0.56%)⬆️
src/spatialdata/_io/io_zarr.py88.46% <87.71%> (+0.68%)⬆️

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@LucaMarconato

Copy link
Copy Markdown
Member

Thanks @aeisenbarth, great PR and very well structured tests! I made some very minor adjustments here: I made a test more strict and replaced shutil.rmtree() with another approach; with the other devs we opted to minimize/avoid calls to rmtree(), for possible safety concerns.

@LucaMarconato
LucaMarconato enabled auto-merge (squash) January 31, 2025 15:00
@LucaMarconato
LucaMarconato merged commit ce88f15 into scverse:mainJan 31, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Tolerance when reading corrupted data

2 participants

@aeisenbarth@LucaMarconato
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Partial reading by aeisenbarth · Pull Request #765 · scverse/spatialdata · GitHub
Skip to content

Partial reading - #765

Merged
LucaMarconato merged 11 commits into
scverse:mainfrom
aeisenbarth:partial-reading
Jan 31, 2025
Merged

Partial reading#765
LucaMarconato merged 11 commits into
scverse:mainfrom
aeisenbarth:partial-reading

Conversation

@aeisenbarth

@aeisenbarthaeisenbarth commented Nov 6, 2024

Copy link
Copy Markdown
Contributor

This PR implements modes for read_zarr that allow to continue reading SpatialData elements even if some elements caused read errors, for example due to corrupted files (crash during write, a race condition or copy errors). This is useful, because with growing datasets (by number of elements), a corruption of a single file can make all other elements inaccessible. The ability to read at least the unaffected data helps for data recovery.

This is implemented by wrapping element-specific reader calls into a context managerhandle_read_errors. If enabled with on_bad_files="warn", it converts exceptions into warnings, and continues execution. Since execution is not aborted, all subsequent code within an iteration must be wrapped in the context manager to avoid that it is executed after a handled error. In order to identify which element failed to be read, the warning message is prepended with its location, e.g.:

shapes/blobs_polygons: JSONDecodeError: Expecting value: line 1 column 1 (char 0)

At the end, a SpatialData object is always returned, containing only the successfully read elements.

The current implementation is at the granularity level of whole elements, that means if anything in an element causes an error, the whole element is skipped. In principle it could be made more fine-granular, e.g. returning an element with just invalid attributes skipped, or returning a table with unreadable columns skipped. Especially for large single tables with annotations for all elements, this would be very beneficial (although they are column-based, so that a corrupted file would destroy a column's values for all annotations). However, the relevant code is not contained in this repository, but requires changes in anndata.read_zarr. To avoid losing annotations due to file corruption, I would recommend not to use a single table, but one table per SpatialData element.

For tests, see test_read_zarr_with_error and test_read_zarr_with_warnings.

Closes#457

Release notes

- Added option `on_bad_files="warn"` to `read_zarr` for partially reading corrupted datasets #457

@codecov

codecovBot commented Nov 6, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 89.89899% with 10 lines in your changes missing coverage. Please review.

Project coverage is 92.00%. Comparing base (a7e9e99) to head (9b1c7a3).
Report is 25 commits behind head on main.

Files with missing linesPatch %Lines
src/spatialdata/_io/io_zarr.py87.71%7 Missing ⚠️
src/spatialdata/_io/io_table.py86.36%3 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #765 +/- ##
==========================================
+ Coverage 91.95% 92.00% +0.05% 
==========================================
Files 47 47 Lines 7293 7331 +38 ==========================================
+ Hits 6706 6745 +39 + Misses 587 586 -1 
Files with missing linesCoverage Δ
src/spatialdata/_io/_utils.py88.44% <100.00%> (+1.21%)⬆️
src/spatialdata/_io/io_table.py94.44% <86.36%> (+0.56%)⬆️
src/spatialdata/_io/io_zarr.py88.46% <87.71%> (+0.68%)⬆️

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@LucaMarconato

Copy link
Copy Markdown
Member

Thanks @aeisenbarth, great PR and very well structured tests! I made some very minor adjustments here: I made a test more strict and replaced shutil.rmtree() with another approach; with the other devs we opted to minimize/avoid calls to rmtree(), for possible safety concerns.

@LucaMarconato
LucaMarconato enabled auto-merge (squash) January 31, 2025 15:00
@LucaMarconato
LucaMarconato merged commit ce88f15 into scverse:mainJan 31, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Tolerance when reading corrupted data

2 participants

@aeisenbarth@LucaMarconato
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Partial reading by aeisenbarth · Pull Request #765 · scverse/spatialdata · GitHub
Skip to content

Partial reading - #765

Merged
LucaMarconato merged 11 commits into
scverse:mainfrom
aeisenbarth:partial-reading
Jan 31, 2025
Merged

Partial reading#765
LucaMarconato merged 11 commits into
scverse:mainfrom
aeisenbarth:partial-reading

Conversation

@aeisenbarth

@aeisenbarthaeisenbarth commented Nov 6, 2024

Copy link
Copy Markdown
Contributor

This PR implements modes for read_zarr that allow to continue reading SpatialData elements even if some elements caused read errors, for example due to corrupted files (crash during write, a race condition or copy errors). This is useful, because with growing datasets (by number of elements), a corruption of a single file can make all other elements inaccessible. The ability to read at least the unaffected data helps for data recovery.

This is implemented by wrapping element-specific reader calls into a context managerhandle_read_errors. If enabled with on_bad_files="warn", it converts exceptions into warnings, and continues execution. Since execution is not aborted, all subsequent code within an iteration must be wrapped in the context manager to avoid that it is executed after a handled error. In order to identify which element failed to be read, the warning message is prepended with its location, e.g.:

shapes/blobs_polygons: JSONDecodeError: Expecting value: line 1 column 1 (char 0)

At the end, a SpatialData object is always returned, containing only the successfully read elements.

The current implementation is at the granularity level of whole elements, that means if anything in an element causes an error, the whole element is skipped. In principle it could be made more fine-granular, e.g. returning an element with just invalid attributes skipped, or returning a table with unreadable columns skipped. Especially for large single tables with annotations for all elements, this would be very beneficial (although they are column-based, so that a corrupted file would destroy a column's values for all annotations). However, the relevant code is not contained in this repository, but requires changes in anndata.read_zarr. To avoid losing annotations due to file corruption, I would recommend not to use a single table, but one table per SpatialData element.

For tests, see test_read_zarr_with_error and test_read_zarr_with_warnings.

Closes#457

Release notes

- Added option `on_bad_files="warn"` to `read_zarr` for partially reading corrupted datasets #457

@codecov

codecovBot commented Nov 6, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 89.89899% with 10 lines in your changes missing coverage. Please review.

Project coverage is 92.00%. Comparing base (a7e9e99) to head (9b1c7a3).
Report is 25 commits behind head on main.

Files with missing linesPatch %Lines
src/spatialdata/_io/io_zarr.py87.71%7 Missing ⚠️
src/spatialdata/_io/io_table.py86.36%3 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #765 +/- ##
==========================================
+ Coverage 91.95% 92.00% +0.05% 
==========================================
Files 47 47 Lines 7293 7331 +38 ==========================================
+ Hits 6706 6745 +39 + Misses 587 586 -1 
Files with missing linesCoverage Δ
src/spatialdata/_io/_utils.py88.44% <100.00%> (+1.21%)⬆️
src/spatialdata/_io/io_table.py94.44% <86.36%> (+0.56%)⬆️
src/spatialdata/_io/io_zarr.py88.46% <87.71%> (+0.68%)⬆️

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@LucaMarconato

Copy link
Copy Markdown
Member

Thanks @aeisenbarth, great PR and very well structured tests! I made some very minor adjustments here: I made a test more strict and replaced shutil.rmtree() with another approach; with the other devs we opted to minimize/avoid calls to rmtree(), for possible safety concerns.

@LucaMarconato
LucaMarconato enabled auto-merge (squash) January 31, 2025 15:00
@LucaMarconato
LucaMarconato merged commit ce88f15 into scverse:mainJan 31, 2025
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Tolerance when reading corrupted data

2 participants

@aeisenbarth@LucaMarconato