Parallelize file metadata inspection in 2-stage reader - #75

Open
Baptiste-Arnould wants to merge 2 commits into
MathEXLab:mainfrom
Baptiste-Arnould:feature/parallel-metadata-inspection
Open

Parallelize file metadata inspection in 2-stage reader#75
Baptiste-Arnould wants to merge 2 commits into
MathEXLab:mainfrom
Baptiste-Arnould:feature/parallel-metadata-inspection

Conversation

@Baptiste-Arnould

Copy link
Copy Markdown

Summary

This PR parallelizes file metadata inspection during the initialization of the 2-stage reader.

Previously, metadata for all input files were inspected sequentially by rank 0. For datasets containing a large number of files, this can result in significant initialization time.

The metadata inspection is now distributed across the available reader ranks. The resulting metadata are gathered on rank 0 and reordered according to the original file order before constructing _file_time, _shape, _is_real, and _files_size.

The data-reading algorithm itself is unchanged.

Changes

  • Add a helper function to inspect file metadata.
  • Distribute metadata inspection across reader ranks.
  • Gather metadata on rank 0.
  • Restore the original file ordering before constructing reader metadata.

Testing

Tested with MPI on a dataset composed of 11800 files (~250TB in total).

Closes#74

@dalcinl

Copy link
Copy Markdown
Collaborator

Your changes are definitely an improvement and the implementation looks reasonable. There is just a minor thing I would like you to reconsider. You used a cyclic distribution to parallelize. That is OK in general, it is very convenient. But when order is important, it forces you to keep track of indices to reconstruct the original ordering, leading to slightly more convoluted code..

Look at pyspod/utils/parallel.py, note the _blockdist function. It is currently "private" (under the conventional initial underscore). However, we could rename and make it "public", and then use it in in the reader to distribute the metadata reading using a block-contiguous distribution. This way you would not need to sort to restore order, and the code may end up being simpler to read. Let me sketch the code below:

local_infos= []
# Distribute files across reader rankscount, start=utils_par.blockdist(len(data_list), comm.size, comm.rank)
foriinrange(start, start+count):
file_name=data_list[i]
file_shape, file_dtype, file_size=_inspect_file_metadata(file_name, variables[0])
local_infos.append((file_name, file_shape, str(file_dtype), file_size))
# Gather metadata on rank 0all_infos=comm.gather(local_infos, root=0)

and then you need to keep updating a few things to remove the sort call, etc.

Would you be willing to explore this approach and update the PR accordingly?

@Baptiste-Arnould

Copy link
Copy Markdown
Author

Thanks for the suggestion. I updated the PR to use a block distribution for metadata inspection and made blockdist public. This removes the explicit file indices and sorting step and simplifies the reconstruction on rank 0.

@dalcinl

Copy link
Copy Markdown
Collaborator

LGTM.

@mengaldo Once tests pass (modulo the MPI failures that no one had taken care to fix 🤦 ) I think this one is good to go.

@codecov

codecovBot commented Aug 30, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 71.42857% with 14 lines in your changes missing coverage. Please review.
✅ Project coverage is 76.95%. Comparing base (19372d1) to head (1ee2f69).

Files with missing linesPatch %Lines
pyspod/utils/reader.py68.88%13 Missing and 1 partial ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #75 +/- ##
==========================================
+ Coverage 76.84% 76.95% +0.11% 
==========================================
Files 16 16 Lines 3532 3549 +17 Branches 465 467 +2 ==========================================
+ Hits 2714 2731 +17 
Misses 660 660 Partials 158 158 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Parallelize file metadata inspection in 2-stage reader

2 participants

@Baptiste-Arnould@dalcinl
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Parallelize file metadata inspection in 2-stage reader - #75

Open
Baptiste-Arnould wants to merge 2 commits into
MathEXLab:mainfrom
Baptiste-Arnould:feature/parallel-metadata-inspection
Open

Parallelize file metadata inspection in 2-stage reader#75
Baptiste-Arnould wants to merge 2 commits into
MathEXLab:mainfrom
Baptiste-Arnould:feature/parallel-metadata-inspection

Conversation

@Baptiste-Arnould

Copy link
Copy Markdown

Summary

This PR parallelizes file metadata inspection during the initialization of the 2-stage reader.

Previously, metadata for all input files were inspected sequentially by rank 0. For datasets containing a large number of files, this can result in significant initialization time.

The metadata inspection is now distributed across the available reader ranks. The resulting metadata are gathered on rank 0 and reordered according to the original file order before constructing _file_time, _shape, _is_real, and _files_size.

The data-reading algorithm itself is unchanged.

Changes

  • Add a helper function to inspect file metadata.
  • Distribute metadata inspection across reader ranks.
  • Gather metadata on rank 0.
  • Restore the original file ordering before constructing reader metadata.

Testing

Tested with MPI on a dataset composed of 11800 files (~250TB in total).

Closes#74

@dalcinl

Copy link
Copy Markdown
Collaborator

Your changes are definitely an improvement and the implementation looks reasonable. There is just a minor thing I would like you to reconsider. You used a cyclic distribution to parallelize. That is OK in general, it is very convenient. But when order is important, it forces you to keep track of indices to reconstruct the original ordering, leading to slightly more convoluted code..

Look at pyspod/utils/parallel.py, note the _blockdist function. It is currently "private" (under the conventional initial underscore). However, we could rename and make it "public", and then use it in in the reader to distribute the metadata reading using a block-contiguous distribution. This way you would not need to sort to restore order, and the code may end up being simpler to read. Let me sketch the code below:

local_infos= []
# Distribute files across reader rankscount, start=utils_par.blockdist(len(data_list), comm.size, comm.rank)
foriinrange(start, start+count):
file_name=data_list[i]
file_shape, file_dtype, file_size=_inspect_file_metadata(file_name, variables[0])
local_infos.append((file_name, file_shape, str(file_dtype), file_size))
# Gather metadata on rank 0all_infos=comm.gather(local_infos, root=0)

and then you need to keep updating a few things to remove the sort call, etc.

Would you be willing to explore this approach and update the PR accordingly?

@Baptiste-Arnould

Copy link
Copy Markdown
Author

Thanks for the suggestion. I updated the PR to use a block distribution for metadata inspection and made blockdist public. This removes the explicit file indices and sorting step and simplifies the reconstruction on rank 0.

@dalcinl

Copy link
Copy Markdown
Collaborator

LGTM.

@mengaldo Once tests pass (modulo the MPI failures that no one had taken care to fix 🤦 ) I think this one is good to go.

@codecov

codecovBot commented Aug 30, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 71.42857% with 14 lines in your changes missing coverage. Please review.
✅ Project coverage is 76.95%. Comparing base (19372d1) to head (1ee2f69).

Files with missing linesPatch %Lines
pyspod/utils/reader.py68.88%13 Missing and 1 partial ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #75 +/- ##
==========================================
+ Coverage 76.84% 76.95% +0.11% 
==========================================
Files 16 16 Lines 3532 3549 +17 Branches 465 467 +2 ==========================================
+ Hits 2714 2731 +17 
Misses 660 660 Partials 158 158 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Parallelize file metadata inspection in 2-stage reader

2 participants

@Baptiste-Arnould@dalcinl
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Parallelize file metadata inspection in 2-stage reader - #75

Open
Baptiste-Arnould wants to merge 2 commits into
MathEXLab:mainfrom
Baptiste-Arnould:feature/parallel-metadata-inspection
Open

Parallelize file metadata inspection in 2-stage reader#75
Baptiste-Arnould wants to merge 2 commits into
MathEXLab:mainfrom
Baptiste-Arnould:feature/parallel-metadata-inspection

Conversation

@Baptiste-Arnould

Copy link
Copy Markdown

Summary

This PR parallelizes file metadata inspection during the initialization of the 2-stage reader.

Previously, metadata for all input files were inspected sequentially by rank 0. For datasets containing a large number of files, this can result in significant initialization time.

The metadata inspection is now distributed across the available reader ranks. The resulting metadata are gathered on rank 0 and reordered according to the original file order before constructing _file_time, _shape, _is_real, and _files_size.

The data-reading algorithm itself is unchanged.

Changes

  • Add a helper function to inspect file metadata.
  • Distribute metadata inspection across reader ranks.
  • Gather metadata on rank 0.
  • Restore the original file ordering before constructing reader metadata.

Testing

Tested with MPI on a dataset composed of 11800 files (~250TB in total).

Closes#74

@dalcinl

Copy link
Copy Markdown
Collaborator

Your changes are definitely an improvement and the implementation looks reasonable. There is just a minor thing I would like you to reconsider. You used a cyclic distribution to parallelize. That is OK in general, it is very convenient. But when order is important, it forces you to keep track of indices to reconstruct the original ordering, leading to slightly more convoluted code..

Look at pyspod/utils/parallel.py, note the _blockdist function. It is currently "private" (under the conventional initial underscore). However, we could rename and make it "public", and then use it in in the reader to distribute the metadata reading using a block-contiguous distribution. This way you would not need to sort to restore order, and the code may end up being simpler to read. Let me sketch the code below:

local_infos= []
# Distribute files across reader rankscount, start=utils_par.blockdist(len(data_list), comm.size, comm.rank)
foriinrange(start, start+count):
file_name=data_list[i]
file_shape, file_dtype, file_size=_inspect_file_metadata(file_name, variables[0])
local_infos.append((file_name, file_shape, str(file_dtype), file_size))
# Gather metadata on rank 0all_infos=comm.gather(local_infos, root=0)

and then you need to keep updating a few things to remove the sort call, etc.

Would you be willing to explore this approach and update the PR accordingly?

@Baptiste-Arnould

Copy link
Copy Markdown
Author

Thanks for the suggestion. I updated the PR to use a block distribution for metadata inspection and made blockdist public. This removes the explicit file indices and sorting step and simplifies the reconstruction on rank 0.

@dalcinl

Copy link
Copy Markdown
Collaborator

LGTM.

@mengaldo Once tests pass (modulo the MPI failures that no one had taken care to fix 🤦 ) I think this one is good to go.

@codecov

codecovBot commented Aug 30, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 71.42857% with 14 lines in your changes missing coverage. Please review.
✅ Project coverage is 76.95%. Comparing base (19372d1) to head (1ee2f69).

Files with missing linesPatch %Lines
pyspod/utils/reader.py68.88%13 Missing and 1 partial ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #75 +/- ##
==========================================
+ Coverage 76.84% 76.95% +0.11% 
==========================================
Files 16 16 Lines 3532 3549 +17 Branches 465 467 +2 ==========================================
+ Hits 2714 2731 +17 
Misses 660 660 Partials 158 158 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Parallelize file metadata inspection in 2-stage reader

2 participants

@Baptiste-Arnould@dalcinl
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Parallelize file metadata inspection in 2-stage reader - #75

Open
Baptiste-Arnould wants to merge 2 commits into
MathEXLab:mainfrom
Baptiste-Arnould:feature/parallel-metadata-inspection
Open

Parallelize file metadata inspection in 2-stage reader#75
Baptiste-Arnould wants to merge 2 commits into
MathEXLab:mainfrom
Baptiste-Arnould:feature/parallel-metadata-inspection

Conversation

@Baptiste-Arnould

Copy link
Copy Markdown

Summary

This PR parallelizes file metadata inspection during the initialization of the 2-stage reader.

Previously, metadata for all input files were inspected sequentially by rank 0. For datasets containing a large number of files, this can result in significant initialization time.

The metadata inspection is now distributed across the available reader ranks. The resulting metadata are gathered on rank 0 and reordered according to the original file order before constructing _file_time, _shape, _is_real, and _files_size.

The data-reading algorithm itself is unchanged.

Changes

  • Add a helper function to inspect file metadata.
  • Distribute metadata inspection across reader ranks.
  • Gather metadata on rank 0.
  • Restore the original file ordering before constructing reader metadata.

Testing

Tested with MPI on a dataset composed of 11800 files (~250TB in total).

Closes#74

@dalcinl

Copy link
Copy Markdown
Collaborator

Your changes are definitely an improvement and the implementation looks reasonable. There is just a minor thing I would like you to reconsider. You used a cyclic distribution to parallelize. That is OK in general, it is very convenient. But when order is important, it forces you to keep track of indices to reconstruct the original ordering, leading to slightly more convoluted code..

Look at pyspod/utils/parallel.py, note the _blockdist function. It is currently "private" (under the conventional initial underscore). However, we could rename and make it "public", and then use it in in the reader to distribute the metadata reading using a block-contiguous distribution. This way you would not need to sort to restore order, and the code may end up being simpler to read. Let me sketch the code below:

local_infos= []
# Distribute files across reader rankscount, start=utils_par.blockdist(len(data_list), comm.size, comm.rank)
foriinrange(start, start+count):
file_name=data_list[i]
file_shape, file_dtype, file_size=_inspect_file_metadata(file_name, variables[0])
local_infos.append((file_name, file_shape, str(file_dtype), file_size))
# Gather metadata on rank 0all_infos=comm.gather(local_infos, root=0)

and then you need to keep updating a few things to remove the sort call, etc.

Would you be willing to explore this approach and update the PR accordingly?

@Baptiste-Arnould

Copy link
Copy Markdown
Author

Thanks for the suggestion. I updated the PR to use a block distribution for metadata inspection and made blockdist public. This removes the explicit file indices and sorting step and simplifies the reconstruction on rank 0.

@dalcinl

Copy link
Copy Markdown
Collaborator

LGTM.

@mengaldo Once tests pass (modulo the MPI failures that no one had taken care to fix 🤦 ) I think this one is good to go.

@codecov

codecovBot commented Aug 30, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 71.42857% with 14 lines in your changes missing coverage. Please review.
✅ Project coverage is 76.95%. Comparing base (19372d1) to head (1ee2f69).

Files with missing linesPatch %Lines
pyspod/utils/reader.py68.88%13 Missing and 1 partial ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #75 +/- ##
==========================================
+ Coverage 76.84% 76.95% +0.11% 
==========================================
Files 16 16 Lines 3532 3549 +17 Branches 465 467 +2 ==========================================
+ Hits 2714 2731 +17 
Misses 660 660 Partials 158 158 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Parallelize file metadata inspection in 2-stage reader

2 participants

@Baptiste-Arnould@dalcinl
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Parallelize file metadata inspection in 2-stage reader - #75

Open
Baptiste-Arnould wants to merge 2 commits into
MathEXLab:mainfrom
Baptiste-Arnould:feature/parallel-metadata-inspection
Open

Parallelize file metadata inspection in 2-stage reader#75
Baptiste-Arnould wants to merge 2 commits into
MathEXLab:mainfrom
Baptiste-Arnould:feature/parallel-metadata-inspection

Conversation

@Baptiste-Arnould

Copy link
Copy Markdown

Summary

This PR parallelizes file metadata inspection during the initialization of the 2-stage reader.

Previously, metadata for all input files were inspected sequentially by rank 0. For datasets containing a large number of files, this can result in significant initialization time.

The metadata inspection is now distributed across the available reader ranks. The resulting metadata are gathered on rank 0 and reordered according to the original file order before constructing _file_time, _shape, _is_real, and _files_size.

The data-reading algorithm itself is unchanged.

Changes

  • Add a helper function to inspect file metadata.
  • Distribute metadata inspection across reader ranks.
  • Gather metadata on rank 0.
  • Restore the original file ordering before constructing reader metadata.

Testing

Tested with MPI on a dataset composed of 11800 files (~250TB in total).

Closes#74

@dalcinl

Copy link
Copy Markdown
Collaborator

Your changes are definitely an improvement and the implementation looks reasonable. There is just a minor thing I would like you to reconsider. You used a cyclic distribution to parallelize. That is OK in general, it is very convenient. But when order is important, it forces you to keep track of indices to reconstruct the original ordering, leading to slightly more convoluted code..

Look at pyspod/utils/parallel.py, note the _blockdist function. It is currently "private" (under the conventional initial underscore). However, we could rename and make it "public", and then use it in in the reader to distribute the metadata reading using a block-contiguous distribution. This way you would not need to sort to restore order, and the code may end up being simpler to read. Let me sketch the code below:

local_infos= []
# Distribute files across reader rankscount, start=utils_par.blockdist(len(data_list), comm.size, comm.rank)
foriinrange(start, start+count):
file_name=data_list[i]
file_shape, file_dtype, file_size=_inspect_file_metadata(file_name, variables[0])
local_infos.append((file_name, file_shape, str(file_dtype), file_size))
# Gather metadata on rank 0all_infos=comm.gather(local_infos, root=0)

and then you need to keep updating a few things to remove the sort call, etc.

Would you be willing to explore this approach and update the PR accordingly?

@Baptiste-Arnould

Copy link
Copy Markdown
Author

Thanks for the suggestion. I updated the PR to use a block distribution for metadata inspection and made blockdist public. This removes the explicit file indices and sorting step and simplifies the reconstruction on rank 0.

@dalcinl

Copy link
Copy Markdown
Collaborator

LGTM.

@mengaldo Once tests pass (modulo the MPI failures that no one had taken care to fix 🤦 ) I think this one is good to go.

@codecov

codecovBot commented Aug 30, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 71.42857% with 14 lines in your changes missing coverage. Please review.
✅ Project coverage is 76.95%. Comparing base (19372d1) to head (1ee2f69).

Files with missing linesPatch %Lines
pyspod/utils/reader.py68.88%13 Missing and 1 partial ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #75 +/- ##
==========================================
+ Coverage 76.84% 76.95% +0.11% 
==========================================
Files 16 16 Lines 3532 3549 +17 Branches 465 467 +2 ==========================================
+ Hits 2714 2731 +17 
Misses 660 660 Partials 158 158 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Parallelize file metadata inspection in 2-stage reader

2 participants

@Baptiste-Arnould@dalcinl
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Parallelize file metadata inspection in 2-stage reader - #75

Open
Baptiste-Arnould wants to merge 2 commits into
MathEXLab:mainfrom
Baptiste-Arnould:feature/parallel-metadata-inspection
Open

Parallelize file metadata inspection in 2-stage reader#75
Baptiste-Arnould wants to merge 2 commits into
MathEXLab:mainfrom
Baptiste-Arnould:feature/parallel-metadata-inspection

Conversation

@Baptiste-Arnould

Copy link
Copy Markdown

Summary

This PR parallelizes file metadata inspection during the initialization of the 2-stage reader.

Previously, metadata for all input files were inspected sequentially by rank 0. For datasets containing a large number of files, this can result in significant initialization time.

The metadata inspection is now distributed across the available reader ranks. The resulting metadata are gathered on rank 0 and reordered according to the original file order before constructing _file_time, _shape, _is_real, and _files_size.

The data-reading algorithm itself is unchanged.

Changes

  • Add a helper function to inspect file metadata.
  • Distribute metadata inspection across reader ranks.
  • Gather metadata on rank 0.
  • Restore the original file ordering before constructing reader metadata.

Testing

Tested with MPI on a dataset composed of 11800 files (~250TB in total).

Closes#74

@dalcinl

Copy link
Copy Markdown
Collaborator

Your changes are definitely an improvement and the implementation looks reasonable. There is just a minor thing I would like you to reconsider. You used a cyclic distribution to parallelize. That is OK in general, it is very convenient. But when order is important, it forces you to keep track of indices to reconstruct the original ordering, leading to slightly more convoluted code..

Look at pyspod/utils/parallel.py, note the _blockdist function. It is currently "private" (under the conventional initial underscore). However, we could rename and make it "public", and then use it in in the reader to distribute the metadata reading using a block-contiguous distribution. This way you would not need to sort to restore order, and the code may end up being simpler to read. Let me sketch the code below:

local_infos= []
# Distribute files across reader rankscount, start=utils_par.blockdist(len(data_list), comm.size, comm.rank)
foriinrange(start, start+count):
file_name=data_list[i]
file_shape, file_dtype, file_size=_inspect_file_metadata(file_name, variables[0])
local_infos.append((file_name, file_shape, str(file_dtype), file_size))
# Gather metadata on rank 0all_infos=comm.gather(local_infos, root=0)

and then you need to keep updating a few things to remove the sort call, etc.

Would you be willing to explore this approach and update the PR accordingly?

@Baptiste-Arnould

Copy link
Copy Markdown
Author

Thanks for the suggestion. I updated the PR to use a block distribution for metadata inspection and made blockdist public. This removes the explicit file indices and sorting step and simplifies the reconstruction on rank 0.

@dalcinl

Copy link
Copy Markdown
Collaborator

LGTM.

@mengaldo Once tests pass (modulo the MPI failures that no one had taken care to fix 🤦 ) I think this one is good to go.

@codecov

codecovBot commented Aug 30, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 71.42857% with 14 lines in your changes missing coverage. Please review.
✅ Project coverage is 76.95%. Comparing base (19372d1) to head (1ee2f69).

Files with missing linesPatch %Lines
pyspod/utils/reader.py68.88%13 Missing and 1 partial ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #75 +/- ##
==========================================
+ Coverage 76.84% 76.95% +0.11% 
==========================================
Files 16 16 Lines 3532 3549 +17 Branches 465 467 +2 ==========================================
+ Hits 2714 2731 +17 
Misses 660 660 Partials 158 158 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Parallelize file metadata inspection in 2-stage reader

2 participants

@Baptiste-Arnould@dalcinl
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Parallelize file metadata inspection in 2-stage reader - #75

Open
Baptiste-Arnould wants to merge 2 commits into
MathEXLab:mainfrom
Baptiste-Arnould:feature/parallel-metadata-inspection
Open

Parallelize file metadata inspection in 2-stage reader#75
Baptiste-Arnould wants to merge 2 commits into
MathEXLab:mainfrom
Baptiste-Arnould:feature/parallel-metadata-inspection

Conversation

@Baptiste-Arnould

Copy link
Copy Markdown

Summary

This PR parallelizes file metadata inspection during the initialization of the 2-stage reader.

Previously, metadata for all input files were inspected sequentially by rank 0. For datasets containing a large number of files, this can result in significant initialization time.

The metadata inspection is now distributed across the available reader ranks. The resulting metadata are gathered on rank 0 and reordered according to the original file order before constructing _file_time, _shape, _is_real, and _files_size.

The data-reading algorithm itself is unchanged.

Changes

  • Add a helper function to inspect file metadata.
  • Distribute metadata inspection across reader ranks.
  • Gather metadata on rank 0.
  • Restore the original file ordering before constructing reader metadata.

Testing

Tested with MPI on a dataset composed of 11800 files (~250TB in total).

Closes#74

@dalcinl

Copy link
Copy Markdown
Collaborator

Your changes are definitely an improvement and the implementation looks reasonable. There is just a minor thing I would like you to reconsider. You used a cyclic distribution to parallelize. That is OK in general, it is very convenient. But when order is important, it forces you to keep track of indices to reconstruct the original ordering, leading to slightly more convoluted code..

Look at pyspod/utils/parallel.py, note the _blockdist function. It is currently "private" (under the conventional initial underscore). However, we could rename and make it "public", and then use it in in the reader to distribute the metadata reading using a block-contiguous distribution. This way you would not need to sort to restore order, and the code may end up being simpler to read. Let me sketch the code below:

local_infos= []
# Distribute files across reader rankscount, start=utils_par.blockdist(len(data_list), comm.size, comm.rank)
foriinrange(start, start+count):
file_name=data_list[i]
file_shape, file_dtype, file_size=_inspect_file_metadata(file_name, variables[0])
local_infos.append((file_name, file_shape, str(file_dtype), file_size))
# Gather metadata on rank 0all_infos=comm.gather(local_infos, root=0)

and then you need to keep updating a few things to remove the sort call, etc.

Would you be willing to explore this approach and update the PR accordingly?

@Baptiste-Arnould

Copy link
Copy Markdown
Author

Thanks for the suggestion. I updated the PR to use a block distribution for metadata inspection and made blockdist public. This removes the explicit file indices and sorting step and simplifies the reconstruction on rank 0.

@dalcinl

Copy link
Copy Markdown
Collaborator

LGTM.

@mengaldo Once tests pass (modulo the MPI failures that no one had taken care to fix 🤦 ) I think this one is good to go.

@codecov

codecovBot commented Aug 30, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 71.42857% with 14 lines in your changes missing coverage. Please review.
✅ Project coverage is 76.95%. Comparing base (19372d1) to head (1ee2f69).

Files with missing linesPatch %Lines
pyspod/utils/reader.py68.88%13 Missing and 1 partial ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #75 +/- ##
==========================================
+ Coverage 76.84% 76.95% +0.11% 
==========================================
Files 16 16 Lines 3532 3549 +17 Branches 465 467 +2 ==========================================
+ Hits 2714 2731 +17 
Misses 660 660 Partials 158 158 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Parallelize file metadata inspection in 2-stage reader

2 participants

@Baptiste-Arnould@dalcinl
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Parallelize file metadata inspection in 2-stage reader - #75

Open
Baptiste-Arnould wants to merge 2 commits into
MathEXLab:mainfrom
Baptiste-Arnould:feature/parallel-metadata-inspection
Open

Parallelize file metadata inspection in 2-stage reader#75
Baptiste-Arnould wants to merge 2 commits into
MathEXLab:mainfrom
Baptiste-Arnould:feature/parallel-metadata-inspection

Conversation

@Baptiste-Arnould

Copy link
Copy Markdown

Summary

This PR parallelizes file metadata inspection during the initialization of the 2-stage reader.

Previously, metadata for all input files were inspected sequentially by rank 0. For datasets containing a large number of files, this can result in significant initialization time.

The metadata inspection is now distributed across the available reader ranks. The resulting metadata are gathered on rank 0 and reordered according to the original file order before constructing _file_time, _shape, _is_real, and _files_size.

The data-reading algorithm itself is unchanged.

Changes

  • Add a helper function to inspect file metadata.
  • Distribute metadata inspection across reader ranks.
  • Gather metadata on rank 0.
  • Restore the original file ordering before constructing reader metadata.

Testing

Tested with MPI on a dataset composed of 11800 files (~250TB in total).

Closes#74

@dalcinl

Copy link
Copy Markdown
Collaborator

Your changes are definitely an improvement and the implementation looks reasonable. There is just a minor thing I would like you to reconsider. You used a cyclic distribution to parallelize. That is OK in general, it is very convenient. But when order is important, it forces you to keep track of indices to reconstruct the original ordering, leading to slightly more convoluted code..

Look at pyspod/utils/parallel.py, note the _blockdist function. It is currently "private" (under the conventional initial underscore). However, we could rename and make it "public", and then use it in in the reader to distribute the metadata reading using a block-contiguous distribution. This way you would not need to sort to restore order, and the code may end up being simpler to read. Let me sketch the code below:

local_infos= []
# Distribute files across reader rankscount, start=utils_par.blockdist(len(data_list), comm.size, comm.rank)
foriinrange(start, start+count):
file_name=data_list[i]
file_shape, file_dtype, file_size=_inspect_file_metadata(file_name, variables[0])
local_infos.append((file_name, file_shape, str(file_dtype), file_size))
# Gather metadata on rank 0all_infos=comm.gather(local_infos, root=0)

and then you need to keep updating a few things to remove the sort call, etc.

Would you be willing to explore this approach and update the PR accordingly?

@Baptiste-Arnould

Copy link
Copy Markdown
Author

Thanks for the suggestion. I updated the PR to use a block distribution for metadata inspection and made blockdist public. This removes the explicit file indices and sorting step and simplifies the reconstruction on rank 0.

@dalcinl

Copy link
Copy Markdown
Collaborator

LGTM.

@mengaldo Once tests pass (modulo the MPI failures that no one had taken care to fix 🤦 ) I think this one is good to go.

@codecov

codecovBot commented Aug 30, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 71.42857% with 14 lines in your changes missing coverage. Please review.
✅ Project coverage is 76.95%. Comparing base (19372d1) to head (1ee2f69).

Files with missing linesPatch %Lines
pyspod/utils/reader.py68.88%13 Missing and 1 partial ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #75 +/- ##
==========================================
+ Coverage 76.84% 76.95% +0.11% 
==========================================
Files 16 16 Lines 3532 3549 +17 Branches 465 467 +2 ==========================================
+ Hits 2714 2731 +17 
Misses 660 660 Partials 158 158 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Parallelize file metadata inspection in 2-stage reader

2 participants

@Baptiste-Arnould@dalcinl