replace openmatrix with h5py - #87

Open
jpn-- wants to merge 9 commits into
ActivitySim:mainfrom
driftlesslabs:codex/replace-openmatrix-with-h5py
Open

replace openmatrix with h5py#87
jpn-- wants to merge 9 commits into
ActivitySim:mainfrom
driftlesslabs:codex/replace-openmatrix-with-h5py

Conversation

@jpn--

@jpn--jpn-- commented Aug 7, 2026

Copy link
Copy Markdown
Member

This pull request removes the use of the openmatrix package throughout the codebase and replaces it with direct usage of the h5py library for reading and writing OMX (Open Matrix) HDF5 files. The change simplifies dependencies, updates documentation and examples, and refactors code to use h5py APIs. Additionally, improvements are made to shared memory handling, including support for custom array storage order and better memory management.

Dependency and Documentation Updates:

  • Removed openmatrix from all dependency files (envs/development.yml, envs/testing.yml, pyproject.toml, .github/workflows/run-tests.yml) and updated code, documentation, and examples to use h5py directly for OMX file operations. [1][2][3][4][5][6][7][8][9]

Code Refactoring:

  • Refactored all code that previously used openmatrix to use h5py, including in sharrow/example_data.py, sharrow/omx.py, and sharrow/tests/test_datasets.py. This includes updating file reading/writing logic and dataset creation. [1][2][3][4][5]

OMX File Handling Improvements:

  • Rewrote the split_omx function in sharrow/omx.py to use h5py for splitting OMX files, added helper functions for copying metadata and datasets, and improved output file organization and lookup management.

Shared Memory Enhancements:

  • Added support for specifying custom array storage order (e.g., "last-axis-first") in shared memory, improved buffer initialization to be more memory efficient, and enhanced cleanup logic for memory-mapped arrays in sharrow/shared_memory.py. [1][2][3][4][5][6][7][8][9][10]

API and Type Annotations:

  • Updated function signatures and docstrings to reflect the removal of openmatrix and clarify the use of h5py objects, especially in sharrow/omx_reader.py. [1][2]

jpn-- added 5 commits August 5, 2026 23:33
Reopen compatible OMX wrappers through h5py while preserving caller ownership.
Add coverage for 2D, 3D, and reload workflows.
Add configurable worker support while preserving last-source precedence and serial fallbacks. Validate worker settings and cover parallel reload behavior.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Replaces OpenMatrix with direct h5py-based OMX handling and adds optimized eager, shared-memory, and memory-mapped loading.

Changes:

  • Refactors OMX reading, writing, splitting, and Zarr conversion.
  • Adds parallel loading and configurable shared-memory array ordering.
  • Removes OpenMatrix dependencies and updates tests and documentation.

Reviewed changes

Copilot reviewed 13 out of 14 changed files in this pull request and generated 3 comments.

Show a summary per file
FileDescription
uv.lockRemoves OpenMatrix and PyTables dependencies.
sharrow/translate.pyUses the shared OMX reader for Zarr conversion.
sharrow/tests/test_omx.pyAdds OMX splitting and conversion tests.
sharrow/tests/test_datasets.pyMigrates tests to h5py and covers loading modes.
sharrow/shared_memory.pyAdds array ordering and memmap cleanup.
sharrow/omx.pyReimplements OMX splitting with h5py.
sharrow/omx_reader.pyNarrows filename handling to h5py inputs.
sharrow/example_data.pyReads example OMX data with h5py.
sharrow/dataset.pyAdds h5py-based loading and parallel storage modes.
pyproject.tomlRemoves the OpenMatrix development dependency.
envs/testing.ymlRemoves OpenMatrix from testing dependencies.
envs/development.ymlRemoves OpenMatrix from development dependencies.
docs/walkthrough/loading-skims.ipynbUpdates the OMX walkthrough to h5py.
.github/workflows/run-tests.ymlRemoves OpenMatrix from CI installation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadsharrow/dataset.py
if load == "shared":
result.shm.release_shared_memory()
else:
result.shm.delete_shared_memory_files(key)
Comment threadsharrow/omx.py
Comment on lines +60 to +67
output_paths = []
for number, matrix_name in enumerate(matrix_names):
output_path = destination / chunk_names[number % len(chunk_names)]
output_paths.append(output_path)
logger.info(f"writing {matrix_name} to {output_path}")
with h5py.File(output_path, "a") as target:
_initialize_omx_file(source, target)
_copy_dataset(source["data"], target["data"], matrix_name)
" omx_file.create_carray(\"/lookup\", \"taz\", obj=zones)\n",
" omx_file.root._v_attrs.SHAPE = np.array([len(zones), len(zones)])"
" omx_file.create_dataset(\"lookup/taz\", data=zones)\n",
" omx_file.attrs[\"SHAPE\"] = np.array([len(zones), len(zones)])"
Truncate split outputs on reruns to remove stale data.
Release failed memmaps before deleting their backing files.
Share a bounded decoder pool across matrix loads and support dtype conversion into noncontiguous outputs.
Infer OMX dimensions when the SHAPE attribute is missing.
Make HDF5 filter plugins optional, use platform-safe parallel loading on Windows, and clean up failed memmap allocations. Add portability tests, documentation, and an isolated OMX memory benchmark.
Compare release and development loaders in isolated environments, tracking runtime and peak memory across loading strategies.
Document setup, caching, source selection, and benchmark options.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

replace openmatrix with h5py - #87

Open
jpn-- wants to merge 9 commits into
ActivitySim:mainfrom
driftlesslabs:codex/replace-openmatrix-with-h5py
Open

replace openmatrix with h5py#87
jpn-- wants to merge 9 commits into
ActivitySim:mainfrom
driftlesslabs:codex/replace-openmatrix-with-h5py

Conversation

@jpn--

@jpn--jpn-- commented Aug 7, 2026

Copy link
Copy Markdown
Member

This pull request removes the use of the openmatrix package throughout the codebase and replaces it with direct usage of the h5py library for reading and writing OMX (Open Matrix) HDF5 files. The change simplifies dependencies, updates documentation and examples, and refactors code to use h5py APIs. Additionally, improvements are made to shared memory handling, including support for custom array storage order and better memory management.

Dependency and Documentation Updates:

  • Removed openmatrix from all dependency files (envs/development.yml, envs/testing.yml, pyproject.toml, .github/workflows/run-tests.yml) and updated code, documentation, and examples to use h5py directly for OMX file operations. [1][2][3][4][5][6][7][8][9]

Code Refactoring:

  • Refactored all code that previously used openmatrix to use h5py, including in sharrow/example_data.py, sharrow/omx.py, and sharrow/tests/test_datasets.py. This includes updating file reading/writing logic and dataset creation. [1][2][3][4][5]

OMX File Handling Improvements:

  • Rewrote the split_omx function in sharrow/omx.py to use h5py for splitting OMX files, added helper functions for copying metadata and datasets, and improved output file organization and lookup management.

Shared Memory Enhancements:

  • Added support for specifying custom array storage order (e.g., "last-axis-first") in shared memory, improved buffer initialization to be more memory efficient, and enhanced cleanup logic for memory-mapped arrays in sharrow/shared_memory.py. [1][2][3][4][5][6][7][8][9][10]

API and Type Annotations:

  • Updated function signatures and docstrings to reflect the removal of openmatrix and clarify the use of h5py objects, especially in sharrow/omx_reader.py. [1][2]

jpn-- added 5 commits August 5, 2026 23:33
Reopen compatible OMX wrappers through h5py while preserving caller ownership.
Add coverage for 2D, 3D, and reload workflows.
Add configurable worker support while preserving last-source precedence and serial fallbacks. Validate worker settings and cover parallel reload behavior.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Replaces OpenMatrix with direct h5py-based OMX handling and adds optimized eager, shared-memory, and memory-mapped loading.

Changes:

  • Refactors OMX reading, writing, splitting, and Zarr conversion.
  • Adds parallel loading and configurable shared-memory array ordering.
  • Removes OpenMatrix dependencies and updates tests and documentation.

Reviewed changes

Copilot reviewed 13 out of 14 changed files in this pull request and generated 3 comments.

Show a summary per file
FileDescription
uv.lockRemoves OpenMatrix and PyTables dependencies.
sharrow/translate.pyUses the shared OMX reader for Zarr conversion.
sharrow/tests/test_omx.pyAdds OMX splitting and conversion tests.
sharrow/tests/test_datasets.pyMigrates tests to h5py and covers loading modes.
sharrow/shared_memory.pyAdds array ordering and memmap cleanup.
sharrow/omx.pyReimplements OMX splitting with h5py.
sharrow/omx_reader.pyNarrows filename handling to h5py inputs.
sharrow/example_data.pyReads example OMX data with h5py.
sharrow/dataset.pyAdds h5py-based loading and parallel storage modes.
pyproject.tomlRemoves the OpenMatrix development dependency.
envs/testing.ymlRemoves OpenMatrix from testing dependencies.
envs/development.ymlRemoves OpenMatrix from development dependencies.
docs/walkthrough/loading-skims.ipynbUpdates the OMX walkthrough to h5py.
.github/workflows/run-tests.ymlRemoves OpenMatrix from CI installation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadsharrow/dataset.py
if load == "shared":
result.shm.release_shared_memory()
else:
result.shm.delete_shared_memory_files(key)
Comment threadsharrow/omx.py
Comment on lines +60 to +67
output_paths = []
for number, matrix_name in enumerate(matrix_names):
output_path = destination / chunk_names[number % len(chunk_names)]
output_paths.append(output_path)
logger.info(f"writing {matrix_name} to {output_path}")
with h5py.File(output_path, "a") as target:
_initialize_omx_file(source, target)
_copy_dataset(source["data"], target["data"], matrix_name)
" omx_file.create_carray(\"/lookup\", \"taz\", obj=zones)\n",
" omx_file.root._v_attrs.SHAPE = np.array([len(zones), len(zones)])"
" omx_file.create_dataset(\"lookup/taz\", data=zones)\n",
" omx_file.attrs[\"SHAPE\"] = np.array([len(zones), len(zones)])"
Truncate split outputs on reruns to remove stale data.
Release failed memmaps before deleting their backing files.
Share a bounded decoder pool across matrix loads and support dtype conversion into noncontiguous outputs.
Infer OMX dimensions when the SHAPE attribute is missing.
Make HDF5 filter plugins optional, use platform-safe parallel loading on Windows, and clean up failed memmap allocations. Add portability tests, documentation, and an isolated OMX memory benchmark.
Compare release and development loaders in isolated environments, tracking runtime and peak memory across loading strategies.
Document setup, caching, source selection, and benchmark options.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

replace openmatrix with h5py - #87

Open
jpn-- wants to merge 9 commits into
ActivitySim:mainfrom
driftlesslabs:codex/replace-openmatrix-with-h5py
Open

replace openmatrix with h5py#87
jpn-- wants to merge 9 commits into
ActivitySim:mainfrom
driftlesslabs:codex/replace-openmatrix-with-h5py

Conversation

@jpn--

@jpn--jpn-- commented Aug 7, 2026

Copy link
Copy Markdown
Member

This pull request removes the use of the openmatrix package throughout the codebase and replaces it with direct usage of the h5py library for reading and writing OMX (Open Matrix) HDF5 files. The change simplifies dependencies, updates documentation and examples, and refactors code to use h5py APIs. Additionally, improvements are made to shared memory handling, including support for custom array storage order and better memory management.

Dependency and Documentation Updates:

  • Removed openmatrix from all dependency files (envs/development.yml, envs/testing.yml, pyproject.toml, .github/workflows/run-tests.yml) and updated code, documentation, and examples to use h5py directly for OMX file operations. [1][2][3][4][5][6][7][8][9]

Code Refactoring:

  • Refactored all code that previously used openmatrix to use h5py, including in sharrow/example_data.py, sharrow/omx.py, and sharrow/tests/test_datasets.py. This includes updating file reading/writing logic and dataset creation. [1][2][3][4][5]

OMX File Handling Improvements:

  • Rewrote the split_omx function in sharrow/omx.py to use h5py for splitting OMX files, added helper functions for copying metadata and datasets, and improved output file organization and lookup management.

Shared Memory Enhancements:

  • Added support for specifying custom array storage order (e.g., "last-axis-first") in shared memory, improved buffer initialization to be more memory efficient, and enhanced cleanup logic for memory-mapped arrays in sharrow/shared_memory.py. [1][2][3][4][5][6][7][8][9][10]

API and Type Annotations:

  • Updated function signatures and docstrings to reflect the removal of openmatrix and clarify the use of h5py objects, especially in sharrow/omx_reader.py. [1][2]

jpn-- added 5 commits August 5, 2026 23:33
Reopen compatible OMX wrappers through h5py while preserving caller ownership.
Add coverage for 2D, 3D, and reload workflows.
Add configurable worker support while preserving last-source precedence and serial fallbacks. Validate worker settings and cover parallel reload behavior.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Replaces OpenMatrix with direct h5py-based OMX handling and adds optimized eager, shared-memory, and memory-mapped loading.

Changes:

  • Refactors OMX reading, writing, splitting, and Zarr conversion.
  • Adds parallel loading and configurable shared-memory array ordering.
  • Removes OpenMatrix dependencies and updates tests and documentation.

Reviewed changes

Copilot reviewed 13 out of 14 changed files in this pull request and generated 3 comments.

Show a summary per file
FileDescription
uv.lockRemoves OpenMatrix and PyTables dependencies.
sharrow/translate.pyUses the shared OMX reader for Zarr conversion.
sharrow/tests/test_omx.pyAdds OMX splitting and conversion tests.
sharrow/tests/test_datasets.pyMigrates tests to h5py and covers loading modes.
sharrow/shared_memory.pyAdds array ordering and memmap cleanup.
sharrow/omx.pyReimplements OMX splitting with h5py.
sharrow/omx_reader.pyNarrows filename handling to h5py inputs.
sharrow/example_data.pyReads example OMX data with h5py.
sharrow/dataset.pyAdds h5py-based loading and parallel storage modes.
pyproject.tomlRemoves the OpenMatrix development dependency.
envs/testing.ymlRemoves OpenMatrix from testing dependencies.
envs/development.ymlRemoves OpenMatrix from development dependencies.
docs/walkthrough/loading-skims.ipynbUpdates the OMX walkthrough to h5py.
.github/workflows/run-tests.ymlRemoves OpenMatrix from CI installation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadsharrow/dataset.py
if load == "shared":
result.shm.release_shared_memory()
else:
result.shm.delete_shared_memory_files(key)
Comment threadsharrow/omx.py
Comment on lines +60 to +67
output_paths = []
for number, matrix_name in enumerate(matrix_names):
output_path = destination / chunk_names[number % len(chunk_names)]
output_paths.append(output_path)
logger.info(f"writing {matrix_name} to {output_path}")
with h5py.File(output_path, "a") as target:
_initialize_omx_file(source, target)
_copy_dataset(source["data"], target["data"], matrix_name)
" omx_file.create_carray(\"/lookup\", \"taz\", obj=zones)\n",
" omx_file.root._v_attrs.SHAPE = np.array([len(zones), len(zones)])"
" omx_file.create_dataset(\"lookup/taz\", data=zones)\n",
" omx_file.attrs[\"SHAPE\"] = np.array([len(zones), len(zones)])"
Truncate split outputs on reruns to remove stale data.
Release failed memmaps before deleting their backing files.
Share a bounded decoder pool across matrix loads and support dtype conversion into noncontiguous outputs.
Infer OMX dimensions when the SHAPE attribute is missing.
Make HDF5 filter plugins optional, use platform-safe parallel loading on Windows, and clean up failed memmap allocations. Add portability tests, documentation, and an isolated OMX memory benchmark.
Compare release and development loaders in isolated environments, tracking runtime and peak memory across loading strategies.
Document setup, caching, source selection, and benchmark options.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

replace openmatrix with h5py - #87

Open
jpn-- wants to merge 9 commits into
ActivitySim:mainfrom
driftlesslabs:codex/replace-openmatrix-with-h5py
Open

replace openmatrix with h5py#87
jpn-- wants to merge 9 commits into
ActivitySim:mainfrom
driftlesslabs:codex/replace-openmatrix-with-h5py

Conversation

@jpn--

@jpn--jpn-- commented Aug 7, 2026

Copy link
Copy Markdown
Member

This pull request removes the use of the openmatrix package throughout the codebase and replaces it with direct usage of the h5py library for reading and writing OMX (Open Matrix) HDF5 files. The change simplifies dependencies, updates documentation and examples, and refactors code to use h5py APIs. Additionally, improvements are made to shared memory handling, including support for custom array storage order and better memory management.

Dependency and Documentation Updates:

  • Removed openmatrix from all dependency files (envs/development.yml, envs/testing.yml, pyproject.toml, .github/workflows/run-tests.yml) and updated code, documentation, and examples to use h5py directly for OMX file operations. [1][2][3][4][5][6][7][8][9]

Code Refactoring:

  • Refactored all code that previously used openmatrix to use h5py, including in sharrow/example_data.py, sharrow/omx.py, and sharrow/tests/test_datasets.py. This includes updating file reading/writing logic and dataset creation. [1][2][3][4][5]

OMX File Handling Improvements:

  • Rewrote the split_omx function in sharrow/omx.py to use h5py for splitting OMX files, added helper functions for copying metadata and datasets, and improved output file organization and lookup management.

Shared Memory Enhancements:

  • Added support for specifying custom array storage order (e.g., "last-axis-first") in shared memory, improved buffer initialization to be more memory efficient, and enhanced cleanup logic for memory-mapped arrays in sharrow/shared_memory.py. [1][2][3][4][5][6][7][8][9][10]

API and Type Annotations:

  • Updated function signatures and docstrings to reflect the removal of openmatrix and clarify the use of h5py objects, especially in sharrow/omx_reader.py. [1][2]

jpn-- added 5 commits August 5, 2026 23:33
Reopen compatible OMX wrappers through h5py while preserving caller ownership.
Add coverage for 2D, 3D, and reload workflows.
Add configurable worker support while preserving last-source precedence and serial fallbacks. Validate worker settings and cover parallel reload behavior.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Replaces OpenMatrix with direct h5py-based OMX handling and adds optimized eager, shared-memory, and memory-mapped loading.

Changes:

  • Refactors OMX reading, writing, splitting, and Zarr conversion.
  • Adds parallel loading and configurable shared-memory array ordering.
  • Removes OpenMatrix dependencies and updates tests and documentation.

Reviewed changes

Copilot reviewed 13 out of 14 changed files in this pull request and generated 3 comments.

Show a summary per file
FileDescription
uv.lockRemoves OpenMatrix and PyTables dependencies.
sharrow/translate.pyUses the shared OMX reader for Zarr conversion.
sharrow/tests/test_omx.pyAdds OMX splitting and conversion tests.
sharrow/tests/test_datasets.pyMigrates tests to h5py and covers loading modes.
sharrow/shared_memory.pyAdds array ordering and memmap cleanup.
sharrow/omx.pyReimplements OMX splitting with h5py.
sharrow/omx_reader.pyNarrows filename handling to h5py inputs.
sharrow/example_data.pyReads example OMX data with h5py.
sharrow/dataset.pyAdds h5py-based loading and parallel storage modes.
pyproject.tomlRemoves the OpenMatrix development dependency.
envs/testing.ymlRemoves OpenMatrix from testing dependencies.
envs/development.ymlRemoves OpenMatrix from development dependencies.
docs/walkthrough/loading-skims.ipynbUpdates the OMX walkthrough to h5py.
.github/workflows/run-tests.ymlRemoves OpenMatrix from CI installation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadsharrow/dataset.py
if load == "shared":
result.shm.release_shared_memory()
else:
result.shm.delete_shared_memory_files(key)
Comment threadsharrow/omx.py
Comment on lines +60 to +67
output_paths = []
for number, matrix_name in enumerate(matrix_names):
output_path = destination / chunk_names[number % len(chunk_names)]
output_paths.append(output_path)
logger.info(f"writing {matrix_name} to {output_path}")
with h5py.File(output_path, "a") as target:
_initialize_omx_file(source, target)
_copy_dataset(source["data"], target["data"], matrix_name)
" omx_file.create_carray(\"/lookup\", \"taz\", obj=zones)\n",
" omx_file.root._v_attrs.SHAPE = np.array([len(zones), len(zones)])"
" omx_file.create_dataset(\"lookup/taz\", data=zones)\n",
" omx_file.attrs[\"SHAPE\"] = np.array([len(zones), len(zones)])"
Truncate split outputs on reruns to remove stale data.
Release failed memmaps before deleting their backing files.
Share a bounded decoder pool across matrix loads and support dtype conversion into noncontiguous outputs.
Infer OMX dimensions when the SHAPE attribute is missing.
Make HDF5 filter plugins optional, use platform-safe parallel loading on Windows, and clean up failed memmap allocations. Add portability tests, documentation, and an isolated OMX memory benchmark.
Compare release and development loaders in isolated environments, tracking runtime and peak memory across loading strategies.
Document setup, caching, source selection, and benchmark options.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

replace openmatrix with h5py - #87

Open
jpn-- wants to merge 9 commits into
ActivitySim:mainfrom
driftlesslabs:codex/replace-openmatrix-with-h5py
Open

replace openmatrix with h5py#87
jpn-- wants to merge 9 commits into
ActivitySim:mainfrom
driftlesslabs:codex/replace-openmatrix-with-h5py

Conversation

@jpn--

@jpn--jpn-- commented Aug 7, 2026

Copy link
Copy Markdown
Member

This pull request removes the use of the openmatrix package throughout the codebase and replaces it with direct usage of the h5py library for reading and writing OMX (Open Matrix) HDF5 files. The change simplifies dependencies, updates documentation and examples, and refactors code to use h5py APIs. Additionally, improvements are made to shared memory handling, including support for custom array storage order and better memory management.

Dependency and Documentation Updates:

  • Removed openmatrix from all dependency files (envs/development.yml, envs/testing.yml, pyproject.toml, .github/workflows/run-tests.yml) and updated code, documentation, and examples to use h5py directly for OMX file operations. [1][2][3][4][5][6][7][8][9]

Code Refactoring:

  • Refactored all code that previously used openmatrix to use h5py, including in sharrow/example_data.py, sharrow/omx.py, and sharrow/tests/test_datasets.py. This includes updating file reading/writing logic and dataset creation. [1][2][3][4][5]

OMX File Handling Improvements:

  • Rewrote the split_omx function in sharrow/omx.py to use h5py for splitting OMX files, added helper functions for copying metadata and datasets, and improved output file organization and lookup management.

Shared Memory Enhancements:

  • Added support for specifying custom array storage order (e.g., "last-axis-first") in shared memory, improved buffer initialization to be more memory efficient, and enhanced cleanup logic for memory-mapped arrays in sharrow/shared_memory.py. [1][2][3][4][5][6][7][8][9][10]

API and Type Annotations:

  • Updated function signatures and docstrings to reflect the removal of openmatrix and clarify the use of h5py objects, especially in sharrow/omx_reader.py. [1][2]

jpn-- added 5 commits August 5, 2026 23:33
Reopen compatible OMX wrappers through h5py while preserving caller ownership.
Add coverage for 2D, 3D, and reload workflows.
Add configurable worker support while preserving last-source precedence and serial fallbacks. Validate worker settings and cover parallel reload behavior.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Replaces OpenMatrix with direct h5py-based OMX handling and adds optimized eager, shared-memory, and memory-mapped loading.

Changes:

  • Refactors OMX reading, writing, splitting, and Zarr conversion.
  • Adds parallel loading and configurable shared-memory array ordering.
  • Removes OpenMatrix dependencies and updates tests and documentation.

Reviewed changes

Copilot reviewed 13 out of 14 changed files in this pull request and generated 3 comments.

Show a summary per file
FileDescription
uv.lockRemoves OpenMatrix and PyTables dependencies.
sharrow/translate.pyUses the shared OMX reader for Zarr conversion.
sharrow/tests/test_omx.pyAdds OMX splitting and conversion tests.
sharrow/tests/test_datasets.pyMigrates tests to h5py and covers loading modes.
sharrow/shared_memory.pyAdds array ordering and memmap cleanup.
sharrow/omx.pyReimplements OMX splitting with h5py.
sharrow/omx_reader.pyNarrows filename handling to h5py inputs.
sharrow/example_data.pyReads example OMX data with h5py.
sharrow/dataset.pyAdds h5py-based loading and parallel storage modes.
pyproject.tomlRemoves the OpenMatrix development dependency.
envs/testing.ymlRemoves OpenMatrix from testing dependencies.
envs/development.ymlRemoves OpenMatrix from development dependencies.
docs/walkthrough/loading-skims.ipynbUpdates the OMX walkthrough to h5py.
.github/workflows/run-tests.ymlRemoves OpenMatrix from CI installation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadsharrow/dataset.py
if load == "shared":
result.shm.release_shared_memory()
else:
result.shm.delete_shared_memory_files(key)
Comment threadsharrow/omx.py
Comment on lines +60 to +67
output_paths = []
for number, matrix_name in enumerate(matrix_names):
output_path = destination / chunk_names[number % len(chunk_names)]
output_paths.append(output_path)
logger.info(f"writing {matrix_name} to {output_path}")
with h5py.File(output_path, "a") as target:
_initialize_omx_file(source, target)
_copy_dataset(source["data"], target["data"], matrix_name)
" omx_file.create_carray(\"/lookup\", \"taz\", obj=zones)\n",
" omx_file.root._v_attrs.SHAPE = np.array([len(zones), len(zones)])"
" omx_file.create_dataset(\"lookup/taz\", data=zones)\n",
" omx_file.attrs[\"SHAPE\"] = np.array([len(zones), len(zones)])"
Truncate split outputs on reruns to remove stale data.
Release failed memmaps before deleting their backing files.
Share a bounded decoder pool across matrix loads and support dtype conversion into noncontiguous outputs.
Infer OMX dimensions when the SHAPE attribute is missing.
Make HDF5 filter plugins optional, use platform-safe parallel loading on Windows, and clean up failed memmap allocations. Add portability tests, documentation, and an isolated OMX memory benchmark.
Compare release and development loaders in isolated environments, tracking runtime and peak memory across loading strategies.
Document setup, caching, source selection, and benchmark options.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

replace openmatrix with h5py - #87

Open
jpn-- wants to merge 9 commits into
ActivitySim:mainfrom
driftlesslabs:codex/replace-openmatrix-with-h5py
Open

replace openmatrix with h5py#87
jpn-- wants to merge 9 commits into
ActivitySim:mainfrom
driftlesslabs:codex/replace-openmatrix-with-h5py

Conversation

@jpn--

@jpn--jpn-- commented Aug 7, 2026

Copy link
Copy Markdown
Member

This pull request removes the use of the openmatrix package throughout the codebase and replaces it with direct usage of the h5py library for reading and writing OMX (Open Matrix) HDF5 files. The change simplifies dependencies, updates documentation and examples, and refactors code to use h5py APIs. Additionally, improvements are made to shared memory handling, including support for custom array storage order and better memory management.

Dependency and Documentation Updates:

  • Removed openmatrix from all dependency files (envs/development.yml, envs/testing.yml, pyproject.toml, .github/workflows/run-tests.yml) and updated code, documentation, and examples to use h5py directly for OMX file operations. [1][2][3][4][5][6][7][8][9]

Code Refactoring:

  • Refactored all code that previously used openmatrix to use h5py, including in sharrow/example_data.py, sharrow/omx.py, and sharrow/tests/test_datasets.py. This includes updating file reading/writing logic and dataset creation. [1][2][3][4][5]

OMX File Handling Improvements:

  • Rewrote the split_omx function in sharrow/omx.py to use h5py for splitting OMX files, added helper functions for copying metadata and datasets, and improved output file organization and lookup management.

Shared Memory Enhancements:

  • Added support for specifying custom array storage order (e.g., "last-axis-first") in shared memory, improved buffer initialization to be more memory efficient, and enhanced cleanup logic for memory-mapped arrays in sharrow/shared_memory.py. [1][2][3][4][5][6][7][8][9][10]

API and Type Annotations:

  • Updated function signatures and docstrings to reflect the removal of openmatrix and clarify the use of h5py objects, especially in sharrow/omx_reader.py. [1][2]

jpn-- added 5 commits August 5, 2026 23:33
Reopen compatible OMX wrappers through h5py while preserving caller ownership.
Add coverage for 2D, 3D, and reload workflows.
Add configurable worker support while preserving last-source precedence and serial fallbacks. Validate worker settings and cover parallel reload behavior.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Replaces OpenMatrix with direct h5py-based OMX handling and adds optimized eager, shared-memory, and memory-mapped loading.

Changes:

  • Refactors OMX reading, writing, splitting, and Zarr conversion.
  • Adds parallel loading and configurable shared-memory array ordering.
  • Removes OpenMatrix dependencies and updates tests and documentation.

Reviewed changes

Copilot reviewed 13 out of 14 changed files in this pull request and generated 3 comments.

Show a summary per file
FileDescription
uv.lockRemoves OpenMatrix and PyTables dependencies.
sharrow/translate.pyUses the shared OMX reader for Zarr conversion.
sharrow/tests/test_omx.pyAdds OMX splitting and conversion tests.
sharrow/tests/test_datasets.pyMigrates tests to h5py and covers loading modes.
sharrow/shared_memory.pyAdds array ordering and memmap cleanup.
sharrow/omx.pyReimplements OMX splitting with h5py.
sharrow/omx_reader.pyNarrows filename handling to h5py inputs.
sharrow/example_data.pyReads example OMX data with h5py.
sharrow/dataset.pyAdds h5py-based loading and parallel storage modes.
pyproject.tomlRemoves the OpenMatrix development dependency.
envs/testing.ymlRemoves OpenMatrix from testing dependencies.
envs/development.ymlRemoves OpenMatrix from development dependencies.
docs/walkthrough/loading-skims.ipynbUpdates the OMX walkthrough to h5py.
.github/workflows/run-tests.ymlRemoves OpenMatrix from CI installation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadsharrow/dataset.py
if load == "shared":
result.shm.release_shared_memory()
else:
result.shm.delete_shared_memory_files(key)
Comment threadsharrow/omx.py
Comment on lines +60 to +67
output_paths = []
for number, matrix_name in enumerate(matrix_names):
output_path = destination / chunk_names[number % len(chunk_names)]
output_paths.append(output_path)
logger.info(f"writing {matrix_name} to {output_path}")
with h5py.File(output_path, "a") as target:
_initialize_omx_file(source, target)
_copy_dataset(source["data"], target["data"], matrix_name)
" omx_file.create_carray(\"/lookup\", \"taz\", obj=zones)\n",
" omx_file.root._v_attrs.SHAPE = np.array([len(zones), len(zones)])"
" omx_file.create_dataset(\"lookup/taz\", data=zones)\n",
" omx_file.attrs[\"SHAPE\"] = np.array([len(zones), len(zones)])"
Truncate split outputs on reruns to remove stale data.
Release failed memmaps before deleting their backing files.
Share a bounded decoder pool across matrix loads and support dtype conversion into noncontiguous outputs.
Infer OMX dimensions when the SHAPE attribute is missing.
Make HDF5 filter plugins optional, use platform-safe parallel loading on Windows, and clean up failed memmap allocations. Add portability tests, documentation, and an isolated OMX memory benchmark.
Compare release and development loaders in isolated environments, tracking runtime and peak memory across loading strategies.
Document setup, caching, source selection, and benchmark options.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

replace openmatrix with h5py - #87

Open
jpn-- wants to merge 9 commits into
ActivitySim:mainfrom
driftlesslabs:codex/replace-openmatrix-with-h5py
Open

replace openmatrix with h5py#87
jpn-- wants to merge 9 commits into
ActivitySim:mainfrom
driftlesslabs:codex/replace-openmatrix-with-h5py

Conversation

@jpn--

@jpn--jpn-- commented Aug 7, 2026

Copy link
Copy Markdown
Member

This pull request removes the use of the openmatrix package throughout the codebase and replaces it with direct usage of the h5py library for reading and writing OMX (Open Matrix) HDF5 files. The change simplifies dependencies, updates documentation and examples, and refactors code to use h5py APIs. Additionally, improvements are made to shared memory handling, including support for custom array storage order and better memory management.

Dependency and Documentation Updates:

  • Removed openmatrix from all dependency files (envs/development.yml, envs/testing.yml, pyproject.toml, .github/workflows/run-tests.yml) and updated code, documentation, and examples to use h5py directly for OMX file operations. [1][2][3][4][5][6][7][8][9]

Code Refactoring:

  • Refactored all code that previously used openmatrix to use h5py, including in sharrow/example_data.py, sharrow/omx.py, and sharrow/tests/test_datasets.py. This includes updating file reading/writing logic and dataset creation. [1][2][3][4][5]

OMX File Handling Improvements:

  • Rewrote the split_omx function in sharrow/omx.py to use h5py for splitting OMX files, added helper functions for copying metadata and datasets, and improved output file organization and lookup management.

Shared Memory Enhancements:

  • Added support for specifying custom array storage order (e.g., "last-axis-first") in shared memory, improved buffer initialization to be more memory efficient, and enhanced cleanup logic for memory-mapped arrays in sharrow/shared_memory.py. [1][2][3][4][5][6][7][8][9][10]

API and Type Annotations:

  • Updated function signatures and docstrings to reflect the removal of openmatrix and clarify the use of h5py objects, especially in sharrow/omx_reader.py. [1][2]

jpn-- added 5 commits August 5, 2026 23:33
Reopen compatible OMX wrappers through h5py while preserving caller ownership.
Add coverage for 2D, 3D, and reload workflows.
Add configurable worker support while preserving last-source precedence and serial fallbacks. Validate worker settings and cover parallel reload behavior.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Replaces OpenMatrix with direct h5py-based OMX handling and adds optimized eager, shared-memory, and memory-mapped loading.

Changes:

  • Refactors OMX reading, writing, splitting, and Zarr conversion.
  • Adds parallel loading and configurable shared-memory array ordering.
  • Removes OpenMatrix dependencies and updates tests and documentation.

Reviewed changes

Copilot reviewed 13 out of 14 changed files in this pull request and generated 3 comments.

Show a summary per file
FileDescription
uv.lockRemoves OpenMatrix and PyTables dependencies.
sharrow/translate.pyUses the shared OMX reader for Zarr conversion.
sharrow/tests/test_omx.pyAdds OMX splitting and conversion tests.
sharrow/tests/test_datasets.pyMigrates tests to h5py and covers loading modes.
sharrow/shared_memory.pyAdds array ordering and memmap cleanup.
sharrow/omx.pyReimplements OMX splitting with h5py.
sharrow/omx_reader.pyNarrows filename handling to h5py inputs.
sharrow/example_data.pyReads example OMX data with h5py.
sharrow/dataset.pyAdds h5py-based loading and parallel storage modes.
pyproject.tomlRemoves the OpenMatrix development dependency.
envs/testing.ymlRemoves OpenMatrix from testing dependencies.
envs/development.ymlRemoves OpenMatrix from development dependencies.
docs/walkthrough/loading-skims.ipynbUpdates the OMX walkthrough to h5py.
.github/workflows/run-tests.ymlRemoves OpenMatrix from CI installation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadsharrow/dataset.py
if load == "shared":
result.shm.release_shared_memory()
else:
result.shm.delete_shared_memory_files(key)
Comment threadsharrow/omx.py
Comment on lines +60 to +67
output_paths = []
for number, matrix_name in enumerate(matrix_names):
output_path = destination / chunk_names[number % len(chunk_names)]
output_paths.append(output_path)
logger.info(f"writing {matrix_name} to {output_path}")
with h5py.File(output_path, "a") as target:
_initialize_omx_file(source, target)
_copy_dataset(source["data"], target["data"], matrix_name)
" omx_file.create_carray(\"/lookup\", \"taz\", obj=zones)\n",
" omx_file.root._v_attrs.SHAPE = np.array([len(zones), len(zones)])"
" omx_file.create_dataset(\"lookup/taz\", data=zones)\n",
" omx_file.attrs[\"SHAPE\"] = np.array([len(zones), len(zones)])"
Truncate split outputs on reruns to remove stale data.
Release failed memmaps before deleting their backing files.
Share a bounded decoder pool across matrix loads and support dtype conversion into noncontiguous outputs.
Infer OMX dimensions when the SHAPE attribute is missing.
Make HDF5 filter plugins optional, use platform-safe parallel loading on Windows, and clean up failed memmap allocations. Add portability tests, documentation, and an isolated OMX memory benchmark.
Compare release and development loaders in isolated environments, tracking runtime and peak memory across loading strategies.
Document setup, caching, source selection, and benchmark options.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jpn--
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

replace openmatrix with h5py - #87

Open
jpn-- wants to merge 9 commits into
ActivitySim:mainfrom
driftlesslabs:codex/replace-openmatrix-with-h5py
Open

replace openmatrix with h5py#87
jpn-- wants to merge 9 commits into
ActivitySim:mainfrom
driftlesslabs:codex/replace-openmatrix-with-h5py

Conversation

@jpn--

@jpn--jpn-- commented Aug 7, 2026

Copy link
Copy Markdown
Member

This pull request removes the use of the openmatrix package throughout the codebase and replaces it with direct usage of the h5py library for reading and writing OMX (Open Matrix) HDF5 files. The change simplifies dependencies, updates documentation and examples, and refactors code to use h5py APIs. Additionally, improvements are made to shared memory handling, including support for custom array storage order and better memory management.

Dependency and Documentation Updates:

  • Removed openmatrix from all dependency files (envs/development.yml, envs/testing.yml, pyproject.toml, .github/workflows/run-tests.yml) and updated code, documentation, and examples to use h5py directly for OMX file operations. [1][2][3][4][5][6][7][8][9]

Code Refactoring:

  • Refactored all code that previously used openmatrix to use h5py, including in sharrow/example_data.py, sharrow/omx.py, and sharrow/tests/test_datasets.py. This includes updating file reading/writing logic and dataset creation. [1][2][3][4][5]

OMX File Handling Improvements:

  • Rewrote the split_omx function in sharrow/omx.py to use h5py for splitting OMX files, added helper functions for copying metadata and datasets, and improved output file organization and lookup management.

Shared Memory Enhancements:

  • Added support for specifying custom array storage order (e.g., "last-axis-first") in shared memory, improved buffer initialization to be more memory efficient, and enhanced cleanup logic for memory-mapped arrays in sharrow/shared_memory.py. [1][2][3][4][5][6][7][8][9][10]

API and Type Annotations:

  • Updated function signatures and docstrings to reflect the removal of openmatrix and clarify the use of h5py objects, especially in sharrow/omx_reader.py. [1][2]

jpn-- added 5 commits August 5, 2026 23:33
Reopen compatible OMX wrappers through h5py while preserving caller ownership.
Add coverage for 2D, 3D, and reload workflows.
Add configurable worker support while preserving last-source precedence and serial fallbacks. Validate worker settings and cover parallel reload behavior.

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Replaces OpenMatrix with direct h5py-based OMX handling and adds optimized eager, shared-memory, and memory-mapped loading.

Changes:

  • Refactors OMX reading, writing, splitting, and Zarr conversion.
  • Adds parallel loading and configurable shared-memory array ordering.
  • Removes OpenMatrix dependencies and updates tests and documentation.

Reviewed changes

Copilot reviewed 13 out of 14 changed files in this pull request and generated 3 comments.

Show a summary per file
FileDescription
uv.lockRemoves OpenMatrix and PyTables dependencies.
sharrow/translate.pyUses the shared OMX reader for Zarr conversion.
sharrow/tests/test_omx.pyAdds OMX splitting and conversion tests.
sharrow/tests/test_datasets.pyMigrates tests to h5py and covers loading modes.
sharrow/shared_memory.pyAdds array ordering and memmap cleanup.
sharrow/omx.pyReimplements OMX splitting with h5py.
sharrow/omx_reader.pyNarrows filename handling to h5py inputs.
sharrow/example_data.pyReads example OMX data with h5py.
sharrow/dataset.pyAdds h5py-based loading and parallel storage modes.
pyproject.tomlRemoves the OpenMatrix development dependency.
envs/testing.ymlRemoves OpenMatrix from testing dependencies.
envs/development.ymlRemoves OpenMatrix from development dependencies.
docs/walkthrough/loading-skims.ipynbUpdates the OMX walkthrough to h5py.
.github/workflows/run-tests.ymlRemoves OpenMatrix from CI installation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment threadsharrow/dataset.py
if load == "shared":
result.shm.release_shared_memory()
else:
result.shm.delete_shared_memory_files(key)
Comment threadsharrow/omx.py
Comment on lines +60 to +67
output_paths = []
for number, matrix_name in enumerate(matrix_names):
output_path = destination / chunk_names[number % len(chunk_names)]
output_paths.append(output_path)
logger.info(f"writing {matrix_name} to {output_path}")
with h5py.File(output_path, "a") as target:
_initialize_omx_file(source, target)
_copy_dataset(source["data"], target["data"], matrix_name)
" omx_file.create_carray(\"/lookup\", \"taz\", obj=zones)\n",
" omx_file.root._v_attrs.SHAPE = np.array([len(zones), len(zones)])"
" omx_file.create_dataset(\"lookup/taz\", data=zones)\n",
" omx_file.attrs[\"SHAPE\"] = np.array([len(zones), len(zones)])"
Truncate split outputs on reruns to remove stale data.
Release failed memmaps before deleting their backing files.
Share a bounded decoder pool across matrix loads and support dtype conversion into noncontiguous outputs.
Infer OMX dimensions when the SHAPE attribute is missing.
Make HDF5 filter plugins optional, use platform-safe parallel loading on Windows, and clean up failed memmap allocations. Add portability tests, documentation, and an isolated OMX memory benchmark.
Compare release and development loaders in isolated environments, tracking runtime and peak memory across loading strategies.
Document setup, caching, source selection, and benchmark options.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@jpn--