Implement .to_chunk_cached_arrays() with ChunkCachedArray - #2860

Open
VeckoTheGecko wants to merge 13 commits into
Parcels-code:mainfrom
VeckoTheGecko:2854-dask-chunk-caching
Open

Implement .to_chunk_cached_arrays() with ChunkCachedArray#2860
VeckoTheGecko wants to merge 13 commits into
Parcels-code:mainfrom
VeckoTheGecko:2854-dask-chunk-caching

Conversation

@VeckoTheGecko

Copy link
Copy Markdown
Contributor

Description

Provides an opt-in optimization that wraps dask-backed field data in
chunk-level LRU caches.
vectorized .isel() calls hit an in-memory cache instead of recomputing
dask task graphs, giving large speedups for particle simulations.

This PR adds ChunkCachedArray according to the approach detailed in #2854 , providing an opt-in optimization wrapping dash-backed field data with (dask)chunk-level LRU caches on the array level. This PR:

  • Adds a subpackage parcels._chunk_cached_array with the implementation of the Array and the needed LRU cache. Note the cache is on the array level with configurable memory limits per array.
  • Exposes a function parcels._chunk_cached_array.wrap_dataset providing the main entrypoint for Parcels for wrapping Xarray dataset objects
  • Adds .to_chunk_cached_arrays to FieldSet and ModelData classes.
  • Adds equivalence testing (test_backends.py) on a NEMO dataset, asserting that the same results are received regardless of backend.

Still needed:

  • (separate PR) Upstream testing
  • (separate PR?) Backend equivalency testing on Unstructured grids?
  • Benchmarking with realworld simulations
  • (pre-merge) Remove temporary testing files (data-generation.py,benchmark_chunk_cache.py,benchmark_chunk_cache.png)

Future work:

  • max_cache_bytes on FieldSet.to_chunk_cached_arrays operates per variable (i.e., per Field). This requires the user to think about how many Fields are in their simulation in order to tune their performance. Down the line, we could tune this to allocate a total cache of (e.g.,) 50% of the system memory (which would allow for other memory overhead of coordinates (not stored in the cache), the Python process, or other user processes). This would require significant refactoring not only on how the chunk cache is handled, but also backends in general.

(cc @erikvansebille, @wyatt-fluidnumerics )

Performance

I have some some preliminary profiling with a 20Gb idealised dataset, and found the following performance profile which looks like a promising first step.

image

I've included data-generation.py and benchmark_chunk_cache.py here for testing in case it helps.

Its not clear to me the performance with real-world simulations and how that scales (I only have access to my laptop, and found working/debugging on Lorenz quite frustrating with my high latency from travels). Let me know if there's anything I can do to help here @erikvansebille .

Checklist

AI Disclosure

  • This PR contains AI-generated content.
    • I have tested any AI-generated content in my PR.
    • I take responsibility for any AI-generated content in my PR.
    • Describe how you used it (e.g., by pasting your prompt): I guided an LLM to explore the Xarray codebase, and to help work on the implementation of the Array class. I also used an LLM to generate and iterate on the profiling script, and write docstrings.

Provides an opt-in optimization that wraps dask-backed field data in
chunk-level LRU caches via the chunk_cached_array package. Repeated
vectorized .isel() calls hit an in-memory cache instead of recomputing
dask task graphs, giving large speedups for particle simulations.
Results on ds_2d_left_agrid.zarr with 10k particles:
- Plain dask: 85.0s (1x)
- Windowed arrays: 16.8s (5.1x)
- Cached chunk arrays: 4.9s (17.3x)
Moves the code to a Parcels subpackage (rather than a separate package).
This was a testing-only convenience method, now handled internally by Xarray
instead of `cached_chunk`
@erikvansebille

Copy link
Copy Markdown
Member

Its not clear to me the performance with real-world simulations and how that scales (I only have access to my laptop, and found working/debugging on Lorenz quite frustrating with my high latency from travels). Let me know if there's anything I can do to help here @erikvansebille .

Thanks for this exciting PR, @VeckoTheGecko! I plan to do some real-world performance testing (also including/comparing #2846 and v3) later this week. Will report back here when I know more!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Backlog

Development

Successfully merging this pull request may close these issues.

2 participants

@VeckoTheGecko@erikvansebille
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Implement .to_chunk_cached_arrays() with ChunkCachedArray - #2860

Open
VeckoTheGecko wants to merge 13 commits into
Parcels-code:mainfrom
VeckoTheGecko:2854-dask-chunk-caching
Open

Implement .to_chunk_cached_arrays() with ChunkCachedArray#2860
VeckoTheGecko wants to merge 13 commits into
Parcels-code:mainfrom
VeckoTheGecko:2854-dask-chunk-caching

Conversation

@VeckoTheGecko

Copy link
Copy Markdown
Contributor

Description

Provides an opt-in optimization that wraps dask-backed field data in
chunk-level LRU caches.
vectorized .isel() calls hit an in-memory cache instead of recomputing
dask task graphs, giving large speedups for particle simulations.

This PR adds ChunkCachedArray according to the approach detailed in #2854 , providing an opt-in optimization wrapping dash-backed field data with (dask)chunk-level LRU caches on the array level. This PR:

  • Adds a subpackage parcels._chunk_cached_array with the implementation of the Array and the needed LRU cache. Note the cache is on the array level with configurable memory limits per array.
  • Exposes a function parcels._chunk_cached_array.wrap_dataset providing the main entrypoint for Parcels for wrapping Xarray dataset objects
  • Adds .to_chunk_cached_arrays to FieldSet and ModelData classes.
  • Adds equivalence testing (test_backends.py) on a NEMO dataset, asserting that the same results are received regardless of backend.

Still needed:

  • (separate PR) Upstream testing
  • (separate PR?) Backend equivalency testing on Unstructured grids?
  • Benchmarking with realworld simulations
  • (pre-merge) Remove temporary testing files (data-generation.py,benchmark_chunk_cache.py,benchmark_chunk_cache.png)

Future work:

  • max_cache_bytes on FieldSet.to_chunk_cached_arrays operates per variable (i.e., per Field). This requires the user to think about how many Fields are in their simulation in order to tune their performance. Down the line, we could tune this to allocate a total cache of (e.g.,) 50% of the system memory (which would allow for other memory overhead of coordinates (not stored in the cache), the Python process, or other user processes). This would require significant refactoring not only on how the chunk cache is handled, but also backends in general.

(cc @erikvansebille, @wyatt-fluidnumerics )

Performance

I have some some preliminary profiling with a 20Gb idealised dataset, and found the following performance profile which looks like a promising first step.

image

I've included data-generation.py and benchmark_chunk_cache.py here for testing in case it helps.

Its not clear to me the performance with real-world simulations and how that scales (I only have access to my laptop, and found working/debugging on Lorenz quite frustrating with my high latency from travels). Let me know if there's anything I can do to help here @erikvansebille .

Checklist

AI Disclosure

  • This PR contains AI-generated content.
    • I have tested any AI-generated content in my PR.
    • I take responsibility for any AI-generated content in my PR.
    • Describe how you used it (e.g., by pasting your prompt): I guided an LLM to explore the Xarray codebase, and to help work on the implementation of the Array class. I also used an LLM to generate and iterate on the profiling script, and write docstrings.

Provides an opt-in optimization that wraps dask-backed field data in
chunk-level LRU caches via the chunk_cached_array package. Repeated
vectorized .isel() calls hit an in-memory cache instead of recomputing
dask task graphs, giving large speedups for particle simulations.
Results on ds_2d_left_agrid.zarr with 10k particles:
- Plain dask: 85.0s (1x)
- Windowed arrays: 16.8s (5.1x)
- Cached chunk arrays: 4.9s (17.3x)
Moves the code to a Parcels subpackage (rather than a separate package).
This was a testing-only convenience method, now handled internally by Xarray
instead of `cached_chunk`
@erikvansebille

Copy link
Copy Markdown
Member

Its not clear to me the performance with real-world simulations and how that scales (I only have access to my laptop, and found working/debugging on Lorenz quite frustrating with my high latency from travels). Let me know if there's anything I can do to help here @erikvansebille .

Thanks for this exciting PR, @VeckoTheGecko! I plan to do some real-world performance testing (also including/comparing #2846 and v3) later this week. Will report back here when I know more!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Backlog

Development

Successfully merging this pull request may close these issues.

2 participants

@VeckoTheGecko@erikvansebille
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Implement .to_chunk_cached_arrays() with ChunkCachedArray - #2860

Open
VeckoTheGecko wants to merge 13 commits into
Parcels-code:mainfrom
VeckoTheGecko:2854-dask-chunk-caching
Open

Implement .to_chunk_cached_arrays() with ChunkCachedArray#2860
VeckoTheGecko wants to merge 13 commits into
Parcels-code:mainfrom
VeckoTheGecko:2854-dask-chunk-caching

Conversation

@VeckoTheGecko

Copy link
Copy Markdown
Contributor

Description

Provides an opt-in optimization that wraps dask-backed field data in
chunk-level LRU caches.
vectorized .isel() calls hit an in-memory cache instead of recomputing
dask task graphs, giving large speedups for particle simulations.

This PR adds ChunkCachedArray according to the approach detailed in #2854 , providing an opt-in optimization wrapping dash-backed field data with (dask)chunk-level LRU caches on the array level. This PR:

  • Adds a subpackage parcels._chunk_cached_array with the implementation of the Array and the needed LRU cache. Note the cache is on the array level with configurable memory limits per array.
  • Exposes a function parcels._chunk_cached_array.wrap_dataset providing the main entrypoint for Parcels for wrapping Xarray dataset objects
  • Adds .to_chunk_cached_arrays to FieldSet and ModelData classes.
  • Adds equivalence testing (test_backends.py) on a NEMO dataset, asserting that the same results are received regardless of backend.

Still needed:

  • (separate PR) Upstream testing
  • (separate PR?) Backend equivalency testing on Unstructured grids?
  • Benchmarking with realworld simulations
  • (pre-merge) Remove temporary testing files (data-generation.py,benchmark_chunk_cache.py,benchmark_chunk_cache.png)

Future work:

  • max_cache_bytes on FieldSet.to_chunk_cached_arrays operates per variable (i.e., per Field). This requires the user to think about how many Fields are in their simulation in order to tune their performance. Down the line, we could tune this to allocate a total cache of (e.g.,) 50% of the system memory (which would allow for other memory overhead of coordinates (not stored in the cache), the Python process, or other user processes). This would require significant refactoring not only on how the chunk cache is handled, but also backends in general.

(cc @erikvansebille, @wyatt-fluidnumerics )

Performance

I have some some preliminary profiling with a 20Gb idealised dataset, and found the following performance profile which looks like a promising first step.

image

I've included data-generation.py and benchmark_chunk_cache.py here for testing in case it helps.

Its not clear to me the performance with real-world simulations and how that scales (I only have access to my laptop, and found working/debugging on Lorenz quite frustrating with my high latency from travels). Let me know if there's anything I can do to help here @erikvansebille .

Checklist

AI Disclosure

  • This PR contains AI-generated content.
    • I have tested any AI-generated content in my PR.
    • I take responsibility for any AI-generated content in my PR.
    • Describe how you used it (e.g., by pasting your prompt): I guided an LLM to explore the Xarray codebase, and to help work on the implementation of the Array class. I also used an LLM to generate and iterate on the profiling script, and write docstrings.

Provides an opt-in optimization that wraps dask-backed field data in
chunk-level LRU caches via the chunk_cached_array package. Repeated
vectorized .isel() calls hit an in-memory cache instead of recomputing
dask task graphs, giving large speedups for particle simulations.
Results on ds_2d_left_agrid.zarr with 10k particles:
- Plain dask: 85.0s (1x)
- Windowed arrays: 16.8s (5.1x)
- Cached chunk arrays: 4.9s (17.3x)
Moves the code to a Parcels subpackage (rather than a separate package).
This was a testing-only convenience method, now handled internally by Xarray
instead of `cached_chunk`
@erikvansebille

Copy link
Copy Markdown
Member

Its not clear to me the performance with real-world simulations and how that scales (I only have access to my laptop, and found working/debugging on Lorenz quite frustrating with my high latency from travels). Let me know if there's anything I can do to help here @erikvansebille .

Thanks for this exciting PR, @VeckoTheGecko! I plan to do some real-world performance testing (also including/comparing #2846 and v3) later this week. Will report back here when I know more!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Backlog

Development

Successfully merging this pull request may close these issues.

2 participants

@VeckoTheGecko@erikvansebille
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Implement .to_chunk_cached_arrays() with ChunkCachedArray - #2860

Open
VeckoTheGecko wants to merge 13 commits into
Parcels-code:mainfrom
VeckoTheGecko:2854-dask-chunk-caching
Open

Implement .to_chunk_cached_arrays() with ChunkCachedArray#2860
VeckoTheGecko wants to merge 13 commits into
Parcels-code:mainfrom
VeckoTheGecko:2854-dask-chunk-caching

Conversation

@VeckoTheGecko

Copy link
Copy Markdown
Contributor

Description

Provides an opt-in optimization that wraps dask-backed field data in
chunk-level LRU caches.
vectorized .isel() calls hit an in-memory cache instead of recomputing
dask task graphs, giving large speedups for particle simulations.

This PR adds ChunkCachedArray according to the approach detailed in #2854 , providing an opt-in optimization wrapping dash-backed field data with (dask)chunk-level LRU caches on the array level. This PR:

  • Adds a subpackage parcels._chunk_cached_array with the implementation of the Array and the needed LRU cache. Note the cache is on the array level with configurable memory limits per array.
  • Exposes a function parcels._chunk_cached_array.wrap_dataset providing the main entrypoint for Parcels for wrapping Xarray dataset objects
  • Adds .to_chunk_cached_arrays to FieldSet and ModelData classes.
  • Adds equivalence testing (test_backends.py) on a NEMO dataset, asserting that the same results are received regardless of backend.

Still needed:

  • (separate PR) Upstream testing
  • (separate PR?) Backend equivalency testing on Unstructured grids?
  • Benchmarking with realworld simulations
  • (pre-merge) Remove temporary testing files (data-generation.py,benchmark_chunk_cache.py,benchmark_chunk_cache.png)

Future work:

  • max_cache_bytes on FieldSet.to_chunk_cached_arrays operates per variable (i.e., per Field). This requires the user to think about how many Fields are in their simulation in order to tune their performance. Down the line, we could tune this to allocate a total cache of (e.g.,) 50% of the system memory (which would allow for other memory overhead of coordinates (not stored in the cache), the Python process, or other user processes). This would require significant refactoring not only on how the chunk cache is handled, but also backends in general.

(cc @erikvansebille, @wyatt-fluidnumerics )

Performance

I have some some preliminary profiling with a 20Gb idealised dataset, and found the following performance profile which looks like a promising first step.

image

I've included data-generation.py and benchmark_chunk_cache.py here for testing in case it helps.

Its not clear to me the performance with real-world simulations and how that scales (I only have access to my laptop, and found working/debugging on Lorenz quite frustrating with my high latency from travels). Let me know if there's anything I can do to help here @erikvansebille .

Checklist

AI Disclosure

  • This PR contains AI-generated content.
    • I have tested any AI-generated content in my PR.
    • I take responsibility for any AI-generated content in my PR.
    • Describe how you used it (e.g., by pasting your prompt): I guided an LLM to explore the Xarray codebase, and to help work on the implementation of the Array class. I also used an LLM to generate and iterate on the profiling script, and write docstrings.

Provides an opt-in optimization that wraps dask-backed field data in
chunk-level LRU caches via the chunk_cached_array package. Repeated
vectorized .isel() calls hit an in-memory cache instead of recomputing
dask task graphs, giving large speedups for particle simulations.
Results on ds_2d_left_agrid.zarr with 10k particles:
- Plain dask: 85.0s (1x)
- Windowed arrays: 16.8s (5.1x)
- Cached chunk arrays: 4.9s (17.3x)
Moves the code to a Parcels subpackage (rather than a separate package).
This was a testing-only convenience method, now handled internally by Xarray
instead of `cached_chunk`
@erikvansebille

Copy link
Copy Markdown
Member

Its not clear to me the performance with real-world simulations and how that scales (I only have access to my laptop, and found working/debugging on Lorenz quite frustrating with my high latency from travels). Let me know if there's anything I can do to help here @erikvansebille .

Thanks for this exciting PR, @VeckoTheGecko! I plan to do some real-world performance testing (also including/comparing #2846 and v3) later this week. Will report back here when I know more!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Backlog

Development

Successfully merging this pull request may close these issues.

2 participants

@VeckoTheGecko@erikvansebille
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Implement .to_chunk_cached_arrays() with ChunkCachedArray - #2860

Open
VeckoTheGecko wants to merge 13 commits into
Parcels-code:mainfrom
VeckoTheGecko:2854-dask-chunk-caching
Open

Implement .to_chunk_cached_arrays() with ChunkCachedArray#2860
VeckoTheGecko wants to merge 13 commits into
Parcels-code:mainfrom
VeckoTheGecko:2854-dask-chunk-caching

Conversation

@VeckoTheGecko

Copy link
Copy Markdown
Contributor

Description

Provides an opt-in optimization that wraps dask-backed field data in
chunk-level LRU caches.
vectorized .isel() calls hit an in-memory cache instead of recomputing
dask task graphs, giving large speedups for particle simulations.

This PR adds ChunkCachedArray according to the approach detailed in #2854 , providing an opt-in optimization wrapping dash-backed field data with (dask)chunk-level LRU caches on the array level. This PR:

  • Adds a subpackage parcels._chunk_cached_array with the implementation of the Array and the needed LRU cache. Note the cache is on the array level with configurable memory limits per array.
  • Exposes a function parcels._chunk_cached_array.wrap_dataset providing the main entrypoint for Parcels for wrapping Xarray dataset objects
  • Adds .to_chunk_cached_arrays to FieldSet and ModelData classes.
  • Adds equivalence testing (test_backends.py) on a NEMO dataset, asserting that the same results are received regardless of backend.

Still needed:

  • (separate PR) Upstream testing
  • (separate PR?) Backend equivalency testing on Unstructured grids?
  • Benchmarking with realworld simulations
  • (pre-merge) Remove temporary testing files (data-generation.py,benchmark_chunk_cache.py,benchmark_chunk_cache.png)

Future work:

  • max_cache_bytes on FieldSet.to_chunk_cached_arrays operates per variable (i.e., per Field). This requires the user to think about how many Fields are in their simulation in order to tune their performance. Down the line, we could tune this to allocate a total cache of (e.g.,) 50% of the system memory (which would allow for other memory overhead of coordinates (not stored in the cache), the Python process, or other user processes). This would require significant refactoring not only on how the chunk cache is handled, but also backends in general.

(cc @erikvansebille, @wyatt-fluidnumerics )

Performance

I have some some preliminary profiling with a 20Gb idealised dataset, and found the following performance profile which looks like a promising first step.

image

I've included data-generation.py and benchmark_chunk_cache.py here for testing in case it helps.

Its not clear to me the performance with real-world simulations and how that scales (I only have access to my laptop, and found working/debugging on Lorenz quite frustrating with my high latency from travels). Let me know if there's anything I can do to help here @erikvansebille .

Checklist

AI Disclosure

  • This PR contains AI-generated content.
    • I have tested any AI-generated content in my PR.
    • I take responsibility for any AI-generated content in my PR.
    • Describe how you used it (e.g., by pasting your prompt): I guided an LLM to explore the Xarray codebase, and to help work on the implementation of the Array class. I also used an LLM to generate and iterate on the profiling script, and write docstrings.

Provides an opt-in optimization that wraps dask-backed field data in
chunk-level LRU caches via the chunk_cached_array package. Repeated
vectorized .isel() calls hit an in-memory cache instead of recomputing
dask task graphs, giving large speedups for particle simulations.
Results on ds_2d_left_agrid.zarr with 10k particles:
- Plain dask: 85.0s (1x)
- Windowed arrays: 16.8s (5.1x)
- Cached chunk arrays: 4.9s (17.3x)
Moves the code to a Parcels subpackage (rather than a separate package).
This was a testing-only convenience method, now handled internally by Xarray
instead of `cached_chunk`
@erikvansebille

Copy link
Copy Markdown
Member

Its not clear to me the performance with real-world simulations and how that scales (I only have access to my laptop, and found working/debugging on Lorenz quite frustrating with my high latency from travels). Let me know if there's anything I can do to help here @erikvansebille .

Thanks for this exciting PR, @VeckoTheGecko! I plan to do some real-world performance testing (also including/comparing #2846 and v3) later this week. Will report back here when I know more!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Backlog

Development

Successfully merging this pull request may close these issues.

2 participants

@VeckoTheGecko@erikvansebille
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Implement .to_chunk_cached_arrays() with ChunkCachedArray - #2860

Open
VeckoTheGecko wants to merge 13 commits into
Parcels-code:mainfrom
VeckoTheGecko:2854-dask-chunk-caching
Open

Implement .to_chunk_cached_arrays() with ChunkCachedArray#2860
VeckoTheGecko wants to merge 13 commits into
Parcels-code:mainfrom
VeckoTheGecko:2854-dask-chunk-caching

Conversation

@VeckoTheGecko

Copy link
Copy Markdown
Contributor

Description

Provides an opt-in optimization that wraps dask-backed field data in
chunk-level LRU caches.
vectorized .isel() calls hit an in-memory cache instead of recomputing
dask task graphs, giving large speedups for particle simulations.

This PR adds ChunkCachedArray according to the approach detailed in #2854 , providing an opt-in optimization wrapping dash-backed field data with (dask)chunk-level LRU caches on the array level. This PR:

  • Adds a subpackage parcels._chunk_cached_array with the implementation of the Array and the needed LRU cache. Note the cache is on the array level with configurable memory limits per array.
  • Exposes a function parcels._chunk_cached_array.wrap_dataset providing the main entrypoint for Parcels for wrapping Xarray dataset objects
  • Adds .to_chunk_cached_arrays to FieldSet and ModelData classes.
  • Adds equivalence testing (test_backends.py) on a NEMO dataset, asserting that the same results are received regardless of backend.

Still needed:

  • (separate PR) Upstream testing
  • (separate PR?) Backend equivalency testing on Unstructured grids?
  • Benchmarking with realworld simulations
  • (pre-merge) Remove temporary testing files (data-generation.py,benchmark_chunk_cache.py,benchmark_chunk_cache.png)

Future work:

  • max_cache_bytes on FieldSet.to_chunk_cached_arrays operates per variable (i.e., per Field). This requires the user to think about how many Fields are in their simulation in order to tune their performance. Down the line, we could tune this to allocate a total cache of (e.g.,) 50% of the system memory (which would allow for other memory overhead of coordinates (not stored in the cache), the Python process, or other user processes). This would require significant refactoring not only on how the chunk cache is handled, but also backends in general.

(cc @erikvansebille, @wyatt-fluidnumerics )

Performance

I have some some preliminary profiling with a 20Gb idealised dataset, and found the following performance profile which looks like a promising first step.

image

I've included data-generation.py and benchmark_chunk_cache.py here for testing in case it helps.

Its not clear to me the performance with real-world simulations and how that scales (I only have access to my laptop, and found working/debugging on Lorenz quite frustrating with my high latency from travels). Let me know if there's anything I can do to help here @erikvansebille .

Checklist

AI Disclosure

  • This PR contains AI-generated content.
    • I have tested any AI-generated content in my PR.
    • I take responsibility for any AI-generated content in my PR.
    • Describe how you used it (e.g., by pasting your prompt): I guided an LLM to explore the Xarray codebase, and to help work on the implementation of the Array class. I also used an LLM to generate and iterate on the profiling script, and write docstrings.

Provides an opt-in optimization that wraps dask-backed field data in
chunk-level LRU caches via the chunk_cached_array package. Repeated
vectorized .isel() calls hit an in-memory cache instead of recomputing
dask task graphs, giving large speedups for particle simulations.
Results on ds_2d_left_agrid.zarr with 10k particles:
- Plain dask: 85.0s (1x)
- Windowed arrays: 16.8s (5.1x)
- Cached chunk arrays: 4.9s (17.3x)
Moves the code to a Parcels subpackage (rather than a separate package).
This was a testing-only convenience method, now handled internally by Xarray
instead of `cached_chunk`
@erikvansebille

Copy link
Copy Markdown
Member

Its not clear to me the performance with real-world simulations and how that scales (I only have access to my laptop, and found working/debugging on Lorenz quite frustrating with my high latency from travels). Let me know if there's anything I can do to help here @erikvansebille .

Thanks for this exciting PR, @VeckoTheGecko! I plan to do some real-world performance testing (also including/comparing #2846 and v3) later this week. Will report back here when I know more!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Backlog

Development

Successfully merging this pull request may close these issues.

2 participants

@VeckoTheGecko@erikvansebille
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Implement .to_chunk_cached_arrays() with ChunkCachedArray - #2860

Open
VeckoTheGecko wants to merge 13 commits into
Parcels-code:mainfrom
VeckoTheGecko:2854-dask-chunk-caching
Open

Implement .to_chunk_cached_arrays() with ChunkCachedArray#2860
VeckoTheGecko wants to merge 13 commits into
Parcels-code:mainfrom
VeckoTheGecko:2854-dask-chunk-caching

Conversation

@VeckoTheGecko

Copy link
Copy Markdown
Contributor

Description

Provides an opt-in optimization that wraps dask-backed field data in
chunk-level LRU caches.
vectorized .isel() calls hit an in-memory cache instead of recomputing
dask task graphs, giving large speedups for particle simulations.

This PR adds ChunkCachedArray according to the approach detailed in #2854 , providing an opt-in optimization wrapping dash-backed field data with (dask)chunk-level LRU caches on the array level. This PR:

  • Adds a subpackage parcels._chunk_cached_array with the implementation of the Array and the needed LRU cache. Note the cache is on the array level with configurable memory limits per array.
  • Exposes a function parcels._chunk_cached_array.wrap_dataset providing the main entrypoint for Parcels for wrapping Xarray dataset objects
  • Adds .to_chunk_cached_arrays to FieldSet and ModelData classes.
  • Adds equivalence testing (test_backends.py) on a NEMO dataset, asserting that the same results are received regardless of backend.

Still needed:

  • (separate PR) Upstream testing
  • (separate PR?) Backend equivalency testing on Unstructured grids?
  • Benchmarking with realworld simulations
  • (pre-merge) Remove temporary testing files (data-generation.py,benchmark_chunk_cache.py,benchmark_chunk_cache.png)

Future work:

  • max_cache_bytes on FieldSet.to_chunk_cached_arrays operates per variable (i.e., per Field). This requires the user to think about how many Fields are in their simulation in order to tune their performance. Down the line, we could tune this to allocate a total cache of (e.g.,) 50% of the system memory (which would allow for other memory overhead of coordinates (not stored in the cache), the Python process, or other user processes). This would require significant refactoring not only on how the chunk cache is handled, but also backends in general.

(cc @erikvansebille, @wyatt-fluidnumerics )

Performance

I have some some preliminary profiling with a 20Gb idealised dataset, and found the following performance profile which looks like a promising first step.

image

I've included data-generation.py and benchmark_chunk_cache.py here for testing in case it helps.

Its not clear to me the performance with real-world simulations and how that scales (I only have access to my laptop, and found working/debugging on Lorenz quite frustrating with my high latency from travels). Let me know if there's anything I can do to help here @erikvansebille .

Checklist

AI Disclosure

  • This PR contains AI-generated content.
    • I have tested any AI-generated content in my PR.
    • I take responsibility for any AI-generated content in my PR.
    • Describe how you used it (e.g., by pasting your prompt): I guided an LLM to explore the Xarray codebase, and to help work on the implementation of the Array class. I also used an LLM to generate and iterate on the profiling script, and write docstrings.

Provides an opt-in optimization that wraps dask-backed field data in
chunk-level LRU caches via the chunk_cached_array package. Repeated
vectorized .isel() calls hit an in-memory cache instead of recomputing
dask task graphs, giving large speedups for particle simulations.
Results on ds_2d_left_agrid.zarr with 10k particles:
- Plain dask: 85.0s (1x)
- Windowed arrays: 16.8s (5.1x)
- Cached chunk arrays: 4.9s (17.3x)
Moves the code to a Parcels subpackage (rather than a separate package).
This was a testing-only convenience method, now handled internally by Xarray
instead of `cached_chunk`
@erikvansebille

Copy link
Copy Markdown
Member

Its not clear to me the performance with real-world simulations and how that scales (I only have access to my laptop, and found working/debugging on Lorenz quite frustrating with my high latency from travels). Let me know if there's anything I can do to help here @erikvansebille .

Thanks for this exciting PR, @VeckoTheGecko! I plan to do some real-world performance testing (also including/comparing #2846 and v3) later this week. Will report back here when I know more!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Backlog

Development

Successfully merging this pull request may close these issues.

2 participants

@VeckoTheGecko@erikvansebille
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Implement .to_chunk_cached_arrays() with ChunkCachedArray - #2860

Open
VeckoTheGecko wants to merge 13 commits into
Parcels-code:mainfrom
VeckoTheGecko:2854-dask-chunk-caching
Open

Implement .to_chunk_cached_arrays() with ChunkCachedArray#2860
VeckoTheGecko wants to merge 13 commits into
Parcels-code:mainfrom
VeckoTheGecko:2854-dask-chunk-caching

Conversation

@VeckoTheGecko

Copy link
Copy Markdown
Contributor

Description

Provides an opt-in optimization that wraps dask-backed field data in
chunk-level LRU caches.
vectorized .isel() calls hit an in-memory cache instead of recomputing
dask task graphs, giving large speedups for particle simulations.

This PR adds ChunkCachedArray according to the approach detailed in #2854 , providing an opt-in optimization wrapping dash-backed field data with (dask)chunk-level LRU caches on the array level. This PR:

  • Adds a subpackage parcels._chunk_cached_array with the implementation of the Array and the needed LRU cache. Note the cache is on the array level with configurable memory limits per array.
  • Exposes a function parcels._chunk_cached_array.wrap_dataset providing the main entrypoint for Parcels for wrapping Xarray dataset objects
  • Adds .to_chunk_cached_arrays to FieldSet and ModelData classes.
  • Adds equivalence testing (test_backends.py) on a NEMO dataset, asserting that the same results are received regardless of backend.

Still needed:

  • (separate PR) Upstream testing
  • (separate PR?) Backend equivalency testing on Unstructured grids?
  • Benchmarking with realworld simulations
  • (pre-merge) Remove temporary testing files (data-generation.py,benchmark_chunk_cache.py,benchmark_chunk_cache.png)

Future work:

  • max_cache_bytes on FieldSet.to_chunk_cached_arrays operates per variable (i.e., per Field). This requires the user to think about how many Fields are in their simulation in order to tune their performance. Down the line, we could tune this to allocate a total cache of (e.g.,) 50% of the system memory (which would allow for other memory overhead of coordinates (not stored in the cache), the Python process, or other user processes). This would require significant refactoring not only on how the chunk cache is handled, but also backends in general.

(cc @erikvansebille, @wyatt-fluidnumerics )

Performance

I have some some preliminary profiling with a 20Gb idealised dataset, and found the following performance profile which looks like a promising first step.

image

I've included data-generation.py and benchmark_chunk_cache.py here for testing in case it helps.

Its not clear to me the performance with real-world simulations and how that scales (I only have access to my laptop, and found working/debugging on Lorenz quite frustrating with my high latency from travels). Let me know if there's anything I can do to help here @erikvansebille .

Checklist

AI Disclosure

  • This PR contains AI-generated content.
    • I have tested any AI-generated content in my PR.
    • I take responsibility for any AI-generated content in my PR.
    • Describe how you used it (e.g., by pasting your prompt): I guided an LLM to explore the Xarray codebase, and to help work on the implementation of the Array class. I also used an LLM to generate and iterate on the profiling script, and write docstrings.

Provides an opt-in optimization that wraps dask-backed field data in
chunk-level LRU caches via the chunk_cached_array package. Repeated
vectorized .isel() calls hit an in-memory cache instead of recomputing
dask task graphs, giving large speedups for particle simulations.
Results on ds_2d_left_agrid.zarr with 10k particles:
- Plain dask: 85.0s (1x)
- Windowed arrays: 16.8s (5.1x)
- Cached chunk arrays: 4.9s (17.3x)
Moves the code to a Parcels subpackage (rather than a separate package).
This was a testing-only convenience method, now handled internally by Xarray
instead of `cached_chunk`
@erikvansebille

Copy link
Copy Markdown
Member

Its not clear to me the performance with real-world simulations and how that scales (I only have access to my laptop, and found working/debugging on Lorenz quite frustrating with my high latency from travels). Let me know if there's anything I can do to help here @erikvansebille .

Thanks for this exciting PR, @VeckoTheGecko! I plan to do some real-world performance testing (also including/comparing #2846 and v3) later this week. Will report back here when I know more!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Backlog

Development

Successfully merging this pull request may close these issues.

2 participants

@VeckoTheGecko@erikvansebille