Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in) #4

Description

@sodre

Follow-up to #3 (Plan F).

Plan F's enroot load docker:// path materializes the merged image into a single ZFS template per image (cached by image config digest). That covers most cases but loses several properties that the per-layer-clone-chain approach (mirroring Docker's own zfs storage driver) would give us. Plan G is the opt-in mode that adds the per-layer path alongside Plan F's single-merge path, gated on a config flag like ENROOT_ZFS_LAYER_CHAIN=y.

Why it's worth doing

For HPC / CI hosts that pull many images sharing common bases, the current single-merge design wastes disk and CPU. Per-layer chains would buy back:

  1. Cross-image layer dedup at the dataset level. Two images sharing a Debian base store the base bytes once instead of twice. Block-level dedup=on recovers this in Plan F's design but at ~5–6 GB RAM per TB indexed; per-layer datasets dedup for free.
  2. Incremental re-pull cost. When alpine:3.21 replaces alpine:3.20, only the changed top layers are re-extracted; lower-layer datasets are reused. Plan F re-merges the whole stack.
  3. Layer-granular cache invalidation. A poisoned layer can be zfs destroyd in isolation; Plan F throws out the whole template.
  4. Native ZFS introspection.zfs list -t all shows the layer chain; zfs send per-layer becomes a sensible cross-host replication primitive.
  5. Quota accounting matches intuition.quota=200G on <store>/.templates reflects shared layers once, not multiplied by the number of images that use them.
  6. Aligns with Docker's zfs storage driver shape, which ops people already know.

What it costs

  • Whiteout/opaque-dir merging in shell. Overlayfs whiteouts are character-device files (mknod c 0 0); opaque-dir markers are trusted.overlay.opaque=y xattrs. Without the kernel's overlay engine doing the merge, we need to apply these manually during each clone-extract step. Real edge-case surface.
  • ~5–15× more dataset objects per image. zfs list clutter; more bookkeeping.
  • Per-layer atomic locks (multiple .tmp datasets racing).
  • zfs promote (or chain-preservation alternative) to flatten the leaf into a standalone template — depends on user delegations.
  • More complex cache invalidation logic (decide what to destroy when a leaf is reaped vs. when a shared lower layer is reaped).

Sketch of design

zfs::docker_install_from_layers (in src/storage_zfs.sh) gains a check: if ENROOT_ZFS_LAYER_CHAIN=y, dispatch to zfs::docker_install_chain instead. The new function:

  1. For each layer in stack order, hash the layer tarball (already in ${ENROOT_CACHE_PATH}/<digest> from _prepare_layers) → cache key.
  2. If <store>/.layers/<digest>@done exists, reuse; else zfs clone parent@done <new> (or zfs create for the base), apply layer's whiteouts + extracted contents, zfs snapshot @done.
  3. Final leaf is the merged image; zfs promote it into <store>/.templates/<image-config-sha> to flatten the chain into a standalone template.
  4. Existing zfs::clone_container then clones the template for the user.

Whiteout application for step 2 needs a helper like zfs::apply_layer_whiteouts <layer_dir> <target>:

  • For each *.wh.foo (AUFS) or 0:0-char-device named foo (overlayfs) in layer_dir: rm -rf "${target}/foo".
  • For each .wh..wh..opq or trusted.overlay.opaque=y dir: clear children of corresponding dir in target.
  • Then cp -a (or tar | tar) the rest of layer_dir over target.

enroot-aufs2ovlfs already converts to overlayfs whiteouts in-place, so the helper only needs the overlayfs forms. Worth confirming the exact char-device format it produces.

Coexistence with Plan F

  • Default behavior unchanged: Plan F's single-merge path stays the default.
  • ENROOT_ZFS_LAYER_CHAIN=y opts into per-layer.
  • Both paths populate <store>/.templates/<sha> — same shape, same zfs::clone_container for the user. Only the fill mechanism differs.
  • A site can switch on/off without migration; existing single-merge templates remain valid.

Open questions

  • Where do layer datasets live? <store>/.layers/<sha> (parallel to .templates) keeps the layer cache separate from per-image templates. Avoids confusion when an admin scans .templates.
  • zfs promote permissions: needs promote in the user's zfs allow. Document in the admin recipe.
  • Should we GC unused layer datasets when no template references them? Plan B's sweep mechanism could be extended to .layers/ with the same warm/cold logic.
  • How does Plan B's warm/cold eviction interact? Templates and layers should probably share the same lifecycle policy, but layers' lifecycle is "no template references" rather than "no clones."
  • xattr propagation under cp -a vs tar: tar --xattrs --xattrs-include='*' --selinux --acls is the safer pipe; enroot-aufs2ovlfs likely emits these correctly already.

Acceptance criteria

  • Two distinct images sharing a base layer (e.g. python:3-slim and node:20-slim, both Debian-based) store the shared layer once on disk under <store>/.layers/<digest>.
  • Re-pulling an image after a top-layer-only update reuses the cached lower-layer datasets (verify with timing + zfs list snapshot before/after).
  • Whiteouts and opaque dirs from real Docker images merge correctly (verified against python:3-slim, nginx, cuda base images).
  • Concurrent enroot load of the same image is race-safe (same .tmp lock pattern as Plan F).
  • ENROOT_ZFS_LAYER_CHAIN= (unset) leaves Plan F's behavior unchanged.

Out of scope

  • Replacing Plan F's single-merge path. Plan G is purely additive.
  • Cross-host layer replication via zfs send. Natural follow-up but tracked separately.
  • Migration tooling between merged-template and per-layer-chain caches.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
       blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in) #4

      Description

      @sodre

      Follow-up to #3 (Plan F).

      Plan F's enroot load docker:// path materializes the merged image into a single ZFS template per image (cached by image config digest). That covers most cases but loses several properties that the per-layer-clone-chain approach (mirroring Docker's own zfs storage driver) would give us. Plan G is the opt-in mode that adds the per-layer path alongside Plan F's single-merge path, gated on a config flag like ENROOT_ZFS_LAYER_CHAIN=y.

      Why it's worth doing

      For HPC / CI hosts that pull many images sharing common bases, the current single-merge design wastes disk and CPU. Per-layer chains would buy back:

      1. Cross-image layer dedup at the dataset level. Two images sharing a Debian base store the base bytes once instead of twice. Block-level dedup=on recovers this in Plan F's design but at ~5–6 GB RAM per TB indexed; per-layer datasets dedup for free.
      2. Incremental re-pull cost. When alpine:3.21 replaces alpine:3.20, only the changed top layers are re-extracted; lower-layer datasets are reused. Plan F re-merges the whole stack.
      3. Layer-granular cache invalidation. A poisoned layer can be zfs destroyd in isolation; Plan F throws out the whole template.
      4. Native ZFS introspection.zfs list -t all shows the layer chain; zfs send per-layer becomes a sensible cross-host replication primitive.
      5. Quota accounting matches intuition.quota=200G on <store>/.templates reflects shared layers once, not multiplied by the number of images that use them.
      6. Aligns with Docker's zfs storage driver shape, which ops people already know.

      What it costs

      • Whiteout/opaque-dir merging in shell. Overlayfs whiteouts are character-device files (mknod c 0 0); opaque-dir markers are trusted.overlay.opaque=y xattrs. Without the kernel's overlay engine doing the merge, we need to apply these manually during each clone-extract step. Real edge-case surface.
      • ~5–15× more dataset objects per image. zfs list clutter; more bookkeeping.
      • Per-layer atomic locks (multiple .tmp datasets racing).
      • zfs promote (or chain-preservation alternative) to flatten the leaf into a standalone template — depends on user delegations.
      • More complex cache invalidation logic (decide what to destroy when a leaf is reaped vs. when a shared lower layer is reaped).

      Sketch of design

      zfs::docker_install_from_layers (in src/storage_zfs.sh) gains a check: if ENROOT_ZFS_LAYER_CHAIN=y, dispatch to zfs::docker_install_chain instead. The new function:

      1. For each layer in stack order, hash the layer tarball (already in ${ENROOT_CACHE_PATH}/<digest> from _prepare_layers) → cache key.
      2. If <store>/.layers/<digest>@done exists, reuse; else zfs clone parent@done <new> (or zfs create for the base), apply layer's whiteouts + extracted contents, zfs snapshot @done.
      3. Final leaf is the merged image; zfs promote it into <store>/.templates/<image-config-sha> to flatten the chain into a standalone template.
      4. Existing zfs::clone_container then clones the template for the user.

      Whiteout application for step 2 needs a helper like zfs::apply_layer_whiteouts <layer_dir> <target>:

      • For each *.wh.foo (AUFS) or 0:0-char-device named foo (overlayfs) in layer_dir: rm -rf "${target}/foo".
      • For each .wh..wh..opq or trusted.overlay.opaque=y dir: clear children of corresponding dir in target.
      • Then cp -a (or tar | tar) the rest of layer_dir over target.

      enroot-aufs2ovlfs already converts to overlayfs whiteouts in-place, so the helper only needs the overlayfs forms. Worth confirming the exact char-device format it produces.

      Coexistence with Plan F

      • Default behavior unchanged: Plan F's single-merge path stays the default.
      • ENROOT_ZFS_LAYER_CHAIN=y opts into per-layer.
      • Both paths populate <store>/.templates/<sha> — same shape, same zfs::clone_container for the user. Only the fill mechanism differs.
      • A site can switch on/off without migration; existing single-merge templates remain valid.

      Open questions

      • Where do layer datasets live? <store>/.layers/<sha> (parallel to .templates) keeps the layer cache separate from per-image templates. Avoids confusion when an admin scans .templates.
      • zfs promote permissions: needs promote in the user's zfs allow. Document in the admin recipe.
      • Should we GC unused layer datasets when no template references them? Plan B's sweep mechanism could be extended to .layers/ with the same warm/cold logic.
      • How does Plan B's warm/cold eviction interact? Templates and layers should probably share the same lifecycle policy, but layers' lifecycle is "no template references" rather than "no clones."
      • xattr propagation under cp -a vs tar: tar --xattrs --xattrs-include='*' --selinux --acls is the safer pipe; enroot-aufs2ovlfs likely emits these correctly already.

      Acceptance criteria

      • Two distinct images sharing a base layer (e.g. python:3-slim and node:20-slim, both Debian-based) store the shared layer once on disk under <store>/.layers/<digest>.
      • Re-pulling an image after a top-layer-only update reuses the cached lower-layer datasets (verify with timing + zfs list snapshot before/after).
      • Whiteouts and opaque dirs from real Docker images merge correctly (verified against python:3-slim, nginx, cuda base images).
      • Concurrent enroot load of the same image is race-safe (same .tmp lock pattern as Plan F).
      • ENROOT_ZFS_LAYER_CHAIN= (unset) leaves Plan F's behavior unchanged.

      Out of scope

      • Replacing Plan F's single-merge path. Plan G is purely additive.
      • Cross-host layer replication via zfs send. Natural follow-up but tracked separately.
      • Migration tooling between merged-template and per-layer-chain caches.

      Activity

      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        No labels
        No labels

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in) #4

          Description

          @sodre

          Follow-up to #3 (Plan F).

          Plan F's enroot load docker:// path materializes the merged image into a single ZFS template per image (cached by image config digest). That covers most cases but loses several properties that the per-layer-clone-chain approach (mirroring Docker's own zfs storage driver) would give us. Plan G is the opt-in mode that adds the per-layer path alongside Plan F's single-merge path, gated on a config flag like ENROOT_ZFS_LAYER_CHAIN=y.

          Why it's worth doing

          For HPC / CI hosts that pull many images sharing common bases, the current single-merge design wastes disk and CPU. Per-layer chains would buy back:

          1. Cross-image layer dedup at the dataset level. Two images sharing a Debian base store the base bytes once instead of twice. Block-level dedup=on recovers this in Plan F's design but at ~5–6 GB RAM per TB indexed; per-layer datasets dedup for free.
          2. Incremental re-pull cost. When alpine:3.21 replaces alpine:3.20, only the changed top layers are re-extracted; lower-layer datasets are reused. Plan F re-merges the whole stack.
          3. Layer-granular cache invalidation. A poisoned layer can be zfs destroyd in isolation; Plan F throws out the whole template.
          4. Native ZFS introspection.zfs list -t all shows the layer chain; zfs send per-layer becomes a sensible cross-host replication primitive.
          5. Quota accounting matches intuition.quota=200G on <store>/.templates reflects shared layers once, not multiplied by the number of images that use them.
          6. Aligns with Docker's zfs storage driver shape, which ops people already know.

          What it costs

          • Whiteout/opaque-dir merging in shell. Overlayfs whiteouts are character-device files (mknod c 0 0); opaque-dir markers are trusted.overlay.opaque=y xattrs. Without the kernel's overlay engine doing the merge, we need to apply these manually during each clone-extract step. Real edge-case surface.
          • ~5–15× more dataset objects per image. zfs list clutter; more bookkeeping.
          • Per-layer atomic locks (multiple .tmp datasets racing).
          • zfs promote (or chain-preservation alternative) to flatten the leaf into a standalone template — depends on user delegations.
          • More complex cache invalidation logic (decide what to destroy when a leaf is reaped vs. when a shared lower layer is reaped).

          Sketch of design

          zfs::docker_install_from_layers (in src/storage_zfs.sh) gains a check: if ENROOT_ZFS_LAYER_CHAIN=y, dispatch to zfs::docker_install_chain instead. The new function:

          1. For each layer in stack order, hash the layer tarball (already in ${ENROOT_CACHE_PATH}/<digest> from _prepare_layers) → cache key.
          2. If <store>/.layers/<digest>@done exists, reuse; else zfs clone parent@done <new> (or zfs create for the base), apply layer's whiteouts + extracted contents, zfs snapshot @done.
          3. Final leaf is the merged image; zfs promote it into <store>/.templates/<image-config-sha> to flatten the chain into a standalone template.
          4. Existing zfs::clone_container then clones the template for the user.

          Whiteout application for step 2 needs a helper like zfs::apply_layer_whiteouts <layer_dir> <target>:

          • For each *.wh.foo (AUFS) or 0:0-char-device named foo (overlayfs) in layer_dir: rm -rf "${target}/foo".
          • For each .wh..wh..opq or trusted.overlay.opaque=y dir: clear children of corresponding dir in target.
          • Then cp -a (or tar | tar) the rest of layer_dir over target.

          enroot-aufs2ovlfs already converts to overlayfs whiteouts in-place, so the helper only needs the overlayfs forms. Worth confirming the exact char-device format it produces.

          Coexistence with Plan F

          • Default behavior unchanged: Plan F's single-merge path stays the default.
          • ENROOT_ZFS_LAYER_CHAIN=y opts into per-layer.
          • Both paths populate <store>/.templates/<sha> — same shape, same zfs::clone_container for the user. Only the fill mechanism differs.
          • A site can switch on/off without migration; existing single-merge templates remain valid.

          Open questions

          • Where do layer datasets live? <store>/.layers/<sha> (parallel to .templates) keeps the layer cache separate from per-image templates. Avoids confusion when an admin scans .templates.
          • zfs promote permissions: needs promote in the user's zfs allow. Document in the admin recipe.
          • Should we GC unused layer datasets when no template references them? Plan B's sweep mechanism could be extended to .layers/ with the same warm/cold logic.
          • How does Plan B's warm/cold eviction interact? Templates and layers should probably share the same lifecycle policy, but layers' lifecycle is "no template references" rather than "no clones."
          • xattr propagation under cp -a vs tar: tar --xattrs --xattrs-include='*' --selinux --acls is the safer pipe; enroot-aufs2ovlfs likely emits these correctly already.

          Acceptance criteria

          • Two distinct images sharing a base layer (e.g. python:3-slim and node:20-slim, both Debian-based) store the shared layer once on disk under <store>/.layers/<digest>.
          • Re-pulling an image after a top-layer-only update reuses the cached lower-layer datasets (verify with timing + zfs list snapshot before/after).
          • Whiteouts and opaque dirs from real Docker images merge correctly (verified against python:3-slim, nginx, cuda base images).
          • Concurrent enroot load of the same image is race-safe (same .tmp lock pattern as Plan F).
          • ENROOT_ZFS_LAYER_CHAIN= (unset) leaves Plan F's behavior unchanged.

          Out of scope

          • Replacing Plan F's single-merge path. Plan G is purely additive.
          • Cross-host layer replication via zfs send. Natural follow-up but tracked separately.
          • Migration tooling between merged-template and per-layer-chain caches.

          Activity

          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            No labels
            No labels

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in) #4

              Description

              @sodre

              Follow-up to #3 (Plan F).

              Plan F's enroot load docker:// path materializes the merged image into a single ZFS template per image (cached by image config digest). That covers most cases but loses several properties that the per-layer-clone-chain approach (mirroring Docker's own zfs storage driver) would give us. Plan G is the opt-in mode that adds the per-layer path alongside Plan F's single-merge path, gated on a config flag like ENROOT_ZFS_LAYER_CHAIN=y.

              Why it's worth doing

              For HPC / CI hosts that pull many images sharing common bases, the current single-merge design wastes disk and CPU. Per-layer chains would buy back:

              1. Cross-image layer dedup at the dataset level. Two images sharing a Debian base store the base bytes once instead of twice. Block-level dedup=on recovers this in Plan F's design but at ~5–6 GB RAM per TB indexed; per-layer datasets dedup for free.
              2. Incremental re-pull cost. When alpine:3.21 replaces alpine:3.20, only the changed top layers are re-extracted; lower-layer datasets are reused. Plan F re-merges the whole stack.
              3. Layer-granular cache invalidation. A poisoned layer can be zfs destroyd in isolation; Plan F throws out the whole template.
              4. Native ZFS introspection.zfs list -t all shows the layer chain; zfs send per-layer becomes a sensible cross-host replication primitive.
              5. Quota accounting matches intuition.quota=200G on <store>/.templates reflects shared layers once, not multiplied by the number of images that use them.
              6. Aligns with Docker's zfs storage driver shape, which ops people already know.

              What it costs

              • Whiteout/opaque-dir merging in shell. Overlayfs whiteouts are character-device files (mknod c 0 0); opaque-dir markers are trusted.overlay.opaque=y xattrs. Without the kernel's overlay engine doing the merge, we need to apply these manually during each clone-extract step. Real edge-case surface.
              • ~5–15× more dataset objects per image. zfs list clutter; more bookkeeping.
              • Per-layer atomic locks (multiple .tmp datasets racing).
              • zfs promote (or chain-preservation alternative) to flatten the leaf into a standalone template — depends on user delegations.
              • More complex cache invalidation logic (decide what to destroy when a leaf is reaped vs. when a shared lower layer is reaped).

              Sketch of design

              zfs::docker_install_from_layers (in src/storage_zfs.sh) gains a check: if ENROOT_ZFS_LAYER_CHAIN=y, dispatch to zfs::docker_install_chain instead. The new function:

              1. For each layer in stack order, hash the layer tarball (already in ${ENROOT_CACHE_PATH}/<digest> from _prepare_layers) → cache key.
              2. If <store>/.layers/<digest>@done exists, reuse; else zfs clone parent@done <new> (or zfs create for the base), apply layer's whiteouts + extracted contents, zfs snapshot @done.
              3. Final leaf is the merged image; zfs promote it into <store>/.templates/<image-config-sha> to flatten the chain into a standalone template.
              4. Existing zfs::clone_container then clones the template for the user.

              Whiteout application for step 2 needs a helper like zfs::apply_layer_whiteouts <layer_dir> <target>:

              • For each *.wh.foo (AUFS) or 0:0-char-device named foo (overlayfs) in layer_dir: rm -rf "${target}/foo".
              • For each .wh..wh..opq or trusted.overlay.opaque=y dir: clear children of corresponding dir in target.
              • Then cp -a (or tar | tar) the rest of layer_dir over target.

              enroot-aufs2ovlfs already converts to overlayfs whiteouts in-place, so the helper only needs the overlayfs forms. Worth confirming the exact char-device format it produces.

              Coexistence with Plan F

              • Default behavior unchanged: Plan F's single-merge path stays the default.
              • ENROOT_ZFS_LAYER_CHAIN=y opts into per-layer.
              • Both paths populate <store>/.templates/<sha> — same shape, same zfs::clone_container for the user. Only the fill mechanism differs.
              • A site can switch on/off without migration; existing single-merge templates remain valid.

              Open questions

              • Where do layer datasets live? <store>/.layers/<sha> (parallel to .templates) keeps the layer cache separate from per-image templates. Avoids confusion when an admin scans .templates.
              • zfs promote permissions: needs promote in the user's zfs allow. Document in the admin recipe.
              • Should we GC unused layer datasets when no template references them? Plan B's sweep mechanism could be extended to .layers/ with the same warm/cold logic.
              • How does Plan B's warm/cold eviction interact? Templates and layers should probably share the same lifecycle policy, but layers' lifecycle is "no template references" rather than "no clones."
              • xattr propagation under cp -a vs tar: tar --xattrs --xattrs-include='*' --selinux --acls is the safer pipe; enroot-aufs2ovlfs likely emits these correctly already.

              Acceptance criteria

              • Two distinct images sharing a base layer (e.g. python:3-slim and node:20-slim, both Debian-based) store the shared layer once on disk under <store>/.layers/<digest>.
              • Re-pulling an image after a top-layer-only update reuses the cached lower-layer datasets (verify with timing + zfs list snapshot before/after).
              • Whiteouts and opaque dirs from real Docker images merge correctly (verified against python:3-slim, nginx, cuda base images).
              • Concurrent enroot load of the same image is race-safe (same .tmp lock pattern as Plan F).
              • ENROOT_ZFS_LAYER_CHAIN= (unset) leaves Plan F's behavior unchanged.

              Out of scope

              • Replacing Plan F's single-merge path. Plan G is purely additive.
              • Cross-host layer replication via zfs send. Natural follow-up but tracked separately.
              • Migration tooling between merged-template and per-layer-chain caches.

              Activity

              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                No labels
                No labels

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in) #4

                  Description

                  @sodre

                  Follow-up to #3 (Plan F).

                  Plan F's enroot load docker:// path materializes the merged image into a single ZFS template per image (cached by image config digest). That covers most cases but loses several properties that the per-layer-clone-chain approach (mirroring Docker's own zfs storage driver) would give us. Plan G is the opt-in mode that adds the per-layer path alongside Plan F's single-merge path, gated on a config flag like ENROOT_ZFS_LAYER_CHAIN=y.

                  Why it's worth doing

                  For HPC / CI hosts that pull many images sharing common bases, the current single-merge design wastes disk and CPU. Per-layer chains would buy back:

                  1. Cross-image layer dedup at the dataset level. Two images sharing a Debian base store the base bytes once instead of twice. Block-level dedup=on recovers this in Plan F's design but at ~5–6 GB RAM per TB indexed; per-layer datasets dedup for free.
                  2. Incremental re-pull cost. When alpine:3.21 replaces alpine:3.20, only the changed top layers are re-extracted; lower-layer datasets are reused. Plan F re-merges the whole stack.
                  3. Layer-granular cache invalidation. A poisoned layer can be zfs destroyd in isolation; Plan F throws out the whole template.
                  4. Native ZFS introspection.zfs list -t all shows the layer chain; zfs send per-layer becomes a sensible cross-host replication primitive.
                  5. Quota accounting matches intuition.quota=200G on <store>/.templates reflects shared layers once, not multiplied by the number of images that use them.
                  6. Aligns with Docker's zfs storage driver shape, which ops people already know.

                  What it costs

                  • Whiteout/opaque-dir merging in shell. Overlayfs whiteouts are character-device files (mknod c 0 0); opaque-dir markers are trusted.overlay.opaque=y xattrs. Without the kernel's overlay engine doing the merge, we need to apply these manually during each clone-extract step. Real edge-case surface.
                  • ~5–15× more dataset objects per image. zfs list clutter; more bookkeeping.
                  • Per-layer atomic locks (multiple .tmp datasets racing).
                  • zfs promote (or chain-preservation alternative) to flatten the leaf into a standalone template — depends on user delegations.
                  • More complex cache invalidation logic (decide what to destroy when a leaf is reaped vs. when a shared lower layer is reaped).

                  Sketch of design

                  zfs::docker_install_from_layers (in src/storage_zfs.sh) gains a check: if ENROOT_ZFS_LAYER_CHAIN=y, dispatch to zfs::docker_install_chain instead. The new function:

                  1. For each layer in stack order, hash the layer tarball (already in ${ENROOT_CACHE_PATH}/<digest> from _prepare_layers) → cache key.
                  2. If <store>/.layers/<digest>@done exists, reuse; else zfs clone parent@done <new> (or zfs create for the base), apply layer's whiteouts + extracted contents, zfs snapshot @done.
                  3. Final leaf is the merged image; zfs promote it into <store>/.templates/<image-config-sha> to flatten the chain into a standalone template.
                  4. Existing zfs::clone_container then clones the template for the user.

                  Whiteout application for step 2 needs a helper like zfs::apply_layer_whiteouts <layer_dir> <target>:

                  • For each *.wh.foo (AUFS) or 0:0-char-device named foo (overlayfs) in layer_dir: rm -rf "${target}/foo".
                  • For each .wh..wh..opq or trusted.overlay.opaque=y dir: clear children of corresponding dir in target.
                  • Then cp -a (or tar | tar) the rest of layer_dir over target.

                  enroot-aufs2ovlfs already converts to overlayfs whiteouts in-place, so the helper only needs the overlayfs forms. Worth confirming the exact char-device format it produces.

                  Coexistence with Plan F

                  • Default behavior unchanged: Plan F's single-merge path stays the default.
                  • ENROOT_ZFS_LAYER_CHAIN=y opts into per-layer.
                  • Both paths populate <store>/.templates/<sha> — same shape, same zfs::clone_container for the user. Only the fill mechanism differs.
                  • A site can switch on/off without migration; existing single-merge templates remain valid.

                  Open questions

                  • Where do layer datasets live? <store>/.layers/<sha> (parallel to .templates) keeps the layer cache separate from per-image templates. Avoids confusion when an admin scans .templates.
                  • zfs promote permissions: needs promote in the user's zfs allow. Document in the admin recipe.
                  • Should we GC unused layer datasets when no template references them? Plan B's sweep mechanism could be extended to .layers/ with the same warm/cold logic.
                  • How does Plan B's warm/cold eviction interact? Templates and layers should probably share the same lifecycle policy, but layers' lifecycle is "no template references" rather than "no clones."
                  • xattr propagation under cp -a vs tar: tar --xattrs --xattrs-include='*' --selinux --acls is the safer pipe; enroot-aufs2ovlfs likely emits these correctly already.

                  Acceptance criteria

                  • Two distinct images sharing a base layer (e.g. python:3-slim and node:20-slim, both Debian-based) store the shared layer once on disk under <store>/.layers/<digest>.
                  • Re-pulling an image after a top-layer-only update reuses the cached lower-layer datasets (verify with timing + zfs list snapshot before/after).
                  • Whiteouts and opaque dirs from real Docker images merge correctly (verified against python:3-slim, nginx, cuda base images).
                  • Concurrent enroot load of the same image is race-safe (same .tmp lock pattern as Plan F).
                  • ENROOT_ZFS_LAYER_CHAIN= (unset) leaves Plan F's behavior unchanged.

                  Out of scope

                  • Replacing Plan F's single-merge path. Plan G is purely additive.
                  • Cross-host layer replication via zfs send. Natural follow-up but tracked separately.
                  • Migration tooling between merged-template and per-layer-chain caches.

                  Activity

                  Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    No labels
                    No labels

                    Type

                    No type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in) #4

                      Description

                      @sodre

                      Follow-up to #3 (Plan F).

                      Plan F's enroot load docker:// path materializes the merged image into a single ZFS template per image (cached by image config digest). That covers most cases but loses several properties that the per-layer-clone-chain approach (mirroring Docker's own zfs storage driver) would give us. Plan G is the opt-in mode that adds the per-layer path alongside Plan F's single-merge path, gated on a config flag like ENROOT_ZFS_LAYER_CHAIN=y.

                      Why it's worth doing

                      For HPC / CI hosts that pull many images sharing common bases, the current single-merge design wastes disk and CPU. Per-layer chains would buy back:

                      1. Cross-image layer dedup at the dataset level. Two images sharing a Debian base store the base bytes once instead of twice. Block-level dedup=on recovers this in Plan F's design but at ~5–6 GB RAM per TB indexed; per-layer datasets dedup for free.
                      2. Incremental re-pull cost. When alpine:3.21 replaces alpine:3.20, only the changed top layers are re-extracted; lower-layer datasets are reused. Plan F re-merges the whole stack.
                      3. Layer-granular cache invalidation. A poisoned layer can be zfs destroyd in isolation; Plan F throws out the whole template.
                      4. Native ZFS introspection.zfs list -t all shows the layer chain; zfs send per-layer becomes a sensible cross-host replication primitive.
                      5. Quota accounting matches intuition.quota=200G on <store>/.templates reflects shared layers once, not multiplied by the number of images that use them.
                      6. Aligns with Docker's zfs storage driver shape, which ops people already know.

                      What it costs

                      • Whiteout/opaque-dir merging in shell. Overlayfs whiteouts are character-device files (mknod c 0 0); opaque-dir markers are trusted.overlay.opaque=y xattrs. Without the kernel's overlay engine doing the merge, we need to apply these manually during each clone-extract step. Real edge-case surface.
                      • ~5–15× more dataset objects per image. zfs list clutter; more bookkeeping.
                      • Per-layer atomic locks (multiple .tmp datasets racing).
                      • zfs promote (or chain-preservation alternative) to flatten the leaf into a standalone template — depends on user delegations.
                      • More complex cache invalidation logic (decide what to destroy when a leaf is reaped vs. when a shared lower layer is reaped).

                      Sketch of design

                      zfs::docker_install_from_layers (in src/storage_zfs.sh) gains a check: if ENROOT_ZFS_LAYER_CHAIN=y, dispatch to zfs::docker_install_chain instead. The new function:

                      1. For each layer in stack order, hash the layer tarball (already in ${ENROOT_CACHE_PATH}/<digest> from _prepare_layers) → cache key.
                      2. If <store>/.layers/<digest>@done exists, reuse; else zfs clone parent@done <new> (or zfs create for the base), apply layer's whiteouts + extracted contents, zfs snapshot @done.
                      3. Final leaf is the merged image; zfs promote it into <store>/.templates/<image-config-sha> to flatten the chain into a standalone template.
                      4. Existing zfs::clone_container then clones the template for the user.

                      Whiteout application for step 2 needs a helper like zfs::apply_layer_whiteouts <layer_dir> <target>:

                      • For each *.wh.foo (AUFS) or 0:0-char-device named foo (overlayfs) in layer_dir: rm -rf "${target}/foo".
                      • For each .wh..wh..opq or trusted.overlay.opaque=y dir: clear children of corresponding dir in target.
                      • Then cp -a (or tar | tar) the rest of layer_dir over target.

                      enroot-aufs2ovlfs already converts to overlayfs whiteouts in-place, so the helper only needs the overlayfs forms. Worth confirming the exact char-device format it produces.

                      Coexistence with Plan F

                      • Default behavior unchanged: Plan F's single-merge path stays the default.
                      • ENROOT_ZFS_LAYER_CHAIN=y opts into per-layer.
                      • Both paths populate <store>/.templates/<sha> — same shape, same zfs::clone_container for the user. Only the fill mechanism differs.
                      • A site can switch on/off without migration; existing single-merge templates remain valid.

                      Open questions

                      • Where do layer datasets live? <store>/.layers/<sha> (parallel to .templates) keeps the layer cache separate from per-image templates. Avoids confusion when an admin scans .templates.
                      • zfs promote permissions: needs promote in the user's zfs allow. Document in the admin recipe.
                      • Should we GC unused layer datasets when no template references them? Plan B's sweep mechanism could be extended to .layers/ with the same warm/cold logic.
                      • How does Plan B's warm/cold eviction interact? Templates and layers should probably share the same lifecycle policy, but layers' lifecycle is "no template references" rather than "no clones."
                      • xattr propagation under cp -a vs tar: tar --xattrs --xattrs-include='*' --selinux --acls is the safer pipe; enroot-aufs2ovlfs likely emits these correctly already.

                      Acceptance criteria

                      • Two distinct images sharing a base layer (e.g. python:3-slim and node:20-slim, both Debian-based) store the shared layer once on disk under <store>/.layers/<digest>.
                      • Re-pulling an image after a top-layer-only update reuses the cached lower-layer datasets (verify with timing + zfs list snapshot before/after).
                      • Whiteouts and opaque dirs from real Docker images merge correctly (verified against python:3-slim, nginx, cuda base images).
                      • Concurrent enroot load of the same image is race-safe (same .tmp lock pattern as Plan F).
                      • ENROOT_ZFS_LAYER_CHAIN= (unset) leaves Plan F's behavior unchanged.

                      Out of scope

                      • Replacing Plan F's single-merge path. Plan G is purely additive.
                      • Cross-host layer replication via zfs send. Natural follow-up but tracked separately.
                      • Migration tooling between merged-template and per-layer-chain caches.

                      Activity

                      Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        No labels
                        No labels

                        Type

                        No type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in) #4

                          Description

                          @sodre

                          Follow-up to #3 (Plan F).

                          Plan F's enroot load docker:// path materializes the merged image into a single ZFS template per image (cached by image config digest). That covers most cases but loses several properties that the per-layer-clone-chain approach (mirroring Docker's own zfs storage driver) would give us. Plan G is the opt-in mode that adds the per-layer path alongside Plan F's single-merge path, gated on a config flag like ENROOT_ZFS_LAYER_CHAIN=y.

                          Why it's worth doing

                          For HPC / CI hosts that pull many images sharing common bases, the current single-merge design wastes disk and CPU. Per-layer chains would buy back:

                          1. Cross-image layer dedup at the dataset level. Two images sharing a Debian base store the base bytes once instead of twice. Block-level dedup=on recovers this in Plan F's design but at ~5–6 GB RAM per TB indexed; per-layer datasets dedup for free.
                          2. Incremental re-pull cost. When alpine:3.21 replaces alpine:3.20, only the changed top layers are re-extracted; lower-layer datasets are reused. Plan F re-merges the whole stack.
                          3. Layer-granular cache invalidation. A poisoned layer can be zfs destroyd in isolation; Plan F throws out the whole template.
                          4. Native ZFS introspection.zfs list -t all shows the layer chain; zfs send per-layer becomes a sensible cross-host replication primitive.
                          5. Quota accounting matches intuition.quota=200G on <store>/.templates reflects shared layers once, not multiplied by the number of images that use them.
                          6. Aligns with Docker's zfs storage driver shape, which ops people already know.

                          What it costs

                          • Whiteout/opaque-dir merging in shell. Overlayfs whiteouts are character-device files (mknod c 0 0); opaque-dir markers are trusted.overlay.opaque=y xattrs. Without the kernel's overlay engine doing the merge, we need to apply these manually during each clone-extract step. Real edge-case surface.
                          • ~5–15× more dataset objects per image. zfs list clutter; more bookkeeping.
                          • Per-layer atomic locks (multiple .tmp datasets racing).
                          • zfs promote (or chain-preservation alternative) to flatten the leaf into a standalone template — depends on user delegations.
                          • More complex cache invalidation logic (decide what to destroy when a leaf is reaped vs. when a shared lower layer is reaped).

                          Sketch of design

                          zfs::docker_install_from_layers (in src/storage_zfs.sh) gains a check: if ENROOT_ZFS_LAYER_CHAIN=y, dispatch to zfs::docker_install_chain instead. The new function:

                          1. For each layer in stack order, hash the layer tarball (already in ${ENROOT_CACHE_PATH}/<digest> from _prepare_layers) → cache key.
                          2. If <store>/.layers/<digest>@done exists, reuse; else zfs clone parent@done <new> (or zfs create for the base), apply layer's whiteouts + extracted contents, zfs snapshot @done.
                          3. Final leaf is the merged image; zfs promote it into <store>/.templates/<image-config-sha> to flatten the chain into a standalone template.
                          4. Existing zfs::clone_container then clones the template for the user.

                          Whiteout application for step 2 needs a helper like zfs::apply_layer_whiteouts <layer_dir> <target>:

                          • For each *.wh.foo (AUFS) or 0:0-char-device named foo (overlayfs) in layer_dir: rm -rf "${target}/foo".
                          • For each .wh..wh..opq or trusted.overlay.opaque=y dir: clear children of corresponding dir in target.
                          • Then cp -a (or tar | tar) the rest of layer_dir over target.

                          enroot-aufs2ovlfs already converts to overlayfs whiteouts in-place, so the helper only needs the overlayfs forms. Worth confirming the exact char-device format it produces.

                          Coexistence with Plan F

                          • Default behavior unchanged: Plan F's single-merge path stays the default.
                          • ENROOT_ZFS_LAYER_CHAIN=y opts into per-layer.
                          • Both paths populate <store>/.templates/<sha> — same shape, same zfs::clone_container for the user. Only the fill mechanism differs.
                          • A site can switch on/off without migration; existing single-merge templates remain valid.

                          Open questions

                          • Where do layer datasets live? <store>/.layers/<sha> (parallel to .templates) keeps the layer cache separate from per-image templates. Avoids confusion when an admin scans .templates.
                          • zfs promote permissions: needs promote in the user's zfs allow. Document in the admin recipe.
                          • Should we GC unused layer datasets when no template references them? Plan B's sweep mechanism could be extended to .layers/ with the same warm/cold logic.
                          • How does Plan B's warm/cold eviction interact? Templates and layers should probably share the same lifecycle policy, but layers' lifecycle is "no template references" rather than "no clones."
                          • xattr propagation under cp -a vs tar: tar --xattrs --xattrs-include='*' --selinux --acls is the safer pipe; enroot-aufs2ovlfs likely emits these correctly already.

                          Acceptance criteria

                          • Two distinct images sharing a base layer (e.g. python:3-slim and node:20-slim, both Debian-based) store the shared layer once on disk under <store>/.layers/<digest>.
                          • Re-pulling an image after a top-layer-only update reuses the cached lower-layer datasets (verify with timing + zfs list snapshot before/after).
                          • Whiteouts and opaque dirs from real Docker images merge correctly (verified against python:3-slim, nginx, cuda base images).
                          • Concurrent enroot load of the same image is race-safe (same .tmp lock pattern as Plan F).
                          • ENROOT_ZFS_LAYER_CHAIN= (unset) leaves Plan F's behavior unchanged.

                          Out of scope

                          • Replacing Plan F's single-merge path. Plan G is purely additive.
                          • Cross-host layer replication via zfs send. Natural follow-up but tracked separately.
                          • Migration tooling between merged-template and per-layer-chain caches.

                          Activity

                          Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            No labels
                            No labels

                            Type

                            No type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in) #4

                              Description

                              @sodre

                              Follow-up to #3 (Plan F).

                              Plan F's enroot load docker:// path materializes the merged image into a single ZFS template per image (cached by image config digest). That covers most cases but loses several properties that the per-layer-clone-chain approach (mirroring Docker's own zfs storage driver) would give us. Plan G is the opt-in mode that adds the per-layer path alongside Plan F's single-merge path, gated on a config flag like ENROOT_ZFS_LAYER_CHAIN=y.

                              Why it's worth doing

                              For HPC / CI hosts that pull many images sharing common bases, the current single-merge design wastes disk and CPU. Per-layer chains would buy back:

                              1. Cross-image layer dedup at the dataset level. Two images sharing a Debian base store the base bytes once instead of twice. Block-level dedup=on recovers this in Plan F's design but at ~5–6 GB RAM per TB indexed; per-layer datasets dedup for free.
                              2. Incremental re-pull cost. When alpine:3.21 replaces alpine:3.20, only the changed top layers are re-extracted; lower-layer datasets are reused. Plan F re-merges the whole stack.
                              3. Layer-granular cache invalidation. A poisoned layer can be zfs destroyd in isolation; Plan F throws out the whole template.
                              4. Native ZFS introspection.zfs list -t all shows the layer chain; zfs send per-layer becomes a sensible cross-host replication primitive.
                              5. Quota accounting matches intuition.quota=200G on <store>/.templates reflects shared layers once, not multiplied by the number of images that use them.
                              6. Aligns with Docker's zfs storage driver shape, which ops people already know.

                              What it costs

                              • Whiteout/opaque-dir merging in shell. Overlayfs whiteouts are character-device files (mknod c 0 0); opaque-dir markers are trusted.overlay.opaque=y xattrs. Without the kernel's overlay engine doing the merge, we need to apply these manually during each clone-extract step. Real edge-case surface.
                              • ~5–15× more dataset objects per image. zfs list clutter; more bookkeeping.
                              • Per-layer atomic locks (multiple .tmp datasets racing).
                              • zfs promote (or chain-preservation alternative) to flatten the leaf into a standalone template — depends on user delegations.
                              • More complex cache invalidation logic (decide what to destroy when a leaf is reaped vs. when a shared lower layer is reaped).

                              Sketch of design

                              zfs::docker_install_from_layers (in src/storage_zfs.sh) gains a check: if ENROOT_ZFS_LAYER_CHAIN=y, dispatch to zfs::docker_install_chain instead. The new function:

                              1. For each layer in stack order, hash the layer tarball (already in ${ENROOT_CACHE_PATH}/<digest> from _prepare_layers) → cache key.
                              2. If <store>/.layers/<digest>@done exists, reuse; else zfs clone parent@done <new> (or zfs create for the base), apply layer's whiteouts + extracted contents, zfs snapshot @done.
                              3. Final leaf is the merged image; zfs promote it into <store>/.templates/<image-config-sha> to flatten the chain into a standalone template.
                              4. Existing zfs::clone_container then clones the template for the user.

                              Whiteout application for step 2 needs a helper like zfs::apply_layer_whiteouts <layer_dir> <target>:

                              • For each *.wh.foo (AUFS) or 0:0-char-device named foo (overlayfs) in layer_dir: rm -rf "${target}/foo".
                              • For each .wh..wh..opq or trusted.overlay.opaque=y dir: clear children of corresponding dir in target.
                              • Then cp -a (or tar | tar) the rest of layer_dir over target.

                              enroot-aufs2ovlfs already converts to overlayfs whiteouts in-place, so the helper only needs the overlayfs forms. Worth confirming the exact char-device format it produces.

                              Coexistence with Plan F

                              • Default behavior unchanged: Plan F's single-merge path stays the default.
                              • ENROOT_ZFS_LAYER_CHAIN=y opts into per-layer.
                              • Both paths populate <store>/.templates/<sha> — same shape, same zfs::clone_container for the user. Only the fill mechanism differs.
                              • A site can switch on/off without migration; existing single-merge templates remain valid.

                              Open questions

                              • Where do layer datasets live? <store>/.layers/<sha> (parallel to .templates) keeps the layer cache separate from per-image templates. Avoids confusion when an admin scans .templates.
                              • zfs promote permissions: needs promote in the user's zfs allow. Document in the admin recipe.
                              • Should we GC unused layer datasets when no template references them? Plan B's sweep mechanism could be extended to .layers/ with the same warm/cold logic.
                              • How does Plan B's warm/cold eviction interact? Templates and layers should probably share the same lifecycle policy, but layers' lifecycle is "no template references" rather than "no clones."
                              • xattr propagation under cp -a vs tar: tar --xattrs --xattrs-include='*' --selinux --acls is the safer pipe; enroot-aufs2ovlfs likely emits these correctly already.

                              Acceptance criteria

                              • Two distinct images sharing a base layer (e.g. python:3-slim and node:20-slim, both Debian-based) store the shared layer once on disk under <store>/.layers/<digest>.
                              • Re-pulling an image after a top-layer-only update reuses the cached lower-layer datasets (verify with timing + zfs list snapshot before/after).
                              • Whiteouts and opaque dirs from real Docker images merge correctly (verified against python:3-slim, nginx, cuda base images).
                              • Concurrent enroot load of the same image is race-safe (same .tmp lock pattern as Plan F).
                              • ENROOT_ZFS_LAYER_CHAIN= (unset) leaves Plan F's behavior unchanged.

                              Out of scope

                              • Replacing Plan F's single-merge path. Plan G is purely additive.
                              • Cross-host layer replication via zfs send. Natural follow-up but tracked separately.
                              • Migration tooling between merged-template and per-layer-chain caches.

                              Activity

                              Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                No labels
                                No labels

                                Type

                                No type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions