Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in) - #17

Merged
sodre merged 14 commits into
zenroot/mainfrom
feature/zfs-g-layer-chain
May 6, 2026
Merged

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in)#17
sodre merged 14 commits into
zenroot/mainfrom
feature/zfs-g-layer-chain

Conversation

@sodre

@sodresodre commented May 2, 2026

Copy link
Copy Markdown
Member

Closes#4.

Adds an opt-in per-layer zfs clone chain mode for the Docker template cache. With ENROOT_ZFS_LAYER_CHAIN=y, two images sharing a registry layer digest physically share the bytes on disk; re-pulling an image after a top-layer-only change reuses the cached lower-layer datasets.

Layout

<store>/.layers/<layer-digest> # one per distinct registry layer (origin)
<store>/.layers/<layer-digest>@done # snapshot taken after layer apply
<store>/.templates/<image-config-sha> # zfs clone of the chain leaf @done
<store>/.templates/<image-config-sha>@pristine

Each layer dataset is zfs cloned from the previous layer's @done, with overlayfs whiteouts (mknod 0:0) and opaque-dir markers (trusted.overlay.opaque=y) replayed in shell on top of the cloned target — overlayfs only does that merge at mount time, but a chain stored at-rest needs it baked in. The chain leaf is then cloned into .templates/<config_sha>, the per-image synthetic 0/ config layer (rc/fstab/environment from docker::configure) is applied on top, and the result is snapshotted as @pristine so the existing zfs::clone_container, pointer-format, eviction-recovery, and zfs:// paths all work unchanged.

Why no zfs promote

The issue mentions promote as one option for flattening the chain. We don't promote — promoting inverts the chain (layers become clones of the template), which works for one image but produces a complex image-private topology that defeats the cross-image sharing goal. Plan G keeps layers as immutable origins; ZFS refuses to destroy a layer while any descendant clone exists, so layer GC is automatic once all referencing templates are evicted.

What's added

FileChange
src/storage_zfs.shzfs::layer_chain_active, zfs::_apply_layer_payload, zfs::_build_layer, zfs::_install_layer_chain; chain-mode dispatch in docker_install_from_layers and _pull_and_install_template.
src/docker.sh_prepare_layers side-emits the ordered layer-digest list to ./.layers in its temp cwd; docker::load's ZFS branch reads it back when chain mode is active.
pkg/deb/controlRecommends attr (provides getfattr, required by chain-mode opaque-dir handling).
doc/zfs.md, CLAUDE.mdDocument the new knob, store layout, and dedup semantics.
doc/plans/2026-05-01-zfs-g-layer-chain.mdImplementation plan, mirrors Plans A–F structure.

Coexistence with Plan F

  • ENROOT_ZFS_LAYER_CHAIN= (unset/empty/anything but y): Plan F's single-merge _install_template_from_layers runs unchanged.
  • ENROOT_ZFS_LAYER_CHAIN=y: chain mode. Same dispatch hits both docker::load (direct create) and _pull_and_install_template (used by pointer-format import + eviction recovery).
  • The fast path "template @pristine already exists, reuse it" runs before the chain dispatch — templates produced under either mode are reused under the other without rebuild.

Smoke results (spark-ctrl, Pi 5 / Debian 13 / OpenZFS 2.4.1, 3.75G test pool)

TestResult
Single-layer alpine, chain mode✓ pointer file written, layer dataset created, template clones leaf, rootfs has os-release + /etc/{rc,fstab,environment}
Multi-layer node:20-alpine✓ 3 layer datasets in BASE→TOP order, leaf REFERs full 69.9M merged tree, /usr/local/bin/node (102M binary) present in container rootfs
Multi-layer python:3.13-alpine✓ 3 layer datasets, clones-of-clones visible in zfs list (3070388042c6 1.03M USED / 19.3M REFER — pure dedup)
Whiteout/opaque sanity✓ no .wh.* AUFS files leak through, no char-device whiteouts in final rootfs
Layer reuse after template eviction✓ destroy templates+containers, keep layers; re-create from pointer takes 1.7s with NO "Building layer" messages, just final clone
Plan F regression (flag unset)✓ no .layers/ namespace created, _install_template_from_layers runs as before

Smoke testing also flagged two bugs that were fixed in 3f7e3af:

  • Inverted chain iteration order (docker::_download reverses the manifest, so digests[0] is the TOP, not the BASE).
  • Missing synthetic 0/ config layer apply on the leaf — Plan F's overlay mount stacks 0:1:…:N with 0/ on top; the chain installer needed an explicit final tar-pipe of 0/ onto the template.

Plus one packaging fix: attr is now Recommended (was Suggested), since getfattr is required for chain-mode opaque-dir handling and Suggests is not auto-installed.

🤖 Generated with Claude Code

sodre added 9 commits May 1, 2026 22:17
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
…ants
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
…ll path
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Smoke testing on a 3-layer image (node:20-alpine) caught two bugs in the
chain installer:
1. Inverted iteration. docker::_download reverses the manifest's layer
order via jq's `reverse`, so digests[0] is the TOP layer and
digests[N-1] is the BASE. The original `for i in 0..N-1` loop treated
digests[0] as the base, building the chain upside-down and producing
a leaf that contained only the top-layer's diffs (e.g. 5.4M for what
should have been a 70M merged node:20-alpine rootfs). Iterating from
N-1 down to 0 puts BASE first in the zfs hierarchy and the TOP at
the leaf.
2. Missing synthetic config layer. docker::_prepare_layers populates a
directory 0/ via docker::configure with the per-image
/etc/{rc,fstab,environment} derived from the image config blob; Plan
F's overlay mount stacks 0:1:2:...:N so 0/ ends up on top. The chain
installer ignored 0/ entirely, so containers created via chain mode
were missing /etc/rc and the merged fstab entries. Now applied as a
final tar-pipe step on top of the leaf clone during template
finalization, before snapshotting @pristine.
Also tighten the apply payload:
- getfattr returns non-zero when no files match the requested xattr;
with set -euo pipefail in the payload that aborted the whole apply on
alpine (no opaque dirs). Capture to a temp file with `|| true`.
- Drop tar's --acls. Default ZFS datasets have acltype=off, which makes
POSIX ACL set/get fail with "Operation not supported" warnings even
when the source has no ACLs. Docker images effectively never depend
on ACLs, and xattrs (overlayfs opaque markers, capability bits,
SELinux labels) are still preserved.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre
sodre marked this pull request as ready for review May 2, 2026 03:12
@sodre

sodre commented May 2, 2026

Copy link
Copy Markdown
MemberAuthor

Cross-image dedup verified on disk (acceptance criterion #1 from issue #4):

Pulled python:3.13-alpine3.21 then node:22-alpine3.21 (both pinned to the same alpine base) on a fresh store. Origin tree (zfs list -r -d 1 -o name,origin):

c2fe130f4aab (alpine 3.21 base, 5.11M USED, ORIGIN)
├── 1c6063f559a3 → 838d25d4769a → 3517b1771ef3 (python:3.13-alpine3.21)
└── 720ee653d3d4 → 055ee03d01c9 → 7f3c333e617d (node:22-alpine3.21)

Both chains branch from c2fe130f4aab…@done. The 5.11M alpine base is stored once on disk; ZFS protects it via "snapshot has clones" so it's automatically immortal as long as either python or node is still around.

@sodre

sodre commented May 2, 2026

Copy link
Copy Markdown
MemberAuthor

Two more acceptance criteria covered.

Concurrent pull of the same image — race-safe

ENROOT_ZFS_LAYER_CHAIN=y enroot import -o /tmp/n1.sqsh docker://node:22-alpine3.21 & p1=$!
ENROOT_ZFS_LAYER_CHAIN=y enroot import -o /tmp/n2.sqsh docker://node:22-alpine3.21 & p2=$!wait$p1$p2

Result: both processes returned 0; only one set of [INFO] Building layer ... messages printed (4 messages, one per node:22-alpine3.21 layer); on disk: exactly 4 layer datasets, no .tmp orphans, both containers usable. The per-layer <digest>.tmp lock collapsed both invocations onto a single builder; the loser waited for @done silently.

Lower-layer reuse on a second pull that shares a base

When node:22-alpine3.21 is imported after python:3.13-alpine3.21, the chain installer logged 3 Building layer messages — not 4 — because alpine 3.21's c2fe130f4aab@done was already in the cache from python. The node chain branched directly off that snapshot via zfs clone, so the second pull paid the layer-build cost only for node-specific content.

This generalizes to the issue's "top-layer-only re-pull" case: when a docker tag is republished with only the top digest changed, every cached lower-layer <digest>@done is a hit and _build_layer skips immediately to the clone-on-top step.

Plan G applies only to registry-pulled docker:// URIs. Daemon-local
URIs (dockerd://, podman://) take a separate path that uses
\`${engine} export | tar -x\` (flat rootfs) instead of layer tarballs,
so there is no per-layer structure for the chain installer to consume.
Spell this out in the plan's Coexistence section and the user-facing
knob description in doc/zfs.md, plus add a future-work note in Out of
scope describing what bringing chain mode to daemon URIs would require
(switching to \`${engine} save\` and parsing manifest.json).
The current code already silently no-ops for daemon URIs; this is a
docs-only commit clarifying the boundary.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre

sodre commented May 6, 2026

Copy link
Copy Markdown
MemberAuthor

Scope clarified: `docker://` only.

`dockerd://` and `podman://` URIs take a separate path (`${engine} export | tar -x` → `zfs::_install_template_from_dir`) which flattens the image into a single rootfs before extraction. There is no per-layer structure for Plan G's chain installer to consume, so `ENROOT_ZFS_LAYER_CHAIN=y` is silently a no-op for daemon URIs (the daemon path's behavior is unchanged either way).

Bringing chain mode to daemon URIs is feasible but is a real follow-up plan: it requires switching from `${engine} export` to `${engine} save` (which preserves per-layer tarballs in a tar archive plus a `manifest.json` describing layer order), parsing that manifest, and constructing a synthetic `0/` from `${engine} inspect` output. Documented in plan's Out of scope and the `ENROOT_ZFS_LAYER_CHAIN` knob description in `doc/zfs.md`.

sodre added 2 commits May 6, 2026 06:52
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
sodre added a commit that referenced this pull request May 6, 2026
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Plan-doc only — depends on Plan G's helpers, which land in PR #17.
The doc/plans/README.md index and doc/zfs.md knob description will
follow when both Plan G and Plan H land on zenroot/main.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre
sodre merged commit c4baea6 into zenroot/mainMay 6, 2026
sodre added a commit that referenced this pull request May 6, 2026
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Plan-doc only — depends on Plan G's helpers, which land in PR #17.
The doc/plans/README.md index and doc/zfs.md knob description will
follow when both Plan G and Plan H land on zenroot/main.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in)

1 participant

@sodre
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in) - #17

Merged
sodre merged 14 commits into
zenroot/mainfrom
feature/zfs-g-layer-chain
May 6, 2026
Merged

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in)#17
sodre merged 14 commits into
zenroot/mainfrom
feature/zfs-g-layer-chain

Conversation

@sodre

@sodresodre commented May 2, 2026

Copy link
Copy Markdown
Member

Closes#4.

Adds an opt-in per-layer zfs clone chain mode for the Docker template cache. With ENROOT_ZFS_LAYER_CHAIN=y, two images sharing a registry layer digest physically share the bytes on disk; re-pulling an image after a top-layer-only change reuses the cached lower-layer datasets.

Layout

<store>/.layers/<layer-digest> # one per distinct registry layer (origin)
<store>/.layers/<layer-digest>@done # snapshot taken after layer apply
<store>/.templates/<image-config-sha> # zfs clone of the chain leaf @done
<store>/.templates/<image-config-sha>@pristine

Each layer dataset is zfs cloned from the previous layer's @done, with overlayfs whiteouts (mknod 0:0) and opaque-dir markers (trusted.overlay.opaque=y) replayed in shell on top of the cloned target — overlayfs only does that merge at mount time, but a chain stored at-rest needs it baked in. The chain leaf is then cloned into .templates/<config_sha>, the per-image synthetic 0/ config layer (rc/fstab/environment from docker::configure) is applied on top, and the result is snapshotted as @pristine so the existing zfs::clone_container, pointer-format, eviction-recovery, and zfs:// paths all work unchanged.

Why no zfs promote

The issue mentions promote as one option for flattening the chain. We don't promote — promoting inverts the chain (layers become clones of the template), which works for one image but produces a complex image-private topology that defeats the cross-image sharing goal. Plan G keeps layers as immutable origins; ZFS refuses to destroy a layer while any descendant clone exists, so layer GC is automatic once all referencing templates are evicted.

What's added

FileChange
src/storage_zfs.shzfs::layer_chain_active, zfs::_apply_layer_payload, zfs::_build_layer, zfs::_install_layer_chain; chain-mode dispatch in docker_install_from_layers and _pull_and_install_template.
src/docker.sh_prepare_layers side-emits the ordered layer-digest list to ./.layers in its temp cwd; docker::load's ZFS branch reads it back when chain mode is active.
pkg/deb/controlRecommends attr (provides getfattr, required by chain-mode opaque-dir handling).
doc/zfs.md, CLAUDE.mdDocument the new knob, store layout, and dedup semantics.
doc/plans/2026-05-01-zfs-g-layer-chain.mdImplementation plan, mirrors Plans A–F structure.

Coexistence with Plan F

  • ENROOT_ZFS_LAYER_CHAIN= (unset/empty/anything but y): Plan F's single-merge _install_template_from_layers runs unchanged.
  • ENROOT_ZFS_LAYER_CHAIN=y: chain mode. Same dispatch hits both docker::load (direct create) and _pull_and_install_template (used by pointer-format import + eviction recovery).
  • The fast path "template @pristine already exists, reuse it" runs before the chain dispatch — templates produced under either mode are reused under the other without rebuild.

Smoke results (spark-ctrl, Pi 5 / Debian 13 / OpenZFS 2.4.1, 3.75G test pool)

TestResult
Single-layer alpine, chain mode✓ pointer file written, layer dataset created, template clones leaf, rootfs has os-release + /etc/{rc,fstab,environment}
Multi-layer node:20-alpine✓ 3 layer datasets in BASE→TOP order, leaf REFERs full 69.9M merged tree, /usr/local/bin/node (102M binary) present in container rootfs
Multi-layer python:3.13-alpine✓ 3 layer datasets, clones-of-clones visible in zfs list (3070388042c6 1.03M USED / 19.3M REFER — pure dedup)
Whiteout/opaque sanity✓ no .wh.* AUFS files leak through, no char-device whiteouts in final rootfs
Layer reuse after template eviction✓ destroy templates+containers, keep layers; re-create from pointer takes 1.7s with NO "Building layer" messages, just final clone
Plan F regression (flag unset)✓ no .layers/ namespace created, _install_template_from_layers runs as before

Smoke testing also flagged two bugs that were fixed in 3f7e3af:

  • Inverted chain iteration order (docker::_download reverses the manifest, so digests[0] is the TOP, not the BASE).
  • Missing synthetic 0/ config layer apply on the leaf — Plan F's overlay mount stacks 0:1:…:N with 0/ on top; the chain installer needed an explicit final tar-pipe of 0/ onto the template.

Plus one packaging fix: attr is now Recommended (was Suggested), since getfattr is required for chain-mode opaque-dir handling and Suggests is not auto-installed.

🤖 Generated with Claude Code

sodre added 9 commits May 1, 2026 22:17
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
…ants
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
…ll path
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Smoke testing on a 3-layer image (node:20-alpine) caught two bugs in the
chain installer:
1. Inverted iteration. docker::_download reverses the manifest's layer
order via jq's `reverse`, so digests[0] is the TOP layer and
digests[N-1] is the BASE. The original `for i in 0..N-1` loop treated
digests[0] as the base, building the chain upside-down and producing
a leaf that contained only the top-layer's diffs (e.g. 5.4M for what
should have been a 70M merged node:20-alpine rootfs). Iterating from
N-1 down to 0 puts BASE first in the zfs hierarchy and the TOP at
the leaf.
2. Missing synthetic config layer. docker::_prepare_layers populates a
directory 0/ via docker::configure with the per-image
/etc/{rc,fstab,environment} derived from the image config blob; Plan
F's overlay mount stacks 0:1:2:...:N so 0/ ends up on top. The chain
installer ignored 0/ entirely, so containers created via chain mode
were missing /etc/rc and the merged fstab entries. Now applied as a
final tar-pipe step on top of the leaf clone during template
finalization, before snapshotting @pristine.
Also tighten the apply payload:
- getfattr returns non-zero when no files match the requested xattr;
with set -euo pipefail in the payload that aborted the whole apply on
alpine (no opaque dirs). Capture to a temp file with `|| true`.
- Drop tar's --acls. Default ZFS datasets have acltype=off, which makes
POSIX ACL set/get fail with "Operation not supported" warnings even
when the source has no ACLs. Docker images effectively never depend
on ACLs, and xattrs (overlayfs opaque markers, capability bits,
SELinux labels) are still preserved.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre
sodre marked this pull request as ready for review May 2, 2026 03:12
@sodre

sodre commented May 2, 2026

Copy link
Copy Markdown
MemberAuthor

Cross-image dedup verified on disk (acceptance criterion #1 from issue #4):

Pulled python:3.13-alpine3.21 then node:22-alpine3.21 (both pinned to the same alpine base) on a fresh store. Origin tree (zfs list -r -d 1 -o name,origin):

c2fe130f4aab (alpine 3.21 base, 5.11M USED, ORIGIN)
├── 1c6063f559a3 → 838d25d4769a → 3517b1771ef3 (python:3.13-alpine3.21)
└── 720ee653d3d4 → 055ee03d01c9 → 7f3c333e617d (node:22-alpine3.21)

Both chains branch from c2fe130f4aab…@done. The 5.11M alpine base is stored once on disk; ZFS protects it via "snapshot has clones" so it's automatically immortal as long as either python or node is still around.

@sodre

sodre commented May 2, 2026

Copy link
Copy Markdown
MemberAuthor

Two more acceptance criteria covered.

Concurrent pull of the same image — race-safe

ENROOT_ZFS_LAYER_CHAIN=y enroot import -o /tmp/n1.sqsh docker://node:22-alpine3.21 & p1=$!
ENROOT_ZFS_LAYER_CHAIN=y enroot import -o /tmp/n2.sqsh docker://node:22-alpine3.21 & p2=$!wait$p1$p2

Result: both processes returned 0; only one set of [INFO] Building layer ... messages printed (4 messages, one per node:22-alpine3.21 layer); on disk: exactly 4 layer datasets, no .tmp orphans, both containers usable. The per-layer <digest>.tmp lock collapsed both invocations onto a single builder; the loser waited for @done silently.

Lower-layer reuse on a second pull that shares a base

When node:22-alpine3.21 is imported after python:3.13-alpine3.21, the chain installer logged 3 Building layer messages — not 4 — because alpine 3.21's c2fe130f4aab@done was already in the cache from python. The node chain branched directly off that snapshot via zfs clone, so the second pull paid the layer-build cost only for node-specific content.

This generalizes to the issue's "top-layer-only re-pull" case: when a docker tag is republished with only the top digest changed, every cached lower-layer <digest>@done is a hit and _build_layer skips immediately to the clone-on-top step.

Plan G applies only to registry-pulled docker:// URIs. Daemon-local
URIs (dockerd://, podman://) take a separate path that uses
\`${engine} export | tar -x\` (flat rootfs) instead of layer tarballs,
so there is no per-layer structure for the chain installer to consume.
Spell this out in the plan's Coexistence section and the user-facing
knob description in doc/zfs.md, plus add a future-work note in Out of
scope describing what bringing chain mode to daemon URIs would require
(switching to \`${engine} save\` and parsing manifest.json).
The current code already silently no-ops for daemon URIs; this is a
docs-only commit clarifying the boundary.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre

sodre commented May 6, 2026

Copy link
Copy Markdown
MemberAuthor

Scope clarified: `docker://` only.

`dockerd://` and `podman://` URIs take a separate path (`${engine} export | tar -x` → `zfs::_install_template_from_dir`) which flattens the image into a single rootfs before extraction. There is no per-layer structure for Plan G's chain installer to consume, so `ENROOT_ZFS_LAYER_CHAIN=y` is silently a no-op for daemon URIs (the daemon path's behavior is unchanged either way).

Bringing chain mode to daemon URIs is feasible but is a real follow-up plan: it requires switching from `${engine} export` to `${engine} save` (which preserves per-layer tarballs in a tar archive plus a `manifest.json` describing layer order), parsing that manifest, and constructing a synthetic `0/` from `${engine} inspect` output. Documented in plan's Out of scope and the `ENROOT_ZFS_LAYER_CHAIN` knob description in `doc/zfs.md`.

sodre added 2 commits May 6, 2026 06:52
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
sodre added a commit that referenced this pull request May 6, 2026
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Plan-doc only — depends on Plan G's helpers, which land in PR #17.
The doc/plans/README.md index and doc/zfs.md knob description will
follow when both Plan G and Plan H land on zenroot/main.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre
sodre merged commit c4baea6 into zenroot/mainMay 6, 2026
sodre added a commit that referenced this pull request May 6, 2026
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Plan-doc only — depends on Plan G's helpers, which land in PR #17.
The doc/plans/README.md index and doc/zfs.md knob description will
follow when both Plan G and Plan H land on zenroot/main.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in)

1 participant

@sodre
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in) - #17

Merged
sodre merged 14 commits into
zenroot/mainfrom
feature/zfs-g-layer-chain
May 6, 2026
Merged

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in)#17
sodre merged 14 commits into
zenroot/mainfrom
feature/zfs-g-layer-chain

Conversation

@sodre

@sodresodre commented May 2, 2026

Copy link
Copy Markdown
Member

Closes#4.

Adds an opt-in per-layer zfs clone chain mode for the Docker template cache. With ENROOT_ZFS_LAYER_CHAIN=y, two images sharing a registry layer digest physically share the bytes on disk; re-pulling an image after a top-layer-only change reuses the cached lower-layer datasets.

Layout

<store>/.layers/<layer-digest> # one per distinct registry layer (origin)
<store>/.layers/<layer-digest>@done # snapshot taken after layer apply
<store>/.templates/<image-config-sha> # zfs clone of the chain leaf @done
<store>/.templates/<image-config-sha>@pristine

Each layer dataset is zfs cloned from the previous layer's @done, with overlayfs whiteouts (mknod 0:0) and opaque-dir markers (trusted.overlay.opaque=y) replayed in shell on top of the cloned target — overlayfs only does that merge at mount time, but a chain stored at-rest needs it baked in. The chain leaf is then cloned into .templates/<config_sha>, the per-image synthetic 0/ config layer (rc/fstab/environment from docker::configure) is applied on top, and the result is snapshotted as @pristine so the existing zfs::clone_container, pointer-format, eviction-recovery, and zfs:// paths all work unchanged.

Why no zfs promote

The issue mentions promote as one option for flattening the chain. We don't promote — promoting inverts the chain (layers become clones of the template), which works for one image but produces a complex image-private topology that defeats the cross-image sharing goal. Plan G keeps layers as immutable origins; ZFS refuses to destroy a layer while any descendant clone exists, so layer GC is automatic once all referencing templates are evicted.

What's added

FileChange
src/storage_zfs.shzfs::layer_chain_active, zfs::_apply_layer_payload, zfs::_build_layer, zfs::_install_layer_chain; chain-mode dispatch in docker_install_from_layers and _pull_and_install_template.
src/docker.sh_prepare_layers side-emits the ordered layer-digest list to ./.layers in its temp cwd; docker::load's ZFS branch reads it back when chain mode is active.
pkg/deb/controlRecommends attr (provides getfattr, required by chain-mode opaque-dir handling).
doc/zfs.md, CLAUDE.mdDocument the new knob, store layout, and dedup semantics.
doc/plans/2026-05-01-zfs-g-layer-chain.mdImplementation plan, mirrors Plans A–F structure.

Coexistence with Plan F

  • ENROOT_ZFS_LAYER_CHAIN= (unset/empty/anything but y): Plan F's single-merge _install_template_from_layers runs unchanged.
  • ENROOT_ZFS_LAYER_CHAIN=y: chain mode. Same dispatch hits both docker::load (direct create) and _pull_and_install_template (used by pointer-format import + eviction recovery).
  • The fast path "template @pristine already exists, reuse it" runs before the chain dispatch — templates produced under either mode are reused under the other without rebuild.

Smoke results (spark-ctrl, Pi 5 / Debian 13 / OpenZFS 2.4.1, 3.75G test pool)

TestResult
Single-layer alpine, chain mode✓ pointer file written, layer dataset created, template clones leaf, rootfs has os-release + /etc/{rc,fstab,environment}
Multi-layer node:20-alpine✓ 3 layer datasets in BASE→TOP order, leaf REFERs full 69.9M merged tree, /usr/local/bin/node (102M binary) present in container rootfs
Multi-layer python:3.13-alpine✓ 3 layer datasets, clones-of-clones visible in zfs list (3070388042c6 1.03M USED / 19.3M REFER — pure dedup)
Whiteout/opaque sanity✓ no .wh.* AUFS files leak through, no char-device whiteouts in final rootfs
Layer reuse after template eviction✓ destroy templates+containers, keep layers; re-create from pointer takes 1.7s with NO "Building layer" messages, just final clone
Plan F regression (flag unset)✓ no .layers/ namespace created, _install_template_from_layers runs as before

Smoke testing also flagged two bugs that were fixed in 3f7e3af:

  • Inverted chain iteration order (docker::_download reverses the manifest, so digests[0] is the TOP, not the BASE).
  • Missing synthetic 0/ config layer apply on the leaf — Plan F's overlay mount stacks 0:1:…:N with 0/ on top; the chain installer needed an explicit final tar-pipe of 0/ onto the template.

Plus one packaging fix: attr is now Recommended (was Suggested), since getfattr is required for chain-mode opaque-dir handling and Suggests is not auto-installed.

🤖 Generated with Claude Code

sodre added 9 commits May 1, 2026 22:17
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
…ants
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
…ll path
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Smoke testing on a 3-layer image (node:20-alpine) caught two bugs in the
chain installer:
1. Inverted iteration. docker::_download reverses the manifest's layer
order via jq's `reverse`, so digests[0] is the TOP layer and
digests[N-1] is the BASE. The original `for i in 0..N-1` loop treated
digests[0] as the base, building the chain upside-down and producing
a leaf that contained only the top-layer's diffs (e.g. 5.4M for what
should have been a 70M merged node:20-alpine rootfs). Iterating from
N-1 down to 0 puts BASE first in the zfs hierarchy and the TOP at
the leaf.
2. Missing synthetic config layer. docker::_prepare_layers populates a
directory 0/ via docker::configure with the per-image
/etc/{rc,fstab,environment} derived from the image config blob; Plan
F's overlay mount stacks 0:1:2:...:N so 0/ ends up on top. The chain
installer ignored 0/ entirely, so containers created via chain mode
were missing /etc/rc and the merged fstab entries. Now applied as a
final tar-pipe step on top of the leaf clone during template
finalization, before snapshotting @pristine.
Also tighten the apply payload:
- getfattr returns non-zero when no files match the requested xattr;
with set -euo pipefail in the payload that aborted the whole apply on
alpine (no opaque dirs). Capture to a temp file with `|| true`.
- Drop tar's --acls. Default ZFS datasets have acltype=off, which makes
POSIX ACL set/get fail with "Operation not supported" warnings even
when the source has no ACLs. Docker images effectively never depend
on ACLs, and xattrs (overlayfs opaque markers, capability bits,
SELinux labels) are still preserved.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre
sodre marked this pull request as ready for review May 2, 2026 03:12
@sodre

sodre commented May 2, 2026

Copy link
Copy Markdown
MemberAuthor

Cross-image dedup verified on disk (acceptance criterion #1 from issue #4):

Pulled python:3.13-alpine3.21 then node:22-alpine3.21 (both pinned to the same alpine base) on a fresh store. Origin tree (zfs list -r -d 1 -o name,origin):

c2fe130f4aab (alpine 3.21 base, 5.11M USED, ORIGIN)
├── 1c6063f559a3 → 838d25d4769a → 3517b1771ef3 (python:3.13-alpine3.21)
└── 720ee653d3d4 → 055ee03d01c9 → 7f3c333e617d (node:22-alpine3.21)

Both chains branch from c2fe130f4aab…@done. The 5.11M alpine base is stored once on disk; ZFS protects it via "snapshot has clones" so it's automatically immortal as long as either python or node is still around.

@sodre

sodre commented May 2, 2026

Copy link
Copy Markdown
MemberAuthor

Two more acceptance criteria covered.

Concurrent pull of the same image — race-safe

ENROOT_ZFS_LAYER_CHAIN=y enroot import -o /tmp/n1.sqsh docker://node:22-alpine3.21 & p1=$!
ENROOT_ZFS_LAYER_CHAIN=y enroot import -o /tmp/n2.sqsh docker://node:22-alpine3.21 & p2=$!wait$p1$p2

Result: both processes returned 0; only one set of [INFO] Building layer ... messages printed (4 messages, one per node:22-alpine3.21 layer); on disk: exactly 4 layer datasets, no .tmp orphans, both containers usable. The per-layer <digest>.tmp lock collapsed both invocations onto a single builder; the loser waited for @done silently.

Lower-layer reuse on a second pull that shares a base

When node:22-alpine3.21 is imported after python:3.13-alpine3.21, the chain installer logged 3 Building layer messages — not 4 — because alpine 3.21's c2fe130f4aab@done was already in the cache from python. The node chain branched directly off that snapshot via zfs clone, so the second pull paid the layer-build cost only for node-specific content.

This generalizes to the issue's "top-layer-only re-pull" case: when a docker tag is republished with only the top digest changed, every cached lower-layer <digest>@done is a hit and _build_layer skips immediately to the clone-on-top step.

Plan G applies only to registry-pulled docker:// URIs. Daemon-local
URIs (dockerd://, podman://) take a separate path that uses
\`${engine} export | tar -x\` (flat rootfs) instead of layer tarballs,
so there is no per-layer structure for the chain installer to consume.
Spell this out in the plan's Coexistence section and the user-facing
knob description in doc/zfs.md, plus add a future-work note in Out of
scope describing what bringing chain mode to daemon URIs would require
(switching to \`${engine} save\` and parsing manifest.json).
The current code already silently no-ops for daemon URIs; this is a
docs-only commit clarifying the boundary.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre

sodre commented May 6, 2026

Copy link
Copy Markdown
MemberAuthor

Scope clarified: `docker://` only.

`dockerd://` and `podman://` URIs take a separate path (`${engine} export | tar -x` → `zfs::_install_template_from_dir`) which flattens the image into a single rootfs before extraction. There is no per-layer structure for Plan G's chain installer to consume, so `ENROOT_ZFS_LAYER_CHAIN=y` is silently a no-op for daemon URIs (the daemon path's behavior is unchanged either way).

Bringing chain mode to daemon URIs is feasible but is a real follow-up plan: it requires switching from `${engine} export` to `${engine} save` (which preserves per-layer tarballs in a tar archive plus a `manifest.json` describing layer order), parsing that manifest, and constructing a synthetic `0/` from `${engine} inspect` output. Documented in plan's Out of scope and the `ENROOT_ZFS_LAYER_CHAIN` knob description in `doc/zfs.md`.

sodre added 2 commits May 6, 2026 06:52
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
sodre added a commit that referenced this pull request May 6, 2026
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Plan-doc only — depends on Plan G's helpers, which land in PR #17.
The doc/plans/README.md index and doc/zfs.md knob description will
follow when both Plan G and Plan H land on zenroot/main.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre
sodre merged commit c4baea6 into zenroot/mainMay 6, 2026
sodre added a commit that referenced this pull request May 6, 2026
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Plan-doc only — depends on Plan G's helpers, which land in PR #17.
The doc/plans/README.md index and doc/zfs.md knob description will
follow when both Plan G and Plan H land on zenroot/main.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in)

1 participant

@sodre
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in) - #17

Merged
sodre merged 14 commits into
zenroot/mainfrom
feature/zfs-g-layer-chain
May 6, 2026
Merged

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in)#17
sodre merged 14 commits into
zenroot/mainfrom
feature/zfs-g-layer-chain

Conversation

@sodre

@sodresodre commented May 2, 2026

Copy link
Copy Markdown
Member

Closes#4.

Adds an opt-in per-layer zfs clone chain mode for the Docker template cache. With ENROOT_ZFS_LAYER_CHAIN=y, two images sharing a registry layer digest physically share the bytes on disk; re-pulling an image after a top-layer-only change reuses the cached lower-layer datasets.

Layout

<store>/.layers/<layer-digest> # one per distinct registry layer (origin)
<store>/.layers/<layer-digest>@done # snapshot taken after layer apply
<store>/.templates/<image-config-sha> # zfs clone of the chain leaf @done
<store>/.templates/<image-config-sha>@pristine

Each layer dataset is zfs cloned from the previous layer's @done, with overlayfs whiteouts (mknod 0:0) and opaque-dir markers (trusted.overlay.opaque=y) replayed in shell on top of the cloned target — overlayfs only does that merge at mount time, but a chain stored at-rest needs it baked in. The chain leaf is then cloned into .templates/<config_sha>, the per-image synthetic 0/ config layer (rc/fstab/environment from docker::configure) is applied on top, and the result is snapshotted as @pristine so the existing zfs::clone_container, pointer-format, eviction-recovery, and zfs:// paths all work unchanged.

Why no zfs promote

The issue mentions promote as one option for flattening the chain. We don't promote — promoting inverts the chain (layers become clones of the template), which works for one image but produces a complex image-private topology that defeats the cross-image sharing goal. Plan G keeps layers as immutable origins; ZFS refuses to destroy a layer while any descendant clone exists, so layer GC is automatic once all referencing templates are evicted.

What's added

FileChange
src/storage_zfs.shzfs::layer_chain_active, zfs::_apply_layer_payload, zfs::_build_layer, zfs::_install_layer_chain; chain-mode dispatch in docker_install_from_layers and _pull_and_install_template.
src/docker.sh_prepare_layers side-emits the ordered layer-digest list to ./.layers in its temp cwd; docker::load's ZFS branch reads it back when chain mode is active.
pkg/deb/controlRecommends attr (provides getfattr, required by chain-mode opaque-dir handling).
doc/zfs.md, CLAUDE.mdDocument the new knob, store layout, and dedup semantics.
doc/plans/2026-05-01-zfs-g-layer-chain.mdImplementation plan, mirrors Plans A–F structure.

Coexistence with Plan F

  • ENROOT_ZFS_LAYER_CHAIN= (unset/empty/anything but y): Plan F's single-merge _install_template_from_layers runs unchanged.
  • ENROOT_ZFS_LAYER_CHAIN=y: chain mode. Same dispatch hits both docker::load (direct create) and _pull_and_install_template (used by pointer-format import + eviction recovery).
  • The fast path "template @pristine already exists, reuse it" runs before the chain dispatch — templates produced under either mode are reused under the other without rebuild.

Smoke results (spark-ctrl, Pi 5 / Debian 13 / OpenZFS 2.4.1, 3.75G test pool)

TestResult
Single-layer alpine, chain mode✓ pointer file written, layer dataset created, template clones leaf, rootfs has os-release + /etc/{rc,fstab,environment}
Multi-layer node:20-alpine✓ 3 layer datasets in BASE→TOP order, leaf REFERs full 69.9M merged tree, /usr/local/bin/node (102M binary) present in container rootfs
Multi-layer python:3.13-alpine✓ 3 layer datasets, clones-of-clones visible in zfs list (3070388042c6 1.03M USED / 19.3M REFER — pure dedup)
Whiteout/opaque sanity✓ no .wh.* AUFS files leak through, no char-device whiteouts in final rootfs
Layer reuse after template eviction✓ destroy templates+containers, keep layers; re-create from pointer takes 1.7s with NO "Building layer" messages, just final clone
Plan F regression (flag unset)✓ no .layers/ namespace created, _install_template_from_layers runs as before

Smoke testing also flagged two bugs that were fixed in 3f7e3af:

  • Inverted chain iteration order (docker::_download reverses the manifest, so digests[0] is the TOP, not the BASE).
  • Missing synthetic 0/ config layer apply on the leaf — Plan F's overlay mount stacks 0:1:…:N with 0/ on top; the chain installer needed an explicit final tar-pipe of 0/ onto the template.

Plus one packaging fix: attr is now Recommended (was Suggested), since getfattr is required for chain-mode opaque-dir handling and Suggests is not auto-installed.

🤖 Generated with Claude Code

sodre added 9 commits May 1, 2026 22:17
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
…ants
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
…ll path
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Smoke testing on a 3-layer image (node:20-alpine) caught two bugs in the
chain installer:
1. Inverted iteration. docker::_download reverses the manifest's layer
order via jq's `reverse`, so digests[0] is the TOP layer and
digests[N-1] is the BASE. The original `for i in 0..N-1` loop treated
digests[0] as the base, building the chain upside-down and producing
a leaf that contained only the top-layer's diffs (e.g. 5.4M for what
should have been a 70M merged node:20-alpine rootfs). Iterating from
N-1 down to 0 puts BASE first in the zfs hierarchy and the TOP at
the leaf.
2. Missing synthetic config layer. docker::_prepare_layers populates a
directory 0/ via docker::configure with the per-image
/etc/{rc,fstab,environment} derived from the image config blob; Plan
F's overlay mount stacks 0:1:2:...:N so 0/ ends up on top. The chain
installer ignored 0/ entirely, so containers created via chain mode
were missing /etc/rc and the merged fstab entries. Now applied as a
final tar-pipe step on top of the leaf clone during template
finalization, before snapshotting @pristine.
Also tighten the apply payload:
- getfattr returns non-zero when no files match the requested xattr;
with set -euo pipefail in the payload that aborted the whole apply on
alpine (no opaque dirs). Capture to a temp file with `|| true`.
- Drop tar's --acls. Default ZFS datasets have acltype=off, which makes
POSIX ACL set/get fail with "Operation not supported" warnings even
when the source has no ACLs. Docker images effectively never depend
on ACLs, and xattrs (overlayfs opaque markers, capability bits,
SELinux labels) are still preserved.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre
sodre marked this pull request as ready for review May 2, 2026 03:12
@sodre

sodre commented May 2, 2026

Copy link
Copy Markdown
MemberAuthor

Cross-image dedup verified on disk (acceptance criterion #1 from issue #4):

Pulled python:3.13-alpine3.21 then node:22-alpine3.21 (both pinned to the same alpine base) on a fresh store. Origin tree (zfs list -r -d 1 -o name,origin):

c2fe130f4aab (alpine 3.21 base, 5.11M USED, ORIGIN)
├── 1c6063f559a3 → 838d25d4769a → 3517b1771ef3 (python:3.13-alpine3.21)
└── 720ee653d3d4 → 055ee03d01c9 → 7f3c333e617d (node:22-alpine3.21)

Both chains branch from c2fe130f4aab…@done. The 5.11M alpine base is stored once on disk; ZFS protects it via "snapshot has clones" so it's automatically immortal as long as either python or node is still around.

@sodre

sodre commented May 2, 2026

Copy link
Copy Markdown
MemberAuthor

Two more acceptance criteria covered.

Concurrent pull of the same image — race-safe

ENROOT_ZFS_LAYER_CHAIN=y enroot import -o /tmp/n1.sqsh docker://node:22-alpine3.21 & p1=$!
ENROOT_ZFS_LAYER_CHAIN=y enroot import -o /tmp/n2.sqsh docker://node:22-alpine3.21 & p2=$!wait$p1$p2

Result: both processes returned 0; only one set of [INFO] Building layer ... messages printed (4 messages, one per node:22-alpine3.21 layer); on disk: exactly 4 layer datasets, no .tmp orphans, both containers usable. The per-layer <digest>.tmp lock collapsed both invocations onto a single builder; the loser waited for @done silently.

Lower-layer reuse on a second pull that shares a base

When node:22-alpine3.21 is imported after python:3.13-alpine3.21, the chain installer logged 3 Building layer messages — not 4 — because alpine 3.21's c2fe130f4aab@done was already in the cache from python. The node chain branched directly off that snapshot via zfs clone, so the second pull paid the layer-build cost only for node-specific content.

This generalizes to the issue's "top-layer-only re-pull" case: when a docker tag is republished with only the top digest changed, every cached lower-layer <digest>@done is a hit and _build_layer skips immediately to the clone-on-top step.

Plan G applies only to registry-pulled docker:// URIs. Daemon-local
URIs (dockerd://, podman://) take a separate path that uses
\`${engine} export | tar -x\` (flat rootfs) instead of layer tarballs,
so there is no per-layer structure for the chain installer to consume.
Spell this out in the plan's Coexistence section and the user-facing
knob description in doc/zfs.md, plus add a future-work note in Out of
scope describing what bringing chain mode to daemon URIs would require
(switching to \`${engine} save\` and parsing manifest.json).
The current code already silently no-ops for daemon URIs; this is a
docs-only commit clarifying the boundary.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre

sodre commented May 6, 2026

Copy link
Copy Markdown
MemberAuthor

Scope clarified: `docker://` only.

`dockerd://` and `podman://` URIs take a separate path (`${engine} export | tar -x` → `zfs::_install_template_from_dir`) which flattens the image into a single rootfs before extraction. There is no per-layer structure for Plan G's chain installer to consume, so `ENROOT_ZFS_LAYER_CHAIN=y` is silently a no-op for daemon URIs (the daemon path's behavior is unchanged either way).

Bringing chain mode to daemon URIs is feasible but is a real follow-up plan: it requires switching from `${engine} export` to `${engine} save` (which preserves per-layer tarballs in a tar archive plus a `manifest.json` describing layer order), parsing that manifest, and constructing a synthetic `0/` from `${engine} inspect` output. Documented in plan's Out of scope and the `ENROOT_ZFS_LAYER_CHAIN` knob description in `doc/zfs.md`.

sodre added 2 commits May 6, 2026 06:52
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
sodre added a commit that referenced this pull request May 6, 2026
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Plan-doc only — depends on Plan G's helpers, which land in PR #17.
The doc/plans/README.md index and doc/zfs.md knob description will
follow when both Plan G and Plan H land on zenroot/main.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre
sodre merged commit c4baea6 into zenroot/mainMay 6, 2026
sodre added a commit that referenced this pull request May 6, 2026
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Plan-doc only — depends on Plan G's helpers, which land in PR #17.
The doc/plans/README.md index and doc/zfs.md knob description will
follow when both Plan G and Plan H land on zenroot/main.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in)

1 participant

@sodre
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in) - #17

Merged
sodre merged 14 commits into
zenroot/mainfrom
feature/zfs-g-layer-chain
May 6, 2026
Merged

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in)#17
sodre merged 14 commits into
zenroot/mainfrom
feature/zfs-g-layer-chain

Conversation

@sodre

@sodresodre commented May 2, 2026

Copy link
Copy Markdown
Member

Closes#4.

Adds an opt-in per-layer zfs clone chain mode for the Docker template cache. With ENROOT_ZFS_LAYER_CHAIN=y, two images sharing a registry layer digest physically share the bytes on disk; re-pulling an image after a top-layer-only change reuses the cached lower-layer datasets.

Layout

<store>/.layers/<layer-digest> # one per distinct registry layer (origin)
<store>/.layers/<layer-digest>@done # snapshot taken after layer apply
<store>/.templates/<image-config-sha> # zfs clone of the chain leaf @done
<store>/.templates/<image-config-sha>@pristine

Each layer dataset is zfs cloned from the previous layer's @done, with overlayfs whiteouts (mknod 0:0) and opaque-dir markers (trusted.overlay.opaque=y) replayed in shell on top of the cloned target — overlayfs only does that merge at mount time, but a chain stored at-rest needs it baked in. The chain leaf is then cloned into .templates/<config_sha>, the per-image synthetic 0/ config layer (rc/fstab/environment from docker::configure) is applied on top, and the result is snapshotted as @pristine so the existing zfs::clone_container, pointer-format, eviction-recovery, and zfs:// paths all work unchanged.

Why no zfs promote

The issue mentions promote as one option for flattening the chain. We don't promote — promoting inverts the chain (layers become clones of the template), which works for one image but produces a complex image-private topology that defeats the cross-image sharing goal. Plan G keeps layers as immutable origins; ZFS refuses to destroy a layer while any descendant clone exists, so layer GC is automatic once all referencing templates are evicted.

What's added

FileChange
src/storage_zfs.shzfs::layer_chain_active, zfs::_apply_layer_payload, zfs::_build_layer, zfs::_install_layer_chain; chain-mode dispatch in docker_install_from_layers and _pull_and_install_template.
src/docker.sh_prepare_layers side-emits the ordered layer-digest list to ./.layers in its temp cwd; docker::load's ZFS branch reads it back when chain mode is active.
pkg/deb/controlRecommends attr (provides getfattr, required by chain-mode opaque-dir handling).
doc/zfs.md, CLAUDE.mdDocument the new knob, store layout, and dedup semantics.
doc/plans/2026-05-01-zfs-g-layer-chain.mdImplementation plan, mirrors Plans A–F structure.

Coexistence with Plan F

  • ENROOT_ZFS_LAYER_CHAIN= (unset/empty/anything but y): Plan F's single-merge _install_template_from_layers runs unchanged.
  • ENROOT_ZFS_LAYER_CHAIN=y: chain mode. Same dispatch hits both docker::load (direct create) and _pull_and_install_template (used by pointer-format import + eviction recovery).
  • The fast path "template @pristine already exists, reuse it" runs before the chain dispatch — templates produced under either mode are reused under the other without rebuild.

Smoke results (spark-ctrl, Pi 5 / Debian 13 / OpenZFS 2.4.1, 3.75G test pool)

TestResult
Single-layer alpine, chain mode✓ pointer file written, layer dataset created, template clones leaf, rootfs has os-release + /etc/{rc,fstab,environment}
Multi-layer node:20-alpine✓ 3 layer datasets in BASE→TOP order, leaf REFERs full 69.9M merged tree, /usr/local/bin/node (102M binary) present in container rootfs
Multi-layer python:3.13-alpine✓ 3 layer datasets, clones-of-clones visible in zfs list (3070388042c6 1.03M USED / 19.3M REFER — pure dedup)
Whiteout/opaque sanity✓ no .wh.* AUFS files leak through, no char-device whiteouts in final rootfs
Layer reuse after template eviction✓ destroy templates+containers, keep layers; re-create from pointer takes 1.7s with NO "Building layer" messages, just final clone
Plan F regression (flag unset)✓ no .layers/ namespace created, _install_template_from_layers runs as before

Smoke testing also flagged two bugs that were fixed in 3f7e3af:

  • Inverted chain iteration order (docker::_download reverses the manifest, so digests[0] is the TOP, not the BASE).
  • Missing synthetic 0/ config layer apply on the leaf — Plan F's overlay mount stacks 0:1:…:N with 0/ on top; the chain installer needed an explicit final tar-pipe of 0/ onto the template.

Plus one packaging fix: attr is now Recommended (was Suggested), since getfattr is required for chain-mode opaque-dir handling and Suggests is not auto-installed.

🤖 Generated with Claude Code

sodre added 9 commits May 1, 2026 22:17
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
…ants
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
…ll path
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Smoke testing on a 3-layer image (node:20-alpine) caught two bugs in the
chain installer:
1. Inverted iteration. docker::_download reverses the manifest's layer
order via jq's `reverse`, so digests[0] is the TOP layer and
digests[N-1] is the BASE. The original `for i in 0..N-1` loop treated
digests[0] as the base, building the chain upside-down and producing
a leaf that contained only the top-layer's diffs (e.g. 5.4M for what
should have been a 70M merged node:20-alpine rootfs). Iterating from
N-1 down to 0 puts BASE first in the zfs hierarchy and the TOP at
the leaf.
2. Missing synthetic config layer. docker::_prepare_layers populates a
directory 0/ via docker::configure with the per-image
/etc/{rc,fstab,environment} derived from the image config blob; Plan
F's overlay mount stacks 0:1:2:...:N so 0/ ends up on top. The chain
installer ignored 0/ entirely, so containers created via chain mode
were missing /etc/rc and the merged fstab entries. Now applied as a
final tar-pipe step on top of the leaf clone during template
finalization, before snapshotting @pristine.
Also tighten the apply payload:
- getfattr returns non-zero when no files match the requested xattr;
with set -euo pipefail in the payload that aborted the whole apply on
alpine (no opaque dirs). Capture to a temp file with `|| true`.
- Drop tar's --acls. Default ZFS datasets have acltype=off, which makes
POSIX ACL set/get fail with "Operation not supported" warnings even
when the source has no ACLs. Docker images effectively never depend
on ACLs, and xattrs (overlayfs opaque markers, capability bits,
SELinux labels) are still preserved.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre
sodre marked this pull request as ready for review May 2, 2026 03:12
@sodre

sodre commented May 2, 2026

Copy link
Copy Markdown
MemberAuthor

Cross-image dedup verified on disk (acceptance criterion #1 from issue #4):

Pulled python:3.13-alpine3.21 then node:22-alpine3.21 (both pinned to the same alpine base) on a fresh store. Origin tree (zfs list -r -d 1 -o name,origin):

c2fe130f4aab (alpine 3.21 base, 5.11M USED, ORIGIN)
├── 1c6063f559a3 → 838d25d4769a → 3517b1771ef3 (python:3.13-alpine3.21)
└── 720ee653d3d4 → 055ee03d01c9 → 7f3c333e617d (node:22-alpine3.21)

Both chains branch from c2fe130f4aab…@done. The 5.11M alpine base is stored once on disk; ZFS protects it via "snapshot has clones" so it's automatically immortal as long as either python or node is still around.

@sodre

sodre commented May 2, 2026

Copy link
Copy Markdown
MemberAuthor

Two more acceptance criteria covered.

Concurrent pull of the same image — race-safe

ENROOT_ZFS_LAYER_CHAIN=y enroot import -o /tmp/n1.sqsh docker://node:22-alpine3.21 & p1=$!
ENROOT_ZFS_LAYER_CHAIN=y enroot import -o /tmp/n2.sqsh docker://node:22-alpine3.21 & p2=$!wait$p1$p2

Result: both processes returned 0; only one set of [INFO] Building layer ... messages printed (4 messages, one per node:22-alpine3.21 layer); on disk: exactly 4 layer datasets, no .tmp orphans, both containers usable. The per-layer <digest>.tmp lock collapsed both invocations onto a single builder; the loser waited for @done silently.

Lower-layer reuse on a second pull that shares a base

When node:22-alpine3.21 is imported after python:3.13-alpine3.21, the chain installer logged 3 Building layer messages — not 4 — because alpine 3.21's c2fe130f4aab@done was already in the cache from python. The node chain branched directly off that snapshot via zfs clone, so the second pull paid the layer-build cost only for node-specific content.

This generalizes to the issue's "top-layer-only re-pull" case: when a docker tag is republished with only the top digest changed, every cached lower-layer <digest>@done is a hit and _build_layer skips immediately to the clone-on-top step.

Plan G applies only to registry-pulled docker:// URIs. Daemon-local
URIs (dockerd://, podman://) take a separate path that uses
\`${engine} export | tar -x\` (flat rootfs) instead of layer tarballs,
so there is no per-layer structure for the chain installer to consume.
Spell this out in the plan's Coexistence section and the user-facing
knob description in doc/zfs.md, plus add a future-work note in Out of
scope describing what bringing chain mode to daemon URIs would require
(switching to \`${engine} save\` and parsing manifest.json).
The current code already silently no-ops for daemon URIs; this is a
docs-only commit clarifying the boundary.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre

sodre commented May 6, 2026

Copy link
Copy Markdown
MemberAuthor

Scope clarified: `docker://` only.

`dockerd://` and `podman://` URIs take a separate path (`${engine} export | tar -x` → `zfs::_install_template_from_dir`) which flattens the image into a single rootfs before extraction. There is no per-layer structure for Plan G's chain installer to consume, so `ENROOT_ZFS_LAYER_CHAIN=y` is silently a no-op for daemon URIs (the daemon path's behavior is unchanged either way).

Bringing chain mode to daemon URIs is feasible but is a real follow-up plan: it requires switching from `${engine} export` to `${engine} save` (which preserves per-layer tarballs in a tar archive plus a `manifest.json` describing layer order), parsing that manifest, and constructing a synthetic `0/` from `${engine} inspect` output. Documented in plan's Out of scope and the `ENROOT_ZFS_LAYER_CHAIN` knob description in `doc/zfs.md`.

sodre added 2 commits May 6, 2026 06:52
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
sodre added a commit that referenced this pull request May 6, 2026
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Plan-doc only — depends on Plan G's helpers, which land in PR #17.
The doc/plans/README.md index and doc/zfs.md knob description will
follow when both Plan G and Plan H land on zenroot/main.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre
sodre merged commit c4baea6 into zenroot/mainMay 6, 2026
sodre added a commit that referenced this pull request May 6, 2026
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Plan-doc only — depends on Plan G's helpers, which land in PR #17.
The doc/plans/README.md index and doc/zfs.md knob description will
follow when both Plan G and Plan H land on zenroot/main.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in)

1 participant

@sodre
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in) - #17

Merged
sodre merged 14 commits into
zenroot/mainfrom
feature/zfs-g-layer-chain
May 6, 2026
Merged

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in)#17
sodre merged 14 commits into
zenroot/mainfrom
feature/zfs-g-layer-chain

Conversation

@sodre

@sodresodre commented May 2, 2026

Copy link
Copy Markdown
Member

Closes#4.

Adds an opt-in per-layer zfs clone chain mode for the Docker template cache. With ENROOT_ZFS_LAYER_CHAIN=y, two images sharing a registry layer digest physically share the bytes on disk; re-pulling an image after a top-layer-only change reuses the cached lower-layer datasets.

Layout

<store>/.layers/<layer-digest> # one per distinct registry layer (origin)
<store>/.layers/<layer-digest>@done # snapshot taken after layer apply
<store>/.templates/<image-config-sha> # zfs clone of the chain leaf @done
<store>/.templates/<image-config-sha>@pristine

Each layer dataset is zfs cloned from the previous layer's @done, with overlayfs whiteouts (mknod 0:0) and opaque-dir markers (trusted.overlay.opaque=y) replayed in shell on top of the cloned target — overlayfs only does that merge at mount time, but a chain stored at-rest needs it baked in. The chain leaf is then cloned into .templates/<config_sha>, the per-image synthetic 0/ config layer (rc/fstab/environment from docker::configure) is applied on top, and the result is snapshotted as @pristine so the existing zfs::clone_container, pointer-format, eviction-recovery, and zfs:// paths all work unchanged.

Why no zfs promote

The issue mentions promote as one option for flattening the chain. We don't promote — promoting inverts the chain (layers become clones of the template), which works for one image but produces a complex image-private topology that defeats the cross-image sharing goal. Plan G keeps layers as immutable origins; ZFS refuses to destroy a layer while any descendant clone exists, so layer GC is automatic once all referencing templates are evicted.

What's added

FileChange
src/storage_zfs.shzfs::layer_chain_active, zfs::_apply_layer_payload, zfs::_build_layer, zfs::_install_layer_chain; chain-mode dispatch in docker_install_from_layers and _pull_and_install_template.
src/docker.sh_prepare_layers side-emits the ordered layer-digest list to ./.layers in its temp cwd; docker::load's ZFS branch reads it back when chain mode is active.
pkg/deb/controlRecommends attr (provides getfattr, required by chain-mode opaque-dir handling).
doc/zfs.md, CLAUDE.mdDocument the new knob, store layout, and dedup semantics.
doc/plans/2026-05-01-zfs-g-layer-chain.mdImplementation plan, mirrors Plans A–F structure.

Coexistence with Plan F

  • ENROOT_ZFS_LAYER_CHAIN= (unset/empty/anything but y): Plan F's single-merge _install_template_from_layers runs unchanged.
  • ENROOT_ZFS_LAYER_CHAIN=y: chain mode. Same dispatch hits both docker::load (direct create) and _pull_and_install_template (used by pointer-format import + eviction recovery).
  • The fast path "template @pristine already exists, reuse it" runs before the chain dispatch — templates produced under either mode are reused under the other without rebuild.

Smoke results (spark-ctrl, Pi 5 / Debian 13 / OpenZFS 2.4.1, 3.75G test pool)

TestResult
Single-layer alpine, chain mode✓ pointer file written, layer dataset created, template clones leaf, rootfs has os-release + /etc/{rc,fstab,environment}
Multi-layer node:20-alpine✓ 3 layer datasets in BASE→TOP order, leaf REFERs full 69.9M merged tree, /usr/local/bin/node (102M binary) present in container rootfs
Multi-layer python:3.13-alpine✓ 3 layer datasets, clones-of-clones visible in zfs list (3070388042c6 1.03M USED / 19.3M REFER — pure dedup)
Whiteout/opaque sanity✓ no .wh.* AUFS files leak through, no char-device whiteouts in final rootfs
Layer reuse after template eviction✓ destroy templates+containers, keep layers; re-create from pointer takes 1.7s with NO "Building layer" messages, just final clone
Plan F regression (flag unset)✓ no .layers/ namespace created, _install_template_from_layers runs as before

Smoke testing also flagged two bugs that were fixed in 3f7e3af:

  • Inverted chain iteration order (docker::_download reverses the manifest, so digests[0] is the TOP, not the BASE).
  • Missing synthetic 0/ config layer apply on the leaf — Plan F's overlay mount stacks 0:1:…:N with 0/ on top; the chain installer needed an explicit final tar-pipe of 0/ onto the template.

Plus one packaging fix: attr is now Recommended (was Suggested), since getfattr is required for chain-mode opaque-dir handling and Suggests is not auto-installed.

🤖 Generated with Claude Code

sodre added 9 commits May 1, 2026 22:17
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
…ants
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
…ll path
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Smoke testing on a 3-layer image (node:20-alpine) caught two bugs in the
chain installer:
1. Inverted iteration. docker::_download reverses the manifest's layer
order via jq's `reverse`, so digests[0] is the TOP layer and
digests[N-1] is the BASE. The original `for i in 0..N-1` loop treated
digests[0] as the base, building the chain upside-down and producing
a leaf that contained only the top-layer's diffs (e.g. 5.4M for what
should have been a 70M merged node:20-alpine rootfs). Iterating from
N-1 down to 0 puts BASE first in the zfs hierarchy and the TOP at
the leaf.
2. Missing synthetic config layer. docker::_prepare_layers populates a
directory 0/ via docker::configure with the per-image
/etc/{rc,fstab,environment} derived from the image config blob; Plan
F's overlay mount stacks 0:1:2:...:N so 0/ ends up on top. The chain
installer ignored 0/ entirely, so containers created via chain mode
were missing /etc/rc and the merged fstab entries. Now applied as a
final tar-pipe step on top of the leaf clone during template
finalization, before snapshotting @pristine.
Also tighten the apply payload:
- getfattr returns non-zero when no files match the requested xattr;
with set -euo pipefail in the payload that aborted the whole apply on
alpine (no opaque dirs). Capture to a temp file with `|| true`.
- Drop tar's --acls. Default ZFS datasets have acltype=off, which makes
POSIX ACL set/get fail with "Operation not supported" warnings even
when the source has no ACLs. Docker images effectively never depend
on ACLs, and xattrs (overlayfs opaque markers, capability bits,
SELinux labels) are still preserved.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre
sodre marked this pull request as ready for review May 2, 2026 03:12
@sodre

sodre commented May 2, 2026

Copy link
Copy Markdown
MemberAuthor

Cross-image dedup verified on disk (acceptance criterion #1 from issue #4):

Pulled python:3.13-alpine3.21 then node:22-alpine3.21 (both pinned to the same alpine base) on a fresh store. Origin tree (zfs list -r -d 1 -o name,origin):

c2fe130f4aab (alpine 3.21 base, 5.11M USED, ORIGIN)
├── 1c6063f559a3 → 838d25d4769a → 3517b1771ef3 (python:3.13-alpine3.21)
└── 720ee653d3d4 → 055ee03d01c9 → 7f3c333e617d (node:22-alpine3.21)

Both chains branch from c2fe130f4aab…@done. The 5.11M alpine base is stored once on disk; ZFS protects it via "snapshot has clones" so it's automatically immortal as long as either python or node is still around.

@sodre

sodre commented May 2, 2026

Copy link
Copy Markdown
MemberAuthor

Two more acceptance criteria covered.

Concurrent pull of the same image — race-safe

ENROOT_ZFS_LAYER_CHAIN=y enroot import -o /tmp/n1.sqsh docker://node:22-alpine3.21 & p1=$!
ENROOT_ZFS_LAYER_CHAIN=y enroot import -o /tmp/n2.sqsh docker://node:22-alpine3.21 & p2=$!wait$p1$p2

Result: both processes returned 0; only one set of [INFO] Building layer ... messages printed (4 messages, one per node:22-alpine3.21 layer); on disk: exactly 4 layer datasets, no .tmp orphans, both containers usable. The per-layer <digest>.tmp lock collapsed both invocations onto a single builder; the loser waited for @done silently.

Lower-layer reuse on a second pull that shares a base

When node:22-alpine3.21 is imported after python:3.13-alpine3.21, the chain installer logged 3 Building layer messages — not 4 — because alpine 3.21's c2fe130f4aab@done was already in the cache from python. The node chain branched directly off that snapshot via zfs clone, so the second pull paid the layer-build cost only for node-specific content.

This generalizes to the issue's "top-layer-only re-pull" case: when a docker tag is republished with only the top digest changed, every cached lower-layer <digest>@done is a hit and _build_layer skips immediately to the clone-on-top step.

Plan G applies only to registry-pulled docker:// URIs. Daemon-local
URIs (dockerd://, podman://) take a separate path that uses
\`${engine} export | tar -x\` (flat rootfs) instead of layer tarballs,
so there is no per-layer structure for the chain installer to consume.
Spell this out in the plan's Coexistence section and the user-facing
knob description in doc/zfs.md, plus add a future-work note in Out of
scope describing what bringing chain mode to daemon URIs would require
(switching to \`${engine} save\` and parsing manifest.json).
The current code already silently no-ops for daemon URIs; this is a
docs-only commit clarifying the boundary.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre

sodre commented May 6, 2026

Copy link
Copy Markdown
MemberAuthor

Scope clarified: `docker://` only.

`dockerd://` and `podman://` URIs take a separate path (`${engine} export | tar -x` → `zfs::_install_template_from_dir`) which flattens the image into a single rootfs before extraction. There is no per-layer structure for Plan G's chain installer to consume, so `ENROOT_ZFS_LAYER_CHAIN=y` is silently a no-op for daemon URIs (the daemon path's behavior is unchanged either way).

Bringing chain mode to daemon URIs is feasible but is a real follow-up plan: it requires switching from `${engine} export` to `${engine} save` (which preserves per-layer tarballs in a tar archive plus a `manifest.json` describing layer order), parsing that manifest, and constructing a synthetic `0/` from `${engine} inspect` output. Documented in plan's Out of scope and the `ENROOT_ZFS_LAYER_CHAIN` knob description in `doc/zfs.md`.

sodre added 2 commits May 6, 2026 06:52
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
sodre added a commit that referenced this pull request May 6, 2026
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Plan-doc only — depends on Plan G's helpers, which land in PR #17.
The doc/plans/README.md index and doc/zfs.md knob description will
follow when both Plan G and Plan H land on zenroot/main.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre
sodre merged commit c4baea6 into zenroot/mainMay 6, 2026
sodre added a commit that referenced this pull request May 6, 2026
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Plan-doc only — depends on Plan G's helpers, which land in PR #17.
The doc/plans/README.md index and doc/zfs.md knob description will
follow when both Plan G and Plan H land on zenroot/main.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in)

1 participant

@sodre
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in) - #17

Merged
sodre merged 14 commits into
zenroot/mainfrom
feature/zfs-g-layer-chain
May 6, 2026
Merged

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in)#17
sodre merged 14 commits into
zenroot/mainfrom
feature/zfs-g-layer-chain

Conversation

@sodre

@sodresodre commented May 2, 2026

Copy link
Copy Markdown
Member

Closes#4.

Adds an opt-in per-layer zfs clone chain mode for the Docker template cache. With ENROOT_ZFS_LAYER_CHAIN=y, two images sharing a registry layer digest physically share the bytes on disk; re-pulling an image after a top-layer-only change reuses the cached lower-layer datasets.

Layout

<store>/.layers/<layer-digest> # one per distinct registry layer (origin)
<store>/.layers/<layer-digest>@done # snapshot taken after layer apply
<store>/.templates/<image-config-sha> # zfs clone of the chain leaf @done
<store>/.templates/<image-config-sha>@pristine

Each layer dataset is zfs cloned from the previous layer's @done, with overlayfs whiteouts (mknod 0:0) and opaque-dir markers (trusted.overlay.opaque=y) replayed in shell on top of the cloned target — overlayfs only does that merge at mount time, but a chain stored at-rest needs it baked in. The chain leaf is then cloned into .templates/<config_sha>, the per-image synthetic 0/ config layer (rc/fstab/environment from docker::configure) is applied on top, and the result is snapshotted as @pristine so the existing zfs::clone_container, pointer-format, eviction-recovery, and zfs:// paths all work unchanged.

Why no zfs promote

The issue mentions promote as one option for flattening the chain. We don't promote — promoting inverts the chain (layers become clones of the template), which works for one image but produces a complex image-private topology that defeats the cross-image sharing goal. Plan G keeps layers as immutable origins; ZFS refuses to destroy a layer while any descendant clone exists, so layer GC is automatic once all referencing templates are evicted.

What's added

FileChange
src/storage_zfs.shzfs::layer_chain_active, zfs::_apply_layer_payload, zfs::_build_layer, zfs::_install_layer_chain; chain-mode dispatch in docker_install_from_layers and _pull_and_install_template.
src/docker.sh_prepare_layers side-emits the ordered layer-digest list to ./.layers in its temp cwd; docker::load's ZFS branch reads it back when chain mode is active.
pkg/deb/controlRecommends attr (provides getfattr, required by chain-mode opaque-dir handling).
doc/zfs.md, CLAUDE.mdDocument the new knob, store layout, and dedup semantics.
doc/plans/2026-05-01-zfs-g-layer-chain.mdImplementation plan, mirrors Plans A–F structure.

Coexistence with Plan F

  • ENROOT_ZFS_LAYER_CHAIN= (unset/empty/anything but y): Plan F's single-merge _install_template_from_layers runs unchanged.
  • ENROOT_ZFS_LAYER_CHAIN=y: chain mode. Same dispatch hits both docker::load (direct create) and _pull_and_install_template (used by pointer-format import + eviction recovery).
  • The fast path "template @pristine already exists, reuse it" runs before the chain dispatch — templates produced under either mode are reused under the other without rebuild.

Smoke results (spark-ctrl, Pi 5 / Debian 13 / OpenZFS 2.4.1, 3.75G test pool)

TestResult
Single-layer alpine, chain mode✓ pointer file written, layer dataset created, template clones leaf, rootfs has os-release + /etc/{rc,fstab,environment}
Multi-layer node:20-alpine✓ 3 layer datasets in BASE→TOP order, leaf REFERs full 69.9M merged tree, /usr/local/bin/node (102M binary) present in container rootfs
Multi-layer python:3.13-alpine✓ 3 layer datasets, clones-of-clones visible in zfs list (3070388042c6 1.03M USED / 19.3M REFER — pure dedup)
Whiteout/opaque sanity✓ no .wh.* AUFS files leak through, no char-device whiteouts in final rootfs
Layer reuse after template eviction✓ destroy templates+containers, keep layers; re-create from pointer takes 1.7s with NO "Building layer" messages, just final clone
Plan F regression (flag unset)✓ no .layers/ namespace created, _install_template_from_layers runs as before

Smoke testing also flagged two bugs that were fixed in 3f7e3af:

  • Inverted chain iteration order (docker::_download reverses the manifest, so digests[0] is the TOP, not the BASE).
  • Missing synthetic 0/ config layer apply on the leaf — Plan F's overlay mount stacks 0:1:…:N with 0/ on top; the chain installer needed an explicit final tar-pipe of 0/ onto the template.

Plus one packaging fix: attr is now Recommended (was Suggested), since getfattr is required for chain-mode opaque-dir handling and Suggests is not auto-installed.

🤖 Generated with Claude Code

sodre added 9 commits May 1, 2026 22:17
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
…ants
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
…ll path
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Smoke testing on a 3-layer image (node:20-alpine) caught two bugs in the
chain installer:
1. Inverted iteration. docker::_download reverses the manifest's layer
order via jq's `reverse`, so digests[0] is the TOP layer and
digests[N-1] is the BASE. The original `for i in 0..N-1` loop treated
digests[0] as the base, building the chain upside-down and producing
a leaf that contained only the top-layer's diffs (e.g. 5.4M for what
should have been a 70M merged node:20-alpine rootfs). Iterating from
N-1 down to 0 puts BASE first in the zfs hierarchy and the TOP at
the leaf.
2. Missing synthetic config layer. docker::_prepare_layers populates a
directory 0/ via docker::configure with the per-image
/etc/{rc,fstab,environment} derived from the image config blob; Plan
F's overlay mount stacks 0:1:2:...:N so 0/ ends up on top. The chain
installer ignored 0/ entirely, so containers created via chain mode
were missing /etc/rc and the merged fstab entries. Now applied as a
final tar-pipe step on top of the leaf clone during template
finalization, before snapshotting @pristine.
Also tighten the apply payload:
- getfattr returns non-zero when no files match the requested xattr;
with set -euo pipefail in the payload that aborted the whole apply on
alpine (no opaque dirs). Capture to a temp file with `|| true`.
- Drop tar's --acls. Default ZFS datasets have acltype=off, which makes
POSIX ACL set/get fail with "Operation not supported" warnings even
when the source has no ACLs. Docker images effectively never depend
on ACLs, and xattrs (overlayfs opaque markers, capability bits,
SELinux labels) are still preserved.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre
sodre marked this pull request as ready for review May 2, 2026 03:12
@sodre

sodre commented May 2, 2026

Copy link
Copy Markdown
MemberAuthor

Cross-image dedup verified on disk (acceptance criterion #1 from issue #4):

Pulled python:3.13-alpine3.21 then node:22-alpine3.21 (both pinned to the same alpine base) on a fresh store. Origin tree (zfs list -r -d 1 -o name,origin):

c2fe130f4aab (alpine 3.21 base, 5.11M USED, ORIGIN)
├── 1c6063f559a3 → 838d25d4769a → 3517b1771ef3 (python:3.13-alpine3.21)
└── 720ee653d3d4 → 055ee03d01c9 → 7f3c333e617d (node:22-alpine3.21)

Both chains branch from c2fe130f4aab…@done. The 5.11M alpine base is stored once on disk; ZFS protects it via "snapshot has clones" so it's automatically immortal as long as either python or node is still around.

@sodre

sodre commented May 2, 2026

Copy link
Copy Markdown
MemberAuthor

Two more acceptance criteria covered.

Concurrent pull of the same image — race-safe

ENROOT_ZFS_LAYER_CHAIN=y enroot import -o /tmp/n1.sqsh docker://node:22-alpine3.21 & p1=$!
ENROOT_ZFS_LAYER_CHAIN=y enroot import -o /tmp/n2.sqsh docker://node:22-alpine3.21 & p2=$!wait$p1$p2

Result: both processes returned 0; only one set of [INFO] Building layer ... messages printed (4 messages, one per node:22-alpine3.21 layer); on disk: exactly 4 layer datasets, no .tmp orphans, both containers usable. The per-layer <digest>.tmp lock collapsed both invocations onto a single builder; the loser waited for @done silently.

Lower-layer reuse on a second pull that shares a base

When node:22-alpine3.21 is imported after python:3.13-alpine3.21, the chain installer logged 3 Building layer messages — not 4 — because alpine 3.21's c2fe130f4aab@done was already in the cache from python. The node chain branched directly off that snapshot via zfs clone, so the second pull paid the layer-build cost only for node-specific content.

This generalizes to the issue's "top-layer-only re-pull" case: when a docker tag is republished with only the top digest changed, every cached lower-layer <digest>@done is a hit and _build_layer skips immediately to the clone-on-top step.

Plan G applies only to registry-pulled docker:// URIs. Daemon-local
URIs (dockerd://, podman://) take a separate path that uses
\`${engine} export | tar -x\` (flat rootfs) instead of layer tarballs,
so there is no per-layer structure for the chain installer to consume.
Spell this out in the plan's Coexistence section and the user-facing
knob description in doc/zfs.md, plus add a future-work note in Out of
scope describing what bringing chain mode to daemon URIs would require
(switching to \`${engine} save\` and parsing manifest.json).
The current code already silently no-ops for daemon URIs; this is a
docs-only commit clarifying the boundary.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre

sodre commented May 6, 2026

Copy link
Copy Markdown
MemberAuthor

Scope clarified: `docker://` only.

`dockerd://` and `podman://` URIs take a separate path (`${engine} export | tar -x` → `zfs::_install_template_from_dir`) which flattens the image into a single rootfs before extraction. There is no per-layer structure for Plan G's chain installer to consume, so `ENROOT_ZFS_LAYER_CHAIN=y` is silently a no-op for daemon URIs (the daemon path's behavior is unchanged either way).

Bringing chain mode to daemon URIs is feasible but is a real follow-up plan: it requires switching from `${engine} export` to `${engine} save` (which preserves per-layer tarballs in a tar archive plus a `manifest.json` describing layer order), parsing that manifest, and constructing a synthetic `0/` from `${engine} inspect` output. Documented in plan's Out of scope and the `ENROOT_ZFS_LAYER_CHAIN` knob description in `doc/zfs.md`.

sodre added 2 commits May 6, 2026 06:52
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
sodre added a commit that referenced this pull request May 6, 2026
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Plan-doc only — depends on Plan G's helpers, which land in PR #17.
The doc/plans/README.md index and doc/zfs.md knob description will
follow when both Plan G and Plan H land on zenroot/main.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre
sodre merged commit c4baea6 into zenroot/mainMay 6, 2026
sodre added a commit that referenced this pull request May 6, 2026
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Plan-doc only — depends on Plan G's helpers, which land in PR #17.
The doc/plans/README.md index and doc/zfs.md knob description will
follow when both Plan G and Plan H land on zenroot/main.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in)

1 participant

@sodre
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in) - #17

Merged
sodre merged 14 commits into
zenroot/mainfrom
feature/zfs-g-layer-chain
May 6, 2026
Merged

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in)#17
sodre merged 14 commits into
zenroot/mainfrom
feature/zfs-g-layer-chain

Conversation

@sodre

@sodresodre commented May 2, 2026

Copy link
Copy Markdown
Member

Closes#4.

Adds an opt-in per-layer zfs clone chain mode for the Docker template cache. With ENROOT_ZFS_LAYER_CHAIN=y, two images sharing a registry layer digest physically share the bytes on disk; re-pulling an image after a top-layer-only change reuses the cached lower-layer datasets.

Layout

<store>/.layers/<layer-digest> # one per distinct registry layer (origin)
<store>/.layers/<layer-digest>@done # snapshot taken after layer apply
<store>/.templates/<image-config-sha> # zfs clone of the chain leaf @done
<store>/.templates/<image-config-sha>@pristine

Each layer dataset is zfs cloned from the previous layer's @done, with overlayfs whiteouts (mknod 0:0) and opaque-dir markers (trusted.overlay.opaque=y) replayed in shell on top of the cloned target — overlayfs only does that merge at mount time, but a chain stored at-rest needs it baked in. The chain leaf is then cloned into .templates/<config_sha>, the per-image synthetic 0/ config layer (rc/fstab/environment from docker::configure) is applied on top, and the result is snapshotted as @pristine so the existing zfs::clone_container, pointer-format, eviction-recovery, and zfs:// paths all work unchanged.

Why no zfs promote

The issue mentions promote as one option for flattening the chain. We don't promote — promoting inverts the chain (layers become clones of the template), which works for one image but produces a complex image-private topology that defeats the cross-image sharing goal. Plan G keeps layers as immutable origins; ZFS refuses to destroy a layer while any descendant clone exists, so layer GC is automatic once all referencing templates are evicted.

What's added

FileChange
src/storage_zfs.shzfs::layer_chain_active, zfs::_apply_layer_payload, zfs::_build_layer, zfs::_install_layer_chain; chain-mode dispatch in docker_install_from_layers and _pull_and_install_template.
src/docker.sh_prepare_layers side-emits the ordered layer-digest list to ./.layers in its temp cwd; docker::load's ZFS branch reads it back when chain mode is active.
pkg/deb/controlRecommends attr (provides getfattr, required by chain-mode opaque-dir handling).
doc/zfs.md, CLAUDE.mdDocument the new knob, store layout, and dedup semantics.
doc/plans/2026-05-01-zfs-g-layer-chain.mdImplementation plan, mirrors Plans A–F structure.

Coexistence with Plan F

  • ENROOT_ZFS_LAYER_CHAIN= (unset/empty/anything but y): Plan F's single-merge _install_template_from_layers runs unchanged.
  • ENROOT_ZFS_LAYER_CHAIN=y: chain mode. Same dispatch hits both docker::load (direct create) and _pull_and_install_template (used by pointer-format import + eviction recovery).
  • The fast path "template @pristine already exists, reuse it" runs before the chain dispatch — templates produced under either mode are reused under the other without rebuild.

Smoke results (spark-ctrl, Pi 5 / Debian 13 / OpenZFS 2.4.1, 3.75G test pool)

TestResult
Single-layer alpine, chain mode✓ pointer file written, layer dataset created, template clones leaf, rootfs has os-release + /etc/{rc,fstab,environment}
Multi-layer node:20-alpine✓ 3 layer datasets in BASE→TOP order, leaf REFERs full 69.9M merged tree, /usr/local/bin/node (102M binary) present in container rootfs
Multi-layer python:3.13-alpine✓ 3 layer datasets, clones-of-clones visible in zfs list (3070388042c6 1.03M USED / 19.3M REFER — pure dedup)
Whiteout/opaque sanity✓ no .wh.* AUFS files leak through, no char-device whiteouts in final rootfs
Layer reuse after template eviction✓ destroy templates+containers, keep layers; re-create from pointer takes 1.7s with NO "Building layer" messages, just final clone
Plan F regression (flag unset)✓ no .layers/ namespace created, _install_template_from_layers runs as before

Smoke testing also flagged two bugs that were fixed in 3f7e3af:

  • Inverted chain iteration order (docker::_download reverses the manifest, so digests[0] is the TOP, not the BASE).
  • Missing synthetic 0/ config layer apply on the leaf — Plan F's overlay mount stacks 0:1:…:N with 0/ on top; the chain installer needed an explicit final tar-pipe of 0/ onto the template.

Plus one packaging fix: attr is now Recommended (was Suggested), since getfattr is required for chain-mode opaque-dir handling and Suggests is not auto-installed.

🤖 Generated with Claude Code

sodre added 9 commits May 1, 2026 22:17
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
…ants
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
…ll path
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Smoke testing on a 3-layer image (node:20-alpine) caught two bugs in the
chain installer:
1. Inverted iteration. docker::_download reverses the manifest's layer
order via jq's `reverse`, so digests[0] is the TOP layer and
digests[N-1] is the BASE. The original `for i in 0..N-1` loop treated
digests[0] as the base, building the chain upside-down and producing
a leaf that contained only the top-layer's diffs (e.g. 5.4M for what
should have been a 70M merged node:20-alpine rootfs). Iterating from
N-1 down to 0 puts BASE first in the zfs hierarchy and the TOP at
the leaf.
2. Missing synthetic config layer. docker::_prepare_layers populates a
directory 0/ via docker::configure with the per-image
/etc/{rc,fstab,environment} derived from the image config blob; Plan
F's overlay mount stacks 0:1:2:...:N so 0/ ends up on top. The chain
installer ignored 0/ entirely, so containers created via chain mode
were missing /etc/rc and the merged fstab entries. Now applied as a
final tar-pipe step on top of the leaf clone during template
finalization, before snapshotting @pristine.
Also tighten the apply payload:
- getfattr returns non-zero when no files match the requested xattr;
with set -euo pipefail in the payload that aborted the whole apply on
alpine (no opaque dirs). Capture to a temp file with `|| true`.
- Drop tar's --acls. Default ZFS datasets have acltype=off, which makes
POSIX ACL set/get fail with "Operation not supported" warnings even
when the source has no ACLs. Docker images effectively never depend
on ACLs, and xattrs (overlayfs opaque markers, capability bits,
SELinux labels) are still preserved.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre
sodre marked this pull request as ready for review May 2, 2026 03:12
@sodre

sodre commented May 2, 2026

Copy link
Copy Markdown
MemberAuthor

Cross-image dedup verified on disk (acceptance criterion #1 from issue #4):

Pulled python:3.13-alpine3.21 then node:22-alpine3.21 (both pinned to the same alpine base) on a fresh store. Origin tree (zfs list -r -d 1 -o name,origin):

c2fe130f4aab (alpine 3.21 base, 5.11M USED, ORIGIN)
├── 1c6063f559a3 → 838d25d4769a → 3517b1771ef3 (python:3.13-alpine3.21)
└── 720ee653d3d4 → 055ee03d01c9 → 7f3c333e617d (node:22-alpine3.21)

Both chains branch from c2fe130f4aab…@done. The 5.11M alpine base is stored once on disk; ZFS protects it via "snapshot has clones" so it's automatically immortal as long as either python or node is still around.

@sodre

sodre commented May 2, 2026

Copy link
Copy Markdown
MemberAuthor

Two more acceptance criteria covered.

Concurrent pull of the same image — race-safe

ENROOT_ZFS_LAYER_CHAIN=y enroot import -o /tmp/n1.sqsh docker://node:22-alpine3.21 & p1=$!
ENROOT_ZFS_LAYER_CHAIN=y enroot import -o /tmp/n2.sqsh docker://node:22-alpine3.21 & p2=$!wait$p1$p2

Result: both processes returned 0; only one set of [INFO] Building layer ... messages printed (4 messages, one per node:22-alpine3.21 layer); on disk: exactly 4 layer datasets, no .tmp orphans, both containers usable. The per-layer <digest>.tmp lock collapsed both invocations onto a single builder; the loser waited for @done silently.

Lower-layer reuse on a second pull that shares a base

When node:22-alpine3.21 is imported after python:3.13-alpine3.21, the chain installer logged 3 Building layer messages — not 4 — because alpine 3.21's c2fe130f4aab@done was already in the cache from python. The node chain branched directly off that snapshot via zfs clone, so the second pull paid the layer-build cost only for node-specific content.

This generalizes to the issue's "top-layer-only re-pull" case: when a docker tag is republished with only the top digest changed, every cached lower-layer <digest>@done is a hit and _build_layer skips immediately to the clone-on-top step.

Plan G applies only to registry-pulled docker:// URIs. Daemon-local
URIs (dockerd://, podman://) take a separate path that uses
\`${engine} export | tar -x\` (flat rootfs) instead of layer tarballs,
so there is no per-layer structure for the chain installer to consume.
Spell this out in the plan's Coexistence section and the user-facing
knob description in doc/zfs.md, plus add a future-work note in Out of
scope describing what bringing chain mode to daemon URIs would require
(switching to \`${engine} save\` and parsing manifest.json).
The current code already silently no-ops for daemon URIs; this is a
docs-only commit clarifying the boundary.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre

sodre commented May 6, 2026

Copy link
Copy Markdown
MemberAuthor

Scope clarified: `docker://` only.

`dockerd://` and `podman://` URIs take a separate path (`${engine} export | tar -x` → `zfs::_install_template_from_dir`) which flattens the image into a single rootfs before extraction. There is no per-layer structure for Plan G's chain installer to consume, so `ENROOT_ZFS_LAYER_CHAIN=y` is silently a no-op for daemon URIs (the daemon path's behavior is unchanged either way).

Bringing chain mode to daemon URIs is feasible but is a real follow-up plan: it requires switching from `${engine} export` to `${engine} save` (which preserves per-layer tarballs in a tar archive plus a `manifest.json` describing layer order), parsing that manifest, and constructing a synthetic `0/` from `${engine} inspect` output. Documented in plan's Out of scope and the `ENROOT_ZFS_LAYER_CHAIN` knob description in `doc/zfs.md`.

sodre added 2 commits May 6, 2026 06:52
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
sodre added a commit that referenced this pull request May 6, 2026
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Plan-doc only — depends on Plan G's helpers, which land in PR #17.
The doc/plans/README.md index and doc/zfs.md knob description will
follow when both Plan G and Plan H land on zenroot/main.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
@sodre
sodre merged commit c4baea6 into zenroot/mainMay 6, 2026
sodre added a commit that referenced this pull request May 6, 2026
Plan H extends Plan G's per-layer zfs clone chain to dockerd:// and
podman:// URIs via \`${engine} save\` (preserves per-layer tarballs +
manifest.json describing layer order) instead of \`${engine} export\`
(flattens). Layer digests are content-addressed via sha256 of each
layer.tar, so the same image won't share .layers/ datasets across
docker:// and dockerd:// sources (registry blobs are compressed; save's
layer.tar is uncompressed — same content, different sha) but multiple
local daemon images with shared base layers DO dedup at the .layers/
level.
Default-off — flag unset leaves the existing flat-export daemon path
byte-for-byte unchanged. Same dispatch shape as Plan G (chain-mode
gate on ENROOT_ZFS_LAYER_CHAIN=y, branched in import_daemon_pointer
and create_from_pointer's recovery arm).
Plan-doc only — depends on Plan G's helpers, which land in PR #17.
The doc/plans/README.md index and doc/zfs.md knob description will
follow when both Plan G and Plan H land on zenroot/main.
Signed-off-by: Patrick Sodré <patrick@zero-ae.com>
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Plan G: per-layer ZFS clone chain for enroot load docker:// (opt-in)

1 participant

@sodre