Let a filesystem operation name a container, and report the bytes a trim discarded - #859

Closed
MayCXC wants to merge 1 commit into
apple:mainfrom
MayCXC:trim-container-rootfs
Closed

Let a filesystem operation name a container, and report the bytes a trim discarded#859
MayCXC wants to merge 1 commit into
apple:mainfrom
MayCXC:trim-container-rootfs

Conversation

@MayCXC

Copy link
Copy Markdown

Summary

A filesystem operation takes a path in the init namespace, and a container's root filesystem is not there: the container mounts it in a namespace of its own, and after the pivot nothing in the init namespace names it. The operations that matter for a root filesystem, trim above all, were out of reach for exactly the filesystems that grow.

A request that names a container reaches its root filesystem through the init process's proc entry: the kernel resolves /proc/<pid>/root through that process's namespace, and the trailing dot component keeps the opened target a directory rather than the magic link itself, which the handler's O_NOFOLLOW refuses. The agent surface takes the container id and answers with the bytes the filesystem reported trimmed, for a single container and for a pod's members alike.

Closes#858. Addresses the requirement described in apple/container#2105.

Relationship to #838

This supersedes #838, which asks for the same thing from issue apple/container#2105. Both add containerID to FilesystemOperationRequest so the operation runs against the intended container. The differences:

  • FiTrimResult.trimmed_bytes is populated here. The field arrived with the trim operation in [vminitd]: api for trim filesystem operations #700 and has never been set: trimFilesystem(fd:) throws away the count that FITRIM writes back into the range, and filesystemOperation returns a bare .init(), so the oneof is never set and callers decode the default. [vminitd]: update filesystem operation to run in new namespace #838 keeps return .init(), so the field stays dead after it. Reporting is the whole reason the field exists, and a caller that loops until a trim returns nothing cannot work without it.
  • containerID stays optional.[vminitd]: update filesystem operation to run in new namespace #838 rejects a request that does not carry one, which breaks the existing path-addressed callers. Here a request names either a path or a container, so existing callers keep working.
  • The host side gains the operation it needs.LinuxContainer.trimRootfs() and LinuxPod.trimContainer(_:) return the bytes discarded, rather than leaving each caller to assemble the request.
  • There is a test for the count.testContainerTrimReportsBytes boots a container, trims, and asserts the answer is not zero, which is what fails today.

One thing #838 does that this does not: it enters the container's mount namespace with setns(CLONE_NEWNS) on a dedicated thread after unshare(CLONE_FS). That is a correct way to do it, and it generalises to any path inside the container. Resolving through /proc/<pid>/root reaches the same filesystem through the same namespace without moving a thread between namespaces. If maintainers prefer the setns approach, the reporting fix and the host-side API here apply on top of it unchanged.

Testing

  • swift build and make check clean.
  • swift test: 603 tests in 83 suites passed.
  • make vminitd builds the guest with the change.
  • Integration: ./bin/containerization-integration --filter trim runs container trim reports bytes and the existing container trim ext4 clone, both passing. The new test fails before the reporting fix with trim reported 0 bytes.

A filesystem operation takes a path in the init namespace, and a
container's root filesystem is not there: the container mounts it in a
namespace of its own, and after the pivot nothing in the init namespace
names it. The operations that matter for a root filesystem, trim above
all, were out of reach for exactly the filesystems that grow.
A request that names a container reaches its root filesystem through
the init process's proc entry: the kernel resolves /proc/<pid>/root
through that process's namespace, and the trailing dot component keeps
the opened target a directory rather than the magic link itself, which
the handler's O_NOFOLLOW refuses. The agent surface takes the container
id and returns the bytes the filesystem reported trimmed, from the
machine's one agent, for a single container and for a pod's members
alike.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: A trim never reports the bytes it discarded

2 participants

@MayCXC@crosbymichael
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Let a filesystem operation name a container, and report the bytes a trim discarded - #859

Closed
MayCXC wants to merge 1 commit into
apple:mainfrom
MayCXC:trim-container-rootfs
Closed

Let a filesystem operation name a container, and report the bytes a trim discarded#859
MayCXC wants to merge 1 commit into
apple:mainfrom
MayCXC:trim-container-rootfs

Conversation

@MayCXC

Copy link
Copy Markdown

Summary

A filesystem operation takes a path in the init namespace, and a container's root filesystem is not there: the container mounts it in a namespace of its own, and after the pivot nothing in the init namespace names it. The operations that matter for a root filesystem, trim above all, were out of reach for exactly the filesystems that grow.

A request that names a container reaches its root filesystem through the init process's proc entry: the kernel resolves /proc/<pid>/root through that process's namespace, and the trailing dot component keeps the opened target a directory rather than the magic link itself, which the handler's O_NOFOLLOW refuses. The agent surface takes the container id and answers with the bytes the filesystem reported trimmed, for a single container and for a pod's members alike.

Closes#858. Addresses the requirement described in apple/container#2105.

Relationship to #838

This supersedes #838, which asks for the same thing from issue apple/container#2105. Both add containerID to FilesystemOperationRequest so the operation runs against the intended container. The differences:

  • FiTrimResult.trimmed_bytes is populated here. The field arrived with the trim operation in [vminitd]: api for trim filesystem operations #700 and has never been set: trimFilesystem(fd:) throws away the count that FITRIM writes back into the range, and filesystemOperation returns a bare .init(), so the oneof is never set and callers decode the default. [vminitd]: update filesystem operation to run in new namespace #838 keeps return .init(), so the field stays dead after it. Reporting is the whole reason the field exists, and a caller that loops until a trim returns nothing cannot work without it.
  • containerID stays optional.[vminitd]: update filesystem operation to run in new namespace #838 rejects a request that does not carry one, which breaks the existing path-addressed callers. Here a request names either a path or a container, so existing callers keep working.
  • The host side gains the operation it needs.LinuxContainer.trimRootfs() and LinuxPod.trimContainer(_:) return the bytes discarded, rather than leaving each caller to assemble the request.
  • There is a test for the count.testContainerTrimReportsBytes boots a container, trims, and asserts the answer is not zero, which is what fails today.

One thing #838 does that this does not: it enters the container's mount namespace with setns(CLONE_NEWNS) on a dedicated thread after unshare(CLONE_FS). That is a correct way to do it, and it generalises to any path inside the container. Resolving through /proc/<pid>/root reaches the same filesystem through the same namespace without moving a thread between namespaces. If maintainers prefer the setns approach, the reporting fix and the host-side API here apply on top of it unchanged.

Testing

  • swift build and make check clean.
  • swift test: 603 tests in 83 suites passed.
  • make vminitd builds the guest with the change.
  • Integration: ./bin/containerization-integration --filter trim runs container trim reports bytes and the existing container trim ext4 clone, both passing. The new test fails before the reporting fix with trim reported 0 bytes.

A filesystem operation takes a path in the init namespace, and a
container's root filesystem is not there: the container mounts it in a
namespace of its own, and after the pivot nothing in the init namespace
names it. The operations that matter for a root filesystem, trim above
all, were out of reach for exactly the filesystems that grow.
A request that names a container reaches its root filesystem through
the init process's proc entry: the kernel resolves /proc/<pid>/root
through that process's namespace, and the trailing dot component keeps
the opened target a directory rather than the magic link itself, which
the handler's O_NOFOLLOW refuses. The agent surface takes the container
id and returns the bytes the filesystem reported trimmed, from the
machine's one agent, for a single container and for a pod's members
alike.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: A trim never reports the bytes it discarded

2 participants

@MayCXC@crosbymichael
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Let a filesystem operation name a container, and report the bytes a trim discarded - #859

Closed
MayCXC wants to merge 1 commit into
apple:mainfrom
MayCXC:trim-container-rootfs
Closed

Let a filesystem operation name a container, and report the bytes a trim discarded#859
MayCXC wants to merge 1 commit into
apple:mainfrom
MayCXC:trim-container-rootfs

Conversation

@MayCXC

Copy link
Copy Markdown

Summary

A filesystem operation takes a path in the init namespace, and a container's root filesystem is not there: the container mounts it in a namespace of its own, and after the pivot nothing in the init namespace names it. The operations that matter for a root filesystem, trim above all, were out of reach for exactly the filesystems that grow.

A request that names a container reaches its root filesystem through the init process's proc entry: the kernel resolves /proc/<pid>/root through that process's namespace, and the trailing dot component keeps the opened target a directory rather than the magic link itself, which the handler's O_NOFOLLOW refuses. The agent surface takes the container id and answers with the bytes the filesystem reported trimmed, for a single container and for a pod's members alike.

Closes#858. Addresses the requirement described in apple/container#2105.

Relationship to #838

This supersedes #838, which asks for the same thing from issue apple/container#2105. Both add containerID to FilesystemOperationRequest so the operation runs against the intended container. The differences:

  • FiTrimResult.trimmed_bytes is populated here. The field arrived with the trim operation in [vminitd]: api for trim filesystem operations #700 and has never been set: trimFilesystem(fd:) throws away the count that FITRIM writes back into the range, and filesystemOperation returns a bare .init(), so the oneof is never set and callers decode the default. [vminitd]: update filesystem operation to run in new namespace #838 keeps return .init(), so the field stays dead after it. Reporting is the whole reason the field exists, and a caller that loops until a trim returns nothing cannot work without it.
  • containerID stays optional.[vminitd]: update filesystem operation to run in new namespace #838 rejects a request that does not carry one, which breaks the existing path-addressed callers. Here a request names either a path or a container, so existing callers keep working.
  • The host side gains the operation it needs.LinuxContainer.trimRootfs() and LinuxPod.trimContainer(_:) return the bytes discarded, rather than leaving each caller to assemble the request.
  • There is a test for the count.testContainerTrimReportsBytes boots a container, trims, and asserts the answer is not zero, which is what fails today.

One thing #838 does that this does not: it enters the container's mount namespace with setns(CLONE_NEWNS) on a dedicated thread after unshare(CLONE_FS). That is a correct way to do it, and it generalises to any path inside the container. Resolving through /proc/<pid>/root reaches the same filesystem through the same namespace without moving a thread between namespaces. If maintainers prefer the setns approach, the reporting fix and the host-side API here apply on top of it unchanged.

Testing

  • swift build and make check clean.
  • swift test: 603 tests in 83 suites passed.
  • make vminitd builds the guest with the change.
  • Integration: ./bin/containerization-integration --filter trim runs container trim reports bytes and the existing container trim ext4 clone, both passing. The new test fails before the reporting fix with trim reported 0 bytes.

A filesystem operation takes a path in the init namespace, and a
container's root filesystem is not there: the container mounts it in a
namespace of its own, and after the pivot nothing in the init namespace
names it. The operations that matter for a root filesystem, trim above
all, were out of reach for exactly the filesystems that grow.
A request that names a container reaches its root filesystem through
the init process's proc entry: the kernel resolves /proc/<pid>/root
through that process's namespace, and the trailing dot component keeps
the opened target a directory rather than the magic link itself, which
the handler's O_NOFOLLOW refuses. The agent surface takes the container
id and returns the bytes the filesystem reported trimmed, from the
machine's one agent, for a single container and for a pod's members
alike.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: A trim never reports the bytes it discarded

2 participants

@MayCXC@crosbymichael
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Let a filesystem operation name a container, and report the bytes a trim discarded - #859

Closed
MayCXC wants to merge 1 commit into
apple:mainfrom
MayCXC:trim-container-rootfs
Closed

Let a filesystem operation name a container, and report the bytes a trim discarded#859
MayCXC wants to merge 1 commit into
apple:mainfrom
MayCXC:trim-container-rootfs

Conversation

@MayCXC

Copy link
Copy Markdown

Summary

A filesystem operation takes a path in the init namespace, and a container's root filesystem is not there: the container mounts it in a namespace of its own, and after the pivot nothing in the init namespace names it. The operations that matter for a root filesystem, trim above all, were out of reach for exactly the filesystems that grow.

A request that names a container reaches its root filesystem through the init process's proc entry: the kernel resolves /proc/<pid>/root through that process's namespace, and the trailing dot component keeps the opened target a directory rather than the magic link itself, which the handler's O_NOFOLLOW refuses. The agent surface takes the container id and answers with the bytes the filesystem reported trimmed, for a single container and for a pod's members alike.

Closes#858. Addresses the requirement described in apple/container#2105.

Relationship to #838

This supersedes #838, which asks for the same thing from issue apple/container#2105. Both add containerID to FilesystemOperationRequest so the operation runs against the intended container. The differences:

  • FiTrimResult.trimmed_bytes is populated here. The field arrived with the trim operation in [vminitd]: api for trim filesystem operations #700 and has never been set: trimFilesystem(fd:) throws away the count that FITRIM writes back into the range, and filesystemOperation returns a bare .init(), so the oneof is never set and callers decode the default. [vminitd]: update filesystem operation to run in new namespace #838 keeps return .init(), so the field stays dead after it. Reporting is the whole reason the field exists, and a caller that loops until a trim returns nothing cannot work without it.
  • containerID stays optional.[vminitd]: update filesystem operation to run in new namespace #838 rejects a request that does not carry one, which breaks the existing path-addressed callers. Here a request names either a path or a container, so existing callers keep working.
  • The host side gains the operation it needs.LinuxContainer.trimRootfs() and LinuxPod.trimContainer(_:) return the bytes discarded, rather than leaving each caller to assemble the request.
  • There is a test for the count.testContainerTrimReportsBytes boots a container, trims, and asserts the answer is not zero, which is what fails today.

One thing #838 does that this does not: it enters the container's mount namespace with setns(CLONE_NEWNS) on a dedicated thread after unshare(CLONE_FS). That is a correct way to do it, and it generalises to any path inside the container. Resolving through /proc/<pid>/root reaches the same filesystem through the same namespace without moving a thread between namespaces. If maintainers prefer the setns approach, the reporting fix and the host-side API here apply on top of it unchanged.

Testing

  • swift build and make check clean.
  • swift test: 603 tests in 83 suites passed.
  • make vminitd builds the guest with the change.
  • Integration: ./bin/containerization-integration --filter trim runs container trim reports bytes and the existing container trim ext4 clone, both passing. The new test fails before the reporting fix with trim reported 0 bytes.

A filesystem operation takes a path in the init namespace, and a
container's root filesystem is not there: the container mounts it in a
namespace of its own, and after the pivot nothing in the init namespace
names it. The operations that matter for a root filesystem, trim above
all, were out of reach for exactly the filesystems that grow.
A request that names a container reaches its root filesystem through
the init process's proc entry: the kernel resolves /proc/<pid>/root
through that process's namespace, and the trailing dot component keeps
the opened target a directory rather than the magic link itself, which
the handler's O_NOFOLLOW refuses. The agent surface takes the container
id and returns the bytes the filesystem reported trimmed, from the
machine's one agent, for a single container and for a pod's members
alike.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: A trim never reports the bytes it discarded

2 participants

@MayCXC@crosbymichael
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Let a filesystem operation name a container, and report the bytes a trim discarded - #859

Closed
MayCXC wants to merge 1 commit into
apple:mainfrom
MayCXC:trim-container-rootfs
Closed

Let a filesystem operation name a container, and report the bytes a trim discarded#859
MayCXC wants to merge 1 commit into
apple:mainfrom
MayCXC:trim-container-rootfs

Conversation

@MayCXC

Copy link
Copy Markdown

Summary

A filesystem operation takes a path in the init namespace, and a container's root filesystem is not there: the container mounts it in a namespace of its own, and after the pivot nothing in the init namespace names it. The operations that matter for a root filesystem, trim above all, were out of reach for exactly the filesystems that grow.

A request that names a container reaches its root filesystem through the init process's proc entry: the kernel resolves /proc/<pid>/root through that process's namespace, and the trailing dot component keeps the opened target a directory rather than the magic link itself, which the handler's O_NOFOLLOW refuses. The agent surface takes the container id and answers with the bytes the filesystem reported trimmed, for a single container and for a pod's members alike.

Closes#858. Addresses the requirement described in apple/container#2105.

Relationship to #838

This supersedes #838, which asks for the same thing from issue apple/container#2105. Both add containerID to FilesystemOperationRequest so the operation runs against the intended container. The differences:

  • FiTrimResult.trimmed_bytes is populated here. The field arrived with the trim operation in [vminitd]: api for trim filesystem operations #700 and has never been set: trimFilesystem(fd:) throws away the count that FITRIM writes back into the range, and filesystemOperation returns a bare .init(), so the oneof is never set and callers decode the default. [vminitd]: update filesystem operation to run in new namespace #838 keeps return .init(), so the field stays dead after it. Reporting is the whole reason the field exists, and a caller that loops until a trim returns nothing cannot work without it.
  • containerID stays optional.[vminitd]: update filesystem operation to run in new namespace #838 rejects a request that does not carry one, which breaks the existing path-addressed callers. Here a request names either a path or a container, so existing callers keep working.
  • The host side gains the operation it needs.LinuxContainer.trimRootfs() and LinuxPod.trimContainer(_:) return the bytes discarded, rather than leaving each caller to assemble the request.
  • There is a test for the count.testContainerTrimReportsBytes boots a container, trims, and asserts the answer is not zero, which is what fails today.

One thing #838 does that this does not: it enters the container's mount namespace with setns(CLONE_NEWNS) on a dedicated thread after unshare(CLONE_FS). That is a correct way to do it, and it generalises to any path inside the container. Resolving through /proc/<pid>/root reaches the same filesystem through the same namespace without moving a thread between namespaces. If maintainers prefer the setns approach, the reporting fix and the host-side API here apply on top of it unchanged.

Testing

  • swift build and make check clean.
  • swift test: 603 tests in 83 suites passed.
  • make vminitd builds the guest with the change.
  • Integration: ./bin/containerization-integration --filter trim runs container trim reports bytes and the existing container trim ext4 clone, both passing. The new test fails before the reporting fix with trim reported 0 bytes.

A filesystem operation takes a path in the init namespace, and a
container's root filesystem is not there: the container mounts it in a
namespace of its own, and after the pivot nothing in the init namespace
names it. The operations that matter for a root filesystem, trim above
all, were out of reach for exactly the filesystems that grow.
A request that names a container reaches its root filesystem through
the init process's proc entry: the kernel resolves /proc/<pid>/root
through that process's namespace, and the trailing dot component keeps
the opened target a directory rather than the magic link itself, which
the handler's O_NOFOLLOW refuses. The agent surface takes the container
id and returns the bytes the filesystem reported trimmed, from the
machine's one agent, for a single container and for a pod's members
alike.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: A trim never reports the bytes it discarded

2 participants

@MayCXC@crosbymichael
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Let a filesystem operation name a container, and report the bytes a trim discarded - #859

Closed
MayCXC wants to merge 1 commit into
apple:mainfrom
MayCXC:trim-container-rootfs
Closed

Let a filesystem operation name a container, and report the bytes a trim discarded#859
MayCXC wants to merge 1 commit into
apple:mainfrom
MayCXC:trim-container-rootfs

Conversation

@MayCXC

Copy link
Copy Markdown

Summary

A filesystem operation takes a path in the init namespace, and a container's root filesystem is not there: the container mounts it in a namespace of its own, and after the pivot nothing in the init namespace names it. The operations that matter for a root filesystem, trim above all, were out of reach for exactly the filesystems that grow.

A request that names a container reaches its root filesystem through the init process's proc entry: the kernel resolves /proc/<pid>/root through that process's namespace, and the trailing dot component keeps the opened target a directory rather than the magic link itself, which the handler's O_NOFOLLOW refuses. The agent surface takes the container id and answers with the bytes the filesystem reported trimmed, for a single container and for a pod's members alike.

Closes#858. Addresses the requirement described in apple/container#2105.

Relationship to #838

This supersedes #838, which asks for the same thing from issue apple/container#2105. Both add containerID to FilesystemOperationRequest so the operation runs against the intended container. The differences:

  • FiTrimResult.trimmed_bytes is populated here. The field arrived with the trim operation in [vminitd]: api for trim filesystem operations #700 and has never been set: trimFilesystem(fd:) throws away the count that FITRIM writes back into the range, and filesystemOperation returns a bare .init(), so the oneof is never set and callers decode the default. [vminitd]: update filesystem operation to run in new namespace #838 keeps return .init(), so the field stays dead after it. Reporting is the whole reason the field exists, and a caller that loops until a trim returns nothing cannot work without it.
  • containerID stays optional.[vminitd]: update filesystem operation to run in new namespace #838 rejects a request that does not carry one, which breaks the existing path-addressed callers. Here a request names either a path or a container, so existing callers keep working.
  • The host side gains the operation it needs.LinuxContainer.trimRootfs() and LinuxPod.trimContainer(_:) return the bytes discarded, rather than leaving each caller to assemble the request.
  • There is a test for the count.testContainerTrimReportsBytes boots a container, trims, and asserts the answer is not zero, which is what fails today.

One thing #838 does that this does not: it enters the container's mount namespace with setns(CLONE_NEWNS) on a dedicated thread after unshare(CLONE_FS). That is a correct way to do it, and it generalises to any path inside the container. Resolving through /proc/<pid>/root reaches the same filesystem through the same namespace without moving a thread between namespaces. If maintainers prefer the setns approach, the reporting fix and the host-side API here apply on top of it unchanged.

Testing

  • swift build and make check clean.
  • swift test: 603 tests in 83 suites passed.
  • make vminitd builds the guest with the change.
  • Integration: ./bin/containerization-integration --filter trim runs container trim reports bytes and the existing container trim ext4 clone, both passing. The new test fails before the reporting fix with trim reported 0 bytes.

A filesystem operation takes a path in the init namespace, and a
container's root filesystem is not there: the container mounts it in a
namespace of its own, and after the pivot nothing in the init namespace
names it. The operations that matter for a root filesystem, trim above
all, were out of reach for exactly the filesystems that grow.
A request that names a container reaches its root filesystem through
the init process's proc entry: the kernel resolves /proc/<pid>/root
through that process's namespace, and the trailing dot component keeps
the opened target a directory rather than the magic link itself, which
the handler's O_NOFOLLOW refuses. The agent surface takes the container
id and returns the bytes the filesystem reported trimmed, from the
machine's one agent, for a single container and for a pod's members
alike.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: A trim never reports the bytes it discarded

2 participants

@MayCXC@crosbymichael
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Let a filesystem operation name a container, and report the bytes a trim discarded - #859

Closed
MayCXC wants to merge 1 commit into
apple:mainfrom
MayCXC:trim-container-rootfs
Closed

Let a filesystem operation name a container, and report the bytes a trim discarded#859
MayCXC wants to merge 1 commit into
apple:mainfrom
MayCXC:trim-container-rootfs

Conversation

@MayCXC

Copy link
Copy Markdown

Summary

A filesystem operation takes a path in the init namespace, and a container's root filesystem is not there: the container mounts it in a namespace of its own, and after the pivot nothing in the init namespace names it. The operations that matter for a root filesystem, trim above all, were out of reach for exactly the filesystems that grow.

A request that names a container reaches its root filesystem through the init process's proc entry: the kernel resolves /proc/<pid>/root through that process's namespace, and the trailing dot component keeps the opened target a directory rather than the magic link itself, which the handler's O_NOFOLLOW refuses. The agent surface takes the container id and answers with the bytes the filesystem reported trimmed, for a single container and for a pod's members alike.

Closes#858. Addresses the requirement described in apple/container#2105.

Relationship to #838

This supersedes #838, which asks for the same thing from issue apple/container#2105. Both add containerID to FilesystemOperationRequest so the operation runs against the intended container. The differences:

  • FiTrimResult.trimmed_bytes is populated here. The field arrived with the trim operation in [vminitd]: api for trim filesystem operations #700 and has never been set: trimFilesystem(fd:) throws away the count that FITRIM writes back into the range, and filesystemOperation returns a bare .init(), so the oneof is never set and callers decode the default. [vminitd]: update filesystem operation to run in new namespace #838 keeps return .init(), so the field stays dead after it. Reporting is the whole reason the field exists, and a caller that loops until a trim returns nothing cannot work without it.
  • containerID stays optional.[vminitd]: update filesystem operation to run in new namespace #838 rejects a request that does not carry one, which breaks the existing path-addressed callers. Here a request names either a path or a container, so existing callers keep working.
  • The host side gains the operation it needs.LinuxContainer.trimRootfs() and LinuxPod.trimContainer(_:) return the bytes discarded, rather than leaving each caller to assemble the request.
  • There is a test for the count.testContainerTrimReportsBytes boots a container, trims, and asserts the answer is not zero, which is what fails today.

One thing #838 does that this does not: it enters the container's mount namespace with setns(CLONE_NEWNS) on a dedicated thread after unshare(CLONE_FS). That is a correct way to do it, and it generalises to any path inside the container. Resolving through /proc/<pid>/root reaches the same filesystem through the same namespace without moving a thread between namespaces. If maintainers prefer the setns approach, the reporting fix and the host-side API here apply on top of it unchanged.

Testing

  • swift build and make check clean.
  • swift test: 603 tests in 83 suites passed.
  • make vminitd builds the guest with the change.
  • Integration: ./bin/containerization-integration --filter trim runs container trim reports bytes and the existing container trim ext4 clone, both passing. The new test fails before the reporting fix with trim reported 0 bytes.

A filesystem operation takes a path in the init namespace, and a
container's root filesystem is not there: the container mounts it in a
namespace of its own, and after the pivot nothing in the init namespace
names it. The operations that matter for a root filesystem, trim above
all, were out of reach for exactly the filesystems that grow.
A request that names a container reaches its root filesystem through
the init process's proc entry: the kernel resolves /proc/<pid>/root
through that process's namespace, and the trailing dot component keeps
the opened target a directory rather than the magic link itself, which
the handler's O_NOFOLLOW refuses. The agent surface takes the container
id and returns the bytes the filesystem reported trimmed, from the
machine's one agent, for a single container and for a pod's members
alike.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: A trim never reports the bytes it discarded

2 participants

@MayCXC@crosbymichael
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Let a filesystem operation name a container, and report the bytes a trim discarded - #859

Closed
MayCXC wants to merge 1 commit into
apple:mainfrom
MayCXC:trim-container-rootfs
Closed

Let a filesystem operation name a container, and report the bytes a trim discarded#859
MayCXC wants to merge 1 commit into
apple:mainfrom
MayCXC:trim-container-rootfs

Conversation

@MayCXC

Copy link
Copy Markdown

Summary

A filesystem operation takes a path in the init namespace, and a container's root filesystem is not there: the container mounts it in a namespace of its own, and after the pivot nothing in the init namespace names it. The operations that matter for a root filesystem, trim above all, were out of reach for exactly the filesystems that grow.

A request that names a container reaches its root filesystem through the init process's proc entry: the kernel resolves /proc/<pid>/root through that process's namespace, and the trailing dot component keeps the opened target a directory rather than the magic link itself, which the handler's O_NOFOLLOW refuses. The agent surface takes the container id and answers with the bytes the filesystem reported trimmed, for a single container and for a pod's members alike.

Closes#858. Addresses the requirement described in apple/container#2105.

Relationship to #838

This supersedes #838, which asks for the same thing from issue apple/container#2105. Both add containerID to FilesystemOperationRequest so the operation runs against the intended container. The differences:

  • FiTrimResult.trimmed_bytes is populated here. The field arrived with the trim operation in [vminitd]: api for trim filesystem operations #700 and has never been set: trimFilesystem(fd:) throws away the count that FITRIM writes back into the range, and filesystemOperation returns a bare .init(), so the oneof is never set and callers decode the default. [vminitd]: update filesystem operation to run in new namespace #838 keeps return .init(), so the field stays dead after it. Reporting is the whole reason the field exists, and a caller that loops until a trim returns nothing cannot work without it.
  • containerID stays optional.[vminitd]: update filesystem operation to run in new namespace #838 rejects a request that does not carry one, which breaks the existing path-addressed callers. Here a request names either a path or a container, so existing callers keep working.
  • The host side gains the operation it needs.LinuxContainer.trimRootfs() and LinuxPod.trimContainer(_:) return the bytes discarded, rather than leaving each caller to assemble the request.
  • There is a test for the count.testContainerTrimReportsBytes boots a container, trims, and asserts the answer is not zero, which is what fails today.

One thing #838 does that this does not: it enters the container's mount namespace with setns(CLONE_NEWNS) on a dedicated thread after unshare(CLONE_FS). That is a correct way to do it, and it generalises to any path inside the container. Resolving through /proc/<pid>/root reaches the same filesystem through the same namespace without moving a thread between namespaces. If maintainers prefer the setns approach, the reporting fix and the host-side API here apply on top of it unchanged.

Testing

  • swift build and make check clean.
  • swift test: 603 tests in 83 suites passed.
  • make vminitd builds the guest with the change.
  • Integration: ./bin/containerization-integration --filter trim runs container trim reports bytes and the existing container trim ext4 clone, both passing. The new test fails before the reporting fix with trim reported 0 bytes.

A filesystem operation takes a path in the init namespace, and a
container's root filesystem is not there: the container mounts it in a
namespace of its own, and after the pivot nothing in the init namespace
names it. The operations that matter for a root filesystem, trim above
all, were out of reach for exactly the filesystems that grow.
A request that names a container reaches its root filesystem through
the init process's proc entry: the kernel resolves /proc/<pid>/root
through that process's namespace, and the trailing dot component keeps
the opened target a directory rather than the magic link itself, which
the handler's O_NOFOLLOW refuses. The agent surface takes the container
id and returns the bytes the filesystem reported trimmed, from the
machine's one agent, for a single container and for a pod's members
alike.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: A trim never reports the bytes it discarded

2 participants

@MayCXC@crosbymichael