Skip to content

Add custom OOM killer for Linux containers - #653

Closed
JaewonHur wants to merge 1 commit into
apple:mainfrom
JaewonHur:oom-kill
Closed

Add custom OOM killer for Linux containers#653
JaewonHur wants to merge 1 commit into
apple:mainfrom
JaewonHur:oom-kill

Conversation

@JaewonHur

Copy link
Copy Markdown
Contributor

This PR implements a custom OOM killer that is spawned as a child process of vmexec.

While Linux kernel also OOM kills a process if it hits cgroup memory limit and the kernel cannot reclaim the memory, kernel often fails to kill the process and left the system hang due to the memory thrashing. Especially, the process is not OOM killed because the kernel still succeeds reclaiming the memory, not meeting the condition for OOM kill (but which takes way longer time, and leads to hang).

Thus, this PR adds a user space OOM killer as a child process of vmexec, which monitors cgroup memory events, and kills the process when max event hits a specified limit. This approach can reliably kills the OOM process as monitoring memory events can be performed in small time window.

This PR needs following more works:

  1. Plumb UI to inform the users that the container has been killed due to the OOM.
  2. Refactor errorPipe to catch errors from (long running) OOM killer process (or any other ways to catch the errors).

@dkovba
dkovba self-requested a review April 6, 2026 20:03
usleep(1_000_000)
let events = try cgroupManager.getMemoryEvents()

if events.max > oomLimit {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

events.max represents the number of events. 1_000_000 seems to be a too large threshold. Would it be appropriate to use a threshold of zero?


while true {
usleep(1_000_000)
let events = try cgroupManager.getMemoryEvents()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we move MemoryMonitor to the Cgroup library and use it instead of pulling memory events with a fixed interval? CC @dcantah

let events = try cgroupManager.getMemoryEvents()

if events.max > oomLimit {
try cgroupManager.kill()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we use App.writeError(error) and exit(code) when try fails?

@dcantah

Copy link
Copy Markdown
Contributor

I'm sort of confused on what this is trying to solve. If the idea is we have some oom kills (likely child processes) that happen but init keeps running, there exists a cgroup toggle that makes it such that every process in the cgroup gets killed if there was an oom condition. Meaning, if the init process for the container is well within its limits, but some child process(es) keep getting oom killed, the kernel would kill the whole cgroup (and thus the whole container).

@dkovba

Copy link
Copy Markdown
Contributor

When we run out of memory, a container hangs. The goal is to make it crash with an OOM error.

@dcantah

dcantah commented Apr 6, 2026

Copy link
Copy Markdown
Contributor

Ok, regardless of what we decide I don't think we should run a separate forked process to do this. We can try and expose an API on LinuxContainer to monitor memory events in a stream like fashion. If we don't want to to do that either, today this whole scheme could be done with the APIs we expose right now. You could call LinuxContainer.statistics every {arbitrary} seconds and check the memoryEvents field.

@crosbymichael

Copy link
Copy Markdown
Contributor

The kernel is doing the correct thing here, it is reclaiming memory without SIGKILL. SIGKILL of any process should always be a last resort because of data loss. We need to look into the root cause for why reclaim is slow. We don't have swap, therefore in low memory situations, reclaim can take a while and we could have a lock up. We should look into overhead for the VM and protect responsiveness for vminitd, keeping the container's memory.max < vm memory.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@JaewonHur@dcantah@dkovba@crosbymichael@jglogan
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Add custom OOM killer for Linux containers by JaewonHur · Pull Request #653 · apple/containerization · GitHub
Skip to content

Add custom OOM killer for Linux containers - #653

Closed
JaewonHur wants to merge 1 commit into
apple:mainfrom
JaewonHur:oom-kill
Closed

Add custom OOM killer for Linux containers#653
JaewonHur wants to merge 1 commit into
apple:mainfrom
JaewonHur:oom-kill

Conversation

@JaewonHur

Copy link
Copy Markdown
Contributor

This PR implements a custom OOM killer that is spawned as a child process of vmexec.

While Linux kernel also OOM kills a process if it hits cgroup memory limit and the kernel cannot reclaim the memory, kernel often fails to kill the process and left the system hang due to the memory thrashing. Especially, the process is not OOM killed because the kernel still succeeds reclaiming the memory, not meeting the condition for OOM kill (but which takes way longer time, and leads to hang).

Thus, this PR adds a user space OOM killer as a child process of vmexec, which monitors cgroup memory events, and kills the process when max event hits a specified limit. This approach can reliably kills the OOM process as monitoring memory events can be performed in small time window.

This PR needs following more works:

  1. Plumb UI to inform the users that the container has been killed due to the OOM.
  2. Refactor errorPipe to catch errors from (long running) OOM killer process (or any other ways to catch the errors).

@dkovba
dkovba self-requested a review April 6, 2026 20:03
usleep(1_000_000)
let events = try cgroupManager.getMemoryEvents()

if events.max > oomLimit {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

events.max represents the number of events. 1_000_000 seems to be a too large threshold. Would it be appropriate to use a threshold of zero?


while true {
usleep(1_000_000)
let events = try cgroupManager.getMemoryEvents()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we move MemoryMonitor to the Cgroup library and use it instead of pulling memory events with a fixed interval? CC @dcantah

let events = try cgroupManager.getMemoryEvents()

if events.max > oomLimit {
try cgroupManager.kill()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we use App.writeError(error) and exit(code) when try fails?

@dcantah

Copy link
Copy Markdown
Contributor

I'm sort of confused on what this is trying to solve. If the idea is we have some oom kills (likely child processes) that happen but init keeps running, there exists a cgroup toggle that makes it such that every process in the cgroup gets killed if there was an oom condition. Meaning, if the init process for the container is well within its limits, but some child process(es) keep getting oom killed, the kernel would kill the whole cgroup (and thus the whole container).

@dkovba

Copy link
Copy Markdown
Contributor

When we run out of memory, a container hangs. The goal is to make it crash with an OOM error.

@dcantah

dcantah commented Apr 6, 2026

Copy link
Copy Markdown
Contributor

Ok, regardless of what we decide I don't think we should run a separate forked process to do this. We can try and expose an API on LinuxContainer to monitor memory events in a stream like fashion. If we don't want to to do that either, today this whole scheme could be done with the APIs we expose right now. You could call LinuxContainer.statistics every {arbitrary} seconds and check the memoryEvents field.

@crosbymichael

Copy link
Copy Markdown
Contributor

The kernel is doing the correct thing here, it is reclaiming memory without SIGKILL. SIGKILL of any process should always be a last resort because of data loss. We need to look into the root cause for why reclaim is slow. We don't have swap, therefore in low memory situations, reclaim can take a while and we could have a lock up. We should look into overhead for the VM and protect responsiveness for vminitd, keeping the container's memory.max < vm memory.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@JaewonHur@dcantah@dkovba@crosbymichael@jglogan
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add custom OOM killer for Linux containers by JaewonHur · Pull Request #653 · apple/containerization · GitHub
Skip to content

Add custom OOM killer for Linux containers - #653

Closed
JaewonHur wants to merge 1 commit into
apple:mainfrom
JaewonHur:oom-kill
Closed

Add custom OOM killer for Linux containers#653
JaewonHur wants to merge 1 commit into
apple:mainfrom
JaewonHur:oom-kill

Conversation

@JaewonHur

Copy link
Copy Markdown
Contributor

This PR implements a custom OOM killer that is spawned as a child process of vmexec.

While Linux kernel also OOM kills a process if it hits cgroup memory limit and the kernel cannot reclaim the memory, kernel often fails to kill the process and left the system hang due to the memory thrashing. Especially, the process is not OOM killed because the kernel still succeeds reclaiming the memory, not meeting the condition for OOM kill (but which takes way longer time, and leads to hang).

Thus, this PR adds a user space OOM killer as a child process of vmexec, which monitors cgroup memory events, and kills the process when max event hits a specified limit. This approach can reliably kills the OOM process as monitoring memory events can be performed in small time window.

This PR needs following more works:

  1. Plumb UI to inform the users that the container has been killed due to the OOM.
  2. Refactor errorPipe to catch errors from (long running) OOM killer process (or any other ways to catch the errors).

@dkovba
dkovba self-requested a review April 6, 2026 20:03
usleep(1_000_000)
let events = try cgroupManager.getMemoryEvents()

if events.max > oomLimit {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

events.max represents the number of events. 1_000_000 seems to be a too large threshold. Would it be appropriate to use a threshold of zero?


while true {
usleep(1_000_000)
let events = try cgroupManager.getMemoryEvents()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we move MemoryMonitor to the Cgroup library and use it instead of pulling memory events with a fixed interval? CC @dcantah

let events = try cgroupManager.getMemoryEvents()

if events.max > oomLimit {
try cgroupManager.kill()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we use App.writeError(error) and exit(code) when try fails?

@dcantah

Copy link
Copy Markdown
Contributor

I'm sort of confused on what this is trying to solve. If the idea is we have some oom kills (likely child processes) that happen but init keeps running, there exists a cgroup toggle that makes it such that every process in the cgroup gets killed if there was an oom condition. Meaning, if the init process for the container is well within its limits, but some child process(es) keep getting oom killed, the kernel would kill the whole cgroup (and thus the whole container).

@dkovba

Copy link
Copy Markdown
Contributor

When we run out of memory, a container hangs. The goal is to make it crash with an OOM error.

@dcantah

dcantah commented Apr 6, 2026

Copy link
Copy Markdown
Contributor

Ok, regardless of what we decide I don't think we should run a separate forked process to do this. We can try and expose an API on LinuxContainer to monitor memory events in a stream like fashion. If we don't want to to do that either, today this whole scheme could be done with the APIs we expose right now. You could call LinuxContainer.statistics every {arbitrary} seconds and check the memoryEvents field.

@crosbymichael

Copy link
Copy Markdown
Contributor

The kernel is doing the correct thing here, it is reclaiming memory without SIGKILL. SIGKILL of any process should always be a last resort because of data loss. We need to look into the root cause for why reclaim is slow. We don't have swap, therefore in low memory situations, reclaim can take a while and we could have a lock up. We should look into overhead for the VM and protect responsiveness for vminitd, keeping the container's memory.max < vm memory.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@JaewonHur@dcantah@dkovba@crosbymichael@jglogan
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add custom OOM killer for Linux containers by JaewonHur · Pull Request #653 · apple/containerization · GitHub
Skip to content

Add custom OOM killer for Linux containers - #653

Closed
JaewonHur wants to merge 1 commit into
apple:mainfrom
JaewonHur:oom-kill
Closed

Add custom OOM killer for Linux containers#653
JaewonHur wants to merge 1 commit into
apple:mainfrom
JaewonHur:oom-kill

Conversation

@JaewonHur

Copy link
Copy Markdown
Contributor

This PR implements a custom OOM killer that is spawned as a child process of vmexec.

While Linux kernel also OOM kills a process if it hits cgroup memory limit and the kernel cannot reclaim the memory, kernel often fails to kill the process and left the system hang due to the memory thrashing. Especially, the process is not OOM killed because the kernel still succeeds reclaiming the memory, not meeting the condition for OOM kill (but which takes way longer time, and leads to hang).

Thus, this PR adds a user space OOM killer as a child process of vmexec, which monitors cgroup memory events, and kills the process when max event hits a specified limit. This approach can reliably kills the OOM process as monitoring memory events can be performed in small time window.

This PR needs following more works:

  1. Plumb UI to inform the users that the container has been killed due to the OOM.
  2. Refactor errorPipe to catch errors from (long running) OOM killer process (or any other ways to catch the errors).

@dkovba
dkovba self-requested a review April 6, 2026 20:03
usleep(1_000_000)
let events = try cgroupManager.getMemoryEvents()

if events.max > oomLimit {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

events.max represents the number of events. 1_000_000 seems to be a too large threshold. Would it be appropriate to use a threshold of zero?


while true {
usleep(1_000_000)
let events = try cgroupManager.getMemoryEvents()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we move MemoryMonitor to the Cgroup library and use it instead of pulling memory events with a fixed interval? CC @dcantah

let events = try cgroupManager.getMemoryEvents()

if events.max > oomLimit {
try cgroupManager.kill()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we use App.writeError(error) and exit(code) when try fails?

@dcantah

Copy link
Copy Markdown
Contributor

I'm sort of confused on what this is trying to solve. If the idea is we have some oom kills (likely child processes) that happen but init keeps running, there exists a cgroup toggle that makes it such that every process in the cgroup gets killed if there was an oom condition. Meaning, if the init process for the container is well within its limits, but some child process(es) keep getting oom killed, the kernel would kill the whole cgroup (and thus the whole container).

@dkovba

Copy link
Copy Markdown
Contributor

When we run out of memory, a container hangs. The goal is to make it crash with an OOM error.

@dcantah

dcantah commented Apr 6, 2026

Copy link
Copy Markdown
Contributor

Ok, regardless of what we decide I don't think we should run a separate forked process to do this. We can try and expose an API on LinuxContainer to monitor memory events in a stream like fashion. If we don't want to to do that either, today this whole scheme could be done with the APIs we expose right now. You could call LinuxContainer.statistics every {arbitrary} seconds and check the memoryEvents field.

@crosbymichael

Copy link
Copy Markdown
Contributor

The kernel is doing the correct thing here, it is reclaiming memory without SIGKILL. SIGKILL of any process should always be a last resort because of data loss. We need to look into the root cause for why reclaim is slow. We don't have swap, therefore in low memory situations, reclaim can take a while and we could have a lock up. We should look into overhead for the VM and protect responsiveness for vminitd, keeping the container's memory.max < vm memory.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@JaewonHur@dcantah@dkovba@crosbymichael@jglogan
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Add custom OOM killer for Linux containers by JaewonHur · Pull Request #653 · apple/containerization · GitHub
Skip to content

Add custom OOM killer for Linux containers - #653

Closed
JaewonHur wants to merge 1 commit into
apple:mainfrom
JaewonHur:oom-kill
Closed

Add custom OOM killer for Linux containers#653
JaewonHur wants to merge 1 commit into
apple:mainfrom
JaewonHur:oom-kill

Conversation

@JaewonHur

Copy link
Copy Markdown
Contributor

This PR implements a custom OOM killer that is spawned as a child process of vmexec.

While Linux kernel also OOM kills a process if it hits cgroup memory limit and the kernel cannot reclaim the memory, kernel often fails to kill the process and left the system hang due to the memory thrashing. Especially, the process is not OOM killed because the kernel still succeeds reclaiming the memory, not meeting the condition for OOM kill (but which takes way longer time, and leads to hang).

Thus, this PR adds a user space OOM killer as a child process of vmexec, which monitors cgroup memory events, and kills the process when max event hits a specified limit. This approach can reliably kills the OOM process as monitoring memory events can be performed in small time window.

This PR needs following more works:

  1. Plumb UI to inform the users that the container has been killed due to the OOM.
  2. Refactor errorPipe to catch errors from (long running) OOM killer process (or any other ways to catch the errors).

@dkovba
dkovba self-requested a review April 6, 2026 20:03
usleep(1_000_000)
let events = try cgroupManager.getMemoryEvents()

if events.max > oomLimit {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

events.max represents the number of events. 1_000_000 seems to be a too large threshold. Would it be appropriate to use a threshold of zero?


while true {
usleep(1_000_000)
let events = try cgroupManager.getMemoryEvents()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we move MemoryMonitor to the Cgroup library and use it instead of pulling memory events with a fixed interval? CC @dcantah

let events = try cgroupManager.getMemoryEvents()

if events.max > oomLimit {
try cgroupManager.kill()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we use App.writeError(error) and exit(code) when try fails?

@dcantah

Copy link
Copy Markdown
Contributor

I'm sort of confused on what this is trying to solve. If the idea is we have some oom kills (likely child processes) that happen but init keeps running, there exists a cgroup toggle that makes it such that every process in the cgroup gets killed if there was an oom condition. Meaning, if the init process for the container is well within its limits, but some child process(es) keep getting oom killed, the kernel would kill the whole cgroup (and thus the whole container).

@dkovba

Copy link
Copy Markdown
Contributor

When we run out of memory, a container hangs. The goal is to make it crash with an OOM error.

@dcantah

dcantah commented Apr 6, 2026

Copy link
Copy Markdown
Contributor

Ok, regardless of what we decide I don't think we should run a separate forked process to do this. We can try and expose an API on LinuxContainer to monitor memory events in a stream like fashion. If we don't want to to do that either, today this whole scheme could be done with the APIs we expose right now. You could call LinuxContainer.statistics every {arbitrary} seconds and check the memoryEvents field.

@crosbymichael

Copy link
Copy Markdown
Contributor

The kernel is doing the correct thing here, it is reclaiming memory without SIGKILL. SIGKILL of any process should always be a last resort because of data loss. We need to look into the root cause for why reclaim is slow. We don't have swap, therefore in low memory situations, reclaim can take a while and we could have a lock up. We should look into overhead for the VM and protect responsiveness for vminitd, keeping the container's memory.max < vm memory.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@JaewonHur@dcantah@dkovba@crosbymichael@jglogan
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Add custom OOM killer for Linux containers by JaewonHur · Pull Request #653 · apple/containerization · GitHub
Skip to content

Add custom OOM killer for Linux containers - #653

Closed
JaewonHur wants to merge 1 commit into
apple:mainfrom
JaewonHur:oom-kill
Closed

Add custom OOM killer for Linux containers#653
JaewonHur wants to merge 1 commit into
apple:mainfrom
JaewonHur:oom-kill

Conversation

@JaewonHur

Copy link
Copy Markdown
Contributor

This PR implements a custom OOM killer that is spawned as a child process of vmexec.

While Linux kernel also OOM kills a process if it hits cgroup memory limit and the kernel cannot reclaim the memory, kernel often fails to kill the process and left the system hang due to the memory thrashing. Especially, the process is not OOM killed because the kernel still succeeds reclaiming the memory, not meeting the condition for OOM kill (but which takes way longer time, and leads to hang).

Thus, this PR adds a user space OOM killer as a child process of vmexec, which monitors cgroup memory events, and kills the process when max event hits a specified limit. This approach can reliably kills the OOM process as monitoring memory events can be performed in small time window.

This PR needs following more works:

  1. Plumb UI to inform the users that the container has been killed due to the OOM.
  2. Refactor errorPipe to catch errors from (long running) OOM killer process (or any other ways to catch the errors).

@dkovba
dkovba self-requested a review April 6, 2026 20:03
usleep(1_000_000)
let events = try cgroupManager.getMemoryEvents()

if events.max > oomLimit {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

events.max represents the number of events. 1_000_000 seems to be a too large threshold. Would it be appropriate to use a threshold of zero?


while true {
usleep(1_000_000)
let events = try cgroupManager.getMemoryEvents()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we move MemoryMonitor to the Cgroup library and use it instead of pulling memory events with a fixed interval? CC @dcantah

let events = try cgroupManager.getMemoryEvents()

if events.max > oomLimit {
try cgroupManager.kill()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we use App.writeError(error) and exit(code) when try fails?

@dcantah

Copy link
Copy Markdown
Contributor

I'm sort of confused on what this is trying to solve. If the idea is we have some oom kills (likely child processes) that happen but init keeps running, there exists a cgroup toggle that makes it such that every process in the cgroup gets killed if there was an oom condition. Meaning, if the init process for the container is well within its limits, but some child process(es) keep getting oom killed, the kernel would kill the whole cgroup (and thus the whole container).

@dkovba

Copy link
Copy Markdown
Contributor

When we run out of memory, a container hangs. The goal is to make it crash with an OOM error.

@dcantah

dcantah commented Apr 6, 2026

Copy link
Copy Markdown
Contributor

Ok, regardless of what we decide I don't think we should run a separate forked process to do this. We can try and expose an API on LinuxContainer to monitor memory events in a stream like fashion. If we don't want to to do that either, today this whole scheme could be done with the APIs we expose right now. You could call LinuxContainer.statistics every {arbitrary} seconds and check the memoryEvents field.

@crosbymichael

Copy link
Copy Markdown
Contributor

The kernel is doing the correct thing here, it is reclaiming memory without SIGKILL. SIGKILL of any process should always be a last resort because of data loss. We need to look into the root cause for why reclaim is slow. We don't have swap, therefore in low memory situations, reclaim can take a while and we could have a lock up. We should look into overhead for the VM and protect responsiveness for vminitd, keeping the container's memory.max < vm memory.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@JaewonHur@dcantah@dkovba@crosbymichael@jglogan
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); })(); Add custom OOM killer for Linux containers by JaewonHur · Pull Request #653 · apple/containerization · GitHub
Skip to content

Add custom OOM killer for Linux containers - #653

Closed
JaewonHur wants to merge 1 commit into
apple:mainfrom
JaewonHur:oom-kill
Closed

Add custom OOM killer for Linux containers#653
JaewonHur wants to merge 1 commit into
apple:mainfrom
JaewonHur:oom-kill

Conversation

@JaewonHur

Copy link
Copy Markdown
Contributor

This PR implements a custom OOM killer that is spawned as a child process of vmexec.

While Linux kernel also OOM kills a process if it hits cgroup memory limit and the kernel cannot reclaim the memory, kernel often fails to kill the process and left the system hang due to the memory thrashing. Especially, the process is not OOM killed because the kernel still succeeds reclaiming the memory, not meeting the condition for OOM kill (but which takes way longer time, and leads to hang).

Thus, this PR adds a user space OOM killer as a child process of vmexec, which monitors cgroup memory events, and kills the process when max event hits a specified limit. This approach can reliably kills the OOM process as monitoring memory events can be performed in small time window.

This PR needs following more works:

  1. Plumb UI to inform the users that the container has been killed due to the OOM.
  2. Refactor errorPipe to catch errors from (long running) OOM killer process (or any other ways to catch the errors).

@dkovba
dkovba self-requested a review April 6, 2026 20:03
usleep(1_000_000)
let events = try cgroupManager.getMemoryEvents()

if events.max > oomLimit {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

events.max represents the number of events. 1_000_000 seems to be a too large threshold. Would it be appropriate to use a threshold of zero?


while true {
usleep(1_000_000)
let events = try cgroupManager.getMemoryEvents()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we move MemoryMonitor to the Cgroup library and use it instead of pulling memory events with a fixed interval? CC @dcantah

let events = try cgroupManager.getMemoryEvents()

if events.max > oomLimit {
try cgroupManager.kill()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we use App.writeError(error) and exit(code) when try fails?

@dcantah

Copy link
Copy Markdown
Contributor

I'm sort of confused on what this is trying to solve. If the idea is we have some oom kills (likely child processes) that happen but init keeps running, there exists a cgroup toggle that makes it such that every process in the cgroup gets killed if there was an oom condition. Meaning, if the init process for the container is well within its limits, but some child process(es) keep getting oom killed, the kernel would kill the whole cgroup (and thus the whole container).

@dkovba

Copy link
Copy Markdown
Contributor

When we run out of memory, a container hangs. The goal is to make it crash with an OOM error.

@dcantah

dcantah commented Apr 6, 2026

Copy link
Copy Markdown
Contributor

Ok, regardless of what we decide I don't think we should run a separate forked process to do this. We can try and expose an API on LinuxContainer to monitor memory events in a stream like fashion. If we don't want to to do that either, today this whole scheme could be done with the APIs we expose right now. You could call LinuxContainer.statistics every {arbitrary} seconds and check the memoryEvents field.

@crosbymichael

Copy link
Copy Markdown
Contributor

The kernel is doing the correct thing here, it is reclaiming memory without SIGKILL. SIGKILL of any process should always be a last resort because of data loss. We need to look into the root cause for why reclaim is slow. We don't have swap, therefore in low memory situations, reclaim can take a while and we could have a lock up. We should look into overhead for the VM and protect responsiveness for vminitd, keeping the container's memory.max < vm memory.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@JaewonHur@dcantah@dkovba@crosbymichael@jglogan