Skip to content

Respect task retries for signal killed tasks - #55767

Merged
amoghrajesh merged 6 commits into
apache:mainfrom
astronomer:handle-signal-task-better
Oct 28, 2025
Merged

Respect task retries for signal killed tasks#55767
amoghrajesh merged 6 commits into
apache:mainfrom
astronomer:handle-signal-task-better

Conversation

@amoghrajesh

@amoghrajeshamoghrajesh commented Sep 17, 2025

Copy link
Copy Markdown
Contributor

closes: #55753

Problem

When tasks are killed by system signals (SIGKILL for OOM, SIGTERM for worker restarts), they immediately go to FAILED state instead of respecting the task retries set and going to UP_FOR_RETRY state. This creates unexpected behavior where exception based failures respect retries but signal based failures don't.

Root Cause

The supervisor's final_state property only checked the exit code and didn't consider retry eligibility for signal based failures. While exception based failures properly checked should_retry which is set by the API server when a task is run for the first time a.k.a, in its run context, signal based failures ignored this logic entirely.

Testing

DAG used:

from airflow import DAG
from datetime import datetime
from airflow.providers.standard.operators.python import PythonOperator
def func():
a = "asd"
while True:
a += a*100000
with DAG(
dag_id="oom_example",
start_date=datetime(2024, 1, 1),
schedule=None,
catchup=False,
doc_md=__doc__,
tags=["oom"],
) as dag:
hello_task = PythonOperator(
task_id="oom_task",
python_callable=func,
retries=3
)

Earlier:

The dag immediately moved into FAILED state without any retries;

image

After:

The dag moves into up for retry state as appropriate:

image

See retries:

image

^ Add meaningful description above
Read the Pull Request Guidelines for more information.
In case of fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
In case of a new dependency, check compliance with the ASF 3rd Party License Policy.
In case of backwards incompatible changes please leave a note in a newsfragment file, named {pr_number}.significant.rst or {issue_number}.significant.rst, in airflow-core/newsfragments.

@ashbashb left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SIGABRT and SIGSEGV should not retry

Why? That's not what I would expect to do (and also I suspect not what Airflow 2 did?)

Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
@kaxil

kaxil commented Sep 17, 2025

Copy link
Copy Markdown
Member

Update: Nvm, thanks to @rawwar figured out that K8s/cgroups has configs to kill only the process


When tasks are killed by system signals (SIGKILL for OOM, SIGTERM for worker restarts), they immediately go to FAILED state instead of respecting the task retries set and going to UP_FOR_RETRY state. This creates unexpected behavior where exception based failures respect retries but signal based failures don't.

How common is this scenario (excluding manually killing task process)? Since the supervisor and task processes are running in the same container, wouldn't an OOM condition typically kill the entire container rather than just the individual task process?

In the more common case where the entire container gets OOM-killed:

  1. The supervisor process would also die
  2. Heartbeat to the scheduler would fail
  3. Scheduler would receive a FAILED executor event and handle retries through the normal process_executor_events()handle_failure() path

@ashb

ashb commented Sep 19, 2025

Copy link
Copy Markdown
Member

@kaxil Also people do occasionally run Airflow outside of Kubernetes you know 😉

@amoghrajesh

Copy link
Copy Markdown
ContributorAuthor

I'll get to the comments on this later, not needed for 3.1, can wait till 3.1.1

@amoghrajeshamoghrajesh added this to the Airflow 3.1.1 milestone Sep 19, 2025
@amoghrajesh

Copy link
Copy Markdown
ContributorAuthor

@ashb replied to some of your comments, could you take a look when possible?

Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
@rawwar

rawwar commented Oct 27, 2025

Copy link
Copy Markdown
Contributor

Verified by running a task that gets killed due to oom and the task went to "up_for_retry"

image

Just wanted to let it run all tries. It did fail at the end:
image

@rawwarrawwar left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

@amoghrajesh
amoghrajeshforce-pushed the handle-signal-task-better branch from 3191733 to 24ef66fCompareOctober 28, 2025 08:36
@amoghrajesh
amoghrajesh merged commit de0c78e into apache:mainOct 28, 2025
82 checks passed
@amoghrajesh
amoghrajesh deleted the handle-signal-task-better branch October 28, 2025 09:38
github-actionsBot pushed a commit that referenced this pull request Oct 28, 2025
(cherry picked from commit de0c78e)
Co-authored-by: Amogh Desai <amoghrajesh1999@gmail.com>
@github-actions

Copy link
Copy Markdown
Contributor

Backport successfully created: v3-1-test

StatusBranchResult
v3-1-testPR Link

github-actionsBot pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Oct 28, 2025
(cherry picked from commit de0c78e)
Co-authored-by: Amogh Desai <amoghrajesh1999@gmail.com>
kaxil pushed a commit that referenced this pull request Oct 31, 2025
kaxil pushed a commit that referenced this pull request Oct 31, 2025
@ephraimbuddyephraimbuddy added the type:bug-fix Changelog: Bug Fixes label Nov 10, 2025
odaneau-astro pushed a commit to odaneau-astro/airflow that referenced this pull request Mar 18, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:task-sdktype:bug-fixChangelog: Bug Fixes

Projects

No open projects

Development

Successfully merging this pull request may close these issues.

Task does not retry when worker is killed due to OOM

6 participants

@amoghrajesh@kaxil@ashb@rawwar@potiuk@ephraimbuddy
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Respect task retries for signal killed tasks by amoghrajesh · Pull Request #55767 · apache/airflow · GitHub
Skip to content

Respect task retries for signal killed tasks - #55767

Merged
amoghrajesh merged 6 commits into
apache:mainfrom
astronomer:handle-signal-task-better
Oct 28, 2025
Merged

Respect task retries for signal killed tasks#55767
amoghrajesh merged 6 commits into
apache:mainfrom
astronomer:handle-signal-task-better

Conversation

@amoghrajesh

@amoghrajeshamoghrajesh commented Sep 17, 2025

Copy link
Copy Markdown
Contributor

closes: #55753

Problem

When tasks are killed by system signals (SIGKILL for OOM, SIGTERM for worker restarts), they immediately go to FAILED state instead of respecting the task retries set and going to UP_FOR_RETRY state. This creates unexpected behavior where exception based failures respect retries but signal based failures don't.

Root Cause

The supervisor's final_state property only checked the exit code and didn't consider retry eligibility for signal based failures. While exception based failures properly checked should_retry which is set by the API server when a task is run for the first time a.k.a, in its run context, signal based failures ignored this logic entirely.

Testing

DAG used:

from airflow import DAG
from datetime import datetime
from airflow.providers.standard.operators.python import PythonOperator
def func():
a = "asd"
while True:
a += a*100000
with DAG(
dag_id="oom_example",
start_date=datetime(2024, 1, 1),
schedule=None,
catchup=False,
doc_md=__doc__,
tags=["oom"],
) as dag:
hello_task = PythonOperator(
task_id="oom_task",
python_callable=func,
retries=3
)

Earlier:

The dag immediately moved into FAILED state without any retries;

image

After:

The dag moves into up for retry state as appropriate:

image

See retries:

image

^ Add meaningful description above
Read the Pull Request Guidelines for more information.
In case of fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
In case of a new dependency, check compliance with the ASF 3rd Party License Policy.
In case of backwards incompatible changes please leave a note in a newsfragment file, named {pr_number}.significant.rst or {issue_number}.significant.rst, in airflow-core/newsfragments.

@ashbashb left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SIGABRT and SIGSEGV should not retry

Why? That's not what I would expect to do (and also I suspect not what Airflow 2 did?)

Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
@kaxil

kaxil commented Sep 17, 2025

Copy link
Copy Markdown
Member

Update: Nvm, thanks to @rawwar figured out that K8s/cgroups has configs to kill only the process


When tasks are killed by system signals (SIGKILL for OOM, SIGTERM for worker restarts), they immediately go to FAILED state instead of respecting the task retries set and going to UP_FOR_RETRY state. This creates unexpected behavior where exception based failures respect retries but signal based failures don't.

How common is this scenario (excluding manually killing task process)? Since the supervisor and task processes are running in the same container, wouldn't an OOM condition typically kill the entire container rather than just the individual task process?

In the more common case where the entire container gets OOM-killed:

  1. The supervisor process would also die
  2. Heartbeat to the scheduler would fail
  3. Scheduler would receive a FAILED executor event and handle retries through the normal process_executor_events()handle_failure() path

@ashb

ashb commented Sep 19, 2025

Copy link
Copy Markdown
Member

@kaxil Also people do occasionally run Airflow outside of Kubernetes you know 😉

@amoghrajesh

Copy link
Copy Markdown
ContributorAuthor

I'll get to the comments on this later, not needed for 3.1, can wait till 3.1.1

@amoghrajeshamoghrajesh added this to the Airflow 3.1.1 milestone Sep 19, 2025
@amoghrajesh

Copy link
Copy Markdown
ContributorAuthor

@ashb replied to some of your comments, could you take a look when possible?

Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
@rawwar

rawwar commented Oct 27, 2025

Copy link
Copy Markdown
Contributor

Verified by running a task that gets killed due to oom and the task went to "up_for_retry"

image

Just wanted to let it run all tries. It did fail at the end:
image

@rawwarrawwar left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

@amoghrajesh
amoghrajeshforce-pushed the handle-signal-task-better branch from 3191733 to 24ef66fCompareOctober 28, 2025 08:36
@amoghrajesh
amoghrajesh merged commit de0c78e into apache:mainOct 28, 2025
82 checks passed
@amoghrajesh
amoghrajesh deleted the handle-signal-task-better branch October 28, 2025 09:38
github-actionsBot pushed a commit that referenced this pull request Oct 28, 2025
(cherry picked from commit de0c78e)
Co-authored-by: Amogh Desai <amoghrajesh1999@gmail.com>
@github-actions

Copy link
Copy Markdown
Contributor

Backport successfully created: v3-1-test

StatusBranchResult
v3-1-testPR Link

github-actionsBot pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Oct 28, 2025
(cherry picked from commit de0c78e)
Co-authored-by: Amogh Desai <amoghrajesh1999@gmail.com>
kaxil pushed a commit that referenced this pull request Oct 31, 2025
kaxil pushed a commit that referenced this pull request Oct 31, 2025
@ephraimbuddyephraimbuddy added the type:bug-fix Changelog: Bug Fixes label Nov 10, 2025
odaneau-astro pushed a commit to odaneau-astro/airflow that referenced this pull request Mar 18, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:task-sdktype:bug-fixChangelog: Bug Fixes

Projects

No open projects

Development

Successfully merging this pull request may close these issues.

Task does not retry when worker is killed due to OOM

6 participants

@amoghrajesh@kaxil@ashb@rawwar@potiuk@ephraimbuddy
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Respect task retries for signal killed tasks by amoghrajesh · Pull Request #55767 · apache/airflow · GitHub
Skip to content

Respect task retries for signal killed tasks - #55767

Merged
amoghrajesh merged 6 commits into
apache:mainfrom
astronomer:handle-signal-task-better
Oct 28, 2025
Merged

Respect task retries for signal killed tasks#55767
amoghrajesh merged 6 commits into
apache:mainfrom
astronomer:handle-signal-task-better

Conversation

@amoghrajesh

@amoghrajeshamoghrajesh commented Sep 17, 2025

Copy link
Copy Markdown
Contributor

closes: #55753

Problem

When tasks are killed by system signals (SIGKILL for OOM, SIGTERM for worker restarts), they immediately go to FAILED state instead of respecting the task retries set and going to UP_FOR_RETRY state. This creates unexpected behavior where exception based failures respect retries but signal based failures don't.

Root Cause

The supervisor's final_state property only checked the exit code and didn't consider retry eligibility for signal based failures. While exception based failures properly checked should_retry which is set by the API server when a task is run for the first time a.k.a, in its run context, signal based failures ignored this logic entirely.

Testing

DAG used:

from airflow import DAG
from datetime import datetime
from airflow.providers.standard.operators.python import PythonOperator
def func():
a = "asd"
while True:
a += a*100000
with DAG(
dag_id="oom_example",
start_date=datetime(2024, 1, 1),
schedule=None,
catchup=False,
doc_md=__doc__,
tags=["oom"],
) as dag:
hello_task = PythonOperator(
task_id="oom_task",
python_callable=func,
retries=3
)

Earlier:

The dag immediately moved into FAILED state without any retries;

image

After:

The dag moves into up for retry state as appropriate:

image

See retries:

image

^ Add meaningful description above
Read the Pull Request Guidelines for more information.
In case of fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
In case of a new dependency, check compliance with the ASF 3rd Party License Policy.
In case of backwards incompatible changes please leave a note in a newsfragment file, named {pr_number}.significant.rst or {issue_number}.significant.rst, in airflow-core/newsfragments.

@ashbashb left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SIGABRT and SIGSEGV should not retry

Why? That's not what I would expect to do (and also I suspect not what Airflow 2 did?)

Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
@kaxil

kaxil commented Sep 17, 2025

Copy link
Copy Markdown
Member

Update: Nvm, thanks to @rawwar figured out that K8s/cgroups has configs to kill only the process


When tasks are killed by system signals (SIGKILL for OOM, SIGTERM for worker restarts), they immediately go to FAILED state instead of respecting the task retries set and going to UP_FOR_RETRY state. This creates unexpected behavior where exception based failures respect retries but signal based failures don't.

How common is this scenario (excluding manually killing task process)? Since the supervisor and task processes are running in the same container, wouldn't an OOM condition typically kill the entire container rather than just the individual task process?

In the more common case where the entire container gets OOM-killed:

  1. The supervisor process would also die
  2. Heartbeat to the scheduler would fail
  3. Scheduler would receive a FAILED executor event and handle retries through the normal process_executor_events()handle_failure() path

@ashb

ashb commented Sep 19, 2025

Copy link
Copy Markdown
Member

@kaxil Also people do occasionally run Airflow outside of Kubernetes you know 😉

@amoghrajesh

Copy link
Copy Markdown
ContributorAuthor

I'll get to the comments on this later, not needed for 3.1, can wait till 3.1.1

@amoghrajeshamoghrajesh added this to the Airflow 3.1.1 milestone Sep 19, 2025
@amoghrajesh

Copy link
Copy Markdown
ContributorAuthor

@ashb replied to some of your comments, could you take a look when possible?

Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
@rawwar

rawwar commented Oct 27, 2025

Copy link
Copy Markdown
Contributor

Verified by running a task that gets killed due to oom and the task went to "up_for_retry"

image

Just wanted to let it run all tries. It did fail at the end:
image

@rawwarrawwar left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

@amoghrajesh
amoghrajeshforce-pushed the handle-signal-task-better branch from 3191733 to 24ef66fCompareOctober 28, 2025 08:36
@amoghrajesh
amoghrajesh merged commit de0c78e into apache:mainOct 28, 2025
82 checks passed
@amoghrajesh
amoghrajesh deleted the handle-signal-task-better branch October 28, 2025 09:38
github-actionsBot pushed a commit that referenced this pull request Oct 28, 2025
(cherry picked from commit de0c78e)
Co-authored-by: Amogh Desai <amoghrajesh1999@gmail.com>
@github-actions

Copy link
Copy Markdown
Contributor

Backport successfully created: v3-1-test

StatusBranchResult
v3-1-testPR Link

github-actionsBot pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Oct 28, 2025
(cherry picked from commit de0c78e)
Co-authored-by: Amogh Desai <amoghrajesh1999@gmail.com>
kaxil pushed a commit that referenced this pull request Oct 31, 2025
kaxil pushed a commit that referenced this pull request Oct 31, 2025
@ephraimbuddyephraimbuddy added the type:bug-fix Changelog: Bug Fixes label Nov 10, 2025
odaneau-astro pushed a commit to odaneau-astro/airflow that referenced this pull request Mar 18, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:task-sdktype:bug-fixChangelog: Bug Fixes

Projects

No open projects

Development

Successfully merging this pull request may close these issues.

Task does not retry when worker is killed due to OOM

6 participants

@amoghrajesh@kaxil@ashb@rawwar@potiuk@ephraimbuddy
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Respect task retries for signal killed tasks by amoghrajesh · Pull Request #55767 · apache/airflow · GitHub
Skip to content

Respect task retries for signal killed tasks - #55767

Merged
amoghrajesh merged 6 commits into
apache:mainfrom
astronomer:handle-signal-task-better
Oct 28, 2025
Merged

Respect task retries for signal killed tasks#55767
amoghrajesh merged 6 commits into
apache:mainfrom
astronomer:handle-signal-task-better

Conversation

@amoghrajesh

@amoghrajeshamoghrajesh commented Sep 17, 2025

Copy link
Copy Markdown
Contributor

closes: #55753

Problem

When tasks are killed by system signals (SIGKILL for OOM, SIGTERM for worker restarts), they immediately go to FAILED state instead of respecting the task retries set and going to UP_FOR_RETRY state. This creates unexpected behavior where exception based failures respect retries but signal based failures don't.

Root Cause

The supervisor's final_state property only checked the exit code and didn't consider retry eligibility for signal based failures. While exception based failures properly checked should_retry which is set by the API server when a task is run for the first time a.k.a, in its run context, signal based failures ignored this logic entirely.

Testing

DAG used:

from airflow import DAG
from datetime import datetime
from airflow.providers.standard.operators.python import PythonOperator
def func():
a = "asd"
while True:
a += a*100000
with DAG(
dag_id="oom_example",
start_date=datetime(2024, 1, 1),
schedule=None,
catchup=False,
doc_md=__doc__,
tags=["oom"],
) as dag:
hello_task = PythonOperator(
task_id="oom_task",
python_callable=func,
retries=3
)

Earlier:

The dag immediately moved into FAILED state without any retries;

image

After:

The dag moves into up for retry state as appropriate:

image

See retries:

image

^ Add meaningful description above
Read the Pull Request Guidelines for more information.
In case of fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
In case of a new dependency, check compliance with the ASF 3rd Party License Policy.
In case of backwards incompatible changes please leave a note in a newsfragment file, named {pr_number}.significant.rst or {issue_number}.significant.rst, in airflow-core/newsfragments.

@ashbashb left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SIGABRT and SIGSEGV should not retry

Why? That's not what I would expect to do (and also I suspect not what Airflow 2 did?)

Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
@kaxil

kaxil commented Sep 17, 2025

Copy link
Copy Markdown
Member

Update: Nvm, thanks to @rawwar figured out that K8s/cgroups has configs to kill only the process


When tasks are killed by system signals (SIGKILL for OOM, SIGTERM for worker restarts), they immediately go to FAILED state instead of respecting the task retries set and going to UP_FOR_RETRY state. This creates unexpected behavior where exception based failures respect retries but signal based failures don't.

How common is this scenario (excluding manually killing task process)? Since the supervisor and task processes are running in the same container, wouldn't an OOM condition typically kill the entire container rather than just the individual task process?

In the more common case where the entire container gets OOM-killed:

  1. The supervisor process would also die
  2. Heartbeat to the scheduler would fail
  3. Scheduler would receive a FAILED executor event and handle retries through the normal process_executor_events()handle_failure() path

@ashb

ashb commented Sep 19, 2025

Copy link
Copy Markdown
Member

@kaxil Also people do occasionally run Airflow outside of Kubernetes you know 😉

@amoghrajesh

Copy link
Copy Markdown
ContributorAuthor

I'll get to the comments on this later, not needed for 3.1, can wait till 3.1.1

@amoghrajeshamoghrajesh added this to the Airflow 3.1.1 milestone Sep 19, 2025
@amoghrajesh

Copy link
Copy Markdown
ContributorAuthor

@ashb replied to some of your comments, could you take a look when possible?

Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
@rawwar

rawwar commented Oct 27, 2025

Copy link
Copy Markdown
Contributor

Verified by running a task that gets killed due to oom and the task went to "up_for_retry"

image

Just wanted to let it run all tries. It did fail at the end:
image

@rawwarrawwar left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

@amoghrajesh
amoghrajeshforce-pushed the handle-signal-task-better branch from 3191733 to 24ef66fCompareOctober 28, 2025 08:36
@amoghrajesh
amoghrajesh merged commit de0c78e into apache:mainOct 28, 2025
82 checks passed
@amoghrajesh
amoghrajesh deleted the handle-signal-task-better branch October 28, 2025 09:38
github-actionsBot pushed a commit that referenced this pull request Oct 28, 2025
(cherry picked from commit de0c78e)
Co-authored-by: Amogh Desai <amoghrajesh1999@gmail.com>
@github-actions

Copy link
Copy Markdown
Contributor

Backport successfully created: v3-1-test

StatusBranchResult
v3-1-testPR Link

github-actionsBot pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Oct 28, 2025
(cherry picked from commit de0c78e)
Co-authored-by: Amogh Desai <amoghrajesh1999@gmail.com>
kaxil pushed a commit that referenced this pull request Oct 31, 2025
kaxil pushed a commit that referenced this pull request Oct 31, 2025
@ephraimbuddyephraimbuddy added the type:bug-fix Changelog: Bug Fixes label Nov 10, 2025
odaneau-astro pushed a commit to odaneau-astro/airflow that referenced this pull request Mar 18, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:task-sdktype:bug-fixChangelog: Bug Fixes

Projects

No open projects

Development

Successfully merging this pull request may close these issues.

Task does not retry when worker is killed due to OOM

6 participants

@amoghrajesh@kaxil@ashb@rawwar@potiuk@ephraimbuddy
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Respect task retries for signal killed tasks by amoghrajesh · Pull Request #55767 · apache/airflow · GitHub
Skip to content

Respect task retries for signal killed tasks - #55767

Merged
amoghrajesh merged 6 commits into
apache:mainfrom
astronomer:handle-signal-task-better
Oct 28, 2025
Merged

Respect task retries for signal killed tasks#55767
amoghrajesh merged 6 commits into
apache:mainfrom
astronomer:handle-signal-task-better

Conversation

@amoghrajesh

@amoghrajeshamoghrajesh commented Sep 17, 2025

Copy link
Copy Markdown
Contributor

closes: #55753

Problem

When tasks are killed by system signals (SIGKILL for OOM, SIGTERM for worker restarts), they immediately go to FAILED state instead of respecting the task retries set and going to UP_FOR_RETRY state. This creates unexpected behavior where exception based failures respect retries but signal based failures don't.

Root Cause

The supervisor's final_state property only checked the exit code and didn't consider retry eligibility for signal based failures. While exception based failures properly checked should_retry which is set by the API server when a task is run for the first time a.k.a, in its run context, signal based failures ignored this logic entirely.

Testing

DAG used:

from airflow import DAG
from datetime import datetime
from airflow.providers.standard.operators.python import PythonOperator
def func():
a = "asd"
while True:
a += a*100000
with DAG(
dag_id="oom_example",
start_date=datetime(2024, 1, 1),
schedule=None,
catchup=False,
doc_md=__doc__,
tags=["oom"],
) as dag:
hello_task = PythonOperator(
task_id="oom_task",
python_callable=func,
retries=3
)

Earlier:

The dag immediately moved into FAILED state without any retries;

image

After:

The dag moves into up for retry state as appropriate:

image

See retries:

image

^ Add meaningful description above
Read the Pull Request Guidelines for more information.
In case of fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
In case of a new dependency, check compliance with the ASF 3rd Party License Policy.
In case of backwards incompatible changes please leave a note in a newsfragment file, named {pr_number}.significant.rst or {issue_number}.significant.rst, in airflow-core/newsfragments.

@ashbashb left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SIGABRT and SIGSEGV should not retry

Why? That's not what I would expect to do (and also I suspect not what Airflow 2 did?)

Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
@kaxil

kaxil commented Sep 17, 2025

Copy link
Copy Markdown
Member

Update: Nvm, thanks to @rawwar figured out that K8s/cgroups has configs to kill only the process


When tasks are killed by system signals (SIGKILL for OOM, SIGTERM for worker restarts), they immediately go to FAILED state instead of respecting the task retries set and going to UP_FOR_RETRY state. This creates unexpected behavior where exception based failures respect retries but signal based failures don't.

How common is this scenario (excluding manually killing task process)? Since the supervisor and task processes are running in the same container, wouldn't an OOM condition typically kill the entire container rather than just the individual task process?

In the more common case where the entire container gets OOM-killed:

  1. The supervisor process would also die
  2. Heartbeat to the scheduler would fail
  3. Scheduler would receive a FAILED executor event and handle retries through the normal process_executor_events()handle_failure() path

@ashb

ashb commented Sep 19, 2025

Copy link
Copy Markdown
Member

@kaxil Also people do occasionally run Airflow outside of Kubernetes you know 😉

@amoghrajesh

Copy link
Copy Markdown
ContributorAuthor

I'll get to the comments on this later, not needed for 3.1, can wait till 3.1.1

@amoghrajeshamoghrajesh added this to the Airflow 3.1.1 milestone Sep 19, 2025
@amoghrajesh

Copy link
Copy Markdown
ContributorAuthor

@ashb replied to some of your comments, could you take a look when possible?

Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
@rawwar

rawwar commented Oct 27, 2025

Copy link
Copy Markdown
Contributor

Verified by running a task that gets killed due to oom and the task went to "up_for_retry"

image

Just wanted to let it run all tries. It did fail at the end:
image

@rawwarrawwar left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

@amoghrajesh
amoghrajeshforce-pushed the handle-signal-task-better branch from 3191733 to 24ef66fCompareOctober 28, 2025 08:36
@amoghrajesh
amoghrajesh merged commit de0c78e into apache:mainOct 28, 2025
82 checks passed
@amoghrajesh
amoghrajesh deleted the handle-signal-task-better branch October 28, 2025 09:38
github-actionsBot pushed a commit that referenced this pull request Oct 28, 2025
(cherry picked from commit de0c78e)
Co-authored-by: Amogh Desai <amoghrajesh1999@gmail.com>
@github-actions

Copy link
Copy Markdown
Contributor

Backport successfully created: v3-1-test

StatusBranchResult
v3-1-testPR Link

github-actionsBot pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Oct 28, 2025
(cherry picked from commit de0c78e)
Co-authored-by: Amogh Desai <amoghrajesh1999@gmail.com>
kaxil pushed a commit that referenced this pull request Oct 31, 2025
kaxil pushed a commit that referenced this pull request Oct 31, 2025
@ephraimbuddyephraimbuddy added the type:bug-fix Changelog: Bug Fixes label Nov 10, 2025
odaneau-astro pushed a commit to odaneau-astro/airflow that referenced this pull request Mar 18, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:task-sdktype:bug-fixChangelog: Bug Fixes

Projects

No open projects

Development

Successfully merging this pull request may close these issues.

Task does not retry when worker is killed due to OOM

6 participants

@amoghrajesh@kaxil@ashb@rawwar@potiuk@ephraimbuddy
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Respect task retries for signal killed tasks by amoghrajesh · Pull Request #55767 · apache/airflow · GitHub
Skip to content

Respect task retries for signal killed tasks - #55767

Merged
amoghrajesh merged 6 commits into
apache:mainfrom
astronomer:handle-signal-task-better
Oct 28, 2025
Merged

Respect task retries for signal killed tasks#55767
amoghrajesh merged 6 commits into
apache:mainfrom
astronomer:handle-signal-task-better

Conversation

@amoghrajesh

@amoghrajeshamoghrajesh commented Sep 17, 2025

Copy link
Copy Markdown
Contributor

closes: #55753

Problem

When tasks are killed by system signals (SIGKILL for OOM, SIGTERM for worker restarts), they immediately go to FAILED state instead of respecting the task retries set and going to UP_FOR_RETRY state. This creates unexpected behavior where exception based failures respect retries but signal based failures don't.

Root Cause

The supervisor's final_state property only checked the exit code and didn't consider retry eligibility for signal based failures. While exception based failures properly checked should_retry which is set by the API server when a task is run for the first time a.k.a, in its run context, signal based failures ignored this logic entirely.

Testing

DAG used:

from airflow import DAG
from datetime import datetime
from airflow.providers.standard.operators.python import PythonOperator
def func():
a = "asd"
while True:
a += a*100000
with DAG(
dag_id="oom_example",
start_date=datetime(2024, 1, 1),
schedule=None,
catchup=False,
doc_md=__doc__,
tags=["oom"],
) as dag:
hello_task = PythonOperator(
task_id="oom_task",
python_callable=func,
retries=3
)

Earlier:

The dag immediately moved into FAILED state without any retries;

image

After:

The dag moves into up for retry state as appropriate:

image

See retries:

image

^ Add meaningful description above
Read the Pull Request Guidelines for more information.
In case of fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
In case of a new dependency, check compliance with the ASF 3rd Party License Policy.
In case of backwards incompatible changes please leave a note in a newsfragment file, named {pr_number}.significant.rst or {issue_number}.significant.rst, in airflow-core/newsfragments.

@ashbashb left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SIGABRT and SIGSEGV should not retry

Why? That's not what I would expect to do (and also I suspect not what Airflow 2 did?)

Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
@kaxil

kaxil commented Sep 17, 2025

Copy link
Copy Markdown
Member

Update: Nvm, thanks to @rawwar figured out that K8s/cgroups has configs to kill only the process


When tasks are killed by system signals (SIGKILL for OOM, SIGTERM for worker restarts), they immediately go to FAILED state instead of respecting the task retries set and going to UP_FOR_RETRY state. This creates unexpected behavior where exception based failures respect retries but signal based failures don't.

How common is this scenario (excluding manually killing task process)? Since the supervisor and task processes are running in the same container, wouldn't an OOM condition typically kill the entire container rather than just the individual task process?

In the more common case where the entire container gets OOM-killed:

  1. The supervisor process would also die
  2. Heartbeat to the scheduler would fail
  3. Scheduler would receive a FAILED executor event and handle retries through the normal process_executor_events()handle_failure() path

@ashb

ashb commented Sep 19, 2025

Copy link
Copy Markdown
Member

@kaxil Also people do occasionally run Airflow outside of Kubernetes you know 😉

@amoghrajesh

Copy link
Copy Markdown
ContributorAuthor

I'll get to the comments on this later, not needed for 3.1, can wait till 3.1.1

@amoghrajeshamoghrajesh added this to the Airflow 3.1.1 milestone Sep 19, 2025
@amoghrajesh

Copy link
Copy Markdown
ContributorAuthor

@ashb replied to some of your comments, could you take a look when possible?

Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
@rawwar

rawwar commented Oct 27, 2025

Copy link
Copy Markdown
Contributor

Verified by running a task that gets killed due to oom and the task went to "up_for_retry"

image

Just wanted to let it run all tries. It did fail at the end:
image

@rawwarrawwar left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

@amoghrajesh
amoghrajeshforce-pushed the handle-signal-task-better branch from 3191733 to 24ef66fCompareOctober 28, 2025 08:36
@amoghrajesh
amoghrajesh merged commit de0c78e into apache:mainOct 28, 2025
82 checks passed
@amoghrajesh
amoghrajesh deleted the handle-signal-task-better branch October 28, 2025 09:38
github-actionsBot pushed a commit that referenced this pull request Oct 28, 2025
(cherry picked from commit de0c78e)
Co-authored-by: Amogh Desai <amoghrajesh1999@gmail.com>
@github-actions

Copy link
Copy Markdown
Contributor

Backport successfully created: v3-1-test

StatusBranchResult
v3-1-testPR Link

github-actionsBot pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Oct 28, 2025
(cherry picked from commit de0c78e)
Co-authored-by: Amogh Desai <amoghrajesh1999@gmail.com>
kaxil pushed a commit that referenced this pull request Oct 31, 2025
kaxil pushed a commit that referenced this pull request Oct 31, 2025
@ephraimbuddyephraimbuddy added the type:bug-fix Changelog: Bug Fixes label Nov 10, 2025
odaneau-astro pushed a commit to odaneau-astro/airflow that referenced this pull request Mar 18, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:task-sdktype:bug-fixChangelog: Bug Fixes

Projects

No open projects

Development

Successfully merging this pull request may close these issues.

Task does not retry when worker is killed due to OOM

6 participants

@amoghrajesh@kaxil@ashb@rawwar@potiuk@ephraimbuddy
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); })(); Respect task retries for signal killed tasks by amoghrajesh · Pull Request #55767 · apache/airflow · GitHub
Skip to content

Respect task retries for signal killed tasks - #55767

Merged
amoghrajesh merged 6 commits into
apache:mainfrom
astronomer:handle-signal-task-better
Oct 28, 2025
Merged

Respect task retries for signal killed tasks#55767
amoghrajesh merged 6 commits into
apache:mainfrom
astronomer:handle-signal-task-better

Conversation

@amoghrajesh

@amoghrajeshamoghrajesh commented Sep 17, 2025

Copy link
Copy Markdown
Contributor

closes: #55753

Problem

When tasks are killed by system signals (SIGKILL for OOM, SIGTERM for worker restarts), they immediately go to FAILED state instead of respecting the task retries set and going to UP_FOR_RETRY state. This creates unexpected behavior where exception based failures respect retries but signal based failures don't.

Root Cause

The supervisor's final_state property only checked the exit code and didn't consider retry eligibility for signal based failures. While exception based failures properly checked should_retry which is set by the API server when a task is run for the first time a.k.a, in its run context, signal based failures ignored this logic entirely.

Testing

DAG used:

from airflow import DAG
from datetime import datetime
from airflow.providers.standard.operators.python import PythonOperator
def func():
a = "asd"
while True:
a += a*100000
with DAG(
dag_id="oom_example",
start_date=datetime(2024, 1, 1),
schedule=None,
catchup=False,
doc_md=__doc__,
tags=["oom"],
) as dag:
hello_task = PythonOperator(
task_id="oom_task",
python_callable=func,
retries=3
)

Earlier:

The dag immediately moved into FAILED state without any retries;

image

After:

The dag moves into up for retry state as appropriate:

image

See retries:

image

^ Add meaningful description above
Read the Pull Request Guidelines for more information.
In case of fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
In case of a new dependency, check compliance with the ASF 3rd Party License Policy.
In case of backwards incompatible changes please leave a note in a newsfragment file, named {pr_number}.significant.rst or {issue_number}.significant.rst, in airflow-core/newsfragments.

@ashbashb left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SIGABRT and SIGSEGV should not retry

Why? That's not what I would expect to do (and also I suspect not what Airflow 2 did?)

Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
@kaxil

kaxil commented Sep 17, 2025

Copy link
Copy Markdown
Member

Update: Nvm, thanks to @rawwar figured out that K8s/cgroups has configs to kill only the process


When tasks are killed by system signals (SIGKILL for OOM, SIGTERM for worker restarts), they immediately go to FAILED state instead of respecting the task retries set and going to UP_FOR_RETRY state. This creates unexpected behavior where exception based failures respect retries but signal based failures don't.

How common is this scenario (excluding manually killing task process)? Since the supervisor and task processes are running in the same container, wouldn't an OOM condition typically kill the entire container rather than just the individual task process?

In the more common case where the entire container gets OOM-killed:

  1. The supervisor process would also die
  2. Heartbeat to the scheduler would fail
  3. Scheduler would receive a FAILED executor event and handle retries through the normal process_executor_events()handle_failure() path

@ashb

ashb commented Sep 19, 2025

Copy link
Copy Markdown
Member

@kaxil Also people do occasionally run Airflow outside of Kubernetes you know 😉

@amoghrajesh

Copy link
Copy Markdown
ContributorAuthor

I'll get to the comments on this later, not needed for 3.1, can wait till 3.1.1

@amoghrajeshamoghrajesh added this to the Airflow 3.1.1 milestone Sep 19, 2025
@amoghrajesh

Copy link
Copy Markdown
ContributorAuthor

@ashb replied to some of your comments, could you take a look when possible?

Comment threadtask-sdk/src/airflow/sdk/execution_time/supervisor.py Outdated
@rawwar

rawwar commented Oct 27, 2025

Copy link
Copy Markdown
Contributor

Verified by running a task that gets killed due to oom and the task went to "up_for_retry"

image

Just wanted to let it run all tries. It did fail at the end:
image

@rawwarrawwar left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

@amoghrajesh
amoghrajeshforce-pushed the handle-signal-task-better branch from 3191733 to 24ef66fCompareOctober 28, 2025 08:36
@amoghrajesh
amoghrajesh merged commit de0c78e into apache:mainOct 28, 2025
82 checks passed
@amoghrajesh
amoghrajesh deleted the handle-signal-task-better branch October 28, 2025 09:38
github-actionsBot pushed a commit that referenced this pull request Oct 28, 2025
(cherry picked from commit de0c78e)
Co-authored-by: Amogh Desai <amoghrajesh1999@gmail.com>
@github-actions

Copy link
Copy Markdown
Contributor

Backport successfully created: v3-1-test

StatusBranchResult
v3-1-testPR Link

github-actionsBot pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Oct 28, 2025
(cherry picked from commit de0c78e)
Co-authored-by: Amogh Desai <amoghrajesh1999@gmail.com>
kaxil pushed a commit that referenced this pull request Oct 31, 2025
kaxil pushed a commit that referenced this pull request Oct 31, 2025
@ephraimbuddyephraimbuddy added the type:bug-fix Changelog: Bug Fixes label Nov 10, 2025
odaneau-astro pushed a commit to odaneau-astro/airflow that referenced this pull request Mar 18, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:task-sdktype:bug-fixChangelog: Bug Fixes

Projects

No open projects

Development

Successfully merging this pull request may close these issues.

Task does not retry when worker is killed due to OOM

6 participants

@amoghrajesh@kaxil@ashb@rawwar@potiuk@ephraimbuddy