Skip to content

Terminate pool when timeout is reached for parallel tests. - #53860

Merged
potiuk merged 1 commit into
apache:mainfrom
potiuk:better-timeout-handling
Aug 1, 2025
Merged

Terminate pool when timeout is reached for parallel tests.#53860
potiuk merged 1 commit into
apache:mainfrom
potiuk:better-timeout-handling

Conversation

@potiuk

@potiukpotiuk commented Jul 29, 2025

Copy link
Copy Markdown
Member

When we reach timeut we kill all the hanging containers already and after the pool has been terminated, we will print all the logs.

However, when the pool had not yet been fully executing (i.e the containers were hanging and some tasks were not started) - without terminating the pool that would kill running containers and the remaining tasks would start new ones.

This PR changes the timeout handler to terminate the pool before attempting to kill all the containers.

This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.

It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:

  • it was waiting in a loop to wait for all task to complete (which
    would never happen)

  • it was trying to retrieve result from ApplyResult without timeout
    where it would hang for ever for terminated tasks

This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.


^ Add meaningful description above
Read the Pull Request Guidelines for more information.
In case of fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
In case of a new dependency, check compliance with the ASF 3rd Party License Policy.
In case of backwards incompatible changes please leave a note in a newsfragment file, named {pr_number}.significant.rst or {issue_number}.significant.rst, in airflow-core/newsfragments.

@potiuk

Copy link
Copy Markdown
MemberAuthor

The origin of this PR - when trying to diagnode Sqlallchemy 2 CI #52233 it turned out that when things timed-out for all tests, the log output has not been printed .

@potiuk
potiukforce-pushed the better-timeout-handling branch from 85291ef to ff08b2bCompareJuly 29, 2025 07:52
@gopidesupavan

Copy link
Copy Markdown
Member

Nice Thanks for the update :) LGTM

aritra24

This comment was marked as resolved.

Comment threaddev/breeze/src/airflow_breeze/utils/parallel.py Outdated
@potiuk
potiukforce-pushed the better-timeout-handling branch 2 times, most recently from 0b167d6 to ad1d84aCompareJuly 29, 2025 11:49

@jscheffljscheffl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool!

@bugraoz93

Copy link
Copy Markdown
Contributor

Great!

When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
@potiuk
potiukforce-pushed the better-timeout-handling branch from ad1d84a to 0364bd0CompareAugust 1, 2025 11:21
@potiuk
potiuk merged commit e8d424e into apache:mainAug 1, 2025
103 checks passed
@potiuk
potiuk deleted the better-timeout-handling branch August 1, 2025 12:06
@github-actions

Copy link
Copy Markdown
Contributor

Backport failed to create: v3-0-test. View the failure log Run details

StatusBranchResult
v3-0-testCommit Link

You can attempt to backport this manually by running:

cherry_picker e8d424e v3-0-test

This should apply the commit to the v3-0-test branch and leave the commit in conflict state marking
the files that need manual conflict resolution.

After you have resolved the conflicts, you can continue the backport process by running:

cherry_picker --continue

ferruzzi pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Aug 7, 2025
)
When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
fweilun pushed a commit to fweilun/airflow that referenced this pull request Aug 11, 2025
)
When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@potiuk@gopidesupavan@bugraoz93@aritra24@jscheffl
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Terminate pool when timeout is reached for parallel tests. by potiuk · Pull Request #53860 · apache/airflow · GitHub
Skip to content

Terminate pool when timeout is reached for parallel tests. - #53860

Merged
potiuk merged 1 commit into
apache:mainfrom
potiuk:better-timeout-handling
Aug 1, 2025
Merged

Terminate pool when timeout is reached for parallel tests.#53860
potiuk merged 1 commit into
apache:mainfrom
potiuk:better-timeout-handling

Conversation

@potiuk

@potiukpotiuk commented Jul 29, 2025

Copy link
Copy Markdown
Member

When we reach timeut we kill all the hanging containers already and after the pool has been terminated, we will print all the logs.

However, when the pool had not yet been fully executing (i.e the containers were hanging and some tasks were not started) - without terminating the pool that would kill running containers and the remaining tasks would start new ones.

This PR changes the timeout handler to terminate the pool before attempting to kill all the containers.

This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.

It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:

  • it was waiting in a loop to wait for all task to complete (which
    would never happen)

  • it was trying to retrieve result from ApplyResult without timeout
    where it would hang for ever for terminated tasks

This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.


^ Add meaningful description above
Read the Pull Request Guidelines for more information.
In case of fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
In case of a new dependency, check compliance with the ASF 3rd Party License Policy.
In case of backwards incompatible changes please leave a note in a newsfragment file, named {pr_number}.significant.rst or {issue_number}.significant.rst, in airflow-core/newsfragments.

@potiuk

Copy link
Copy Markdown
MemberAuthor

The origin of this PR - when trying to diagnode Sqlallchemy 2 CI #52233 it turned out that when things timed-out for all tests, the log output has not been printed .

@potiuk
potiukforce-pushed the better-timeout-handling branch from 85291ef to ff08b2bCompareJuly 29, 2025 07:52
@gopidesupavan

Copy link
Copy Markdown
Member

Nice Thanks for the update :) LGTM

aritra24

This comment was marked as resolved.

Comment threaddev/breeze/src/airflow_breeze/utils/parallel.py Outdated
@potiuk
potiukforce-pushed the better-timeout-handling branch 2 times, most recently from 0b167d6 to ad1d84aCompareJuly 29, 2025 11:49

@jscheffljscheffl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool!

@bugraoz93

Copy link
Copy Markdown
Contributor

Great!

When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
@potiuk
potiukforce-pushed the better-timeout-handling branch from ad1d84a to 0364bd0CompareAugust 1, 2025 11:21
@potiuk
potiuk merged commit e8d424e into apache:mainAug 1, 2025
103 checks passed
@potiuk
potiuk deleted the better-timeout-handling branch August 1, 2025 12:06
@github-actions

Copy link
Copy Markdown
Contributor

Backport failed to create: v3-0-test. View the failure log Run details

StatusBranchResult
v3-0-testCommit Link

You can attempt to backport this manually by running:

cherry_picker e8d424e v3-0-test

This should apply the commit to the v3-0-test branch and leave the commit in conflict state marking
the files that need manual conflict resolution.

After you have resolved the conflicts, you can continue the backport process by running:

cherry_picker --continue

ferruzzi pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Aug 7, 2025
)
When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
fweilun pushed a commit to fweilun/airflow that referenced this pull request Aug 11, 2025
)
When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@potiuk@gopidesupavan@bugraoz93@aritra24@jscheffl
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Terminate pool when timeout is reached for parallel tests. by potiuk · Pull Request #53860 · apache/airflow · GitHub
Skip to content

Terminate pool when timeout is reached for parallel tests. - #53860

Merged
potiuk merged 1 commit into
apache:mainfrom
potiuk:better-timeout-handling
Aug 1, 2025
Merged

Terminate pool when timeout is reached for parallel tests.#53860
potiuk merged 1 commit into
apache:mainfrom
potiuk:better-timeout-handling

Conversation

@potiuk

@potiukpotiuk commented Jul 29, 2025

Copy link
Copy Markdown
Member

When we reach timeut we kill all the hanging containers already and after the pool has been terminated, we will print all the logs.

However, when the pool had not yet been fully executing (i.e the containers were hanging and some tasks were not started) - without terminating the pool that would kill running containers and the remaining tasks would start new ones.

This PR changes the timeout handler to terminate the pool before attempting to kill all the containers.

This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.

It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:

  • it was waiting in a loop to wait for all task to complete (which
    would never happen)

  • it was trying to retrieve result from ApplyResult without timeout
    where it would hang for ever for terminated tasks

This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.


^ Add meaningful description above
Read the Pull Request Guidelines for more information.
In case of fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
In case of a new dependency, check compliance with the ASF 3rd Party License Policy.
In case of backwards incompatible changes please leave a note in a newsfragment file, named {pr_number}.significant.rst or {issue_number}.significant.rst, in airflow-core/newsfragments.

@potiuk

Copy link
Copy Markdown
MemberAuthor

The origin of this PR - when trying to diagnode Sqlallchemy 2 CI #52233 it turned out that when things timed-out for all tests, the log output has not been printed .

@potiuk
potiukforce-pushed the better-timeout-handling branch from 85291ef to ff08b2bCompareJuly 29, 2025 07:52
@gopidesupavan

Copy link
Copy Markdown
Member

Nice Thanks for the update :) LGTM

aritra24

This comment was marked as resolved.

Comment threaddev/breeze/src/airflow_breeze/utils/parallel.py Outdated
@potiuk
potiukforce-pushed the better-timeout-handling branch 2 times, most recently from 0b167d6 to ad1d84aCompareJuly 29, 2025 11:49

@jscheffljscheffl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool!

@bugraoz93

Copy link
Copy Markdown
Contributor

Great!

When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
@potiuk
potiukforce-pushed the better-timeout-handling branch from ad1d84a to 0364bd0CompareAugust 1, 2025 11:21
@potiuk
potiuk merged commit e8d424e into apache:mainAug 1, 2025
103 checks passed
@potiuk
potiuk deleted the better-timeout-handling branch August 1, 2025 12:06
@github-actions

Copy link
Copy Markdown
Contributor

Backport failed to create: v3-0-test. View the failure log Run details

StatusBranchResult
v3-0-testCommit Link

You can attempt to backport this manually by running:

cherry_picker e8d424e v3-0-test

This should apply the commit to the v3-0-test branch and leave the commit in conflict state marking
the files that need manual conflict resolution.

After you have resolved the conflicts, you can continue the backport process by running:

cherry_picker --continue

ferruzzi pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Aug 7, 2025
)
When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
fweilun pushed a commit to fweilun/airflow that referenced this pull request Aug 11, 2025
)
When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@potiuk@gopidesupavan@bugraoz93@aritra24@jscheffl
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Terminate pool when timeout is reached for parallel tests. by potiuk · Pull Request #53860 · apache/airflow · GitHub
Skip to content

Terminate pool when timeout is reached for parallel tests. - #53860

Merged
potiuk merged 1 commit into
apache:mainfrom
potiuk:better-timeout-handling
Aug 1, 2025
Merged

Terminate pool when timeout is reached for parallel tests.#53860
potiuk merged 1 commit into
apache:mainfrom
potiuk:better-timeout-handling

Conversation

@potiuk

@potiukpotiuk commented Jul 29, 2025

Copy link
Copy Markdown
Member

When we reach timeut we kill all the hanging containers already and after the pool has been terminated, we will print all the logs.

However, when the pool had not yet been fully executing (i.e the containers were hanging and some tasks were not started) - without terminating the pool that would kill running containers and the remaining tasks would start new ones.

This PR changes the timeout handler to terminate the pool before attempting to kill all the containers.

This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.

It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:

  • it was waiting in a loop to wait for all task to complete (which
    would never happen)

  • it was trying to retrieve result from ApplyResult without timeout
    where it would hang for ever for terminated tasks

This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.


^ Add meaningful description above
Read the Pull Request Guidelines for more information.
In case of fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
In case of a new dependency, check compliance with the ASF 3rd Party License Policy.
In case of backwards incompatible changes please leave a note in a newsfragment file, named {pr_number}.significant.rst or {issue_number}.significant.rst, in airflow-core/newsfragments.

@potiuk

Copy link
Copy Markdown
MemberAuthor

The origin of this PR - when trying to diagnode Sqlallchemy 2 CI #52233 it turned out that when things timed-out for all tests, the log output has not been printed .

@potiuk
potiukforce-pushed the better-timeout-handling branch from 85291ef to ff08b2bCompareJuly 29, 2025 07:52
@gopidesupavan

Copy link
Copy Markdown
Member

Nice Thanks for the update :) LGTM

aritra24

This comment was marked as resolved.

Comment threaddev/breeze/src/airflow_breeze/utils/parallel.py Outdated
@potiuk
potiukforce-pushed the better-timeout-handling branch 2 times, most recently from 0b167d6 to ad1d84aCompareJuly 29, 2025 11:49

@jscheffljscheffl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool!

@bugraoz93

Copy link
Copy Markdown
Contributor

Great!

When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
@potiuk
potiukforce-pushed the better-timeout-handling branch from ad1d84a to 0364bd0CompareAugust 1, 2025 11:21
@potiuk
potiuk merged commit e8d424e into apache:mainAug 1, 2025
103 checks passed
@potiuk
potiuk deleted the better-timeout-handling branch August 1, 2025 12:06
@github-actions

Copy link
Copy Markdown
Contributor

Backport failed to create: v3-0-test. View the failure log Run details

StatusBranchResult
v3-0-testCommit Link

You can attempt to backport this manually by running:

cherry_picker e8d424e v3-0-test

This should apply the commit to the v3-0-test branch and leave the commit in conflict state marking
the files that need manual conflict resolution.

After you have resolved the conflicts, you can continue the backport process by running:

cherry_picker --continue

ferruzzi pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Aug 7, 2025
)
When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
fweilun pushed a commit to fweilun/airflow that referenced this pull request Aug 11, 2025
)
When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@potiuk@gopidesupavan@bugraoz93@aritra24@jscheffl
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Terminate pool when timeout is reached for parallel tests. by potiuk · Pull Request #53860 · apache/airflow · GitHub
Skip to content

Terminate pool when timeout is reached for parallel tests. - #53860

Merged
potiuk merged 1 commit into
apache:mainfrom
potiuk:better-timeout-handling
Aug 1, 2025
Merged

Terminate pool when timeout is reached for parallel tests.#53860
potiuk merged 1 commit into
apache:mainfrom
potiuk:better-timeout-handling

Conversation

@potiuk

@potiukpotiuk commented Jul 29, 2025

Copy link
Copy Markdown
Member

When we reach timeut we kill all the hanging containers already and after the pool has been terminated, we will print all the logs.

However, when the pool had not yet been fully executing (i.e the containers were hanging and some tasks were not started) - without terminating the pool that would kill running containers and the remaining tasks would start new ones.

This PR changes the timeout handler to terminate the pool before attempting to kill all the containers.

This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.

It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:

  • it was waiting in a loop to wait for all task to complete (which
    would never happen)

  • it was trying to retrieve result from ApplyResult without timeout
    where it would hang for ever for terminated tasks

This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.


^ Add meaningful description above
Read the Pull Request Guidelines for more information.
In case of fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
In case of a new dependency, check compliance with the ASF 3rd Party License Policy.
In case of backwards incompatible changes please leave a note in a newsfragment file, named {pr_number}.significant.rst or {issue_number}.significant.rst, in airflow-core/newsfragments.

@potiuk

Copy link
Copy Markdown
MemberAuthor

The origin of this PR - when trying to diagnode Sqlallchemy 2 CI #52233 it turned out that when things timed-out for all tests, the log output has not been printed .

@potiuk
potiukforce-pushed the better-timeout-handling branch from 85291ef to ff08b2bCompareJuly 29, 2025 07:52
@gopidesupavan

Copy link
Copy Markdown
Member

Nice Thanks for the update :) LGTM

aritra24

This comment was marked as resolved.

Comment threaddev/breeze/src/airflow_breeze/utils/parallel.py Outdated
@potiuk
potiukforce-pushed the better-timeout-handling branch 2 times, most recently from 0b167d6 to ad1d84aCompareJuly 29, 2025 11:49

@jscheffljscheffl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool!

@bugraoz93

Copy link
Copy Markdown
Contributor

Great!

When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
@potiuk
potiukforce-pushed the better-timeout-handling branch from ad1d84a to 0364bd0CompareAugust 1, 2025 11:21
@potiuk
potiuk merged commit e8d424e into apache:mainAug 1, 2025
103 checks passed
@potiuk
potiuk deleted the better-timeout-handling branch August 1, 2025 12:06
@github-actions

Copy link
Copy Markdown
Contributor

Backport failed to create: v3-0-test. View the failure log Run details

StatusBranchResult
v3-0-testCommit Link

You can attempt to backport this manually by running:

cherry_picker e8d424e v3-0-test

This should apply the commit to the v3-0-test branch and leave the commit in conflict state marking
the files that need manual conflict resolution.

After you have resolved the conflicts, you can continue the backport process by running:

cherry_picker --continue

ferruzzi pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Aug 7, 2025
)
When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
fweilun pushed a commit to fweilun/airflow that referenced this pull request Aug 11, 2025
)
When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@potiuk@gopidesupavan@bugraoz93@aritra24@jscheffl
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Terminate pool when timeout is reached for parallel tests. by potiuk · Pull Request #53860 · apache/airflow · GitHub
Skip to content

Terminate pool when timeout is reached for parallel tests. - #53860

Merged
potiuk merged 1 commit into
apache:mainfrom
potiuk:better-timeout-handling
Aug 1, 2025
Merged

Terminate pool when timeout is reached for parallel tests.#53860
potiuk merged 1 commit into
apache:mainfrom
potiuk:better-timeout-handling

Conversation

@potiuk

@potiukpotiuk commented Jul 29, 2025

Copy link
Copy Markdown
Member

When we reach timeut we kill all the hanging containers already and after the pool has been terminated, we will print all the logs.

However, when the pool had not yet been fully executing (i.e the containers were hanging and some tasks were not started) - without terminating the pool that would kill running containers and the remaining tasks would start new ones.

This PR changes the timeout handler to terminate the pool before attempting to kill all the containers.

This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.

It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:

  • it was waiting in a loop to wait for all task to complete (which
    would never happen)

  • it was trying to retrieve result from ApplyResult without timeout
    where it would hang for ever for terminated tasks

This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.


^ Add meaningful description above
Read the Pull Request Guidelines for more information.
In case of fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
In case of a new dependency, check compliance with the ASF 3rd Party License Policy.
In case of backwards incompatible changes please leave a note in a newsfragment file, named {pr_number}.significant.rst or {issue_number}.significant.rst, in airflow-core/newsfragments.

@potiuk

Copy link
Copy Markdown
MemberAuthor

The origin of this PR - when trying to diagnode Sqlallchemy 2 CI #52233 it turned out that when things timed-out for all tests, the log output has not been printed .

@potiuk
potiukforce-pushed the better-timeout-handling branch from 85291ef to ff08b2bCompareJuly 29, 2025 07:52
@gopidesupavan

Copy link
Copy Markdown
Member

Nice Thanks for the update :) LGTM

aritra24

This comment was marked as resolved.

Comment threaddev/breeze/src/airflow_breeze/utils/parallel.py Outdated
@potiuk
potiukforce-pushed the better-timeout-handling branch 2 times, most recently from 0b167d6 to ad1d84aCompareJuly 29, 2025 11:49

@jscheffljscheffl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool!

@bugraoz93

Copy link
Copy Markdown
Contributor

Great!

When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
@potiuk
potiukforce-pushed the better-timeout-handling branch from ad1d84a to 0364bd0CompareAugust 1, 2025 11:21
@potiuk
potiuk merged commit e8d424e into apache:mainAug 1, 2025
103 checks passed
@potiuk
potiuk deleted the better-timeout-handling branch August 1, 2025 12:06
@github-actions

Copy link
Copy Markdown
Contributor

Backport failed to create: v3-0-test. View the failure log Run details

StatusBranchResult
v3-0-testCommit Link

You can attempt to backport this manually by running:

cherry_picker e8d424e v3-0-test

This should apply the commit to the v3-0-test branch and leave the commit in conflict state marking
the files that need manual conflict resolution.

After you have resolved the conflicts, you can continue the backport process by running:

cherry_picker --continue

ferruzzi pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Aug 7, 2025
)
When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
fweilun pushed a commit to fweilun/airflow that referenced this pull request Aug 11, 2025
)
When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@potiuk@gopidesupavan@bugraoz93@aritra24@jscheffl
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Terminate pool when timeout is reached for parallel tests. by potiuk · Pull Request #53860 · apache/airflow · GitHub
Skip to content

Terminate pool when timeout is reached for parallel tests. - #53860

Merged
potiuk merged 1 commit into
apache:mainfrom
potiuk:better-timeout-handling
Aug 1, 2025
Merged

Terminate pool when timeout is reached for parallel tests.#53860
potiuk merged 1 commit into
apache:mainfrom
potiuk:better-timeout-handling

Conversation

@potiuk

@potiukpotiuk commented Jul 29, 2025

Copy link
Copy Markdown
Member

When we reach timeut we kill all the hanging containers already and after the pool has been terminated, we will print all the logs.

However, when the pool had not yet been fully executing (i.e the containers were hanging and some tasks were not started) - without terminating the pool that would kill running containers and the remaining tasks would start new ones.

This PR changes the timeout handler to terminate the pool before attempting to kill all the containers.

This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.

It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:

  • it was waiting in a loop to wait for all task to complete (which
    would never happen)

  • it was trying to retrieve result from ApplyResult without timeout
    where it would hang for ever for terminated tasks

This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.


^ Add meaningful description above
Read the Pull Request Guidelines for more information.
In case of fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
In case of a new dependency, check compliance with the ASF 3rd Party License Policy.
In case of backwards incompatible changes please leave a note in a newsfragment file, named {pr_number}.significant.rst or {issue_number}.significant.rst, in airflow-core/newsfragments.

@potiuk

Copy link
Copy Markdown
MemberAuthor

The origin of this PR - when trying to diagnode Sqlallchemy 2 CI #52233 it turned out that when things timed-out for all tests, the log output has not been printed .

@potiuk
potiukforce-pushed the better-timeout-handling branch from 85291ef to ff08b2bCompareJuly 29, 2025 07:52
@gopidesupavan

Copy link
Copy Markdown
Member

Nice Thanks for the update :) LGTM

aritra24

This comment was marked as resolved.

Comment threaddev/breeze/src/airflow_breeze/utils/parallel.py Outdated
@potiuk
potiukforce-pushed the better-timeout-handling branch 2 times, most recently from 0b167d6 to ad1d84aCompareJuly 29, 2025 11:49

@jscheffljscheffl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool!

@bugraoz93

Copy link
Copy Markdown
Contributor

Great!

When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
@potiuk
potiukforce-pushed the better-timeout-handling branch from ad1d84a to 0364bd0CompareAugust 1, 2025 11:21
@potiuk
potiuk merged commit e8d424e into apache:mainAug 1, 2025
103 checks passed
@potiuk
potiuk deleted the better-timeout-handling branch August 1, 2025 12:06
@github-actions

Copy link
Copy Markdown
Contributor

Backport failed to create: v3-0-test. View the failure log Run details

StatusBranchResult
v3-0-testCommit Link

You can attempt to backport this manually by running:

cherry_picker e8d424e v3-0-test

This should apply the commit to the v3-0-test branch and leave the commit in conflict state marking
the files that need manual conflict resolution.

After you have resolved the conflicts, you can continue the backport process by running:

cherry_picker --continue

ferruzzi pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Aug 7, 2025
)
When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
fweilun pushed a commit to fweilun/airflow that referenced this pull request Aug 11, 2025
)
When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@potiuk@gopidesupavan@bugraoz93@aritra24@jscheffl
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Terminate pool when timeout is reached for parallel tests. by potiuk · Pull Request #53860 · apache/airflow · GitHub
Skip to content

Terminate pool when timeout is reached for parallel tests. - #53860

Merged
potiuk merged 1 commit into
apache:mainfrom
potiuk:better-timeout-handling
Aug 1, 2025
Merged

Terminate pool when timeout is reached for parallel tests.#53860
potiuk merged 1 commit into
apache:mainfrom
potiuk:better-timeout-handling

Conversation

@potiuk

@potiukpotiuk commented Jul 29, 2025

Copy link
Copy Markdown
Member

When we reach timeut we kill all the hanging containers already and after the pool has been terminated, we will print all the logs.

However, when the pool had not yet been fully executing (i.e the containers were hanging and some tasks were not started) - without terminating the pool that would kill running containers and the remaining tasks would start new ones.

This PR changes the timeout handler to terminate the pool before attempting to kill all the containers.

This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.

It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:

  • it was waiting in a loop to wait for all task to complete (which
    would never happen)

  • it was trying to retrieve result from ApplyResult without timeout
    where it would hang for ever for terminated tasks

This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.


^ Add meaningful description above
Read the Pull Request Guidelines for more information.
In case of fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
In case of a new dependency, check compliance with the ASF 3rd Party License Policy.
In case of backwards incompatible changes please leave a note in a newsfragment file, named {pr_number}.significant.rst or {issue_number}.significant.rst, in airflow-core/newsfragments.

@potiuk

Copy link
Copy Markdown
MemberAuthor

The origin of this PR - when trying to diagnode Sqlallchemy 2 CI #52233 it turned out that when things timed-out for all tests, the log output has not been printed .

@potiuk
potiukforce-pushed the better-timeout-handling branch from 85291ef to ff08b2bCompareJuly 29, 2025 07:52
@gopidesupavan

Copy link
Copy Markdown
Member

Nice Thanks for the update :) LGTM

aritra24

This comment was marked as resolved.

Comment threaddev/breeze/src/airflow_breeze/utils/parallel.py Outdated
@potiuk
potiukforce-pushed the better-timeout-handling branch 2 times, most recently from 0b167d6 to ad1d84aCompareJuly 29, 2025 11:49

@jscheffljscheffl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cool!

@bugraoz93

Copy link
Copy Markdown
Contributor

Great!

When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
@potiuk
potiukforce-pushed the better-timeout-handling branch from ad1d84a to 0364bd0CompareAugust 1, 2025 11:21
@potiuk
potiuk merged commit e8d424e into apache:mainAug 1, 2025
103 checks passed
@potiuk
potiuk deleted the better-timeout-handling branch August 1, 2025 12:06
@github-actions

Copy link
Copy Markdown
Contributor

Backport failed to create: v3-0-test. View the failure log Run details

StatusBranchResult
v3-0-testCommit Link

You can attempt to backport this manually by running:

cherry_picker e8d424e v3-0-test

This should apply the commit to the v3-0-test branch and leave the commit in conflict state marking
the files that need manual conflict resolution.

After you have resolved the conflicts, you can continue the backport process by running:

cherry_picker --continue

ferruzzi pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Aug 7, 2025
)
When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
fweilun pushed a commit to fweilun/airflow that referenced this pull request Aug 11, 2025
)
When we reach timeut we kill all the hanging containers already
and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the
containers were hanging and some tasks were not started) - without
terminating the pool that would kill running containers and the
remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
* it was waiting in a loop to wait for all task to complete (which
would never happen)
* it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@potiuk@gopidesupavan@bugraoz93@aritra24@jscheffl