Uh oh!
There was an error while loading. Please reload this page.
Terminate pool when timeout is reached for parallel tests. - #53860
Conversation
potiuk
commented
Jul 29, 2025
The origin of this PR - when trying to diagnode Sqlallchemy 2 CI #52233 it turned out that when things timed-out for all tests, the log output has not been printed . |
85291ef to
ff08b2bComparegopidesupavan
commented
Jul 29, 2025
Nice Thanks for the update :) LGTM |
This comment was marked as resolved.
This comment was marked as resolved.
Sorry, something went wrong.
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
0b167d6 to
ad1d84aComparebugraoz93
commented
Jul 30, 2025
Great! |
When we reach timeut we kill all the hanging containers already and after the pool has been terminated, we will print all the logs. However, when the pool had not yet been fully executing (i.e the containers were hanging and some tasks were not started) - without terminating the pool that would kill running containers and the remaining tasks would start new ones. This PR changes the timeout handler to terminate the pool before attempting to kill all the containers. It also turned out that exit handling by the main thread monitorint the tests in this case would hang rather than print logs: * it was waiting in a loop to wait for all task to complete (which would never happen) * it was trying to retrieve result from ApplyResult without timeout where it would hang for ever for terminated tasks This PR introduces a separate path to handle timeout, which does not wait for those two and handles timeout immediately. It also refactors the whole "end of tests" method splitting it into several methods to make it easier to reason and read.
ad1d84a to
0364bd0CompareUh oh!
There was an error while loading. Please reload this page.
Backport failed to create: v3-0-test. View the failure log Run details
You can attempt to backport this manually by running: cherry_picker e8d424e v3-0-testThis should apply the commit to the v3-0-test branch and leave the commit in conflict state marking After you have resolved the conflicts, you can continue the backport process by running: cherry_picker --continue |
) When we reach timeut we kill all the hanging containers already and after the pool has been terminated, we will print all the logs. However, when the pool had not yet been fully executing (i.e the containers were hanging and some tasks were not started) - without terminating the pool that would kill running containers and the remaining tasks would start new ones. This PR changes the timeout handler to terminate the pool before attempting to kill all the containers. It also turned out that exit handling by the main thread monitorint the tests in this case would hang rather than print logs: * it was waiting in a loop to wait for all task to complete (which would never happen) * it was trying to retrieve result from ApplyResult without timeout where it would hang for ever for terminated tasks This PR introduces a separate path to handle timeout, which does not wait for those two and handles timeout immediately. It also refactors the whole "end of tests" method splitting it into several methods to make it easier to reason and read.
) When we reach timeut we kill all the hanging containers already and after the pool has been terminated, we will print all the logs. However, when the pool had not yet been fully executing (i.e the containers were hanging and some tasks were not started) - without terminating the pool that would kill running containers and the remaining tasks would start new ones. This PR changes the timeout handler to terminate the pool before attempting to kill all the containers. It also turned out that exit handling by the main thread monitorint the tests in this case would hang rather than print logs: * it was waiting in a loop to wait for all task to complete (which would never happen) * it was trying to retrieve result from ApplyResult without timeout where it would hang for ever for terminated tasks This PR introduces a separate path to handle timeout, which does not wait for those two and handles timeout immediately. It also refactors the whole "end of tests" method splitting it into several methods to make it easier to reason and read.
When we reach timeut we kill all the hanging containers already and after the pool has been terminated, we will print all the logs.
However, when the pool had not yet been fully executing (i.e the containers were hanging and some tasks were not started) - without terminating the pool that would kill running containers and the remaining tasks would start new ones.
This PR changes the timeout handler to terminate the pool before attempting to kill all the containers.
This PR changes the timeout handler to terminate the pool before
attempting to kill all the containers.
It also turned out that exit handling by the main thread monitorint
the tests in this case would hang rather than print logs:
it was waiting in a loop to wait for all task to complete (which
would never happen)
it was trying to retrieve result from ApplyResult without timeout
where it would hang for ever for terminated tasks
This PR introduces a separate path to handle timeout, which does
not wait for those two and handles timeout immediately. It also
refactors the whole "end of tests" method splitting it into several
methods to make it easier to reason and read.
^ Add meaningful description above
Read the Pull Request Guidelines for more information.
In case of fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
In case of a new dependency, check compliance with the ASF 3rd Party License Policy.
In case of backwards incompatible changes please leave a note in a newsfragment file, named
{pr_number}.significant.rstor{issue_number}.significant.rst, in airflow-core/newsfragments.