Rerun flaky SSHRemoteJobOperator kill test on process-group races - #69384

Merged
potiuk merged 1 commit into
apache:mainfrom
shahar1:fix-flaky-ssh-remote-job-kill-test
Jul 5, 2026
Merged

Rerun flaky SSHRemoteJobOperator kill test on process-group races#69384
potiuk merged 1 commit into
apache:mainfrom
shahar1:fix-flaky-ssh-remote-job-kill-test

Conversation

@shahar1

Copy link
Copy Markdown
Contributor

TestPosixKillBehaviour::test_kill_terminates_whole_job_tree (added in #68644) intermittently fails its pre-kill assertion on CI:

AssertionError: job tree should be running before kill
assert False = _group_alive(326)

The wrapper records $! expecting it to equal the job PGID, but setsid only skips forking when the launching shell is not a process-group leader. On some runners it forks, so $! is the short-lived setsid parent and the job's real process group is already empty by the time pgrep -g runs — failing the check before the kill is even exercised. Passes locally (15/15, and under xdist), fails only on certain CI runners, so it's environment-dependent and lands on unrelated PRs.

Marks the test with @pytest.mark.flaky(reruns=5) (the convention already used elsewhere in this provider) so a fresh re-launch clears the race.

Note: if setsid forks in a real SSH session too, $! would be the wrong PID there and on_kill could orphan the job — a possible latent issue in the wrapper itself, left for a separate change.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 4.8)

Generated-by: Claude Code (Opus 4.8) following the guidelines

test_kill_terminates_whole_job_tree intermittently fails its pre-kill
check on CI: the wrapper records $! expecting it to equal the job PGID,
but setsid only skips forking when the launching shell is not a
process-group leader. On some runners it forks, so $! is the short-lived
setsid parent and the job's real group is empty by the time pgrep -g
runs. Rerun on a fresh draw so the environment-dependent race does not
fail unrelated PRs.
@potiuk
potiuk merged commit ea10f4f into apache:mainJul 5, 2026
81 checks passed
@shahar1
shahar1 deleted the fix-flaky-ssh-remote-job-kill-test branch July 5, 2026 20:35
potiuk added a commit to potiuk/airflow that referenced this pull request Jul 27, 2026
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - apache#69384 had them on the
sibling test until apache#69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
potiuk added a commit that referenced this pull request Jul 27, 2026
…70562)
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - #69384 had them on the
sibling test until #69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
dabla pushed a commit to dabla/airflow that referenced this pull request Aug 14, 2026
…pache#70562)
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - apache#69384 had them on the
sibling test until apache#69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@shahar1@potiuk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Rerun flaky SSHRemoteJobOperator kill test on process-group races - #69384

Merged
potiuk merged 1 commit into
apache:mainfrom
shahar1:fix-flaky-ssh-remote-job-kill-test
Jul 5, 2026
Merged

Rerun flaky SSHRemoteJobOperator kill test on process-group races#69384
potiuk merged 1 commit into
apache:mainfrom
shahar1:fix-flaky-ssh-remote-job-kill-test

Conversation

@shahar1

Copy link
Copy Markdown
Contributor

TestPosixKillBehaviour::test_kill_terminates_whole_job_tree (added in #68644) intermittently fails its pre-kill assertion on CI:

AssertionError: job tree should be running before kill
assert False = _group_alive(326)

The wrapper records $! expecting it to equal the job PGID, but setsid only skips forking when the launching shell is not a process-group leader. On some runners it forks, so $! is the short-lived setsid parent and the job's real process group is already empty by the time pgrep -g runs — failing the check before the kill is even exercised. Passes locally (15/15, and under xdist), fails only on certain CI runners, so it's environment-dependent and lands on unrelated PRs.

Marks the test with @pytest.mark.flaky(reruns=5) (the convention already used elsewhere in this provider) so a fresh re-launch clears the race.

Note: if setsid forks in a real SSH session too, $! would be the wrong PID there and on_kill could orphan the job — a possible latent issue in the wrapper itself, left for a separate change.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 4.8)

Generated-by: Claude Code (Opus 4.8) following the guidelines

test_kill_terminates_whole_job_tree intermittently fails its pre-kill
check on CI: the wrapper records $! expecting it to equal the job PGID,
but setsid only skips forking when the launching shell is not a
process-group leader. On some runners it forks, so $! is the short-lived
setsid parent and the job's real group is empty by the time pgrep -g
runs. Rerun on a fresh draw so the environment-dependent race does not
fail unrelated PRs.
@potiuk
potiuk merged commit ea10f4f into apache:mainJul 5, 2026
81 checks passed
@shahar1
shahar1 deleted the fix-flaky-ssh-remote-job-kill-test branch July 5, 2026 20:35
potiuk added a commit to potiuk/airflow that referenced this pull request Jul 27, 2026
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - apache#69384 had them on the
sibling test until apache#69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
potiuk added a commit that referenced this pull request Jul 27, 2026
…70562)
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - #69384 had them on the
sibling test until #69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
dabla pushed a commit to dabla/airflow that referenced this pull request Aug 14, 2026
…pache#70562)
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - apache#69384 had them on the
sibling test until apache#69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@shahar1@potiuk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Rerun flaky SSHRemoteJobOperator kill test on process-group races - #69384

Merged
potiuk merged 1 commit into
apache:mainfrom
shahar1:fix-flaky-ssh-remote-job-kill-test
Jul 5, 2026
Merged

Rerun flaky SSHRemoteJobOperator kill test on process-group races#69384
potiuk merged 1 commit into
apache:mainfrom
shahar1:fix-flaky-ssh-remote-job-kill-test

Conversation

@shahar1

Copy link
Copy Markdown
Contributor

TestPosixKillBehaviour::test_kill_terminates_whole_job_tree (added in #68644) intermittently fails its pre-kill assertion on CI:

AssertionError: job tree should be running before kill
assert False = _group_alive(326)

The wrapper records $! expecting it to equal the job PGID, but setsid only skips forking when the launching shell is not a process-group leader. On some runners it forks, so $! is the short-lived setsid parent and the job's real process group is already empty by the time pgrep -g runs — failing the check before the kill is even exercised. Passes locally (15/15, and under xdist), fails only on certain CI runners, so it's environment-dependent and lands on unrelated PRs.

Marks the test with @pytest.mark.flaky(reruns=5) (the convention already used elsewhere in this provider) so a fresh re-launch clears the race.

Note: if setsid forks in a real SSH session too, $! would be the wrong PID there and on_kill could orphan the job — a possible latent issue in the wrapper itself, left for a separate change.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 4.8)

Generated-by: Claude Code (Opus 4.8) following the guidelines

test_kill_terminates_whole_job_tree intermittently fails its pre-kill
check on CI: the wrapper records $! expecting it to equal the job PGID,
but setsid only skips forking when the launching shell is not a
process-group leader. On some runners it forks, so $! is the short-lived
setsid parent and the job's real group is empty by the time pgrep -g
runs. Rerun on a fresh draw so the environment-dependent race does not
fail unrelated PRs.
@potiuk
potiuk merged commit ea10f4f into apache:mainJul 5, 2026
81 checks passed
@shahar1
shahar1 deleted the fix-flaky-ssh-remote-job-kill-test branch July 5, 2026 20:35
potiuk added a commit to potiuk/airflow that referenced this pull request Jul 27, 2026
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - apache#69384 had them on the
sibling test until apache#69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
potiuk added a commit that referenced this pull request Jul 27, 2026
…70562)
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - #69384 had them on the
sibling test until #69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
dabla pushed a commit to dabla/airflow that referenced this pull request Aug 14, 2026
…pache#70562)
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - apache#69384 had them on the
sibling test until apache#69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@shahar1@potiuk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Rerun flaky SSHRemoteJobOperator kill test on process-group races - #69384

Merged
potiuk merged 1 commit into
apache:mainfrom
shahar1:fix-flaky-ssh-remote-job-kill-test
Jul 5, 2026
Merged

Rerun flaky SSHRemoteJobOperator kill test on process-group races#69384
potiuk merged 1 commit into
apache:mainfrom
shahar1:fix-flaky-ssh-remote-job-kill-test

Conversation

@shahar1

Copy link
Copy Markdown
Contributor

TestPosixKillBehaviour::test_kill_terminates_whole_job_tree (added in #68644) intermittently fails its pre-kill assertion on CI:

AssertionError: job tree should be running before kill
assert False = _group_alive(326)

The wrapper records $! expecting it to equal the job PGID, but setsid only skips forking when the launching shell is not a process-group leader. On some runners it forks, so $! is the short-lived setsid parent and the job's real process group is already empty by the time pgrep -g runs — failing the check before the kill is even exercised. Passes locally (15/15, and under xdist), fails only on certain CI runners, so it's environment-dependent and lands on unrelated PRs.

Marks the test with @pytest.mark.flaky(reruns=5) (the convention already used elsewhere in this provider) so a fresh re-launch clears the race.

Note: if setsid forks in a real SSH session too, $! would be the wrong PID there and on_kill could orphan the job — a possible latent issue in the wrapper itself, left for a separate change.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 4.8)

Generated-by: Claude Code (Opus 4.8) following the guidelines

test_kill_terminates_whole_job_tree intermittently fails its pre-kill
check on CI: the wrapper records $! expecting it to equal the job PGID,
but setsid only skips forking when the launching shell is not a
process-group leader. On some runners it forks, so $! is the short-lived
setsid parent and the job's real group is empty by the time pgrep -g
runs. Rerun on a fresh draw so the environment-dependent race does not
fail unrelated PRs.
@potiuk
potiuk merged commit ea10f4f into apache:mainJul 5, 2026
81 checks passed
@shahar1
shahar1 deleted the fix-flaky-ssh-remote-job-kill-test branch July 5, 2026 20:35
potiuk added a commit to potiuk/airflow that referenced this pull request Jul 27, 2026
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - apache#69384 had them on the
sibling test until apache#69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
potiuk added a commit that referenced this pull request Jul 27, 2026
…70562)
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - #69384 had them on the
sibling test until #69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
dabla pushed a commit to dabla/airflow that referenced this pull request Aug 14, 2026
…pache#70562)
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - apache#69384 had them on the
sibling test until apache#69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@shahar1@potiuk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Rerun flaky SSHRemoteJobOperator kill test on process-group races - #69384

Merged
potiuk merged 1 commit into
apache:mainfrom
shahar1:fix-flaky-ssh-remote-job-kill-test
Jul 5, 2026
Merged

Rerun flaky SSHRemoteJobOperator kill test on process-group races#69384
potiuk merged 1 commit into
apache:mainfrom
shahar1:fix-flaky-ssh-remote-job-kill-test

Conversation

@shahar1

Copy link
Copy Markdown
Contributor

TestPosixKillBehaviour::test_kill_terminates_whole_job_tree (added in #68644) intermittently fails its pre-kill assertion on CI:

AssertionError: job tree should be running before kill
assert False = _group_alive(326)

The wrapper records $! expecting it to equal the job PGID, but setsid only skips forking when the launching shell is not a process-group leader. On some runners it forks, so $! is the short-lived setsid parent and the job's real process group is already empty by the time pgrep -g runs — failing the check before the kill is even exercised. Passes locally (15/15, and under xdist), fails only on certain CI runners, so it's environment-dependent and lands on unrelated PRs.

Marks the test with @pytest.mark.flaky(reruns=5) (the convention already used elsewhere in this provider) so a fresh re-launch clears the race.

Note: if setsid forks in a real SSH session too, $! would be the wrong PID there and on_kill could orphan the job — a possible latent issue in the wrapper itself, left for a separate change.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 4.8)

Generated-by: Claude Code (Opus 4.8) following the guidelines

test_kill_terminates_whole_job_tree intermittently fails its pre-kill
check on CI: the wrapper records $! expecting it to equal the job PGID,
but setsid only skips forking when the launching shell is not a
process-group leader. On some runners it forks, so $! is the short-lived
setsid parent and the job's real group is empty by the time pgrep -g
runs. Rerun on a fresh draw so the environment-dependent race does not
fail unrelated PRs.
@potiuk
potiuk merged commit ea10f4f into apache:mainJul 5, 2026
81 checks passed
@shahar1
shahar1 deleted the fix-flaky-ssh-remote-job-kill-test branch July 5, 2026 20:35
potiuk added a commit to potiuk/airflow that referenced this pull request Jul 27, 2026
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - apache#69384 had them on the
sibling test until apache#69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
potiuk added a commit that referenced this pull request Jul 27, 2026
…70562)
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - #69384 had them on the
sibling test until #69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
dabla pushed a commit to dabla/airflow that referenced this pull request Aug 14, 2026
…pache#70562)
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - apache#69384 had them on the
sibling test until apache#69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@shahar1@potiuk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Rerun flaky SSHRemoteJobOperator kill test on process-group races - #69384

Merged
potiuk merged 1 commit into
apache:mainfrom
shahar1:fix-flaky-ssh-remote-job-kill-test
Jul 5, 2026
Merged

Rerun flaky SSHRemoteJobOperator kill test on process-group races#69384
potiuk merged 1 commit into
apache:mainfrom
shahar1:fix-flaky-ssh-remote-job-kill-test

Conversation

@shahar1

Copy link
Copy Markdown
Contributor

TestPosixKillBehaviour::test_kill_terminates_whole_job_tree (added in #68644) intermittently fails its pre-kill assertion on CI:

AssertionError: job tree should be running before kill
assert False = _group_alive(326)

The wrapper records $! expecting it to equal the job PGID, but setsid only skips forking when the launching shell is not a process-group leader. On some runners it forks, so $! is the short-lived setsid parent and the job's real process group is already empty by the time pgrep -g runs — failing the check before the kill is even exercised. Passes locally (15/15, and under xdist), fails only on certain CI runners, so it's environment-dependent and lands on unrelated PRs.

Marks the test with @pytest.mark.flaky(reruns=5) (the convention already used elsewhere in this provider) so a fresh re-launch clears the race.

Note: if setsid forks in a real SSH session too, $! would be the wrong PID there and on_kill could orphan the job — a possible latent issue in the wrapper itself, left for a separate change.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 4.8)

Generated-by: Claude Code (Opus 4.8) following the guidelines

test_kill_terminates_whole_job_tree intermittently fails its pre-kill
check on CI: the wrapper records $! expecting it to equal the job PGID,
but setsid only skips forking when the launching shell is not a
process-group leader. On some runners it forks, so $! is the short-lived
setsid parent and the job's real group is empty by the time pgrep -g
runs. Rerun on a fresh draw so the environment-dependent race does not
fail unrelated PRs.
@potiuk
potiuk merged commit ea10f4f into apache:mainJul 5, 2026
81 checks passed
@shahar1
shahar1 deleted the fix-flaky-ssh-remote-job-kill-test branch July 5, 2026 20:35
potiuk added a commit to potiuk/airflow that referenced this pull request Jul 27, 2026
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - apache#69384 had them on the
sibling test until apache#69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
potiuk added a commit that referenced this pull request Jul 27, 2026
…70562)
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - #69384 had them on the
sibling test until #69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
dabla pushed a commit to dabla/airflow that referenced this pull request Aug 14, 2026
…pache#70562)
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - apache#69384 had them on the
sibling test until apache#69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@shahar1@potiuk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Rerun flaky SSHRemoteJobOperator kill test on process-group races - #69384

Merged
potiuk merged 1 commit into
apache:mainfrom
shahar1:fix-flaky-ssh-remote-job-kill-test
Jul 5, 2026
Merged

Rerun flaky SSHRemoteJobOperator kill test on process-group races#69384
potiuk merged 1 commit into
apache:mainfrom
shahar1:fix-flaky-ssh-remote-job-kill-test

Conversation

@shahar1

Copy link
Copy Markdown
Contributor

TestPosixKillBehaviour::test_kill_terminates_whole_job_tree (added in #68644) intermittently fails its pre-kill assertion on CI:

AssertionError: job tree should be running before kill
assert False = _group_alive(326)

The wrapper records $! expecting it to equal the job PGID, but setsid only skips forking when the launching shell is not a process-group leader. On some runners it forks, so $! is the short-lived setsid parent and the job's real process group is already empty by the time pgrep -g runs — failing the check before the kill is even exercised. Passes locally (15/15, and under xdist), fails only on certain CI runners, so it's environment-dependent and lands on unrelated PRs.

Marks the test with @pytest.mark.flaky(reruns=5) (the convention already used elsewhere in this provider) so a fresh re-launch clears the race.

Note: if setsid forks in a real SSH session too, $! would be the wrong PID there and on_kill could orphan the job — a possible latent issue in the wrapper itself, left for a separate change.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 4.8)

Generated-by: Claude Code (Opus 4.8) following the guidelines

test_kill_terminates_whole_job_tree intermittently fails its pre-kill
check on CI: the wrapper records $! expecting it to equal the job PGID,
but setsid only skips forking when the launching shell is not a
process-group leader. On some runners it forks, so $! is the short-lived
setsid parent and the job's real group is empty by the time pgrep -g
runs. Rerun on a fresh draw so the environment-dependent race does not
fail unrelated PRs.
@potiuk
potiuk merged commit ea10f4f into apache:mainJul 5, 2026
81 checks passed
@shahar1
shahar1 deleted the fix-flaky-ssh-remote-job-kill-test branch July 5, 2026 20:35
potiuk added a commit to potiuk/airflow that referenced this pull request Jul 27, 2026
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - apache#69384 had them on the
sibling test until apache#69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
potiuk added a commit that referenced this pull request Jul 27, 2026
…70562)
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - #69384 had them on the
sibling test until #69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
dabla pushed a commit to dabla/airflow that referenced this pull request Aug 14, 2026
…pache#70562)
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - apache#69384 had them on the
sibling test until apache#69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@shahar1@potiuk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Rerun flaky SSHRemoteJobOperator kill test on process-group races - #69384

Merged
potiuk merged 1 commit into
apache:mainfrom
shahar1:fix-flaky-ssh-remote-job-kill-test
Jul 5, 2026
Merged

Rerun flaky SSHRemoteJobOperator kill test on process-group races#69384
potiuk merged 1 commit into
apache:mainfrom
shahar1:fix-flaky-ssh-remote-job-kill-test

Conversation

@shahar1

Copy link
Copy Markdown
Contributor

TestPosixKillBehaviour::test_kill_terminates_whole_job_tree (added in #68644) intermittently fails its pre-kill assertion on CI:

AssertionError: job tree should be running before kill
assert False = _group_alive(326)

The wrapper records $! expecting it to equal the job PGID, but setsid only skips forking when the launching shell is not a process-group leader. On some runners it forks, so $! is the short-lived setsid parent and the job's real process group is already empty by the time pgrep -g runs — failing the check before the kill is even exercised. Passes locally (15/15, and under xdist), fails only on certain CI runners, so it's environment-dependent and lands on unrelated PRs.

Marks the test with @pytest.mark.flaky(reruns=5) (the convention already used elsewhere in this provider) so a fresh re-launch clears the race.

Note: if setsid forks in a real SSH session too, $! would be the wrong PID there and on_kill could orphan the job — a possible latent issue in the wrapper itself, left for a separate change.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 4.8)

Generated-by: Claude Code (Opus 4.8) following the guidelines

test_kill_terminates_whole_job_tree intermittently fails its pre-kill
check on CI: the wrapper records $! expecting it to equal the job PGID,
but setsid only skips forking when the launching shell is not a
process-group leader. On some runners it forks, so $! is the short-lived
setsid parent and the job's real group is empty by the time pgrep -g
runs. Rerun on a fresh draw so the environment-dependent race does not
fail unrelated PRs.
@potiuk
potiuk merged commit ea10f4f into apache:mainJul 5, 2026
81 checks passed
@shahar1
shahar1 deleted the fix-flaky-ssh-remote-job-kill-test branch July 5, 2026 20:35
potiuk added a commit to potiuk/airflow that referenced this pull request Jul 27, 2026
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - apache#69384 had them on the
sibling test until apache#69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
potiuk added a commit that referenced this pull request Jul 27, 2026
…70562)
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - #69384 had them on the
sibling test until #69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
dabla pushed a commit to dabla/airflow that referenced this pull request Aug 14, 2026
…pache#70562)
test_kill_terminates_whole_job_tree_under_job_control failed on main with "job
never wrote its pid file". The run took 5.16s, so SUBMIT_DONE arrived promptly and
the whole budget went on polling for a pid file that never appeared.
That symptom has two very different causes and the test cannot tell them apart:
_run_bash_mc_under_pty returns silently when the marker never arrives, so a
launcher that died immediately (EOF, no marker, returns at once) and a job that
died after being launched both surface later as the same empty pid file - and both
produce the same ~5s runtime. Two changes separate them:
- assert the marker was actually seen, quoting what the pty did produce
- report the job directory contents when the pid file stays empty; the wrapper
creates that directory and the log file before it backgrounds anything, so a
missing directory means the launcher never got there and an empty one means the
job was launched and died before its first statement
The teardown also has a real ordering hazard, closed here: the marker only says the
launcher returned, which it does the moment it backgrounds setsid - before that
child has forked, called setsid(2) and exec'd into the job. Closing the pty master
hangs up the terminal, and pty.fork() makes the launcher the session leader, so
hanging up inside that window could take the job down with the session. The pty is
now held open until the job proves it left the session by recording its own pid.
I could not reproduce the failure. macOS has no setsid(1) so the class skips there;
under Linux in Docker the real wrapper ran through this exact harness 65 times, 25
idle and 40 with the container throttled to 0.35 CPU against six busy loops, and
recorded its pid every time with and without the hold. So the hold is hardening, not
a demonstrated fix. Reruns come back for that reason - apache#69384 had them on the
sibling test until apache#69490 dropped the marker while writing this one - and the first
attempt's assertion text still reaches the CI log, so the next occurrence should
identify which half broke.
Generated-by: Claude Opus 5 (1M context)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@shahar1@potiuk