Keep SSH remote job running when its PTY session hangs up - #70573

Merged
potiuk merged 1 commit into
apache:mainfrom
jason810496:fix/ssh-remote-job-detach-stdin
Jul 28, 2026
Merged

Keep SSH remote job running when its PTY session hangs up#70573
potiuk merged 1 commit into
apache:mainfrom
jason810496:fix/ssh-remote-job-detach-stdin

Conversation

@jason810496

Copy link
Copy Markdown
Member

SSHRemoteJobOperator launches the remote job detached under setsid so it survives the SSH connection dropping. The wrapper redirected the job's stdout/stderr to /dev/null but left its stdin on the launching terminal. A fresh setsid session leader that still holds a terminal on any fd re-adopts it as its controlling terminal, so when an SSH session that allocated a PTY (get_pty=True) hangs up, the detached job receives SIGHUP and dies — orphaning the work the operator exists to keep alive.

Redirecting stdin from /dev/null as well (standard daemonization) leaves the job in a session with no controlling terminal, immune to the hangup.

This is also the root cause of the flaky test_kill_terminates_whole_job_tree_under_job_control: its PTY harness hangs up the terminal and then asserts the job is still running. The job was being killed by that hangup, which the @pytest.mark.flaky(reruns=5) marker only masked. With the wrapper fixed the test is deterministic, so the marker and its stale "process-group race" comment are removed.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 4.8)

Generated-by: Claude Code (Opus 4.8) following the guidelines

SSHRemoteJobOperator launches the remote job detached under setsid so it
survives the SSH connection dropping. The launcher redirected the job's stdout
and stderr to /dev/null but left its stdin on the launching terminal. A fresh
setsid session leader that still holds a terminal on any file descriptor
re-adopts it as its controlling terminal, so when an SSH session that allocated
a PTY hangs up, the job received SIGHUP and died -- orphaning the work the
operator exists to keep alive. Detaching stdin as well leaves the job in a
session with no controlling terminal, immune to the hangup.

@jason810496jason810496 left a comment

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some PRs ( #69933 ) still encountered the CI failure of SSH provider even after rebasing on top of #70562.

cc @potiuk

@potiuk
potiuk merged commit c04b93b into apache:mainJul 28, 2026
83 checks passed
@1fanwang

1fanwang commented Jul 28, 2026

Copy link
Copy Markdown
Contributor

Thanks for tracking this down, @jason810496. landed on the same RC while chasing this:

The detached job keeps the launcher's terminal on fd 0, so the fresh setsid session leader re-acquires it as its controlling terminal. When the pty hangs up, the job gets SIGHUP and dies. Redirecting stdin from /dev/null is what fully detaches it, exactly as you have it on both the setsid and nohup paths.

I reproduced it in a Linux container. The controlling-terminal timing makes it intermittent on CI but deterministic in a slower Docker VM, which gives a clean before/after.

  1. Baseline (current main): the real test fails with the CI assertion, all 5 reruns exhausted

Linux 6.10.14-linuxkit aarch64
$ pytest test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control --noconftest

 job_pids = subprocess.run(
["pgrep", "-f", marker], capture_output=True, text=True, check=False
).stdout.split()
 assert job_pids, "job never started"

E AssertionError: job never started
E assert []

test_remote_job.py:426: AssertionError
=========================== short test summary info ============================
FAILED test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control
========================== 1 failed, 5 rerun in 0.32s ==========================

  1. Mechanism: the detached job holds the pty as its controlling terminal

 ps on the job right before the pty hangup. pid == sid == pgid confirms setsid(2) gave it a new session; tty = pts/0 shows it still holds the launcher's pty as its controlling terminal, which is the path the SIGHUP travels down:

recorded pid: 31
ps (pid sid pgid tty) BEFORE hangup:
31 31 31 pts/0
job alive AFTER hangup: False

  1. With this PR's change: the test passes deterministically, no regression

Reproduce

From an apache/airflow checkout. The test skips on macOS, so run it under Linux/Docker:

docker run --rm -v "$PWD:/repo" -w /tmp python:3.10 bash -c '
apt-get update -qq && apt-get install -y -qq util-linux procps >/dev/null
pip install -q pytest pytest-rerunfailures
cp /repo/providers/ssh/src/airflow/providers/ssh/utils/remote_job.py .
cp /repo/providers/ssh/tests/unit/ssh/utils/test_remote_job.py .
sed -i "s/from airflow.providers.ssh.utils.remote_job import (/from remote_job import (/" test_remote_job.py
pytest test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control --noconftest'

Dropping the @pytest.mark.flaky(reruns=5) marker is the right call too: with stdin detached the test passed 20/20 for me, so it no longer needs reruns. +1 from me.

@jason810496

Copy link
Copy Markdown
MemberAuthor

Yeah, I verified on Linux container as well. Having standard daemonization solves the flaky test issue. Thanks for double checking this!

dabla pushed a commit to dabla/airflow that referenced this pull request Aug 14, 2026
SSHRemoteJobOperator launches the remote job detached under setsid so it
survives the SSH connection dropping. The launcher redirected the job's stdout
and stderr to /dev/null but left its stdin on the launching terminal. A fresh
setsid session leader that still holds a terminal on any file descriptor
re-adopts it as its controlling terminal, so when an SSH session that allocated
a PTY hangs up, the job received SIGHUP and died -- orphaning the work the
operator exists to keep alive. Detaching stdin as well leaves the job in a
session with no controlling terminal, immune to the hangup.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jason810496@1fanwang@potiuk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Keep SSH remote job running when its PTY session hangs up - #70573

Merged
potiuk merged 1 commit into
apache:mainfrom
jason810496:fix/ssh-remote-job-detach-stdin
Jul 28, 2026
Merged

Keep SSH remote job running when its PTY session hangs up#70573
potiuk merged 1 commit into
apache:mainfrom
jason810496:fix/ssh-remote-job-detach-stdin

Conversation

@jason810496

Copy link
Copy Markdown
Member

SSHRemoteJobOperator launches the remote job detached under setsid so it survives the SSH connection dropping. The wrapper redirected the job's stdout/stderr to /dev/null but left its stdin on the launching terminal. A fresh setsid session leader that still holds a terminal on any fd re-adopts it as its controlling terminal, so when an SSH session that allocated a PTY (get_pty=True) hangs up, the detached job receives SIGHUP and dies — orphaning the work the operator exists to keep alive.

Redirecting stdin from /dev/null as well (standard daemonization) leaves the job in a session with no controlling terminal, immune to the hangup.

This is also the root cause of the flaky test_kill_terminates_whole_job_tree_under_job_control: its PTY harness hangs up the terminal and then asserts the job is still running. The job was being killed by that hangup, which the @pytest.mark.flaky(reruns=5) marker only masked. With the wrapper fixed the test is deterministic, so the marker and its stale "process-group race" comment are removed.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 4.8)

Generated-by: Claude Code (Opus 4.8) following the guidelines

SSHRemoteJobOperator launches the remote job detached under setsid so it
survives the SSH connection dropping. The launcher redirected the job's stdout
and stderr to /dev/null but left its stdin on the launching terminal. A fresh
setsid session leader that still holds a terminal on any file descriptor
re-adopts it as its controlling terminal, so when an SSH session that allocated
a PTY hangs up, the job received SIGHUP and died -- orphaning the work the
operator exists to keep alive. Detaching stdin as well leaves the job in a
session with no controlling terminal, immune to the hangup.

@jason810496jason810496 left a comment

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some PRs ( #69933 ) still encountered the CI failure of SSH provider even after rebasing on top of #70562.

cc @potiuk

@potiuk
potiuk merged commit c04b93b into apache:mainJul 28, 2026
83 checks passed
@1fanwang

1fanwang commented Jul 28, 2026

Copy link
Copy Markdown
Contributor

Thanks for tracking this down, @jason810496. landed on the same RC while chasing this:

The detached job keeps the launcher's terminal on fd 0, so the fresh setsid session leader re-acquires it as its controlling terminal. When the pty hangs up, the job gets SIGHUP and dies. Redirecting stdin from /dev/null is what fully detaches it, exactly as you have it on both the setsid and nohup paths.

I reproduced it in a Linux container. The controlling-terminal timing makes it intermittent on CI but deterministic in a slower Docker VM, which gives a clean before/after.

  1. Baseline (current main): the real test fails with the CI assertion, all 5 reruns exhausted

Linux 6.10.14-linuxkit aarch64
$ pytest test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control --noconftest

 job_pids = subprocess.run(
["pgrep", "-f", marker], capture_output=True, text=True, check=False
).stdout.split()
 assert job_pids, "job never started"

E AssertionError: job never started
E assert []

test_remote_job.py:426: AssertionError
=========================== short test summary info ============================
FAILED test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control
========================== 1 failed, 5 rerun in 0.32s ==========================

  1. Mechanism: the detached job holds the pty as its controlling terminal

 ps on the job right before the pty hangup. pid == sid == pgid confirms setsid(2) gave it a new session; tty = pts/0 shows it still holds the launcher's pty as its controlling terminal, which is the path the SIGHUP travels down:

recorded pid: 31
ps (pid sid pgid tty) BEFORE hangup:
31 31 31 pts/0
job alive AFTER hangup: False

  1. With this PR's change: the test passes deterministically, no regression

Reproduce

From an apache/airflow checkout. The test skips on macOS, so run it under Linux/Docker:

docker run --rm -v "$PWD:/repo" -w /tmp python:3.10 bash -c '
apt-get update -qq && apt-get install -y -qq util-linux procps >/dev/null
pip install -q pytest pytest-rerunfailures
cp /repo/providers/ssh/src/airflow/providers/ssh/utils/remote_job.py .
cp /repo/providers/ssh/tests/unit/ssh/utils/test_remote_job.py .
sed -i "s/from airflow.providers.ssh.utils.remote_job import (/from remote_job import (/" test_remote_job.py
pytest test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control --noconftest'

Dropping the @pytest.mark.flaky(reruns=5) marker is the right call too: with stdin detached the test passed 20/20 for me, so it no longer needs reruns. +1 from me.

@jason810496

Copy link
Copy Markdown
MemberAuthor

Yeah, I verified on Linux container as well. Having standard daemonization solves the flaky test issue. Thanks for double checking this!

dabla pushed a commit to dabla/airflow that referenced this pull request Aug 14, 2026
SSHRemoteJobOperator launches the remote job detached under setsid so it
survives the SSH connection dropping. The launcher redirected the job's stdout
and stderr to /dev/null but left its stdin on the launching terminal. A fresh
setsid session leader that still holds a terminal on any file descriptor
re-adopts it as its controlling terminal, so when an SSH session that allocated
a PTY hangs up, the job received SIGHUP and died -- orphaning the work the
operator exists to keep alive. Detaching stdin as well leaves the job in a
session with no controlling terminal, immune to the hangup.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jason810496@1fanwang@potiuk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Keep SSH remote job running when its PTY session hangs up - #70573

Merged
potiuk merged 1 commit into
apache:mainfrom
jason810496:fix/ssh-remote-job-detach-stdin
Jul 28, 2026
Merged

Keep SSH remote job running when its PTY session hangs up#70573
potiuk merged 1 commit into
apache:mainfrom
jason810496:fix/ssh-remote-job-detach-stdin

Conversation

@jason810496

Copy link
Copy Markdown
Member

SSHRemoteJobOperator launches the remote job detached under setsid so it survives the SSH connection dropping. The wrapper redirected the job's stdout/stderr to /dev/null but left its stdin on the launching terminal. A fresh setsid session leader that still holds a terminal on any fd re-adopts it as its controlling terminal, so when an SSH session that allocated a PTY (get_pty=True) hangs up, the detached job receives SIGHUP and dies — orphaning the work the operator exists to keep alive.

Redirecting stdin from /dev/null as well (standard daemonization) leaves the job in a session with no controlling terminal, immune to the hangup.

This is also the root cause of the flaky test_kill_terminates_whole_job_tree_under_job_control: its PTY harness hangs up the terminal and then asserts the job is still running. The job was being killed by that hangup, which the @pytest.mark.flaky(reruns=5) marker only masked. With the wrapper fixed the test is deterministic, so the marker and its stale "process-group race" comment are removed.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 4.8)

Generated-by: Claude Code (Opus 4.8) following the guidelines

SSHRemoteJobOperator launches the remote job detached under setsid so it
survives the SSH connection dropping. The launcher redirected the job's stdout
and stderr to /dev/null but left its stdin on the launching terminal. A fresh
setsid session leader that still holds a terminal on any file descriptor
re-adopts it as its controlling terminal, so when an SSH session that allocated
a PTY hangs up, the job received SIGHUP and died -- orphaning the work the
operator exists to keep alive. Detaching stdin as well leaves the job in a
session with no controlling terminal, immune to the hangup.

@jason810496jason810496 left a comment

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some PRs ( #69933 ) still encountered the CI failure of SSH provider even after rebasing on top of #70562.

cc @potiuk

@potiuk
potiuk merged commit c04b93b into apache:mainJul 28, 2026
83 checks passed
@1fanwang

1fanwang commented Jul 28, 2026

Copy link
Copy Markdown
Contributor

Thanks for tracking this down, @jason810496. landed on the same RC while chasing this:

The detached job keeps the launcher's terminal on fd 0, so the fresh setsid session leader re-acquires it as its controlling terminal. When the pty hangs up, the job gets SIGHUP and dies. Redirecting stdin from /dev/null is what fully detaches it, exactly as you have it on both the setsid and nohup paths.

I reproduced it in a Linux container. The controlling-terminal timing makes it intermittent on CI but deterministic in a slower Docker VM, which gives a clean before/after.

  1. Baseline (current main): the real test fails with the CI assertion, all 5 reruns exhausted

Linux 6.10.14-linuxkit aarch64
$ pytest test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control --noconftest

 job_pids = subprocess.run(
["pgrep", "-f", marker], capture_output=True, text=True, check=False
).stdout.split()
 assert job_pids, "job never started"

E AssertionError: job never started
E assert []

test_remote_job.py:426: AssertionError
=========================== short test summary info ============================
FAILED test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control
========================== 1 failed, 5 rerun in 0.32s ==========================

  1. Mechanism: the detached job holds the pty as its controlling terminal

 ps on the job right before the pty hangup. pid == sid == pgid confirms setsid(2) gave it a new session; tty = pts/0 shows it still holds the launcher's pty as its controlling terminal, which is the path the SIGHUP travels down:

recorded pid: 31
ps (pid sid pgid tty) BEFORE hangup:
31 31 31 pts/0
job alive AFTER hangup: False

  1. With this PR's change: the test passes deterministically, no regression

Reproduce

From an apache/airflow checkout. The test skips on macOS, so run it under Linux/Docker:

docker run --rm -v "$PWD:/repo" -w /tmp python:3.10 bash -c '
apt-get update -qq && apt-get install -y -qq util-linux procps >/dev/null
pip install -q pytest pytest-rerunfailures
cp /repo/providers/ssh/src/airflow/providers/ssh/utils/remote_job.py .
cp /repo/providers/ssh/tests/unit/ssh/utils/test_remote_job.py .
sed -i "s/from airflow.providers.ssh.utils.remote_job import (/from remote_job import (/" test_remote_job.py
pytest test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control --noconftest'

Dropping the @pytest.mark.flaky(reruns=5) marker is the right call too: with stdin detached the test passed 20/20 for me, so it no longer needs reruns. +1 from me.

@jason810496

Copy link
Copy Markdown
MemberAuthor

Yeah, I verified on Linux container as well. Having standard daemonization solves the flaky test issue. Thanks for double checking this!

dabla pushed a commit to dabla/airflow that referenced this pull request Aug 14, 2026
SSHRemoteJobOperator launches the remote job detached under setsid so it
survives the SSH connection dropping. The launcher redirected the job's stdout
and stderr to /dev/null but left its stdin on the launching terminal. A fresh
setsid session leader that still holds a terminal on any file descriptor
re-adopts it as its controlling terminal, so when an SSH session that allocated
a PTY hangs up, the job received SIGHUP and died -- orphaning the work the
operator exists to keep alive. Detaching stdin as well leaves the job in a
session with no controlling terminal, immune to the hangup.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jason810496@1fanwang@potiuk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Keep SSH remote job running when its PTY session hangs up - #70573

Merged
potiuk merged 1 commit into
apache:mainfrom
jason810496:fix/ssh-remote-job-detach-stdin
Jul 28, 2026
Merged

Keep SSH remote job running when its PTY session hangs up#70573
potiuk merged 1 commit into
apache:mainfrom
jason810496:fix/ssh-remote-job-detach-stdin

Conversation

@jason810496

Copy link
Copy Markdown
Member

SSHRemoteJobOperator launches the remote job detached under setsid so it survives the SSH connection dropping. The wrapper redirected the job's stdout/stderr to /dev/null but left its stdin on the launching terminal. A fresh setsid session leader that still holds a terminal on any fd re-adopts it as its controlling terminal, so when an SSH session that allocated a PTY (get_pty=True) hangs up, the detached job receives SIGHUP and dies — orphaning the work the operator exists to keep alive.

Redirecting stdin from /dev/null as well (standard daemonization) leaves the job in a session with no controlling terminal, immune to the hangup.

This is also the root cause of the flaky test_kill_terminates_whole_job_tree_under_job_control: its PTY harness hangs up the terminal and then asserts the job is still running. The job was being killed by that hangup, which the @pytest.mark.flaky(reruns=5) marker only masked. With the wrapper fixed the test is deterministic, so the marker and its stale "process-group race" comment are removed.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 4.8)

Generated-by: Claude Code (Opus 4.8) following the guidelines

SSHRemoteJobOperator launches the remote job detached under setsid so it
survives the SSH connection dropping. The launcher redirected the job's stdout
and stderr to /dev/null but left its stdin on the launching terminal. A fresh
setsid session leader that still holds a terminal on any file descriptor
re-adopts it as its controlling terminal, so when an SSH session that allocated
a PTY hangs up, the job received SIGHUP and died -- orphaning the work the
operator exists to keep alive. Detaching stdin as well leaves the job in a
session with no controlling terminal, immune to the hangup.

@jason810496jason810496 left a comment

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some PRs ( #69933 ) still encountered the CI failure of SSH provider even after rebasing on top of #70562.

cc @potiuk

@potiuk
potiuk merged commit c04b93b into apache:mainJul 28, 2026
83 checks passed
@1fanwang

1fanwang commented Jul 28, 2026

Copy link
Copy Markdown
Contributor

Thanks for tracking this down, @jason810496. landed on the same RC while chasing this:

The detached job keeps the launcher's terminal on fd 0, so the fresh setsid session leader re-acquires it as its controlling terminal. When the pty hangs up, the job gets SIGHUP and dies. Redirecting stdin from /dev/null is what fully detaches it, exactly as you have it on both the setsid and nohup paths.

I reproduced it in a Linux container. The controlling-terminal timing makes it intermittent on CI but deterministic in a slower Docker VM, which gives a clean before/after.

  1. Baseline (current main): the real test fails with the CI assertion, all 5 reruns exhausted

Linux 6.10.14-linuxkit aarch64
$ pytest test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control --noconftest

 job_pids = subprocess.run(
["pgrep", "-f", marker], capture_output=True, text=True, check=False
).stdout.split()
 assert job_pids, "job never started"

E AssertionError: job never started
E assert []

test_remote_job.py:426: AssertionError
=========================== short test summary info ============================
FAILED test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control
========================== 1 failed, 5 rerun in 0.32s ==========================

  1. Mechanism: the detached job holds the pty as its controlling terminal

 ps on the job right before the pty hangup. pid == sid == pgid confirms setsid(2) gave it a new session; tty = pts/0 shows it still holds the launcher's pty as its controlling terminal, which is the path the SIGHUP travels down:

recorded pid: 31
ps (pid sid pgid tty) BEFORE hangup:
31 31 31 pts/0
job alive AFTER hangup: False

  1. With this PR's change: the test passes deterministically, no regression

Reproduce

From an apache/airflow checkout. The test skips on macOS, so run it under Linux/Docker:

docker run --rm -v "$PWD:/repo" -w /tmp python:3.10 bash -c '
apt-get update -qq && apt-get install -y -qq util-linux procps >/dev/null
pip install -q pytest pytest-rerunfailures
cp /repo/providers/ssh/src/airflow/providers/ssh/utils/remote_job.py .
cp /repo/providers/ssh/tests/unit/ssh/utils/test_remote_job.py .
sed -i "s/from airflow.providers.ssh.utils.remote_job import (/from remote_job import (/" test_remote_job.py
pytest test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control --noconftest'

Dropping the @pytest.mark.flaky(reruns=5) marker is the right call too: with stdin detached the test passed 20/20 for me, so it no longer needs reruns. +1 from me.

@jason810496

Copy link
Copy Markdown
MemberAuthor

Yeah, I verified on Linux container as well. Having standard daemonization solves the flaky test issue. Thanks for double checking this!

dabla pushed a commit to dabla/airflow that referenced this pull request Aug 14, 2026
SSHRemoteJobOperator launches the remote job detached under setsid so it
survives the SSH connection dropping. The launcher redirected the job's stdout
and stderr to /dev/null but left its stdin on the launching terminal. A fresh
setsid session leader that still holds a terminal on any file descriptor
re-adopts it as its controlling terminal, so when an SSH session that allocated
a PTY hangs up, the job received SIGHUP and died -- orphaning the work the
operator exists to keep alive. Detaching stdin as well leaves the job in a
session with no controlling terminal, immune to the hangup.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jason810496@1fanwang@potiuk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Keep SSH remote job running when its PTY session hangs up - #70573

Merged
potiuk merged 1 commit into
apache:mainfrom
jason810496:fix/ssh-remote-job-detach-stdin
Jul 28, 2026
Merged

Keep SSH remote job running when its PTY session hangs up#70573
potiuk merged 1 commit into
apache:mainfrom
jason810496:fix/ssh-remote-job-detach-stdin

Conversation

@jason810496

Copy link
Copy Markdown
Member

SSHRemoteJobOperator launches the remote job detached under setsid so it survives the SSH connection dropping. The wrapper redirected the job's stdout/stderr to /dev/null but left its stdin on the launching terminal. A fresh setsid session leader that still holds a terminal on any fd re-adopts it as its controlling terminal, so when an SSH session that allocated a PTY (get_pty=True) hangs up, the detached job receives SIGHUP and dies — orphaning the work the operator exists to keep alive.

Redirecting stdin from /dev/null as well (standard daemonization) leaves the job in a session with no controlling terminal, immune to the hangup.

This is also the root cause of the flaky test_kill_terminates_whole_job_tree_under_job_control: its PTY harness hangs up the terminal and then asserts the job is still running. The job was being killed by that hangup, which the @pytest.mark.flaky(reruns=5) marker only masked. With the wrapper fixed the test is deterministic, so the marker and its stale "process-group race" comment are removed.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 4.8)

Generated-by: Claude Code (Opus 4.8) following the guidelines

SSHRemoteJobOperator launches the remote job detached under setsid so it
survives the SSH connection dropping. The launcher redirected the job's stdout
and stderr to /dev/null but left its stdin on the launching terminal. A fresh
setsid session leader that still holds a terminal on any file descriptor
re-adopts it as its controlling terminal, so when an SSH session that allocated
a PTY hangs up, the job received SIGHUP and died -- orphaning the work the
operator exists to keep alive. Detaching stdin as well leaves the job in a
session with no controlling terminal, immune to the hangup.

@jason810496jason810496 left a comment

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some PRs ( #69933 ) still encountered the CI failure of SSH provider even after rebasing on top of #70562.

cc @potiuk

@potiuk
potiuk merged commit c04b93b into apache:mainJul 28, 2026
83 checks passed
@1fanwang

1fanwang commented Jul 28, 2026

Copy link
Copy Markdown
Contributor

Thanks for tracking this down, @jason810496. landed on the same RC while chasing this:

The detached job keeps the launcher's terminal on fd 0, so the fresh setsid session leader re-acquires it as its controlling terminal. When the pty hangs up, the job gets SIGHUP and dies. Redirecting stdin from /dev/null is what fully detaches it, exactly as you have it on both the setsid and nohup paths.

I reproduced it in a Linux container. The controlling-terminal timing makes it intermittent on CI but deterministic in a slower Docker VM, which gives a clean before/after.

  1. Baseline (current main): the real test fails with the CI assertion, all 5 reruns exhausted

Linux 6.10.14-linuxkit aarch64
$ pytest test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control --noconftest

 job_pids = subprocess.run(
["pgrep", "-f", marker], capture_output=True, text=True, check=False
).stdout.split()
 assert job_pids, "job never started"

E AssertionError: job never started
E assert []

test_remote_job.py:426: AssertionError
=========================== short test summary info ============================
FAILED test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control
========================== 1 failed, 5 rerun in 0.32s ==========================

  1. Mechanism: the detached job holds the pty as its controlling terminal

 ps on the job right before the pty hangup. pid == sid == pgid confirms setsid(2) gave it a new session; tty = pts/0 shows it still holds the launcher's pty as its controlling terminal, which is the path the SIGHUP travels down:

recorded pid: 31
ps (pid sid pgid tty) BEFORE hangup:
31 31 31 pts/0
job alive AFTER hangup: False

  1. With this PR's change: the test passes deterministically, no regression

Reproduce

From an apache/airflow checkout. The test skips on macOS, so run it under Linux/Docker:

docker run --rm -v "$PWD:/repo" -w /tmp python:3.10 bash -c '
apt-get update -qq && apt-get install -y -qq util-linux procps >/dev/null
pip install -q pytest pytest-rerunfailures
cp /repo/providers/ssh/src/airflow/providers/ssh/utils/remote_job.py .
cp /repo/providers/ssh/tests/unit/ssh/utils/test_remote_job.py .
sed -i "s/from airflow.providers.ssh.utils.remote_job import (/from remote_job import (/" test_remote_job.py
pytest test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control --noconftest'

Dropping the @pytest.mark.flaky(reruns=5) marker is the right call too: with stdin detached the test passed 20/20 for me, so it no longer needs reruns. +1 from me.

@jason810496

Copy link
Copy Markdown
MemberAuthor

Yeah, I verified on Linux container as well. Having standard daemonization solves the flaky test issue. Thanks for double checking this!

dabla pushed a commit to dabla/airflow that referenced this pull request Aug 14, 2026
SSHRemoteJobOperator launches the remote job detached under setsid so it
survives the SSH connection dropping. The launcher redirected the job's stdout
and stderr to /dev/null but left its stdin on the launching terminal. A fresh
setsid session leader that still holds a terminal on any file descriptor
re-adopts it as its controlling terminal, so when an SSH session that allocated
a PTY hangs up, the job received SIGHUP and died -- orphaning the work the
operator exists to keep alive. Detaching stdin as well leaves the job in a
session with no controlling terminal, immune to the hangup.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jason810496@1fanwang@potiuk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Keep SSH remote job running when its PTY session hangs up - #70573

Merged
potiuk merged 1 commit into
apache:mainfrom
jason810496:fix/ssh-remote-job-detach-stdin
Jul 28, 2026
Merged

Keep SSH remote job running when its PTY session hangs up#70573
potiuk merged 1 commit into
apache:mainfrom
jason810496:fix/ssh-remote-job-detach-stdin

Conversation

@jason810496

Copy link
Copy Markdown
Member

SSHRemoteJobOperator launches the remote job detached under setsid so it survives the SSH connection dropping. The wrapper redirected the job's stdout/stderr to /dev/null but left its stdin on the launching terminal. A fresh setsid session leader that still holds a terminal on any fd re-adopts it as its controlling terminal, so when an SSH session that allocated a PTY (get_pty=True) hangs up, the detached job receives SIGHUP and dies — orphaning the work the operator exists to keep alive.

Redirecting stdin from /dev/null as well (standard daemonization) leaves the job in a session with no controlling terminal, immune to the hangup.

This is also the root cause of the flaky test_kill_terminates_whole_job_tree_under_job_control: its PTY harness hangs up the terminal and then asserts the job is still running. The job was being killed by that hangup, which the @pytest.mark.flaky(reruns=5) marker only masked. With the wrapper fixed the test is deterministic, so the marker and its stale "process-group race" comment are removed.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 4.8)

Generated-by: Claude Code (Opus 4.8) following the guidelines

SSHRemoteJobOperator launches the remote job detached under setsid so it
survives the SSH connection dropping. The launcher redirected the job's stdout
and stderr to /dev/null but left its stdin on the launching terminal. A fresh
setsid session leader that still holds a terminal on any file descriptor
re-adopts it as its controlling terminal, so when an SSH session that allocated
a PTY hangs up, the job received SIGHUP and died -- orphaning the work the
operator exists to keep alive. Detaching stdin as well leaves the job in a
session with no controlling terminal, immune to the hangup.

@jason810496jason810496 left a comment

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some PRs ( #69933 ) still encountered the CI failure of SSH provider even after rebasing on top of #70562.

cc @potiuk

@potiuk
potiuk merged commit c04b93b into apache:mainJul 28, 2026
83 checks passed
@1fanwang

1fanwang commented Jul 28, 2026

Copy link
Copy Markdown
Contributor

Thanks for tracking this down, @jason810496. landed on the same RC while chasing this:

The detached job keeps the launcher's terminal on fd 0, so the fresh setsid session leader re-acquires it as its controlling terminal. When the pty hangs up, the job gets SIGHUP and dies. Redirecting stdin from /dev/null is what fully detaches it, exactly as you have it on both the setsid and nohup paths.

I reproduced it in a Linux container. The controlling-terminal timing makes it intermittent on CI but deterministic in a slower Docker VM, which gives a clean before/after.

  1. Baseline (current main): the real test fails with the CI assertion, all 5 reruns exhausted

Linux 6.10.14-linuxkit aarch64
$ pytest test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control --noconftest

 job_pids = subprocess.run(
["pgrep", "-f", marker], capture_output=True, text=True, check=False
).stdout.split()
 assert job_pids, "job never started"

E AssertionError: job never started
E assert []

test_remote_job.py:426: AssertionError
=========================== short test summary info ============================
FAILED test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control
========================== 1 failed, 5 rerun in 0.32s ==========================

  1. Mechanism: the detached job holds the pty as its controlling terminal

 ps on the job right before the pty hangup. pid == sid == pgid confirms setsid(2) gave it a new session; tty = pts/0 shows it still holds the launcher's pty as its controlling terminal, which is the path the SIGHUP travels down:

recorded pid: 31
ps (pid sid pgid tty) BEFORE hangup:
31 31 31 pts/0
job alive AFTER hangup: False

  1. With this PR's change: the test passes deterministically, no regression

Reproduce

From an apache/airflow checkout. The test skips on macOS, so run it under Linux/Docker:

docker run --rm -v "$PWD:/repo" -w /tmp python:3.10 bash -c '
apt-get update -qq && apt-get install -y -qq util-linux procps >/dev/null
pip install -q pytest pytest-rerunfailures
cp /repo/providers/ssh/src/airflow/providers/ssh/utils/remote_job.py .
cp /repo/providers/ssh/tests/unit/ssh/utils/test_remote_job.py .
sed -i "s/from airflow.providers.ssh.utils.remote_job import (/from remote_job import (/" test_remote_job.py
pytest test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control --noconftest'

Dropping the @pytest.mark.flaky(reruns=5) marker is the right call too: with stdin detached the test passed 20/20 for me, so it no longer needs reruns. +1 from me.

@jason810496

Copy link
Copy Markdown
MemberAuthor

Yeah, I verified on Linux container as well. Having standard daemonization solves the flaky test issue. Thanks for double checking this!

dabla pushed a commit to dabla/airflow that referenced this pull request Aug 14, 2026
SSHRemoteJobOperator launches the remote job detached under setsid so it
survives the SSH connection dropping. The launcher redirected the job's stdout
and stderr to /dev/null but left its stdin on the launching terminal. A fresh
setsid session leader that still holds a terminal on any file descriptor
re-adopts it as its controlling terminal, so when an SSH session that allocated
a PTY hangs up, the job received SIGHUP and died -- orphaning the work the
operator exists to keep alive. Detaching stdin as well leaves the job in a
session with no controlling terminal, immune to the hangup.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jason810496@1fanwang@potiuk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Keep SSH remote job running when its PTY session hangs up - #70573

Merged
potiuk merged 1 commit into
apache:mainfrom
jason810496:fix/ssh-remote-job-detach-stdin
Jul 28, 2026
Merged

Keep SSH remote job running when its PTY session hangs up#70573
potiuk merged 1 commit into
apache:mainfrom
jason810496:fix/ssh-remote-job-detach-stdin

Conversation

@jason810496

Copy link
Copy Markdown
Member

SSHRemoteJobOperator launches the remote job detached under setsid so it survives the SSH connection dropping. The wrapper redirected the job's stdout/stderr to /dev/null but left its stdin on the launching terminal. A fresh setsid session leader that still holds a terminal on any fd re-adopts it as its controlling terminal, so when an SSH session that allocated a PTY (get_pty=True) hangs up, the detached job receives SIGHUP and dies — orphaning the work the operator exists to keep alive.

Redirecting stdin from /dev/null as well (standard daemonization) leaves the job in a session with no controlling terminal, immune to the hangup.

This is also the root cause of the flaky test_kill_terminates_whole_job_tree_under_job_control: its PTY harness hangs up the terminal and then asserts the job is still running. The job was being killed by that hangup, which the @pytest.mark.flaky(reruns=5) marker only masked. With the wrapper fixed the test is deterministic, so the marker and its stale "process-group race" comment are removed.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 4.8)

Generated-by: Claude Code (Opus 4.8) following the guidelines

SSHRemoteJobOperator launches the remote job detached under setsid so it
survives the SSH connection dropping. The launcher redirected the job's stdout
and stderr to /dev/null but left its stdin on the launching terminal. A fresh
setsid session leader that still holds a terminal on any file descriptor
re-adopts it as its controlling terminal, so when an SSH session that allocated
a PTY hangs up, the job received SIGHUP and died -- orphaning the work the
operator exists to keep alive. Detaching stdin as well leaves the job in a
session with no controlling terminal, immune to the hangup.

@jason810496jason810496 left a comment

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some PRs ( #69933 ) still encountered the CI failure of SSH provider even after rebasing on top of #70562.

cc @potiuk

@potiuk
potiuk merged commit c04b93b into apache:mainJul 28, 2026
83 checks passed
@1fanwang

1fanwang commented Jul 28, 2026

Copy link
Copy Markdown
Contributor

Thanks for tracking this down, @jason810496. landed on the same RC while chasing this:

The detached job keeps the launcher's terminal on fd 0, so the fresh setsid session leader re-acquires it as its controlling terminal. When the pty hangs up, the job gets SIGHUP and dies. Redirecting stdin from /dev/null is what fully detaches it, exactly as you have it on both the setsid and nohup paths.

I reproduced it in a Linux container. The controlling-terminal timing makes it intermittent on CI but deterministic in a slower Docker VM, which gives a clean before/after.

  1. Baseline (current main): the real test fails with the CI assertion, all 5 reruns exhausted

Linux 6.10.14-linuxkit aarch64
$ pytest test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control --noconftest

 job_pids = subprocess.run(
["pgrep", "-f", marker], capture_output=True, text=True, check=False
).stdout.split()
 assert job_pids, "job never started"

E AssertionError: job never started
E assert []

test_remote_job.py:426: AssertionError
=========================== short test summary info ============================
FAILED test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control
========================== 1 failed, 5 rerun in 0.32s ==========================

  1. Mechanism: the detached job holds the pty as its controlling terminal

 ps on the job right before the pty hangup. pid == sid == pgid confirms setsid(2) gave it a new session; tty = pts/0 shows it still holds the launcher's pty as its controlling terminal, which is the path the SIGHUP travels down:

recorded pid: 31
ps (pid sid pgid tty) BEFORE hangup:
31 31 31 pts/0
job alive AFTER hangup: False

  1. With this PR's change: the test passes deterministically, no regression

Reproduce

From an apache/airflow checkout. The test skips on macOS, so run it under Linux/Docker:

docker run --rm -v "$PWD:/repo" -w /tmp python:3.10 bash -c '
apt-get update -qq && apt-get install -y -qq util-linux procps >/dev/null
pip install -q pytest pytest-rerunfailures
cp /repo/providers/ssh/src/airflow/providers/ssh/utils/remote_job.py .
cp /repo/providers/ssh/tests/unit/ssh/utils/test_remote_job.py .
sed -i "s/from airflow.providers.ssh.utils.remote_job import (/from remote_job import (/" test_remote_job.py
pytest test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control --noconftest'

Dropping the @pytest.mark.flaky(reruns=5) marker is the right call too: with stdin detached the test passed 20/20 for me, so it no longer needs reruns. +1 from me.

@jason810496

Copy link
Copy Markdown
MemberAuthor

Yeah, I verified on Linux container as well. Having standard daemonization solves the flaky test issue. Thanks for double checking this!

dabla pushed a commit to dabla/airflow that referenced this pull request Aug 14, 2026
SSHRemoteJobOperator launches the remote job detached under setsid so it
survives the SSH connection dropping. The launcher redirected the job's stdout
and stderr to /dev/null but left its stdin on the launching terminal. A fresh
setsid session leader that still holds a terminal on any file descriptor
re-adopts it as its controlling terminal, so when an SSH session that allocated
a PTY hangs up, the job received SIGHUP and died -- orphaning the work the
operator exists to keep alive. Detaching stdin as well leaves the job in a
session with no controlling terminal, immune to the hangup.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jason810496@1fanwang@potiuk
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Keep SSH remote job running when its PTY session hangs up - #70573

Merged
potiuk merged 1 commit into
apache:mainfrom
jason810496:fix/ssh-remote-job-detach-stdin
Jul 28, 2026
Merged

Keep SSH remote job running when its PTY session hangs up#70573
potiuk merged 1 commit into
apache:mainfrom
jason810496:fix/ssh-remote-job-detach-stdin

Conversation

@jason810496

Copy link
Copy Markdown
Member

SSHRemoteJobOperator launches the remote job detached under setsid so it survives the SSH connection dropping. The wrapper redirected the job's stdout/stderr to /dev/null but left its stdin on the launching terminal. A fresh setsid session leader that still holds a terminal on any fd re-adopts it as its controlling terminal, so when an SSH session that allocated a PTY (get_pty=True) hangs up, the detached job receives SIGHUP and dies — orphaning the work the operator exists to keep alive.

Redirecting stdin from /dev/null as well (standard daemonization) leaves the job in a session with no controlling terminal, immune to the hangup.

This is also the root cause of the flaky test_kill_terminates_whole_job_tree_under_job_control: its PTY harness hangs up the terminal and then asserts the job is still running. The job was being killed by that hangup, which the @pytest.mark.flaky(reruns=5) marker only masked. With the wrapper fixed the test is deterministic, so the marker and its stale "process-group race" comment are removed.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 4.8)

Generated-by: Claude Code (Opus 4.8) following the guidelines

SSHRemoteJobOperator launches the remote job detached under setsid so it
survives the SSH connection dropping. The launcher redirected the job's stdout
and stderr to /dev/null but left its stdin on the launching terminal. A fresh
setsid session leader that still holds a terminal on any file descriptor
re-adopts it as its controlling terminal, so when an SSH session that allocated
a PTY hangs up, the job received SIGHUP and died -- orphaning the work the
operator exists to keep alive. Detaching stdin as well leaves the job in a
session with no controlling terminal, immune to the hangup.

@jason810496jason810496 left a comment

Copy link
Copy Markdown
MemberAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some PRs ( #69933 ) still encountered the CI failure of SSH provider even after rebasing on top of #70562.

cc @potiuk

@potiuk
potiuk merged commit c04b93b into apache:mainJul 28, 2026
83 checks passed
@1fanwang

1fanwang commented Jul 28, 2026

Copy link
Copy Markdown
Contributor

Thanks for tracking this down, @jason810496. landed on the same RC while chasing this:

The detached job keeps the launcher's terminal on fd 0, so the fresh setsid session leader re-acquires it as its controlling terminal. When the pty hangs up, the job gets SIGHUP and dies. Redirecting stdin from /dev/null is what fully detaches it, exactly as you have it on both the setsid and nohup paths.

I reproduced it in a Linux container. The controlling-terminal timing makes it intermittent on CI but deterministic in a slower Docker VM, which gives a clean before/after.

  1. Baseline (current main): the real test fails with the CI assertion, all 5 reruns exhausted

Linux 6.10.14-linuxkit aarch64
$ pytest test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control --noconftest

 job_pids = subprocess.run(
["pgrep", "-f", marker], capture_output=True, text=True, check=False
).stdout.split()
 assert job_pids, "job never started"

E AssertionError: job never started
E assert []

test_remote_job.py:426: AssertionError
=========================== short test summary info ============================
FAILED test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control
========================== 1 failed, 5 rerun in 0.32s ==========================

  1. Mechanism: the detached job holds the pty as its controlling terminal

 ps on the job right before the pty hangup. pid == sid == pgid confirms setsid(2) gave it a new session; tty = pts/0 shows it still holds the launcher's pty as its controlling terminal, which is the path the SIGHUP travels down:

recorded pid: 31
ps (pid sid pgid tty) BEFORE hangup:
31 31 31 pts/0
job alive AFTER hangup: False

  1. With this PR's change: the test passes deterministically, no regression

Reproduce

From an apache/airflow checkout. The test skips on macOS, so run it under Linux/Docker:

docker run --rm -v "$PWD:/repo" -w /tmp python:3.10 bash -c '
apt-get update -qq && apt-get install -y -qq util-linux procps >/dev/null
pip install -q pytest pytest-rerunfailures
cp /repo/providers/ssh/src/airflow/providers/ssh/utils/remote_job.py .
cp /repo/providers/ssh/tests/unit/ssh/utils/test_remote_job.py .
sed -i "s/from airflow.providers.ssh.utils.remote_job import (/from remote_job import (/" test_remote_job.py
pytest test_remote_job.py::TestPosixKillBehaviour::test_kill_terminates_whole_job_tree_under_job_control --noconftest'

Dropping the @pytest.mark.flaky(reruns=5) marker is the right call too: with stdin detached the test passed 20/20 for me, so it no longer needs reruns. +1 from me.

@jason810496

Copy link
Copy Markdown
MemberAuthor

Yeah, I verified on Linux container as well. Having standard daemonization solves the flaky test issue. Thanks for double checking this!

dabla pushed a commit to dabla/airflow that referenced this pull request Aug 14, 2026
SSHRemoteJobOperator launches the remote job detached under setsid so it
survives the SSH connection dropping. The launcher redirected the job's stdout
and stderr to /dev/null but left its stdin on the launching terminal. A fresh
setsid session leader that still holds a terminal on any file descriptor
re-adopts it as its controlling terminal, so when an SSH session that allocated
a PTY hangs up, the job received SIGHUP and died -- orphaning the work the
operator exists to keep alive. Detaching stdin as well leaves the job in a
session with no controlling terminal, immune to the hangup.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jason810496@1fanwang@potiuk