Spawner: Confirm startup was successful - #810

Merged
jan-janssen merged 2 commits into
mainfrom
spawner_success
Sep 7, 2025
Merged

Spawner: Confirm startup was successful#810
jan-janssen merged 2 commits into
mainfrom
spawner_success

Conversation

@jan-janssen

@jan-janssenjan-janssen commented Sep 7, 2025

Copy link
Copy Markdown
Member

Summary by CodeRabbit

  • New Features
    • bootup now returns a boolean success flag across interactive interfaces/spawners, enabling immediate startup status checks.
  • Documentation
    • Updated method docs to describe the boolean return value and expected behavior.
  • Refactor
    • Unified bootup behavior to propagate running status from underlying processes for consistent API semantics.
  • Tests
    • Updated existing tests to assert the success flag.
    • Added an error-handling test that verifies graceful termination when given an invalid command.

@coderabbitai

coderabbitaiBot commented Sep 7, 2025

Copy link
Copy Markdown
Contributor

Walkthrough

Bootup methods now return a boolean across interface and spawner components. SocketInterface.bootup returns the underlying spawner’s status. SubprocessSpawner.bootup and FluxPythonSpawner.bootup return poll() results post-start. Tests updated to assert the boolean success flag and include a new error-handling test for a failing serial command.

Changes

Cohort / File(s)Summary of Changes
Standalone interactive: bootup return propagation
executorlib/standalone/interactive/communication.py, executorlib/standalone/interactive/spawner.py
SocketInterface.bootup annotated to return bool and now returns the spawner’s bootup result. SubprocessSpawner.bootup returns poll() after launching, converting prior no-return into a boolean status. Docstrings/signatures updated accordingly.
Task scheduler Flux spawner
executorlib/task_scheduler/interactive/spawner_flux.py
FluxPythonSpawner.bootup annotated to return bool and now returns poll() after scheduling the job. Docstring updated. No other functional changes.
Tests updated for boolean bootup and error path
tests/test_standalone_interactive_communication.py
Tests now capture and assert the bootup boolean. Added test_interface_serial_with_error to validate behavior with a failing command and graceful shutdown. Minor imports and waits added.

Sequence Diagram(s)

sequenceDiagram
autonumber
participant T as Test
participant I as SocketInterface
participant S as Spawner (Subprocess/Flux)
participant P as Process/Job
T->>I: bootup(command_lst)
I->>S: bootup(command_lst)
S->>P: start/schedule
S->>S: poll() running?
S-->>I: bool (running status)
I-->>T: bool (propagated)
rect rgba(230,255,230,0.5)
note right of T: Caller can immediately know if process/job is running
end
Loading

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

  • Rename Interface to Spawner #398 — Also modifies Spawner/SocketInterface pathways, renaming/rewiring Interface → Spawner and adjusting types; closely related to return-type changes here.

Poem

A rabbit taps the bootup drum—thump, thump—
Now truth returns as bool, no guesswork slump.
Spawners poll, interfaces tell,
Tests nod yes: “All is well.”
If a job hops out of run, we still shut down—done! 🐇✅

Warning

Review ran into problems

🔥 Problems

Git: Failed to clone repository. Please run the @coderabbitai full review command to re-trigger a full review. If the issue persists, set path_filters to include or exclude specific files.

✨ Finishing Touches
  • 📝 Generate Docstrings
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch spawner_success

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@codecov

codecovBot commented Sep 7, 2025

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 97.77%. Comparing base (f1847a0) to head (8cfc8fa).
⚠️ Report is 2 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #810 +/- ##
=======================================
Coverage 97.76% 97.77% =======================================
Files 32 32 Lines 1479 1481 +2 =======================================
+ Hits 1446 1448 +2 
Misses 33 33 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
executorlib/standalone/interactive/spawner.py (1)

87-105: Make bootup success non-racy and type it as bool

Returning self.poll() immediately after Popen can intermittently be False for very short-lived commands, making startup detection flaky. Treat a successful Popen as “startup succeeded” and annotate the return type.

- def bootup(+ def bootup(
self,
command_lst: list[str],
- ):+ ) -> bool:
@@
- return self.poll()+ # Successful process creation == successful bootup; ongoing liveness is tracked via poll().+ return True

Outside this hunk (optional, for consistency across the hierarchy), consider updating the abstract method to advertise the contract:

# In BaseSpawner@abstractmethoddefbootup(self, command_lst: list[str]) ->bool: ...
🧹 Nitpick comments (2)
executorlib/standalone/interactive/communication.py (1)

109-121: Nit: docstring type should match annotation

Tighten the docstring type to list[str] for consistency.

- Args:- command_lst (list): list of strings to start the client process+ Args:+ command_lst (list[str]): list of strings to start the client process
tests/test_standalone_interactive_communication.py (1)

122-134: Stabilize the failing-command test and add a timeout

["bash", "exit"] relies on a file named “exit” and may behave differently across systems. Use bash -c 'exit 1' and add a bounded wait to avoid potential infinite loops.

- success_flag = interface.bootup(command_lst=["bash", "exit"])+ success_flag = interface.bootup(command_lst=["bash", "-c", "exit 1"])
self.assertTrue(success_flag)
- while interface._spawner.poll():- sleep(0.1)+ # Wait up to ~20s for the child to exit+ for _ in range(200):+ if not interface._spawner.poll():+ break+ sleep(0.1)
self.assertFalse(interface._spawner.poll())
interface.shutdown(wait=True)
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between e05297e and cf5088f.

📒 Files selected for processing (4)
  • executorlib/standalone/interactive/communication.py (1 hunks)
  • executorlib/standalone/interactive/spawner.py (1 hunks)
  • executorlib/task_scheduler/interactive/spawner_flux.py (2 hunks)
  • tests/test_standalone_interactive_communication.py (8 hunks)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (7)
  • GitHub Check: unittest_win
  • GitHub Check: notebooks_integration
  • GitHub Check: unittest_slurm_mpich
  • GitHub Check: unittest_openmpi (macos-latest, 3.13)
  • GitHub Check: benchmark (ubuntu-latest, 3.13, .ci_support/environment-openmpi.yml)
  • GitHub Check: notebooks
  • GitHub Check: benchmark (ubuntu-latest, 3.13, .ci_support/environment-mpich.yml)
🔇 Additional comments (2)
executorlib/standalone/interactive/communication.py (1)

109-121: LGTM: bootup now propagates spawner’s status

Returning the spawner’s bootup result from SocketInterface is the right API surface.

tests/test_standalone_interactive_communication.py (1)

39-57: LGTM: tests assert bootup’s boolean contract

Capturing and asserting the bootup success flag across scenarios looks good.

Also applies to: 69-87, 99-117, 142-160

Comment on lines 75 to 90
def bootup(
self,
command_lst: list[str],
):
) -> bool:
"""
Boot up the client process to connect to the SocketInterface.

Args:
command_lst (list[str]): List of strings to start the client process.

Raises:
ValueError: If oversubscribing is not supported for the Flux adapter or if conda environments are not supported.

Returns:
bool: Whether the interface was successfully started.
"""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

Align Flux bootup semantics to avoid race with fast failures

Same flakiness risk: a job that fails near-instantly can make poll() return False at bootup. Return True on successful submit and keep poll() for liveness.

- def bootup(+ def bootup(
self,
command_lst: list[str],
) -> bool:
@@
- return self.poll()+ # Submit succeeded → startup succeeded. Use poll() for liveness checks.+ return True

Also applies to: 133-134

🤖 Prompt for AI Agents
In executorlib/task_scheduler/interactive/spawner_flux.py around lines 75-90
(and likewise adjust lines 133-134), the bootup method currently treats an
immediate poll() False as startup failure which races with very fast job
failures; instead, after a successful job submission return True to indicate
boot completed and rely on poll() only for later liveness checks. Modify bootup
so it verifies submission succeeded (no exceptions and valid job handle),
returns True immediately on submit success even if poll() currently returns
False, and keep poll() used elsewhere to detect runtime termination; apply the
same semantics to the similar logic at lines 133-134.

@jan-janssen
jan-janssen merged commit 3524942 into mainSep 7, 2025
109 of 119 checks passed
@jan-janssen
jan-janssen deleted the spawner_success branch September 7, 2025 11:49
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jan-janssen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Spawner: Confirm startup was successful - #810

Merged
jan-janssen merged 2 commits into
mainfrom
spawner_success
Sep 7, 2025
Merged

Spawner: Confirm startup was successful#810
jan-janssen merged 2 commits into
mainfrom
spawner_success

Conversation

@jan-janssen

@jan-janssenjan-janssen commented Sep 7, 2025

Copy link
Copy Markdown
Member

Summary by CodeRabbit

  • New Features
    • bootup now returns a boolean success flag across interactive interfaces/spawners, enabling immediate startup status checks.
  • Documentation
    • Updated method docs to describe the boolean return value and expected behavior.
  • Refactor
    • Unified bootup behavior to propagate running status from underlying processes for consistent API semantics.
  • Tests
    • Updated existing tests to assert the success flag.
    • Added an error-handling test that verifies graceful termination when given an invalid command.

@coderabbitai

coderabbitaiBot commented Sep 7, 2025

Copy link
Copy Markdown
Contributor

Walkthrough

Bootup methods now return a boolean across interface and spawner components. SocketInterface.bootup returns the underlying spawner’s status. SubprocessSpawner.bootup and FluxPythonSpawner.bootup return poll() results post-start. Tests updated to assert the boolean success flag and include a new error-handling test for a failing serial command.

Changes

Cohort / File(s)Summary of Changes
Standalone interactive: bootup return propagation
executorlib/standalone/interactive/communication.py, executorlib/standalone/interactive/spawner.py
SocketInterface.bootup annotated to return bool and now returns the spawner’s bootup result. SubprocessSpawner.bootup returns poll() after launching, converting prior no-return into a boolean status. Docstrings/signatures updated accordingly.
Task scheduler Flux spawner
executorlib/task_scheduler/interactive/spawner_flux.py
FluxPythonSpawner.bootup annotated to return bool and now returns poll() after scheduling the job. Docstring updated. No other functional changes.
Tests updated for boolean bootup and error path
tests/test_standalone_interactive_communication.py
Tests now capture and assert the bootup boolean. Added test_interface_serial_with_error to validate behavior with a failing command and graceful shutdown. Minor imports and waits added.

Sequence Diagram(s)

sequenceDiagram
autonumber
participant T as Test
participant I as SocketInterface
participant S as Spawner (Subprocess/Flux)
participant P as Process/Job
T->>I: bootup(command_lst)
I->>S: bootup(command_lst)
S->>P: start/schedule
S->>S: poll() running?
S-->>I: bool (running status)
I-->>T: bool (propagated)
rect rgba(230,255,230,0.5)
note right of T: Caller can immediately know if process/job is running
end
Loading

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

  • Rename Interface to Spawner #398 — Also modifies Spawner/SocketInterface pathways, renaming/rewiring Interface → Spawner and adjusting types; closely related to return-type changes here.

Poem

A rabbit taps the bootup drum—thump, thump—
Now truth returns as bool, no guesswork slump.
Spawners poll, interfaces tell,
Tests nod yes: “All is well.”
If a job hops out of run, we still shut down—done! 🐇✅

Warning

Review ran into problems

🔥 Problems

Git: Failed to clone repository. Please run the @coderabbitai full review command to re-trigger a full review. If the issue persists, set path_filters to include or exclude specific files.

✨ Finishing Touches
  • 📝 Generate Docstrings
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch spawner_success

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@codecov

codecovBot commented Sep 7, 2025

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 97.77%. Comparing base (f1847a0) to head (8cfc8fa).
⚠️ Report is 2 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #810 +/- ##
=======================================
Coverage 97.76% 97.77% =======================================
Files 32 32 Lines 1479 1481 +2 =======================================
+ Hits 1446 1448 +2 
Misses 33 33 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
executorlib/standalone/interactive/spawner.py (1)

87-105: Make bootup success non-racy and type it as bool

Returning self.poll() immediately after Popen can intermittently be False for very short-lived commands, making startup detection flaky. Treat a successful Popen as “startup succeeded” and annotate the return type.

- def bootup(+ def bootup(
self,
command_lst: list[str],
- ):+ ) -> bool:
@@
- return self.poll()+ # Successful process creation == successful bootup; ongoing liveness is tracked via poll().+ return True

Outside this hunk (optional, for consistency across the hierarchy), consider updating the abstract method to advertise the contract:

# In BaseSpawner@abstractmethoddefbootup(self, command_lst: list[str]) ->bool: ...
🧹 Nitpick comments (2)
executorlib/standalone/interactive/communication.py (1)

109-121: Nit: docstring type should match annotation

Tighten the docstring type to list[str] for consistency.

- Args:- command_lst (list): list of strings to start the client process+ Args:+ command_lst (list[str]): list of strings to start the client process
tests/test_standalone_interactive_communication.py (1)

122-134: Stabilize the failing-command test and add a timeout

["bash", "exit"] relies on a file named “exit” and may behave differently across systems. Use bash -c 'exit 1' and add a bounded wait to avoid potential infinite loops.

- success_flag = interface.bootup(command_lst=["bash", "exit"])+ success_flag = interface.bootup(command_lst=["bash", "-c", "exit 1"])
self.assertTrue(success_flag)
- while interface._spawner.poll():- sleep(0.1)+ # Wait up to ~20s for the child to exit+ for _ in range(200):+ if not interface._spawner.poll():+ break+ sleep(0.1)
self.assertFalse(interface._spawner.poll())
interface.shutdown(wait=True)
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between e05297e and cf5088f.

📒 Files selected for processing (4)
  • executorlib/standalone/interactive/communication.py (1 hunks)
  • executorlib/standalone/interactive/spawner.py (1 hunks)
  • executorlib/task_scheduler/interactive/spawner_flux.py (2 hunks)
  • tests/test_standalone_interactive_communication.py (8 hunks)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (7)
  • GitHub Check: unittest_win
  • GitHub Check: notebooks_integration
  • GitHub Check: unittest_slurm_mpich
  • GitHub Check: unittest_openmpi (macos-latest, 3.13)
  • GitHub Check: benchmark (ubuntu-latest, 3.13, .ci_support/environment-openmpi.yml)
  • GitHub Check: notebooks
  • GitHub Check: benchmark (ubuntu-latest, 3.13, .ci_support/environment-mpich.yml)
🔇 Additional comments (2)
executorlib/standalone/interactive/communication.py (1)

109-121: LGTM: bootup now propagates spawner’s status

Returning the spawner’s bootup result from SocketInterface is the right API surface.

tests/test_standalone_interactive_communication.py (1)

39-57: LGTM: tests assert bootup’s boolean contract

Capturing and asserting the bootup success flag across scenarios looks good.

Also applies to: 69-87, 99-117, 142-160

Comment on lines 75 to 90
def bootup(
self,
command_lst: list[str],
):
) -> bool:
"""
Boot up the client process to connect to the SocketInterface.

Args:
command_lst (list[str]): List of strings to start the client process.

Raises:
ValueError: If oversubscribing is not supported for the Flux adapter or if conda environments are not supported.

Returns:
bool: Whether the interface was successfully started.
"""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

Align Flux bootup semantics to avoid race with fast failures

Same flakiness risk: a job that fails near-instantly can make poll() return False at bootup. Return True on successful submit and keep poll() for liveness.

- def bootup(+ def bootup(
self,
command_lst: list[str],
) -> bool:
@@
- return self.poll()+ # Submit succeeded → startup succeeded. Use poll() for liveness checks.+ return True

Also applies to: 133-134

🤖 Prompt for AI Agents
In executorlib/task_scheduler/interactive/spawner_flux.py around lines 75-90
(and likewise adjust lines 133-134), the bootup method currently treats an
immediate poll() False as startup failure which races with very fast job
failures; instead, after a successful job submission return True to indicate
boot completed and rely on poll() only for later liveness checks. Modify bootup
so it verifies submission succeeded (no exceptions and valid job handle),
returns True immediately on submit success even if poll() currently returns
False, and keep poll() used elsewhere to detect runtime termination; apply the
same semantics to the similar logic at lines 133-134.

@jan-janssen
jan-janssen merged commit 3524942 into mainSep 7, 2025
109 of 119 checks passed
@jan-janssen
jan-janssen deleted the spawner_success branch September 7, 2025 11:49
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jan-janssen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Spawner: Confirm startup was successful - #810

Merged
jan-janssen merged 2 commits into
mainfrom
spawner_success
Sep 7, 2025
Merged

Spawner: Confirm startup was successful#810
jan-janssen merged 2 commits into
mainfrom
spawner_success

Conversation

@jan-janssen

@jan-janssenjan-janssen commented Sep 7, 2025

Copy link
Copy Markdown
Member

Summary by CodeRabbit

  • New Features
    • bootup now returns a boolean success flag across interactive interfaces/spawners, enabling immediate startup status checks.
  • Documentation
    • Updated method docs to describe the boolean return value and expected behavior.
  • Refactor
    • Unified bootup behavior to propagate running status from underlying processes for consistent API semantics.
  • Tests
    • Updated existing tests to assert the success flag.
    • Added an error-handling test that verifies graceful termination when given an invalid command.

@coderabbitai

coderabbitaiBot commented Sep 7, 2025

Copy link
Copy Markdown
Contributor

Walkthrough

Bootup methods now return a boolean across interface and spawner components. SocketInterface.bootup returns the underlying spawner’s status. SubprocessSpawner.bootup and FluxPythonSpawner.bootup return poll() results post-start. Tests updated to assert the boolean success flag and include a new error-handling test for a failing serial command.

Changes

Cohort / File(s)Summary of Changes
Standalone interactive: bootup return propagation
executorlib/standalone/interactive/communication.py, executorlib/standalone/interactive/spawner.py
SocketInterface.bootup annotated to return bool and now returns the spawner’s bootup result. SubprocessSpawner.bootup returns poll() after launching, converting prior no-return into a boolean status. Docstrings/signatures updated accordingly.
Task scheduler Flux spawner
executorlib/task_scheduler/interactive/spawner_flux.py
FluxPythonSpawner.bootup annotated to return bool and now returns poll() after scheduling the job. Docstring updated. No other functional changes.
Tests updated for boolean bootup and error path
tests/test_standalone_interactive_communication.py
Tests now capture and assert the bootup boolean. Added test_interface_serial_with_error to validate behavior with a failing command and graceful shutdown. Minor imports and waits added.

Sequence Diagram(s)

sequenceDiagram
autonumber
participant T as Test
participant I as SocketInterface
participant S as Spawner (Subprocess/Flux)
participant P as Process/Job
T->>I: bootup(command_lst)
I->>S: bootup(command_lst)
S->>P: start/schedule
S->>S: poll() running?
S-->>I: bool (running status)
I-->>T: bool (propagated)
rect rgba(230,255,230,0.5)
note right of T: Caller can immediately know if process/job is running
end
Loading

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

  • Rename Interface to Spawner #398 — Also modifies Spawner/SocketInterface pathways, renaming/rewiring Interface → Spawner and adjusting types; closely related to return-type changes here.

Poem

A rabbit taps the bootup drum—thump, thump—
Now truth returns as bool, no guesswork slump.
Spawners poll, interfaces tell,
Tests nod yes: “All is well.”
If a job hops out of run, we still shut down—done! 🐇✅

Warning

Review ran into problems

🔥 Problems

Git: Failed to clone repository. Please run the @coderabbitai full review command to re-trigger a full review. If the issue persists, set path_filters to include or exclude specific files.

✨ Finishing Touches
  • 📝 Generate Docstrings
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch spawner_success

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@codecov

codecovBot commented Sep 7, 2025

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 97.77%. Comparing base (f1847a0) to head (8cfc8fa).
⚠️ Report is 2 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #810 +/- ##
=======================================
Coverage 97.76% 97.77% =======================================
Files 32 32 Lines 1479 1481 +2 =======================================
+ Hits 1446 1448 +2 
Misses 33 33 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
executorlib/standalone/interactive/spawner.py (1)

87-105: Make bootup success non-racy and type it as bool

Returning self.poll() immediately after Popen can intermittently be False for very short-lived commands, making startup detection flaky. Treat a successful Popen as “startup succeeded” and annotate the return type.

- def bootup(+ def bootup(
self,
command_lst: list[str],
- ):+ ) -> bool:
@@
- return self.poll()+ # Successful process creation == successful bootup; ongoing liveness is tracked via poll().+ return True

Outside this hunk (optional, for consistency across the hierarchy), consider updating the abstract method to advertise the contract:

# In BaseSpawner@abstractmethoddefbootup(self, command_lst: list[str]) ->bool: ...
🧹 Nitpick comments (2)
executorlib/standalone/interactive/communication.py (1)

109-121: Nit: docstring type should match annotation

Tighten the docstring type to list[str] for consistency.

- Args:- command_lst (list): list of strings to start the client process+ Args:+ command_lst (list[str]): list of strings to start the client process
tests/test_standalone_interactive_communication.py (1)

122-134: Stabilize the failing-command test and add a timeout

["bash", "exit"] relies on a file named “exit” and may behave differently across systems. Use bash -c 'exit 1' and add a bounded wait to avoid potential infinite loops.

- success_flag = interface.bootup(command_lst=["bash", "exit"])+ success_flag = interface.bootup(command_lst=["bash", "-c", "exit 1"])
self.assertTrue(success_flag)
- while interface._spawner.poll():- sleep(0.1)+ # Wait up to ~20s for the child to exit+ for _ in range(200):+ if not interface._spawner.poll():+ break+ sleep(0.1)
self.assertFalse(interface._spawner.poll())
interface.shutdown(wait=True)
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between e05297e and cf5088f.

📒 Files selected for processing (4)
  • executorlib/standalone/interactive/communication.py (1 hunks)
  • executorlib/standalone/interactive/spawner.py (1 hunks)
  • executorlib/task_scheduler/interactive/spawner_flux.py (2 hunks)
  • tests/test_standalone_interactive_communication.py (8 hunks)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (7)
  • GitHub Check: unittest_win
  • GitHub Check: notebooks_integration
  • GitHub Check: unittest_slurm_mpich
  • GitHub Check: unittest_openmpi (macos-latest, 3.13)
  • GitHub Check: benchmark (ubuntu-latest, 3.13, .ci_support/environment-openmpi.yml)
  • GitHub Check: notebooks
  • GitHub Check: benchmark (ubuntu-latest, 3.13, .ci_support/environment-mpich.yml)
🔇 Additional comments (2)
executorlib/standalone/interactive/communication.py (1)

109-121: LGTM: bootup now propagates spawner’s status

Returning the spawner’s bootup result from SocketInterface is the right API surface.

tests/test_standalone_interactive_communication.py (1)

39-57: LGTM: tests assert bootup’s boolean contract

Capturing and asserting the bootup success flag across scenarios looks good.

Also applies to: 69-87, 99-117, 142-160

Comment on lines 75 to 90
def bootup(
self,
command_lst: list[str],
):
) -> bool:
"""
Boot up the client process to connect to the SocketInterface.

Args:
command_lst (list[str]): List of strings to start the client process.

Raises:
ValueError: If oversubscribing is not supported for the Flux adapter or if conda environments are not supported.

Returns:
bool: Whether the interface was successfully started.
"""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

Align Flux bootup semantics to avoid race with fast failures

Same flakiness risk: a job that fails near-instantly can make poll() return False at bootup. Return True on successful submit and keep poll() for liveness.

- def bootup(+ def bootup(
self,
command_lst: list[str],
) -> bool:
@@
- return self.poll()+ # Submit succeeded → startup succeeded. Use poll() for liveness checks.+ return True

Also applies to: 133-134

🤖 Prompt for AI Agents
In executorlib/task_scheduler/interactive/spawner_flux.py around lines 75-90
(and likewise adjust lines 133-134), the bootup method currently treats an
immediate poll() False as startup failure which races with very fast job
failures; instead, after a successful job submission return True to indicate
boot completed and rely on poll() only for later liveness checks. Modify bootup
so it verifies submission succeeded (no exceptions and valid job handle),
returns True immediately on submit success even if poll() currently returns
False, and keep poll() used elsewhere to detect runtime termination; apply the
same semantics to the similar logic at lines 133-134.

@jan-janssen
jan-janssen merged commit 3524942 into mainSep 7, 2025
109 of 119 checks passed
@jan-janssen
jan-janssen deleted the spawner_success branch September 7, 2025 11:49
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jan-janssen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Spawner: Confirm startup was successful - #810

Merged
jan-janssen merged 2 commits into
mainfrom
spawner_success
Sep 7, 2025
Merged

Spawner: Confirm startup was successful#810
jan-janssen merged 2 commits into
mainfrom
spawner_success

Conversation

@jan-janssen

@jan-janssenjan-janssen commented Sep 7, 2025

Copy link
Copy Markdown
Member

Summary by CodeRabbit

  • New Features
    • bootup now returns a boolean success flag across interactive interfaces/spawners, enabling immediate startup status checks.
  • Documentation
    • Updated method docs to describe the boolean return value and expected behavior.
  • Refactor
    • Unified bootup behavior to propagate running status from underlying processes for consistent API semantics.
  • Tests
    • Updated existing tests to assert the success flag.
    • Added an error-handling test that verifies graceful termination when given an invalid command.

@coderabbitai

coderabbitaiBot commented Sep 7, 2025

Copy link
Copy Markdown
Contributor

Walkthrough

Bootup methods now return a boolean across interface and spawner components. SocketInterface.bootup returns the underlying spawner’s status. SubprocessSpawner.bootup and FluxPythonSpawner.bootup return poll() results post-start. Tests updated to assert the boolean success flag and include a new error-handling test for a failing serial command.

Changes

Cohort / File(s)Summary of Changes
Standalone interactive: bootup return propagation
executorlib/standalone/interactive/communication.py, executorlib/standalone/interactive/spawner.py
SocketInterface.bootup annotated to return bool and now returns the spawner’s bootup result. SubprocessSpawner.bootup returns poll() after launching, converting prior no-return into a boolean status. Docstrings/signatures updated accordingly.
Task scheduler Flux spawner
executorlib/task_scheduler/interactive/spawner_flux.py
FluxPythonSpawner.bootup annotated to return bool and now returns poll() after scheduling the job. Docstring updated. No other functional changes.
Tests updated for boolean bootup and error path
tests/test_standalone_interactive_communication.py
Tests now capture and assert the bootup boolean. Added test_interface_serial_with_error to validate behavior with a failing command and graceful shutdown. Minor imports and waits added.

Sequence Diagram(s)

sequenceDiagram
autonumber
participant T as Test
participant I as SocketInterface
participant S as Spawner (Subprocess/Flux)
participant P as Process/Job
T->>I: bootup(command_lst)
I->>S: bootup(command_lst)
S->>P: start/schedule
S->>S: poll() running?
S-->>I: bool (running status)
I-->>T: bool (propagated)
rect rgba(230,255,230,0.5)
note right of T: Caller can immediately know if process/job is running
end
Loading

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

  • Rename Interface to Spawner #398 — Also modifies Spawner/SocketInterface pathways, renaming/rewiring Interface → Spawner and adjusting types; closely related to return-type changes here.

Poem

A rabbit taps the bootup drum—thump, thump—
Now truth returns as bool, no guesswork slump.
Spawners poll, interfaces tell,
Tests nod yes: “All is well.”
If a job hops out of run, we still shut down—done! 🐇✅

Warning

Review ran into problems

🔥 Problems

Git: Failed to clone repository. Please run the @coderabbitai full review command to re-trigger a full review. If the issue persists, set path_filters to include or exclude specific files.

✨ Finishing Touches
  • 📝 Generate Docstrings
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch spawner_success

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@codecov

codecovBot commented Sep 7, 2025

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 97.77%. Comparing base (f1847a0) to head (8cfc8fa).
⚠️ Report is 2 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #810 +/- ##
=======================================
Coverage 97.76% 97.77% =======================================
Files 32 32 Lines 1479 1481 +2 =======================================
+ Hits 1446 1448 +2 
Misses 33 33 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
executorlib/standalone/interactive/spawner.py (1)

87-105: Make bootup success non-racy and type it as bool

Returning self.poll() immediately after Popen can intermittently be False for very short-lived commands, making startup detection flaky. Treat a successful Popen as “startup succeeded” and annotate the return type.

- def bootup(+ def bootup(
self,
command_lst: list[str],
- ):+ ) -> bool:
@@
- return self.poll()+ # Successful process creation == successful bootup; ongoing liveness is tracked via poll().+ return True

Outside this hunk (optional, for consistency across the hierarchy), consider updating the abstract method to advertise the contract:

# In BaseSpawner@abstractmethoddefbootup(self, command_lst: list[str]) ->bool: ...
🧹 Nitpick comments (2)
executorlib/standalone/interactive/communication.py (1)

109-121: Nit: docstring type should match annotation

Tighten the docstring type to list[str] for consistency.

- Args:- command_lst (list): list of strings to start the client process+ Args:+ command_lst (list[str]): list of strings to start the client process
tests/test_standalone_interactive_communication.py (1)

122-134: Stabilize the failing-command test and add a timeout

["bash", "exit"] relies on a file named “exit” and may behave differently across systems. Use bash -c 'exit 1' and add a bounded wait to avoid potential infinite loops.

- success_flag = interface.bootup(command_lst=["bash", "exit"])+ success_flag = interface.bootup(command_lst=["bash", "-c", "exit 1"])
self.assertTrue(success_flag)
- while interface._spawner.poll():- sleep(0.1)+ # Wait up to ~20s for the child to exit+ for _ in range(200):+ if not interface._spawner.poll():+ break+ sleep(0.1)
self.assertFalse(interface._spawner.poll())
interface.shutdown(wait=True)
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between e05297e and cf5088f.

📒 Files selected for processing (4)
  • executorlib/standalone/interactive/communication.py (1 hunks)
  • executorlib/standalone/interactive/spawner.py (1 hunks)
  • executorlib/task_scheduler/interactive/spawner_flux.py (2 hunks)
  • tests/test_standalone_interactive_communication.py (8 hunks)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (7)
  • GitHub Check: unittest_win
  • GitHub Check: notebooks_integration
  • GitHub Check: unittest_slurm_mpich
  • GitHub Check: unittest_openmpi (macos-latest, 3.13)
  • GitHub Check: benchmark (ubuntu-latest, 3.13, .ci_support/environment-openmpi.yml)
  • GitHub Check: notebooks
  • GitHub Check: benchmark (ubuntu-latest, 3.13, .ci_support/environment-mpich.yml)
🔇 Additional comments (2)
executorlib/standalone/interactive/communication.py (1)

109-121: LGTM: bootup now propagates spawner’s status

Returning the spawner’s bootup result from SocketInterface is the right API surface.

tests/test_standalone_interactive_communication.py (1)

39-57: LGTM: tests assert bootup’s boolean contract

Capturing and asserting the bootup success flag across scenarios looks good.

Also applies to: 69-87, 99-117, 142-160

Comment on lines 75 to 90
def bootup(
self,
command_lst: list[str],
):
) -> bool:
"""
Boot up the client process to connect to the SocketInterface.

Args:
command_lst (list[str]): List of strings to start the client process.

Raises:
ValueError: If oversubscribing is not supported for the Flux adapter or if conda environments are not supported.

Returns:
bool: Whether the interface was successfully started.
"""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

Align Flux bootup semantics to avoid race with fast failures

Same flakiness risk: a job that fails near-instantly can make poll() return False at bootup. Return True on successful submit and keep poll() for liveness.

- def bootup(+ def bootup(
self,
command_lst: list[str],
) -> bool:
@@
- return self.poll()+ # Submit succeeded → startup succeeded. Use poll() for liveness checks.+ return True

Also applies to: 133-134

🤖 Prompt for AI Agents
In executorlib/task_scheduler/interactive/spawner_flux.py around lines 75-90
(and likewise adjust lines 133-134), the bootup method currently treats an
immediate poll() False as startup failure which races with very fast job
failures; instead, after a successful job submission return True to indicate
boot completed and rely on poll() only for later liveness checks. Modify bootup
so it verifies submission succeeded (no exceptions and valid job handle),
returns True immediately on submit success even if poll() currently returns
False, and keep poll() used elsewhere to detect runtime termination; apply the
same semantics to the similar logic at lines 133-134.

@jan-janssen
jan-janssen merged commit 3524942 into mainSep 7, 2025
109 of 119 checks passed
@jan-janssen
jan-janssen deleted the spawner_success branch September 7, 2025 11:49
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jan-janssen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Spawner: Confirm startup was successful - #810

Merged
jan-janssen merged 2 commits into
mainfrom
spawner_success
Sep 7, 2025
Merged

Spawner: Confirm startup was successful#810
jan-janssen merged 2 commits into
mainfrom
spawner_success

Conversation

@jan-janssen

@jan-janssenjan-janssen commented Sep 7, 2025

Copy link
Copy Markdown
Member

Summary by CodeRabbit

  • New Features
    • bootup now returns a boolean success flag across interactive interfaces/spawners, enabling immediate startup status checks.
  • Documentation
    • Updated method docs to describe the boolean return value and expected behavior.
  • Refactor
    • Unified bootup behavior to propagate running status from underlying processes for consistent API semantics.
  • Tests
    • Updated existing tests to assert the success flag.
    • Added an error-handling test that verifies graceful termination when given an invalid command.

@coderabbitai

coderabbitaiBot commented Sep 7, 2025

Copy link
Copy Markdown
Contributor

Walkthrough

Bootup methods now return a boolean across interface and spawner components. SocketInterface.bootup returns the underlying spawner’s status. SubprocessSpawner.bootup and FluxPythonSpawner.bootup return poll() results post-start. Tests updated to assert the boolean success flag and include a new error-handling test for a failing serial command.

Changes

Cohort / File(s)Summary of Changes
Standalone interactive: bootup return propagation
executorlib/standalone/interactive/communication.py, executorlib/standalone/interactive/spawner.py
SocketInterface.bootup annotated to return bool and now returns the spawner’s bootup result. SubprocessSpawner.bootup returns poll() after launching, converting prior no-return into a boolean status. Docstrings/signatures updated accordingly.
Task scheduler Flux spawner
executorlib/task_scheduler/interactive/spawner_flux.py
FluxPythonSpawner.bootup annotated to return bool and now returns poll() after scheduling the job. Docstring updated. No other functional changes.
Tests updated for boolean bootup and error path
tests/test_standalone_interactive_communication.py
Tests now capture and assert the bootup boolean. Added test_interface_serial_with_error to validate behavior with a failing command and graceful shutdown. Minor imports and waits added.

Sequence Diagram(s)

sequenceDiagram
autonumber
participant T as Test
participant I as SocketInterface
participant S as Spawner (Subprocess/Flux)
participant P as Process/Job
T->>I: bootup(command_lst)
I->>S: bootup(command_lst)
S->>P: start/schedule
S->>S: poll() running?
S-->>I: bool (running status)
I-->>T: bool (propagated)
rect rgba(230,255,230,0.5)
note right of T: Caller can immediately know if process/job is running
end
Loading

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

  • Rename Interface to Spawner #398 — Also modifies Spawner/SocketInterface pathways, renaming/rewiring Interface → Spawner and adjusting types; closely related to return-type changes here.

Poem

A rabbit taps the bootup drum—thump, thump—
Now truth returns as bool, no guesswork slump.
Spawners poll, interfaces tell,
Tests nod yes: “All is well.”
If a job hops out of run, we still shut down—done! 🐇✅

Warning

Review ran into problems

🔥 Problems

Git: Failed to clone repository. Please run the @coderabbitai full review command to re-trigger a full review. If the issue persists, set path_filters to include or exclude specific files.

✨ Finishing Touches
  • 📝 Generate Docstrings
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch spawner_success

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@codecov

codecovBot commented Sep 7, 2025

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 97.77%. Comparing base (f1847a0) to head (8cfc8fa).
⚠️ Report is 2 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #810 +/- ##
=======================================
Coverage 97.76% 97.77% =======================================
Files 32 32 Lines 1479 1481 +2 =======================================
+ Hits 1446 1448 +2 
Misses 33 33 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
executorlib/standalone/interactive/spawner.py (1)

87-105: Make bootup success non-racy and type it as bool

Returning self.poll() immediately after Popen can intermittently be False for very short-lived commands, making startup detection flaky. Treat a successful Popen as “startup succeeded” and annotate the return type.

- def bootup(+ def bootup(
self,
command_lst: list[str],
- ):+ ) -> bool:
@@
- return self.poll()+ # Successful process creation == successful bootup; ongoing liveness is tracked via poll().+ return True

Outside this hunk (optional, for consistency across the hierarchy), consider updating the abstract method to advertise the contract:

# In BaseSpawner@abstractmethoddefbootup(self, command_lst: list[str]) ->bool: ...
🧹 Nitpick comments (2)
executorlib/standalone/interactive/communication.py (1)

109-121: Nit: docstring type should match annotation

Tighten the docstring type to list[str] for consistency.

- Args:- command_lst (list): list of strings to start the client process+ Args:+ command_lst (list[str]): list of strings to start the client process
tests/test_standalone_interactive_communication.py (1)

122-134: Stabilize the failing-command test and add a timeout

["bash", "exit"] relies on a file named “exit” and may behave differently across systems. Use bash -c 'exit 1' and add a bounded wait to avoid potential infinite loops.

- success_flag = interface.bootup(command_lst=["bash", "exit"])+ success_flag = interface.bootup(command_lst=["bash", "-c", "exit 1"])
self.assertTrue(success_flag)
- while interface._spawner.poll():- sleep(0.1)+ # Wait up to ~20s for the child to exit+ for _ in range(200):+ if not interface._spawner.poll():+ break+ sleep(0.1)
self.assertFalse(interface._spawner.poll())
interface.shutdown(wait=True)
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between e05297e and cf5088f.

📒 Files selected for processing (4)
  • executorlib/standalone/interactive/communication.py (1 hunks)
  • executorlib/standalone/interactive/spawner.py (1 hunks)
  • executorlib/task_scheduler/interactive/spawner_flux.py (2 hunks)
  • tests/test_standalone_interactive_communication.py (8 hunks)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (7)
  • GitHub Check: unittest_win
  • GitHub Check: notebooks_integration
  • GitHub Check: unittest_slurm_mpich
  • GitHub Check: unittest_openmpi (macos-latest, 3.13)
  • GitHub Check: benchmark (ubuntu-latest, 3.13, .ci_support/environment-openmpi.yml)
  • GitHub Check: notebooks
  • GitHub Check: benchmark (ubuntu-latest, 3.13, .ci_support/environment-mpich.yml)
🔇 Additional comments (2)
executorlib/standalone/interactive/communication.py (1)

109-121: LGTM: bootup now propagates spawner’s status

Returning the spawner’s bootup result from SocketInterface is the right API surface.

tests/test_standalone_interactive_communication.py (1)

39-57: LGTM: tests assert bootup’s boolean contract

Capturing and asserting the bootup success flag across scenarios looks good.

Also applies to: 69-87, 99-117, 142-160

Comment on lines 75 to 90
def bootup(
self,
command_lst: list[str],
):
) -> bool:
"""
Boot up the client process to connect to the SocketInterface.

Args:
command_lst (list[str]): List of strings to start the client process.

Raises:
ValueError: If oversubscribing is not supported for the Flux adapter or if conda environments are not supported.

Returns:
bool: Whether the interface was successfully started.
"""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

Align Flux bootup semantics to avoid race with fast failures

Same flakiness risk: a job that fails near-instantly can make poll() return False at bootup. Return True on successful submit and keep poll() for liveness.

- def bootup(+ def bootup(
self,
command_lst: list[str],
) -> bool:
@@
- return self.poll()+ # Submit succeeded → startup succeeded. Use poll() for liveness checks.+ return True

Also applies to: 133-134

🤖 Prompt for AI Agents
In executorlib/task_scheduler/interactive/spawner_flux.py around lines 75-90
(and likewise adjust lines 133-134), the bootup method currently treats an
immediate poll() False as startup failure which races with very fast job
failures; instead, after a successful job submission return True to indicate
boot completed and rely on poll() only for later liveness checks. Modify bootup
so it verifies submission succeeded (no exceptions and valid job handle),
returns True immediately on submit success even if poll() currently returns
False, and keep poll() used elsewhere to detect runtime termination; apply the
same semantics to the similar logic at lines 133-134.

@jan-janssen
jan-janssen merged commit 3524942 into mainSep 7, 2025
109 of 119 checks passed
@jan-janssen
jan-janssen deleted the spawner_success branch September 7, 2025 11:49
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jan-janssen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Spawner: Confirm startup was successful - #810

Merged
jan-janssen merged 2 commits into
mainfrom
spawner_success
Sep 7, 2025
Merged

Spawner: Confirm startup was successful#810
jan-janssen merged 2 commits into
mainfrom
spawner_success

Conversation

@jan-janssen

@jan-janssenjan-janssen commented Sep 7, 2025

Copy link
Copy Markdown
Member

Summary by CodeRabbit

  • New Features
    • bootup now returns a boolean success flag across interactive interfaces/spawners, enabling immediate startup status checks.
  • Documentation
    • Updated method docs to describe the boolean return value and expected behavior.
  • Refactor
    • Unified bootup behavior to propagate running status from underlying processes for consistent API semantics.
  • Tests
    • Updated existing tests to assert the success flag.
    • Added an error-handling test that verifies graceful termination when given an invalid command.

@coderabbitai

coderabbitaiBot commented Sep 7, 2025

Copy link
Copy Markdown
Contributor

Walkthrough

Bootup methods now return a boolean across interface and spawner components. SocketInterface.bootup returns the underlying spawner’s status. SubprocessSpawner.bootup and FluxPythonSpawner.bootup return poll() results post-start. Tests updated to assert the boolean success flag and include a new error-handling test for a failing serial command.

Changes

Cohort / File(s)Summary of Changes
Standalone interactive: bootup return propagation
executorlib/standalone/interactive/communication.py, executorlib/standalone/interactive/spawner.py
SocketInterface.bootup annotated to return bool and now returns the spawner’s bootup result. SubprocessSpawner.bootup returns poll() after launching, converting prior no-return into a boolean status. Docstrings/signatures updated accordingly.
Task scheduler Flux spawner
executorlib/task_scheduler/interactive/spawner_flux.py
FluxPythonSpawner.bootup annotated to return bool and now returns poll() after scheduling the job. Docstring updated. No other functional changes.
Tests updated for boolean bootup and error path
tests/test_standalone_interactive_communication.py
Tests now capture and assert the bootup boolean. Added test_interface_serial_with_error to validate behavior with a failing command and graceful shutdown. Minor imports and waits added.

Sequence Diagram(s)

sequenceDiagram
autonumber
participant T as Test
participant I as SocketInterface
participant S as Spawner (Subprocess/Flux)
participant P as Process/Job
T->>I: bootup(command_lst)
I->>S: bootup(command_lst)
S->>P: start/schedule
S->>S: poll() running?
S-->>I: bool (running status)
I-->>T: bool (propagated)
rect rgba(230,255,230,0.5)
note right of T: Caller can immediately know if process/job is running
end
Loading

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

  • Rename Interface to Spawner #398 — Also modifies Spawner/SocketInterface pathways, renaming/rewiring Interface → Spawner and adjusting types; closely related to return-type changes here.

Poem

A rabbit taps the bootup drum—thump, thump—
Now truth returns as bool, no guesswork slump.
Spawners poll, interfaces tell,
Tests nod yes: “All is well.”
If a job hops out of run, we still shut down—done! 🐇✅

Warning

Review ran into problems

🔥 Problems

Git: Failed to clone repository. Please run the @coderabbitai full review command to re-trigger a full review. If the issue persists, set path_filters to include or exclude specific files.

✨ Finishing Touches
  • 📝 Generate Docstrings
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch spawner_success

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@codecov

codecovBot commented Sep 7, 2025

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 97.77%. Comparing base (f1847a0) to head (8cfc8fa).
⚠️ Report is 2 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #810 +/- ##
=======================================
Coverage 97.76% 97.77% =======================================
Files 32 32 Lines 1479 1481 +2 =======================================
+ Hits 1446 1448 +2 
Misses 33 33 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
executorlib/standalone/interactive/spawner.py (1)

87-105: Make bootup success non-racy and type it as bool

Returning self.poll() immediately after Popen can intermittently be False for very short-lived commands, making startup detection flaky. Treat a successful Popen as “startup succeeded” and annotate the return type.

- def bootup(+ def bootup(
self,
command_lst: list[str],
- ):+ ) -> bool:
@@
- return self.poll()+ # Successful process creation == successful bootup; ongoing liveness is tracked via poll().+ return True

Outside this hunk (optional, for consistency across the hierarchy), consider updating the abstract method to advertise the contract:

# In BaseSpawner@abstractmethoddefbootup(self, command_lst: list[str]) ->bool: ...
🧹 Nitpick comments (2)
executorlib/standalone/interactive/communication.py (1)

109-121: Nit: docstring type should match annotation

Tighten the docstring type to list[str] for consistency.

- Args:- command_lst (list): list of strings to start the client process+ Args:+ command_lst (list[str]): list of strings to start the client process
tests/test_standalone_interactive_communication.py (1)

122-134: Stabilize the failing-command test and add a timeout

["bash", "exit"] relies on a file named “exit” and may behave differently across systems. Use bash -c 'exit 1' and add a bounded wait to avoid potential infinite loops.

- success_flag = interface.bootup(command_lst=["bash", "exit"])+ success_flag = interface.bootup(command_lst=["bash", "-c", "exit 1"])
self.assertTrue(success_flag)
- while interface._spawner.poll():- sleep(0.1)+ # Wait up to ~20s for the child to exit+ for _ in range(200):+ if not interface._spawner.poll():+ break+ sleep(0.1)
self.assertFalse(interface._spawner.poll())
interface.shutdown(wait=True)
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between e05297e and cf5088f.

📒 Files selected for processing (4)
  • executorlib/standalone/interactive/communication.py (1 hunks)
  • executorlib/standalone/interactive/spawner.py (1 hunks)
  • executorlib/task_scheduler/interactive/spawner_flux.py (2 hunks)
  • tests/test_standalone_interactive_communication.py (8 hunks)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (7)
  • GitHub Check: unittest_win
  • GitHub Check: notebooks_integration
  • GitHub Check: unittest_slurm_mpich
  • GitHub Check: unittest_openmpi (macos-latest, 3.13)
  • GitHub Check: benchmark (ubuntu-latest, 3.13, .ci_support/environment-openmpi.yml)
  • GitHub Check: notebooks
  • GitHub Check: benchmark (ubuntu-latest, 3.13, .ci_support/environment-mpich.yml)
🔇 Additional comments (2)
executorlib/standalone/interactive/communication.py (1)

109-121: LGTM: bootup now propagates spawner’s status

Returning the spawner’s bootup result from SocketInterface is the right API surface.

tests/test_standalone_interactive_communication.py (1)

39-57: LGTM: tests assert bootup’s boolean contract

Capturing and asserting the bootup success flag across scenarios looks good.

Also applies to: 69-87, 99-117, 142-160

Comment on lines 75 to 90
def bootup(
self,
command_lst: list[str],
):
) -> bool:
"""
Boot up the client process to connect to the SocketInterface.

Args:
command_lst (list[str]): List of strings to start the client process.

Raises:
ValueError: If oversubscribing is not supported for the Flux adapter or if conda environments are not supported.

Returns:
bool: Whether the interface was successfully started.
"""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

Align Flux bootup semantics to avoid race with fast failures

Same flakiness risk: a job that fails near-instantly can make poll() return False at bootup. Return True on successful submit and keep poll() for liveness.

- def bootup(+ def bootup(
self,
command_lst: list[str],
) -> bool:
@@
- return self.poll()+ # Submit succeeded → startup succeeded. Use poll() for liveness checks.+ return True

Also applies to: 133-134

🤖 Prompt for AI Agents
In executorlib/task_scheduler/interactive/spawner_flux.py around lines 75-90
(and likewise adjust lines 133-134), the bootup method currently treats an
immediate poll() False as startup failure which races with very fast job
failures; instead, after a successful job submission return True to indicate
boot completed and rely on poll() only for later liveness checks. Modify bootup
so it verifies submission succeeded (no exceptions and valid job handle),
returns True immediately on submit success even if poll() currently returns
False, and keep poll() used elsewhere to detect runtime termination; apply the
same semantics to the similar logic at lines 133-134.

@jan-janssen
jan-janssen merged commit 3524942 into mainSep 7, 2025
109 of 119 checks passed
@jan-janssen
jan-janssen deleted the spawner_success branch September 7, 2025 11:49
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jan-janssen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Spawner: Confirm startup was successful - #810

Merged
jan-janssen merged 2 commits into
mainfrom
spawner_success
Sep 7, 2025
Merged

Spawner: Confirm startup was successful#810
jan-janssen merged 2 commits into
mainfrom
spawner_success

Conversation

@jan-janssen

@jan-janssenjan-janssen commented Sep 7, 2025

Copy link
Copy Markdown
Member

Summary by CodeRabbit

  • New Features
    • bootup now returns a boolean success flag across interactive interfaces/spawners, enabling immediate startup status checks.
  • Documentation
    • Updated method docs to describe the boolean return value and expected behavior.
  • Refactor
    • Unified bootup behavior to propagate running status from underlying processes for consistent API semantics.
  • Tests
    • Updated existing tests to assert the success flag.
    • Added an error-handling test that verifies graceful termination when given an invalid command.

@coderabbitai

coderabbitaiBot commented Sep 7, 2025

Copy link
Copy Markdown
Contributor

Walkthrough

Bootup methods now return a boolean across interface and spawner components. SocketInterface.bootup returns the underlying spawner’s status. SubprocessSpawner.bootup and FluxPythonSpawner.bootup return poll() results post-start. Tests updated to assert the boolean success flag and include a new error-handling test for a failing serial command.

Changes

Cohort / File(s)Summary of Changes
Standalone interactive: bootup return propagation
executorlib/standalone/interactive/communication.py, executorlib/standalone/interactive/spawner.py
SocketInterface.bootup annotated to return bool and now returns the spawner’s bootup result. SubprocessSpawner.bootup returns poll() after launching, converting prior no-return into a boolean status. Docstrings/signatures updated accordingly.
Task scheduler Flux spawner
executorlib/task_scheduler/interactive/spawner_flux.py
FluxPythonSpawner.bootup annotated to return bool and now returns poll() after scheduling the job. Docstring updated. No other functional changes.
Tests updated for boolean bootup and error path
tests/test_standalone_interactive_communication.py
Tests now capture and assert the bootup boolean. Added test_interface_serial_with_error to validate behavior with a failing command and graceful shutdown. Minor imports and waits added.

Sequence Diagram(s)

sequenceDiagram
autonumber
participant T as Test
participant I as SocketInterface
participant S as Spawner (Subprocess/Flux)
participant P as Process/Job
T->>I: bootup(command_lst)
I->>S: bootup(command_lst)
S->>P: start/schedule
S->>S: poll() running?
S-->>I: bool (running status)
I-->>T: bool (propagated)
rect rgba(230,255,230,0.5)
note right of T: Caller can immediately know if process/job is running
end
Loading

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

  • Rename Interface to Spawner #398 — Also modifies Spawner/SocketInterface pathways, renaming/rewiring Interface → Spawner and adjusting types; closely related to return-type changes here.

Poem

A rabbit taps the bootup drum—thump, thump—
Now truth returns as bool, no guesswork slump.
Spawners poll, interfaces tell,
Tests nod yes: “All is well.”
If a job hops out of run, we still shut down—done! 🐇✅

Warning

Review ran into problems

🔥 Problems

Git: Failed to clone repository. Please run the @coderabbitai full review command to re-trigger a full review. If the issue persists, set path_filters to include or exclude specific files.

✨ Finishing Touches
  • 📝 Generate Docstrings
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch spawner_success

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@codecov

codecovBot commented Sep 7, 2025

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 97.77%. Comparing base (f1847a0) to head (8cfc8fa).
⚠️ Report is 2 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #810 +/- ##
=======================================
Coverage 97.76% 97.77% =======================================
Files 32 32 Lines 1479 1481 +2 =======================================
+ Hits 1446 1448 +2 
Misses 33 33 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
executorlib/standalone/interactive/spawner.py (1)

87-105: Make bootup success non-racy and type it as bool

Returning self.poll() immediately after Popen can intermittently be False for very short-lived commands, making startup detection flaky. Treat a successful Popen as “startup succeeded” and annotate the return type.

- def bootup(+ def bootup(
self,
command_lst: list[str],
- ):+ ) -> bool:
@@
- return self.poll()+ # Successful process creation == successful bootup; ongoing liveness is tracked via poll().+ return True

Outside this hunk (optional, for consistency across the hierarchy), consider updating the abstract method to advertise the contract:

# In BaseSpawner@abstractmethoddefbootup(self, command_lst: list[str]) ->bool: ...
🧹 Nitpick comments (2)
executorlib/standalone/interactive/communication.py (1)

109-121: Nit: docstring type should match annotation

Tighten the docstring type to list[str] for consistency.

- Args:- command_lst (list): list of strings to start the client process+ Args:+ command_lst (list[str]): list of strings to start the client process
tests/test_standalone_interactive_communication.py (1)

122-134: Stabilize the failing-command test and add a timeout

["bash", "exit"] relies on a file named “exit” and may behave differently across systems. Use bash -c 'exit 1' and add a bounded wait to avoid potential infinite loops.

- success_flag = interface.bootup(command_lst=["bash", "exit"])+ success_flag = interface.bootup(command_lst=["bash", "-c", "exit 1"])
self.assertTrue(success_flag)
- while interface._spawner.poll():- sleep(0.1)+ # Wait up to ~20s for the child to exit+ for _ in range(200):+ if not interface._spawner.poll():+ break+ sleep(0.1)
self.assertFalse(interface._spawner.poll())
interface.shutdown(wait=True)
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between e05297e and cf5088f.

📒 Files selected for processing (4)
  • executorlib/standalone/interactive/communication.py (1 hunks)
  • executorlib/standalone/interactive/spawner.py (1 hunks)
  • executorlib/task_scheduler/interactive/spawner_flux.py (2 hunks)
  • tests/test_standalone_interactive_communication.py (8 hunks)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (7)
  • GitHub Check: unittest_win
  • GitHub Check: notebooks_integration
  • GitHub Check: unittest_slurm_mpich
  • GitHub Check: unittest_openmpi (macos-latest, 3.13)
  • GitHub Check: benchmark (ubuntu-latest, 3.13, .ci_support/environment-openmpi.yml)
  • GitHub Check: notebooks
  • GitHub Check: benchmark (ubuntu-latest, 3.13, .ci_support/environment-mpich.yml)
🔇 Additional comments (2)
executorlib/standalone/interactive/communication.py (1)

109-121: LGTM: bootup now propagates spawner’s status

Returning the spawner’s bootup result from SocketInterface is the right API surface.

tests/test_standalone_interactive_communication.py (1)

39-57: LGTM: tests assert bootup’s boolean contract

Capturing and asserting the bootup success flag across scenarios looks good.

Also applies to: 69-87, 99-117, 142-160

Comment on lines 75 to 90
def bootup(
self,
command_lst: list[str],
):
) -> bool:
"""
Boot up the client process to connect to the SocketInterface.

Args:
command_lst (list[str]): List of strings to start the client process.

Raises:
ValueError: If oversubscribing is not supported for the Flux adapter or if conda environments are not supported.

Returns:
bool: Whether the interface was successfully started.
"""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

Align Flux bootup semantics to avoid race with fast failures

Same flakiness risk: a job that fails near-instantly can make poll() return False at bootup. Return True on successful submit and keep poll() for liveness.

- def bootup(+ def bootup(
self,
command_lst: list[str],
) -> bool:
@@
- return self.poll()+ # Submit succeeded → startup succeeded. Use poll() for liveness checks.+ return True

Also applies to: 133-134

🤖 Prompt for AI Agents
In executorlib/task_scheduler/interactive/spawner_flux.py around lines 75-90
(and likewise adjust lines 133-134), the bootup method currently treats an
immediate poll() False as startup failure which races with very fast job
failures; instead, after a successful job submission return True to indicate
boot completed and rely on poll() only for later liveness checks. Modify bootup
so it verifies submission succeeded (no exceptions and valid job handle),
returns True immediately on submit success even if poll() currently returns
False, and keep poll() used elsewhere to detect runtime termination; apply the
same semantics to the similar logic at lines 133-134.

@jan-janssen
jan-janssen merged commit 3524942 into mainSep 7, 2025
109 of 119 checks passed
@jan-janssen
jan-janssen deleted the spawner_success branch September 7, 2025 11:49
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jan-janssen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Spawner: Confirm startup was successful - #810

Merged
jan-janssen merged 2 commits into
mainfrom
spawner_success
Sep 7, 2025
Merged

Spawner: Confirm startup was successful#810
jan-janssen merged 2 commits into
mainfrom
spawner_success

Conversation

@jan-janssen

@jan-janssenjan-janssen commented Sep 7, 2025

Copy link
Copy Markdown
Member

Summary by CodeRabbit

  • New Features
    • bootup now returns a boolean success flag across interactive interfaces/spawners, enabling immediate startup status checks.
  • Documentation
    • Updated method docs to describe the boolean return value and expected behavior.
  • Refactor
    • Unified bootup behavior to propagate running status from underlying processes for consistent API semantics.
  • Tests
    • Updated existing tests to assert the success flag.
    • Added an error-handling test that verifies graceful termination when given an invalid command.

@coderabbitai

coderabbitaiBot commented Sep 7, 2025

Copy link
Copy Markdown
Contributor

Walkthrough

Bootup methods now return a boolean across interface and spawner components. SocketInterface.bootup returns the underlying spawner’s status. SubprocessSpawner.bootup and FluxPythonSpawner.bootup return poll() results post-start. Tests updated to assert the boolean success flag and include a new error-handling test for a failing serial command.

Changes

Cohort / File(s)Summary of Changes
Standalone interactive: bootup return propagation
executorlib/standalone/interactive/communication.py, executorlib/standalone/interactive/spawner.py
SocketInterface.bootup annotated to return bool and now returns the spawner’s bootup result. SubprocessSpawner.bootup returns poll() after launching, converting prior no-return into a boolean status. Docstrings/signatures updated accordingly.
Task scheduler Flux spawner
executorlib/task_scheduler/interactive/spawner_flux.py
FluxPythonSpawner.bootup annotated to return bool and now returns poll() after scheduling the job. Docstring updated. No other functional changes.
Tests updated for boolean bootup and error path
tests/test_standalone_interactive_communication.py
Tests now capture and assert the bootup boolean. Added test_interface_serial_with_error to validate behavior with a failing command and graceful shutdown. Minor imports and waits added.

Sequence Diagram(s)

sequenceDiagram
autonumber
participant T as Test
participant I as SocketInterface
participant S as Spawner (Subprocess/Flux)
participant P as Process/Job
T->>I: bootup(command_lst)
I->>S: bootup(command_lst)
S->>P: start/schedule
S->>S: poll() running?
S-->>I: bool (running status)
I-->>T: bool (propagated)
rect rgba(230,255,230,0.5)
note right of T: Caller can immediately know if process/job is running
end
Loading

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

  • Rename Interface to Spawner #398 — Also modifies Spawner/SocketInterface pathways, renaming/rewiring Interface → Spawner and adjusting types; closely related to return-type changes here.

Poem

A rabbit taps the bootup drum—thump, thump—
Now truth returns as bool, no guesswork slump.
Spawners poll, interfaces tell,
Tests nod yes: “All is well.”
If a job hops out of run, we still shut down—done! 🐇✅

Warning

Review ran into problems

🔥 Problems

Git: Failed to clone repository. Please run the @coderabbitai full review command to re-trigger a full review. If the issue persists, set path_filters to include or exclude specific files.

✨ Finishing Touches
  • 📝 Generate Docstrings
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch spawner_success

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@codecov

codecovBot commented Sep 7, 2025

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 97.77%. Comparing base (f1847a0) to head (8cfc8fa).
⚠️ Report is 2 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #810 +/- ##
=======================================
Coverage 97.76% 97.77% =======================================
Files 32 32 Lines 1479 1481 +2 =======================================
+ Hits 1446 1448 +2 
Misses 33 33 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
executorlib/standalone/interactive/spawner.py (1)

87-105: Make bootup success non-racy and type it as bool

Returning self.poll() immediately after Popen can intermittently be False for very short-lived commands, making startup detection flaky. Treat a successful Popen as “startup succeeded” and annotate the return type.

- def bootup(+ def bootup(
self,
command_lst: list[str],
- ):+ ) -> bool:
@@
- return self.poll()+ # Successful process creation == successful bootup; ongoing liveness is tracked via poll().+ return True

Outside this hunk (optional, for consistency across the hierarchy), consider updating the abstract method to advertise the contract:

# In BaseSpawner@abstractmethoddefbootup(self, command_lst: list[str]) ->bool: ...
🧹 Nitpick comments (2)
executorlib/standalone/interactive/communication.py (1)

109-121: Nit: docstring type should match annotation

Tighten the docstring type to list[str] for consistency.

- Args:- command_lst (list): list of strings to start the client process+ Args:+ command_lst (list[str]): list of strings to start the client process
tests/test_standalone_interactive_communication.py (1)

122-134: Stabilize the failing-command test and add a timeout

["bash", "exit"] relies on a file named “exit” and may behave differently across systems. Use bash -c 'exit 1' and add a bounded wait to avoid potential infinite loops.

- success_flag = interface.bootup(command_lst=["bash", "exit"])+ success_flag = interface.bootup(command_lst=["bash", "-c", "exit 1"])
self.assertTrue(success_flag)
- while interface._spawner.poll():- sleep(0.1)+ # Wait up to ~20s for the child to exit+ for _ in range(200):+ if not interface._spawner.poll():+ break+ sleep(0.1)
self.assertFalse(interface._spawner.poll())
interface.shutdown(wait=True)
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between e05297e and cf5088f.

📒 Files selected for processing (4)
  • executorlib/standalone/interactive/communication.py (1 hunks)
  • executorlib/standalone/interactive/spawner.py (1 hunks)
  • executorlib/task_scheduler/interactive/spawner_flux.py (2 hunks)
  • tests/test_standalone_interactive_communication.py (8 hunks)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (7)
  • GitHub Check: unittest_win
  • GitHub Check: notebooks_integration
  • GitHub Check: unittest_slurm_mpich
  • GitHub Check: unittest_openmpi (macos-latest, 3.13)
  • GitHub Check: benchmark (ubuntu-latest, 3.13, .ci_support/environment-openmpi.yml)
  • GitHub Check: notebooks
  • GitHub Check: benchmark (ubuntu-latest, 3.13, .ci_support/environment-mpich.yml)
🔇 Additional comments (2)
executorlib/standalone/interactive/communication.py (1)

109-121: LGTM: bootup now propagates spawner’s status

Returning the spawner’s bootup result from SocketInterface is the right API surface.

tests/test_standalone_interactive_communication.py (1)

39-57: LGTM: tests assert bootup’s boolean contract

Capturing and asserting the bootup success flag across scenarios looks good.

Also applies to: 69-87, 99-117, 142-160

Comment on lines 75 to 90
def bootup(
self,
command_lst: list[str],
):
) -> bool:
"""
Boot up the client process to connect to the SocketInterface.

Args:
command_lst (list[str]): List of strings to start the client process.

Raises:
ValueError: If oversubscribing is not supported for the Flux adapter or if conda environments are not supported.

Returns:
bool: Whether the interface was successfully started.
"""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

Align Flux bootup semantics to avoid race with fast failures

Same flakiness risk: a job that fails near-instantly can make poll() return False at bootup. Return True on successful submit and keep poll() for liveness.

- def bootup(+ def bootup(
self,
command_lst: list[str],
) -> bool:
@@
- return self.poll()+ # Submit succeeded → startup succeeded. Use poll() for liveness checks.+ return True

Also applies to: 133-134

🤖 Prompt for AI Agents
In executorlib/task_scheduler/interactive/spawner_flux.py around lines 75-90
(and likewise adjust lines 133-134), the bootup method currently treats an
immediate poll() False as startup failure which races with very fast job
failures; instead, after a successful job submission return True to indicate
boot completed and rely on poll() only for later liveness checks. Modify bootup
so it verifies submission succeeded (no exceptions and valid job handle),
returns True immediately on submit success even if poll() currently returns
False, and keep poll() used elsewhere to detect runtime termination; apply the
same semantics to the similar logic at lines 133-134.

@jan-janssen
jan-janssen merged commit 3524942 into mainSep 7, 2025
109 of 119 checks passed
@jan-janssen
jan-janssen deleted the spawner_success branch September 7, 2025 11:49
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@jan-janssen