[Fix] serialize worker job submissions to preserve worker_id ordering - #935

Merged
jan-janssen merged 5 commits into
pyiron:mainfrom
IlgarBaghishov:fix/ordered-worker-submission
Feb 22, 2026
Merged

[Fix] serialize worker job submissions to preserve worker_id ordering#935
jan-janssen merged 5 commits into
pyiron:mainfrom
IlgarBaghishov:fix/ordered-worker-submission

Conversation

@IlgarBaghishov

@IlgarBaghishovIlgarBaghishov commented Feb 21, 2026

Copy link
Copy Markdown
Contributor

Worker threads were submitting Flux jobs concurrently, causing
executorlib_worker_id to not correspond to the Flux scheduling order.
This made worker_id unreliable for resource mapping (e.g. GPU assignment).

Two changes:

  • Add threading.Event chain in BlockAllocationTaskScheduler so each worker waits for the previous worker to finish submitting before starting its own submission.
  • Call self._future.jobid() after FluxExecutor.submit() to block until the job is actually registered with the Flux broker, not just queued in the async FluxExecutor.

Summary by CodeRabbit

  • Improvements
    • Enforced ordered worker startup so workers initialize sequentially, improving predictability and reducing race conditions during startup.
    • Immediately capture job identifiers after submission to improve launch observability and make initial job tracking more reliable.

 Worker threads were submitting Flux jobs concurrently, causing
executorlib_worker_id to not correspond to the Flux scheduling order.
This made worker_id unreliable for resource mapping (e.g. GPU assignment).
Two changes:
- Add threading.Event chain in BlockAllocationTaskScheduler so each
worker waits for the previous worker to finish submitting before
starting its own submission.
- Call self._future.jobid() after FluxExecutor.submit() to block until
the job is actually registered with the Flux broker, not just queued
in the async FluxExecutor.
@coderabbitai

coderabbitaiBot commented Feb 21, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Adds ordered per-worker boot-up synchronization using threading.Event objects and updates _execute_multiple_tasks signature to accept bootup_event/next_bootup_event. Also calls self._future.jobid() immediately after submitting the jobspec in the Flux spawner bootup.

Changes

Cohort / File(s)Summary
Worker Boot-up Synchronization
src/executorlib/task_scheduler/interactive/blockallocation.py
Add per-worker bootup_events, pass bootup_event / next_bootup_event via executor kwargs, extend _execute_multiple_tasks signature to accept these events, wait on bootup_event before creating the worker interface, and set next_bootup_event after initialization (docstring updated).
Flux Spawner Bootup
src/executorlib/task_scheduler/interactive/spawner_flux.py
After submitting the jobspec during bootup, call self._future.jobid() conditionally when self._future is not None; no other control flow or public API changes.
Metadata
manifest_file, pyproject.toml
Minor packaging/manifest line edits (few-line changes).

Sequence Diagram(s)

sequenceDiagram
participant Scheduler as Scheduler (main)
participant WorkerA as Worker-0 thread
participant WorkerB as Worker-1 thread
participant FluxSpawner as FluxSpawner
participant Future as Future
Scheduler->>WorkerA: start thread with bootup_event A, next_bootup_event B
Scheduler->>WorkerB: start thread with bootup_event B, next_bootup_event C
WorkerA-->>Scheduler: bootup_event A.wait()
WorkerA->>WorkerA: create interface / submit tasks
WorkerA->>Scheduler: signal next_bootup_event B (set)
WorkerA->>FluxSpawner: submit jobspec
FluxSpawner->>Future: (if present) call jobid()
WorkerB-->>Scheduler: bootup_event B.wait()
WorkerB->>WorkerB: create interface / submit tasks
WorkerB->>Scheduler: signal next_bootup_event C (set)
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Poem

🐰 I lined the threads in tidy rows,

One hops up, then the next one goes.
A gentle nudge, a tiny ping,
The bootup choir begins to sing.
Hops in order — what a thing!

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 75.00% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title 'serialize worker job submissions to preserve worker_id ordering' accurately captures the main objective of the PR—implementing sequential worker submission to maintain reliable worker_id ordering in Flux scheduling.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

254-267: ⚠️ Potential issue | 🔴 Critical

Deadlock if interface_bootup raises: subsequent workers hang forever.

If interface_bootup (line 256) throws an exception for worker N, the thread terminates without calling next_bootup_event.set(). Workers N+1, N+2, … remain blocked on bootup_event.wait() indefinitely, and shutdown(wait=True) will hang on process.join().

Wrap the bootup + signal in try/finally so the chain always progresses:

🔒 Proposed fix using try/finally
 if bootup_event is not None:
bootup_event.wait()
- interface = interface_bootup(- command_lst=get_interactive_execute_command(- cores=cores,- ),- connections=spawner(cores=cores, worker_id=worker_id, **kwargs),- hostname_localhost=hostname_localhost,- log_obj_size=log_obj_size,- worker_id=worker_id,- stop_function=stop_function,- )- if next_bootup_event is not None:- next_bootup_event.set()+ try:+ interface = interface_bootup(+ command_lst=get_interactive_execute_command(+ cores=cores,+ ),+ connections=spawner(cores=cores, worker_id=worker_id, **kwargs),+ hostname_localhost=hostname_localhost,+ log_obj_size=log_obj_size,+ worker_id=worker_id,+ stop_function=stop_function,+ )+ finally:+ if next_bootup_event is not None:+ next_bootup_event.set()

Note: if bootup fails and the chain unblocks the next worker, subsequent workers will also likely fail. But that's preferable to a silent deadlock. You may also want to consider using bootup_event.wait(timeout=...) as additional protection.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
254 - 267, The chain of bootstrapped workers can deadlock if interface_bootup
raises because next_bootup_event.set() is skipped; modify the block around the
call to interface_bootup (the code that waits on bootup_event, calls
interface_bootup with get_interactive_execute_command and spawner, and then
calls next_bootup_event.set()) to ensure next_bootup_event.set() is executed in
a finally clause: keep the bootup_event.wait() before trying to create the
interface, perform the interface_bootup call (using
get_interactive_execute_command and spawner) inside a try, and always call
next_bootup_event.set() in the finally so the next worker is unblocked even if
interface_bootup raises; optionally re-raise or log the exception inside the
try/except around interface_bootup if you want error visibility.
🧹 Nitpick comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

116-126: Dynamic worker addition skips boot-up ordering (and worker_id).

The max_workers setter creates new threads using only self._process_kwargs, which doesn't include bootup_event, next_bootup_event, or worker_id. The None defaults prevent a crash, but dynamically added workers won't have ordered boot-up or correct worker_id assignment.

This is pre-existing behavior and not a regression from this PR, but worth noting if the setter is expected to be used in contexts where ordering matters.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
116 - 126, The dynamic worker-add path in the max_workers setter builds new
Thread instances using only self._process_kwargs which omits bootup_event,
next_bootup_event, and worker_id, so newly added threads skip boot-up ordering
and get wrong worker IDs; update the branch that creates new_process_lst (and
references _execute_multiple_tasks and self._process_kwargs) to construct
per-thread kwargs mirroring the initial worker creation: set worker_id =
self._max_workers + i for each new thread, pass the correct bootup_event and
next_bootup_event (chain them the same way the initial spawn logic does), and
then start and append those threads to self._process so dynamically added
workers follow the same boot ordering and ID assignment as existing workers.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 254-267: The chain of bootstrapped workers can deadlock if
interface_bootup raises because next_bootup_event.set() is skipped; modify the
block around the call to interface_bootup (the code that waits on bootup_event,
calls interface_bootup with get_interactive_execute_command and spawner, and
then calls next_bootup_event.set()) to ensure next_bootup_event.set() is
executed in a finally clause: keep the bootup_event.wait() before trying to
create the interface, perform the interface_bootup call (using
get_interactive_execute_command and spawner) inside a try, and always call
next_bootup_event.set() in the finally so the next worker is unblocked even if
interface_bootup raises; optionally re-raise or log the exception inside the
try/except around interface_bootup if you want error visibility.
---
Nitpick comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 116-126: The dynamic worker-add path in the max_workers setter
builds new Thread instances using only self._process_kwargs which omits
bootup_event, next_bootup_event, and worker_id, so newly added threads skip
boot-up ordering and get wrong worker IDs; update the branch that creates
new_process_lst (and references _execute_multiple_tasks and
self._process_kwargs) to construct per-thread kwargs mirroring the initial
worker creation: set worker_id = self._max_workers + i for each new thread, pass
the correct bootup_event and next_bootup_event (chain them the same way the
initial spawn logic does), and then start and append those threads to
self._process so dynamically added workers follow the same boot ordering and ID
assignment as existing workers.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

256-269: ⚠️ Potential issue | 🔴 Critical

Deadlock if interface_bootup raises an exception.

If interface_bootup throws (e.g., spawner failure, connection timeout), next_bootup_event.set() on line 268 is never reached. Every subsequent worker will block forever on bootup_event.wait(), and shutdown(wait=True) will also hang because Thread.join() blocks on those stuck threads.

Wrap the boot-up + signal in try/finally so the chain always advances.

🔒 Proposed fix
 if bootup_event is not None:
bootup_event.wait()
- interface = interface_bootup(- command_lst=get_interactive_execute_command(- cores=cores,- ),- connections=spawner(cores=cores, worker_id=worker_id, **kwargs),- hostname_localhost=hostname_localhost,- log_obj_size=log_obj_size,- worker_id=worker_id,- stop_function=stop_function,- )- if next_bootup_event is not None:- next_bootup_event.set()+ try:+ interface = interface_bootup(+ command_lst=get_interactive_execute_command(+ cores=cores,+ ),+ connections=spawner(cores=cores, worker_id=worker_id, **kwargs),+ hostname_localhost=hostname_localhost,+ log_obj_size=log_obj_size,+ worker_id=worker_id,+ stop_function=stop_function,+ )+ finally:+ if next_bootup_event is not None:+ next_bootup_event.set()

Note: after the finally block, if interface_bootup did raise, the exception will propagate and terminate this worker thread. The rest of the function (init_function, task loop) will be skipped — but at least the remaining workers won't deadlock.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
256 - 269, The boot-up chain can deadlock if interface_bootup (called with
get_interactive_execute_command(...) and connections=spawner(...)) raises,
because next_bootup_event.set() is never reached; wrap the call to
interface_bootup and the subsequent signaling in a try/finally so that
next_bootup_event.set() is always executed in the finally block (still allowing
the exception to propagate and terminate this worker thread), ensuring
bootup_event.wait()/next_bootup_event.set() chain always advances.
🧹 Nitpick comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

256-257: bootup_event.wait() is not interruptible by shutdown().

If shutdown() is called while workers are still queued behind bootup_event.wait(), those threads cannot observe the stop signal — stop_function is only evaluated inside interface_bootup, not here. With the try/finally fix above, the chain will unblock naturally once the currently booting worker finishes (or fails), so this is mostly a latency concern rather than a hard deadlock. However, for a fast-exit guarantee you could add a timeout loop:

ifbootup_eventisnotNone:
whilenotbootup_event.wait(timeout=1.0):
ifstop_functionisnotNoneandstop_function():
return

This is optional and can be deferred.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
256 - 257, The current unconditional bootup_event.wait() call can block threads
from observing shutdown; change it to a timeout loop that repeatedly waits with
a short timeout and checks stop_function between waits so threads can return
early on shutdown. Replace the direct call to bootup_event.wait() (in
blockallocation where bootup_event is used alongside interface_bootup) with a
loop that calls bootup_event.wait(timeout=...) and after each timeout calls
stop_function() (if provided) and returns if it indicates shutdown; keep
behavior of eventually proceeding when bootup_event is set. Ensure you reference
bootup_event and stop_function and preserve existing cleanup/try/finally
semantics around interface_bootup.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 256-269: The boot-up chain can deadlock if interface_bootup
(called with get_interactive_execute_command(...) and connections=spawner(...))
raises, because next_bootup_event.set() is never reached; wrap the call to
interface_bootup and the subsequent signaling in a try/finally so that
next_bootup_event.set() is always executed in the finally block (still allowing
the exception to propagate and terminate this worker thread), ensuring
bootup_event.wait()/next_bootup_event.set() chain always advances.
---
Nitpick comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 256-257: The current unconditional bootup_event.wait() call can
block threads from observing shutdown; change it to a timeout loop that
repeatedly waits with a short timeout and checks stop_function between waits so
threads can return early on shutdown. Replace the direct call to
bootup_event.wait() (in blockallocation where bootup_event is used alongside
interface_bootup) with a loop that calls bootup_event.wait(timeout=...) and
after each timeout calls stop_function() (if provided) and returns if it
indicates shutdown; keep behavior of eventually proceeding when bootup_event is
set. Ensure you reference bootup_event and stop_function and preserve existing
cleanup/try/finally semantics around interface_bootup.

@codecov

codecovBot commented Feb 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.09%. Comparing base (2fa06c8) to head (e099b31).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #935 +/- ##
==========================================
+ Coverage 94.07% 94.09% +0.02% 
==========================================
Files 39 39 Lines 2041 2049 +8 ==========================================
+ Hits 1920 1928 +8 
Misses 121 121 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jan-janssenjan-janssen left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me

@jan-janssenjan-janssen changed the title fix: serialize worker job submissions to preserve worker_id ordering[Fix] serialize worker job submissions to preserve worker_id orderingFeb 22, 2026
@jan-janssen
jan-janssen merged commit 693ca9e into pyiron:mainFeb 22, 2026
59 of 62 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@IlgarBaghishov@jan-janssen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[Fix] serialize worker job submissions to preserve worker_id ordering - #935

Merged
jan-janssen merged 5 commits into
pyiron:mainfrom
IlgarBaghishov:fix/ordered-worker-submission
Feb 22, 2026
Merged

[Fix] serialize worker job submissions to preserve worker_id ordering#935
jan-janssen merged 5 commits into
pyiron:mainfrom
IlgarBaghishov:fix/ordered-worker-submission

Conversation

@IlgarBaghishov

@IlgarBaghishovIlgarBaghishov commented Feb 21, 2026

Copy link
Copy Markdown
Contributor

Worker threads were submitting Flux jobs concurrently, causing
executorlib_worker_id to not correspond to the Flux scheduling order.
This made worker_id unreliable for resource mapping (e.g. GPU assignment).

Two changes:

  • Add threading.Event chain in BlockAllocationTaskScheduler so each worker waits for the previous worker to finish submitting before starting its own submission.
  • Call self._future.jobid() after FluxExecutor.submit() to block until the job is actually registered with the Flux broker, not just queued in the async FluxExecutor.

Summary by CodeRabbit

  • Improvements
    • Enforced ordered worker startup so workers initialize sequentially, improving predictability and reducing race conditions during startup.
    • Immediately capture job identifiers after submission to improve launch observability and make initial job tracking more reliable.

 Worker threads were submitting Flux jobs concurrently, causing
executorlib_worker_id to not correspond to the Flux scheduling order.
This made worker_id unreliable for resource mapping (e.g. GPU assignment).
Two changes:
- Add threading.Event chain in BlockAllocationTaskScheduler so each
worker waits for the previous worker to finish submitting before
starting its own submission.
- Call self._future.jobid() after FluxExecutor.submit() to block until
the job is actually registered with the Flux broker, not just queued
in the async FluxExecutor.
@coderabbitai

coderabbitaiBot commented Feb 21, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Adds ordered per-worker boot-up synchronization using threading.Event objects and updates _execute_multiple_tasks signature to accept bootup_event/next_bootup_event. Also calls self._future.jobid() immediately after submitting the jobspec in the Flux spawner bootup.

Changes

Cohort / File(s)Summary
Worker Boot-up Synchronization
src/executorlib/task_scheduler/interactive/blockallocation.py
Add per-worker bootup_events, pass bootup_event / next_bootup_event via executor kwargs, extend _execute_multiple_tasks signature to accept these events, wait on bootup_event before creating the worker interface, and set next_bootup_event after initialization (docstring updated).
Flux Spawner Bootup
src/executorlib/task_scheduler/interactive/spawner_flux.py
After submitting the jobspec during bootup, call self._future.jobid() conditionally when self._future is not None; no other control flow or public API changes.
Metadata
manifest_file, pyproject.toml
Minor packaging/manifest line edits (few-line changes).

Sequence Diagram(s)

sequenceDiagram
participant Scheduler as Scheduler (main)
participant WorkerA as Worker-0 thread
participant WorkerB as Worker-1 thread
participant FluxSpawner as FluxSpawner
participant Future as Future
Scheduler->>WorkerA: start thread with bootup_event A, next_bootup_event B
Scheduler->>WorkerB: start thread with bootup_event B, next_bootup_event C
WorkerA-->>Scheduler: bootup_event A.wait()
WorkerA->>WorkerA: create interface / submit tasks
WorkerA->>Scheduler: signal next_bootup_event B (set)
WorkerA->>FluxSpawner: submit jobspec
FluxSpawner->>Future: (if present) call jobid()
WorkerB-->>Scheduler: bootup_event B.wait()
WorkerB->>WorkerB: create interface / submit tasks
WorkerB->>Scheduler: signal next_bootup_event C (set)
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Poem

🐰 I lined the threads in tidy rows,

One hops up, then the next one goes.
A gentle nudge, a tiny ping,
The bootup choir begins to sing.
Hops in order — what a thing!

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 75.00% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title 'serialize worker job submissions to preserve worker_id ordering' accurately captures the main objective of the PR—implementing sequential worker submission to maintain reliable worker_id ordering in Flux scheduling.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

254-267: ⚠️ Potential issue | 🔴 Critical

Deadlock if interface_bootup raises: subsequent workers hang forever.

If interface_bootup (line 256) throws an exception for worker N, the thread terminates without calling next_bootup_event.set(). Workers N+1, N+2, … remain blocked on bootup_event.wait() indefinitely, and shutdown(wait=True) will hang on process.join().

Wrap the bootup + signal in try/finally so the chain always progresses:

🔒 Proposed fix using try/finally
 if bootup_event is not None:
bootup_event.wait()
- interface = interface_bootup(- command_lst=get_interactive_execute_command(- cores=cores,- ),- connections=spawner(cores=cores, worker_id=worker_id, **kwargs),- hostname_localhost=hostname_localhost,- log_obj_size=log_obj_size,- worker_id=worker_id,- stop_function=stop_function,- )- if next_bootup_event is not None:- next_bootup_event.set()+ try:+ interface = interface_bootup(+ command_lst=get_interactive_execute_command(+ cores=cores,+ ),+ connections=spawner(cores=cores, worker_id=worker_id, **kwargs),+ hostname_localhost=hostname_localhost,+ log_obj_size=log_obj_size,+ worker_id=worker_id,+ stop_function=stop_function,+ )+ finally:+ if next_bootup_event is not None:+ next_bootup_event.set()

Note: if bootup fails and the chain unblocks the next worker, subsequent workers will also likely fail. But that's preferable to a silent deadlock. You may also want to consider using bootup_event.wait(timeout=...) as additional protection.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
254 - 267, The chain of bootstrapped workers can deadlock if interface_bootup
raises because next_bootup_event.set() is skipped; modify the block around the
call to interface_bootup (the code that waits on bootup_event, calls
interface_bootup with get_interactive_execute_command and spawner, and then
calls next_bootup_event.set()) to ensure next_bootup_event.set() is executed in
a finally clause: keep the bootup_event.wait() before trying to create the
interface, perform the interface_bootup call (using
get_interactive_execute_command and spawner) inside a try, and always call
next_bootup_event.set() in the finally so the next worker is unblocked even if
interface_bootup raises; optionally re-raise or log the exception inside the
try/except around interface_bootup if you want error visibility.
🧹 Nitpick comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

116-126: Dynamic worker addition skips boot-up ordering (and worker_id).

The max_workers setter creates new threads using only self._process_kwargs, which doesn't include bootup_event, next_bootup_event, or worker_id. The None defaults prevent a crash, but dynamically added workers won't have ordered boot-up or correct worker_id assignment.

This is pre-existing behavior and not a regression from this PR, but worth noting if the setter is expected to be used in contexts where ordering matters.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
116 - 126, The dynamic worker-add path in the max_workers setter builds new
Thread instances using only self._process_kwargs which omits bootup_event,
next_bootup_event, and worker_id, so newly added threads skip boot-up ordering
and get wrong worker IDs; update the branch that creates new_process_lst (and
references _execute_multiple_tasks and self._process_kwargs) to construct
per-thread kwargs mirroring the initial worker creation: set worker_id =
self._max_workers + i for each new thread, pass the correct bootup_event and
next_bootup_event (chain them the same way the initial spawn logic does), and
then start and append those threads to self._process so dynamically added
workers follow the same boot ordering and ID assignment as existing workers.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 254-267: The chain of bootstrapped workers can deadlock if
interface_bootup raises because next_bootup_event.set() is skipped; modify the
block around the call to interface_bootup (the code that waits on bootup_event,
calls interface_bootup with get_interactive_execute_command and spawner, and
then calls next_bootup_event.set()) to ensure next_bootup_event.set() is
executed in a finally clause: keep the bootup_event.wait() before trying to
create the interface, perform the interface_bootup call (using
get_interactive_execute_command and spawner) inside a try, and always call
next_bootup_event.set() in the finally so the next worker is unblocked even if
interface_bootup raises; optionally re-raise or log the exception inside the
try/except around interface_bootup if you want error visibility.
---
Nitpick comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 116-126: The dynamic worker-add path in the max_workers setter
builds new Thread instances using only self._process_kwargs which omits
bootup_event, next_bootup_event, and worker_id, so newly added threads skip
boot-up ordering and get wrong worker IDs; update the branch that creates
new_process_lst (and references _execute_multiple_tasks and
self._process_kwargs) to construct per-thread kwargs mirroring the initial
worker creation: set worker_id = self._max_workers + i for each new thread, pass
the correct bootup_event and next_bootup_event (chain them the same way the
initial spawn logic does), and then start and append those threads to
self._process so dynamically added workers follow the same boot ordering and ID
assignment as existing workers.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

256-269: ⚠️ Potential issue | 🔴 Critical

Deadlock if interface_bootup raises an exception.

If interface_bootup throws (e.g., spawner failure, connection timeout), next_bootup_event.set() on line 268 is never reached. Every subsequent worker will block forever on bootup_event.wait(), and shutdown(wait=True) will also hang because Thread.join() blocks on those stuck threads.

Wrap the boot-up + signal in try/finally so the chain always advances.

🔒 Proposed fix
 if bootup_event is not None:
bootup_event.wait()
- interface = interface_bootup(- command_lst=get_interactive_execute_command(- cores=cores,- ),- connections=spawner(cores=cores, worker_id=worker_id, **kwargs),- hostname_localhost=hostname_localhost,- log_obj_size=log_obj_size,- worker_id=worker_id,- stop_function=stop_function,- )- if next_bootup_event is not None:- next_bootup_event.set()+ try:+ interface = interface_bootup(+ command_lst=get_interactive_execute_command(+ cores=cores,+ ),+ connections=spawner(cores=cores, worker_id=worker_id, **kwargs),+ hostname_localhost=hostname_localhost,+ log_obj_size=log_obj_size,+ worker_id=worker_id,+ stop_function=stop_function,+ )+ finally:+ if next_bootup_event is not None:+ next_bootup_event.set()

Note: after the finally block, if interface_bootup did raise, the exception will propagate and terminate this worker thread. The rest of the function (init_function, task loop) will be skipped — but at least the remaining workers won't deadlock.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
256 - 269, The boot-up chain can deadlock if interface_bootup (called with
get_interactive_execute_command(...) and connections=spawner(...)) raises,
because next_bootup_event.set() is never reached; wrap the call to
interface_bootup and the subsequent signaling in a try/finally so that
next_bootup_event.set() is always executed in the finally block (still allowing
the exception to propagate and terminate this worker thread), ensuring
bootup_event.wait()/next_bootup_event.set() chain always advances.
🧹 Nitpick comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

256-257: bootup_event.wait() is not interruptible by shutdown().

If shutdown() is called while workers are still queued behind bootup_event.wait(), those threads cannot observe the stop signal — stop_function is only evaluated inside interface_bootup, not here. With the try/finally fix above, the chain will unblock naturally once the currently booting worker finishes (or fails), so this is mostly a latency concern rather than a hard deadlock. However, for a fast-exit guarantee you could add a timeout loop:

ifbootup_eventisnotNone:
whilenotbootup_event.wait(timeout=1.0):
ifstop_functionisnotNoneandstop_function():
return

This is optional and can be deferred.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
256 - 257, The current unconditional bootup_event.wait() call can block threads
from observing shutdown; change it to a timeout loop that repeatedly waits with
a short timeout and checks stop_function between waits so threads can return
early on shutdown. Replace the direct call to bootup_event.wait() (in
blockallocation where bootup_event is used alongside interface_bootup) with a
loop that calls bootup_event.wait(timeout=...) and after each timeout calls
stop_function() (if provided) and returns if it indicates shutdown; keep
behavior of eventually proceeding when bootup_event is set. Ensure you reference
bootup_event and stop_function and preserve existing cleanup/try/finally
semantics around interface_bootup.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 256-269: The boot-up chain can deadlock if interface_bootup
(called with get_interactive_execute_command(...) and connections=spawner(...))
raises, because next_bootup_event.set() is never reached; wrap the call to
interface_bootup and the subsequent signaling in a try/finally so that
next_bootup_event.set() is always executed in the finally block (still allowing
the exception to propagate and terminate this worker thread), ensuring
bootup_event.wait()/next_bootup_event.set() chain always advances.
---
Nitpick comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 256-257: The current unconditional bootup_event.wait() call can
block threads from observing shutdown; change it to a timeout loop that
repeatedly waits with a short timeout and checks stop_function between waits so
threads can return early on shutdown. Replace the direct call to
bootup_event.wait() (in blockallocation where bootup_event is used alongside
interface_bootup) with a loop that calls bootup_event.wait(timeout=...) and
after each timeout calls stop_function() (if provided) and returns if it
indicates shutdown; keep behavior of eventually proceeding when bootup_event is
set. Ensure you reference bootup_event and stop_function and preserve existing
cleanup/try/finally semantics around interface_bootup.

@codecov

codecovBot commented Feb 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.09%. Comparing base (2fa06c8) to head (e099b31).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #935 +/- ##
==========================================
+ Coverage 94.07% 94.09% +0.02% 
==========================================
Files 39 39 Lines 2041 2049 +8 ==========================================
+ Hits 1920 1928 +8 
Misses 121 121 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jan-janssenjan-janssen left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me

@jan-janssenjan-janssen changed the title fix: serialize worker job submissions to preserve worker_id ordering[Fix] serialize worker job submissions to preserve worker_id orderingFeb 22, 2026
@jan-janssen
jan-janssen merged commit 693ca9e into pyiron:mainFeb 22, 2026
59 of 62 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@IlgarBaghishov@jan-janssen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Fix] serialize worker job submissions to preserve worker_id ordering - #935

Merged
jan-janssen merged 5 commits into
pyiron:mainfrom
IlgarBaghishov:fix/ordered-worker-submission
Feb 22, 2026
Merged

[Fix] serialize worker job submissions to preserve worker_id ordering#935
jan-janssen merged 5 commits into
pyiron:mainfrom
IlgarBaghishov:fix/ordered-worker-submission

Conversation

@IlgarBaghishov

@IlgarBaghishovIlgarBaghishov commented Feb 21, 2026

Copy link
Copy Markdown
Contributor

Worker threads were submitting Flux jobs concurrently, causing
executorlib_worker_id to not correspond to the Flux scheduling order.
This made worker_id unreliable for resource mapping (e.g. GPU assignment).

Two changes:

  • Add threading.Event chain in BlockAllocationTaskScheduler so each worker waits for the previous worker to finish submitting before starting its own submission.
  • Call self._future.jobid() after FluxExecutor.submit() to block until the job is actually registered with the Flux broker, not just queued in the async FluxExecutor.

Summary by CodeRabbit

  • Improvements
    • Enforced ordered worker startup so workers initialize sequentially, improving predictability and reducing race conditions during startup.
    • Immediately capture job identifiers after submission to improve launch observability and make initial job tracking more reliable.

 Worker threads were submitting Flux jobs concurrently, causing
executorlib_worker_id to not correspond to the Flux scheduling order.
This made worker_id unreliable for resource mapping (e.g. GPU assignment).
Two changes:
- Add threading.Event chain in BlockAllocationTaskScheduler so each
worker waits for the previous worker to finish submitting before
starting its own submission.
- Call self._future.jobid() after FluxExecutor.submit() to block until
the job is actually registered with the Flux broker, not just queued
in the async FluxExecutor.
@coderabbitai

coderabbitaiBot commented Feb 21, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Adds ordered per-worker boot-up synchronization using threading.Event objects and updates _execute_multiple_tasks signature to accept bootup_event/next_bootup_event. Also calls self._future.jobid() immediately after submitting the jobspec in the Flux spawner bootup.

Changes

Cohort / File(s)Summary
Worker Boot-up Synchronization
src/executorlib/task_scheduler/interactive/blockallocation.py
Add per-worker bootup_events, pass bootup_event / next_bootup_event via executor kwargs, extend _execute_multiple_tasks signature to accept these events, wait on bootup_event before creating the worker interface, and set next_bootup_event after initialization (docstring updated).
Flux Spawner Bootup
src/executorlib/task_scheduler/interactive/spawner_flux.py
After submitting the jobspec during bootup, call self._future.jobid() conditionally when self._future is not None; no other control flow or public API changes.
Metadata
manifest_file, pyproject.toml
Minor packaging/manifest line edits (few-line changes).

Sequence Diagram(s)

sequenceDiagram
participant Scheduler as Scheduler (main)
participant WorkerA as Worker-0 thread
participant WorkerB as Worker-1 thread
participant FluxSpawner as FluxSpawner
participant Future as Future
Scheduler->>WorkerA: start thread with bootup_event A, next_bootup_event B
Scheduler->>WorkerB: start thread with bootup_event B, next_bootup_event C
WorkerA-->>Scheduler: bootup_event A.wait()
WorkerA->>WorkerA: create interface / submit tasks
WorkerA->>Scheduler: signal next_bootup_event B (set)
WorkerA->>FluxSpawner: submit jobspec
FluxSpawner->>Future: (if present) call jobid()
WorkerB-->>Scheduler: bootup_event B.wait()
WorkerB->>WorkerB: create interface / submit tasks
WorkerB->>Scheduler: signal next_bootup_event C (set)
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Poem

🐰 I lined the threads in tidy rows,

One hops up, then the next one goes.
A gentle nudge, a tiny ping,
The bootup choir begins to sing.
Hops in order — what a thing!

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 75.00% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title 'serialize worker job submissions to preserve worker_id ordering' accurately captures the main objective of the PR—implementing sequential worker submission to maintain reliable worker_id ordering in Flux scheduling.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

254-267: ⚠️ Potential issue | 🔴 Critical

Deadlock if interface_bootup raises: subsequent workers hang forever.

If interface_bootup (line 256) throws an exception for worker N, the thread terminates without calling next_bootup_event.set(). Workers N+1, N+2, … remain blocked on bootup_event.wait() indefinitely, and shutdown(wait=True) will hang on process.join().

Wrap the bootup + signal in try/finally so the chain always progresses:

🔒 Proposed fix using try/finally
 if bootup_event is not None:
bootup_event.wait()
- interface = interface_bootup(- command_lst=get_interactive_execute_command(- cores=cores,- ),- connections=spawner(cores=cores, worker_id=worker_id, **kwargs),- hostname_localhost=hostname_localhost,- log_obj_size=log_obj_size,- worker_id=worker_id,- stop_function=stop_function,- )- if next_bootup_event is not None:- next_bootup_event.set()+ try:+ interface = interface_bootup(+ command_lst=get_interactive_execute_command(+ cores=cores,+ ),+ connections=spawner(cores=cores, worker_id=worker_id, **kwargs),+ hostname_localhost=hostname_localhost,+ log_obj_size=log_obj_size,+ worker_id=worker_id,+ stop_function=stop_function,+ )+ finally:+ if next_bootup_event is not None:+ next_bootup_event.set()

Note: if bootup fails and the chain unblocks the next worker, subsequent workers will also likely fail. But that's preferable to a silent deadlock. You may also want to consider using bootup_event.wait(timeout=...) as additional protection.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
254 - 267, The chain of bootstrapped workers can deadlock if interface_bootup
raises because next_bootup_event.set() is skipped; modify the block around the
call to interface_bootup (the code that waits on bootup_event, calls
interface_bootup with get_interactive_execute_command and spawner, and then
calls next_bootup_event.set()) to ensure next_bootup_event.set() is executed in
a finally clause: keep the bootup_event.wait() before trying to create the
interface, perform the interface_bootup call (using
get_interactive_execute_command and spawner) inside a try, and always call
next_bootup_event.set() in the finally so the next worker is unblocked even if
interface_bootup raises; optionally re-raise or log the exception inside the
try/except around interface_bootup if you want error visibility.
🧹 Nitpick comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

116-126: Dynamic worker addition skips boot-up ordering (and worker_id).

The max_workers setter creates new threads using only self._process_kwargs, which doesn't include bootup_event, next_bootup_event, or worker_id. The None defaults prevent a crash, but dynamically added workers won't have ordered boot-up or correct worker_id assignment.

This is pre-existing behavior and not a regression from this PR, but worth noting if the setter is expected to be used in contexts where ordering matters.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
116 - 126, The dynamic worker-add path in the max_workers setter builds new
Thread instances using only self._process_kwargs which omits bootup_event,
next_bootup_event, and worker_id, so newly added threads skip boot-up ordering
and get wrong worker IDs; update the branch that creates new_process_lst (and
references _execute_multiple_tasks and self._process_kwargs) to construct
per-thread kwargs mirroring the initial worker creation: set worker_id =
self._max_workers + i for each new thread, pass the correct bootup_event and
next_bootup_event (chain them the same way the initial spawn logic does), and
then start and append those threads to self._process so dynamically added
workers follow the same boot ordering and ID assignment as existing workers.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 254-267: The chain of bootstrapped workers can deadlock if
interface_bootup raises because next_bootup_event.set() is skipped; modify the
block around the call to interface_bootup (the code that waits on bootup_event,
calls interface_bootup with get_interactive_execute_command and spawner, and
then calls next_bootup_event.set()) to ensure next_bootup_event.set() is
executed in a finally clause: keep the bootup_event.wait() before trying to
create the interface, perform the interface_bootup call (using
get_interactive_execute_command and spawner) inside a try, and always call
next_bootup_event.set() in the finally so the next worker is unblocked even if
interface_bootup raises; optionally re-raise or log the exception inside the
try/except around interface_bootup if you want error visibility.
---
Nitpick comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 116-126: The dynamic worker-add path in the max_workers setter
builds new Thread instances using only self._process_kwargs which omits
bootup_event, next_bootup_event, and worker_id, so newly added threads skip
boot-up ordering and get wrong worker IDs; update the branch that creates
new_process_lst (and references _execute_multiple_tasks and
self._process_kwargs) to construct per-thread kwargs mirroring the initial
worker creation: set worker_id = self._max_workers + i for each new thread, pass
the correct bootup_event and next_bootup_event (chain them the same way the
initial spawn logic does), and then start and append those threads to
self._process so dynamically added workers follow the same boot ordering and ID
assignment as existing workers.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

256-269: ⚠️ Potential issue | 🔴 Critical

Deadlock if interface_bootup raises an exception.

If interface_bootup throws (e.g., spawner failure, connection timeout), next_bootup_event.set() on line 268 is never reached. Every subsequent worker will block forever on bootup_event.wait(), and shutdown(wait=True) will also hang because Thread.join() blocks on those stuck threads.

Wrap the boot-up + signal in try/finally so the chain always advances.

🔒 Proposed fix
 if bootup_event is not None:
bootup_event.wait()
- interface = interface_bootup(- command_lst=get_interactive_execute_command(- cores=cores,- ),- connections=spawner(cores=cores, worker_id=worker_id, **kwargs),- hostname_localhost=hostname_localhost,- log_obj_size=log_obj_size,- worker_id=worker_id,- stop_function=stop_function,- )- if next_bootup_event is not None:- next_bootup_event.set()+ try:+ interface = interface_bootup(+ command_lst=get_interactive_execute_command(+ cores=cores,+ ),+ connections=spawner(cores=cores, worker_id=worker_id, **kwargs),+ hostname_localhost=hostname_localhost,+ log_obj_size=log_obj_size,+ worker_id=worker_id,+ stop_function=stop_function,+ )+ finally:+ if next_bootup_event is not None:+ next_bootup_event.set()

Note: after the finally block, if interface_bootup did raise, the exception will propagate and terminate this worker thread. The rest of the function (init_function, task loop) will be skipped — but at least the remaining workers won't deadlock.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
256 - 269, The boot-up chain can deadlock if interface_bootup (called with
get_interactive_execute_command(...) and connections=spawner(...)) raises,
because next_bootup_event.set() is never reached; wrap the call to
interface_bootup and the subsequent signaling in a try/finally so that
next_bootup_event.set() is always executed in the finally block (still allowing
the exception to propagate and terminate this worker thread), ensuring
bootup_event.wait()/next_bootup_event.set() chain always advances.
🧹 Nitpick comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

256-257: bootup_event.wait() is not interruptible by shutdown().

If shutdown() is called while workers are still queued behind bootup_event.wait(), those threads cannot observe the stop signal — stop_function is only evaluated inside interface_bootup, not here. With the try/finally fix above, the chain will unblock naturally once the currently booting worker finishes (or fails), so this is mostly a latency concern rather than a hard deadlock. However, for a fast-exit guarantee you could add a timeout loop:

ifbootup_eventisnotNone:
whilenotbootup_event.wait(timeout=1.0):
ifstop_functionisnotNoneandstop_function():
return

This is optional and can be deferred.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
256 - 257, The current unconditional bootup_event.wait() call can block threads
from observing shutdown; change it to a timeout loop that repeatedly waits with
a short timeout and checks stop_function between waits so threads can return
early on shutdown. Replace the direct call to bootup_event.wait() (in
blockallocation where bootup_event is used alongside interface_bootup) with a
loop that calls bootup_event.wait(timeout=...) and after each timeout calls
stop_function() (if provided) and returns if it indicates shutdown; keep
behavior of eventually proceeding when bootup_event is set. Ensure you reference
bootup_event and stop_function and preserve existing cleanup/try/finally
semantics around interface_bootup.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 256-269: The boot-up chain can deadlock if interface_bootup
(called with get_interactive_execute_command(...) and connections=spawner(...))
raises, because next_bootup_event.set() is never reached; wrap the call to
interface_bootup and the subsequent signaling in a try/finally so that
next_bootup_event.set() is always executed in the finally block (still allowing
the exception to propagate and terminate this worker thread), ensuring
bootup_event.wait()/next_bootup_event.set() chain always advances.
---
Nitpick comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 256-257: The current unconditional bootup_event.wait() call can
block threads from observing shutdown; change it to a timeout loop that
repeatedly waits with a short timeout and checks stop_function between waits so
threads can return early on shutdown. Replace the direct call to
bootup_event.wait() (in blockallocation where bootup_event is used alongside
interface_bootup) with a loop that calls bootup_event.wait(timeout=...) and
after each timeout calls stop_function() (if provided) and returns if it
indicates shutdown; keep behavior of eventually proceeding when bootup_event is
set. Ensure you reference bootup_event and stop_function and preserve existing
cleanup/try/finally semantics around interface_bootup.

@codecov

codecovBot commented Feb 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.09%. Comparing base (2fa06c8) to head (e099b31).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #935 +/- ##
==========================================
+ Coverage 94.07% 94.09% +0.02% 
==========================================
Files 39 39 Lines 2041 2049 +8 ==========================================
+ Hits 1920 1928 +8 
Misses 121 121 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jan-janssenjan-janssen left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me

@jan-janssenjan-janssen changed the title fix: serialize worker job submissions to preserve worker_id ordering[Fix] serialize worker job submissions to preserve worker_id orderingFeb 22, 2026
@jan-janssen
jan-janssen merged commit 693ca9e into pyiron:mainFeb 22, 2026
59 of 62 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@IlgarBaghishov@jan-janssen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Fix] serialize worker job submissions to preserve worker_id ordering - #935

Merged
jan-janssen merged 5 commits into
pyiron:mainfrom
IlgarBaghishov:fix/ordered-worker-submission
Feb 22, 2026
Merged

[Fix] serialize worker job submissions to preserve worker_id ordering#935
jan-janssen merged 5 commits into
pyiron:mainfrom
IlgarBaghishov:fix/ordered-worker-submission

Conversation

@IlgarBaghishov

@IlgarBaghishovIlgarBaghishov commented Feb 21, 2026

Copy link
Copy Markdown
Contributor

Worker threads were submitting Flux jobs concurrently, causing
executorlib_worker_id to not correspond to the Flux scheduling order.
This made worker_id unreliable for resource mapping (e.g. GPU assignment).

Two changes:

  • Add threading.Event chain in BlockAllocationTaskScheduler so each worker waits for the previous worker to finish submitting before starting its own submission.
  • Call self._future.jobid() after FluxExecutor.submit() to block until the job is actually registered with the Flux broker, not just queued in the async FluxExecutor.

Summary by CodeRabbit

  • Improvements
    • Enforced ordered worker startup so workers initialize sequentially, improving predictability and reducing race conditions during startup.
    • Immediately capture job identifiers after submission to improve launch observability and make initial job tracking more reliable.

 Worker threads were submitting Flux jobs concurrently, causing
executorlib_worker_id to not correspond to the Flux scheduling order.
This made worker_id unreliable for resource mapping (e.g. GPU assignment).
Two changes:
- Add threading.Event chain in BlockAllocationTaskScheduler so each
worker waits for the previous worker to finish submitting before
starting its own submission.
- Call self._future.jobid() after FluxExecutor.submit() to block until
the job is actually registered with the Flux broker, not just queued
in the async FluxExecutor.
@coderabbitai

coderabbitaiBot commented Feb 21, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Adds ordered per-worker boot-up synchronization using threading.Event objects and updates _execute_multiple_tasks signature to accept bootup_event/next_bootup_event. Also calls self._future.jobid() immediately after submitting the jobspec in the Flux spawner bootup.

Changes

Cohort / File(s)Summary
Worker Boot-up Synchronization
src/executorlib/task_scheduler/interactive/blockallocation.py
Add per-worker bootup_events, pass bootup_event / next_bootup_event via executor kwargs, extend _execute_multiple_tasks signature to accept these events, wait on bootup_event before creating the worker interface, and set next_bootup_event after initialization (docstring updated).
Flux Spawner Bootup
src/executorlib/task_scheduler/interactive/spawner_flux.py
After submitting the jobspec during bootup, call self._future.jobid() conditionally when self._future is not None; no other control flow or public API changes.
Metadata
manifest_file, pyproject.toml
Minor packaging/manifest line edits (few-line changes).

Sequence Diagram(s)

sequenceDiagram
participant Scheduler as Scheduler (main)
participant WorkerA as Worker-0 thread
participant WorkerB as Worker-1 thread
participant FluxSpawner as FluxSpawner
participant Future as Future
Scheduler->>WorkerA: start thread with bootup_event A, next_bootup_event B
Scheduler->>WorkerB: start thread with bootup_event B, next_bootup_event C
WorkerA-->>Scheduler: bootup_event A.wait()
WorkerA->>WorkerA: create interface / submit tasks
WorkerA->>Scheduler: signal next_bootup_event B (set)
WorkerA->>FluxSpawner: submit jobspec
FluxSpawner->>Future: (if present) call jobid()
WorkerB-->>Scheduler: bootup_event B.wait()
WorkerB->>WorkerB: create interface / submit tasks
WorkerB->>Scheduler: signal next_bootup_event C (set)
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Poem

🐰 I lined the threads in tidy rows,

One hops up, then the next one goes.
A gentle nudge, a tiny ping,
The bootup choir begins to sing.
Hops in order — what a thing!

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 75.00% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title 'serialize worker job submissions to preserve worker_id ordering' accurately captures the main objective of the PR—implementing sequential worker submission to maintain reliable worker_id ordering in Flux scheduling.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

254-267: ⚠️ Potential issue | 🔴 Critical

Deadlock if interface_bootup raises: subsequent workers hang forever.

If interface_bootup (line 256) throws an exception for worker N, the thread terminates without calling next_bootup_event.set(). Workers N+1, N+2, … remain blocked on bootup_event.wait() indefinitely, and shutdown(wait=True) will hang on process.join().

Wrap the bootup + signal in try/finally so the chain always progresses:

🔒 Proposed fix using try/finally
 if bootup_event is not None:
bootup_event.wait()
- interface = interface_bootup(- command_lst=get_interactive_execute_command(- cores=cores,- ),- connections=spawner(cores=cores, worker_id=worker_id, **kwargs),- hostname_localhost=hostname_localhost,- log_obj_size=log_obj_size,- worker_id=worker_id,- stop_function=stop_function,- )- if next_bootup_event is not None:- next_bootup_event.set()+ try:+ interface = interface_bootup(+ command_lst=get_interactive_execute_command(+ cores=cores,+ ),+ connections=spawner(cores=cores, worker_id=worker_id, **kwargs),+ hostname_localhost=hostname_localhost,+ log_obj_size=log_obj_size,+ worker_id=worker_id,+ stop_function=stop_function,+ )+ finally:+ if next_bootup_event is not None:+ next_bootup_event.set()

Note: if bootup fails and the chain unblocks the next worker, subsequent workers will also likely fail. But that's preferable to a silent deadlock. You may also want to consider using bootup_event.wait(timeout=...) as additional protection.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
254 - 267, The chain of bootstrapped workers can deadlock if interface_bootup
raises because next_bootup_event.set() is skipped; modify the block around the
call to interface_bootup (the code that waits on bootup_event, calls
interface_bootup with get_interactive_execute_command and spawner, and then
calls next_bootup_event.set()) to ensure next_bootup_event.set() is executed in
a finally clause: keep the bootup_event.wait() before trying to create the
interface, perform the interface_bootup call (using
get_interactive_execute_command and spawner) inside a try, and always call
next_bootup_event.set() in the finally so the next worker is unblocked even if
interface_bootup raises; optionally re-raise or log the exception inside the
try/except around interface_bootup if you want error visibility.
🧹 Nitpick comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

116-126: Dynamic worker addition skips boot-up ordering (and worker_id).

The max_workers setter creates new threads using only self._process_kwargs, which doesn't include bootup_event, next_bootup_event, or worker_id. The None defaults prevent a crash, but dynamically added workers won't have ordered boot-up or correct worker_id assignment.

This is pre-existing behavior and not a regression from this PR, but worth noting if the setter is expected to be used in contexts where ordering matters.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
116 - 126, The dynamic worker-add path in the max_workers setter builds new
Thread instances using only self._process_kwargs which omits bootup_event,
next_bootup_event, and worker_id, so newly added threads skip boot-up ordering
and get wrong worker IDs; update the branch that creates new_process_lst (and
references _execute_multiple_tasks and self._process_kwargs) to construct
per-thread kwargs mirroring the initial worker creation: set worker_id =
self._max_workers + i for each new thread, pass the correct bootup_event and
next_bootup_event (chain them the same way the initial spawn logic does), and
then start and append those threads to self._process so dynamically added
workers follow the same boot ordering and ID assignment as existing workers.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 254-267: The chain of bootstrapped workers can deadlock if
interface_bootup raises because next_bootup_event.set() is skipped; modify the
block around the call to interface_bootup (the code that waits on bootup_event,
calls interface_bootup with get_interactive_execute_command and spawner, and
then calls next_bootup_event.set()) to ensure next_bootup_event.set() is
executed in a finally clause: keep the bootup_event.wait() before trying to
create the interface, perform the interface_bootup call (using
get_interactive_execute_command and spawner) inside a try, and always call
next_bootup_event.set() in the finally so the next worker is unblocked even if
interface_bootup raises; optionally re-raise or log the exception inside the
try/except around interface_bootup if you want error visibility.
---
Nitpick comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 116-126: The dynamic worker-add path in the max_workers setter
builds new Thread instances using only self._process_kwargs which omits
bootup_event, next_bootup_event, and worker_id, so newly added threads skip
boot-up ordering and get wrong worker IDs; update the branch that creates
new_process_lst (and references _execute_multiple_tasks and
self._process_kwargs) to construct per-thread kwargs mirroring the initial
worker creation: set worker_id = self._max_workers + i for each new thread, pass
the correct bootup_event and next_bootup_event (chain them the same way the
initial spawn logic does), and then start and append those threads to
self._process so dynamically added workers follow the same boot ordering and ID
assignment as existing workers.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

256-269: ⚠️ Potential issue | 🔴 Critical

Deadlock if interface_bootup raises an exception.

If interface_bootup throws (e.g., spawner failure, connection timeout), next_bootup_event.set() on line 268 is never reached. Every subsequent worker will block forever on bootup_event.wait(), and shutdown(wait=True) will also hang because Thread.join() blocks on those stuck threads.

Wrap the boot-up + signal in try/finally so the chain always advances.

🔒 Proposed fix
 if bootup_event is not None:
bootup_event.wait()
- interface = interface_bootup(- command_lst=get_interactive_execute_command(- cores=cores,- ),- connections=spawner(cores=cores, worker_id=worker_id, **kwargs),- hostname_localhost=hostname_localhost,- log_obj_size=log_obj_size,- worker_id=worker_id,- stop_function=stop_function,- )- if next_bootup_event is not None:- next_bootup_event.set()+ try:+ interface = interface_bootup(+ command_lst=get_interactive_execute_command(+ cores=cores,+ ),+ connections=spawner(cores=cores, worker_id=worker_id, **kwargs),+ hostname_localhost=hostname_localhost,+ log_obj_size=log_obj_size,+ worker_id=worker_id,+ stop_function=stop_function,+ )+ finally:+ if next_bootup_event is not None:+ next_bootup_event.set()

Note: after the finally block, if interface_bootup did raise, the exception will propagate and terminate this worker thread. The rest of the function (init_function, task loop) will be skipped — but at least the remaining workers won't deadlock.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
256 - 269, The boot-up chain can deadlock if interface_bootup (called with
get_interactive_execute_command(...) and connections=spawner(...)) raises,
because next_bootup_event.set() is never reached; wrap the call to
interface_bootup and the subsequent signaling in a try/finally so that
next_bootup_event.set() is always executed in the finally block (still allowing
the exception to propagate and terminate this worker thread), ensuring
bootup_event.wait()/next_bootup_event.set() chain always advances.
🧹 Nitpick comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

256-257: bootup_event.wait() is not interruptible by shutdown().

If shutdown() is called while workers are still queued behind bootup_event.wait(), those threads cannot observe the stop signal — stop_function is only evaluated inside interface_bootup, not here. With the try/finally fix above, the chain will unblock naturally once the currently booting worker finishes (or fails), so this is mostly a latency concern rather than a hard deadlock. However, for a fast-exit guarantee you could add a timeout loop:

ifbootup_eventisnotNone:
whilenotbootup_event.wait(timeout=1.0):
ifstop_functionisnotNoneandstop_function():
return

This is optional and can be deferred.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
256 - 257, The current unconditional bootup_event.wait() call can block threads
from observing shutdown; change it to a timeout loop that repeatedly waits with
a short timeout and checks stop_function between waits so threads can return
early on shutdown. Replace the direct call to bootup_event.wait() (in
blockallocation where bootup_event is used alongside interface_bootup) with a
loop that calls bootup_event.wait(timeout=...) and after each timeout calls
stop_function() (if provided) and returns if it indicates shutdown; keep
behavior of eventually proceeding when bootup_event is set. Ensure you reference
bootup_event and stop_function and preserve existing cleanup/try/finally
semantics around interface_bootup.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 256-269: The boot-up chain can deadlock if interface_bootup
(called with get_interactive_execute_command(...) and connections=spawner(...))
raises, because next_bootup_event.set() is never reached; wrap the call to
interface_bootup and the subsequent signaling in a try/finally so that
next_bootup_event.set() is always executed in the finally block (still allowing
the exception to propagate and terminate this worker thread), ensuring
bootup_event.wait()/next_bootup_event.set() chain always advances.
---
Nitpick comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 256-257: The current unconditional bootup_event.wait() call can
block threads from observing shutdown; change it to a timeout loop that
repeatedly waits with a short timeout and checks stop_function between waits so
threads can return early on shutdown. Replace the direct call to
bootup_event.wait() (in blockallocation where bootup_event is used alongside
interface_bootup) with a loop that calls bootup_event.wait(timeout=...) and
after each timeout calls stop_function() (if provided) and returns if it
indicates shutdown; keep behavior of eventually proceeding when bootup_event is
set. Ensure you reference bootup_event and stop_function and preserve existing
cleanup/try/finally semantics around interface_bootup.

@codecov

codecovBot commented Feb 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.09%. Comparing base (2fa06c8) to head (e099b31).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #935 +/- ##
==========================================
+ Coverage 94.07% 94.09% +0.02% 
==========================================
Files 39 39 Lines 2041 2049 +8 ==========================================
+ Hits 1920 1928 +8 
Misses 121 121 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jan-janssenjan-janssen left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me

@jan-janssenjan-janssen changed the title fix: serialize worker job submissions to preserve worker_id ordering[Fix] serialize worker job submissions to preserve worker_id orderingFeb 22, 2026
@jan-janssen
jan-janssen merged commit 693ca9e into pyiron:mainFeb 22, 2026
59 of 62 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@IlgarBaghishov@jan-janssen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[Fix] serialize worker job submissions to preserve worker_id ordering - #935

Merged
jan-janssen merged 5 commits into
pyiron:mainfrom
IlgarBaghishov:fix/ordered-worker-submission
Feb 22, 2026
Merged

[Fix] serialize worker job submissions to preserve worker_id ordering#935
jan-janssen merged 5 commits into
pyiron:mainfrom
IlgarBaghishov:fix/ordered-worker-submission

Conversation

@IlgarBaghishov

@IlgarBaghishovIlgarBaghishov commented Feb 21, 2026

Copy link
Copy Markdown
Contributor

Worker threads were submitting Flux jobs concurrently, causing
executorlib_worker_id to not correspond to the Flux scheduling order.
This made worker_id unreliable for resource mapping (e.g. GPU assignment).

Two changes:

  • Add threading.Event chain in BlockAllocationTaskScheduler so each worker waits for the previous worker to finish submitting before starting its own submission.
  • Call self._future.jobid() after FluxExecutor.submit() to block until the job is actually registered with the Flux broker, not just queued in the async FluxExecutor.

Summary by CodeRabbit

  • Improvements
    • Enforced ordered worker startup so workers initialize sequentially, improving predictability and reducing race conditions during startup.
    • Immediately capture job identifiers after submission to improve launch observability and make initial job tracking more reliable.

 Worker threads were submitting Flux jobs concurrently, causing
executorlib_worker_id to not correspond to the Flux scheduling order.
This made worker_id unreliable for resource mapping (e.g. GPU assignment).
Two changes:
- Add threading.Event chain in BlockAllocationTaskScheduler so each
worker waits for the previous worker to finish submitting before
starting its own submission.
- Call self._future.jobid() after FluxExecutor.submit() to block until
the job is actually registered with the Flux broker, not just queued
in the async FluxExecutor.
@coderabbitai

coderabbitaiBot commented Feb 21, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Adds ordered per-worker boot-up synchronization using threading.Event objects and updates _execute_multiple_tasks signature to accept bootup_event/next_bootup_event. Also calls self._future.jobid() immediately after submitting the jobspec in the Flux spawner bootup.

Changes

Cohort / File(s)Summary
Worker Boot-up Synchronization
src/executorlib/task_scheduler/interactive/blockallocation.py
Add per-worker bootup_events, pass bootup_event / next_bootup_event via executor kwargs, extend _execute_multiple_tasks signature to accept these events, wait on bootup_event before creating the worker interface, and set next_bootup_event after initialization (docstring updated).
Flux Spawner Bootup
src/executorlib/task_scheduler/interactive/spawner_flux.py
After submitting the jobspec during bootup, call self._future.jobid() conditionally when self._future is not None; no other control flow or public API changes.
Metadata
manifest_file, pyproject.toml
Minor packaging/manifest line edits (few-line changes).

Sequence Diagram(s)

sequenceDiagram
participant Scheduler as Scheduler (main)
participant WorkerA as Worker-0 thread
participant WorkerB as Worker-1 thread
participant FluxSpawner as FluxSpawner
participant Future as Future
Scheduler->>WorkerA: start thread with bootup_event A, next_bootup_event B
Scheduler->>WorkerB: start thread with bootup_event B, next_bootup_event C
WorkerA-->>Scheduler: bootup_event A.wait()
WorkerA->>WorkerA: create interface / submit tasks
WorkerA->>Scheduler: signal next_bootup_event B (set)
WorkerA->>FluxSpawner: submit jobspec
FluxSpawner->>Future: (if present) call jobid()
WorkerB-->>Scheduler: bootup_event B.wait()
WorkerB->>WorkerB: create interface / submit tasks
WorkerB->>Scheduler: signal next_bootup_event C (set)
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Poem

🐰 I lined the threads in tidy rows,

One hops up, then the next one goes.
A gentle nudge, a tiny ping,
The bootup choir begins to sing.
Hops in order — what a thing!

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 75.00% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title 'serialize worker job submissions to preserve worker_id ordering' accurately captures the main objective of the PR—implementing sequential worker submission to maintain reliable worker_id ordering in Flux scheduling.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

254-267: ⚠️ Potential issue | 🔴 Critical

Deadlock if interface_bootup raises: subsequent workers hang forever.

If interface_bootup (line 256) throws an exception for worker N, the thread terminates without calling next_bootup_event.set(). Workers N+1, N+2, … remain blocked on bootup_event.wait() indefinitely, and shutdown(wait=True) will hang on process.join().

Wrap the bootup + signal in try/finally so the chain always progresses:

🔒 Proposed fix using try/finally
 if bootup_event is not None:
bootup_event.wait()
- interface = interface_bootup(- command_lst=get_interactive_execute_command(- cores=cores,- ),- connections=spawner(cores=cores, worker_id=worker_id, **kwargs),- hostname_localhost=hostname_localhost,- log_obj_size=log_obj_size,- worker_id=worker_id,- stop_function=stop_function,- )- if next_bootup_event is not None:- next_bootup_event.set()+ try:+ interface = interface_bootup(+ command_lst=get_interactive_execute_command(+ cores=cores,+ ),+ connections=spawner(cores=cores, worker_id=worker_id, **kwargs),+ hostname_localhost=hostname_localhost,+ log_obj_size=log_obj_size,+ worker_id=worker_id,+ stop_function=stop_function,+ )+ finally:+ if next_bootup_event is not None:+ next_bootup_event.set()

Note: if bootup fails and the chain unblocks the next worker, subsequent workers will also likely fail. But that's preferable to a silent deadlock. You may also want to consider using bootup_event.wait(timeout=...) as additional protection.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
254 - 267, The chain of bootstrapped workers can deadlock if interface_bootup
raises because next_bootup_event.set() is skipped; modify the block around the
call to interface_bootup (the code that waits on bootup_event, calls
interface_bootup with get_interactive_execute_command and spawner, and then
calls next_bootup_event.set()) to ensure next_bootup_event.set() is executed in
a finally clause: keep the bootup_event.wait() before trying to create the
interface, perform the interface_bootup call (using
get_interactive_execute_command and spawner) inside a try, and always call
next_bootup_event.set() in the finally so the next worker is unblocked even if
interface_bootup raises; optionally re-raise or log the exception inside the
try/except around interface_bootup if you want error visibility.
🧹 Nitpick comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

116-126: Dynamic worker addition skips boot-up ordering (and worker_id).

The max_workers setter creates new threads using only self._process_kwargs, which doesn't include bootup_event, next_bootup_event, or worker_id. The None defaults prevent a crash, but dynamically added workers won't have ordered boot-up or correct worker_id assignment.

This is pre-existing behavior and not a regression from this PR, but worth noting if the setter is expected to be used in contexts where ordering matters.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
116 - 126, The dynamic worker-add path in the max_workers setter builds new
Thread instances using only self._process_kwargs which omits bootup_event,
next_bootup_event, and worker_id, so newly added threads skip boot-up ordering
and get wrong worker IDs; update the branch that creates new_process_lst (and
references _execute_multiple_tasks and self._process_kwargs) to construct
per-thread kwargs mirroring the initial worker creation: set worker_id =
self._max_workers + i for each new thread, pass the correct bootup_event and
next_bootup_event (chain them the same way the initial spawn logic does), and
then start and append those threads to self._process so dynamically added
workers follow the same boot ordering and ID assignment as existing workers.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 254-267: The chain of bootstrapped workers can deadlock if
interface_bootup raises because next_bootup_event.set() is skipped; modify the
block around the call to interface_bootup (the code that waits on bootup_event,
calls interface_bootup with get_interactive_execute_command and spawner, and
then calls next_bootup_event.set()) to ensure next_bootup_event.set() is
executed in a finally clause: keep the bootup_event.wait() before trying to
create the interface, perform the interface_bootup call (using
get_interactive_execute_command and spawner) inside a try, and always call
next_bootup_event.set() in the finally so the next worker is unblocked even if
interface_bootup raises; optionally re-raise or log the exception inside the
try/except around interface_bootup if you want error visibility.
---
Nitpick comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 116-126: The dynamic worker-add path in the max_workers setter
builds new Thread instances using only self._process_kwargs which omits
bootup_event, next_bootup_event, and worker_id, so newly added threads skip
boot-up ordering and get wrong worker IDs; update the branch that creates
new_process_lst (and references _execute_multiple_tasks and
self._process_kwargs) to construct per-thread kwargs mirroring the initial
worker creation: set worker_id = self._max_workers + i for each new thread, pass
the correct bootup_event and next_bootup_event (chain them the same way the
initial spawn logic does), and then start and append those threads to
self._process so dynamically added workers follow the same boot ordering and ID
assignment as existing workers.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

256-269: ⚠️ Potential issue | 🔴 Critical

Deadlock if interface_bootup raises an exception.

If interface_bootup throws (e.g., spawner failure, connection timeout), next_bootup_event.set() on line 268 is never reached. Every subsequent worker will block forever on bootup_event.wait(), and shutdown(wait=True) will also hang because Thread.join() blocks on those stuck threads.

Wrap the boot-up + signal in try/finally so the chain always advances.

🔒 Proposed fix
 if bootup_event is not None:
bootup_event.wait()
- interface = interface_bootup(- command_lst=get_interactive_execute_command(- cores=cores,- ),- connections=spawner(cores=cores, worker_id=worker_id, **kwargs),- hostname_localhost=hostname_localhost,- log_obj_size=log_obj_size,- worker_id=worker_id,- stop_function=stop_function,- )- if next_bootup_event is not None:- next_bootup_event.set()+ try:+ interface = interface_bootup(+ command_lst=get_interactive_execute_command(+ cores=cores,+ ),+ connections=spawner(cores=cores, worker_id=worker_id, **kwargs),+ hostname_localhost=hostname_localhost,+ log_obj_size=log_obj_size,+ worker_id=worker_id,+ stop_function=stop_function,+ )+ finally:+ if next_bootup_event is not None:+ next_bootup_event.set()

Note: after the finally block, if interface_bootup did raise, the exception will propagate and terminate this worker thread. The rest of the function (init_function, task loop) will be skipped — but at least the remaining workers won't deadlock.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
256 - 269, The boot-up chain can deadlock if interface_bootup (called with
get_interactive_execute_command(...) and connections=spawner(...)) raises,
because next_bootup_event.set() is never reached; wrap the call to
interface_bootup and the subsequent signaling in a try/finally so that
next_bootup_event.set() is always executed in the finally block (still allowing
the exception to propagate and terminate this worker thread), ensuring
bootup_event.wait()/next_bootup_event.set() chain always advances.
🧹 Nitpick comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

256-257: bootup_event.wait() is not interruptible by shutdown().

If shutdown() is called while workers are still queued behind bootup_event.wait(), those threads cannot observe the stop signal — stop_function is only evaluated inside interface_bootup, not here. With the try/finally fix above, the chain will unblock naturally once the currently booting worker finishes (or fails), so this is mostly a latency concern rather than a hard deadlock. However, for a fast-exit guarantee you could add a timeout loop:

ifbootup_eventisnotNone:
whilenotbootup_event.wait(timeout=1.0):
ifstop_functionisnotNoneandstop_function():
return

This is optional and can be deferred.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
256 - 257, The current unconditional bootup_event.wait() call can block threads
from observing shutdown; change it to a timeout loop that repeatedly waits with
a short timeout and checks stop_function between waits so threads can return
early on shutdown. Replace the direct call to bootup_event.wait() (in
blockallocation where bootup_event is used alongside interface_bootup) with a
loop that calls bootup_event.wait(timeout=...) and after each timeout calls
stop_function() (if provided) and returns if it indicates shutdown; keep
behavior of eventually proceeding when bootup_event is set. Ensure you reference
bootup_event and stop_function and preserve existing cleanup/try/finally
semantics around interface_bootup.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 256-269: The boot-up chain can deadlock if interface_bootup
(called with get_interactive_execute_command(...) and connections=spawner(...))
raises, because next_bootup_event.set() is never reached; wrap the call to
interface_bootup and the subsequent signaling in a try/finally so that
next_bootup_event.set() is always executed in the finally block (still allowing
the exception to propagate and terminate this worker thread), ensuring
bootup_event.wait()/next_bootup_event.set() chain always advances.
---
Nitpick comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 256-257: The current unconditional bootup_event.wait() call can
block threads from observing shutdown; change it to a timeout loop that
repeatedly waits with a short timeout and checks stop_function between waits so
threads can return early on shutdown. Replace the direct call to
bootup_event.wait() (in blockallocation where bootup_event is used alongside
interface_bootup) with a loop that calls bootup_event.wait(timeout=...) and
after each timeout calls stop_function() (if provided) and returns if it
indicates shutdown; keep behavior of eventually proceeding when bootup_event is
set. Ensure you reference bootup_event and stop_function and preserve existing
cleanup/try/finally semantics around interface_bootup.

@codecov

codecovBot commented Feb 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.09%. Comparing base (2fa06c8) to head (e099b31).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #935 +/- ##
==========================================
+ Coverage 94.07% 94.09% +0.02% 
==========================================
Files 39 39 Lines 2041 2049 +8 ==========================================
+ Hits 1920 1928 +8 
Misses 121 121 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jan-janssenjan-janssen left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me

@jan-janssenjan-janssen changed the title fix: serialize worker job submissions to preserve worker_id ordering[Fix] serialize worker job submissions to preserve worker_id orderingFeb 22, 2026
@jan-janssen
jan-janssen merged commit 693ca9e into pyiron:mainFeb 22, 2026
59 of 62 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@IlgarBaghishov@jan-janssen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Fix] serialize worker job submissions to preserve worker_id ordering - #935

Merged
jan-janssen merged 5 commits into
pyiron:mainfrom
IlgarBaghishov:fix/ordered-worker-submission
Feb 22, 2026
Merged

[Fix] serialize worker job submissions to preserve worker_id ordering#935
jan-janssen merged 5 commits into
pyiron:mainfrom
IlgarBaghishov:fix/ordered-worker-submission

Conversation

@IlgarBaghishov

@IlgarBaghishovIlgarBaghishov commented Feb 21, 2026

Copy link
Copy Markdown
Contributor

Worker threads were submitting Flux jobs concurrently, causing
executorlib_worker_id to not correspond to the Flux scheduling order.
This made worker_id unreliable for resource mapping (e.g. GPU assignment).

Two changes:

  • Add threading.Event chain in BlockAllocationTaskScheduler so each worker waits for the previous worker to finish submitting before starting its own submission.
  • Call self._future.jobid() after FluxExecutor.submit() to block until the job is actually registered with the Flux broker, not just queued in the async FluxExecutor.

Summary by CodeRabbit

  • Improvements
    • Enforced ordered worker startup so workers initialize sequentially, improving predictability and reducing race conditions during startup.
    • Immediately capture job identifiers after submission to improve launch observability and make initial job tracking more reliable.

 Worker threads were submitting Flux jobs concurrently, causing
executorlib_worker_id to not correspond to the Flux scheduling order.
This made worker_id unreliable for resource mapping (e.g. GPU assignment).
Two changes:
- Add threading.Event chain in BlockAllocationTaskScheduler so each
worker waits for the previous worker to finish submitting before
starting its own submission.
- Call self._future.jobid() after FluxExecutor.submit() to block until
the job is actually registered with the Flux broker, not just queued
in the async FluxExecutor.
@coderabbitai

coderabbitaiBot commented Feb 21, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Adds ordered per-worker boot-up synchronization using threading.Event objects and updates _execute_multiple_tasks signature to accept bootup_event/next_bootup_event. Also calls self._future.jobid() immediately after submitting the jobspec in the Flux spawner bootup.

Changes

Cohort / File(s)Summary
Worker Boot-up Synchronization
src/executorlib/task_scheduler/interactive/blockallocation.py
Add per-worker bootup_events, pass bootup_event / next_bootup_event via executor kwargs, extend _execute_multiple_tasks signature to accept these events, wait on bootup_event before creating the worker interface, and set next_bootup_event after initialization (docstring updated).
Flux Spawner Bootup
src/executorlib/task_scheduler/interactive/spawner_flux.py
After submitting the jobspec during bootup, call self._future.jobid() conditionally when self._future is not None; no other control flow or public API changes.
Metadata
manifest_file, pyproject.toml
Minor packaging/manifest line edits (few-line changes).

Sequence Diagram(s)

sequenceDiagram
participant Scheduler as Scheduler (main)
participant WorkerA as Worker-0 thread
participant WorkerB as Worker-1 thread
participant FluxSpawner as FluxSpawner
participant Future as Future
Scheduler->>WorkerA: start thread with bootup_event A, next_bootup_event B
Scheduler->>WorkerB: start thread with bootup_event B, next_bootup_event C
WorkerA-->>Scheduler: bootup_event A.wait()
WorkerA->>WorkerA: create interface / submit tasks
WorkerA->>Scheduler: signal next_bootup_event B (set)
WorkerA->>FluxSpawner: submit jobspec
FluxSpawner->>Future: (if present) call jobid()
WorkerB-->>Scheduler: bootup_event B.wait()
WorkerB->>WorkerB: create interface / submit tasks
WorkerB->>Scheduler: signal next_bootup_event C (set)
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Poem

🐰 I lined the threads in tidy rows,

One hops up, then the next one goes.
A gentle nudge, a tiny ping,
The bootup choir begins to sing.
Hops in order — what a thing!

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 75.00% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title 'serialize worker job submissions to preserve worker_id ordering' accurately captures the main objective of the PR—implementing sequential worker submission to maintain reliable worker_id ordering in Flux scheduling.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

254-267: ⚠️ Potential issue | 🔴 Critical

Deadlock if interface_bootup raises: subsequent workers hang forever.

If interface_bootup (line 256) throws an exception for worker N, the thread terminates without calling next_bootup_event.set(). Workers N+1, N+2, … remain blocked on bootup_event.wait() indefinitely, and shutdown(wait=True) will hang on process.join().

Wrap the bootup + signal in try/finally so the chain always progresses:

🔒 Proposed fix using try/finally
 if bootup_event is not None:
bootup_event.wait()
- interface = interface_bootup(- command_lst=get_interactive_execute_command(- cores=cores,- ),- connections=spawner(cores=cores, worker_id=worker_id, **kwargs),- hostname_localhost=hostname_localhost,- log_obj_size=log_obj_size,- worker_id=worker_id,- stop_function=stop_function,- )- if next_bootup_event is not None:- next_bootup_event.set()+ try:+ interface = interface_bootup(+ command_lst=get_interactive_execute_command(+ cores=cores,+ ),+ connections=spawner(cores=cores, worker_id=worker_id, **kwargs),+ hostname_localhost=hostname_localhost,+ log_obj_size=log_obj_size,+ worker_id=worker_id,+ stop_function=stop_function,+ )+ finally:+ if next_bootup_event is not None:+ next_bootup_event.set()

Note: if bootup fails and the chain unblocks the next worker, subsequent workers will also likely fail. But that's preferable to a silent deadlock. You may also want to consider using bootup_event.wait(timeout=...) as additional protection.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
254 - 267, The chain of bootstrapped workers can deadlock if interface_bootup
raises because next_bootup_event.set() is skipped; modify the block around the
call to interface_bootup (the code that waits on bootup_event, calls
interface_bootup with get_interactive_execute_command and spawner, and then
calls next_bootup_event.set()) to ensure next_bootup_event.set() is executed in
a finally clause: keep the bootup_event.wait() before trying to create the
interface, perform the interface_bootup call (using
get_interactive_execute_command and spawner) inside a try, and always call
next_bootup_event.set() in the finally so the next worker is unblocked even if
interface_bootup raises; optionally re-raise or log the exception inside the
try/except around interface_bootup if you want error visibility.
🧹 Nitpick comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

116-126: Dynamic worker addition skips boot-up ordering (and worker_id).

The max_workers setter creates new threads using only self._process_kwargs, which doesn't include bootup_event, next_bootup_event, or worker_id. The None defaults prevent a crash, but dynamically added workers won't have ordered boot-up or correct worker_id assignment.

This is pre-existing behavior and not a regression from this PR, but worth noting if the setter is expected to be used in contexts where ordering matters.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
116 - 126, The dynamic worker-add path in the max_workers setter builds new
Thread instances using only self._process_kwargs which omits bootup_event,
next_bootup_event, and worker_id, so newly added threads skip boot-up ordering
and get wrong worker IDs; update the branch that creates new_process_lst (and
references _execute_multiple_tasks and self._process_kwargs) to construct
per-thread kwargs mirroring the initial worker creation: set worker_id =
self._max_workers + i for each new thread, pass the correct bootup_event and
next_bootup_event (chain them the same way the initial spawn logic does), and
then start and append those threads to self._process so dynamically added
workers follow the same boot ordering and ID assignment as existing workers.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 254-267: The chain of bootstrapped workers can deadlock if
interface_bootup raises because next_bootup_event.set() is skipped; modify the
block around the call to interface_bootup (the code that waits on bootup_event,
calls interface_bootup with get_interactive_execute_command and spawner, and
then calls next_bootup_event.set()) to ensure next_bootup_event.set() is
executed in a finally clause: keep the bootup_event.wait() before trying to
create the interface, perform the interface_bootup call (using
get_interactive_execute_command and spawner) inside a try, and always call
next_bootup_event.set() in the finally so the next worker is unblocked even if
interface_bootup raises; optionally re-raise or log the exception inside the
try/except around interface_bootup if you want error visibility.
---
Nitpick comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 116-126: The dynamic worker-add path in the max_workers setter
builds new Thread instances using only self._process_kwargs which omits
bootup_event, next_bootup_event, and worker_id, so newly added threads skip
boot-up ordering and get wrong worker IDs; update the branch that creates
new_process_lst (and references _execute_multiple_tasks and
self._process_kwargs) to construct per-thread kwargs mirroring the initial
worker creation: set worker_id = self._max_workers + i for each new thread, pass
the correct bootup_event and next_bootup_event (chain them the same way the
initial spawn logic does), and then start and append those threads to
self._process so dynamically added workers follow the same boot ordering and ID
assignment as existing workers.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

256-269: ⚠️ Potential issue | 🔴 Critical

Deadlock if interface_bootup raises an exception.

If interface_bootup throws (e.g., spawner failure, connection timeout), next_bootup_event.set() on line 268 is never reached. Every subsequent worker will block forever on bootup_event.wait(), and shutdown(wait=True) will also hang because Thread.join() blocks on those stuck threads.

Wrap the boot-up + signal in try/finally so the chain always advances.

🔒 Proposed fix
 if bootup_event is not None:
bootup_event.wait()
- interface = interface_bootup(- command_lst=get_interactive_execute_command(- cores=cores,- ),- connections=spawner(cores=cores, worker_id=worker_id, **kwargs),- hostname_localhost=hostname_localhost,- log_obj_size=log_obj_size,- worker_id=worker_id,- stop_function=stop_function,- )- if next_bootup_event is not None:- next_bootup_event.set()+ try:+ interface = interface_bootup(+ command_lst=get_interactive_execute_command(+ cores=cores,+ ),+ connections=spawner(cores=cores, worker_id=worker_id, **kwargs),+ hostname_localhost=hostname_localhost,+ log_obj_size=log_obj_size,+ worker_id=worker_id,+ stop_function=stop_function,+ )+ finally:+ if next_bootup_event is not None:+ next_bootup_event.set()

Note: after the finally block, if interface_bootup did raise, the exception will propagate and terminate this worker thread. The rest of the function (init_function, task loop) will be skipped — but at least the remaining workers won't deadlock.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
256 - 269, The boot-up chain can deadlock if interface_bootup (called with
get_interactive_execute_command(...) and connections=spawner(...)) raises,
because next_bootup_event.set() is never reached; wrap the call to
interface_bootup and the subsequent signaling in a try/finally so that
next_bootup_event.set() is always executed in the finally block (still allowing
the exception to propagate and terminate this worker thread), ensuring
bootup_event.wait()/next_bootup_event.set() chain always advances.
🧹 Nitpick comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

256-257: bootup_event.wait() is not interruptible by shutdown().

If shutdown() is called while workers are still queued behind bootup_event.wait(), those threads cannot observe the stop signal — stop_function is only evaluated inside interface_bootup, not here. With the try/finally fix above, the chain will unblock naturally once the currently booting worker finishes (or fails), so this is mostly a latency concern rather than a hard deadlock. However, for a fast-exit guarantee you could add a timeout loop:

ifbootup_eventisnotNone:
whilenotbootup_event.wait(timeout=1.0):
ifstop_functionisnotNoneandstop_function():
return

This is optional and can be deferred.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
256 - 257, The current unconditional bootup_event.wait() call can block threads
from observing shutdown; change it to a timeout loop that repeatedly waits with
a short timeout and checks stop_function between waits so threads can return
early on shutdown. Replace the direct call to bootup_event.wait() (in
blockallocation where bootup_event is used alongside interface_bootup) with a
loop that calls bootup_event.wait(timeout=...) and after each timeout calls
stop_function() (if provided) and returns if it indicates shutdown; keep
behavior of eventually proceeding when bootup_event is set. Ensure you reference
bootup_event and stop_function and preserve existing cleanup/try/finally
semantics around interface_bootup.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 256-269: The boot-up chain can deadlock if interface_bootup
(called with get_interactive_execute_command(...) and connections=spawner(...))
raises, because next_bootup_event.set() is never reached; wrap the call to
interface_bootup and the subsequent signaling in a try/finally so that
next_bootup_event.set() is always executed in the finally block (still allowing
the exception to propagate and terminate this worker thread), ensuring
bootup_event.wait()/next_bootup_event.set() chain always advances.
---
Nitpick comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 256-257: The current unconditional bootup_event.wait() call can
block threads from observing shutdown; change it to a timeout loop that
repeatedly waits with a short timeout and checks stop_function between waits so
threads can return early on shutdown. Replace the direct call to
bootup_event.wait() (in blockallocation where bootup_event is used alongside
interface_bootup) with a loop that calls bootup_event.wait(timeout=...) and
after each timeout calls stop_function() (if provided) and returns if it
indicates shutdown; keep behavior of eventually proceeding when bootup_event is
set. Ensure you reference bootup_event and stop_function and preserve existing
cleanup/try/finally semantics around interface_bootup.

@codecov

codecovBot commented Feb 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.09%. Comparing base (2fa06c8) to head (e099b31).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #935 +/- ##
==========================================
+ Coverage 94.07% 94.09% +0.02% 
==========================================
Files 39 39 Lines 2041 2049 +8 ==========================================
+ Hits 1920 1928 +8 
Misses 121 121 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jan-janssenjan-janssen left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me

@jan-janssenjan-janssen changed the title fix: serialize worker job submissions to preserve worker_id ordering[Fix] serialize worker job submissions to preserve worker_id orderingFeb 22, 2026
@jan-janssen
jan-janssen merged commit 693ca9e into pyiron:mainFeb 22, 2026
59 of 62 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@IlgarBaghishov@jan-janssen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Fix] serialize worker job submissions to preserve worker_id ordering - #935

Merged
jan-janssen merged 5 commits into
pyiron:mainfrom
IlgarBaghishov:fix/ordered-worker-submission
Feb 22, 2026
Merged

[Fix] serialize worker job submissions to preserve worker_id ordering#935
jan-janssen merged 5 commits into
pyiron:mainfrom
IlgarBaghishov:fix/ordered-worker-submission

Conversation

@IlgarBaghishov

@IlgarBaghishovIlgarBaghishov commented Feb 21, 2026

Copy link
Copy Markdown
Contributor

Worker threads were submitting Flux jobs concurrently, causing
executorlib_worker_id to not correspond to the Flux scheduling order.
This made worker_id unreliable for resource mapping (e.g. GPU assignment).

Two changes:

  • Add threading.Event chain in BlockAllocationTaskScheduler so each worker waits for the previous worker to finish submitting before starting its own submission.
  • Call self._future.jobid() after FluxExecutor.submit() to block until the job is actually registered with the Flux broker, not just queued in the async FluxExecutor.

Summary by CodeRabbit

  • Improvements
    • Enforced ordered worker startup so workers initialize sequentially, improving predictability and reducing race conditions during startup.
    • Immediately capture job identifiers after submission to improve launch observability and make initial job tracking more reliable.

 Worker threads were submitting Flux jobs concurrently, causing
executorlib_worker_id to not correspond to the Flux scheduling order.
This made worker_id unreliable for resource mapping (e.g. GPU assignment).
Two changes:
- Add threading.Event chain in BlockAllocationTaskScheduler so each
worker waits for the previous worker to finish submitting before
starting its own submission.
- Call self._future.jobid() after FluxExecutor.submit() to block until
the job is actually registered with the Flux broker, not just queued
in the async FluxExecutor.
@coderabbitai

coderabbitaiBot commented Feb 21, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Adds ordered per-worker boot-up synchronization using threading.Event objects and updates _execute_multiple_tasks signature to accept bootup_event/next_bootup_event. Also calls self._future.jobid() immediately after submitting the jobspec in the Flux spawner bootup.

Changes

Cohort / File(s)Summary
Worker Boot-up Synchronization
src/executorlib/task_scheduler/interactive/blockallocation.py
Add per-worker bootup_events, pass bootup_event / next_bootup_event via executor kwargs, extend _execute_multiple_tasks signature to accept these events, wait on bootup_event before creating the worker interface, and set next_bootup_event after initialization (docstring updated).
Flux Spawner Bootup
src/executorlib/task_scheduler/interactive/spawner_flux.py
After submitting the jobspec during bootup, call self._future.jobid() conditionally when self._future is not None; no other control flow or public API changes.
Metadata
manifest_file, pyproject.toml
Minor packaging/manifest line edits (few-line changes).

Sequence Diagram(s)

sequenceDiagram
participant Scheduler as Scheduler (main)
participant WorkerA as Worker-0 thread
participant WorkerB as Worker-1 thread
participant FluxSpawner as FluxSpawner
participant Future as Future
Scheduler->>WorkerA: start thread with bootup_event A, next_bootup_event B
Scheduler->>WorkerB: start thread with bootup_event B, next_bootup_event C
WorkerA-->>Scheduler: bootup_event A.wait()
WorkerA->>WorkerA: create interface / submit tasks
WorkerA->>Scheduler: signal next_bootup_event B (set)
WorkerA->>FluxSpawner: submit jobspec
FluxSpawner->>Future: (if present) call jobid()
WorkerB-->>Scheduler: bootup_event B.wait()
WorkerB->>WorkerB: create interface / submit tasks
WorkerB->>Scheduler: signal next_bootup_event C (set)
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Poem

🐰 I lined the threads in tidy rows,

One hops up, then the next one goes.
A gentle nudge, a tiny ping,
The bootup choir begins to sing.
Hops in order — what a thing!

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 75.00% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title 'serialize worker job submissions to preserve worker_id ordering' accurately captures the main objective of the PR—implementing sequential worker submission to maintain reliable worker_id ordering in Flux scheduling.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

254-267: ⚠️ Potential issue | 🔴 Critical

Deadlock if interface_bootup raises: subsequent workers hang forever.

If interface_bootup (line 256) throws an exception for worker N, the thread terminates without calling next_bootup_event.set(). Workers N+1, N+2, … remain blocked on bootup_event.wait() indefinitely, and shutdown(wait=True) will hang on process.join().

Wrap the bootup + signal in try/finally so the chain always progresses:

🔒 Proposed fix using try/finally
 if bootup_event is not None:
bootup_event.wait()
- interface = interface_bootup(- command_lst=get_interactive_execute_command(- cores=cores,- ),- connections=spawner(cores=cores, worker_id=worker_id, **kwargs),- hostname_localhost=hostname_localhost,- log_obj_size=log_obj_size,- worker_id=worker_id,- stop_function=stop_function,- )- if next_bootup_event is not None:- next_bootup_event.set()+ try:+ interface = interface_bootup(+ command_lst=get_interactive_execute_command(+ cores=cores,+ ),+ connections=spawner(cores=cores, worker_id=worker_id, **kwargs),+ hostname_localhost=hostname_localhost,+ log_obj_size=log_obj_size,+ worker_id=worker_id,+ stop_function=stop_function,+ )+ finally:+ if next_bootup_event is not None:+ next_bootup_event.set()

Note: if bootup fails and the chain unblocks the next worker, subsequent workers will also likely fail. But that's preferable to a silent deadlock. You may also want to consider using bootup_event.wait(timeout=...) as additional protection.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
254 - 267, The chain of bootstrapped workers can deadlock if interface_bootup
raises because next_bootup_event.set() is skipped; modify the block around the
call to interface_bootup (the code that waits on bootup_event, calls
interface_bootup with get_interactive_execute_command and spawner, and then
calls next_bootup_event.set()) to ensure next_bootup_event.set() is executed in
a finally clause: keep the bootup_event.wait() before trying to create the
interface, perform the interface_bootup call (using
get_interactive_execute_command and spawner) inside a try, and always call
next_bootup_event.set() in the finally so the next worker is unblocked even if
interface_bootup raises; optionally re-raise or log the exception inside the
try/except around interface_bootup if you want error visibility.
🧹 Nitpick comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

116-126: Dynamic worker addition skips boot-up ordering (and worker_id).

The max_workers setter creates new threads using only self._process_kwargs, which doesn't include bootup_event, next_bootup_event, or worker_id. The None defaults prevent a crash, but dynamically added workers won't have ordered boot-up or correct worker_id assignment.

This is pre-existing behavior and not a regression from this PR, but worth noting if the setter is expected to be used in contexts where ordering matters.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
116 - 126, The dynamic worker-add path in the max_workers setter builds new
Thread instances using only self._process_kwargs which omits bootup_event,
next_bootup_event, and worker_id, so newly added threads skip boot-up ordering
and get wrong worker IDs; update the branch that creates new_process_lst (and
references _execute_multiple_tasks and self._process_kwargs) to construct
per-thread kwargs mirroring the initial worker creation: set worker_id =
self._max_workers + i for each new thread, pass the correct bootup_event and
next_bootup_event (chain them the same way the initial spawn logic does), and
then start and append those threads to self._process so dynamically added
workers follow the same boot ordering and ID assignment as existing workers.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 254-267: The chain of bootstrapped workers can deadlock if
interface_bootup raises because next_bootup_event.set() is skipped; modify the
block around the call to interface_bootup (the code that waits on bootup_event,
calls interface_bootup with get_interactive_execute_command and spawner, and
then calls next_bootup_event.set()) to ensure next_bootup_event.set() is
executed in a finally clause: keep the bootup_event.wait() before trying to
create the interface, perform the interface_bootup call (using
get_interactive_execute_command and spawner) inside a try, and always call
next_bootup_event.set() in the finally so the next worker is unblocked even if
interface_bootup raises; optionally re-raise or log the exception inside the
try/except around interface_bootup if you want error visibility.
---
Nitpick comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 116-126: The dynamic worker-add path in the max_workers setter
builds new Thread instances using only self._process_kwargs which omits
bootup_event, next_bootup_event, and worker_id, so newly added threads skip
boot-up ordering and get wrong worker IDs; update the branch that creates
new_process_lst (and references _execute_multiple_tasks and
self._process_kwargs) to construct per-thread kwargs mirroring the initial
worker creation: set worker_id = self._max_workers + i for each new thread, pass
the correct bootup_event and next_bootup_event (chain them the same way the
initial spawn logic does), and then start and append those threads to
self._process so dynamically added workers follow the same boot ordering and ID
assignment as existing workers.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

256-269: ⚠️ Potential issue | 🔴 Critical

Deadlock if interface_bootup raises an exception.

If interface_bootup throws (e.g., spawner failure, connection timeout), next_bootup_event.set() on line 268 is never reached. Every subsequent worker will block forever on bootup_event.wait(), and shutdown(wait=True) will also hang because Thread.join() blocks on those stuck threads.

Wrap the boot-up + signal in try/finally so the chain always advances.

🔒 Proposed fix
 if bootup_event is not None:
bootup_event.wait()
- interface = interface_bootup(- command_lst=get_interactive_execute_command(- cores=cores,- ),- connections=spawner(cores=cores, worker_id=worker_id, **kwargs),- hostname_localhost=hostname_localhost,- log_obj_size=log_obj_size,- worker_id=worker_id,- stop_function=stop_function,- )- if next_bootup_event is not None:- next_bootup_event.set()+ try:+ interface = interface_bootup(+ command_lst=get_interactive_execute_command(+ cores=cores,+ ),+ connections=spawner(cores=cores, worker_id=worker_id, **kwargs),+ hostname_localhost=hostname_localhost,+ log_obj_size=log_obj_size,+ worker_id=worker_id,+ stop_function=stop_function,+ )+ finally:+ if next_bootup_event is not None:+ next_bootup_event.set()

Note: after the finally block, if interface_bootup did raise, the exception will propagate and terminate this worker thread. The rest of the function (init_function, task loop) will be skipped — but at least the remaining workers won't deadlock.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
256 - 269, The boot-up chain can deadlock if interface_bootup (called with
get_interactive_execute_command(...) and connections=spawner(...)) raises,
because next_bootup_event.set() is never reached; wrap the call to
interface_bootup and the subsequent signaling in a try/finally so that
next_bootup_event.set() is always executed in the finally block (still allowing
the exception to propagate and terminate this worker thread), ensuring
bootup_event.wait()/next_bootup_event.set() chain always advances.
🧹 Nitpick comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

256-257: bootup_event.wait() is not interruptible by shutdown().

If shutdown() is called while workers are still queued behind bootup_event.wait(), those threads cannot observe the stop signal — stop_function is only evaluated inside interface_bootup, not here. With the try/finally fix above, the chain will unblock naturally once the currently booting worker finishes (or fails), so this is mostly a latency concern rather than a hard deadlock. However, for a fast-exit guarantee you could add a timeout loop:

ifbootup_eventisnotNone:
whilenotbootup_event.wait(timeout=1.0):
ifstop_functionisnotNoneandstop_function():
return

This is optional and can be deferred.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
256 - 257, The current unconditional bootup_event.wait() call can block threads
from observing shutdown; change it to a timeout loop that repeatedly waits with
a short timeout and checks stop_function between waits so threads can return
early on shutdown. Replace the direct call to bootup_event.wait() (in
blockallocation where bootup_event is used alongside interface_bootup) with a
loop that calls bootup_event.wait(timeout=...) and after each timeout calls
stop_function() (if provided) and returns if it indicates shutdown; keep
behavior of eventually proceeding when bootup_event is set. Ensure you reference
bootup_event and stop_function and preserve existing cleanup/try/finally
semantics around interface_bootup.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 256-269: The boot-up chain can deadlock if interface_bootup
(called with get_interactive_execute_command(...) and connections=spawner(...))
raises, because next_bootup_event.set() is never reached; wrap the call to
interface_bootup and the subsequent signaling in a try/finally so that
next_bootup_event.set() is always executed in the finally block (still allowing
the exception to propagate and terminate this worker thread), ensuring
bootup_event.wait()/next_bootup_event.set() chain always advances.
---
Nitpick comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 256-257: The current unconditional bootup_event.wait() call can
block threads from observing shutdown; change it to a timeout loop that
repeatedly waits with a short timeout and checks stop_function between waits so
threads can return early on shutdown. Replace the direct call to
bootup_event.wait() (in blockallocation where bootup_event is used alongside
interface_bootup) with a loop that calls bootup_event.wait(timeout=...) and
after each timeout calls stop_function() (if provided) and returns if it
indicates shutdown; keep behavior of eventually proceeding when bootup_event is
set. Ensure you reference bootup_event and stop_function and preserve existing
cleanup/try/finally semantics around interface_bootup.

@codecov

codecovBot commented Feb 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.09%. Comparing base (2fa06c8) to head (e099b31).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #935 +/- ##
==========================================
+ Coverage 94.07% 94.09% +0.02% 
==========================================
Files 39 39 Lines 2041 2049 +8 ==========================================
+ Hits 1920 1928 +8 
Misses 121 121 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jan-janssenjan-janssen left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me

@jan-janssenjan-janssen changed the title fix: serialize worker job submissions to preserve worker_id ordering[Fix] serialize worker job submissions to preserve worker_id orderingFeb 22, 2026
@jan-janssen
jan-janssen merged commit 693ca9e into pyiron:mainFeb 22, 2026
59 of 62 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@IlgarBaghishov@jan-janssen
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[Fix] serialize worker job submissions to preserve worker_id ordering - #935

Merged
jan-janssen merged 5 commits into
pyiron:mainfrom
IlgarBaghishov:fix/ordered-worker-submission
Feb 22, 2026
Merged

[Fix] serialize worker job submissions to preserve worker_id ordering#935
jan-janssen merged 5 commits into
pyiron:mainfrom
IlgarBaghishov:fix/ordered-worker-submission

Conversation

@IlgarBaghishov

@IlgarBaghishovIlgarBaghishov commented Feb 21, 2026

Copy link
Copy Markdown
Contributor

Worker threads were submitting Flux jobs concurrently, causing
executorlib_worker_id to not correspond to the Flux scheduling order.
This made worker_id unreliable for resource mapping (e.g. GPU assignment).

Two changes:

  • Add threading.Event chain in BlockAllocationTaskScheduler so each worker waits for the previous worker to finish submitting before starting its own submission.
  • Call self._future.jobid() after FluxExecutor.submit() to block until the job is actually registered with the Flux broker, not just queued in the async FluxExecutor.

Summary by CodeRabbit

  • Improvements
    • Enforced ordered worker startup so workers initialize sequentially, improving predictability and reducing race conditions during startup.
    • Immediately capture job identifiers after submission to improve launch observability and make initial job tracking more reliable.

 Worker threads were submitting Flux jobs concurrently, causing
executorlib_worker_id to not correspond to the Flux scheduling order.
This made worker_id unreliable for resource mapping (e.g. GPU assignment).
Two changes:
- Add threading.Event chain in BlockAllocationTaskScheduler so each
worker waits for the previous worker to finish submitting before
starting its own submission.
- Call self._future.jobid() after FluxExecutor.submit() to block until
the job is actually registered with the Flux broker, not just queued
in the async FluxExecutor.
@coderabbitai

coderabbitaiBot commented Feb 21, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Adds ordered per-worker boot-up synchronization using threading.Event objects and updates _execute_multiple_tasks signature to accept bootup_event/next_bootup_event. Also calls self._future.jobid() immediately after submitting the jobspec in the Flux spawner bootup.

Changes

Cohort / File(s)Summary
Worker Boot-up Synchronization
src/executorlib/task_scheduler/interactive/blockallocation.py
Add per-worker bootup_events, pass bootup_event / next_bootup_event via executor kwargs, extend _execute_multiple_tasks signature to accept these events, wait on bootup_event before creating the worker interface, and set next_bootup_event after initialization (docstring updated).
Flux Spawner Bootup
src/executorlib/task_scheduler/interactive/spawner_flux.py
After submitting the jobspec during bootup, call self._future.jobid() conditionally when self._future is not None; no other control flow or public API changes.
Metadata
manifest_file, pyproject.toml
Minor packaging/manifest line edits (few-line changes).

Sequence Diagram(s)

sequenceDiagram
participant Scheduler as Scheduler (main)
participant WorkerA as Worker-0 thread
participant WorkerB as Worker-1 thread
participant FluxSpawner as FluxSpawner
participant Future as Future
Scheduler->>WorkerA: start thread with bootup_event A, next_bootup_event B
Scheduler->>WorkerB: start thread with bootup_event B, next_bootup_event C
WorkerA-->>Scheduler: bootup_event A.wait()
WorkerA->>WorkerA: create interface / submit tasks
WorkerA->>Scheduler: signal next_bootup_event B (set)
WorkerA->>FluxSpawner: submit jobspec
FluxSpawner->>Future: (if present) call jobid()
WorkerB-->>Scheduler: bootup_event B.wait()
WorkerB->>WorkerB: create interface / submit tasks
WorkerB->>Scheduler: signal next_bootup_event C (set)
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Poem

🐰 I lined the threads in tidy rows,

One hops up, then the next one goes.
A gentle nudge, a tiny ping,
The bootup choir begins to sing.
Hops in order — what a thing!

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 75.00% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title 'serialize worker job submissions to preserve worker_id ordering' accurately captures the main objective of the PR—implementing sequential worker submission to maintain reliable worker_id ordering in Flux scheduling.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

254-267: ⚠️ Potential issue | 🔴 Critical

Deadlock if interface_bootup raises: subsequent workers hang forever.

If interface_bootup (line 256) throws an exception for worker N, the thread terminates without calling next_bootup_event.set(). Workers N+1, N+2, … remain blocked on bootup_event.wait() indefinitely, and shutdown(wait=True) will hang on process.join().

Wrap the bootup + signal in try/finally so the chain always progresses:

🔒 Proposed fix using try/finally
 if bootup_event is not None:
bootup_event.wait()
- interface = interface_bootup(- command_lst=get_interactive_execute_command(- cores=cores,- ),- connections=spawner(cores=cores, worker_id=worker_id, **kwargs),- hostname_localhost=hostname_localhost,- log_obj_size=log_obj_size,- worker_id=worker_id,- stop_function=stop_function,- )- if next_bootup_event is not None:- next_bootup_event.set()+ try:+ interface = interface_bootup(+ command_lst=get_interactive_execute_command(+ cores=cores,+ ),+ connections=spawner(cores=cores, worker_id=worker_id, **kwargs),+ hostname_localhost=hostname_localhost,+ log_obj_size=log_obj_size,+ worker_id=worker_id,+ stop_function=stop_function,+ )+ finally:+ if next_bootup_event is not None:+ next_bootup_event.set()

Note: if bootup fails and the chain unblocks the next worker, subsequent workers will also likely fail. But that's preferable to a silent deadlock. You may also want to consider using bootup_event.wait(timeout=...) as additional protection.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
254 - 267, The chain of bootstrapped workers can deadlock if interface_bootup
raises because next_bootup_event.set() is skipped; modify the block around the
call to interface_bootup (the code that waits on bootup_event, calls
interface_bootup with get_interactive_execute_command and spawner, and then
calls next_bootup_event.set()) to ensure next_bootup_event.set() is executed in
a finally clause: keep the bootup_event.wait() before trying to create the
interface, perform the interface_bootup call (using
get_interactive_execute_command and spawner) inside a try, and always call
next_bootup_event.set() in the finally so the next worker is unblocked even if
interface_bootup raises; optionally re-raise or log the exception inside the
try/except around interface_bootup if you want error visibility.
🧹 Nitpick comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

116-126: Dynamic worker addition skips boot-up ordering (and worker_id).

The max_workers setter creates new threads using only self._process_kwargs, which doesn't include bootup_event, next_bootup_event, or worker_id. The None defaults prevent a crash, but dynamically added workers won't have ordered boot-up or correct worker_id assignment.

This is pre-existing behavior and not a regression from this PR, but worth noting if the setter is expected to be used in contexts where ordering matters.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
116 - 126, The dynamic worker-add path in the max_workers setter builds new
Thread instances using only self._process_kwargs which omits bootup_event,
next_bootup_event, and worker_id, so newly added threads skip boot-up ordering
and get wrong worker IDs; update the branch that creates new_process_lst (and
references _execute_multiple_tasks and self._process_kwargs) to construct
per-thread kwargs mirroring the initial worker creation: set worker_id =
self._max_workers + i for each new thread, pass the correct bootup_event and
next_bootup_event (chain them the same way the initial spawn logic does), and
then start and append those threads to self._process so dynamically added
workers follow the same boot ordering and ID assignment as existing workers.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 254-267: The chain of bootstrapped workers can deadlock if
interface_bootup raises because next_bootup_event.set() is skipped; modify the
block around the call to interface_bootup (the code that waits on bootup_event,
calls interface_bootup with get_interactive_execute_command and spawner, and
then calls next_bootup_event.set()) to ensure next_bootup_event.set() is
executed in a finally clause: keep the bootup_event.wait() before trying to
create the interface, perform the interface_bootup call (using
get_interactive_execute_command and spawner) inside a try, and always call
next_bootup_event.set() in the finally so the next worker is unblocked even if
interface_bootup raises; optionally re-raise or log the exception inside the
try/except around interface_bootup if you want error visibility.
---
Nitpick comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 116-126: The dynamic worker-add path in the max_workers setter
builds new Thread instances using only self._process_kwargs which omits
bootup_event, next_bootup_event, and worker_id, so newly added threads skip
boot-up ordering and get wrong worker IDs; update the branch that creates
new_process_lst (and references _execute_multiple_tasks and
self._process_kwargs) to construct per-thread kwargs mirroring the initial
worker creation: set worker_id = self._max_workers + i for each new thread, pass
the correct bootup_event and next_bootup_event (chain them the same way the
initial spawn logic does), and then start and append those threads to
self._process so dynamically added workers follow the same boot ordering and ID
assignment as existing workers.

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

256-269: ⚠️ Potential issue | 🔴 Critical

Deadlock if interface_bootup raises an exception.

If interface_bootup throws (e.g., spawner failure, connection timeout), next_bootup_event.set() on line 268 is never reached. Every subsequent worker will block forever on bootup_event.wait(), and shutdown(wait=True) will also hang because Thread.join() blocks on those stuck threads.

Wrap the boot-up + signal in try/finally so the chain always advances.

🔒 Proposed fix
 if bootup_event is not None:
bootup_event.wait()
- interface = interface_bootup(- command_lst=get_interactive_execute_command(- cores=cores,- ),- connections=spawner(cores=cores, worker_id=worker_id, **kwargs),- hostname_localhost=hostname_localhost,- log_obj_size=log_obj_size,- worker_id=worker_id,- stop_function=stop_function,- )- if next_bootup_event is not None:- next_bootup_event.set()+ try:+ interface = interface_bootup(+ command_lst=get_interactive_execute_command(+ cores=cores,+ ),+ connections=spawner(cores=cores, worker_id=worker_id, **kwargs),+ hostname_localhost=hostname_localhost,+ log_obj_size=log_obj_size,+ worker_id=worker_id,+ stop_function=stop_function,+ )+ finally:+ if next_bootup_event is not None:+ next_bootup_event.set()

Note: after the finally block, if interface_bootup did raise, the exception will propagate and terminate this worker thread. The rest of the function (init_function, task loop) will be skipped — but at least the remaining workers won't deadlock.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
256 - 269, The boot-up chain can deadlock if interface_bootup (called with
get_interactive_execute_command(...) and connections=spawner(...)) raises,
because next_bootup_event.set() is never reached; wrap the call to
interface_bootup and the subsequent signaling in a try/finally so that
next_bootup_event.set() is always executed in the finally block (still allowing
the exception to propagate and terminate this worker thread), ensuring
bootup_event.wait()/next_bootup_event.set() chain always advances.
🧹 Nitpick comments (1)
src/executorlib/task_scheduler/interactive/blockallocation.py (1)

256-257: bootup_event.wait() is not interruptible by shutdown().

If shutdown() is called while workers are still queued behind bootup_event.wait(), those threads cannot observe the stop signal — stop_function is only evaluated inside interface_bootup, not here. With the try/finally fix above, the chain will unblock naturally once the currently booting worker finishes (or fails), so this is mostly a latency concern rather than a hard deadlock. However, for a fast-exit guarantee you could add a timeout loop:

ifbootup_eventisnotNone:
whilenotbootup_event.wait(timeout=1.0):
ifstop_functionisnotNoneandstop_function():
return

This is optional and can be deferred.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/interactive/blockallocation.py` around lines
256 - 257, The current unconditional bootup_event.wait() call can block threads
from observing shutdown; change it to a timeout loop that repeatedly waits with
a short timeout and checks stop_function between waits so threads can return
early on shutdown. Replace the direct call to bootup_event.wait() (in
blockallocation where bootup_event is used alongside interface_bootup) with a
loop that calls bootup_event.wait(timeout=...) and after each timeout calls
stop_function() (if provided) and returns if it indicates shutdown; keep
behavior of eventually proceeding when bootup_event is set. Ensure you reference
bootup_event and stop_function and preserve existing cleanup/try/finally
semantics around interface_bootup.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Outside diff comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 256-269: The boot-up chain can deadlock if interface_bootup
(called with get_interactive_execute_command(...) and connections=spawner(...))
raises, because next_bootup_event.set() is never reached; wrap the call to
interface_bootup and the subsequent signaling in a try/finally so that
next_bootup_event.set() is always executed in the finally block (still allowing
the exception to propagate and terminate this worker thread), ensuring
bootup_event.wait()/next_bootup_event.set() chain always advances.
---
Nitpick comments:
In `@src/executorlib/task_scheduler/interactive/blockallocation.py`:
- Around line 256-257: The current unconditional bootup_event.wait() call can
block threads from observing shutdown; change it to a timeout loop that
repeatedly waits with a short timeout and checks stop_function between waits so
threads can return early on shutdown. Replace the direct call to
bootup_event.wait() (in blockallocation where bootup_event is used alongside
interface_bootup) with a loop that calls bootup_event.wait(timeout=...) and
after each timeout calls stop_function() (if provided) and returns if it
indicates shutdown; keep behavior of eventually proceeding when bootup_event is
set. Ensure you reference bootup_event and stop_function and preserve existing
cleanup/try/finally semantics around interface_bootup.

@codecov

codecovBot commented Feb 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.09%. Comparing base (2fa06c8) to head (e099b31).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #935 +/- ##
==========================================
+ Coverage 94.07% 94.09% +0.02% 
==========================================
Files 39 39 Lines 2041 2049 +8 ==========================================
+ Hits 1920 1928 +8 
Misses 121 121 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jan-janssenjan-janssen left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me

@jan-janssenjan-janssen changed the title fix: serialize worker job submissions to preserve worker_id ordering[Fix] serialize worker job submissions to preserve worker_id orderingFeb 22, 2026
@jan-janssen
jan-janssen merged commit 693ca9e into pyiron:mainFeb 22, 2026
59 of 62 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@IlgarBaghishov@jan-janssen