[Bug] Add error handling for submitting jobs using pysqa - #961

Merged
jan-janssen merged 3 commits into
mainfrom
pysqa_error_handling
Apr 15, 2026
Merged

[Bug] Add error handling for submitting jobs using pysqa#961
jan-janssen merged 3 commits into
mainfrom
pysqa_error_handling

Conversation

@jan-janssen

@jan-janssenjan-janssen commented Apr 10, 2026

Copy link
Copy Markdown
Member

Summary by CodeRabbit

  • Bug Fixes

    • Improved error handling for job submission failures, with exceptions now properly recorded in output files.
  • Tests

    • Added test coverage for queue name validation in cluster executor initialization.

CopilotAI review requested due to automatic review settings April 10, 2026 12:26
@coderabbitai

coderabbitaiBot commented Apr 10, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Error handling is added to the execute_with_pysqa function to catch job submission failures (ValueError and CalledProcessError), dump exceptions to HDF5, and rename intermediate cache files accordingly. Additionally, a new unit test validates that FluxClusterExecutor raises ValueError when initialized with an invalid queue name.

Changes

Cohort / File(s)Summary
Error Handling for Job Submission
src/executorlib/task_scheduler/file/spawner_pysqa.py
Wraps QueueAdapter.submit_job() in try/except/else block. On ValueError or CalledProcessError, dumps the exception to HDF5 and renames input cache file from *_i.h5 to *_o.h5. On success, dumps queue_id as before.
Queue Validation Test
tests/unit/executor/test_flux_cluster.py
Adds reusable Flux submission bash template and introduces test_executor_wrong_queue_name test case asserting that FluxClusterExecutor raises ValueError when constructed with invalid queue name "test".

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Poem

🐰 A job submission wrapped in care,
With try-except blocks placed with flair,
When errors hop, we catch and file,
Rename caches with garden style,
Queue names checked—no wrong turns here! 🌿

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 20.00% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title accurately summarizes the main change: adding error handling for pysqa job submission. The changeset shows execute_with_pysqa wrapping the job submission in try/except/else logic to handle ValueError and CalledProcessError, with a corresponding test validating error handling for invalid queue names.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch pysqa_error_handling

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds explicit error handling around pysqa job submission so that submission failures are captured and reflected in the cache files instead of leaving tasks in an indeterminate state.

Changes:

  • Catch ValueError / CalledProcessError from qa.submit_job(...) during task submission.
  • On submission failure, persist the error into the task HDF5 file and rename the input file to an output file to signal completion.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
@codecov

codecovBot commented Apr 10, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.15%. Comparing base (3c0dec7) to head (eaad8ae).
⚠️ Report is 4 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #961 +/- ##
==========================================
+ Coverage 94.14% 94.15% +0.01% 
==========================================
Files 39 39 Lines 2083 2089 +6 ==========================================
+ Hits 1961 1967 +6 
Misses 122 122 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jan-janssen
jan-janssen marked this pull request as ready for review April 10, 2026 13:49

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@src/executorlib/task_scheduler/file/spawner_pysqa.py`:
- Around line 96-102: The except block that catches (ValueError,
CalledProcessError) must clear any stale queue_id to avoid returning an old job
id; inside that except (in spawner_pysqa.py where queue_id is used after calling
qa.submit_job()), set queue_id = None (or another explicit sentinel) before
dumping the error and renaming files so the function returns a cleared value on
failure instead of a previous id.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: db7cdd8b-e2f6-4054-92d6-0e2797b048af

📥 Commits

Reviewing files that changed from the base of the PR and between 3c0dec7 and eaad8ae.

📒 Files selected for processing (2)
  • src/executorlib/task_scheduler/file/spawner_pysqa.py
  • tests/unit/executor/test_flux_cluster.py

Comment on lines +96 to 102
except (ValueError, CalledProcessError) as error:
dump(file_name=file_name, data_dict={"error": error})
file_name_out = os.path.splitext(file_name)[0][:-2]
os.rename(file_name_out + "_i.h5", file_name_out + "_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
return queue_id

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Reset queue_id on submit failure to avoid returning stale IDs.

If file_name already contained a previous queue_id, a failed qa.submit_job() keeps that old value and the function returns it on Line 102. This can propagate an invalid job id to downstream cleanup/termination flows.

💡 Proposed fix
 try:
queue_id = qa.submit_job(**submit_kwargs)
except (ValueError, CalledProcessError) as error:
+ queue_id = None
dump(file_name=file_name, data_dict={"error": error})
file_name_out = os.path.splitext(file_name)[0][:-2]
- os.rename(file_name_out + "_i.h5", file_name_out + "_o.h5")+ os.replace(file_name_out + "_i.h5", file_name_out + "_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
except (ValueError, CalledProcessError) aserror:
dump(file_name=file_name, data_dict={"error": error})
file_name_out=os.path.splitext(file_name)[0][:-2]
os.rename(file_name_out+"_i.h5", file_name_out+"_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
returnqueue_id
except (ValueError, CalledProcessError) aserror:
queue_id=None
dump(file_name=file_name, data_dict={"error": error})
file_name_out=os.path.splitext(file_name)[0][:-2]
os.replace(file_name_out+"_i.h5", file_name_out+"_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
returnqueue_id
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/file/spawner_pysqa.py` around lines 96 - 102,
The except block that catches (ValueError, CalledProcessError) must clear any
stale queue_id to avoid returning an old job id; inside that except (in
spawner_pysqa.py where queue_id is used after calling qa.submit_job()), set
queue_id = None (or another explicit sentinel) before dumping the error and
renaming files so the function returns a cleared value on failure instead of a
previous id.

@pmrvpmrv left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm as I wrote in #959 I actually think having the exception is the most general thing. I just didn't know before that the sbatch output is available on it.

@jan-janssen
jan-janssen merged commit df8a12a into mainApr 15, 2026
36 checks passed
@jan-janssen
jan-janssen deleted the pysqa_error_handling branch April 15, 2026 19:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Documentation] sbatch errors hang executor - explain CalledProcessError output for debugging

3 participants

@jan-janssen@pmrv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

[Bug] Add error handling for submitting jobs using pysqa - #961

Merged
jan-janssen merged 3 commits into
mainfrom
pysqa_error_handling
Apr 15, 2026
Merged

[Bug] Add error handling for submitting jobs using pysqa#961
jan-janssen merged 3 commits into
mainfrom
pysqa_error_handling

Conversation

@jan-janssen

@jan-janssenjan-janssen commented Apr 10, 2026

Copy link
Copy Markdown
Member

Summary by CodeRabbit

  • Bug Fixes

    • Improved error handling for job submission failures, with exceptions now properly recorded in output files.
  • Tests

    • Added test coverage for queue name validation in cluster executor initialization.

CopilotAI review requested due to automatic review settings April 10, 2026 12:26
@coderabbitai

coderabbitaiBot commented Apr 10, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Error handling is added to the execute_with_pysqa function to catch job submission failures (ValueError and CalledProcessError), dump exceptions to HDF5, and rename intermediate cache files accordingly. Additionally, a new unit test validates that FluxClusterExecutor raises ValueError when initialized with an invalid queue name.

Changes

Cohort / File(s)Summary
Error Handling for Job Submission
src/executorlib/task_scheduler/file/spawner_pysqa.py
Wraps QueueAdapter.submit_job() in try/except/else block. On ValueError or CalledProcessError, dumps the exception to HDF5 and renames input cache file from *_i.h5 to *_o.h5. On success, dumps queue_id as before.
Queue Validation Test
tests/unit/executor/test_flux_cluster.py
Adds reusable Flux submission bash template and introduces test_executor_wrong_queue_name test case asserting that FluxClusterExecutor raises ValueError when constructed with invalid queue name "test".

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Poem

🐰 A job submission wrapped in care,
With try-except blocks placed with flair,
When errors hop, we catch and file,
Rename caches with garden style,
Queue names checked—no wrong turns here! 🌿

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 20.00% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title accurately summarizes the main change: adding error handling for pysqa job submission. The changeset shows execute_with_pysqa wrapping the job submission in try/except/else logic to handle ValueError and CalledProcessError, with a corresponding test validating error handling for invalid queue names.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch pysqa_error_handling

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds explicit error handling around pysqa job submission so that submission failures are captured and reflected in the cache files instead of leaving tasks in an indeterminate state.

Changes:

  • Catch ValueError / CalledProcessError from qa.submit_job(...) during task submission.
  • On submission failure, persist the error into the task HDF5 file and rename the input file to an output file to signal completion.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
@codecov

codecovBot commented Apr 10, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.15%. Comparing base (3c0dec7) to head (eaad8ae).
⚠️ Report is 4 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #961 +/- ##
==========================================
+ Coverage 94.14% 94.15% +0.01% 
==========================================
Files 39 39 Lines 2083 2089 +6 ==========================================
+ Hits 1961 1967 +6 
Misses 122 122 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jan-janssen
jan-janssen marked this pull request as ready for review April 10, 2026 13:49

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@src/executorlib/task_scheduler/file/spawner_pysqa.py`:
- Around line 96-102: The except block that catches (ValueError,
CalledProcessError) must clear any stale queue_id to avoid returning an old job
id; inside that except (in spawner_pysqa.py where queue_id is used after calling
qa.submit_job()), set queue_id = None (or another explicit sentinel) before
dumping the error and renaming files so the function returns a cleared value on
failure instead of a previous id.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: db7cdd8b-e2f6-4054-92d6-0e2797b048af

📥 Commits

Reviewing files that changed from the base of the PR and between 3c0dec7 and eaad8ae.

📒 Files selected for processing (2)
  • src/executorlib/task_scheduler/file/spawner_pysqa.py
  • tests/unit/executor/test_flux_cluster.py

Comment on lines +96 to 102
except (ValueError, CalledProcessError) as error:
dump(file_name=file_name, data_dict={"error": error})
file_name_out = os.path.splitext(file_name)[0][:-2]
os.rename(file_name_out + "_i.h5", file_name_out + "_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
return queue_id

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Reset queue_id on submit failure to avoid returning stale IDs.

If file_name already contained a previous queue_id, a failed qa.submit_job() keeps that old value and the function returns it on Line 102. This can propagate an invalid job id to downstream cleanup/termination flows.

💡 Proposed fix
 try:
queue_id = qa.submit_job(**submit_kwargs)
except (ValueError, CalledProcessError) as error:
+ queue_id = None
dump(file_name=file_name, data_dict={"error": error})
file_name_out = os.path.splitext(file_name)[0][:-2]
- os.rename(file_name_out + "_i.h5", file_name_out + "_o.h5")+ os.replace(file_name_out + "_i.h5", file_name_out + "_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
except (ValueError, CalledProcessError) aserror:
dump(file_name=file_name, data_dict={"error": error})
file_name_out=os.path.splitext(file_name)[0][:-2]
os.rename(file_name_out+"_i.h5", file_name_out+"_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
returnqueue_id
except (ValueError, CalledProcessError) aserror:
queue_id=None
dump(file_name=file_name, data_dict={"error": error})
file_name_out=os.path.splitext(file_name)[0][:-2]
os.replace(file_name_out+"_i.h5", file_name_out+"_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
returnqueue_id
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/file/spawner_pysqa.py` around lines 96 - 102,
The except block that catches (ValueError, CalledProcessError) must clear any
stale queue_id to avoid returning an old job id; inside that except (in
spawner_pysqa.py where queue_id is used after calling qa.submit_job()), set
queue_id = None (or another explicit sentinel) before dumping the error and
renaming files so the function returns a cleared value on failure instead of a
previous id.

@pmrvpmrv left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm as I wrote in #959 I actually think having the exception is the most general thing. I just didn't know before that the sbatch output is available on it.

@jan-janssen
jan-janssen merged commit df8a12a into mainApr 15, 2026
36 checks passed
@jan-janssen
jan-janssen deleted the pysqa_error_handling branch April 15, 2026 19:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Documentation] sbatch errors hang executor - explain CalledProcessError output for debugging

3 participants

@jan-janssen@pmrv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Bug] Add error handling for submitting jobs using pysqa - #961

Merged
jan-janssen merged 3 commits into
mainfrom
pysqa_error_handling
Apr 15, 2026
Merged

[Bug] Add error handling for submitting jobs using pysqa#961
jan-janssen merged 3 commits into
mainfrom
pysqa_error_handling

Conversation

@jan-janssen

@jan-janssenjan-janssen commented Apr 10, 2026

Copy link
Copy Markdown
Member

Summary by CodeRabbit

  • Bug Fixes

    • Improved error handling for job submission failures, with exceptions now properly recorded in output files.
  • Tests

    • Added test coverage for queue name validation in cluster executor initialization.

CopilotAI review requested due to automatic review settings April 10, 2026 12:26
@coderabbitai

coderabbitaiBot commented Apr 10, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Error handling is added to the execute_with_pysqa function to catch job submission failures (ValueError and CalledProcessError), dump exceptions to HDF5, and rename intermediate cache files accordingly. Additionally, a new unit test validates that FluxClusterExecutor raises ValueError when initialized with an invalid queue name.

Changes

Cohort / File(s)Summary
Error Handling for Job Submission
src/executorlib/task_scheduler/file/spawner_pysqa.py
Wraps QueueAdapter.submit_job() in try/except/else block. On ValueError or CalledProcessError, dumps the exception to HDF5 and renames input cache file from *_i.h5 to *_o.h5. On success, dumps queue_id as before.
Queue Validation Test
tests/unit/executor/test_flux_cluster.py
Adds reusable Flux submission bash template and introduces test_executor_wrong_queue_name test case asserting that FluxClusterExecutor raises ValueError when constructed with invalid queue name "test".

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Poem

🐰 A job submission wrapped in care,
With try-except blocks placed with flair,
When errors hop, we catch and file,
Rename caches with garden style,
Queue names checked—no wrong turns here! 🌿

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 20.00% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title accurately summarizes the main change: adding error handling for pysqa job submission. The changeset shows execute_with_pysqa wrapping the job submission in try/except/else logic to handle ValueError and CalledProcessError, with a corresponding test validating error handling for invalid queue names.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch pysqa_error_handling

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds explicit error handling around pysqa job submission so that submission failures are captured and reflected in the cache files instead of leaving tasks in an indeterminate state.

Changes:

  • Catch ValueError / CalledProcessError from qa.submit_job(...) during task submission.
  • On submission failure, persist the error into the task HDF5 file and rename the input file to an output file to signal completion.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
@codecov

codecovBot commented Apr 10, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.15%. Comparing base (3c0dec7) to head (eaad8ae).
⚠️ Report is 4 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #961 +/- ##
==========================================
+ Coverage 94.14% 94.15% +0.01% 
==========================================
Files 39 39 Lines 2083 2089 +6 ==========================================
+ Hits 1961 1967 +6 
Misses 122 122 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jan-janssen
jan-janssen marked this pull request as ready for review April 10, 2026 13:49

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@src/executorlib/task_scheduler/file/spawner_pysqa.py`:
- Around line 96-102: The except block that catches (ValueError,
CalledProcessError) must clear any stale queue_id to avoid returning an old job
id; inside that except (in spawner_pysqa.py where queue_id is used after calling
qa.submit_job()), set queue_id = None (or another explicit sentinel) before
dumping the error and renaming files so the function returns a cleared value on
failure instead of a previous id.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: db7cdd8b-e2f6-4054-92d6-0e2797b048af

📥 Commits

Reviewing files that changed from the base of the PR and between 3c0dec7 and eaad8ae.

📒 Files selected for processing (2)
  • src/executorlib/task_scheduler/file/spawner_pysqa.py
  • tests/unit/executor/test_flux_cluster.py

Comment on lines +96 to 102
except (ValueError, CalledProcessError) as error:
dump(file_name=file_name, data_dict={"error": error})
file_name_out = os.path.splitext(file_name)[0][:-2]
os.rename(file_name_out + "_i.h5", file_name_out + "_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
return queue_id

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Reset queue_id on submit failure to avoid returning stale IDs.

If file_name already contained a previous queue_id, a failed qa.submit_job() keeps that old value and the function returns it on Line 102. This can propagate an invalid job id to downstream cleanup/termination flows.

💡 Proposed fix
 try:
queue_id = qa.submit_job(**submit_kwargs)
except (ValueError, CalledProcessError) as error:
+ queue_id = None
dump(file_name=file_name, data_dict={"error": error})
file_name_out = os.path.splitext(file_name)[0][:-2]
- os.rename(file_name_out + "_i.h5", file_name_out + "_o.h5")+ os.replace(file_name_out + "_i.h5", file_name_out + "_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
except (ValueError, CalledProcessError) aserror:
dump(file_name=file_name, data_dict={"error": error})
file_name_out=os.path.splitext(file_name)[0][:-2]
os.rename(file_name_out+"_i.h5", file_name_out+"_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
returnqueue_id
except (ValueError, CalledProcessError) aserror:
queue_id=None
dump(file_name=file_name, data_dict={"error": error})
file_name_out=os.path.splitext(file_name)[0][:-2]
os.replace(file_name_out+"_i.h5", file_name_out+"_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
returnqueue_id
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/file/spawner_pysqa.py` around lines 96 - 102,
The except block that catches (ValueError, CalledProcessError) must clear any
stale queue_id to avoid returning an old job id; inside that except (in
spawner_pysqa.py where queue_id is used after calling qa.submit_job()), set
queue_id = None (or another explicit sentinel) before dumping the error and
renaming files so the function returns a cleared value on failure instead of a
previous id.

@pmrvpmrv left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm as I wrote in #959 I actually think having the exception is the most general thing. I just didn't know before that the sbatch output is available on it.

@jan-janssen
jan-janssen merged commit df8a12a into mainApr 15, 2026
36 checks passed
@jan-janssen
jan-janssen deleted the pysqa_error_handling branch April 15, 2026 19:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Documentation] sbatch errors hang executor - explain CalledProcessError output for debugging

3 participants

@jan-janssen@pmrv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Bug] Add error handling for submitting jobs using pysqa - #961

Merged
jan-janssen merged 3 commits into
mainfrom
pysqa_error_handling
Apr 15, 2026
Merged

[Bug] Add error handling for submitting jobs using pysqa#961
jan-janssen merged 3 commits into
mainfrom
pysqa_error_handling

Conversation

@jan-janssen

@jan-janssenjan-janssen commented Apr 10, 2026

Copy link
Copy Markdown
Member

Summary by CodeRabbit

  • Bug Fixes

    • Improved error handling for job submission failures, with exceptions now properly recorded in output files.
  • Tests

    • Added test coverage for queue name validation in cluster executor initialization.

CopilotAI review requested due to automatic review settings April 10, 2026 12:26
@coderabbitai

coderabbitaiBot commented Apr 10, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Error handling is added to the execute_with_pysqa function to catch job submission failures (ValueError and CalledProcessError), dump exceptions to HDF5, and rename intermediate cache files accordingly. Additionally, a new unit test validates that FluxClusterExecutor raises ValueError when initialized with an invalid queue name.

Changes

Cohort / File(s)Summary
Error Handling for Job Submission
src/executorlib/task_scheduler/file/spawner_pysqa.py
Wraps QueueAdapter.submit_job() in try/except/else block. On ValueError or CalledProcessError, dumps the exception to HDF5 and renames input cache file from *_i.h5 to *_o.h5. On success, dumps queue_id as before.
Queue Validation Test
tests/unit/executor/test_flux_cluster.py
Adds reusable Flux submission bash template and introduces test_executor_wrong_queue_name test case asserting that FluxClusterExecutor raises ValueError when constructed with invalid queue name "test".

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Poem

🐰 A job submission wrapped in care,
With try-except blocks placed with flair,
When errors hop, we catch and file,
Rename caches with garden style,
Queue names checked—no wrong turns here! 🌿

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 20.00% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title accurately summarizes the main change: adding error handling for pysqa job submission. The changeset shows execute_with_pysqa wrapping the job submission in try/except/else logic to handle ValueError and CalledProcessError, with a corresponding test validating error handling for invalid queue names.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch pysqa_error_handling

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds explicit error handling around pysqa job submission so that submission failures are captured and reflected in the cache files instead of leaving tasks in an indeterminate state.

Changes:

  • Catch ValueError / CalledProcessError from qa.submit_job(...) during task submission.
  • On submission failure, persist the error into the task HDF5 file and rename the input file to an output file to signal completion.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
@codecov

codecovBot commented Apr 10, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.15%. Comparing base (3c0dec7) to head (eaad8ae).
⚠️ Report is 4 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #961 +/- ##
==========================================
+ Coverage 94.14% 94.15% +0.01% 
==========================================
Files 39 39 Lines 2083 2089 +6 ==========================================
+ Hits 1961 1967 +6 
Misses 122 122 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jan-janssen
jan-janssen marked this pull request as ready for review April 10, 2026 13:49

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@src/executorlib/task_scheduler/file/spawner_pysqa.py`:
- Around line 96-102: The except block that catches (ValueError,
CalledProcessError) must clear any stale queue_id to avoid returning an old job
id; inside that except (in spawner_pysqa.py where queue_id is used after calling
qa.submit_job()), set queue_id = None (or another explicit sentinel) before
dumping the error and renaming files so the function returns a cleared value on
failure instead of a previous id.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: db7cdd8b-e2f6-4054-92d6-0e2797b048af

📥 Commits

Reviewing files that changed from the base of the PR and between 3c0dec7 and eaad8ae.

📒 Files selected for processing (2)
  • src/executorlib/task_scheduler/file/spawner_pysqa.py
  • tests/unit/executor/test_flux_cluster.py

Comment on lines +96 to 102
except (ValueError, CalledProcessError) as error:
dump(file_name=file_name, data_dict={"error": error})
file_name_out = os.path.splitext(file_name)[0][:-2]
os.rename(file_name_out + "_i.h5", file_name_out + "_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
return queue_id

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Reset queue_id on submit failure to avoid returning stale IDs.

If file_name already contained a previous queue_id, a failed qa.submit_job() keeps that old value and the function returns it on Line 102. This can propagate an invalid job id to downstream cleanup/termination flows.

💡 Proposed fix
 try:
queue_id = qa.submit_job(**submit_kwargs)
except (ValueError, CalledProcessError) as error:
+ queue_id = None
dump(file_name=file_name, data_dict={"error": error})
file_name_out = os.path.splitext(file_name)[0][:-2]
- os.rename(file_name_out + "_i.h5", file_name_out + "_o.h5")+ os.replace(file_name_out + "_i.h5", file_name_out + "_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
except (ValueError, CalledProcessError) aserror:
dump(file_name=file_name, data_dict={"error": error})
file_name_out=os.path.splitext(file_name)[0][:-2]
os.rename(file_name_out+"_i.h5", file_name_out+"_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
returnqueue_id
except (ValueError, CalledProcessError) aserror:
queue_id=None
dump(file_name=file_name, data_dict={"error": error})
file_name_out=os.path.splitext(file_name)[0][:-2]
os.replace(file_name_out+"_i.h5", file_name_out+"_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
returnqueue_id
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/file/spawner_pysqa.py` around lines 96 - 102,
The except block that catches (ValueError, CalledProcessError) must clear any
stale queue_id to avoid returning an old job id; inside that except (in
spawner_pysqa.py where queue_id is used after calling qa.submit_job()), set
queue_id = None (or another explicit sentinel) before dumping the error and
renaming files so the function returns a cleared value on failure instead of a
previous id.

@pmrvpmrv left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm as I wrote in #959 I actually think having the exception is the most general thing. I just didn't know before that the sbatch output is available on it.

@jan-janssen
jan-janssen merged commit df8a12a into mainApr 15, 2026
36 checks passed
@jan-janssen
jan-janssen deleted the pysqa_error_handling branch April 15, 2026 19:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Documentation] sbatch errors hang executor - explain CalledProcessError output for debugging

3 participants

@jan-janssen@pmrv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

[Bug] Add error handling for submitting jobs using pysqa - #961

Merged
jan-janssen merged 3 commits into
mainfrom
pysqa_error_handling
Apr 15, 2026
Merged

[Bug] Add error handling for submitting jobs using pysqa#961
jan-janssen merged 3 commits into
mainfrom
pysqa_error_handling

Conversation

@jan-janssen

@jan-janssenjan-janssen commented Apr 10, 2026

Copy link
Copy Markdown
Member

Summary by CodeRabbit

  • Bug Fixes

    • Improved error handling for job submission failures, with exceptions now properly recorded in output files.
  • Tests

    • Added test coverage for queue name validation in cluster executor initialization.

CopilotAI review requested due to automatic review settings April 10, 2026 12:26
@coderabbitai

coderabbitaiBot commented Apr 10, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Error handling is added to the execute_with_pysqa function to catch job submission failures (ValueError and CalledProcessError), dump exceptions to HDF5, and rename intermediate cache files accordingly. Additionally, a new unit test validates that FluxClusterExecutor raises ValueError when initialized with an invalid queue name.

Changes

Cohort / File(s)Summary
Error Handling for Job Submission
src/executorlib/task_scheduler/file/spawner_pysqa.py
Wraps QueueAdapter.submit_job() in try/except/else block. On ValueError or CalledProcessError, dumps the exception to HDF5 and renames input cache file from *_i.h5 to *_o.h5. On success, dumps queue_id as before.
Queue Validation Test
tests/unit/executor/test_flux_cluster.py
Adds reusable Flux submission bash template and introduces test_executor_wrong_queue_name test case asserting that FluxClusterExecutor raises ValueError when constructed with invalid queue name "test".

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Poem

🐰 A job submission wrapped in care,
With try-except blocks placed with flair,
When errors hop, we catch and file,
Rename caches with garden style,
Queue names checked—no wrong turns here! 🌿

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 20.00% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title accurately summarizes the main change: adding error handling for pysqa job submission. The changeset shows execute_with_pysqa wrapping the job submission in try/except/else logic to handle ValueError and CalledProcessError, with a corresponding test validating error handling for invalid queue names.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch pysqa_error_handling

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds explicit error handling around pysqa job submission so that submission failures are captured and reflected in the cache files instead of leaving tasks in an indeterminate state.

Changes:

  • Catch ValueError / CalledProcessError from qa.submit_job(...) during task submission.
  • On submission failure, persist the error into the task HDF5 file and rename the input file to an output file to signal completion.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
@codecov

codecovBot commented Apr 10, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.15%. Comparing base (3c0dec7) to head (eaad8ae).
⚠️ Report is 4 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #961 +/- ##
==========================================
+ Coverage 94.14% 94.15% +0.01% 
==========================================
Files 39 39 Lines 2083 2089 +6 ==========================================
+ Hits 1961 1967 +6 
Misses 122 122 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jan-janssen
jan-janssen marked this pull request as ready for review April 10, 2026 13:49

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@src/executorlib/task_scheduler/file/spawner_pysqa.py`:
- Around line 96-102: The except block that catches (ValueError,
CalledProcessError) must clear any stale queue_id to avoid returning an old job
id; inside that except (in spawner_pysqa.py where queue_id is used after calling
qa.submit_job()), set queue_id = None (or another explicit sentinel) before
dumping the error and renaming files so the function returns a cleared value on
failure instead of a previous id.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: db7cdd8b-e2f6-4054-92d6-0e2797b048af

📥 Commits

Reviewing files that changed from the base of the PR and between 3c0dec7 and eaad8ae.

📒 Files selected for processing (2)
  • src/executorlib/task_scheduler/file/spawner_pysqa.py
  • tests/unit/executor/test_flux_cluster.py

Comment on lines +96 to 102
except (ValueError, CalledProcessError) as error:
dump(file_name=file_name, data_dict={"error": error})
file_name_out = os.path.splitext(file_name)[0][:-2]
os.rename(file_name_out + "_i.h5", file_name_out + "_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
return queue_id

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Reset queue_id on submit failure to avoid returning stale IDs.

If file_name already contained a previous queue_id, a failed qa.submit_job() keeps that old value and the function returns it on Line 102. This can propagate an invalid job id to downstream cleanup/termination flows.

💡 Proposed fix
 try:
queue_id = qa.submit_job(**submit_kwargs)
except (ValueError, CalledProcessError) as error:
+ queue_id = None
dump(file_name=file_name, data_dict={"error": error})
file_name_out = os.path.splitext(file_name)[0][:-2]
- os.rename(file_name_out + "_i.h5", file_name_out + "_o.h5")+ os.replace(file_name_out + "_i.h5", file_name_out + "_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
except (ValueError, CalledProcessError) aserror:
dump(file_name=file_name, data_dict={"error": error})
file_name_out=os.path.splitext(file_name)[0][:-2]
os.rename(file_name_out+"_i.h5", file_name_out+"_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
returnqueue_id
except (ValueError, CalledProcessError) aserror:
queue_id=None
dump(file_name=file_name, data_dict={"error": error})
file_name_out=os.path.splitext(file_name)[0][:-2]
os.replace(file_name_out+"_i.h5", file_name_out+"_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
returnqueue_id
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/file/spawner_pysqa.py` around lines 96 - 102,
The except block that catches (ValueError, CalledProcessError) must clear any
stale queue_id to avoid returning an old job id; inside that except (in
spawner_pysqa.py where queue_id is used after calling qa.submit_job()), set
queue_id = None (or another explicit sentinel) before dumping the error and
renaming files so the function returns a cleared value on failure instead of a
previous id.

@pmrvpmrv left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm as I wrote in #959 I actually think having the exception is the most general thing. I just didn't know before that the sbatch output is available on it.

@jan-janssen
jan-janssen merged commit df8a12a into mainApr 15, 2026
36 checks passed
@jan-janssen
jan-janssen deleted the pysqa_error_handling branch April 15, 2026 19:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Documentation] sbatch errors hang executor - explain CalledProcessError output for debugging

3 participants

@jan-janssen@pmrv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Bug] Add error handling for submitting jobs using pysqa - #961

Merged
jan-janssen merged 3 commits into
mainfrom
pysqa_error_handling
Apr 15, 2026
Merged

[Bug] Add error handling for submitting jobs using pysqa#961
jan-janssen merged 3 commits into
mainfrom
pysqa_error_handling

Conversation

@jan-janssen

@jan-janssenjan-janssen commented Apr 10, 2026

Copy link
Copy Markdown
Member

Summary by CodeRabbit

  • Bug Fixes

    • Improved error handling for job submission failures, with exceptions now properly recorded in output files.
  • Tests

    • Added test coverage for queue name validation in cluster executor initialization.

CopilotAI review requested due to automatic review settings April 10, 2026 12:26
@coderabbitai

coderabbitaiBot commented Apr 10, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Error handling is added to the execute_with_pysqa function to catch job submission failures (ValueError and CalledProcessError), dump exceptions to HDF5, and rename intermediate cache files accordingly. Additionally, a new unit test validates that FluxClusterExecutor raises ValueError when initialized with an invalid queue name.

Changes

Cohort / File(s)Summary
Error Handling for Job Submission
src/executorlib/task_scheduler/file/spawner_pysqa.py
Wraps QueueAdapter.submit_job() in try/except/else block. On ValueError or CalledProcessError, dumps the exception to HDF5 and renames input cache file from *_i.h5 to *_o.h5. On success, dumps queue_id as before.
Queue Validation Test
tests/unit/executor/test_flux_cluster.py
Adds reusable Flux submission bash template and introduces test_executor_wrong_queue_name test case asserting that FluxClusterExecutor raises ValueError when constructed with invalid queue name "test".

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Poem

🐰 A job submission wrapped in care,
With try-except blocks placed with flair,
When errors hop, we catch and file,
Rename caches with garden style,
Queue names checked—no wrong turns here! 🌿

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 20.00% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title accurately summarizes the main change: adding error handling for pysqa job submission. The changeset shows execute_with_pysqa wrapping the job submission in try/except/else logic to handle ValueError and CalledProcessError, with a corresponding test validating error handling for invalid queue names.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch pysqa_error_handling

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds explicit error handling around pysqa job submission so that submission failures are captured and reflected in the cache files instead of leaving tasks in an indeterminate state.

Changes:

  • Catch ValueError / CalledProcessError from qa.submit_job(...) during task submission.
  • On submission failure, persist the error into the task HDF5 file and rename the input file to an output file to signal completion.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
@codecov

codecovBot commented Apr 10, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.15%. Comparing base (3c0dec7) to head (eaad8ae).
⚠️ Report is 4 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #961 +/- ##
==========================================
+ Coverage 94.14% 94.15% +0.01% 
==========================================
Files 39 39 Lines 2083 2089 +6 ==========================================
+ Hits 1961 1967 +6 
Misses 122 122 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jan-janssen
jan-janssen marked this pull request as ready for review April 10, 2026 13:49

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@src/executorlib/task_scheduler/file/spawner_pysqa.py`:
- Around line 96-102: The except block that catches (ValueError,
CalledProcessError) must clear any stale queue_id to avoid returning an old job
id; inside that except (in spawner_pysqa.py where queue_id is used after calling
qa.submit_job()), set queue_id = None (or another explicit sentinel) before
dumping the error and renaming files so the function returns a cleared value on
failure instead of a previous id.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: db7cdd8b-e2f6-4054-92d6-0e2797b048af

📥 Commits

Reviewing files that changed from the base of the PR and between 3c0dec7 and eaad8ae.

📒 Files selected for processing (2)
  • src/executorlib/task_scheduler/file/spawner_pysqa.py
  • tests/unit/executor/test_flux_cluster.py

Comment on lines +96 to 102
except (ValueError, CalledProcessError) as error:
dump(file_name=file_name, data_dict={"error": error})
file_name_out = os.path.splitext(file_name)[0][:-2]
os.rename(file_name_out + "_i.h5", file_name_out + "_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
return queue_id

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Reset queue_id on submit failure to avoid returning stale IDs.

If file_name already contained a previous queue_id, a failed qa.submit_job() keeps that old value and the function returns it on Line 102. This can propagate an invalid job id to downstream cleanup/termination flows.

💡 Proposed fix
 try:
queue_id = qa.submit_job(**submit_kwargs)
except (ValueError, CalledProcessError) as error:
+ queue_id = None
dump(file_name=file_name, data_dict={"error": error})
file_name_out = os.path.splitext(file_name)[0][:-2]
- os.rename(file_name_out + "_i.h5", file_name_out + "_o.h5")+ os.replace(file_name_out + "_i.h5", file_name_out + "_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
except (ValueError, CalledProcessError) aserror:
dump(file_name=file_name, data_dict={"error": error})
file_name_out=os.path.splitext(file_name)[0][:-2]
os.rename(file_name_out+"_i.h5", file_name_out+"_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
returnqueue_id
except (ValueError, CalledProcessError) aserror:
queue_id=None
dump(file_name=file_name, data_dict={"error": error})
file_name_out=os.path.splitext(file_name)[0][:-2]
os.replace(file_name_out+"_i.h5", file_name_out+"_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
returnqueue_id
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/file/spawner_pysqa.py` around lines 96 - 102,
The except block that catches (ValueError, CalledProcessError) must clear any
stale queue_id to avoid returning an old job id; inside that except (in
spawner_pysqa.py where queue_id is used after calling qa.submit_job()), set
queue_id = None (or another explicit sentinel) before dumping the error and
renaming files so the function returns a cleared value on failure instead of a
previous id.

@pmrvpmrv left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm as I wrote in #959 I actually think having the exception is the most general thing. I just didn't know before that the sbatch output is available on it.

@jan-janssen
jan-janssen merged commit df8a12a into mainApr 15, 2026
36 checks passed
@jan-janssen
jan-janssen deleted the pysqa_error_handling branch April 15, 2026 19:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Documentation] sbatch errors hang executor - explain CalledProcessError output for debugging

3 participants

@jan-janssen@pmrv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

[Bug] Add error handling for submitting jobs using pysqa - #961

Merged
jan-janssen merged 3 commits into
mainfrom
pysqa_error_handling
Apr 15, 2026
Merged

[Bug] Add error handling for submitting jobs using pysqa#961
jan-janssen merged 3 commits into
mainfrom
pysqa_error_handling

Conversation

@jan-janssen

@jan-janssenjan-janssen commented Apr 10, 2026

Copy link
Copy Markdown
Member

Summary by CodeRabbit

  • Bug Fixes

    • Improved error handling for job submission failures, with exceptions now properly recorded in output files.
  • Tests

    • Added test coverage for queue name validation in cluster executor initialization.

CopilotAI review requested due to automatic review settings April 10, 2026 12:26
@coderabbitai

coderabbitaiBot commented Apr 10, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Error handling is added to the execute_with_pysqa function to catch job submission failures (ValueError and CalledProcessError), dump exceptions to HDF5, and rename intermediate cache files accordingly. Additionally, a new unit test validates that FluxClusterExecutor raises ValueError when initialized with an invalid queue name.

Changes

Cohort / File(s)Summary
Error Handling for Job Submission
src/executorlib/task_scheduler/file/spawner_pysqa.py
Wraps QueueAdapter.submit_job() in try/except/else block. On ValueError or CalledProcessError, dumps the exception to HDF5 and renames input cache file from *_i.h5 to *_o.h5. On success, dumps queue_id as before.
Queue Validation Test
tests/unit/executor/test_flux_cluster.py
Adds reusable Flux submission bash template and introduces test_executor_wrong_queue_name test case asserting that FluxClusterExecutor raises ValueError when constructed with invalid queue name "test".

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Poem

🐰 A job submission wrapped in care,
With try-except blocks placed with flair,
When errors hop, we catch and file,
Rename caches with garden style,
Queue names checked—no wrong turns here! 🌿

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 20.00% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title accurately summarizes the main change: adding error handling for pysqa job submission. The changeset shows execute_with_pysqa wrapping the job submission in try/except/else logic to handle ValueError and CalledProcessError, with a corresponding test validating error handling for invalid queue names.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch pysqa_error_handling

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds explicit error handling around pysqa job submission so that submission failures are captured and reflected in the cache files instead of leaving tasks in an indeterminate state.

Changes:

  • Catch ValueError / CalledProcessError from qa.submit_job(...) during task submission.
  • On submission failure, persist the error into the task HDF5 file and rename the input file to an output file to signal completion.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
@codecov

codecovBot commented Apr 10, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.15%. Comparing base (3c0dec7) to head (eaad8ae).
⚠️ Report is 4 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #961 +/- ##
==========================================
+ Coverage 94.14% 94.15% +0.01% 
==========================================
Files 39 39 Lines 2083 2089 +6 ==========================================
+ Hits 1961 1967 +6 
Misses 122 122 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jan-janssen
jan-janssen marked this pull request as ready for review April 10, 2026 13:49

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@src/executorlib/task_scheduler/file/spawner_pysqa.py`:
- Around line 96-102: The except block that catches (ValueError,
CalledProcessError) must clear any stale queue_id to avoid returning an old job
id; inside that except (in spawner_pysqa.py where queue_id is used after calling
qa.submit_job()), set queue_id = None (or another explicit sentinel) before
dumping the error and renaming files so the function returns a cleared value on
failure instead of a previous id.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: db7cdd8b-e2f6-4054-92d6-0e2797b048af

📥 Commits

Reviewing files that changed from the base of the PR and between 3c0dec7 and eaad8ae.

📒 Files selected for processing (2)
  • src/executorlib/task_scheduler/file/spawner_pysqa.py
  • tests/unit/executor/test_flux_cluster.py

Comment on lines +96 to 102
except (ValueError, CalledProcessError) as error:
dump(file_name=file_name, data_dict={"error": error})
file_name_out = os.path.splitext(file_name)[0][:-2]
os.rename(file_name_out + "_i.h5", file_name_out + "_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
return queue_id

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Reset queue_id on submit failure to avoid returning stale IDs.

If file_name already contained a previous queue_id, a failed qa.submit_job() keeps that old value and the function returns it on Line 102. This can propagate an invalid job id to downstream cleanup/termination flows.

💡 Proposed fix
 try:
queue_id = qa.submit_job(**submit_kwargs)
except (ValueError, CalledProcessError) as error:
+ queue_id = None
dump(file_name=file_name, data_dict={"error": error})
file_name_out = os.path.splitext(file_name)[0][:-2]
- os.rename(file_name_out + "_i.h5", file_name_out + "_o.h5")+ os.replace(file_name_out + "_i.h5", file_name_out + "_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
except (ValueError, CalledProcessError) aserror:
dump(file_name=file_name, data_dict={"error": error})
file_name_out=os.path.splitext(file_name)[0][:-2]
os.rename(file_name_out+"_i.h5", file_name_out+"_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
returnqueue_id
except (ValueError, CalledProcessError) aserror:
queue_id=None
dump(file_name=file_name, data_dict={"error": error})
file_name_out=os.path.splitext(file_name)[0][:-2]
os.replace(file_name_out+"_i.h5", file_name_out+"_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
returnqueue_id
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/file/spawner_pysqa.py` around lines 96 - 102,
The except block that catches (ValueError, CalledProcessError) must clear any
stale queue_id to avoid returning an old job id; inside that except (in
spawner_pysqa.py where queue_id is used after calling qa.submit_job()), set
queue_id = None (or another explicit sentinel) before dumping the error and
renaming files so the function returns a cleared value on failure instead of a
previous id.

@pmrvpmrv left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm as I wrote in #959 I actually think having the exception is the most general thing. I just didn't know before that the sbatch output is available on it.

@jan-janssen
jan-janssen merged commit df8a12a into mainApr 15, 2026
36 checks passed
@jan-janssen
jan-janssen deleted the pysqa_error_handling branch April 15, 2026 19:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Documentation] sbatch errors hang executor - explain CalledProcessError output for debugging

3 participants

@jan-janssen@pmrv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

[Bug] Add error handling for submitting jobs using pysqa - #961

Merged
jan-janssen merged 3 commits into
mainfrom
pysqa_error_handling
Apr 15, 2026
Merged

[Bug] Add error handling for submitting jobs using pysqa#961
jan-janssen merged 3 commits into
mainfrom
pysqa_error_handling

Conversation

@jan-janssen

@jan-janssenjan-janssen commented Apr 10, 2026

Copy link
Copy Markdown
Member

Summary by CodeRabbit

  • Bug Fixes

    • Improved error handling for job submission failures, with exceptions now properly recorded in output files.
  • Tests

    • Added test coverage for queue name validation in cluster executor initialization.

CopilotAI review requested due to automatic review settings April 10, 2026 12:26
@coderabbitai

coderabbitaiBot commented Apr 10, 2026

Copy link
Copy Markdown
Contributor
📝 Walkthrough

Walkthrough

Error handling is added to the execute_with_pysqa function to catch job submission failures (ValueError and CalledProcessError), dump exceptions to HDF5, and rename intermediate cache files accordingly. Additionally, a new unit test validates that FluxClusterExecutor raises ValueError when initialized with an invalid queue name.

Changes

Cohort / File(s)Summary
Error Handling for Job Submission
src/executorlib/task_scheduler/file/spawner_pysqa.py
Wraps QueueAdapter.submit_job() in try/except/else block. On ValueError or CalledProcessError, dumps the exception to HDF5 and renames input cache file from *_i.h5 to *_o.h5. On success, dumps queue_id as before.
Queue Validation Test
tests/unit/executor/test_flux_cluster.py
Adds reusable Flux submission bash template and introduces test_executor_wrong_queue_name test case asserting that FluxClusterExecutor raises ValueError when constructed with invalid queue name "test".

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Poem

🐰 A job submission wrapped in care,
With try-except blocks placed with flair,
When errors hop, we catch and file,
Rename caches with garden style,
Queue names checked—no wrong turns here! 🌿

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check nameStatusExplanationResolution
Docstring Coverage⚠️ WarningDocstring coverage is 20.00% which is insufficient. The required threshold is 80.00%.Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check nameStatusExplanation
Description Check✅ PassedCheck skipped - CodeRabbit’s high-level summary is enabled.
Title check✅ PassedThe title accurately summarizes the main change: adding error handling for pysqa job submission. The changeset shows execute_with_pysqa wrapping the job submission in try/except/else logic to handle ValueError and CalledProcessError, with a corresponding test validating error handling for invalid queue names.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch pysqa_error_handling

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

CopilotAI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds explicit error handling around pysqa job submission so that submission failures are captured and reflected in the cache files instead of leaving tasks in an indeterminate state.

Changes:

  • Catch ValueError / CalledProcessError from qa.submit_job(...) during task submission.
  • On submission failure, persist the error into the task HDF5 file and rename the input file to an output file to signal completion.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
Comment threadsrc/executorlib/task_scheduler/file/spawner_pysqa.py
@codecov

codecovBot commented Apr 10, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.15%. Comparing base (3c0dec7) to head (eaad8ae).
⚠️ Report is 4 commits behind head on main.

Additional details and impacted files
@@ Coverage Diff @@## main #961 +/- ##
==========================================
+ Coverage 94.14% 94.15% +0.01% 
==========================================
Files 39 39 Lines 2083 2089 +6 ==========================================
+ Hits 1961 1967 +6 
Misses 122 122 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@jan-janssen
jan-janssen marked this pull request as ready for review April 10, 2026 13:49

@coderabbitaicoderabbitaiBot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@src/executorlib/task_scheduler/file/spawner_pysqa.py`:
- Around line 96-102: The except block that catches (ValueError,
CalledProcessError) must clear any stale queue_id to avoid returning an old job
id; inside that except (in spawner_pysqa.py where queue_id is used after calling
qa.submit_job()), set queue_id = None (or another explicit sentinel) before
dumping the error and renaming files so the function returns a cleared value on
failure instead of a previous id.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: db7cdd8b-e2f6-4054-92d6-0e2797b048af

📥 Commits

Reviewing files that changed from the base of the PR and between 3c0dec7 and eaad8ae.

📒 Files selected for processing (2)
  • src/executorlib/task_scheduler/file/spawner_pysqa.py
  • tests/unit/executor/test_flux_cluster.py

Comment on lines +96 to 102
except (ValueError, CalledProcessError) as error:
dump(file_name=file_name, data_dict={"error": error})
file_name_out = os.path.splitext(file_name)[0][:-2]
os.rename(file_name_out + "_i.h5", file_name_out + "_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
return queue_id

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Reset queue_id on submit failure to avoid returning stale IDs.

If file_name already contained a previous queue_id, a failed qa.submit_job() keeps that old value and the function returns it on Line 102. This can propagate an invalid job id to downstream cleanup/termination flows.

💡 Proposed fix
 try:
queue_id = qa.submit_job(**submit_kwargs)
except (ValueError, CalledProcessError) as error:
+ queue_id = None
dump(file_name=file_name, data_dict={"error": error})
file_name_out = os.path.splitext(file_name)[0][:-2]
- os.rename(file_name_out + "_i.h5", file_name_out + "_o.h5")+ os.replace(file_name_out + "_i.h5", file_name_out + "_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
📝 Committable suggestion

‼️IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
except (ValueError, CalledProcessError) aserror:
dump(file_name=file_name, data_dict={"error": error})
file_name_out=os.path.splitext(file_name)[0][:-2]
os.rename(file_name_out+"_i.h5", file_name_out+"_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
returnqueue_id
except (ValueError, CalledProcessError) aserror:
queue_id=None
dump(file_name=file_name, data_dict={"error": error})
file_name_out=os.path.splitext(file_name)[0][:-2]
os.replace(file_name_out+"_i.h5", file_name_out+"_o.h5")
else:
dump(file_name=file_name, data_dict={"queue_id": queue_id})
returnqueue_id
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.
In `@src/executorlib/task_scheduler/file/spawner_pysqa.py` around lines 96 - 102,
The except block that catches (ValueError, CalledProcessError) must clear any
stale queue_id to avoid returning an old job id; inside that except (in
spawner_pysqa.py where queue_id is used after calling qa.submit_job()), set
queue_id = None (or another explicit sentinel) before dumping the error and
renaming files so the function returns a cleared value on failure instead of a
previous id.

@pmrvpmrv left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm as I wrote in #959 I actually think having the exception is the most general thing. I just didn't know before that the sbatch output is available on it.

@jan-janssen
jan-janssen merged commit df8a12a into mainApr 15, 2026
36 checks passed
@jan-janssen
jan-janssen deleted the pysqa_error_handling branch April 15, 2026 19:55
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Documentation] sbatch errors hang executor - explain CalledProcessError output for debugging

3 participants

@jan-janssen@pmrv