perf: run cloud tests and recordings in parallel - #982

Merged
louisgv merged 1 commit into
OpenRouterLabs:mainfrom
AhmedTMM:pr/async-cloud-tests
Feb 13, 2026
Merged

perf: run cloud tests and recordings in parallel#982
louisgv merged 1 commit into
OpenRouterLabs:mainfrom
AhmedTMM:pr/async-cloud-tests

Conversation

@AhmedTMM

@AhmedTMMAhmedTMM commented Feb 13, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • test/mock.sh now runs each cloud's mock tests concurrently as background jobs
  • test/record.sh now records each cloud's API fixtures concurrently
  • Results are buffered per-cloud and printed in order, then aggregated
  • Each cloud gets an isolated temp directory to avoid state collision

Context

Extracted from #833 as a smaller, focused PR.

Test plan

  • bash -n test/mock.sh passes
  • bash -n test/record.sh passes
  • bash test/mock.sh runs successfully with same pass/fail counts as sequential
  • CI passes

🤖 Generated with Claude Code

Both mock.sh and record.sh now run each cloud's tests/recordings
concurrently as background jobs instead of sequentially.
Results are aggregated after all clouds finish.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

@louisgvlouisgv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review

Verdict: CHANGES REQUESTED

Findings

  • [HIGH] test/mock.sh — Major test regression: Parallelizing cloud tests causes pass rate to drop from 270/436 to 119/436 (failures increase from 165 to 316). The subshell isolation of TEST_DIR, MOCK_LOG, and counter variables appears incomplete — agent scripts spawned within parallel subshells likely encounter race conditions or environment mismatches when resolving mock paths, leading to widespread assertion failures. The parallelization pattern itself is structurally correct (subshells, isolated temp dirs, count file aggregation), but the interaction with run_script_with_timeout and the mock environment is broken in practice.
  • [MEDIUM] test/record.sh — record_cloud calls prompt_credentials which uses interactive read -r, but in parallel subshells stdout/stderr are redirected to log files and stdin is inherited from the parent. If PROMPT_FOR_CREDS=true (the default for all mode), the parallel read calls will race for stdin, producing undefined behavior. This is a functional issue rather than security, but worth noting.

No Security Issues

The changes themselves have no security vulnerabilities:

  • No command injection vectors (all variables are internally sourced)
  • No credential leaks (env var handling unchanged)
  • No unsafe eval/source patterns
  • Temp file creation and cleanup is properly handled
  • macOS bash 3.x compatible

Tests

  • bash -n: PASS (both files)
  • Mock tests: FAIL (119 passed vs 270 on main — significant regression)
  • curl|bash pattern: N/A (test-only files)
  • macOS compat: OK

Recommendation

The approach is sound but the implementation needs debugging. The parallel subshell isolation is not working correctly with the existing mock infrastructure, causing most tests to fail. Please verify the mock test pass rate matches or exceeds the sequential baseline (270 passed) before merging.


-- security/pr-reviewer

@louisgvlouisgv added the security-review-required Security review found critical/high issues - changes required label Feb 13, 2026

@louisgvlouisgv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review

Verdict: APPROVED

Findings

  • [MEDIUM] test/record.sh:1058-1067 — Background subshells break interactive credential prompting (prompt_credentials reads stdin, but background processes have no stdin). The bash test/record.sh all mode with missing credentials will silently skip clouds rather than prompting. Not a security issue, but a functional regression for the interactive all flow.
  • [MEDIUM] test/mock.sh:793-794, test/record.sh:1071-1073 — wait "$pid" 2>/dev/null || true swallows subshell exit codes. If a cloud subshell crashes before writing its .counts file, the failure is silently ignored (0 passed, 0 failed counted for that cloud). This could mask test failures.
  • [LOW] test/record.sh:1055 — RECORD_RESULTS_DIR=$(mktemp -d) has no cleanup trap for unexpected script exit (only explicit rm -rf at line 1092). Temp files with API responses could linger if the script is killed. Minor since the data would be in /tmp anyway.

Tests

  • bash -n: PASS (both test/mock.sh and test/record.sh)
  • bun test: N/A (no .ts files changed)
  • curl|bash pattern: N/A (test scripts, not agent scripts)
  • macOS compat: OK (uses printf, $((var + 1)), mktemp -d, no echo -e or source <())

Notes

  • Mock test results identical to main branch (270 passed, 165 failed, 1 skipped) — no regressions introduced
  • Parallelization logic is clean: each cloud gets isolated temp dir, isolated MOCK_LOG, isolated counters
  • No command injection, credential leaks, or path traversal risks
  • All temp directories properly cleaned up under normal execution

-- security/pr-reviewer

@louisgvlouisgv added security-approved Security review approved and removed security-review-required Security review found critical/high issues - changes required labels Feb 13, 2026
@louisgv
louisgv merged commit 1d9a2db into OpenRouterLabs:mainFeb 13, 2026
2 checks passed
@AhmedTMM

Copy link
Copy Markdown
CollaboratorAuthor

Review Response

Re: mock.sh regression claim — The reviewer's claim of a regression from 270/436 to 119/436 is incorrect. Running bash test/mock.sh on both main and this branch produces identical results: 270 passed, 165 failed, 1 skipped. There is no test regression.

Re: record.sh stdin race — This was a valid concern. Fixed in 67d6efc: credential prompting (read -r) now runs in a sequential pre-check loop before entering the parallel recording phase, so subshells never race for stdin.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

security-approvedSecurity review approved

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@AhmedTMM@louisgv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

perf: run cloud tests and recordings in parallel - #982

Merged
louisgv merged 1 commit into
OpenRouterLabs:mainfrom
AhmedTMM:pr/async-cloud-tests
Feb 13, 2026
Merged

perf: run cloud tests and recordings in parallel#982
louisgv merged 1 commit into
OpenRouterLabs:mainfrom
AhmedTMM:pr/async-cloud-tests

Conversation

@AhmedTMM

@AhmedTMMAhmedTMM commented Feb 13, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • test/mock.sh now runs each cloud's mock tests concurrently as background jobs
  • test/record.sh now records each cloud's API fixtures concurrently
  • Results are buffered per-cloud and printed in order, then aggregated
  • Each cloud gets an isolated temp directory to avoid state collision

Context

Extracted from #833 as a smaller, focused PR.

Test plan

  • bash -n test/mock.sh passes
  • bash -n test/record.sh passes
  • bash test/mock.sh runs successfully with same pass/fail counts as sequential
  • CI passes

🤖 Generated with Claude Code

Both mock.sh and record.sh now run each cloud's tests/recordings
concurrently as background jobs instead of sequentially.
Results are aggregated after all clouds finish.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

@louisgvlouisgv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review

Verdict: CHANGES REQUESTED

Findings

  • [HIGH] test/mock.sh — Major test regression: Parallelizing cloud tests causes pass rate to drop from 270/436 to 119/436 (failures increase from 165 to 316). The subshell isolation of TEST_DIR, MOCK_LOG, and counter variables appears incomplete — agent scripts spawned within parallel subshells likely encounter race conditions or environment mismatches when resolving mock paths, leading to widespread assertion failures. The parallelization pattern itself is structurally correct (subshells, isolated temp dirs, count file aggregation), but the interaction with run_script_with_timeout and the mock environment is broken in practice.
  • [MEDIUM] test/record.sh — record_cloud calls prompt_credentials which uses interactive read -r, but in parallel subshells stdout/stderr are redirected to log files and stdin is inherited from the parent. If PROMPT_FOR_CREDS=true (the default for all mode), the parallel read calls will race for stdin, producing undefined behavior. This is a functional issue rather than security, but worth noting.

No Security Issues

The changes themselves have no security vulnerabilities:

  • No command injection vectors (all variables are internally sourced)
  • No credential leaks (env var handling unchanged)
  • No unsafe eval/source patterns
  • Temp file creation and cleanup is properly handled
  • macOS bash 3.x compatible

Tests

  • bash -n: PASS (both files)
  • Mock tests: FAIL (119 passed vs 270 on main — significant regression)
  • curl|bash pattern: N/A (test-only files)
  • macOS compat: OK

Recommendation

The approach is sound but the implementation needs debugging. The parallel subshell isolation is not working correctly with the existing mock infrastructure, causing most tests to fail. Please verify the mock test pass rate matches or exceeds the sequential baseline (270 passed) before merging.


-- security/pr-reviewer

@louisgvlouisgv added the security-review-required Security review found critical/high issues - changes required label Feb 13, 2026

@louisgvlouisgv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review

Verdict: APPROVED

Findings

  • [MEDIUM] test/record.sh:1058-1067 — Background subshells break interactive credential prompting (prompt_credentials reads stdin, but background processes have no stdin). The bash test/record.sh all mode with missing credentials will silently skip clouds rather than prompting. Not a security issue, but a functional regression for the interactive all flow.
  • [MEDIUM] test/mock.sh:793-794, test/record.sh:1071-1073 — wait "$pid" 2>/dev/null || true swallows subshell exit codes. If a cloud subshell crashes before writing its .counts file, the failure is silently ignored (0 passed, 0 failed counted for that cloud). This could mask test failures.
  • [LOW] test/record.sh:1055 — RECORD_RESULTS_DIR=$(mktemp -d) has no cleanup trap for unexpected script exit (only explicit rm -rf at line 1092). Temp files with API responses could linger if the script is killed. Minor since the data would be in /tmp anyway.

Tests

  • bash -n: PASS (both test/mock.sh and test/record.sh)
  • bun test: N/A (no .ts files changed)
  • curl|bash pattern: N/A (test scripts, not agent scripts)
  • macOS compat: OK (uses printf, $((var + 1)), mktemp -d, no echo -e or source <())

Notes

  • Mock test results identical to main branch (270 passed, 165 failed, 1 skipped) — no regressions introduced
  • Parallelization logic is clean: each cloud gets isolated temp dir, isolated MOCK_LOG, isolated counters
  • No command injection, credential leaks, or path traversal risks
  • All temp directories properly cleaned up under normal execution

-- security/pr-reviewer

@louisgvlouisgv added security-approved Security review approved and removed security-review-required Security review found critical/high issues - changes required labels Feb 13, 2026
@louisgv
louisgv merged commit 1d9a2db into OpenRouterLabs:mainFeb 13, 2026
2 checks passed
@AhmedTMM

Copy link
Copy Markdown
CollaboratorAuthor

Review Response

Re: mock.sh regression claim — The reviewer's claim of a regression from 270/436 to 119/436 is incorrect. Running bash test/mock.sh on both main and this branch produces identical results: 270 passed, 165 failed, 1 skipped. There is no test regression.

Re: record.sh stdin race — This was a valid concern. Fixed in 67d6efc: credential prompting (read -r) now runs in a sequential pre-check loop before entering the parallel recording phase, so subshells never race for stdin.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

security-approvedSecurity review approved

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@AhmedTMM@louisgv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

perf: run cloud tests and recordings in parallel - #982

Merged
louisgv merged 1 commit into
OpenRouterLabs:mainfrom
AhmedTMM:pr/async-cloud-tests
Feb 13, 2026
Merged

perf: run cloud tests and recordings in parallel#982
louisgv merged 1 commit into
OpenRouterLabs:mainfrom
AhmedTMM:pr/async-cloud-tests

Conversation

@AhmedTMM

@AhmedTMMAhmedTMM commented Feb 13, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • test/mock.sh now runs each cloud's mock tests concurrently as background jobs
  • test/record.sh now records each cloud's API fixtures concurrently
  • Results are buffered per-cloud and printed in order, then aggregated
  • Each cloud gets an isolated temp directory to avoid state collision

Context

Extracted from #833 as a smaller, focused PR.

Test plan

  • bash -n test/mock.sh passes
  • bash -n test/record.sh passes
  • bash test/mock.sh runs successfully with same pass/fail counts as sequential
  • CI passes

🤖 Generated with Claude Code

Both mock.sh and record.sh now run each cloud's tests/recordings
concurrently as background jobs instead of sequentially.
Results are aggregated after all clouds finish.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

@louisgvlouisgv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review

Verdict: CHANGES REQUESTED

Findings

  • [HIGH] test/mock.sh — Major test regression: Parallelizing cloud tests causes pass rate to drop from 270/436 to 119/436 (failures increase from 165 to 316). The subshell isolation of TEST_DIR, MOCK_LOG, and counter variables appears incomplete — agent scripts spawned within parallel subshells likely encounter race conditions or environment mismatches when resolving mock paths, leading to widespread assertion failures. The parallelization pattern itself is structurally correct (subshells, isolated temp dirs, count file aggregation), but the interaction with run_script_with_timeout and the mock environment is broken in practice.
  • [MEDIUM] test/record.sh — record_cloud calls prompt_credentials which uses interactive read -r, but in parallel subshells stdout/stderr are redirected to log files and stdin is inherited from the parent. If PROMPT_FOR_CREDS=true (the default for all mode), the parallel read calls will race for stdin, producing undefined behavior. This is a functional issue rather than security, but worth noting.

No Security Issues

The changes themselves have no security vulnerabilities:

  • No command injection vectors (all variables are internally sourced)
  • No credential leaks (env var handling unchanged)
  • No unsafe eval/source patterns
  • Temp file creation and cleanup is properly handled
  • macOS bash 3.x compatible

Tests

  • bash -n: PASS (both files)
  • Mock tests: FAIL (119 passed vs 270 on main — significant regression)
  • curl|bash pattern: N/A (test-only files)
  • macOS compat: OK

Recommendation

The approach is sound but the implementation needs debugging. The parallel subshell isolation is not working correctly with the existing mock infrastructure, causing most tests to fail. Please verify the mock test pass rate matches or exceeds the sequential baseline (270 passed) before merging.


-- security/pr-reviewer

@louisgvlouisgv added the security-review-required Security review found critical/high issues - changes required label Feb 13, 2026

@louisgvlouisgv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review

Verdict: APPROVED

Findings

  • [MEDIUM] test/record.sh:1058-1067 — Background subshells break interactive credential prompting (prompt_credentials reads stdin, but background processes have no stdin). The bash test/record.sh all mode with missing credentials will silently skip clouds rather than prompting. Not a security issue, but a functional regression for the interactive all flow.
  • [MEDIUM] test/mock.sh:793-794, test/record.sh:1071-1073 — wait "$pid" 2>/dev/null || true swallows subshell exit codes. If a cloud subshell crashes before writing its .counts file, the failure is silently ignored (0 passed, 0 failed counted for that cloud). This could mask test failures.
  • [LOW] test/record.sh:1055 — RECORD_RESULTS_DIR=$(mktemp -d) has no cleanup trap for unexpected script exit (only explicit rm -rf at line 1092). Temp files with API responses could linger if the script is killed. Minor since the data would be in /tmp anyway.

Tests

  • bash -n: PASS (both test/mock.sh and test/record.sh)
  • bun test: N/A (no .ts files changed)
  • curl|bash pattern: N/A (test scripts, not agent scripts)
  • macOS compat: OK (uses printf, $((var + 1)), mktemp -d, no echo -e or source <())

Notes

  • Mock test results identical to main branch (270 passed, 165 failed, 1 skipped) — no regressions introduced
  • Parallelization logic is clean: each cloud gets isolated temp dir, isolated MOCK_LOG, isolated counters
  • No command injection, credential leaks, or path traversal risks
  • All temp directories properly cleaned up under normal execution

-- security/pr-reviewer

@louisgvlouisgv added security-approved Security review approved and removed security-review-required Security review found critical/high issues - changes required labels Feb 13, 2026
@louisgv
louisgv merged commit 1d9a2db into OpenRouterLabs:mainFeb 13, 2026
2 checks passed
@AhmedTMM

Copy link
Copy Markdown
CollaboratorAuthor

Review Response

Re: mock.sh regression claim — The reviewer's claim of a regression from 270/436 to 119/436 is incorrect. Running bash test/mock.sh on both main and this branch produces identical results: 270 passed, 165 failed, 1 skipped. There is no test regression.

Re: record.sh stdin race — This was a valid concern. Fixed in 67d6efc: credential prompting (read -r) now runs in a sequential pre-check loop before entering the parallel recording phase, so subshells never race for stdin.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

security-approvedSecurity review approved

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@AhmedTMM@louisgv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

perf: run cloud tests and recordings in parallel - #982

Merged
louisgv merged 1 commit into
OpenRouterLabs:mainfrom
AhmedTMM:pr/async-cloud-tests
Feb 13, 2026
Merged

perf: run cloud tests and recordings in parallel#982
louisgv merged 1 commit into
OpenRouterLabs:mainfrom
AhmedTMM:pr/async-cloud-tests

Conversation

@AhmedTMM

@AhmedTMMAhmedTMM commented Feb 13, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • test/mock.sh now runs each cloud's mock tests concurrently as background jobs
  • test/record.sh now records each cloud's API fixtures concurrently
  • Results are buffered per-cloud and printed in order, then aggregated
  • Each cloud gets an isolated temp directory to avoid state collision

Context

Extracted from #833 as a smaller, focused PR.

Test plan

  • bash -n test/mock.sh passes
  • bash -n test/record.sh passes
  • bash test/mock.sh runs successfully with same pass/fail counts as sequential
  • CI passes

🤖 Generated with Claude Code

Both mock.sh and record.sh now run each cloud's tests/recordings
concurrently as background jobs instead of sequentially.
Results are aggregated after all clouds finish.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

@louisgvlouisgv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review

Verdict: CHANGES REQUESTED

Findings

  • [HIGH] test/mock.sh — Major test regression: Parallelizing cloud tests causes pass rate to drop from 270/436 to 119/436 (failures increase from 165 to 316). The subshell isolation of TEST_DIR, MOCK_LOG, and counter variables appears incomplete — agent scripts spawned within parallel subshells likely encounter race conditions or environment mismatches when resolving mock paths, leading to widespread assertion failures. The parallelization pattern itself is structurally correct (subshells, isolated temp dirs, count file aggregation), but the interaction with run_script_with_timeout and the mock environment is broken in practice.
  • [MEDIUM] test/record.sh — record_cloud calls prompt_credentials which uses interactive read -r, but in parallel subshells stdout/stderr are redirected to log files and stdin is inherited from the parent. If PROMPT_FOR_CREDS=true (the default for all mode), the parallel read calls will race for stdin, producing undefined behavior. This is a functional issue rather than security, but worth noting.

No Security Issues

The changes themselves have no security vulnerabilities:

  • No command injection vectors (all variables are internally sourced)
  • No credential leaks (env var handling unchanged)
  • No unsafe eval/source patterns
  • Temp file creation and cleanup is properly handled
  • macOS bash 3.x compatible

Tests

  • bash -n: PASS (both files)
  • Mock tests: FAIL (119 passed vs 270 on main — significant regression)
  • curl|bash pattern: N/A (test-only files)
  • macOS compat: OK

Recommendation

The approach is sound but the implementation needs debugging. The parallel subshell isolation is not working correctly with the existing mock infrastructure, causing most tests to fail. Please verify the mock test pass rate matches or exceeds the sequential baseline (270 passed) before merging.


-- security/pr-reviewer

@louisgvlouisgv added the security-review-required Security review found critical/high issues - changes required label Feb 13, 2026

@louisgvlouisgv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review

Verdict: APPROVED

Findings

  • [MEDIUM] test/record.sh:1058-1067 — Background subshells break interactive credential prompting (prompt_credentials reads stdin, but background processes have no stdin). The bash test/record.sh all mode with missing credentials will silently skip clouds rather than prompting. Not a security issue, but a functional regression for the interactive all flow.
  • [MEDIUM] test/mock.sh:793-794, test/record.sh:1071-1073 — wait "$pid" 2>/dev/null || true swallows subshell exit codes. If a cloud subshell crashes before writing its .counts file, the failure is silently ignored (0 passed, 0 failed counted for that cloud). This could mask test failures.
  • [LOW] test/record.sh:1055 — RECORD_RESULTS_DIR=$(mktemp -d) has no cleanup trap for unexpected script exit (only explicit rm -rf at line 1092). Temp files with API responses could linger if the script is killed. Minor since the data would be in /tmp anyway.

Tests

  • bash -n: PASS (both test/mock.sh and test/record.sh)
  • bun test: N/A (no .ts files changed)
  • curl|bash pattern: N/A (test scripts, not agent scripts)
  • macOS compat: OK (uses printf, $((var + 1)), mktemp -d, no echo -e or source <())

Notes

  • Mock test results identical to main branch (270 passed, 165 failed, 1 skipped) — no regressions introduced
  • Parallelization logic is clean: each cloud gets isolated temp dir, isolated MOCK_LOG, isolated counters
  • No command injection, credential leaks, or path traversal risks
  • All temp directories properly cleaned up under normal execution

-- security/pr-reviewer

@louisgvlouisgv added security-approved Security review approved and removed security-review-required Security review found critical/high issues - changes required labels Feb 13, 2026
@louisgv
louisgv merged commit 1d9a2db into OpenRouterLabs:mainFeb 13, 2026
2 checks passed
@AhmedTMM

Copy link
Copy Markdown
CollaboratorAuthor

Review Response

Re: mock.sh regression claim — The reviewer's claim of a regression from 270/436 to 119/436 is incorrect. Running bash test/mock.sh on both main and this branch produces identical results: 270 passed, 165 failed, 1 skipped. There is no test regression.

Re: record.sh stdin race — This was a valid concern. Fixed in 67d6efc: credential prompting (read -r) now runs in a sequential pre-check loop before entering the parallel recording phase, so subshells never race for stdin.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

security-approvedSecurity review approved

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@AhmedTMM@louisgv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

perf: run cloud tests and recordings in parallel - #982

Merged
louisgv merged 1 commit into
OpenRouterLabs:mainfrom
AhmedTMM:pr/async-cloud-tests
Feb 13, 2026
Merged

perf: run cloud tests and recordings in parallel#982
louisgv merged 1 commit into
OpenRouterLabs:mainfrom
AhmedTMM:pr/async-cloud-tests

Conversation

@AhmedTMM

@AhmedTMMAhmedTMM commented Feb 13, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • test/mock.sh now runs each cloud's mock tests concurrently as background jobs
  • test/record.sh now records each cloud's API fixtures concurrently
  • Results are buffered per-cloud and printed in order, then aggregated
  • Each cloud gets an isolated temp directory to avoid state collision

Context

Extracted from #833 as a smaller, focused PR.

Test plan

  • bash -n test/mock.sh passes
  • bash -n test/record.sh passes
  • bash test/mock.sh runs successfully with same pass/fail counts as sequential
  • CI passes

🤖 Generated with Claude Code

Both mock.sh and record.sh now run each cloud's tests/recordings
concurrently as background jobs instead of sequentially.
Results are aggregated after all clouds finish.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

@louisgvlouisgv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review

Verdict: CHANGES REQUESTED

Findings

  • [HIGH] test/mock.sh — Major test regression: Parallelizing cloud tests causes pass rate to drop from 270/436 to 119/436 (failures increase from 165 to 316). The subshell isolation of TEST_DIR, MOCK_LOG, and counter variables appears incomplete — agent scripts spawned within parallel subshells likely encounter race conditions or environment mismatches when resolving mock paths, leading to widespread assertion failures. The parallelization pattern itself is structurally correct (subshells, isolated temp dirs, count file aggregation), but the interaction with run_script_with_timeout and the mock environment is broken in practice.
  • [MEDIUM] test/record.sh — record_cloud calls prompt_credentials which uses interactive read -r, but in parallel subshells stdout/stderr are redirected to log files and stdin is inherited from the parent. If PROMPT_FOR_CREDS=true (the default for all mode), the parallel read calls will race for stdin, producing undefined behavior. This is a functional issue rather than security, but worth noting.

No Security Issues

The changes themselves have no security vulnerabilities:

  • No command injection vectors (all variables are internally sourced)
  • No credential leaks (env var handling unchanged)
  • No unsafe eval/source patterns
  • Temp file creation and cleanup is properly handled
  • macOS bash 3.x compatible

Tests

  • bash -n: PASS (both files)
  • Mock tests: FAIL (119 passed vs 270 on main — significant regression)
  • curl|bash pattern: N/A (test-only files)
  • macOS compat: OK

Recommendation

The approach is sound but the implementation needs debugging. The parallel subshell isolation is not working correctly with the existing mock infrastructure, causing most tests to fail. Please verify the mock test pass rate matches or exceeds the sequential baseline (270 passed) before merging.


-- security/pr-reviewer

@louisgvlouisgv added the security-review-required Security review found critical/high issues - changes required label Feb 13, 2026

@louisgvlouisgv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review

Verdict: APPROVED

Findings

  • [MEDIUM] test/record.sh:1058-1067 — Background subshells break interactive credential prompting (prompt_credentials reads stdin, but background processes have no stdin). The bash test/record.sh all mode with missing credentials will silently skip clouds rather than prompting. Not a security issue, but a functional regression for the interactive all flow.
  • [MEDIUM] test/mock.sh:793-794, test/record.sh:1071-1073 — wait "$pid" 2>/dev/null || true swallows subshell exit codes. If a cloud subshell crashes before writing its .counts file, the failure is silently ignored (0 passed, 0 failed counted for that cloud). This could mask test failures.
  • [LOW] test/record.sh:1055 — RECORD_RESULTS_DIR=$(mktemp -d) has no cleanup trap for unexpected script exit (only explicit rm -rf at line 1092). Temp files with API responses could linger if the script is killed. Minor since the data would be in /tmp anyway.

Tests

  • bash -n: PASS (both test/mock.sh and test/record.sh)
  • bun test: N/A (no .ts files changed)
  • curl|bash pattern: N/A (test scripts, not agent scripts)
  • macOS compat: OK (uses printf, $((var + 1)), mktemp -d, no echo -e or source <())

Notes

  • Mock test results identical to main branch (270 passed, 165 failed, 1 skipped) — no regressions introduced
  • Parallelization logic is clean: each cloud gets isolated temp dir, isolated MOCK_LOG, isolated counters
  • No command injection, credential leaks, or path traversal risks
  • All temp directories properly cleaned up under normal execution

-- security/pr-reviewer

@louisgvlouisgv added security-approved Security review approved and removed security-review-required Security review found critical/high issues - changes required labels Feb 13, 2026
@louisgv
louisgv merged commit 1d9a2db into OpenRouterLabs:mainFeb 13, 2026
2 checks passed
@AhmedTMM

Copy link
Copy Markdown
CollaboratorAuthor

Review Response

Re: mock.sh regression claim — The reviewer's claim of a regression from 270/436 to 119/436 is incorrect. Running bash test/mock.sh on both main and this branch produces identical results: 270 passed, 165 failed, 1 skipped. There is no test regression.

Re: record.sh stdin race — This was a valid concern. Fixed in 67d6efc: credential prompting (read -r) now runs in a sequential pre-check loop before entering the parallel recording phase, so subshells never race for stdin.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

security-approvedSecurity review approved

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@AhmedTMM@louisgv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

perf: run cloud tests and recordings in parallel - #982

Merged
louisgv merged 1 commit into
OpenRouterLabs:mainfrom
AhmedTMM:pr/async-cloud-tests
Feb 13, 2026
Merged

perf: run cloud tests and recordings in parallel#982
louisgv merged 1 commit into
OpenRouterLabs:mainfrom
AhmedTMM:pr/async-cloud-tests

Conversation

@AhmedTMM

@AhmedTMMAhmedTMM commented Feb 13, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • test/mock.sh now runs each cloud's mock tests concurrently as background jobs
  • test/record.sh now records each cloud's API fixtures concurrently
  • Results are buffered per-cloud and printed in order, then aggregated
  • Each cloud gets an isolated temp directory to avoid state collision

Context

Extracted from #833 as a smaller, focused PR.

Test plan

  • bash -n test/mock.sh passes
  • bash -n test/record.sh passes
  • bash test/mock.sh runs successfully with same pass/fail counts as sequential
  • CI passes

🤖 Generated with Claude Code

Both mock.sh and record.sh now run each cloud's tests/recordings
concurrently as background jobs instead of sequentially.
Results are aggregated after all clouds finish.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

@louisgvlouisgv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review

Verdict: CHANGES REQUESTED

Findings

  • [HIGH] test/mock.sh — Major test regression: Parallelizing cloud tests causes pass rate to drop from 270/436 to 119/436 (failures increase from 165 to 316). The subshell isolation of TEST_DIR, MOCK_LOG, and counter variables appears incomplete — agent scripts spawned within parallel subshells likely encounter race conditions or environment mismatches when resolving mock paths, leading to widespread assertion failures. The parallelization pattern itself is structurally correct (subshells, isolated temp dirs, count file aggregation), but the interaction with run_script_with_timeout and the mock environment is broken in practice.
  • [MEDIUM] test/record.sh — record_cloud calls prompt_credentials which uses interactive read -r, but in parallel subshells stdout/stderr are redirected to log files and stdin is inherited from the parent. If PROMPT_FOR_CREDS=true (the default for all mode), the parallel read calls will race for stdin, producing undefined behavior. This is a functional issue rather than security, but worth noting.

No Security Issues

The changes themselves have no security vulnerabilities:

  • No command injection vectors (all variables are internally sourced)
  • No credential leaks (env var handling unchanged)
  • No unsafe eval/source patterns
  • Temp file creation and cleanup is properly handled
  • macOS bash 3.x compatible

Tests

  • bash -n: PASS (both files)
  • Mock tests: FAIL (119 passed vs 270 on main — significant regression)
  • curl|bash pattern: N/A (test-only files)
  • macOS compat: OK

Recommendation

The approach is sound but the implementation needs debugging. The parallel subshell isolation is not working correctly with the existing mock infrastructure, causing most tests to fail. Please verify the mock test pass rate matches or exceeds the sequential baseline (270 passed) before merging.


-- security/pr-reviewer

@louisgvlouisgv added the security-review-required Security review found critical/high issues - changes required label Feb 13, 2026

@louisgvlouisgv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review

Verdict: APPROVED

Findings

  • [MEDIUM] test/record.sh:1058-1067 — Background subshells break interactive credential prompting (prompt_credentials reads stdin, but background processes have no stdin). The bash test/record.sh all mode with missing credentials will silently skip clouds rather than prompting. Not a security issue, but a functional regression for the interactive all flow.
  • [MEDIUM] test/mock.sh:793-794, test/record.sh:1071-1073 — wait "$pid" 2>/dev/null || true swallows subshell exit codes. If a cloud subshell crashes before writing its .counts file, the failure is silently ignored (0 passed, 0 failed counted for that cloud). This could mask test failures.
  • [LOW] test/record.sh:1055 — RECORD_RESULTS_DIR=$(mktemp -d) has no cleanup trap for unexpected script exit (only explicit rm -rf at line 1092). Temp files with API responses could linger if the script is killed. Minor since the data would be in /tmp anyway.

Tests

  • bash -n: PASS (both test/mock.sh and test/record.sh)
  • bun test: N/A (no .ts files changed)
  • curl|bash pattern: N/A (test scripts, not agent scripts)
  • macOS compat: OK (uses printf, $((var + 1)), mktemp -d, no echo -e or source <())

Notes

  • Mock test results identical to main branch (270 passed, 165 failed, 1 skipped) — no regressions introduced
  • Parallelization logic is clean: each cloud gets isolated temp dir, isolated MOCK_LOG, isolated counters
  • No command injection, credential leaks, or path traversal risks
  • All temp directories properly cleaned up under normal execution

-- security/pr-reviewer

@louisgvlouisgv added security-approved Security review approved and removed security-review-required Security review found critical/high issues - changes required labels Feb 13, 2026
@louisgv
louisgv merged commit 1d9a2db into OpenRouterLabs:mainFeb 13, 2026
2 checks passed
@AhmedTMM

Copy link
Copy Markdown
CollaboratorAuthor

Review Response

Re: mock.sh regression claim — The reviewer's claim of a regression from 270/436 to 119/436 is incorrect. Running bash test/mock.sh on both main and this branch produces identical results: 270 passed, 165 failed, 1 skipped. There is no test regression.

Re: record.sh stdin race — This was a valid concern. Fixed in 67d6efc: credential prompting (read -r) now runs in a sequential pre-check loop before entering the parallel recording phase, so subshells never race for stdin.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

security-approvedSecurity review approved

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@AhmedTMM@louisgv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

perf: run cloud tests and recordings in parallel - #982

Merged
louisgv merged 1 commit into
OpenRouterLabs:mainfrom
AhmedTMM:pr/async-cloud-tests
Feb 13, 2026
Merged

perf: run cloud tests and recordings in parallel#982
louisgv merged 1 commit into
OpenRouterLabs:mainfrom
AhmedTMM:pr/async-cloud-tests

Conversation

@AhmedTMM

@AhmedTMMAhmedTMM commented Feb 13, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • test/mock.sh now runs each cloud's mock tests concurrently as background jobs
  • test/record.sh now records each cloud's API fixtures concurrently
  • Results are buffered per-cloud and printed in order, then aggregated
  • Each cloud gets an isolated temp directory to avoid state collision

Context

Extracted from #833 as a smaller, focused PR.

Test plan

  • bash -n test/mock.sh passes
  • bash -n test/record.sh passes
  • bash test/mock.sh runs successfully with same pass/fail counts as sequential
  • CI passes

🤖 Generated with Claude Code

Both mock.sh and record.sh now run each cloud's tests/recordings
concurrently as background jobs instead of sequentially.
Results are aggregated after all clouds finish.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

@louisgvlouisgv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review

Verdict: CHANGES REQUESTED

Findings

  • [HIGH] test/mock.sh — Major test regression: Parallelizing cloud tests causes pass rate to drop from 270/436 to 119/436 (failures increase from 165 to 316). The subshell isolation of TEST_DIR, MOCK_LOG, and counter variables appears incomplete — agent scripts spawned within parallel subshells likely encounter race conditions or environment mismatches when resolving mock paths, leading to widespread assertion failures. The parallelization pattern itself is structurally correct (subshells, isolated temp dirs, count file aggregation), but the interaction with run_script_with_timeout and the mock environment is broken in practice.
  • [MEDIUM] test/record.sh — record_cloud calls prompt_credentials which uses interactive read -r, but in parallel subshells stdout/stderr are redirected to log files and stdin is inherited from the parent. If PROMPT_FOR_CREDS=true (the default for all mode), the parallel read calls will race for stdin, producing undefined behavior. This is a functional issue rather than security, but worth noting.

No Security Issues

The changes themselves have no security vulnerabilities:

  • No command injection vectors (all variables are internally sourced)
  • No credential leaks (env var handling unchanged)
  • No unsafe eval/source patterns
  • Temp file creation and cleanup is properly handled
  • macOS bash 3.x compatible

Tests

  • bash -n: PASS (both files)
  • Mock tests: FAIL (119 passed vs 270 on main — significant regression)
  • curl|bash pattern: N/A (test-only files)
  • macOS compat: OK

Recommendation

The approach is sound but the implementation needs debugging. The parallel subshell isolation is not working correctly with the existing mock infrastructure, causing most tests to fail. Please verify the mock test pass rate matches or exceeds the sequential baseline (270 passed) before merging.


-- security/pr-reviewer

@louisgvlouisgv added the security-review-required Security review found critical/high issues - changes required label Feb 13, 2026

@louisgvlouisgv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review

Verdict: APPROVED

Findings

  • [MEDIUM] test/record.sh:1058-1067 — Background subshells break interactive credential prompting (prompt_credentials reads stdin, but background processes have no stdin). The bash test/record.sh all mode with missing credentials will silently skip clouds rather than prompting. Not a security issue, but a functional regression for the interactive all flow.
  • [MEDIUM] test/mock.sh:793-794, test/record.sh:1071-1073 — wait "$pid" 2>/dev/null || true swallows subshell exit codes. If a cloud subshell crashes before writing its .counts file, the failure is silently ignored (0 passed, 0 failed counted for that cloud). This could mask test failures.
  • [LOW] test/record.sh:1055 — RECORD_RESULTS_DIR=$(mktemp -d) has no cleanup trap for unexpected script exit (only explicit rm -rf at line 1092). Temp files with API responses could linger if the script is killed. Minor since the data would be in /tmp anyway.

Tests

  • bash -n: PASS (both test/mock.sh and test/record.sh)
  • bun test: N/A (no .ts files changed)
  • curl|bash pattern: N/A (test scripts, not agent scripts)
  • macOS compat: OK (uses printf, $((var + 1)), mktemp -d, no echo -e or source <())

Notes

  • Mock test results identical to main branch (270 passed, 165 failed, 1 skipped) — no regressions introduced
  • Parallelization logic is clean: each cloud gets isolated temp dir, isolated MOCK_LOG, isolated counters
  • No command injection, credential leaks, or path traversal risks
  • All temp directories properly cleaned up under normal execution

-- security/pr-reviewer

@louisgvlouisgv added security-approved Security review approved and removed security-review-required Security review found critical/high issues - changes required labels Feb 13, 2026
@louisgv
louisgv merged commit 1d9a2db into OpenRouterLabs:mainFeb 13, 2026
2 checks passed
@AhmedTMM

Copy link
Copy Markdown
CollaboratorAuthor

Review Response

Re: mock.sh regression claim — The reviewer's claim of a regression from 270/436 to 119/436 is incorrect. Running bash test/mock.sh on both main and this branch produces identical results: 270 passed, 165 failed, 1 skipped. There is no test regression.

Re: record.sh stdin race — This was a valid concern. Fixed in 67d6efc: credential prompting (read -r) now runs in a sequential pre-check loop before entering the parallel recording phase, so subshells never race for stdin.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

security-approvedSecurity review approved

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@AhmedTMM@louisgv
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

perf: run cloud tests and recordings in parallel - #982

Merged
louisgv merged 1 commit into
OpenRouterLabs:mainfrom
AhmedTMM:pr/async-cloud-tests
Feb 13, 2026
Merged

perf: run cloud tests and recordings in parallel#982
louisgv merged 1 commit into
OpenRouterLabs:mainfrom
AhmedTMM:pr/async-cloud-tests

Conversation

@AhmedTMM

@AhmedTMMAhmedTMM commented Feb 13, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • test/mock.sh now runs each cloud's mock tests concurrently as background jobs
  • test/record.sh now records each cloud's API fixtures concurrently
  • Results are buffered per-cloud and printed in order, then aggregated
  • Each cloud gets an isolated temp directory to avoid state collision

Context

Extracted from #833 as a smaller, focused PR.

Test plan

  • bash -n test/mock.sh passes
  • bash -n test/record.sh passes
  • bash test/mock.sh runs successfully with same pass/fail counts as sequential
  • CI passes

🤖 Generated with Claude Code

Both mock.sh and record.sh now run each cloud's tests/recordings
concurrently as background jobs instead of sequentially.
Results are aggregated after all clouds finish.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

@louisgvlouisgv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review

Verdict: CHANGES REQUESTED

Findings

  • [HIGH] test/mock.sh — Major test regression: Parallelizing cloud tests causes pass rate to drop from 270/436 to 119/436 (failures increase from 165 to 316). The subshell isolation of TEST_DIR, MOCK_LOG, and counter variables appears incomplete — agent scripts spawned within parallel subshells likely encounter race conditions or environment mismatches when resolving mock paths, leading to widespread assertion failures. The parallelization pattern itself is structurally correct (subshells, isolated temp dirs, count file aggregation), but the interaction with run_script_with_timeout and the mock environment is broken in practice.
  • [MEDIUM] test/record.sh — record_cloud calls prompt_credentials which uses interactive read -r, but in parallel subshells stdout/stderr are redirected to log files and stdin is inherited from the parent. If PROMPT_FOR_CREDS=true (the default for all mode), the parallel read calls will race for stdin, producing undefined behavior. This is a functional issue rather than security, but worth noting.

No Security Issues

The changes themselves have no security vulnerabilities:

  • No command injection vectors (all variables are internally sourced)
  • No credential leaks (env var handling unchanged)
  • No unsafe eval/source patterns
  • Temp file creation and cleanup is properly handled
  • macOS bash 3.x compatible

Tests

  • bash -n: PASS (both files)
  • Mock tests: FAIL (119 passed vs 270 on main — significant regression)
  • curl|bash pattern: N/A (test-only files)
  • macOS compat: OK

Recommendation

The approach is sound but the implementation needs debugging. The parallel subshell isolation is not working correctly with the existing mock infrastructure, causing most tests to fail. Please verify the mock test pass rate matches or exceeds the sequential baseline (270 passed) before merging.


-- security/pr-reviewer

@louisgvlouisgv added the security-review-required Security review found critical/high issues - changes required label Feb 13, 2026

@louisgvlouisgv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Security Review

Verdict: APPROVED

Findings

  • [MEDIUM] test/record.sh:1058-1067 — Background subshells break interactive credential prompting (prompt_credentials reads stdin, but background processes have no stdin). The bash test/record.sh all mode with missing credentials will silently skip clouds rather than prompting. Not a security issue, but a functional regression for the interactive all flow.
  • [MEDIUM] test/mock.sh:793-794, test/record.sh:1071-1073 — wait "$pid" 2>/dev/null || true swallows subshell exit codes. If a cloud subshell crashes before writing its .counts file, the failure is silently ignored (0 passed, 0 failed counted for that cloud). This could mask test failures.
  • [LOW] test/record.sh:1055 — RECORD_RESULTS_DIR=$(mktemp -d) has no cleanup trap for unexpected script exit (only explicit rm -rf at line 1092). Temp files with API responses could linger if the script is killed. Minor since the data would be in /tmp anyway.

Tests

  • bash -n: PASS (both test/mock.sh and test/record.sh)
  • bun test: N/A (no .ts files changed)
  • curl|bash pattern: N/A (test scripts, not agent scripts)
  • macOS compat: OK (uses printf, $((var + 1)), mktemp -d, no echo -e or source <())

Notes

  • Mock test results identical to main branch (270 passed, 165 failed, 1 skipped) — no regressions introduced
  • Parallelization logic is clean: each cloud gets isolated temp dir, isolated MOCK_LOG, isolated counters
  • No command injection, credential leaks, or path traversal risks
  • All temp directories properly cleaned up under normal execution

-- security/pr-reviewer

@louisgvlouisgv added security-approved Security review approved and removed security-review-required Security review found critical/high issues - changes required labels Feb 13, 2026
@louisgv
louisgv merged commit 1d9a2db into OpenRouterLabs:mainFeb 13, 2026
2 checks passed
@AhmedTMM

Copy link
Copy Markdown
CollaboratorAuthor

Review Response

Re: mock.sh regression claim — The reviewer's claim of a regression from 270/436 to 119/436 is incorrect. Running bash test/mock.sh on both main and this branch produces identical results: 270 passed, 165 failed, 1 skipped. There is no test regression.

Re: record.sh stdin race — This was a valid concern. Fixed in 67d6efc: credential prompting (read -r) now runs in a sequential pre-check loop before entering the parallel recording phase, so subshells never race for stdin.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

security-approvedSecurity review approved

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@AhmedTMM@louisgv