feat(eval): batch simulate — --ingestion-wait-ms + per-example failures/sessions - #2098

Merged
jariy17 merged 7 commits into
refactorfrom
feat/eval-simulate-followups
Aug 27, 2026
Merged

feat(eval): batch simulate — --ingestion-wait-ms + per-example failures/sessions#2098
jariy17 merged 7 commits into
refactorfrom
feat/eval-simulate-followups

Conversation

@jariy17

@jariy17jariy17 commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

What

Batch simulate review follow-ups (base: refactor):

  1. --ingestion-wait-ms flag (default 180000, 0 skips) → InvokeDatasetInput.waitIngestionMs. Drops the SIMULATE_INGESTION_WAIT_MS env var — value flows through the input only.
  2. Per-example failures.runExamples returns failures: {item,error}[]; invokeDataset surfaces failures: [{exampleId, error}] — names which examples dropped + why, not just a count.
  3. Output now renders sessions[] + failures[] (below).
  4. Uses Golden Test Pattern for simulate instead of dedicated unit tests for invokeDataset

Output

{
"batchEvaluationId": "batch-eval-test", "status": "RUNNING",
"examplesInvoked": 1, "examplesFailed": 1,
"sessions": [ { "exampleId": "ok1", "sessionId": "s1" } ],
"failures": [ { "exampleId": "bad", "error": "HTTP 500" } ]
}
  • sessions[]exampleId ↔ sessionId. Join key: a later eval batch-evaluation get returns results[] keyed by sessionId, mapping each score back to its dataset row.
  • failures[] — always present ([] on a clean run); names which examples dropped + why.

…ailures/sessions
Follow-ups from the batch-evaluation simulate review:
- Add --ingestion-wait-ms (default 180000, 0 skips); thread via InvokeDatasetInput.waitIngestionMs.
Removes the SIMULATE_INGESTION_WAIT_MS env var — tests pass the value through the input.
- runExamples now returns per-item failures (item + error), not a bare count + firstError; invokeDataset
surfaces failures: [{ exampleId, error }] so a partial failure names which examples dropped and why.
- batch simulate output renders sessions[] (exampleId <-> sessionId join key for a later get) and
failures[] (omitted when empty).
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added agentcore-harness-reviewing AgentCore Harness review in progress claude-security-reviewing Claude Code /security-review in progress and removed agentcore-harness-reviewing AgentCore Harness review in progress labels Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@codecov-commenter

codecov-commenter commented Aug 25, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.85714% with 3 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.23%. Comparing base (343d213) to head (ebc604e).
⚠️ Report is 12 commits behind head on refactor.

Files with missing linesPatch %Lines
src/core/eval.tsx86.36%3 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #2098 +/- ##
============================================
- Coverage 97.33% 97.23% -0.11% 
============================================
Files 417 417 Lines 25250 25269 +19 ============================================
- Hits 24578 24570 -8 - Misses 672 699 +27 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 26, 2026
- inject newSessionId into EvalClient (default randomUUID) so replay fixtures + goldens are deterministic
- teach makeRecordingSend to freeze/revive a streaming SDK response (InvokeAgentRuntime), which stringify couldn't serialize
- add simulate fixture-golden case; move handler edges to batch-evaluation.test.tsx
- split invokeDataset.test.ts into run.test.ts (pool) + load.test.ts (parse + GT-shape); delete it and simulate.test.tsx
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/m PR size: M labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@jariy17
jariy17 marked this pull request as ready for review August 26, 2026 22:34
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot added claude-security-reviewing Claude Code /security-review in progress and removed claude-security-reviewing Claude Code /security-review in progress labels Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
);
if (r.invoked === 0) {
const detail = r.firstError ? `; first error: ${r.firstError.message}` : "";
const first = r.failures[0];

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense to just join all the errors? Or are these already communicated to the customer somewhere else?

Just wanted to make sure they'd be aware of individual datasets failing even if the overall command doesn't fail.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes all the errors would show in the debug logs. I assumed that if all dataset examples have failed, it would be the same error. However, maybe we should convene these errors directly to the tui or cli. I'll look into this in a follow up


const STREAM_TAG = "$stream";

async function freezeStream(response: unknown): Promise<unknown> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

took me a sec, but this makes sense and feels like a simple solution.

correct me if I misunderstood, but for streaming apis the response is not json serializable, so we convert the response via the stream tag. In the process, we resolve the whole stream, so we need to convert the data back to a stream before passing it downstream to consumers who are expecting a stream.

nice!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep ur correct. The reviveStream name seems a little weird to me, maybe Ill use mimicStream.

@nborges-awsnborges-aws left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PR overall LGTM. One concern with a gap in testing. Tests are all either exercise happy path, or mock invokeDataset with a mocked failure object in the response. Nowhere is exercising actual invocation failures assemble that failure object correctly.

The deleted invokeDataset test is actually the only place that exercised this.

@notgitikanotgitika left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Left a comment, could be followup

if (typeof stream?.transformToString !== "function") return response;
return {
...(response as Record<string, unknown>),
response: { [STREAM_TAG]: await stream.transformToString() },

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we serialize transformToByteArray() as base64 instead? transformToString() decodes as UTF-8, so binary or invalid UTF-8 runtime responses are changed during replay.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can do in a follow up.

@jariy17
jariy17 merged commit b78d396 into refactorAug 27, 2026
26 checks passed
@jariy17
jariy17 deleted the feat/eval-simulate-followups branch August 27, 2026 16:57
@jariy17

Copy link
Copy Markdown
ContributorAuthor

I'll address the missing bad paths in a follow up

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@jariy17@codecov-commenter@Hweinstock@notgitika@nborges-aws
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(eval): batch simulate — --ingestion-wait-ms + per-example failures/sessions - #2098

Merged
jariy17 merged 7 commits into
refactorfrom
feat/eval-simulate-followups
Aug 27, 2026
Merged

feat(eval): batch simulate — --ingestion-wait-ms + per-example failures/sessions#2098
jariy17 merged 7 commits into
refactorfrom
feat/eval-simulate-followups

Conversation

@jariy17

@jariy17jariy17 commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

What

Batch simulate review follow-ups (base: refactor):

  1. --ingestion-wait-ms flag (default 180000, 0 skips) → InvokeDatasetInput.waitIngestionMs. Drops the SIMULATE_INGESTION_WAIT_MS env var — value flows through the input only.
  2. Per-example failures.runExamples returns failures: {item,error}[]; invokeDataset surfaces failures: [{exampleId, error}] — names which examples dropped + why, not just a count.
  3. Output now renders sessions[] + failures[] (below).
  4. Uses Golden Test Pattern for simulate instead of dedicated unit tests for invokeDataset

Output

{
"batchEvaluationId": "batch-eval-test", "status": "RUNNING",
"examplesInvoked": 1, "examplesFailed": 1,
"sessions": [ { "exampleId": "ok1", "sessionId": "s1" } ],
"failures": [ { "exampleId": "bad", "error": "HTTP 500" } ]
}
  • sessions[]exampleId ↔ sessionId. Join key: a later eval batch-evaluation get returns results[] keyed by sessionId, mapping each score back to its dataset row.
  • failures[] — always present ([] on a clean run); names which examples dropped + why.

…ailures/sessions
Follow-ups from the batch-evaluation simulate review:
- Add --ingestion-wait-ms (default 180000, 0 skips); thread via InvokeDatasetInput.waitIngestionMs.
Removes the SIMULATE_INGESTION_WAIT_MS env var — tests pass the value through the input.
- runExamples now returns per-item failures (item + error), not a bare count + firstError; invokeDataset
surfaces failures: [{ exampleId, error }] so a partial failure names which examples dropped and why.
- batch simulate output renders sessions[] (exampleId <-> sessionId join key for a later get) and
failures[] (omitted when empty).
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added agentcore-harness-reviewing AgentCore Harness review in progress claude-security-reviewing Claude Code /security-review in progress and removed agentcore-harness-reviewing AgentCore Harness review in progress labels Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@codecov-commenter

codecov-commenter commented Aug 25, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.85714% with 3 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.23%. Comparing base (343d213) to head (ebc604e).
⚠️ Report is 12 commits behind head on refactor.

Files with missing linesPatch %Lines
src/core/eval.tsx86.36%3 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #2098 +/- ##
============================================
- Coverage 97.33% 97.23% -0.11% 
============================================
Files 417 417 Lines 25250 25269 +19 ============================================
- Hits 24578 24570 -8 - Misses 672 699 +27 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 26, 2026
- inject newSessionId into EvalClient (default randomUUID) so replay fixtures + goldens are deterministic
- teach makeRecordingSend to freeze/revive a streaming SDK response (InvokeAgentRuntime), which stringify couldn't serialize
- add simulate fixture-golden case; move handler edges to batch-evaluation.test.tsx
- split invokeDataset.test.ts into run.test.ts (pool) + load.test.ts (parse + GT-shape); delete it and simulate.test.tsx
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/m PR size: M labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@jariy17
jariy17 marked this pull request as ready for review August 26, 2026 22:34
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot added claude-security-reviewing Claude Code /security-review in progress and removed claude-security-reviewing Claude Code /security-review in progress labels Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
);
if (r.invoked === 0) {
const detail = r.firstError ? `; first error: ${r.firstError.message}` : "";
const first = r.failures[0];

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense to just join all the errors? Or are these already communicated to the customer somewhere else?

Just wanted to make sure they'd be aware of individual datasets failing even if the overall command doesn't fail.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes all the errors would show in the debug logs. I assumed that if all dataset examples have failed, it would be the same error. However, maybe we should convene these errors directly to the tui or cli. I'll look into this in a follow up


const STREAM_TAG = "$stream";

async function freezeStream(response: unknown): Promise<unknown> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

took me a sec, but this makes sense and feels like a simple solution.

correct me if I misunderstood, but for streaming apis the response is not json serializable, so we convert the response via the stream tag. In the process, we resolve the whole stream, so we need to convert the data back to a stream before passing it downstream to consumers who are expecting a stream.

nice!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep ur correct. The reviveStream name seems a little weird to me, maybe Ill use mimicStream.

@nborges-awsnborges-aws left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PR overall LGTM. One concern with a gap in testing. Tests are all either exercise happy path, or mock invokeDataset with a mocked failure object in the response. Nowhere is exercising actual invocation failures assemble that failure object correctly.

The deleted invokeDataset test is actually the only place that exercised this.

@notgitikanotgitika left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Left a comment, could be followup

if (typeof stream?.transformToString !== "function") return response;
return {
...(response as Record<string, unknown>),
response: { [STREAM_TAG]: await stream.transformToString() },

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we serialize transformToByteArray() as base64 instead? transformToString() decodes as UTF-8, so binary or invalid UTF-8 runtime responses are changed during replay.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can do in a follow up.

@jariy17
jariy17 merged commit b78d396 into refactorAug 27, 2026
26 checks passed
@jariy17
jariy17 deleted the feat/eval-simulate-followups branch August 27, 2026 16:57
@jariy17

Copy link
Copy Markdown
ContributorAuthor

I'll address the missing bad paths in a follow up

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@jariy17@codecov-commenter@Hweinstock@notgitika@nborges-aws
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(eval): batch simulate — --ingestion-wait-ms + per-example failures/sessions - #2098

Merged
jariy17 merged 7 commits into
refactorfrom
feat/eval-simulate-followups
Aug 27, 2026
Merged

feat(eval): batch simulate — --ingestion-wait-ms + per-example failures/sessions#2098
jariy17 merged 7 commits into
refactorfrom
feat/eval-simulate-followups

Conversation

@jariy17

@jariy17jariy17 commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

What

Batch simulate review follow-ups (base: refactor):

  1. --ingestion-wait-ms flag (default 180000, 0 skips) → InvokeDatasetInput.waitIngestionMs. Drops the SIMULATE_INGESTION_WAIT_MS env var — value flows through the input only.
  2. Per-example failures.runExamples returns failures: {item,error}[]; invokeDataset surfaces failures: [{exampleId, error}] — names which examples dropped + why, not just a count.
  3. Output now renders sessions[] + failures[] (below).
  4. Uses Golden Test Pattern for simulate instead of dedicated unit tests for invokeDataset

Output

{
"batchEvaluationId": "batch-eval-test", "status": "RUNNING",
"examplesInvoked": 1, "examplesFailed": 1,
"sessions": [ { "exampleId": "ok1", "sessionId": "s1" } ],
"failures": [ { "exampleId": "bad", "error": "HTTP 500" } ]
}
  • sessions[]exampleId ↔ sessionId. Join key: a later eval batch-evaluation get returns results[] keyed by sessionId, mapping each score back to its dataset row.
  • failures[] — always present ([] on a clean run); names which examples dropped + why.

…ailures/sessions
Follow-ups from the batch-evaluation simulate review:
- Add --ingestion-wait-ms (default 180000, 0 skips); thread via InvokeDatasetInput.waitIngestionMs.
Removes the SIMULATE_INGESTION_WAIT_MS env var — tests pass the value through the input.
- runExamples now returns per-item failures (item + error), not a bare count + firstError; invokeDataset
surfaces failures: [{ exampleId, error }] so a partial failure names which examples dropped and why.
- batch simulate output renders sessions[] (exampleId <-> sessionId join key for a later get) and
failures[] (omitted when empty).
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added agentcore-harness-reviewing AgentCore Harness review in progress claude-security-reviewing Claude Code /security-review in progress and removed agentcore-harness-reviewing AgentCore Harness review in progress labels Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@codecov-commenter

codecov-commenter commented Aug 25, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.85714% with 3 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.23%. Comparing base (343d213) to head (ebc604e).
⚠️ Report is 12 commits behind head on refactor.

Files with missing linesPatch %Lines
src/core/eval.tsx86.36%3 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #2098 +/- ##
============================================
- Coverage 97.33% 97.23% -0.11% 
============================================
Files 417 417 Lines 25250 25269 +19 ============================================
- Hits 24578 24570 -8 - Misses 672 699 +27 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 26, 2026
- inject newSessionId into EvalClient (default randomUUID) so replay fixtures + goldens are deterministic
- teach makeRecordingSend to freeze/revive a streaming SDK response (InvokeAgentRuntime), which stringify couldn't serialize
- add simulate fixture-golden case; move handler edges to batch-evaluation.test.tsx
- split invokeDataset.test.ts into run.test.ts (pool) + load.test.ts (parse + GT-shape); delete it and simulate.test.tsx
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/m PR size: M labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@jariy17
jariy17 marked this pull request as ready for review August 26, 2026 22:34
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot added claude-security-reviewing Claude Code /security-review in progress and removed claude-security-reviewing Claude Code /security-review in progress labels Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
);
if (r.invoked === 0) {
const detail = r.firstError ? `; first error: ${r.firstError.message}` : "";
const first = r.failures[0];

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense to just join all the errors? Or are these already communicated to the customer somewhere else?

Just wanted to make sure they'd be aware of individual datasets failing even if the overall command doesn't fail.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes all the errors would show in the debug logs. I assumed that if all dataset examples have failed, it would be the same error. However, maybe we should convene these errors directly to the tui or cli. I'll look into this in a follow up


const STREAM_TAG = "$stream";

async function freezeStream(response: unknown): Promise<unknown> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

took me a sec, but this makes sense and feels like a simple solution.

correct me if I misunderstood, but for streaming apis the response is not json serializable, so we convert the response via the stream tag. In the process, we resolve the whole stream, so we need to convert the data back to a stream before passing it downstream to consumers who are expecting a stream.

nice!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep ur correct. The reviveStream name seems a little weird to me, maybe Ill use mimicStream.

@nborges-awsnborges-aws left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PR overall LGTM. One concern with a gap in testing. Tests are all either exercise happy path, or mock invokeDataset with a mocked failure object in the response. Nowhere is exercising actual invocation failures assemble that failure object correctly.

The deleted invokeDataset test is actually the only place that exercised this.

@notgitikanotgitika left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Left a comment, could be followup

if (typeof stream?.transformToString !== "function") return response;
return {
...(response as Record<string, unknown>),
response: { [STREAM_TAG]: await stream.transformToString() },

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we serialize transformToByteArray() as base64 instead? transformToString() decodes as UTF-8, so binary or invalid UTF-8 runtime responses are changed during replay.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can do in a follow up.

@jariy17
jariy17 merged commit b78d396 into refactorAug 27, 2026
26 checks passed
@jariy17
jariy17 deleted the feat/eval-simulate-followups branch August 27, 2026 16:57
@jariy17

Copy link
Copy Markdown
ContributorAuthor

I'll address the missing bad paths in a follow up

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@jariy17@codecov-commenter@Hweinstock@notgitika@nborges-aws
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(eval): batch simulate — --ingestion-wait-ms + per-example failures/sessions - #2098

Merged
jariy17 merged 7 commits into
refactorfrom
feat/eval-simulate-followups
Aug 27, 2026
Merged

feat(eval): batch simulate — --ingestion-wait-ms + per-example failures/sessions#2098
jariy17 merged 7 commits into
refactorfrom
feat/eval-simulate-followups

Conversation

@jariy17

@jariy17jariy17 commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

What

Batch simulate review follow-ups (base: refactor):

  1. --ingestion-wait-ms flag (default 180000, 0 skips) → InvokeDatasetInput.waitIngestionMs. Drops the SIMULATE_INGESTION_WAIT_MS env var — value flows through the input only.
  2. Per-example failures.runExamples returns failures: {item,error}[]; invokeDataset surfaces failures: [{exampleId, error}] — names which examples dropped + why, not just a count.
  3. Output now renders sessions[] + failures[] (below).
  4. Uses Golden Test Pattern for simulate instead of dedicated unit tests for invokeDataset

Output

{
"batchEvaluationId": "batch-eval-test", "status": "RUNNING",
"examplesInvoked": 1, "examplesFailed": 1,
"sessions": [ { "exampleId": "ok1", "sessionId": "s1" } ],
"failures": [ { "exampleId": "bad", "error": "HTTP 500" } ]
}
  • sessions[]exampleId ↔ sessionId. Join key: a later eval batch-evaluation get returns results[] keyed by sessionId, mapping each score back to its dataset row.
  • failures[] — always present ([] on a clean run); names which examples dropped + why.

…ailures/sessions
Follow-ups from the batch-evaluation simulate review:
- Add --ingestion-wait-ms (default 180000, 0 skips); thread via InvokeDatasetInput.waitIngestionMs.
Removes the SIMULATE_INGESTION_WAIT_MS env var — tests pass the value through the input.
- runExamples now returns per-item failures (item + error), not a bare count + firstError; invokeDataset
surfaces failures: [{ exampleId, error }] so a partial failure names which examples dropped and why.
- batch simulate output renders sessions[] (exampleId <-> sessionId join key for a later get) and
failures[] (omitted when empty).
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added agentcore-harness-reviewing AgentCore Harness review in progress claude-security-reviewing Claude Code /security-review in progress and removed agentcore-harness-reviewing AgentCore Harness review in progress labels Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@codecov-commenter

codecov-commenter commented Aug 25, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.85714% with 3 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.23%. Comparing base (343d213) to head (ebc604e).
⚠️ Report is 12 commits behind head on refactor.

Files with missing linesPatch %Lines
src/core/eval.tsx86.36%3 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #2098 +/- ##
============================================
- Coverage 97.33% 97.23% -0.11% 
============================================
Files 417 417 Lines 25250 25269 +19 ============================================
- Hits 24578 24570 -8 - Misses 672 699 +27 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 26, 2026
- inject newSessionId into EvalClient (default randomUUID) so replay fixtures + goldens are deterministic
- teach makeRecordingSend to freeze/revive a streaming SDK response (InvokeAgentRuntime), which stringify couldn't serialize
- add simulate fixture-golden case; move handler edges to batch-evaluation.test.tsx
- split invokeDataset.test.ts into run.test.ts (pool) + load.test.ts (parse + GT-shape); delete it and simulate.test.tsx
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/m PR size: M labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@jariy17
jariy17 marked this pull request as ready for review August 26, 2026 22:34
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot added claude-security-reviewing Claude Code /security-review in progress and removed claude-security-reviewing Claude Code /security-review in progress labels Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
);
if (r.invoked === 0) {
const detail = r.firstError ? `; first error: ${r.firstError.message}` : "";
const first = r.failures[0];

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense to just join all the errors? Or are these already communicated to the customer somewhere else?

Just wanted to make sure they'd be aware of individual datasets failing even if the overall command doesn't fail.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes all the errors would show in the debug logs. I assumed that if all dataset examples have failed, it would be the same error. However, maybe we should convene these errors directly to the tui or cli. I'll look into this in a follow up


const STREAM_TAG = "$stream";

async function freezeStream(response: unknown): Promise<unknown> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

took me a sec, but this makes sense and feels like a simple solution.

correct me if I misunderstood, but for streaming apis the response is not json serializable, so we convert the response via the stream tag. In the process, we resolve the whole stream, so we need to convert the data back to a stream before passing it downstream to consumers who are expecting a stream.

nice!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep ur correct. The reviveStream name seems a little weird to me, maybe Ill use mimicStream.

@nborges-awsnborges-aws left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PR overall LGTM. One concern with a gap in testing. Tests are all either exercise happy path, or mock invokeDataset with a mocked failure object in the response. Nowhere is exercising actual invocation failures assemble that failure object correctly.

The deleted invokeDataset test is actually the only place that exercised this.

@notgitikanotgitika left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Left a comment, could be followup

if (typeof stream?.transformToString !== "function") return response;
return {
...(response as Record<string, unknown>),
response: { [STREAM_TAG]: await stream.transformToString() },

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we serialize transformToByteArray() as base64 instead? transformToString() decodes as UTF-8, so binary or invalid UTF-8 runtime responses are changed during replay.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can do in a follow up.

@jariy17
jariy17 merged commit b78d396 into refactorAug 27, 2026
26 checks passed
@jariy17
jariy17 deleted the feat/eval-simulate-followups branch August 27, 2026 16:57
@jariy17

Copy link
Copy Markdown
ContributorAuthor

I'll address the missing bad paths in a follow up

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@jariy17@codecov-commenter@Hweinstock@notgitika@nborges-aws
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(eval): batch simulate — --ingestion-wait-ms + per-example failures/sessions - #2098

Merged
jariy17 merged 7 commits into
refactorfrom
feat/eval-simulate-followups
Aug 27, 2026
Merged

feat(eval): batch simulate — --ingestion-wait-ms + per-example failures/sessions#2098
jariy17 merged 7 commits into
refactorfrom
feat/eval-simulate-followups

Conversation

@jariy17

@jariy17jariy17 commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

What

Batch simulate review follow-ups (base: refactor):

  1. --ingestion-wait-ms flag (default 180000, 0 skips) → InvokeDatasetInput.waitIngestionMs. Drops the SIMULATE_INGESTION_WAIT_MS env var — value flows through the input only.
  2. Per-example failures.runExamples returns failures: {item,error}[]; invokeDataset surfaces failures: [{exampleId, error}] — names which examples dropped + why, not just a count.
  3. Output now renders sessions[] + failures[] (below).
  4. Uses Golden Test Pattern for simulate instead of dedicated unit tests for invokeDataset

Output

{
"batchEvaluationId": "batch-eval-test", "status": "RUNNING",
"examplesInvoked": 1, "examplesFailed": 1,
"sessions": [ { "exampleId": "ok1", "sessionId": "s1" } ],
"failures": [ { "exampleId": "bad", "error": "HTTP 500" } ]
}
  • sessions[]exampleId ↔ sessionId. Join key: a later eval batch-evaluation get returns results[] keyed by sessionId, mapping each score back to its dataset row.
  • failures[] — always present ([] on a clean run); names which examples dropped + why.

…ailures/sessions
Follow-ups from the batch-evaluation simulate review:
- Add --ingestion-wait-ms (default 180000, 0 skips); thread via InvokeDatasetInput.waitIngestionMs.
Removes the SIMULATE_INGESTION_WAIT_MS env var — tests pass the value through the input.
- runExamples now returns per-item failures (item + error), not a bare count + firstError; invokeDataset
surfaces failures: [{ exampleId, error }] so a partial failure names which examples dropped and why.
- batch simulate output renders sessions[] (exampleId <-> sessionId join key for a later get) and
failures[] (omitted when empty).
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added agentcore-harness-reviewing AgentCore Harness review in progress claude-security-reviewing Claude Code /security-review in progress and removed agentcore-harness-reviewing AgentCore Harness review in progress labels Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@codecov-commenter

codecov-commenter commented Aug 25, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.85714% with 3 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.23%. Comparing base (343d213) to head (ebc604e).
⚠️ Report is 12 commits behind head on refactor.

Files with missing linesPatch %Lines
src/core/eval.tsx86.36%3 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #2098 +/- ##
============================================
- Coverage 97.33% 97.23% -0.11% 
============================================
Files 417 417 Lines 25250 25269 +19 ============================================
- Hits 24578 24570 -8 - Misses 672 699 +27 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 26, 2026
- inject newSessionId into EvalClient (default randomUUID) so replay fixtures + goldens are deterministic
- teach makeRecordingSend to freeze/revive a streaming SDK response (InvokeAgentRuntime), which stringify couldn't serialize
- add simulate fixture-golden case; move handler edges to batch-evaluation.test.tsx
- split invokeDataset.test.ts into run.test.ts (pool) + load.test.ts (parse + GT-shape); delete it and simulate.test.tsx
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/m PR size: M labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@jariy17
jariy17 marked this pull request as ready for review August 26, 2026 22:34
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot added claude-security-reviewing Claude Code /security-review in progress and removed claude-security-reviewing Claude Code /security-review in progress labels Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
);
if (r.invoked === 0) {
const detail = r.firstError ? `; first error: ${r.firstError.message}` : "";
const first = r.failures[0];

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense to just join all the errors? Or are these already communicated to the customer somewhere else?

Just wanted to make sure they'd be aware of individual datasets failing even if the overall command doesn't fail.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes all the errors would show in the debug logs. I assumed that if all dataset examples have failed, it would be the same error. However, maybe we should convene these errors directly to the tui or cli. I'll look into this in a follow up


const STREAM_TAG = "$stream";

async function freezeStream(response: unknown): Promise<unknown> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

took me a sec, but this makes sense and feels like a simple solution.

correct me if I misunderstood, but for streaming apis the response is not json serializable, so we convert the response via the stream tag. In the process, we resolve the whole stream, so we need to convert the data back to a stream before passing it downstream to consumers who are expecting a stream.

nice!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep ur correct. The reviveStream name seems a little weird to me, maybe Ill use mimicStream.

@nborges-awsnborges-aws left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PR overall LGTM. One concern with a gap in testing. Tests are all either exercise happy path, or mock invokeDataset with a mocked failure object in the response. Nowhere is exercising actual invocation failures assemble that failure object correctly.

The deleted invokeDataset test is actually the only place that exercised this.

@notgitikanotgitika left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Left a comment, could be followup

if (typeof stream?.transformToString !== "function") return response;
return {
...(response as Record<string, unknown>),
response: { [STREAM_TAG]: await stream.transformToString() },

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we serialize transformToByteArray() as base64 instead? transformToString() decodes as UTF-8, so binary or invalid UTF-8 runtime responses are changed during replay.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can do in a follow up.

@jariy17
jariy17 merged commit b78d396 into refactorAug 27, 2026
26 checks passed
@jariy17
jariy17 deleted the feat/eval-simulate-followups branch August 27, 2026 16:57
@jariy17

Copy link
Copy Markdown
ContributorAuthor

I'll address the missing bad paths in a follow up

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@jariy17@codecov-commenter@Hweinstock@notgitika@nborges-aws
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(eval): batch simulate — --ingestion-wait-ms + per-example failures/sessions - #2098

Merged
jariy17 merged 7 commits into
refactorfrom
feat/eval-simulate-followups
Aug 27, 2026
Merged

feat(eval): batch simulate — --ingestion-wait-ms + per-example failures/sessions#2098
jariy17 merged 7 commits into
refactorfrom
feat/eval-simulate-followups

Conversation

@jariy17

@jariy17jariy17 commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

What

Batch simulate review follow-ups (base: refactor):

  1. --ingestion-wait-ms flag (default 180000, 0 skips) → InvokeDatasetInput.waitIngestionMs. Drops the SIMULATE_INGESTION_WAIT_MS env var — value flows through the input only.
  2. Per-example failures.runExamples returns failures: {item,error}[]; invokeDataset surfaces failures: [{exampleId, error}] — names which examples dropped + why, not just a count.
  3. Output now renders sessions[] + failures[] (below).
  4. Uses Golden Test Pattern for simulate instead of dedicated unit tests for invokeDataset

Output

{
"batchEvaluationId": "batch-eval-test", "status": "RUNNING",
"examplesInvoked": 1, "examplesFailed": 1,
"sessions": [ { "exampleId": "ok1", "sessionId": "s1" } ],
"failures": [ { "exampleId": "bad", "error": "HTTP 500" } ]
}
  • sessions[]exampleId ↔ sessionId. Join key: a later eval batch-evaluation get returns results[] keyed by sessionId, mapping each score back to its dataset row.
  • failures[] — always present ([] on a clean run); names which examples dropped + why.

…ailures/sessions
Follow-ups from the batch-evaluation simulate review:
- Add --ingestion-wait-ms (default 180000, 0 skips); thread via InvokeDatasetInput.waitIngestionMs.
Removes the SIMULATE_INGESTION_WAIT_MS env var — tests pass the value through the input.
- runExamples now returns per-item failures (item + error), not a bare count + firstError; invokeDataset
surfaces failures: [{ exampleId, error }] so a partial failure names which examples dropped and why.
- batch simulate output renders sessions[] (exampleId <-> sessionId join key for a later get) and
failures[] (omitted when empty).
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added agentcore-harness-reviewing AgentCore Harness review in progress claude-security-reviewing Claude Code /security-review in progress and removed agentcore-harness-reviewing AgentCore Harness review in progress labels Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@codecov-commenter

codecov-commenter commented Aug 25, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.85714% with 3 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.23%. Comparing base (343d213) to head (ebc604e).
⚠️ Report is 12 commits behind head on refactor.

Files with missing linesPatch %Lines
src/core/eval.tsx86.36%3 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #2098 +/- ##
============================================
- Coverage 97.33% 97.23% -0.11% 
============================================
Files 417 417 Lines 25250 25269 +19 ============================================
- Hits 24578 24570 -8 - Misses 672 699 +27 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 26, 2026
- inject newSessionId into EvalClient (default randomUUID) so replay fixtures + goldens are deterministic
- teach makeRecordingSend to freeze/revive a streaming SDK response (InvokeAgentRuntime), which stringify couldn't serialize
- add simulate fixture-golden case; move handler edges to batch-evaluation.test.tsx
- split invokeDataset.test.ts into run.test.ts (pool) + load.test.ts (parse + GT-shape); delete it and simulate.test.tsx
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/m PR size: M labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@jariy17
jariy17 marked this pull request as ready for review August 26, 2026 22:34
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot added claude-security-reviewing Claude Code /security-review in progress and removed claude-security-reviewing Claude Code /security-review in progress labels Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
);
if (r.invoked === 0) {
const detail = r.firstError ? `; first error: ${r.firstError.message}` : "";
const first = r.failures[0];

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense to just join all the errors? Or are these already communicated to the customer somewhere else?

Just wanted to make sure they'd be aware of individual datasets failing even if the overall command doesn't fail.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes all the errors would show in the debug logs. I assumed that if all dataset examples have failed, it would be the same error. However, maybe we should convene these errors directly to the tui or cli. I'll look into this in a follow up


const STREAM_TAG = "$stream";

async function freezeStream(response: unknown): Promise<unknown> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

took me a sec, but this makes sense and feels like a simple solution.

correct me if I misunderstood, but for streaming apis the response is not json serializable, so we convert the response via the stream tag. In the process, we resolve the whole stream, so we need to convert the data back to a stream before passing it downstream to consumers who are expecting a stream.

nice!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep ur correct. The reviveStream name seems a little weird to me, maybe Ill use mimicStream.

@nborges-awsnborges-aws left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PR overall LGTM. One concern with a gap in testing. Tests are all either exercise happy path, or mock invokeDataset with a mocked failure object in the response. Nowhere is exercising actual invocation failures assemble that failure object correctly.

The deleted invokeDataset test is actually the only place that exercised this.

@notgitikanotgitika left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Left a comment, could be followup

if (typeof stream?.transformToString !== "function") return response;
return {
...(response as Record<string, unknown>),
response: { [STREAM_TAG]: await stream.transformToString() },

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we serialize transformToByteArray() as base64 instead? transformToString() decodes as UTF-8, so binary or invalid UTF-8 runtime responses are changed during replay.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can do in a follow up.

@jariy17
jariy17 merged commit b78d396 into refactorAug 27, 2026
26 checks passed
@jariy17
jariy17 deleted the feat/eval-simulate-followups branch August 27, 2026 16:57
@jariy17

Copy link
Copy Markdown
ContributorAuthor

I'll address the missing bad paths in a follow up

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@jariy17@codecov-commenter@Hweinstock@notgitika@nborges-aws
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(eval): batch simulate — --ingestion-wait-ms + per-example failures/sessions - #2098

Merged
jariy17 merged 7 commits into
refactorfrom
feat/eval-simulate-followups
Aug 27, 2026
Merged

feat(eval): batch simulate — --ingestion-wait-ms + per-example failures/sessions#2098
jariy17 merged 7 commits into
refactorfrom
feat/eval-simulate-followups

Conversation

@jariy17

@jariy17jariy17 commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

What

Batch simulate review follow-ups (base: refactor):

  1. --ingestion-wait-ms flag (default 180000, 0 skips) → InvokeDatasetInput.waitIngestionMs. Drops the SIMULATE_INGESTION_WAIT_MS env var — value flows through the input only.
  2. Per-example failures.runExamples returns failures: {item,error}[]; invokeDataset surfaces failures: [{exampleId, error}] — names which examples dropped + why, not just a count.
  3. Output now renders sessions[] + failures[] (below).
  4. Uses Golden Test Pattern for simulate instead of dedicated unit tests for invokeDataset

Output

{
"batchEvaluationId": "batch-eval-test", "status": "RUNNING",
"examplesInvoked": 1, "examplesFailed": 1,
"sessions": [ { "exampleId": "ok1", "sessionId": "s1" } ],
"failures": [ { "exampleId": "bad", "error": "HTTP 500" } ]
}
  • sessions[]exampleId ↔ sessionId. Join key: a later eval batch-evaluation get returns results[] keyed by sessionId, mapping each score back to its dataset row.
  • failures[] — always present ([] on a clean run); names which examples dropped + why.

…ailures/sessions
Follow-ups from the batch-evaluation simulate review:
- Add --ingestion-wait-ms (default 180000, 0 skips); thread via InvokeDatasetInput.waitIngestionMs.
Removes the SIMULATE_INGESTION_WAIT_MS env var — tests pass the value through the input.
- runExamples now returns per-item failures (item + error), not a bare count + firstError; invokeDataset
surfaces failures: [{ exampleId, error }] so a partial failure names which examples dropped and why.
- batch simulate output renders sessions[] (exampleId <-> sessionId join key for a later get) and
failures[] (omitted when empty).
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added agentcore-harness-reviewing AgentCore Harness review in progress claude-security-reviewing Claude Code /security-review in progress and removed agentcore-harness-reviewing AgentCore Harness review in progress labels Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@codecov-commenter

codecov-commenter commented Aug 25, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.85714% with 3 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.23%. Comparing base (343d213) to head (ebc604e).
⚠️ Report is 12 commits behind head on refactor.

Files with missing linesPatch %Lines
src/core/eval.tsx86.36%3 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #2098 +/- ##
============================================
- Coverage 97.33% 97.23% -0.11% 
============================================
Files 417 417 Lines 25250 25269 +19 ============================================
- Hits 24578 24570 -8 - Misses 672 699 +27 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 26, 2026
- inject newSessionId into EvalClient (default randomUUID) so replay fixtures + goldens are deterministic
- teach makeRecordingSend to freeze/revive a streaming SDK response (InvokeAgentRuntime), which stringify couldn't serialize
- add simulate fixture-golden case; move handler edges to batch-evaluation.test.tsx
- split invokeDataset.test.ts into run.test.ts (pool) + load.test.ts (parse + GT-shape); delete it and simulate.test.tsx
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/m PR size: M labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@jariy17
jariy17 marked this pull request as ready for review August 26, 2026 22:34
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot added claude-security-reviewing Claude Code /security-review in progress and removed claude-security-reviewing Claude Code /security-review in progress labels Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
);
if (r.invoked === 0) {
const detail = r.firstError ? `; first error: ${r.firstError.message}` : "";
const first = r.failures[0];

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense to just join all the errors? Or are these already communicated to the customer somewhere else?

Just wanted to make sure they'd be aware of individual datasets failing even if the overall command doesn't fail.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes all the errors would show in the debug logs. I assumed that if all dataset examples have failed, it would be the same error. However, maybe we should convene these errors directly to the tui or cli. I'll look into this in a follow up


const STREAM_TAG = "$stream";

async function freezeStream(response: unknown): Promise<unknown> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

took me a sec, but this makes sense and feels like a simple solution.

correct me if I misunderstood, but for streaming apis the response is not json serializable, so we convert the response via the stream tag. In the process, we resolve the whole stream, so we need to convert the data back to a stream before passing it downstream to consumers who are expecting a stream.

nice!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep ur correct. The reviveStream name seems a little weird to me, maybe Ill use mimicStream.

@nborges-awsnborges-aws left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PR overall LGTM. One concern with a gap in testing. Tests are all either exercise happy path, or mock invokeDataset with a mocked failure object in the response. Nowhere is exercising actual invocation failures assemble that failure object correctly.

The deleted invokeDataset test is actually the only place that exercised this.

@notgitikanotgitika left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Left a comment, could be followup

if (typeof stream?.transformToString !== "function") return response;
return {
...(response as Record<string, unknown>),
response: { [STREAM_TAG]: await stream.transformToString() },

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we serialize transformToByteArray() as base64 instead? transformToString() decodes as UTF-8, so binary or invalid UTF-8 runtime responses are changed during replay.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can do in a follow up.

@jariy17
jariy17 merged commit b78d396 into refactorAug 27, 2026
26 checks passed
@jariy17
jariy17 deleted the feat/eval-simulate-followups branch August 27, 2026 16:57
@jariy17

Copy link
Copy Markdown
ContributorAuthor

I'll address the missing bad paths in a follow up

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@jariy17@codecov-commenter@Hweinstock@notgitika@nborges-aws
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(eval): batch simulate — --ingestion-wait-ms + per-example failures/sessions - #2098

Merged
jariy17 merged 7 commits into
refactorfrom
feat/eval-simulate-followups
Aug 27, 2026
Merged

feat(eval): batch simulate — --ingestion-wait-ms + per-example failures/sessions#2098
jariy17 merged 7 commits into
refactorfrom
feat/eval-simulate-followups

Conversation

@jariy17

@jariy17jariy17 commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

What

Batch simulate review follow-ups (base: refactor):

  1. --ingestion-wait-ms flag (default 180000, 0 skips) → InvokeDatasetInput.waitIngestionMs. Drops the SIMULATE_INGESTION_WAIT_MS env var — value flows through the input only.
  2. Per-example failures.runExamples returns failures: {item,error}[]; invokeDataset surfaces failures: [{exampleId, error}] — names which examples dropped + why, not just a count.
  3. Output now renders sessions[] + failures[] (below).
  4. Uses Golden Test Pattern for simulate instead of dedicated unit tests for invokeDataset

Output

{
"batchEvaluationId": "batch-eval-test", "status": "RUNNING",
"examplesInvoked": 1, "examplesFailed": 1,
"sessions": [ { "exampleId": "ok1", "sessionId": "s1" } ],
"failures": [ { "exampleId": "bad", "error": "HTTP 500" } ]
}
  • sessions[]exampleId ↔ sessionId. Join key: a later eval batch-evaluation get returns results[] keyed by sessionId, mapping each score back to its dataset row.
  • failures[] — always present ([] on a clean run); names which examples dropped + why.

…ailures/sessions
Follow-ups from the batch-evaluation simulate review:
- Add --ingestion-wait-ms (default 180000, 0 skips); thread via InvokeDatasetInput.waitIngestionMs.
Removes the SIMULATE_INGESTION_WAIT_MS env var — tests pass the value through the input.
- runExamples now returns per-item failures (item + error), not a bare count + firstError; invokeDataset
surfaces failures: [{ exampleId, error }] so a partial failure names which examples dropped and why.
- batch simulate output renders sessions[] (exampleId <-> sessionId join key for a later get) and
failures[] (omitted when empty).
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added agentcore-harness-reviewing AgentCore Harness review in progress claude-security-reviewing Claude Code /security-review in progress and removed agentcore-harness-reviewing AgentCore Harness review in progress labels Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@codecov-commenter

codecov-commenter commented Aug 25, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.85714% with 3 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.23%. Comparing base (343d213) to head (ebc604e).
⚠️ Report is 12 commits behind head on refactor.

Files with missing linesPatch %Lines
src/core/eval.tsx86.36%3 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #2098 +/- ##
============================================
- Coverage 97.33% 97.23% -0.11% 
============================================
Files 417 417 Lines 25250 25269 +19 ============================================
- Hits 24578 24570 -8 - Misses 672 699 +27 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 25, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 25, 2026
@github-actionsgithub-actionsBot added size/m PR size: M and removed size/m PR size: M labels Aug 26, 2026
- inject newSessionId into EvalClient (default randomUUID) so replay fixtures + goldens are deterministic
- teach makeRecordingSend to freeze/revive a streaming SDK response (InvokeAgentRuntime), which stringify couldn't serialize
- add simulate fixture-golden case; move handler edges to batch-evaluation.test.tsx
- split invokeDataset.test.ts into run.test.ts (pool) + load.test.ts (parse + GT-shape); delete it and simulate.test.tsx
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/m PR size: M labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Aug 26, 2026
@jariy17
jariy17 marked this pull request as ready for review August 26, 2026 22:34
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot added claude-security-reviewing Claude Code /security-review in progress and removed claude-security-reviewing Claude Code /security-review in progress labels Aug 26, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 26, 2026
);
if (r.invoked === 0) {
const detail = r.firstError ? `; first error: ${r.firstError.message}` : "";
const first = r.failures[0];

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it make sense to just join all the errors? Or are these already communicated to the customer somewhere else?

Just wanted to make sure they'd be aware of individual datasets failing even if the overall command doesn't fail.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes all the errors would show in the debug logs. I assumed that if all dataset examples have failed, it would be the same error. However, maybe we should convene these errors directly to the tui or cli. I'll look into this in a follow up


const STREAM_TAG = "$stream";

async function freezeStream(response: unknown): Promise<unknown> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

took me a sec, but this makes sense and feels like a simple solution.

correct me if I misunderstood, but for streaming apis the response is not json serializable, so we convert the response via the stream tag. In the process, we resolve the whole stream, so we need to convert the data back to a stream before passing it downstream to consumers who are expecting a stream.

nice!

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep ur correct. The reviveStream name seems a little weird to me, maybe Ill use mimicStream.

@nborges-awsnborges-aws left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PR overall LGTM. One concern with a gap in testing. Tests are all either exercise happy path, or mock invokeDataset with a mocked failure object in the response. Nowhere is exercising actual invocation failures assemble that failure object correctly.

The deleted invokeDataset test is actually the only place that exercised this.

@notgitikanotgitika left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Left a comment, could be followup

if (typeof stream?.transformToString !== "function") return response;
return {
...(response as Record<string, unknown>),
response: { [STREAM_TAG]: await stream.transformToString() },

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we serialize transformToByteArray() as base64 instead? transformToString() decodes as UTF-8, so binary or invalid UTF-8 runtime responses are changed during replay.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can do in a follow up.

@jariy17
jariy17 merged commit b78d396 into refactorAug 27, 2026
26 checks passed
@jariy17
jariy17 deleted the feat/eval-simulate-followups branch August 27, 2026 16:57
@jariy17

Copy link
Copy Markdown
ContributorAuthor

I'll address the missing bad paths in a follow up

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@jariy17@codecov-commenter@Hweinstock@notgitika@nborges-aws