feat: add ground truth reference inputs for on-demand evaluation - #732

Merged
jariy17 merged 3 commits into
mainfrom
feat/GT_Support
Mar 30, 2026
Merged

feat: add ground truth reference inputs for on-demand evaluation#732
jariy17 merged 3 commits into
mainfrom
feat/GT_Support

Conversation

@jariy17

Copy link
Copy Markdown
Contributor

Summary

Adds ground truth reference inputs to on-demand evaluation (run eval), allowing users to specify expected agent behavior for more precise evaluation.

New CLI flags

  • -A, --assertion <text...> — Assertion the agent should satisfy (repeatable)
  • --expected-trajectory <names> — Expected tool calls in order (comma-separated)
  • --expected-response <text> — Expected agent response text

TUI changes

  • New Ground Truth wizard step in run eval flow (appears when exactly 1 session is selected)
  • Evaluator creation screen now shows reference input placeholders ({assertions}, {expected_tool_trajectory}, {actual_tool_trajectory}, {expected_response})
  • Updated default evaluator instructions to include GT placeholders

Implementation

  • evaluationReferenceInputs passed through to the evaluate() API call
  • Results include referenceInputs summary in both CLI and TUI output
  • Note: The second commit (temp: use middleware...) is a temporary workaround — the SDK does not yet include evaluationReferenceInputs in its model, so we inject it via Smithy middleware. Remove when SDK is updated.

Test plan

  • npx tsc --noEmit — clean
  • npm run lint — 0 errors
  • 5 new unit tests for GT reference inputs (all pass)
  • Existing 33 run-eval tests continue to pass

@jariy17
jariy17 requested a review from a teamMarch 30, 2026 19:48
@github-actionsgithub-actionsBot added the size/l PR size: L label Mar 30, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Package Tarball

aws-agentcore-0.4.0.tgz

How to install

npm install https://github.com/aws/agentcore-cli/releases/download/pr-732-tarball/aws-agentcore-0.4.0.tgz

@github-actions

github-actionsBot commented Mar 30, 2026

Copy link
Copy Markdown
Contributor

Coverage Report

StatusCategoryPercentageCovered / Total
🔵Lines45.98%6575 / 14297
🔵Statements45.56%6986 / 15333
🔵Functions44.68%1177 / 2634
🔵Branches46.18%4362 / 9444
Generated in workflow #1538 for commit d5e84ae by the Vitest Coverage Report Action

Bump @aws-sdk/client-bedrock-agentcore and
@aws-sdk/client-bedrock-agentcore-control from ^3.893.0 to ^3.1020.0.
The SDK now includes EvaluationReferenceInput in its model,
so we pass it directly to EvaluateCommand instead of injecting
via middleware.
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Mar 30, 2026
Comment threadsrc/cli/aws/agentcore.ts Outdated
@jariy17

Copy link
Copy Markdown
ContributorAuthor

Testing Results

Verified all new GT features against a deployed runtime (TempAgent in EvalCheck project):

CLI flags tested

  • -A, --assertion — single and multiple assertions with Builtin.GoalSuccessRate
  • --expected-trajectory — comma-separated tool names with Builtin.TrajectoryExactOrderMatch
  • --expected-response — with and without explicit -t trace ID, using Builtin.Correctness
  • All three combined in a single run across multiple evaluators
  • Baseline run (no GT inputs) — confirms referenceInputs is absent from output

Custom evaluators tested

  • MyEvaluatorTraceGT (TRACE level) — {expected_response} placeholder populated correctly
  • CustomEvaluatorSessionGT (SESSION level) — {assertions}, {expected_tool_trajectory}, {actual_tool_trajectory} placeholders populated correctly

SDK upgrade verified

  • @aws-sdk/client-bedrock-agentcore@^3.1020.0evaluationReferenceInputs passed natively via EvaluateCommand (no middleware workaround needed)

Other checks

  • npx tsc --noEmit — clean
  • npm run lint — 0 errors
  • All 38 unit tests pass (including 5 new GT tests)
  • TUI run eval wizard shows Ground Truth step when exactly 1 session is selected

@HweinstockHweinstock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some nits and question on service behavior. Otherwise lgtm.

Comment threadsrc/cli/commands/run/command.tsx
Comment threadsrc/cli/commands/run/command.tsx
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/tui/screens/evaluator/types.ts
Hweinstock
Hweinstock previously approved these changes Mar 30, 2026

@HweinstockHweinstock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Slightly confused on the customer experience, but thats likely because I'm new to evals. The changes themselves lgtm.

@awsaws deleted a comment from HweinstockMar 30, 2026
@jariy17
jariy17 merged commit 01623ff into mainMar 30, 2026
22 checks passed
@jariy17
jariy17 deleted the feat/GT_Support branch March 30, 2026 23:03
aidandaly24 added a commit that referenced this pull request Apr 8, 2026
Cover 13 PRs merged to main since the initial docs audit (March 28):
- import subcommands: runtime, memory, evaluator, online-eval (#763, #780)
- invoke/dev exec mode with --exec and --timeout (#750)
- code-based evaluator support with --type, --lambda-arn, --timeout (#739)
- ground truth eval inputs: --assertion, --expected-trajectory, --expected-response (#732)
- memory record streaming: --data-stream-arn, --stream-content-level (#531, #534)
- create --skip-install flag (#782)
- fetch access --identity-name option (#774)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jariy17@Hweinstock@awsswarnim
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat: add ground truth reference inputs for on-demand evaluation - #732

Merged
jariy17 merged 3 commits into
mainfrom
feat/GT_Support
Mar 30, 2026
Merged

feat: add ground truth reference inputs for on-demand evaluation#732
jariy17 merged 3 commits into
mainfrom
feat/GT_Support

Conversation

@jariy17

Copy link
Copy Markdown
Contributor

Summary

Adds ground truth reference inputs to on-demand evaluation (run eval), allowing users to specify expected agent behavior for more precise evaluation.

New CLI flags

  • -A, --assertion <text...> — Assertion the agent should satisfy (repeatable)
  • --expected-trajectory <names> — Expected tool calls in order (comma-separated)
  • --expected-response <text> — Expected agent response text

TUI changes

  • New Ground Truth wizard step in run eval flow (appears when exactly 1 session is selected)
  • Evaluator creation screen now shows reference input placeholders ({assertions}, {expected_tool_trajectory}, {actual_tool_trajectory}, {expected_response})
  • Updated default evaluator instructions to include GT placeholders

Implementation

  • evaluationReferenceInputs passed through to the evaluate() API call
  • Results include referenceInputs summary in both CLI and TUI output
  • Note: The second commit (temp: use middleware...) is a temporary workaround — the SDK does not yet include evaluationReferenceInputs in its model, so we inject it via Smithy middleware. Remove when SDK is updated.

Test plan

  • npx tsc --noEmit — clean
  • npm run lint — 0 errors
  • 5 new unit tests for GT reference inputs (all pass)
  • Existing 33 run-eval tests continue to pass

@jariy17
jariy17 requested a review from a teamMarch 30, 2026 19:48
@github-actionsgithub-actionsBot added the size/l PR size: L label Mar 30, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Package Tarball

aws-agentcore-0.4.0.tgz

How to install

npm install https://github.com/aws/agentcore-cli/releases/download/pr-732-tarball/aws-agentcore-0.4.0.tgz

@github-actions

github-actionsBot commented Mar 30, 2026

Copy link
Copy Markdown
Contributor

Coverage Report

StatusCategoryPercentageCovered / Total
🔵Lines45.98%6575 / 14297
🔵Statements45.56%6986 / 15333
🔵Functions44.68%1177 / 2634
🔵Branches46.18%4362 / 9444
Generated in workflow #1538 for commit d5e84ae by the Vitest Coverage Report Action

Bump @aws-sdk/client-bedrock-agentcore and
@aws-sdk/client-bedrock-agentcore-control from ^3.893.0 to ^3.1020.0.
The SDK now includes EvaluationReferenceInput in its model,
so we pass it directly to EvaluateCommand instead of injecting
via middleware.
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Mar 30, 2026
Comment threadsrc/cli/aws/agentcore.ts Outdated
@jariy17

Copy link
Copy Markdown
ContributorAuthor

Testing Results

Verified all new GT features against a deployed runtime (TempAgent in EvalCheck project):

CLI flags tested

  • -A, --assertion — single and multiple assertions with Builtin.GoalSuccessRate
  • --expected-trajectory — comma-separated tool names with Builtin.TrajectoryExactOrderMatch
  • --expected-response — with and without explicit -t trace ID, using Builtin.Correctness
  • All three combined in a single run across multiple evaluators
  • Baseline run (no GT inputs) — confirms referenceInputs is absent from output

Custom evaluators tested

  • MyEvaluatorTraceGT (TRACE level) — {expected_response} placeholder populated correctly
  • CustomEvaluatorSessionGT (SESSION level) — {assertions}, {expected_tool_trajectory}, {actual_tool_trajectory} placeholders populated correctly

SDK upgrade verified

  • @aws-sdk/client-bedrock-agentcore@^3.1020.0evaluationReferenceInputs passed natively via EvaluateCommand (no middleware workaround needed)

Other checks

  • npx tsc --noEmit — clean
  • npm run lint — 0 errors
  • All 38 unit tests pass (including 5 new GT tests)
  • TUI run eval wizard shows Ground Truth step when exactly 1 session is selected

@HweinstockHweinstock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some nits and question on service behavior. Otherwise lgtm.

Comment threadsrc/cli/commands/run/command.tsx
Comment threadsrc/cli/commands/run/command.tsx
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/tui/screens/evaluator/types.ts
Hweinstock
Hweinstock previously approved these changes Mar 30, 2026

@HweinstockHweinstock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Slightly confused on the customer experience, but thats likely because I'm new to evals. The changes themselves lgtm.

@awsaws deleted a comment from HweinstockMar 30, 2026
@jariy17
jariy17 merged commit 01623ff into mainMar 30, 2026
22 checks passed
@jariy17
jariy17 deleted the feat/GT_Support branch March 30, 2026 23:03
aidandaly24 added a commit that referenced this pull request Apr 8, 2026
Cover 13 PRs merged to main since the initial docs audit (March 28):
- import subcommands: runtime, memory, evaluator, online-eval (#763, #780)
- invoke/dev exec mode with --exec and --timeout (#750)
- code-based evaluator support with --type, --lambda-arn, --timeout (#739)
- ground truth eval inputs: --assertion, --expected-trajectory, --expected-response (#732)
- memory record streaming: --data-stream-arn, --stream-content-level (#531, #534)
- create --skip-install flag (#782)
- fetch access --identity-name option (#774)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jariy17@Hweinstock@awsswarnim
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: add ground truth reference inputs for on-demand evaluation - #732

Merged
jariy17 merged 3 commits into
mainfrom
feat/GT_Support
Mar 30, 2026
Merged

feat: add ground truth reference inputs for on-demand evaluation#732
jariy17 merged 3 commits into
mainfrom
feat/GT_Support

Conversation

@jariy17

Copy link
Copy Markdown
Contributor

Summary

Adds ground truth reference inputs to on-demand evaluation (run eval), allowing users to specify expected agent behavior for more precise evaluation.

New CLI flags

  • -A, --assertion <text...> — Assertion the agent should satisfy (repeatable)
  • --expected-trajectory <names> — Expected tool calls in order (comma-separated)
  • --expected-response <text> — Expected agent response text

TUI changes

  • New Ground Truth wizard step in run eval flow (appears when exactly 1 session is selected)
  • Evaluator creation screen now shows reference input placeholders ({assertions}, {expected_tool_trajectory}, {actual_tool_trajectory}, {expected_response})
  • Updated default evaluator instructions to include GT placeholders

Implementation

  • evaluationReferenceInputs passed through to the evaluate() API call
  • Results include referenceInputs summary in both CLI and TUI output
  • Note: The second commit (temp: use middleware...) is a temporary workaround — the SDK does not yet include evaluationReferenceInputs in its model, so we inject it via Smithy middleware. Remove when SDK is updated.

Test plan

  • npx tsc --noEmit — clean
  • npm run lint — 0 errors
  • 5 new unit tests for GT reference inputs (all pass)
  • Existing 33 run-eval tests continue to pass

@jariy17
jariy17 requested a review from a teamMarch 30, 2026 19:48
@github-actionsgithub-actionsBot added the size/l PR size: L label Mar 30, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Package Tarball

aws-agentcore-0.4.0.tgz

How to install

npm install https://github.com/aws/agentcore-cli/releases/download/pr-732-tarball/aws-agentcore-0.4.0.tgz

@github-actions

github-actionsBot commented Mar 30, 2026

Copy link
Copy Markdown
Contributor

Coverage Report

StatusCategoryPercentageCovered / Total
🔵Lines45.98%6575 / 14297
🔵Statements45.56%6986 / 15333
🔵Functions44.68%1177 / 2634
🔵Branches46.18%4362 / 9444
Generated in workflow #1538 for commit d5e84ae by the Vitest Coverage Report Action

Bump @aws-sdk/client-bedrock-agentcore and
@aws-sdk/client-bedrock-agentcore-control from ^3.893.0 to ^3.1020.0.
The SDK now includes EvaluationReferenceInput in its model,
so we pass it directly to EvaluateCommand instead of injecting
via middleware.
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Mar 30, 2026
Comment threadsrc/cli/aws/agentcore.ts Outdated
@jariy17

Copy link
Copy Markdown
ContributorAuthor

Testing Results

Verified all new GT features against a deployed runtime (TempAgent in EvalCheck project):

CLI flags tested

  • -A, --assertion — single and multiple assertions with Builtin.GoalSuccessRate
  • --expected-trajectory — comma-separated tool names with Builtin.TrajectoryExactOrderMatch
  • --expected-response — with and without explicit -t trace ID, using Builtin.Correctness
  • All three combined in a single run across multiple evaluators
  • Baseline run (no GT inputs) — confirms referenceInputs is absent from output

Custom evaluators tested

  • MyEvaluatorTraceGT (TRACE level) — {expected_response} placeholder populated correctly
  • CustomEvaluatorSessionGT (SESSION level) — {assertions}, {expected_tool_trajectory}, {actual_tool_trajectory} placeholders populated correctly

SDK upgrade verified

  • @aws-sdk/client-bedrock-agentcore@^3.1020.0evaluationReferenceInputs passed natively via EvaluateCommand (no middleware workaround needed)

Other checks

  • npx tsc --noEmit — clean
  • npm run lint — 0 errors
  • All 38 unit tests pass (including 5 new GT tests)
  • TUI run eval wizard shows Ground Truth step when exactly 1 session is selected

@HweinstockHweinstock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some nits and question on service behavior. Otherwise lgtm.

Comment threadsrc/cli/commands/run/command.tsx
Comment threadsrc/cli/commands/run/command.tsx
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/tui/screens/evaluator/types.ts
Hweinstock
Hweinstock previously approved these changes Mar 30, 2026

@HweinstockHweinstock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Slightly confused on the customer experience, but thats likely because I'm new to evals. The changes themselves lgtm.

@awsaws deleted a comment from HweinstockMar 30, 2026
@jariy17
jariy17 merged commit 01623ff into mainMar 30, 2026
22 checks passed
@jariy17
jariy17 deleted the feat/GT_Support branch March 30, 2026 23:03
aidandaly24 added a commit that referenced this pull request Apr 8, 2026
Cover 13 PRs merged to main since the initial docs audit (March 28):
- import subcommands: runtime, memory, evaluator, online-eval (#763, #780)
- invoke/dev exec mode with --exec and --timeout (#750)
- code-based evaluator support with --type, --lambda-arn, --timeout (#739)
- ground truth eval inputs: --assertion, --expected-trajectory, --expected-response (#732)
- memory record streaming: --data-stream-arn, --stream-content-level (#531, #534)
- create --skip-install flag (#782)
- fetch access --identity-name option (#774)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jariy17@Hweinstock@awsswarnim
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: add ground truth reference inputs for on-demand evaluation - #732

Merged
jariy17 merged 3 commits into
mainfrom
feat/GT_Support
Mar 30, 2026
Merged

feat: add ground truth reference inputs for on-demand evaluation#732
jariy17 merged 3 commits into
mainfrom
feat/GT_Support

Conversation

@jariy17

Copy link
Copy Markdown
Contributor

Summary

Adds ground truth reference inputs to on-demand evaluation (run eval), allowing users to specify expected agent behavior for more precise evaluation.

New CLI flags

  • -A, --assertion <text...> — Assertion the agent should satisfy (repeatable)
  • --expected-trajectory <names> — Expected tool calls in order (comma-separated)
  • --expected-response <text> — Expected agent response text

TUI changes

  • New Ground Truth wizard step in run eval flow (appears when exactly 1 session is selected)
  • Evaluator creation screen now shows reference input placeholders ({assertions}, {expected_tool_trajectory}, {actual_tool_trajectory}, {expected_response})
  • Updated default evaluator instructions to include GT placeholders

Implementation

  • evaluationReferenceInputs passed through to the evaluate() API call
  • Results include referenceInputs summary in both CLI and TUI output
  • Note: The second commit (temp: use middleware...) is a temporary workaround — the SDK does not yet include evaluationReferenceInputs in its model, so we inject it via Smithy middleware. Remove when SDK is updated.

Test plan

  • npx tsc --noEmit — clean
  • npm run lint — 0 errors
  • 5 new unit tests for GT reference inputs (all pass)
  • Existing 33 run-eval tests continue to pass

@jariy17
jariy17 requested a review from a teamMarch 30, 2026 19:48
@github-actionsgithub-actionsBot added the size/l PR size: L label Mar 30, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Package Tarball

aws-agentcore-0.4.0.tgz

How to install

npm install https://github.com/aws/agentcore-cli/releases/download/pr-732-tarball/aws-agentcore-0.4.0.tgz

@github-actions

github-actionsBot commented Mar 30, 2026

Copy link
Copy Markdown
Contributor

Coverage Report

StatusCategoryPercentageCovered / Total
🔵Lines45.98%6575 / 14297
🔵Statements45.56%6986 / 15333
🔵Functions44.68%1177 / 2634
🔵Branches46.18%4362 / 9444
Generated in workflow #1538 for commit d5e84ae by the Vitest Coverage Report Action

Bump @aws-sdk/client-bedrock-agentcore and
@aws-sdk/client-bedrock-agentcore-control from ^3.893.0 to ^3.1020.0.
The SDK now includes EvaluationReferenceInput in its model,
so we pass it directly to EvaluateCommand instead of injecting
via middleware.
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Mar 30, 2026
Comment threadsrc/cli/aws/agentcore.ts Outdated
@jariy17

Copy link
Copy Markdown
ContributorAuthor

Testing Results

Verified all new GT features against a deployed runtime (TempAgent in EvalCheck project):

CLI flags tested

  • -A, --assertion — single and multiple assertions with Builtin.GoalSuccessRate
  • --expected-trajectory — comma-separated tool names with Builtin.TrajectoryExactOrderMatch
  • --expected-response — with and without explicit -t trace ID, using Builtin.Correctness
  • All three combined in a single run across multiple evaluators
  • Baseline run (no GT inputs) — confirms referenceInputs is absent from output

Custom evaluators tested

  • MyEvaluatorTraceGT (TRACE level) — {expected_response} placeholder populated correctly
  • CustomEvaluatorSessionGT (SESSION level) — {assertions}, {expected_tool_trajectory}, {actual_tool_trajectory} placeholders populated correctly

SDK upgrade verified

  • @aws-sdk/client-bedrock-agentcore@^3.1020.0evaluationReferenceInputs passed natively via EvaluateCommand (no middleware workaround needed)

Other checks

  • npx tsc --noEmit — clean
  • npm run lint — 0 errors
  • All 38 unit tests pass (including 5 new GT tests)
  • TUI run eval wizard shows Ground Truth step when exactly 1 session is selected

@HweinstockHweinstock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some nits and question on service behavior. Otherwise lgtm.

Comment threadsrc/cli/commands/run/command.tsx
Comment threadsrc/cli/commands/run/command.tsx
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/tui/screens/evaluator/types.ts
Hweinstock
Hweinstock previously approved these changes Mar 30, 2026

@HweinstockHweinstock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Slightly confused on the customer experience, but thats likely because I'm new to evals. The changes themselves lgtm.

@awsaws deleted a comment from HweinstockMar 30, 2026
@jariy17
jariy17 merged commit 01623ff into mainMar 30, 2026
22 checks passed
@jariy17
jariy17 deleted the feat/GT_Support branch March 30, 2026 23:03
aidandaly24 added a commit that referenced this pull request Apr 8, 2026
Cover 13 PRs merged to main since the initial docs audit (March 28):
- import subcommands: runtime, memory, evaluator, online-eval (#763, #780)
- invoke/dev exec mode with --exec and --timeout (#750)
- code-based evaluator support with --type, --lambda-arn, --timeout (#739)
- ground truth eval inputs: --assertion, --expected-trajectory, --expected-response (#732)
- memory record streaming: --data-stream-arn, --stream-content-level (#531, #534)
- create --skip-install flag (#782)
- fetch access --identity-name option (#774)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jariy17@Hweinstock@awsswarnim
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat: add ground truth reference inputs for on-demand evaluation - #732

Merged
jariy17 merged 3 commits into
mainfrom
feat/GT_Support
Mar 30, 2026
Merged

feat: add ground truth reference inputs for on-demand evaluation#732
jariy17 merged 3 commits into
mainfrom
feat/GT_Support

Conversation

@jariy17

Copy link
Copy Markdown
Contributor

Summary

Adds ground truth reference inputs to on-demand evaluation (run eval), allowing users to specify expected agent behavior for more precise evaluation.

New CLI flags

  • -A, --assertion <text...> — Assertion the agent should satisfy (repeatable)
  • --expected-trajectory <names> — Expected tool calls in order (comma-separated)
  • --expected-response <text> — Expected agent response text

TUI changes

  • New Ground Truth wizard step in run eval flow (appears when exactly 1 session is selected)
  • Evaluator creation screen now shows reference input placeholders ({assertions}, {expected_tool_trajectory}, {actual_tool_trajectory}, {expected_response})
  • Updated default evaluator instructions to include GT placeholders

Implementation

  • evaluationReferenceInputs passed through to the evaluate() API call
  • Results include referenceInputs summary in both CLI and TUI output
  • Note: The second commit (temp: use middleware...) is a temporary workaround — the SDK does not yet include evaluationReferenceInputs in its model, so we inject it via Smithy middleware. Remove when SDK is updated.

Test plan

  • npx tsc --noEmit — clean
  • npm run lint — 0 errors
  • 5 new unit tests for GT reference inputs (all pass)
  • Existing 33 run-eval tests continue to pass

@jariy17
jariy17 requested a review from a teamMarch 30, 2026 19:48
@github-actionsgithub-actionsBot added the size/l PR size: L label Mar 30, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Package Tarball

aws-agentcore-0.4.0.tgz

How to install

npm install https://github.com/aws/agentcore-cli/releases/download/pr-732-tarball/aws-agentcore-0.4.0.tgz

@github-actions

github-actionsBot commented Mar 30, 2026

Copy link
Copy Markdown
Contributor

Coverage Report

StatusCategoryPercentageCovered / Total
🔵Lines45.98%6575 / 14297
🔵Statements45.56%6986 / 15333
🔵Functions44.68%1177 / 2634
🔵Branches46.18%4362 / 9444
Generated in workflow #1538 for commit d5e84ae by the Vitest Coverage Report Action

Bump @aws-sdk/client-bedrock-agentcore and
@aws-sdk/client-bedrock-agentcore-control from ^3.893.0 to ^3.1020.0.
The SDK now includes EvaluationReferenceInput in its model,
so we pass it directly to EvaluateCommand instead of injecting
via middleware.
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Mar 30, 2026
Comment threadsrc/cli/aws/agentcore.ts Outdated
@jariy17

Copy link
Copy Markdown
ContributorAuthor

Testing Results

Verified all new GT features against a deployed runtime (TempAgent in EvalCheck project):

CLI flags tested

  • -A, --assertion — single and multiple assertions with Builtin.GoalSuccessRate
  • --expected-trajectory — comma-separated tool names with Builtin.TrajectoryExactOrderMatch
  • --expected-response — with and without explicit -t trace ID, using Builtin.Correctness
  • All three combined in a single run across multiple evaluators
  • Baseline run (no GT inputs) — confirms referenceInputs is absent from output

Custom evaluators tested

  • MyEvaluatorTraceGT (TRACE level) — {expected_response} placeholder populated correctly
  • CustomEvaluatorSessionGT (SESSION level) — {assertions}, {expected_tool_trajectory}, {actual_tool_trajectory} placeholders populated correctly

SDK upgrade verified

  • @aws-sdk/client-bedrock-agentcore@^3.1020.0evaluationReferenceInputs passed natively via EvaluateCommand (no middleware workaround needed)

Other checks

  • npx tsc --noEmit — clean
  • npm run lint — 0 errors
  • All 38 unit tests pass (including 5 new GT tests)
  • TUI run eval wizard shows Ground Truth step when exactly 1 session is selected

@HweinstockHweinstock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some nits and question on service behavior. Otherwise lgtm.

Comment threadsrc/cli/commands/run/command.tsx
Comment threadsrc/cli/commands/run/command.tsx
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/tui/screens/evaluator/types.ts
Hweinstock
Hweinstock previously approved these changes Mar 30, 2026

@HweinstockHweinstock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Slightly confused on the customer experience, but thats likely because I'm new to evals. The changes themselves lgtm.

@awsaws deleted a comment from HweinstockMar 30, 2026
@jariy17
jariy17 merged commit 01623ff into mainMar 30, 2026
22 checks passed
@jariy17
jariy17 deleted the feat/GT_Support branch March 30, 2026 23:03
aidandaly24 added a commit that referenced this pull request Apr 8, 2026
Cover 13 PRs merged to main since the initial docs audit (March 28):
- import subcommands: runtime, memory, evaluator, online-eval (#763, #780)
- invoke/dev exec mode with --exec and --timeout (#750)
- code-based evaluator support with --type, --lambda-arn, --timeout (#739)
- ground truth eval inputs: --assertion, --expected-trajectory, --expected-response (#732)
- memory record streaming: --data-stream-arn, --stream-content-level (#531, #534)
- create --skip-install flag (#782)
- fetch access --identity-name option (#774)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jariy17@Hweinstock@awsswarnim
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: add ground truth reference inputs for on-demand evaluation - #732

Merged
jariy17 merged 3 commits into
mainfrom
feat/GT_Support
Mar 30, 2026
Merged

feat: add ground truth reference inputs for on-demand evaluation#732
jariy17 merged 3 commits into
mainfrom
feat/GT_Support

Conversation

@jariy17

Copy link
Copy Markdown
Contributor

Summary

Adds ground truth reference inputs to on-demand evaluation (run eval), allowing users to specify expected agent behavior for more precise evaluation.

New CLI flags

  • -A, --assertion <text...> — Assertion the agent should satisfy (repeatable)
  • --expected-trajectory <names> — Expected tool calls in order (comma-separated)
  • --expected-response <text> — Expected agent response text

TUI changes

  • New Ground Truth wizard step in run eval flow (appears when exactly 1 session is selected)
  • Evaluator creation screen now shows reference input placeholders ({assertions}, {expected_tool_trajectory}, {actual_tool_trajectory}, {expected_response})
  • Updated default evaluator instructions to include GT placeholders

Implementation

  • evaluationReferenceInputs passed through to the evaluate() API call
  • Results include referenceInputs summary in both CLI and TUI output
  • Note: The second commit (temp: use middleware...) is a temporary workaround — the SDK does not yet include evaluationReferenceInputs in its model, so we inject it via Smithy middleware. Remove when SDK is updated.

Test plan

  • npx tsc --noEmit — clean
  • npm run lint — 0 errors
  • 5 new unit tests for GT reference inputs (all pass)
  • Existing 33 run-eval tests continue to pass

@jariy17
jariy17 requested a review from a teamMarch 30, 2026 19:48
@github-actionsgithub-actionsBot added the size/l PR size: L label Mar 30, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Package Tarball

aws-agentcore-0.4.0.tgz

How to install

npm install https://github.com/aws/agentcore-cli/releases/download/pr-732-tarball/aws-agentcore-0.4.0.tgz

@github-actions

github-actionsBot commented Mar 30, 2026

Copy link
Copy Markdown
Contributor

Coverage Report

StatusCategoryPercentageCovered / Total
🔵Lines45.98%6575 / 14297
🔵Statements45.56%6986 / 15333
🔵Functions44.68%1177 / 2634
🔵Branches46.18%4362 / 9444
Generated in workflow #1538 for commit d5e84ae by the Vitest Coverage Report Action

Bump @aws-sdk/client-bedrock-agentcore and
@aws-sdk/client-bedrock-agentcore-control from ^3.893.0 to ^3.1020.0.
The SDK now includes EvaluationReferenceInput in its model,
so we pass it directly to EvaluateCommand instead of injecting
via middleware.
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Mar 30, 2026
Comment threadsrc/cli/aws/agentcore.ts Outdated
@jariy17

Copy link
Copy Markdown
ContributorAuthor

Testing Results

Verified all new GT features against a deployed runtime (TempAgent in EvalCheck project):

CLI flags tested

  • -A, --assertion — single and multiple assertions with Builtin.GoalSuccessRate
  • --expected-trajectory — comma-separated tool names with Builtin.TrajectoryExactOrderMatch
  • --expected-response — with and without explicit -t trace ID, using Builtin.Correctness
  • All three combined in a single run across multiple evaluators
  • Baseline run (no GT inputs) — confirms referenceInputs is absent from output

Custom evaluators tested

  • MyEvaluatorTraceGT (TRACE level) — {expected_response} placeholder populated correctly
  • CustomEvaluatorSessionGT (SESSION level) — {assertions}, {expected_tool_trajectory}, {actual_tool_trajectory} placeholders populated correctly

SDK upgrade verified

  • @aws-sdk/client-bedrock-agentcore@^3.1020.0evaluationReferenceInputs passed natively via EvaluateCommand (no middleware workaround needed)

Other checks

  • npx tsc --noEmit — clean
  • npm run lint — 0 errors
  • All 38 unit tests pass (including 5 new GT tests)
  • TUI run eval wizard shows Ground Truth step when exactly 1 session is selected

@HweinstockHweinstock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some nits and question on service behavior. Otherwise lgtm.

Comment threadsrc/cli/commands/run/command.tsx
Comment threadsrc/cli/commands/run/command.tsx
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/tui/screens/evaluator/types.ts
Hweinstock
Hweinstock previously approved these changes Mar 30, 2026

@HweinstockHweinstock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Slightly confused on the customer experience, but thats likely because I'm new to evals. The changes themselves lgtm.

@awsaws deleted a comment from HweinstockMar 30, 2026
@jariy17
jariy17 merged commit 01623ff into mainMar 30, 2026
22 checks passed
@jariy17
jariy17 deleted the feat/GT_Support branch March 30, 2026 23:03
aidandaly24 added a commit that referenced this pull request Apr 8, 2026
Cover 13 PRs merged to main since the initial docs audit (March 28):
- import subcommands: runtime, memory, evaluator, online-eval (#763, #780)
- invoke/dev exec mode with --exec and --timeout (#750)
- code-based evaluator support with --type, --lambda-arn, --timeout (#739)
- ground truth eval inputs: --assertion, --expected-trajectory, --expected-response (#732)
- memory record streaming: --data-stream-arn, --stream-content-level (#531, #534)
- create --skip-install flag (#782)
- fetch access --identity-name option (#774)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jariy17@Hweinstock@awsswarnim
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: add ground truth reference inputs for on-demand evaluation - #732

Merged
jariy17 merged 3 commits into
mainfrom
feat/GT_Support
Mar 30, 2026
Merged

feat: add ground truth reference inputs for on-demand evaluation#732
jariy17 merged 3 commits into
mainfrom
feat/GT_Support

Conversation

@jariy17

Copy link
Copy Markdown
Contributor

Summary

Adds ground truth reference inputs to on-demand evaluation (run eval), allowing users to specify expected agent behavior for more precise evaluation.

New CLI flags

  • -A, --assertion <text...> — Assertion the agent should satisfy (repeatable)
  • --expected-trajectory <names> — Expected tool calls in order (comma-separated)
  • --expected-response <text> — Expected agent response text

TUI changes

  • New Ground Truth wizard step in run eval flow (appears when exactly 1 session is selected)
  • Evaluator creation screen now shows reference input placeholders ({assertions}, {expected_tool_trajectory}, {actual_tool_trajectory}, {expected_response})
  • Updated default evaluator instructions to include GT placeholders

Implementation

  • evaluationReferenceInputs passed through to the evaluate() API call
  • Results include referenceInputs summary in both CLI and TUI output
  • Note: The second commit (temp: use middleware...) is a temporary workaround — the SDK does not yet include evaluationReferenceInputs in its model, so we inject it via Smithy middleware. Remove when SDK is updated.

Test plan

  • npx tsc --noEmit — clean
  • npm run lint — 0 errors
  • 5 new unit tests for GT reference inputs (all pass)
  • Existing 33 run-eval tests continue to pass

@jariy17
jariy17 requested a review from a teamMarch 30, 2026 19:48
@github-actionsgithub-actionsBot added the size/l PR size: L label Mar 30, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Package Tarball

aws-agentcore-0.4.0.tgz

How to install

npm install https://github.com/aws/agentcore-cli/releases/download/pr-732-tarball/aws-agentcore-0.4.0.tgz

@github-actions

github-actionsBot commented Mar 30, 2026

Copy link
Copy Markdown
Contributor

Coverage Report

StatusCategoryPercentageCovered / Total
🔵Lines45.98%6575 / 14297
🔵Statements45.56%6986 / 15333
🔵Functions44.68%1177 / 2634
🔵Branches46.18%4362 / 9444
Generated in workflow #1538 for commit d5e84ae by the Vitest Coverage Report Action

Bump @aws-sdk/client-bedrock-agentcore and
@aws-sdk/client-bedrock-agentcore-control from ^3.893.0 to ^3.1020.0.
The SDK now includes EvaluationReferenceInput in its model,
so we pass it directly to EvaluateCommand instead of injecting
via middleware.
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Mar 30, 2026
Comment threadsrc/cli/aws/agentcore.ts Outdated
@jariy17

Copy link
Copy Markdown
ContributorAuthor

Testing Results

Verified all new GT features against a deployed runtime (TempAgent in EvalCheck project):

CLI flags tested

  • -A, --assertion — single and multiple assertions with Builtin.GoalSuccessRate
  • --expected-trajectory — comma-separated tool names with Builtin.TrajectoryExactOrderMatch
  • --expected-response — with and without explicit -t trace ID, using Builtin.Correctness
  • All three combined in a single run across multiple evaluators
  • Baseline run (no GT inputs) — confirms referenceInputs is absent from output

Custom evaluators tested

  • MyEvaluatorTraceGT (TRACE level) — {expected_response} placeholder populated correctly
  • CustomEvaluatorSessionGT (SESSION level) — {assertions}, {expected_tool_trajectory}, {actual_tool_trajectory} placeholders populated correctly

SDK upgrade verified

  • @aws-sdk/client-bedrock-agentcore@^3.1020.0evaluationReferenceInputs passed natively via EvaluateCommand (no middleware workaround needed)

Other checks

  • npx tsc --noEmit — clean
  • npm run lint — 0 errors
  • All 38 unit tests pass (including 5 new GT tests)
  • TUI run eval wizard shows Ground Truth step when exactly 1 session is selected

@HweinstockHweinstock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some nits and question on service behavior. Otherwise lgtm.

Comment threadsrc/cli/commands/run/command.tsx
Comment threadsrc/cli/commands/run/command.tsx
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/tui/screens/evaluator/types.ts
Hweinstock
Hweinstock previously approved these changes Mar 30, 2026

@HweinstockHweinstock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Slightly confused on the customer experience, but thats likely because I'm new to evals. The changes themselves lgtm.

@awsaws deleted a comment from HweinstockMar 30, 2026
@jariy17
jariy17 merged commit 01623ff into mainMar 30, 2026
22 checks passed
@jariy17
jariy17 deleted the feat/GT_Support branch March 30, 2026 23:03
aidandaly24 added a commit that referenced this pull request Apr 8, 2026
Cover 13 PRs merged to main since the initial docs audit (March 28):
- import subcommands: runtime, memory, evaluator, online-eval (#763, #780)
- invoke/dev exec mode with --exec and --timeout (#750)
- code-based evaluator support with --type, --lambda-arn, --timeout (#739)
- ground truth eval inputs: --assertion, --expected-trajectory, --expected-response (#732)
- memory record streaming: --data-stream-arn, --stream-content-level (#531, #534)
- create --skip-install flag (#782)
- fetch access --identity-name option (#774)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jariy17@Hweinstock@awsswarnim
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat: add ground truth reference inputs for on-demand evaluation - #732

Merged
jariy17 merged 3 commits into
mainfrom
feat/GT_Support
Mar 30, 2026
Merged

feat: add ground truth reference inputs for on-demand evaluation#732
jariy17 merged 3 commits into
mainfrom
feat/GT_Support

Conversation

@jariy17

Copy link
Copy Markdown
Contributor

Summary

Adds ground truth reference inputs to on-demand evaluation (run eval), allowing users to specify expected agent behavior for more precise evaluation.

New CLI flags

  • -A, --assertion <text...> — Assertion the agent should satisfy (repeatable)
  • --expected-trajectory <names> — Expected tool calls in order (comma-separated)
  • --expected-response <text> — Expected agent response text

TUI changes

  • New Ground Truth wizard step in run eval flow (appears when exactly 1 session is selected)
  • Evaluator creation screen now shows reference input placeholders ({assertions}, {expected_tool_trajectory}, {actual_tool_trajectory}, {expected_response})
  • Updated default evaluator instructions to include GT placeholders

Implementation

  • evaluationReferenceInputs passed through to the evaluate() API call
  • Results include referenceInputs summary in both CLI and TUI output
  • Note: The second commit (temp: use middleware...) is a temporary workaround — the SDK does not yet include evaluationReferenceInputs in its model, so we inject it via Smithy middleware. Remove when SDK is updated.

Test plan

  • npx tsc --noEmit — clean
  • npm run lint — 0 errors
  • 5 new unit tests for GT reference inputs (all pass)
  • Existing 33 run-eval tests continue to pass

@jariy17
jariy17 requested a review from a teamMarch 30, 2026 19:48
@github-actionsgithub-actionsBot added the size/l PR size: L label Mar 30, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Package Tarball

aws-agentcore-0.4.0.tgz

How to install

npm install https://github.com/aws/agentcore-cli/releases/download/pr-732-tarball/aws-agentcore-0.4.0.tgz

@github-actions

github-actionsBot commented Mar 30, 2026

Copy link
Copy Markdown
Contributor

Coverage Report

StatusCategoryPercentageCovered / Total
🔵Lines45.98%6575 / 14297
🔵Statements45.56%6986 / 15333
🔵Functions44.68%1177 / 2634
🔵Branches46.18%4362 / 9444
Generated in workflow #1538 for commit d5e84ae by the Vitest Coverage Report Action

Bump @aws-sdk/client-bedrock-agentcore and
@aws-sdk/client-bedrock-agentcore-control from ^3.893.0 to ^3.1020.0.
The SDK now includes EvaluationReferenceInput in its model,
so we pass it directly to EvaluateCommand instead of injecting
via middleware.
@github-actionsgithub-actionsBot added size/l PR size: L and removed size/l PR size: L labels Mar 30, 2026
Comment threadsrc/cli/aws/agentcore.ts Outdated
@jariy17

Copy link
Copy Markdown
ContributorAuthor

Testing Results

Verified all new GT features against a deployed runtime (TempAgent in EvalCheck project):

CLI flags tested

  • -A, --assertion — single and multiple assertions with Builtin.GoalSuccessRate
  • --expected-trajectory — comma-separated tool names with Builtin.TrajectoryExactOrderMatch
  • --expected-response — with and without explicit -t trace ID, using Builtin.Correctness
  • All three combined in a single run across multiple evaluators
  • Baseline run (no GT inputs) — confirms referenceInputs is absent from output

Custom evaluators tested

  • MyEvaluatorTraceGT (TRACE level) — {expected_response} placeholder populated correctly
  • CustomEvaluatorSessionGT (SESSION level) — {assertions}, {expected_tool_trajectory}, {actual_tool_trajectory} placeholders populated correctly

SDK upgrade verified

  • @aws-sdk/client-bedrock-agentcore@^3.1020.0evaluationReferenceInputs passed natively via EvaluateCommand (no middleware workaround needed)

Other checks

  • npx tsc --noEmit — clean
  • npm run lint — 0 errors
  • All 38 unit tests pass (including 5 new GT tests)
  • TUI run eval wizard shows Ground Truth step when exactly 1 session is selected

@HweinstockHweinstock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some nits and question on service behavior. Otherwise lgtm.

Comment threadsrc/cli/commands/run/command.tsx
Comment threadsrc/cli/commands/run/command.tsx
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/operations/eval/run-eval.ts
Comment threadsrc/cli/tui/screens/evaluator/types.ts
Hweinstock
Hweinstock previously approved these changes Mar 30, 2026

@HweinstockHweinstock left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Slightly confused on the customer experience, but thats likely because I'm new to evals. The changes themselves lgtm.

@awsaws deleted a comment from HweinstockMar 30, 2026
@jariy17
jariy17 merged commit 01623ff into mainMar 30, 2026
22 checks passed
@jariy17
jariy17 deleted the feat/GT_Support branch March 30, 2026 23:03
aidandaly24 added a commit that referenced this pull request Apr 8, 2026
Cover 13 PRs merged to main since the initial docs audit (March 28):
- import subcommands: runtime, memory, evaluator, online-eval (#763, #780)
- invoke/dev exec mode with --exec and --timeout (#750)
- code-based evaluator support with --type, --lambda-arn, --timeout (#739)
- ground truth eval inputs: --assertion, --expected-trajectory, --expected-response (#732)
- memory record streaming: --data-stream-arn, --stream-content-level (#531, #534)
- create --skip-install flag (#782)
- fetch access --identity-name option (#774)
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@jariy17@Hweinstock@awsswarnim