feat: support imperative eval recommendation command - #2111

Merged
jariy17 merged 3 commits into
refactorfrom
eval-recommendation
Aug 27, 2026
Merged

feat: support imperative eval recommendation command#2111
jariy17 merged 3 commits into
refactorfrom
eval-recommendation

Conversation

@nborges-aws

Copy link
Copy Markdown
Contributor

Description

Adds imperative command-line support for AgentCore evaluation recommendations:

  • eval recommendation start
  • eval recommendation get
  • eval recommendation list
  • eval recommendation delete

Summary of Changes

  • Supports SYSTEM_PROMPT_RECOMMENDATION and TOOL_DESCRIPTION_RECOMMENDATION recommendation types
  • start accepts recommendation configuration as inline JSON, a file:// path, or stdin (-) through SourceResolver class
  • Supports the service API fields for description, KMS key ARN, and tags
  • Adds the recommendation operations to EvalClient, the shared handler types, and TestCoreClient
  • Adds recorded golden fixtures covering the complete start, get, list, and delete flow

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Documentation update
  • Other (please describe):

Testing

How have you tested the change?

Added unit tests and golden fixture tests exercising recommendation logic.

  • bun run test (2052 pass, 0 fail)
  • I ran npm run test:unit and npm run test:integ
  • I ran npm run typecheck
  • I ran npm run lint
  • If I modified src/assets/, I ran npm run test:update-snapshots and committed the updated snapshots

Checklist

  • I have read the CONTRIBUTING document
  • I have added any necessary tests that prove my fix is effective or my feature works
  • I have updated the documentation accordingly
  • I have added an appropriate example to the documentation to outline the feature, or no new docs are needed
  • My changes generate no new warnings
  • Any dependent changes have been merged and published

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the
terms of your choice.

description: "get a recommendation by id",
flags: [flag("id", "the ID of the recommendation", z.string().optional())],
handle: async (ctx, flags) => {
if (!flags["id"]) throw new InputValidationError("required option '--id <id>' not specified");

@jariy17jariy17Aug 26, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can definitely abstracted this so we have a common error message for this type of thing. OBO

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah I think marking flags as required and supporting it directly in the framework would be ideal, but likely OOS here.

Comment threadsrc/core/eval.tsx Outdated
}

async startRecommendation(
request: StartRecommendationRequest,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We create our own types here because handlers define the interface for the core client. Even if it's a one-to-one mapping, we should still define our own type.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sure will update

Comment threadsrc/core/recommendation.test.tsx Outdated
@@ -0,0 +1,119 @@
import { describe, expect, test } from "bun:test";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't need this. Use Golden tests please

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are the handler tests sufficient? I think these are testing the same things at a lower level.

flag("tags", "tags as key=value (repeatable) or JSON object", z.array(z.string()).optional()),
],
handle: async (ctx, flags) => {
if (!flags["name"]) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

const requiredFlags: Array<[string, string]> = [
["name", "--name <name>"],
["type", "--type <type>"],
["recommendation-config", "--recommendation-config <recommendation-config>"],
];
for (const [key, usage] of requiredFlags) {
if (!flags[key]) {
throw new InputValidationError(`required option '${usage}' not specified`);
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: we should use something like this.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure will add something like this

return call.args;
}

describe("eval recommendation command hierarchy", () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need these tests?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, I've seen these tests in other places, but they feel pretty low value. The tests on the handlers would fail if the hierarchy is incorrect. Also, what coverage do we get on the tests below that we don't get on the fixture based tests?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added these tests to cover behavior the fixtures don't exercise. E.g. malformed json, exercising source resolver, etc. But consensus is against unit test so I've removed this + the other one entirely

description: "get a recommendation by id",
flags: [flag("id", "the ID of the recommendation", z.string().optional())],
handle: async (ctx, flags) => {
if (!flags["id"]) throw new InputValidationError("required option '--id <id>' not specified");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah I think marking flags as required and supporting it directly in the framework would be ideal, but likely OOS here.

return (error as Error).name === "ResourceNotFoundException";
}

async function waitForTerminal(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is there a way to leverage

exportasyncfunctionwaitFor(
for these?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good call out I'll switch to use this util

return call.args;
}

describe("eval recommendation command hierarchy", () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, I've seen these tests in other places, but they feel pretty low value. The tests on the handlers would fail if the hierarchy is incorrect. Also, what coverage do we get on the tests below that we don't get on the fixture based tests?

status: "PENDING",
createdAt: new Date("2026-08-26T12:00:00.000Z"),
updatedAt: new Date("2026-08-26T12:00:00.000Z"),
} as StartRecommendationResponse;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why do we need to cast here?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

N/A since test file was removed altogether

Comment threadsrc/core/recommendation.test.tsx Outdated
@@ -0,0 +1,119 @@
import { describe, expect, test } from "bun:test";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are the handler tests sufficient? I think these are testing the same things at a lower level.

@github-actionsgithub-actionsBot added the size/l PR size: L label Aug 27, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.70115% with 4 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.41%. Comparing base (a1f2f32) to head (2783292).

Files with missing linesPatch %Lines
src/handlers/eval/recommendation/start/index.tsx94.20%4 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #2111 +/- ##
==========================================
Coverage 97.41% 97.41% ==========================================
Files 453 458 +5 Lines 27637 27810 +173 ==========================================
+ Hits 26922 27091 +169 - Misses 715 719 +4 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@jariy17
jariy17 merged commit c105d4f into refactorAug 27, 2026
22 checks passed
@jariy17
jariy17 deleted the eval-recommendation branch August 27, 2026 16:29
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@nborges-aws@codecov-commenter@Hweinstock@jariy17
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat: support imperative eval recommendation command - #2111

Merged
jariy17 merged 3 commits into
refactorfrom
eval-recommendation
Aug 27, 2026
Merged

feat: support imperative eval recommendation command#2111
jariy17 merged 3 commits into
refactorfrom
eval-recommendation

Conversation

@nborges-aws

Copy link
Copy Markdown
Contributor

Description

Adds imperative command-line support for AgentCore evaluation recommendations:

  • eval recommendation start
  • eval recommendation get
  • eval recommendation list
  • eval recommendation delete

Summary of Changes

  • Supports SYSTEM_PROMPT_RECOMMENDATION and TOOL_DESCRIPTION_RECOMMENDATION recommendation types
  • start accepts recommendation configuration as inline JSON, a file:// path, or stdin (-) through SourceResolver class
  • Supports the service API fields for description, KMS key ARN, and tags
  • Adds the recommendation operations to EvalClient, the shared handler types, and TestCoreClient
  • Adds recorded golden fixtures covering the complete start, get, list, and delete flow

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Documentation update
  • Other (please describe):

Testing

How have you tested the change?

Added unit tests and golden fixture tests exercising recommendation logic.

  • bun run test (2052 pass, 0 fail)
  • I ran npm run test:unit and npm run test:integ
  • I ran npm run typecheck
  • I ran npm run lint
  • If I modified src/assets/, I ran npm run test:update-snapshots and committed the updated snapshots

Checklist

  • I have read the CONTRIBUTING document
  • I have added any necessary tests that prove my fix is effective or my feature works
  • I have updated the documentation accordingly
  • I have added an appropriate example to the documentation to outline the feature, or no new docs are needed
  • My changes generate no new warnings
  • Any dependent changes have been merged and published

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the
terms of your choice.

description: "get a recommendation by id",
flags: [flag("id", "the ID of the recommendation", z.string().optional())],
handle: async (ctx, flags) => {
if (!flags["id"]) throw new InputValidationError("required option '--id <id>' not specified");

@jariy17jariy17Aug 26, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can definitely abstracted this so we have a common error message for this type of thing. OBO

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah I think marking flags as required and supporting it directly in the framework would be ideal, but likely OOS here.

Comment threadsrc/core/eval.tsx Outdated
}

async startRecommendation(
request: StartRecommendationRequest,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We create our own types here because handlers define the interface for the core client. Even if it's a one-to-one mapping, we should still define our own type.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sure will update

Comment threadsrc/core/recommendation.test.tsx Outdated
@@ -0,0 +1,119 @@
import { describe, expect, test } from "bun:test";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't need this. Use Golden tests please

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are the handler tests sufficient? I think these are testing the same things at a lower level.

flag("tags", "tags as key=value (repeatable) or JSON object", z.array(z.string()).optional()),
],
handle: async (ctx, flags) => {
if (!flags["name"]) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

const requiredFlags: Array<[string, string]> = [
["name", "--name <name>"],
["type", "--type <type>"],
["recommendation-config", "--recommendation-config <recommendation-config>"],
];
for (const [key, usage] of requiredFlags) {
if (!flags[key]) {
throw new InputValidationError(`required option '${usage}' not specified`);
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: we should use something like this.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure will add something like this

return call.args;
}

describe("eval recommendation command hierarchy", () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need these tests?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, I've seen these tests in other places, but they feel pretty low value. The tests on the handlers would fail if the hierarchy is incorrect. Also, what coverage do we get on the tests below that we don't get on the fixture based tests?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added these tests to cover behavior the fixtures don't exercise. E.g. malformed json, exercising source resolver, etc. But consensus is against unit test so I've removed this + the other one entirely

description: "get a recommendation by id",
flags: [flag("id", "the ID of the recommendation", z.string().optional())],
handle: async (ctx, flags) => {
if (!flags["id"]) throw new InputValidationError("required option '--id <id>' not specified");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah I think marking flags as required and supporting it directly in the framework would be ideal, but likely OOS here.

return (error as Error).name === "ResourceNotFoundException";
}

async function waitForTerminal(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is there a way to leverage

exportasyncfunctionwaitFor(
for these?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good call out I'll switch to use this util

return call.args;
}

describe("eval recommendation command hierarchy", () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, I've seen these tests in other places, but they feel pretty low value. The tests on the handlers would fail if the hierarchy is incorrect. Also, what coverage do we get on the tests below that we don't get on the fixture based tests?

status: "PENDING",
createdAt: new Date("2026-08-26T12:00:00.000Z"),
updatedAt: new Date("2026-08-26T12:00:00.000Z"),
} as StartRecommendationResponse;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why do we need to cast here?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

N/A since test file was removed altogether

Comment threadsrc/core/recommendation.test.tsx Outdated
@@ -0,0 +1,119 @@
import { describe, expect, test } from "bun:test";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are the handler tests sufficient? I think these are testing the same things at a lower level.

@github-actionsgithub-actionsBot added the size/l PR size: L label Aug 27, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.70115% with 4 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.41%. Comparing base (a1f2f32) to head (2783292).

Files with missing linesPatch %Lines
src/handlers/eval/recommendation/start/index.tsx94.20%4 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #2111 +/- ##
==========================================
Coverage 97.41% 97.41% ==========================================
Files 453 458 +5 Lines 27637 27810 +173 ==========================================
+ Hits 26922 27091 +169 - Misses 715 719 +4 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@jariy17
jariy17 merged commit c105d4f into refactorAug 27, 2026
22 checks passed
@jariy17
jariy17 deleted the eval-recommendation branch August 27, 2026 16:29
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@nborges-aws@codecov-commenter@Hweinstock@jariy17
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: support imperative eval recommendation command - #2111

Merged
jariy17 merged 3 commits into
refactorfrom
eval-recommendation
Aug 27, 2026
Merged

feat: support imperative eval recommendation command#2111
jariy17 merged 3 commits into
refactorfrom
eval-recommendation

Conversation

@nborges-aws

Copy link
Copy Markdown
Contributor

Description

Adds imperative command-line support for AgentCore evaluation recommendations:

  • eval recommendation start
  • eval recommendation get
  • eval recommendation list
  • eval recommendation delete

Summary of Changes

  • Supports SYSTEM_PROMPT_RECOMMENDATION and TOOL_DESCRIPTION_RECOMMENDATION recommendation types
  • start accepts recommendation configuration as inline JSON, a file:// path, or stdin (-) through SourceResolver class
  • Supports the service API fields for description, KMS key ARN, and tags
  • Adds the recommendation operations to EvalClient, the shared handler types, and TestCoreClient
  • Adds recorded golden fixtures covering the complete start, get, list, and delete flow

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Documentation update
  • Other (please describe):

Testing

How have you tested the change?

Added unit tests and golden fixture tests exercising recommendation logic.

  • bun run test (2052 pass, 0 fail)
  • I ran npm run test:unit and npm run test:integ
  • I ran npm run typecheck
  • I ran npm run lint
  • If I modified src/assets/, I ran npm run test:update-snapshots and committed the updated snapshots

Checklist

  • I have read the CONTRIBUTING document
  • I have added any necessary tests that prove my fix is effective or my feature works
  • I have updated the documentation accordingly
  • I have added an appropriate example to the documentation to outline the feature, or no new docs are needed
  • My changes generate no new warnings
  • Any dependent changes have been merged and published

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the
terms of your choice.

description: "get a recommendation by id",
flags: [flag("id", "the ID of the recommendation", z.string().optional())],
handle: async (ctx, flags) => {
if (!flags["id"]) throw new InputValidationError("required option '--id <id>' not specified");

@jariy17jariy17Aug 26, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can definitely abstracted this so we have a common error message for this type of thing. OBO

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah I think marking flags as required and supporting it directly in the framework would be ideal, but likely OOS here.

Comment threadsrc/core/eval.tsx Outdated
}

async startRecommendation(
request: StartRecommendationRequest,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We create our own types here because handlers define the interface for the core client. Even if it's a one-to-one mapping, we should still define our own type.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sure will update

Comment threadsrc/core/recommendation.test.tsx Outdated
@@ -0,0 +1,119 @@
import { describe, expect, test } from "bun:test";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't need this. Use Golden tests please

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are the handler tests sufficient? I think these are testing the same things at a lower level.

flag("tags", "tags as key=value (repeatable) or JSON object", z.array(z.string()).optional()),
],
handle: async (ctx, flags) => {
if (!flags["name"]) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

const requiredFlags: Array<[string, string]> = [
["name", "--name <name>"],
["type", "--type <type>"],
["recommendation-config", "--recommendation-config <recommendation-config>"],
];
for (const [key, usage] of requiredFlags) {
if (!flags[key]) {
throw new InputValidationError(`required option '${usage}' not specified`);
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: we should use something like this.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure will add something like this

return call.args;
}

describe("eval recommendation command hierarchy", () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need these tests?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, I've seen these tests in other places, but they feel pretty low value. The tests on the handlers would fail if the hierarchy is incorrect. Also, what coverage do we get on the tests below that we don't get on the fixture based tests?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added these tests to cover behavior the fixtures don't exercise. E.g. malformed json, exercising source resolver, etc. But consensus is against unit test so I've removed this + the other one entirely

description: "get a recommendation by id",
flags: [flag("id", "the ID of the recommendation", z.string().optional())],
handle: async (ctx, flags) => {
if (!flags["id"]) throw new InputValidationError("required option '--id <id>' not specified");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah I think marking flags as required and supporting it directly in the framework would be ideal, but likely OOS here.

return (error as Error).name === "ResourceNotFoundException";
}

async function waitForTerminal(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is there a way to leverage

exportasyncfunctionwaitFor(
for these?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good call out I'll switch to use this util

return call.args;
}

describe("eval recommendation command hierarchy", () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, I've seen these tests in other places, but they feel pretty low value. The tests on the handlers would fail if the hierarchy is incorrect. Also, what coverage do we get on the tests below that we don't get on the fixture based tests?

status: "PENDING",
createdAt: new Date("2026-08-26T12:00:00.000Z"),
updatedAt: new Date("2026-08-26T12:00:00.000Z"),
} as StartRecommendationResponse;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why do we need to cast here?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

N/A since test file was removed altogether

Comment threadsrc/core/recommendation.test.tsx Outdated
@@ -0,0 +1,119 @@
import { describe, expect, test } from "bun:test";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are the handler tests sufficient? I think these are testing the same things at a lower level.

@github-actionsgithub-actionsBot added the size/l PR size: L label Aug 27, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.70115% with 4 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.41%. Comparing base (a1f2f32) to head (2783292).

Files with missing linesPatch %Lines
src/handlers/eval/recommendation/start/index.tsx94.20%4 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #2111 +/- ##
==========================================
Coverage 97.41% 97.41% ==========================================
Files 453 458 +5 Lines 27637 27810 +173 ==========================================
+ Hits 26922 27091 +169 - Misses 715 719 +4 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@jariy17
jariy17 merged commit c105d4f into refactorAug 27, 2026
22 checks passed
@jariy17
jariy17 deleted the eval-recommendation branch August 27, 2026 16:29
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@nborges-aws@codecov-commenter@Hweinstock@jariy17
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: support imperative eval recommendation command - #2111

Merged
jariy17 merged 3 commits into
refactorfrom
eval-recommendation
Aug 27, 2026
Merged

feat: support imperative eval recommendation command#2111
jariy17 merged 3 commits into
refactorfrom
eval-recommendation

Conversation

@nborges-aws

Copy link
Copy Markdown
Contributor

Description

Adds imperative command-line support for AgentCore evaluation recommendations:

  • eval recommendation start
  • eval recommendation get
  • eval recommendation list
  • eval recommendation delete

Summary of Changes

  • Supports SYSTEM_PROMPT_RECOMMENDATION and TOOL_DESCRIPTION_RECOMMENDATION recommendation types
  • start accepts recommendation configuration as inline JSON, a file:// path, or stdin (-) through SourceResolver class
  • Supports the service API fields for description, KMS key ARN, and tags
  • Adds the recommendation operations to EvalClient, the shared handler types, and TestCoreClient
  • Adds recorded golden fixtures covering the complete start, get, list, and delete flow

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Documentation update
  • Other (please describe):

Testing

How have you tested the change?

Added unit tests and golden fixture tests exercising recommendation logic.

  • bun run test (2052 pass, 0 fail)
  • I ran npm run test:unit and npm run test:integ
  • I ran npm run typecheck
  • I ran npm run lint
  • If I modified src/assets/, I ran npm run test:update-snapshots and committed the updated snapshots

Checklist

  • I have read the CONTRIBUTING document
  • I have added any necessary tests that prove my fix is effective or my feature works
  • I have updated the documentation accordingly
  • I have added an appropriate example to the documentation to outline the feature, or no new docs are needed
  • My changes generate no new warnings
  • Any dependent changes have been merged and published

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the
terms of your choice.

description: "get a recommendation by id",
flags: [flag("id", "the ID of the recommendation", z.string().optional())],
handle: async (ctx, flags) => {
if (!flags["id"]) throw new InputValidationError("required option '--id <id>' not specified");

@jariy17jariy17Aug 26, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can definitely abstracted this so we have a common error message for this type of thing. OBO

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah I think marking flags as required and supporting it directly in the framework would be ideal, but likely OOS here.

Comment threadsrc/core/eval.tsx Outdated
}

async startRecommendation(
request: StartRecommendationRequest,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We create our own types here because handlers define the interface for the core client. Even if it's a one-to-one mapping, we should still define our own type.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sure will update

Comment threadsrc/core/recommendation.test.tsx Outdated
@@ -0,0 +1,119 @@
import { describe, expect, test } from "bun:test";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't need this. Use Golden tests please

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are the handler tests sufficient? I think these are testing the same things at a lower level.

flag("tags", "tags as key=value (repeatable) or JSON object", z.array(z.string()).optional()),
],
handle: async (ctx, flags) => {
if (!flags["name"]) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

const requiredFlags: Array<[string, string]> = [
["name", "--name <name>"],
["type", "--type <type>"],
["recommendation-config", "--recommendation-config <recommendation-config>"],
];
for (const [key, usage] of requiredFlags) {
if (!flags[key]) {
throw new InputValidationError(`required option '${usage}' not specified`);
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: we should use something like this.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure will add something like this

return call.args;
}

describe("eval recommendation command hierarchy", () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need these tests?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, I've seen these tests in other places, but they feel pretty low value. The tests on the handlers would fail if the hierarchy is incorrect. Also, what coverage do we get on the tests below that we don't get on the fixture based tests?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added these tests to cover behavior the fixtures don't exercise. E.g. malformed json, exercising source resolver, etc. But consensus is against unit test so I've removed this + the other one entirely

description: "get a recommendation by id",
flags: [flag("id", "the ID of the recommendation", z.string().optional())],
handle: async (ctx, flags) => {
if (!flags["id"]) throw new InputValidationError("required option '--id <id>' not specified");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah I think marking flags as required and supporting it directly in the framework would be ideal, but likely OOS here.

return (error as Error).name === "ResourceNotFoundException";
}

async function waitForTerminal(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is there a way to leverage

exportasyncfunctionwaitFor(
for these?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good call out I'll switch to use this util

return call.args;
}

describe("eval recommendation command hierarchy", () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, I've seen these tests in other places, but they feel pretty low value. The tests on the handlers would fail if the hierarchy is incorrect. Also, what coverage do we get on the tests below that we don't get on the fixture based tests?

status: "PENDING",
createdAt: new Date("2026-08-26T12:00:00.000Z"),
updatedAt: new Date("2026-08-26T12:00:00.000Z"),
} as StartRecommendationResponse;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why do we need to cast here?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

N/A since test file was removed altogether

Comment threadsrc/core/recommendation.test.tsx Outdated
@@ -0,0 +1,119 @@
import { describe, expect, test } from "bun:test";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are the handler tests sufficient? I think these are testing the same things at a lower level.

@github-actionsgithub-actionsBot added the size/l PR size: L label Aug 27, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.70115% with 4 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.41%. Comparing base (a1f2f32) to head (2783292).

Files with missing linesPatch %Lines
src/handlers/eval/recommendation/start/index.tsx94.20%4 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #2111 +/- ##
==========================================
Coverage 97.41% 97.41% ==========================================
Files 453 458 +5 Lines 27637 27810 +173 ==========================================
+ Hits 26922 27091 +169 - Misses 715 719 +4 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@jariy17
jariy17 merged commit c105d4f into refactorAug 27, 2026
22 checks passed
@jariy17
jariy17 deleted the eval-recommendation branch August 27, 2026 16:29
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@nborges-aws@codecov-commenter@Hweinstock@jariy17
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat: support imperative eval recommendation command - #2111

Merged
jariy17 merged 3 commits into
refactorfrom
eval-recommendation
Aug 27, 2026
Merged

feat: support imperative eval recommendation command#2111
jariy17 merged 3 commits into
refactorfrom
eval-recommendation

Conversation

@nborges-aws

Copy link
Copy Markdown
Contributor

Description

Adds imperative command-line support for AgentCore evaluation recommendations:

  • eval recommendation start
  • eval recommendation get
  • eval recommendation list
  • eval recommendation delete

Summary of Changes

  • Supports SYSTEM_PROMPT_RECOMMENDATION and TOOL_DESCRIPTION_RECOMMENDATION recommendation types
  • start accepts recommendation configuration as inline JSON, a file:// path, or stdin (-) through SourceResolver class
  • Supports the service API fields for description, KMS key ARN, and tags
  • Adds the recommendation operations to EvalClient, the shared handler types, and TestCoreClient
  • Adds recorded golden fixtures covering the complete start, get, list, and delete flow

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Documentation update
  • Other (please describe):

Testing

How have you tested the change?

Added unit tests and golden fixture tests exercising recommendation logic.

  • bun run test (2052 pass, 0 fail)
  • I ran npm run test:unit and npm run test:integ
  • I ran npm run typecheck
  • I ran npm run lint
  • If I modified src/assets/, I ran npm run test:update-snapshots and committed the updated snapshots

Checklist

  • I have read the CONTRIBUTING document
  • I have added any necessary tests that prove my fix is effective or my feature works
  • I have updated the documentation accordingly
  • I have added an appropriate example to the documentation to outline the feature, or no new docs are needed
  • My changes generate no new warnings
  • Any dependent changes have been merged and published

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the
terms of your choice.

description: "get a recommendation by id",
flags: [flag("id", "the ID of the recommendation", z.string().optional())],
handle: async (ctx, flags) => {
if (!flags["id"]) throw new InputValidationError("required option '--id <id>' not specified");

@jariy17jariy17Aug 26, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can definitely abstracted this so we have a common error message for this type of thing. OBO

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah I think marking flags as required and supporting it directly in the framework would be ideal, but likely OOS here.

Comment threadsrc/core/eval.tsx Outdated
}

async startRecommendation(
request: StartRecommendationRequest,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We create our own types here because handlers define the interface for the core client. Even if it's a one-to-one mapping, we should still define our own type.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sure will update

Comment threadsrc/core/recommendation.test.tsx Outdated
@@ -0,0 +1,119 @@
import { describe, expect, test } from "bun:test";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't need this. Use Golden tests please

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are the handler tests sufficient? I think these are testing the same things at a lower level.

flag("tags", "tags as key=value (repeatable) or JSON object", z.array(z.string()).optional()),
],
handle: async (ctx, flags) => {
if (!flags["name"]) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

const requiredFlags: Array<[string, string]> = [
["name", "--name <name>"],
["type", "--type <type>"],
["recommendation-config", "--recommendation-config <recommendation-config>"],
];
for (const [key, usage] of requiredFlags) {
if (!flags[key]) {
throw new InputValidationError(`required option '${usage}' not specified`);
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: we should use something like this.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure will add something like this

return call.args;
}

describe("eval recommendation command hierarchy", () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need these tests?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, I've seen these tests in other places, but they feel pretty low value. The tests on the handlers would fail if the hierarchy is incorrect. Also, what coverage do we get on the tests below that we don't get on the fixture based tests?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added these tests to cover behavior the fixtures don't exercise. E.g. malformed json, exercising source resolver, etc. But consensus is against unit test so I've removed this + the other one entirely

description: "get a recommendation by id",
flags: [flag("id", "the ID of the recommendation", z.string().optional())],
handle: async (ctx, flags) => {
if (!flags["id"]) throw new InputValidationError("required option '--id <id>' not specified");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah I think marking flags as required and supporting it directly in the framework would be ideal, but likely OOS here.

return (error as Error).name === "ResourceNotFoundException";
}

async function waitForTerminal(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is there a way to leverage

exportasyncfunctionwaitFor(
for these?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good call out I'll switch to use this util

return call.args;
}

describe("eval recommendation command hierarchy", () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, I've seen these tests in other places, but they feel pretty low value. The tests on the handlers would fail if the hierarchy is incorrect. Also, what coverage do we get on the tests below that we don't get on the fixture based tests?

status: "PENDING",
createdAt: new Date("2026-08-26T12:00:00.000Z"),
updatedAt: new Date("2026-08-26T12:00:00.000Z"),
} as StartRecommendationResponse;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why do we need to cast here?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

N/A since test file was removed altogether

Comment threadsrc/core/recommendation.test.tsx Outdated
@@ -0,0 +1,119 @@
import { describe, expect, test } from "bun:test";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are the handler tests sufficient? I think these are testing the same things at a lower level.

@github-actionsgithub-actionsBot added the size/l PR size: L label Aug 27, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.70115% with 4 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.41%. Comparing base (a1f2f32) to head (2783292).

Files with missing linesPatch %Lines
src/handlers/eval/recommendation/start/index.tsx94.20%4 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #2111 +/- ##
==========================================
Coverage 97.41% 97.41% ==========================================
Files 453 458 +5 Lines 27637 27810 +173 ==========================================
+ Hits 26922 27091 +169 - Misses 715 719 +4 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@jariy17
jariy17 merged commit c105d4f into refactorAug 27, 2026
22 checks passed
@jariy17
jariy17 deleted the eval-recommendation branch August 27, 2026 16:29
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@nborges-aws@codecov-commenter@Hweinstock@jariy17
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: support imperative eval recommendation command - #2111

Merged
jariy17 merged 3 commits into
refactorfrom
eval-recommendation
Aug 27, 2026
Merged

feat: support imperative eval recommendation command#2111
jariy17 merged 3 commits into
refactorfrom
eval-recommendation

Conversation

@nborges-aws

Copy link
Copy Markdown
Contributor

Description

Adds imperative command-line support for AgentCore evaluation recommendations:

  • eval recommendation start
  • eval recommendation get
  • eval recommendation list
  • eval recommendation delete

Summary of Changes

  • Supports SYSTEM_PROMPT_RECOMMENDATION and TOOL_DESCRIPTION_RECOMMENDATION recommendation types
  • start accepts recommendation configuration as inline JSON, a file:// path, or stdin (-) through SourceResolver class
  • Supports the service API fields for description, KMS key ARN, and tags
  • Adds the recommendation operations to EvalClient, the shared handler types, and TestCoreClient
  • Adds recorded golden fixtures covering the complete start, get, list, and delete flow

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Documentation update
  • Other (please describe):

Testing

How have you tested the change?

Added unit tests and golden fixture tests exercising recommendation logic.

  • bun run test (2052 pass, 0 fail)
  • I ran npm run test:unit and npm run test:integ
  • I ran npm run typecheck
  • I ran npm run lint
  • If I modified src/assets/, I ran npm run test:update-snapshots and committed the updated snapshots

Checklist

  • I have read the CONTRIBUTING document
  • I have added any necessary tests that prove my fix is effective or my feature works
  • I have updated the documentation accordingly
  • I have added an appropriate example to the documentation to outline the feature, or no new docs are needed
  • My changes generate no new warnings
  • Any dependent changes have been merged and published

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the
terms of your choice.

description: "get a recommendation by id",
flags: [flag("id", "the ID of the recommendation", z.string().optional())],
handle: async (ctx, flags) => {
if (!flags["id"]) throw new InputValidationError("required option '--id <id>' not specified");

@jariy17jariy17Aug 26, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can definitely abstracted this so we have a common error message for this type of thing. OBO

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah I think marking flags as required and supporting it directly in the framework would be ideal, but likely OOS here.

Comment threadsrc/core/eval.tsx Outdated
}

async startRecommendation(
request: StartRecommendationRequest,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We create our own types here because handlers define the interface for the core client. Even if it's a one-to-one mapping, we should still define our own type.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sure will update

Comment threadsrc/core/recommendation.test.tsx Outdated
@@ -0,0 +1,119 @@
import { describe, expect, test } from "bun:test";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't need this. Use Golden tests please

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are the handler tests sufficient? I think these are testing the same things at a lower level.

flag("tags", "tags as key=value (repeatable) or JSON object", z.array(z.string()).optional()),
],
handle: async (ctx, flags) => {
if (!flags["name"]) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

const requiredFlags: Array<[string, string]> = [
["name", "--name <name>"],
["type", "--type <type>"],
["recommendation-config", "--recommendation-config <recommendation-config>"],
];
for (const [key, usage] of requiredFlags) {
if (!flags[key]) {
throw new InputValidationError(`required option '${usage}' not specified`);
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: we should use something like this.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure will add something like this

return call.args;
}

describe("eval recommendation command hierarchy", () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need these tests?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, I've seen these tests in other places, but they feel pretty low value. The tests on the handlers would fail if the hierarchy is incorrect. Also, what coverage do we get on the tests below that we don't get on the fixture based tests?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added these tests to cover behavior the fixtures don't exercise. E.g. malformed json, exercising source resolver, etc. But consensus is against unit test so I've removed this + the other one entirely

description: "get a recommendation by id",
flags: [flag("id", "the ID of the recommendation", z.string().optional())],
handle: async (ctx, flags) => {
if (!flags["id"]) throw new InputValidationError("required option '--id <id>' not specified");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah I think marking flags as required and supporting it directly in the framework would be ideal, but likely OOS here.

return (error as Error).name === "ResourceNotFoundException";
}

async function waitForTerminal(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is there a way to leverage

exportasyncfunctionwaitFor(
for these?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good call out I'll switch to use this util

return call.args;
}

describe("eval recommendation command hierarchy", () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, I've seen these tests in other places, but they feel pretty low value. The tests on the handlers would fail if the hierarchy is incorrect. Also, what coverage do we get on the tests below that we don't get on the fixture based tests?

status: "PENDING",
createdAt: new Date("2026-08-26T12:00:00.000Z"),
updatedAt: new Date("2026-08-26T12:00:00.000Z"),
} as StartRecommendationResponse;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why do we need to cast here?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

N/A since test file was removed altogether

Comment threadsrc/core/recommendation.test.tsx Outdated
@@ -0,0 +1,119 @@
import { describe, expect, test } from "bun:test";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are the handler tests sufficient? I think these are testing the same things at a lower level.

@github-actionsgithub-actionsBot added the size/l PR size: L label Aug 27, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.70115% with 4 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.41%. Comparing base (a1f2f32) to head (2783292).

Files with missing linesPatch %Lines
src/handlers/eval/recommendation/start/index.tsx94.20%4 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #2111 +/- ##
==========================================
Coverage 97.41% 97.41% ==========================================
Files 453 458 +5 Lines 27637 27810 +173 ==========================================
+ Hits 26922 27091 +169 - Misses 715 719 +4 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@jariy17
jariy17 merged commit c105d4f into refactorAug 27, 2026
22 checks passed
@jariy17
jariy17 deleted the eval-recommendation branch August 27, 2026 16:29
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@nborges-aws@codecov-commenter@Hweinstock@jariy17
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: support imperative eval recommendation command - #2111

Merged
jariy17 merged 3 commits into
refactorfrom
eval-recommendation
Aug 27, 2026
Merged

feat: support imperative eval recommendation command#2111
jariy17 merged 3 commits into
refactorfrom
eval-recommendation

Conversation

@nborges-aws

Copy link
Copy Markdown
Contributor

Description

Adds imperative command-line support for AgentCore evaluation recommendations:

  • eval recommendation start
  • eval recommendation get
  • eval recommendation list
  • eval recommendation delete

Summary of Changes

  • Supports SYSTEM_PROMPT_RECOMMENDATION and TOOL_DESCRIPTION_RECOMMENDATION recommendation types
  • start accepts recommendation configuration as inline JSON, a file:// path, or stdin (-) through SourceResolver class
  • Supports the service API fields for description, KMS key ARN, and tags
  • Adds the recommendation operations to EvalClient, the shared handler types, and TestCoreClient
  • Adds recorded golden fixtures covering the complete start, get, list, and delete flow

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Documentation update
  • Other (please describe):

Testing

How have you tested the change?

Added unit tests and golden fixture tests exercising recommendation logic.

  • bun run test (2052 pass, 0 fail)
  • I ran npm run test:unit and npm run test:integ
  • I ran npm run typecheck
  • I ran npm run lint
  • If I modified src/assets/, I ran npm run test:update-snapshots and committed the updated snapshots

Checklist

  • I have read the CONTRIBUTING document
  • I have added any necessary tests that prove my fix is effective or my feature works
  • I have updated the documentation accordingly
  • I have added an appropriate example to the documentation to outline the feature, or no new docs are needed
  • My changes generate no new warnings
  • Any dependent changes have been merged and published

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the
terms of your choice.

description: "get a recommendation by id",
flags: [flag("id", "the ID of the recommendation", z.string().optional())],
handle: async (ctx, flags) => {
if (!flags["id"]) throw new InputValidationError("required option '--id <id>' not specified");

@jariy17jariy17Aug 26, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can definitely abstracted this so we have a common error message for this type of thing. OBO

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah I think marking flags as required and supporting it directly in the framework would be ideal, but likely OOS here.

Comment threadsrc/core/eval.tsx Outdated
}

async startRecommendation(
request: StartRecommendationRequest,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We create our own types here because handlers define the interface for the core client. Even if it's a one-to-one mapping, we should still define our own type.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sure will update

Comment threadsrc/core/recommendation.test.tsx Outdated
@@ -0,0 +1,119 @@
import { describe, expect, test } from "bun:test";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't need this. Use Golden tests please

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are the handler tests sufficient? I think these are testing the same things at a lower level.

flag("tags", "tags as key=value (repeatable) or JSON object", z.array(z.string()).optional()),
],
handle: async (ctx, flags) => {
if (!flags["name"]) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

const requiredFlags: Array<[string, string]> = [
["name", "--name <name>"],
["type", "--type <type>"],
["recommendation-config", "--recommendation-config <recommendation-config>"],
];
for (const [key, usage] of requiredFlags) {
if (!flags[key]) {
throw new InputValidationError(`required option '${usage}' not specified`);
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: we should use something like this.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure will add something like this

return call.args;
}

describe("eval recommendation command hierarchy", () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need these tests?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, I've seen these tests in other places, but they feel pretty low value. The tests on the handlers would fail if the hierarchy is incorrect. Also, what coverage do we get on the tests below that we don't get on the fixture based tests?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added these tests to cover behavior the fixtures don't exercise. E.g. malformed json, exercising source resolver, etc. But consensus is against unit test so I've removed this + the other one entirely

description: "get a recommendation by id",
flags: [flag("id", "the ID of the recommendation", z.string().optional())],
handle: async (ctx, flags) => {
if (!flags["id"]) throw new InputValidationError("required option '--id <id>' not specified");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah I think marking flags as required and supporting it directly in the framework would be ideal, but likely OOS here.

return (error as Error).name === "ResourceNotFoundException";
}

async function waitForTerminal(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is there a way to leverage

exportasyncfunctionwaitFor(
for these?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good call out I'll switch to use this util

return call.args;
}

describe("eval recommendation command hierarchy", () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, I've seen these tests in other places, but they feel pretty low value. The tests on the handlers would fail if the hierarchy is incorrect. Also, what coverage do we get on the tests below that we don't get on the fixture based tests?

status: "PENDING",
createdAt: new Date("2026-08-26T12:00:00.000Z"),
updatedAt: new Date("2026-08-26T12:00:00.000Z"),
} as StartRecommendationResponse;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why do we need to cast here?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

N/A since test file was removed altogether

Comment threadsrc/core/recommendation.test.tsx Outdated
@@ -0,0 +1,119 @@
import { describe, expect, test } from "bun:test";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are the handler tests sufficient? I think these are testing the same things at a lower level.

@github-actionsgithub-actionsBot added the size/l PR size: L label Aug 27, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.70115% with 4 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.41%. Comparing base (a1f2f32) to head (2783292).

Files with missing linesPatch %Lines
src/handlers/eval/recommendation/start/index.tsx94.20%4 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #2111 +/- ##
==========================================
Coverage 97.41% 97.41% ==========================================
Files 453 458 +5 Lines 27637 27810 +173 ==========================================
+ Hits 26922 27091 +169 - Misses 715 719 +4 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@jariy17
jariy17 merged commit c105d4f into refactorAug 27, 2026
22 checks passed
@jariy17
jariy17 deleted the eval-recommendation branch August 27, 2026 16:29
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@nborges-aws@codecov-commenter@Hweinstock@jariy17
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat: support imperative eval recommendation command - #2111

Merged
jariy17 merged 3 commits into
refactorfrom
eval-recommendation
Aug 27, 2026
Merged

feat: support imperative eval recommendation command#2111
jariy17 merged 3 commits into
refactorfrom
eval-recommendation

Conversation

@nborges-aws

Copy link
Copy Markdown
Contributor

Description

Adds imperative command-line support for AgentCore evaluation recommendations:

  • eval recommendation start
  • eval recommendation get
  • eval recommendation list
  • eval recommendation delete

Summary of Changes

  • Supports SYSTEM_PROMPT_RECOMMENDATION and TOOL_DESCRIPTION_RECOMMENDATION recommendation types
  • start accepts recommendation configuration as inline JSON, a file:// path, or stdin (-) through SourceResolver class
  • Supports the service API fields for description, KMS key ARN, and tags
  • Adds the recommendation operations to EvalClient, the shared handler types, and TestCoreClient
  • Adds recorded golden fixtures covering the complete start, get, list, and delete flow

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Documentation update
  • Other (please describe):

Testing

How have you tested the change?

Added unit tests and golden fixture tests exercising recommendation logic.

  • bun run test (2052 pass, 0 fail)
  • I ran npm run test:unit and npm run test:integ
  • I ran npm run typecheck
  • I ran npm run lint
  • If I modified src/assets/, I ran npm run test:update-snapshots and committed the updated snapshots

Checklist

  • I have read the CONTRIBUTING document
  • I have added any necessary tests that prove my fix is effective or my feature works
  • I have updated the documentation accordingly
  • I have added an appropriate example to the documentation to outline the feature, or no new docs are needed
  • My changes generate no new warnings
  • Any dependent changes have been merged and published

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the
terms of your choice.

description: "get a recommendation by id",
flags: [flag("id", "the ID of the recommendation", z.string().optional())],
handle: async (ctx, flags) => {
if (!flags["id"]) throw new InputValidationError("required option '--id <id>' not specified");

@jariy17jariy17Aug 26, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can definitely abstracted this so we have a common error message for this type of thing. OBO

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah I think marking flags as required and supporting it directly in the framework would be ideal, but likely OOS here.

Comment threadsrc/core/eval.tsx Outdated
}

async startRecommendation(
request: StartRecommendationRequest,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We create our own types here because handlers define the interface for the core client. Even if it's a one-to-one mapping, we should still define our own type.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

sure will update

Comment threadsrc/core/recommendation.test.tsx Outdated
@@ -0,0 +1,119 @@
import { describe, expect, test } from "bun:test";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't need this. Use Golden tests please

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are the handler tests sufficient? I think these are testing the same things at a lower level.

flag("tags", "tags as key=value (repeatable) or JSON object", z.array(z.string()).optional()),
],
handle: async (ctx, flags) => {
if (!flags["name"]) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

const requiredFlags: Array<[string, string]> = [
["name", "--name <name>"],
["type", "--type <type>"],
["recommendation-config", "--recommendation-config <recommendation-config>"],
];
for (const [key, usage] of requiredFlags) {
if (!flags[key]) {
throw new InputValidationError(`required option '${usage}' not specified`);
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: we should use something like this.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure will add something like this

return call.args;
}

describe("eval recommendation command hierarchy", () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need these tests?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, I've seen these tests in other places, but they feel pretty low value. The tests on the handlers would fail if the hierarchy is incorrect. Also, what coverage do we get on the tests below that we don't get on the fixture based tests?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added these tests to cover behavior the fixtures don't exercise. E.g. malformed json, exercising source resolver, etc. But consensus is against unit test so I've removed this + the other one entirely

description: "get a recommendation by id",
flags: [flag("id", "the ID of the recommendation", z.string().optional())],
handle: async (ctx, flags) => {
if (!flags["id"]) throw new InputValidationError("required option '--id <id>' not specified");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yeah I think marking flags as required and supporting it directly in the framework would be ideal, but likely OOS here.

return (error as Error).name === "ResourceNotFoundException";
}

async function waitForTerminal(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is there a way to leverage

exportasyncfunctionwaitFor(
for these?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good call out I'll switch to use this util

return call.args;
}

describe("eval recommendation command hierarchy", () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, I've seen these tests in other places, but they feel pretty low value. The tests on the handlers would fail if the hierarchy is incorrect. Also, what coverage do we get on the tests below that we don't get on the fixture based tests?

status: "PENDING",
createdAt: new Date("2026-08-26T12:00:00.000Z"),
updatedAt: new Date("2026-08-26T12:00:00.000Z"),
} as StartRecommendationResponse;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why do we need to cast here?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

N/A since test file was removed altogether

Comment threadsrc/core/recommendation.test.tsx Outdated
@@ -0,0 +1,119 @@
import { describe, expect, test } from "bun:test";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are the handler tests sufficient? I think these are testing the same things at a lower level.

@github-actionsgithub-actionsBot added the size/l PR size: L label Aug 27, 2026
@agentcore-devx-automationagentcore-devx-automationBot added the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automationagentcore-devx-automationBot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.70115% with 4 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.41%. Comparing base (a1f2f32) to head (2783292).

Files with missing linesPatch %Lines
src/handlers/eval/recommendation/start/index.tsx94.20%4 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #2111 +/- ##
==========================================
Coverage 97.41% 97.41% ==========================================
Files 453 458 +5 Lines 27637 27810 +173 ==========================================
+ Hits 26922 27091 +169 - Misses 715 719 +4 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@jariy17
jariy17 merged commit c105d4f into refactorAug 27, 2026
22 checks passed
@jariy17
jariy17 deleted the eval-recommendation branch August 27, 2026 16:29
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/lPR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants

@nborges-aws@codecov-commenter@Hweinstock@jariy17