feat(tel): implement otel sink - #1888

Merged
Hweinstock merged 14 commits into
aws:refactorfrom
Hweinstock:feat/otel-sink
Aug 14, 2026
Merged

feat(tel): implement otel sink#1888
Hweinstock merged 14 commits into
aws:refactorfrom
Hweinstock:feat/otel-sink

Conversation

@Hweinstock

@HweinstockHweinstock commented Jul 31, 2026

Copy link
Copy Markdown
Contributor

Problem

The AgentCore CLI is not currently publishing telemetry to our collector.

Solution

  • otel sdk requires the resource attributes separated from the metric attributes (even though they are merged on the request) so we explicitly inject those into each sink.
  • implement an otel histogram metric sink. We use histogram because our values represent duration, so histogram gives us count, and distributions over the values in a single instrument.

Testing

  • run our collector locally, and override local endpoint to point to it. Then verified metrics hit cloudwatch in dev account. (Requires one small backend change that is already in the pipeline).
  • added a unit test that spins up a local server, and verifies it only receives request when telemetry is enabled, and requests match otel format.

@github-actionsgithub-actionsBot added the agentcore-harness-reviewing AgentCore Harness review in progress label Jul 31, 2026
@github-actionsgithub-actionsBot removed the agentcore-harness-reviewing AgentCore Harness review in progress label Jul 31, 2026
@codecov-commenter

codecov-commenter commented Jul 31, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.48485% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 96.70%. Comparing base (7e29eda) to head (6806436).
⚠️ Report is 12 commits behind head on refactor.

Files with missing linesPatch %Lines
src/telemetry/otelSink.tsx98.07%1 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #1888 +/- ##
=========================================
Coverage 96.69% 96.70% =========================================
Files 291 292 +1 Lines 16012 16068 +56 =========================================
+ Hits 15483 15538 +55 - Misses 529 530 +1 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@Hweinstock

Copy link
Copy Markdown
ContributorAuthor

^ the line missing coverage is getName() on the collector sink.

@HweinstockHweinstock changed the title 0feat(tel): implement otel sinkfeat(tel): implement otel sinkJul 31, 2026
@Hweinstock
Hweinstock marked this pull request as ready for review July 31, 2026 19:22

async shutdown(): Promise<void> {
try {
await this.meterProvider.forceFlush({ timeoutMillis: this.flushTimeoutMs });

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

outside this try block, there is meterProvider.shutdown() which also flushes the metric again.

https://opentelemetry.io/docs/specs/otel/metrics/sdk/#shutdown

can we ensure one export somehow?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This method provides a way for provider to do any cleanup required.
Shutdown MUST be called only once for each MeterProvider instance. After the call to Shutdown, subsequent attempts to get a Meter are not allowed. SDKs SHOULD return a valid no-op Meter for these calls, if possible.
Shutdown SHOULD provide a way to let the caller know whether it succeeded, failed or timed out.
Shutdown SHOULD complete or abort within some timeout. Shutdown MAY be implemented as a blocking API or an asynchronous API which notifies the caller via a callback or an event. [OpenTelemetry SDK](https://opentelemetry.io/docs/specs/otel/overview/#sdk) authors MAY decide if they want to make the shutdown timeout configurable.
Shutdown MUST be implemented at least by invoking Shutdown on all registered [MetricReader](https://opentelemetry.io/docs/specs/otel/metrics/sdk/#metricreader) and [MetricExporter](https://opentelemetry.io/docs/specs/otel/metrics/sdk/#metricexporter) instances.

from https://opentelemetry.io/docs/specs/otel/metrics/sdk/#shutdown.

I don't see any explicit lines in the protocol linked for shutdown that it also flushes. Based on some testing, I think it does internally, but I think its safer to make that behavior explicit. If we flush twice, its a no-op anyway.


if (globalConfig.telemetry.enabled)
metricSinks.push(
new OtelHistogramSink({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like a wrong/malformed endpoint makes getMetricSinks() reject, and shutdown() propagates that rejection, erroring out in the CLI command. Is that understanding correct? can we make it best-effort?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we (the dev team) should be the only ones modifying the endpoint for testing purposes. In which case, I think the ideal behavior is that we reject early.

If a user decides to go into the global config and add an invalid override, I think rejecting is reasonable.

notgitika
notgitika previously approved these changes Aug 1, 2026

@notgitikanotgitika left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGMT

Comment threadsrc/telemetry/otelSink.tsx Outdated
try {
await this.meterProvider.forceFlush({ timeoutMillis: this.flushTimeoutMs });
} catch (e) {
const error = e instanceof Error ? e : new Error(String(e));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to use our AgentCoreError.fromError() here instead?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah I think that makes sense to ensure its consistent.

@Hweinstock

Copy link
Copy Markdown
ContributorAuthor

slack failure appears unrelated.

Configure AWS credentials (central secrets reader)
1m 0s
Run aws-actions/configure-aws-credentials@v6
Retry validateCredentials: attempt 1 of 12 failed: Credentials could not be loaded, please check your action inputs: Could not load credentials from any providers. Retrying after 23ms.

notgitika
notgitika previously approved these changes Aug 12, 2026
resource: resourceFromAttributes(config.resourceAttributes),
readers: [
new PeriodicExportingMetricReader({
exporter: new OTLPMetricExporter({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I may be missing whether inheriting the caller's OTEL configuration is intentional, but OTLPMetricExporter reads OTEL_EXPORTER_OTLP_HEADERS and OTEL_EXPORTER_OTLP_METRICS_HEADERS when headers is omitted. I set dummy authorization and x-api-key values and confirmed both were forwarded to the configured collector. A user running the CLI in an environment configured for another OTLP backend could therefore send those credentials to telemetry.agentcore.aws.dev. Would passing headers: {} here make sense?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nice catch, was not aware of these environment variables. Was going to add this as a follow-up, but let me inject the right header now to avoid this behavior.

.child({ metricName, metricValue: value, metricAttributes: attributes })
.info(`sending telemetry metric to collector`);

this.getHistogram(metricName).record(value, attributes);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the resource attributes are still ending up in both places. InMemoryMetricEvent.emit() merges resourceAttributes into the sink attributes, then this line records that merged map after the provider already registered the same fields with resourceFromAttributes. In a local OTLP capture, all eight resource keys appeared under both resource.attributes and histogram.dataPoints[].attributes. Since the goal here is to separate resource and metric attributes, would it make sense for the sink contract to receive only metric attributes, with FileSystemSink composing its JSONL entry from its configured resource attributes? The test could also assert that resource keys are absent from the datapoint. This is really a low finding and could always be changed in the future.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good catch, I think removing is the right call to simplify.

aidandaly24
aidandaly24 previously approved these changes Aug 12, 2026
@Hweinstock
Hweinstock dismissed stale reviews from aidandaly24 and notgitika via 7e9a97eAugust 13, 2026 13:05
@notgitika

Copy link
Copy Markdown
Contributor

looks good to me!

@Hweinstock
Hweinstock merged commit 01eb60b into aws:refactorAug 14, 2026
8 of 11 checks passed
@Hweinstock
Hweinstock deleted the feat/otel-sink branch August 14, 2026 19:09
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@Hweinstock@codecov-commenter@notgitika@AlexanderRichey@aidandaly24
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(tel): implement otel sink - #1888

Merged
Hweinstock merged 14 commits into
aws:refactorfrom
Hweinstock:feat/otel-sink
Aug 14, 2026
Merged

feat(tel): implement otel sink#1888
Hweinstock merged 14 commits into
aws:refactorfrom
Hweinstock:feat/otel-sink

Conversation

@Hweinstock

@HweinstockHweinstock commented Jul 31, 2026

Copy link
Copy Markdown
Contributor

Problem

The AgentCore CLI is not currently publishing telemetry to our collector.

Solution

  • otel sdk requires the resource attributes separated from the metric attributes (even though they are merged on the request) so we explicitly inject those into each sink.
  • implement an otel histogram metric sink. We use histogram because our values represent duration, so histogram gives us count, and distributions over the values in a single instrument.

Testing

  • run our collector locally, and override local endpoint to point to it. Then verified metrics hit cloudwatch in dev account. (Requires one small backend change that is already in the pipeline).
  • added a unit test that spins up a local server, and verifies it only receives request when telemetry is enabled, and requests match otel format.

@github-actionsgithub-actionsBot added the agentcore-harness-reviewing AgentCore Harness review in progress label Jul 31, 2026
@github-actionsgithub-actionsBot removed the agentcore-harness-reviewing AgentCore Harness review in progress label Jul 31, 2026
@codecov-commenter

codecov-commenter commented Jul 31, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.48485% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 96.70%. Comparing base (7e29eda) to head (6806436).
⚠️ Report is 12 commits behind head on refactor.

Files with missing linesPatch %Lines
src/telemetry/otelSink.tsx98.07%1 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #1888 +/- ##
=========================================
Coverage 96.69% 96.70% =========================================
Files 291 292 +1 Lines 16012 16068 +56 =========================================
+ Hits 15483 15538 +55 - Misses 529 530 +1 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@Hweinstock

Copy link
Copy Markdown
ContributorAuthor

^ the line missing coverage is getName() on the collector sink.

@HweinstockHweinstock changed the title 0feat(tel): implement otel sinkfeat(tel): implement otel sinkJul 31, 2026
@Hweinstock
Hweinstock marked this pull request as ready for review July 31, 2026 19:22

async shutdown(): Promise<void> {
try {
await this.meterProvider.forceFlush({ timeoutMillis: this.flushTimeoutMs });

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

outside this try block, there is meterProvider.shutdown() which also flushes the metric again.

https://opentelemetry.io/docs/specs/otel/metrics/sdk/#shutdown

can we ensure one export somehow?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This method provides a way for provider to do any cleanup required.
Shutdown MUST be called only once for each MeterProvider instance. After the call to Shutdown, subsequent attempts to get a Meter are not allowed. SDKs SHOULD return a valid no-op Meter for these calls, if possible.
Shutdown SHOULD provide a way to let the caller know whether it succeeded, failed or timed out.
Shutdown SHOULD complete or abort within some timeout. Shutdown MAY be implemented as a blocking API or an asynchronous API which notifies the caller via a callback or an event. [OpenTelemetry SDK](https://opentelemetry.io/docs/specs/otel/overview/#sdk) authors MAY decide if they want to make the shutdown timeout configurable.
Shutdown MUST be implemented at least by invoking Shutdown on all registered [MetricReader](https://opentelemetry.io/docs/specs/otel/metrics/sdk/#metricreader) and [MetricExporter](https://opentelemetry.io/docs/specs/otel/metrics/sdk/#metricexporter) instances.

from https://opentelemetry.io/docs/specs/otel/metrics/sdk/#shutdown.

I don't see any explicit lines in the protocol linked for shutdown that it also flushes. Based on some testing, I think it does internally, but I think its safer to make that behavior explicit. If we flush twice, its a no-op anyway.


if (globalConfig.telemetry.enabled)
metricSinks.push(
new OtelHistogramSink({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like a wrong/malformed endpoint makes getMetricSinks() reject, and shutdown() propagates that rejection, erroring out in the CLI command. Is that understanding correct? can we make it best-effort?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we (the dev team) should be the only ones modifying the endpoint for testing purposes. In which case, I think the ideal behavior is that we reject early.

If a user decides to go into the global config and add an invalid override, I think rejecting is reasonable.

notgitika
notgitika previously approved these changes Aug 1, 2026

@notgitikanotgitika left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGMT

Comment threadsrc/telemetry/otelSink.tsx Outdated
try {
await this.meterProvider.forceFlush({ timeoutMillis: this.flushTimeoutMs });
} catch (e) {
const error = e instanceof Error ? e : new Error(String(e));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to use our AgentCoreError.fromError() here instead?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah I think that makes sense to ensure its consistent.

@Hweinstock

Copy link
Copy Markdown
ContributorAuthor

slack failure appears unrelated.

Configure AWS credentials (central secrets reader)
1m 0s
Run aws-actions/configure-aws-credentials@v6
Retry validateCredentials: attempt 1 of 12 failed: Credentials could not be loaded, please check your action inputs: Could not load credentials from any providers. Retrying after 23ms.

notgitika
notgitika previously approved these changes Aug 12, 2026
resource: resourceFromAttributes(config.resourceAttributes),
readers: [
new PeriodicExportingMetricReader({
exporter: new OTLPMetricExporter({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I may be missing whether inheriting the caller's OTEL configuration is intentional, but OTLPMetricExporter reads OTEL_EXPORTER_OTLP_HEADERS and OTEL_EXPORTER_OTLP_METRICS_HEADERS when headers is omitted. I set dummy authorization and x-api-key values and confirmed both were forwarded to the configured collector. A user running the CLI in an environment configured for another OTLP backend could therefore send those credentials to telemetry.agentcore.aws.dev. Would passing headers: {} here make sense?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nice catch, was not aware of these environment variables. Was going to add this as a follow-up, but let me inject the right header now to avoid this behavior.

.child({ metricName, metricValue: value, metricAttributes: attributes })
.info(`sending telemetry metric to collector`);

this.getHistogram(metricName).record(value, attributes);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the resource attributes are still ending up in both places. InMemoryMetricEvent.emit() merges resourceAttributes into the sink attributes, then this line records that merged map after the provider already registered the same fields with resourceFromAttributes. In a local OTLP capture, all eight resource keys appeared under both resource.attributes and histogram.dataPoints[].attributes. Since the goal here is to separate resource and metric attributes, would it make sense for the sink contract to receive only metric attributes, with FileSystemSink composing its JSONL entry from its configured resource attributes? The test could also assert that resource keys are absent from the datapoint. This is really a low finding and could always be changed in the future.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good catch, I think removing is the right call to simplify.

aidandaly24
aidandaly24 previously approved these changes Aug 12, 2026
@Hweinstock
Hweinstock dismissed stale reviews from aidandaly24 and notgitika via 7e9a97eAugust 13, 2026 13:05
@notgitika

Copy link
Copy Markdown
Contributor

looks good to me!

@Hweinstock
Hweinstock merged commit 01eb60b into aws:refactorAug 14, 2026
8 of 11 checks passed
@Hweinstock
Hweinstock deleted the feat/otel-sink branch August 14, 2026 19:09
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@Hweinstock@codecov-commenter@notgitika@AlexanderRichey@aidandaly24
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(tel): implement otel sink - #1888

Merged
Hweinstock merged 14 commits into
aws:refactorfrom
Hweinstock:feat/otel-sink
Aug 14, 2026
Merged

feat(tel): implement otel sink#1888
Hweinstock merged 14 commits into
aws:refactorfrom
Hweinstock:feat/otel-sink

Conversation

@Hweinstock

@HweinstockHweinstock commented Jul 31, 2026

Copy link
Copy Markdown
Contributor

Problem

The AgentCore CLI is not currently publishing telemetry to our collector.

Solution

  • otel sdk requires the resource attributes separated from the metric attributes (even though they are merged on the request) so we explicitly inject those into each sink.
  • implement an otel histogram metric sink. We use histogram because our values represent duration, so histogram gives us count, and distributions over the values in a single instrument.

Testing

  • run our collector locally, and override local endpoint to point to it. Then verified metrics hit cloudwatch in dev account. (Requires one small backend change that is already in the pipeline).
  • added a unit test that spins up a local server, and verifies it only receives request when telemetry is enabled, and requests match otel format.

@github-actionsgithub-actionsBot added the agentcore-harness-reviewing AgentCore Harness review in progress label Jul 31, 2026
@github-actionsgithub-actionsBot removed the agentcore-harness-reviewing AgentCore Harness review in progress label Jul 31, 2026
@codecov-commenter

codecov-commenter commented Jul 31, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.48485% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 96.70%. Comparing base (7e29eda) to head (6806436).
⚠️ Report is 12 commits behind head on refactor.

Files with missing linesPatch %Lines
src/telemetry/otelSink.tsx98.07%1 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #1888 +/- ##
=========================================
Coverage 96.69% 96.70% =========================================
Files 291 292 +1 Lines 16012 16068 +56 =========================================
+ Hits 15483 15538 +55 - Misses 529 530 +1 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@Hweinstock

Copy link
Copy Markdown
ContributorAuthor

^ the line missing coverage is getName() on the collector sink.

@HweinstockHweinstock changed the title 0feat(tel): implement otel sinkfeat(tel): implement otel sinkJul 31, 2026
@Hweinstock
Hweinstock marked this pull request as ready for review July 31, 2026 19:22

async shutdown(): Promise<void> {
try {
await this.meterProvider.forceFlush({ timeoutMillis: this.flushTimeoutMs });

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

outside this try block, there is meterProvider.shutdown() which also flushes the metric again.

https://opentelemetry.io/docs/specs/otel/metrics/sdk/#shutdown

can we ensure one export somehow?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This method provides a way for provider to do any cleanup required.
Shutdown MUST be called only once for each MeterProvider instance. After the call to Shutdown, subsequent attempts to get a Meter are not allowed. SDKs SHOULD return a valid no-op Meter for these calls, if possible.
Shutdown SHOULD provide a way to let the caller know whether it succeeded, failed or timed out.
Shutdown SHOULD complete or abort within some timeout. Shutdown MAY be implemented as a blocking API or an asynchronous API which notifies the caller via a callback or an event. [OpenTelemetry SDK](https://opentelemetry.io/docs/specs/otel/overview/#sdk) authors MAY decide if they want to make the shutdown timeout configurable.
Shutdown MUST be implemented at least by invoking Shutdown on all registered [MetricReader](https://opentelemetry.io/docs/specs/otel/metrics/sdk/#metricreader) and [MetricExporter](https://opentelemetry.io/docs/specs/otel/metrics/sdk/#metricexporter) instances.

from https://opentelemetry.io/docs/specs/otel/metrics/sdk/#shutdown.

I don't see any explicit lines in the protocol linked for shutdown that it also flushes. Based on some testing, I think it does internally, but I think its safer to make that behavior explicit. If we flush twice, its a no-op anyway.


if (globalConfig.telemetry.enabled)
metricSinks.push(
new OtelHistogramSink({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like a wrong/malformed endpoint makes getMetricSinks() reject, and shutdown() propagates that rejection, erroring out in the CLI command. Is that understanding correct? can we make it best-effort?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we (the dev team) should be the only ones modifying the endpoint for testing purposes. In which case, I think the ideal behavior is that we reject early.

If a user decides to go into the global config and add an invalid override, I think rejecting is reasonable.

notgitika
notgitika previously approved these changes Aug 1, 2026

@notgitikanotgitika left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGMT

Comment threadsrc/telemetry/otelSink.tsx Outdated
try {
await this.meterProvider.forceFlush({ timeoutMillis: this.flushTimeoutMs });
} catch (e) {
const error = e instanceof Error ? e : new Error(String(e));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to use our AgentCoreError.fromError() here instead?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah I think that makes sense to ensure its consistent.

@Hweinstock

Copy link
Copy Markdown
ContributorAuthor

slack failure appears unrelated.

Configure AWS credentials (central secrets reader)
1m 0s
Run aws-actions/configure-aws-credentials@v6
Retry validateCredentials: attempt 1 of 12 failed: Credentials could not be loaded, please check your action inputs: Could not load credentials from any providers. Retrying after 23ms.

notgitika
notgitika previously approved these changes Aug 12, 2026
resource: resourceFromAttributes(config.resourceAttributes),
readers: [
new PeriodicExportingMetricReader({
exporter: new OTLPMetricExporter({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I may be missing whether inheriting the caller's OTEL configuration is intentional, but OTLPMetricExporter reads OTEL_EXPORTER_OTLP_HEADERS and OTEL_EXPORTER_OTLP_METRICS_HEADERS when headers is omitted. I set dummy authorization and x-api-key values and confirmed both were forwarded to the configured collector. A user running the CLI in an environment configured for another OTLP backend could therefore send those credentials to telemetry.agentcore.aws.dev. Would passing headers: {} here make sense?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nice catch, was not aware of these environment variables. Was going to add this as a follow-up, but let me inject the right header now to avoid this behavior.

.child({ metricName, metricValue: value, metricAttributes: attributes })
.info(`sending telemetry metric to collector`);

this.getHistogram(metricName).record(value, attributes);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the resource attributes are still ending up in both places. InMemoryMetricEvent.emit() merges resourceAttributes into the sink attributes, then this line records that merged map after the provider already registered the same fields with resourceFromAttributes. In a local OTLP capture, all eight resource keys appeared under both resource.attributes and histogram.dataPoints[].attributes. Since the goal here is to separate resource and metric attributes, would it make sense for the sink contract to receive only metric attributes, with FileSystemSink composing its JSONL entry from its configured resource attributes? The test could also assert that resource keys are absent from the datapoint. This is really a low finding and could always be changed in the future.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good catch, I think removing is the right call to simplify.

aidandaly24
aidandaly24 previously approved these changes Aug 12, 2026
@Hweinstock
Hweinstock dismissed stale reviews from aidandaly24 and notgitika via 7e9a97eAugust 13, 2026 13:05
@notgitika

Copy link
Copy Markdown
Contributor

looks good to me!

@Hweinstock
Hweinstock merged commit 01eb60b into aws:refactorAug 14, 2026
8 of 11 checks passed
@Hweinstock
Hweinstock deleted the feat/otel-sink branch August 14, 2026 19:09
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@Hweinstock@codecov-commenter@notgitika@AlexanderRichey@aidandaly24
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(tel): implement otel sink - #1888

Merged
Hweinstock merged 14 commits into
aws:refactorfrom
Hweinstock:feat/otel-sink
Aug 14, 2026
Merged

feat(tel): implement otel sink#1888
Hweinstock merged 14 commits into
aws:refactorfrom
Hweinstock:feat/otel-sink

Conversation

@Hweinstock

@HweinstockHweinstock commented Jul 31, 2026

Copy link
Copy Markdown
Contributor

Problem

The AgentCore CLI is not currently publishing telemetry to our collector.

Solution

  • otel sdk requires the resource attributes separated from the metric attributes (even though they are merged on the request) so we explicitly inject those into each sink.
  • implement an otel histogram metric sink. We use histogram because our values represent duration, so histogram gives us count, and distributions over the values in a single instrument.

Testing

  • run our collector locally, and override local endpoint to point to it. Then verified metrics hit cloudwatch in dev account. (Requires one small backend change that is already in the pipeline).
  • added a unit test that spins up a local server, and verifies it only receives request when telemetry is enabled, and requests match otel format.

@github-actionsgithub-actionsBot added the agentcore-harness-reviewing AgentCore Harness review in progress label Jul 31, 2026
@github-actionsgithub-actionsBot removed the agentcore-harness-reviewing AgentCore Harness review in progress label Jul 31, 2026
@codecov-commenter

codecov-commenter commented Jul 31, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.48485% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 96.70%. Comparing base (7e29eda) to head (6806436).
⚠️ Report is 12 commits behind head on refactor.

Files with missing linesPatch %Lines
src/telemetry/otelSink.tsx98.07%1 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #1888 +/- ##
=========================================
Coverage 96.69% 96.70% =========================================
Files 291 292 +1 Lines 16012 16068 +56 =========================================
+ Hits 15483 15538 +55 - Misses 529 530 +1 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@Hweinstock

Copy link
Copy Markdown
ContributorAuthor

^ the line missing coverage is getName() on the collector sink.

@HweinstockHweinstock changed the title 0feat(tel): implement otel sinkfeat(tel): implement otel sinkJul 31, 2026
@Hweinstock
Hweinstock marked this pull request as ready for review July 31, 2026 19:22

async shutdown(): Promise<void> {
try {
await this.meterProvider.forceFlush({ timeoutMillis: this.flushTimeoutMs });

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

outside this try block, there is meterProvider.shutdown() which also flushes the metric again.

https://opentelemetry.io/docs/specs/otel/metrics/sdk/#shutdown

can we ensure one export somehow?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This method provides a way for provider to do any cleanup required.
Shutdown MUST be called only once for each MeterProvider instance. After the call to Shutdown, subsequent attempts to get a Meter are not allowed. SDKs SHOULD return a valid no-op Meter for these calls, if possible.
Shutdown SHOULD provide a way to let the caller know whether it succeeded, failed or timed out.
Shutdown SHOULD complete or abort within some timeout. Shutdown MAY be implemented as a blocking API or an asynchronous API which notifies the caller via a callback or an event. [OpenTelemetry SDK](https://opentelemetry.io/docs/specs/otel/overview/#sdk) authors MAY decide if they want to make the shutdown timeout configurable.
Shutdown MUST be implemented at least by invoking Shutdown on all registered [MetricReader](https://opentelemetry.io/docs/specs/otel/metrics/sdk/#metricreader) and [MetricExporter](https://opentelemetry.io/docs/specs/otel/metrics/sdk/#metricexporter) instances.

from https://opentelemetry.io/docs/specs/otel/metrics/sdk/#shutdown.

I don't see any explicit lines in the protocol linked for shutdown that it also flushes. Based on some testing, I think it does internally, but I think its safer to make that behavior explicit. If we flush twice, its a no-op anyway.


if (globalConfig.telemetry.enabled)
metricSinks.push(
new OtelHistogramSink({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like a wrong/malformed endpoint makes getMetricSinks() reject, and shutdown() propagates that rejection, erroring out in the CLI command. Is that understanding correct? can we make it best-effort?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we (the dev team) should be the only ones modifying the endpoint for testing purposes. In which case, I think the ideal behavior is that we reject early.

If a user decides to go into the global config and add an invalid override, I think rejecting is reasonable.

notgitika
notgitika previously approved these changes Aug 1, 2026

@notgitikanotgitika left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGMT

Comment threadsrc/telemetry/otelSink.tsx Outdated
try {
await this.meterProvider.forceFlush({ timeoutMillis: this.flushTimeoutMs });
} catch (e) {
const error = e instanceof Error ? e : new Error(String(e));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to use our AgentCoreError.fromError() here instead?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah I think that makes sense to ensure its consistent.

@Hweinstock

Copy link
Copy Markdown
ContributorAuthor

slack failure appears unrelated.

Configure AWS credentials (central secrets reader)
1m 0s
Run aws-actions/configure-aws-credentials@v6
Retry validateCredentials: attempt 1 of 12 failed: Credentials could not be loaded, please check your action inputs: Could not load credentials from any providers. Retrying after 23ms.

notgitika
notgitika previously approved these changes Aug 12, 2026
resource: resourceFromAttributes(config.resourceAttributes),
readers: [
new PeriodicExportingMetricReader({
exporter: new OTLPMetricExporter({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I may be missing whether inheriting the caller's OTEL configuration is intentional, but OTLPMetricExporter reads OTEL_EXPORTER_OTLP_HEADERS and OTEL_EXPORTER_OTLP_METRICS_HEADERS when headers is omitted. I set dummy authorization and x-api-key values and confirmed both were forwarded to the configured collector. A user running the CLI in an environment configured for another OTLP backend could therefore send those credentials to telemetry.agentcore.aws.dev. Would passing headers: {} here make sense?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nice catch, was not aware of these environment variables. Was going to add this as a follow-up, but let me inject the right header now to avoid this behavior.

.child({ metricName, metricValue: value, metricAttributes: attributes })
.info(`sending telemetry metric to collector`);

this.getHistogram(metricName).record(value, attributes);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the resource attributes are still ending up in both places. InMemoryMetricEvent.emit() merges resourceAttributes into the sink attributes, then this line records that merged map after the provider already registered the same fields with resourceFromAttributes. In a local OTLP capture, all eight resource keys appeared under both resource.attributes and histogram.dataPoints[].attributes. Since the goal here is to separate resource and metric attributes, would it make sense for the sink contract to receive only metric attributes, with FileSystemSink composing its JSONL entry from its configured resource attributes? The test could also assert that resource keys are absent from the datapoint. This is really a low finding and could always be changed in the future.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good catch, I think removing is the right call to simplify.

aidandaly24
aidandaly24 previously approved these changes Aug 12, 2026
@Hweinstock
Hweinstock dismissed stale reviews from aidandaly24 and notgitika via 7e9a97eAugust 13, 2026 13:05
@notgitika

Copy link
Copy Markdown
Contributor

looks good to me!

@Hweinstock
Hweinstock merged commit 01eb60b into aws:refactorAug 14, 2026
8 of 11 checks passed
@Hweinstock
Hweinstock deleted the feat/otel-sink branch August 14, 2026 19:09
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@Hweinstock@codecov-commenter@notgitika@AlexanderRichey@aidandaly24
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(tel): implement otel sink - #1888

Merged
Hweinstock merged 14 commits into
aws:refactorfrom
Hweinstock:feat/otel-sink
Aug 14, 2026
Merged

feat(tel): implement otel sink#1888
Hweinstock merged 14 commits into
aws:refactorfrom
Hweinstock:feat/otel-sink

Conversation

@Hweinstock

@HweinstockHweinstock commented Jul 31, 2026

Copy link
Copy Markdown
Contributor

Problem

The AgentCore CLI is not currently publishing telemetry to our collector.

Solution

  • otel sdk requires the resource attributes separated from the metric attributes (even though they are merged on the request) so we explicitly inject those into each sink.
  • implement an otel histogram metric sink. We use histogram because our values represent duration, so histogram gives us count, and distributions over the values in a single instrument.

Testing

  • run our collector locally, and override local endpoint to point to it. Then verified metrics hit cloudwatch in dev account. (Requires one small backend change that is already in the pipeline).
  • added a unit test that spins up a local server, and verifies it only receives request when telemetry is enabled, and requests match otel format.

@github-actionsgithub-actionsBot added the agentcore-harness-reviewing AgentCore Harness review in progress label Jul 31, 2026
@github-actionsgithub-actionsBot removed the agentcore-harness-reviewing AgentCore Harness review in progress label Jul 31, 2026
@codecov-commenter

codecov-commenter commented Jul 31, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.48485% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 96.70%. Comparing base (7e29eda) to head (6806436).
⚠️ Report is 12 commits behind head on refactor.

Files with missing linesPatch %Lines
src/telemetry/otelSink.tsx98.07%1 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #1888 +/- ##
=========================================
Coverage 96.69% 96.70% =========================================
Files 291 292 +1 Lines 16012 16068 +56 =========================================
+ Hits 15483 15538 +55 - Misses 529 530 +1 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@Hweinstock

Copy link
Copy Markdown
ContributorAuthor

^ the line missing coverage is getName() on the collector sink.

@HweinstockHweinstock changed the title 0feat(tel): implement otel sinkfeat(tel): implement otel sinkJul 31, 2026
@Hweinstock
Hweinstock marked this pull request as ready for review July 31, 2026 19:22

async shutdown(): Promise<void> {
try {
await this.meterProvider.forceFlush({ timeoutMillis: this.flushTimeoutMs });

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

outside this try block, there is meterProvider.shutdown() which also flushes the metric again.

https://opentelemetry.io/docs/specs/otel/metrics/sdk/#shutdown

can we ensure one export somehow?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This method provides a way for provider to do any cleanup required.
Shutdown MUST be called only once for each MeterProvider instance. After the call to Shutdown, subsequent attempts to get a Meter are not allowed. SDKs SHOULD return a valid no-op Meter for these calls, if possible.
Shutdown SHOULD provide a way to let the caller know whether it succeeded, failed or timed out.
Shutdown SHOULD complete or abort within some timeout. Shutdown MAY be implemented as a blocking API or an asynchronous API which notifies the caller via a callback or an event. [OpenTelemetry SDK](https://opentelemetry.io/docs/specs/otel/overview/#sdk) authors MAY decide if they want to make the shutdown timeout configurable.
Shutdown MUST be implemented at least by invoking Shutdown on all registered [MetricReader](https://opentelemetry.io/docs/specs/otel/metrics/sdk/#metricreader) and [MetricExporter](https://opentelemetry.io/docs/specs/otel/metrics/sdk/#metricexporter) instances.

from https://opentelemetry.io/docs/specs/otel/metrics/sdk/#shutdown.

I don't see any explicit lines in the protocol linked for shutdown that it also flushes. Based on some testing, I think it does internally, but I think its safer to make that behavior explicit. If we flush twice, its a no-op anyway.


if (globalConfig.telemetry.enabled)
metricSinks.push(
new OtelHistogramSink({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like a wrong/malformed endpoint makes getMetricSinks() reject, and shutdown() propagates that rejection, erroring out in the CLI command. Is that understanding correct? can we make it best-effort?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we (the dev team) should be the only ones modifying the endpoint for testing purposes. In which case, I think the ideal behavior is that we reject early.

If a user decides to go into the global config and add an invalid override, I think rejecting is reasonable.

notgitika
notgitika previously approved these changes Aug 1, 2026

@notgitikanotgitika left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGMT

Comment threadsrc/telemetry/otelSink.tsx Outdated
try {
await this.meterProvider.forceFlush({ timeoutMillis: this.flushTimeoutMs });
} catch (e) {
const error = e instanceof Error ? e : new Error(String(e));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to use our AgentCoreError.fromError() here instead?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah I think that makes sense to ensure its consistent.

@Hweinstock

Copy link
Copy Markdown
ContributorAuthor

slack failure appears unrelated.

Configure AWS credentials (central secrets reader)
1m 0s
Run aws-actions/configure-aws-credentials@v6
Retry validateCredentials: attempt 1 of 12 failed: Credentials could not be loaded, please check your action inputs: Could not load credentials from any providers. Retrying after 23ms.

notgitika
notgitika previously approved these changes Aug 12, 2026
resource: resourceFromAttributes(config.resourceAttributes),
readers: [
new PeriodicExportingMetricReader({
exporter: new OTLPMetricExporter({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I may be missing whether inheriting the caller's OTEL configuration is intentional, but OTLPMetricExporter reads OTEL_EXPORTER_OTLP_HEADERS and OTEL_EXPORTER_OTLP_METRICS_HEADERS when headers is omitted. I set dummy authorization and x-api-key values and confirmed both were forwarded to the configured collector. A user running the CLI in an environment configured for another OTLP backend could therefore send those credentials to telemetry.agentcore.aws.dev. Would passing headers: {} here make sense?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nice catch, was not aware of these environment variables. Was going to add this as a follow-up, but let me inject the right header now to avoid this behavior.

.child({ metricName, metricValue: value, metricAttributes: attributes })
.info(`sending telemetry metric to collector`);

this.getHistogram(metricName).record(value, attributes);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the resource attributes are still ending up in both places. InMemoryMetricEvent.emit() merges resourceAttributes into the sink attributes, then this line records that merged map after the provider already registered the same fields with resourceFromAttributes. In a local OTLP capture, all eight resource keys appeared under both resource.attributes and histogram.dataPoints[].attributes. Since the goal here is to separate resource and metric attributes, would it make sense for the sink contract to receive only metric attributes, with FileSystemSink composing its JSONL entry from its configured resource attributes? The test could also assert that resource keys are absent from the datapoint. This is really a low finding and could always be changed in the future.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good catch, I think removing is the right call to simplify.

aidandaly24
aidandaly24 previously approved these changes Aug 12, 2026
@Hweinstock
Hweinstock dismissed stale reviews from aidandaly24 and notgitika via 7e9a97eAugust 13, 2026 13:05
@notgitika

Copy link
Copy Markdown
Contributor

looks good to me!

@Hweinstock
Hweinstock merged commit 01eb60b into aws:refactorAug 14, 2026
8 of 11 checks passed
@Hweinstock
Hweinstock deleted the feat/otel-sink branch August 14, 2026 19:09
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@Hweinstock@codecov-commenter@notgitika@AlexanderRichey@aidandaly24
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(tel): implement otel sink - #1888

Merged
Hweinstock merged 14 commits into
aws:refactorfrom
Hweinstock:feat/otel-sink
Aug 14, 2026
Merged

feat(tel): implement otel sink#1888
Hweinstock merged 14 commits into
aws:refactorfrom
Hweinstock:feat/otel-sink

Conversation

@Hweinstock

@HweinstockHweinstock commented Jul 31, 2026

Copy link
Copy Markdown
Contributor

Problem

The AgentCore CLI is not currently publishing telemetry to our collector.

Solution

  • otel sdk requires the resource attributes separated from the metric attributes (even though they are merged on the request) so we explicitly inject those into each sink.
  • implement an otel histogram metric sink. We use histogram because our values represent duration, so histogram gives us count, and distributions over the values in a single instrument.

Testing

  • run our collector locally, and override local endpoint to point to it. Then verified metrics hit cloudwatch in dev account. (Requires one small backend change that is already in the pipeline).
  • added a unit test that spins up a local server, and verifies it only receives request when telemetry is enabled, and requests match otel format.

@github-actionsgithub-actionsBot added the agentcore-harness-reviewing AgentCore Harness review in progress label Jul 31, 2026
@github-actionsgithub-actionsBot removed the agentcore-harness-reviewing AgentCore Harness review in progress label Jul 31, 2026
@codecov-commenter

codecov-commenter commented Jul 31, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.48485% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 96.70%. Comparing base (7e29eda) to head (6806436).
⚠️ Report is 12 commits behind head on refactor.

Files with missing linesPatch %Lines
src/telemetry/otelSink.tsx98.07%1 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #1888 +/- ##
=========================================
Coverage 96.69% 96.70% =========================================
Files 291 292 +1 Lines 16012 16068 +56 =========================================
+ Hits 15483 15538 +55 - Misses 529 530 +1 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@Hweinstock

Copy link
Copy Markdown
ContributorAuthor

^ the line missing coverage is getName() on the collector sink.

@HweinstockHweinstock changed the title 0feat(tel): implement otel sinkfeat(tel): implement otel sinkJul 31, 2026
@Hweinstock
Hweinstock marked this pull request as ready for review July 31, 2026 19:22

async shutdown(): Promise<void> {
try {
await this.meterProvider.forceFlush({ timeoutMillis: this.flushTimeoutMs });

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

outside this try block, there is meterProvider.shutdown() which also flushes the metric again.

https://opentelemetry.io/docs/specs/otel/metrics/sdk/#shutdown

can we ensure one export somehow?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This method provides a way for provider to do any cleanup required.
Shutdown MUST be called only once for each MeterProvider instance. After the call to Shutdown, subsequent attempts to get a Meter are not allowed. SDKs SHOULD return a valid no-op Meter for these calls, if possible.
Shutdown SHOULD provide a way to let the caller know whether it succeeded, failed or timed out.
Shutdown SHOULD complete or abort within some timeout. Shutdown MAY be implemented as a blocking API or an asynchronous API which notifies the caller via a callback or an event. [OpenTelemetry SDK](https://opentelemetry.io/docs/specs/otel/overview/#sdk) authors MAY decide if they want to make the shutdown timeout configurable.
Shutdown MUST be implemented at least by invoking Shutdown on all registered [MetricReader](https://opentelemetry.io/docs/specs/otel/metrics/sdk/#metricreader) and [MetricExporter](https://opentelemetry.io/docs/specs/otel/metrics/sdk/#metricexporter) instances.

from https://opentelemetry.io/docs/specs/otel/metrics/sdk/#shutdown.

I don't see any explicit lines in the protocol linked for shutdown that it also flushes. Based on some testing, I think it does internally, but I think its safer to make that behavior explicit. If we flush twice, its a no-op anyway.


if (globalConfig.telemetry.enabled)
metricSinks.push(
new OtelHistogramSink({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like a wrong/malformed endpoint makes getMetricSinks() reject, and shutdown() propagates that rejection, erroring out in the CLI command. Is that understanding correct? can we make it best-effort?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we (the dev team) should be the only ones modifying the endpoint for testing purposes. In which case, I think the ideal behavior is that we reject early.

If a user decides to go into the global config and add an invalid override, I think rejecting is reasonable.

notgitika
notgitika previously approved these changes Aug 1, 2026

@notgitikanotgitika left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGMT

Comment threadsrc/telemetry/otelSink.tsx Outdated
try {
await this.meterProvider.forceFlush({ timeoutMillis: this.flushTimeoutMs });
} catch (e) {
const error = e instanceof Error ? e : new Error(String(e));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to use our AgentCoreError.fromError() here instead?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah I think that makes sense to ensure its consistent.

@Hweinstock

Copy link
Copy Markdown
ContributorAuthor

slack failure appears unrelated.

Configure AWS credentials (central secrets reader)
1m 0s
Run aws-actions/configure-aws-credentials@v6
Retry validateCredentials: attempt 1 of 12 failed: Credentials could not be loaded, please check your action inputs: Could not load credentials from any providers. Retrying after 23ms.

notgitika
notgitika previously approved these changes Aug 12, 2026
resource: resourceFromAttributes(config.resourceAttributes),
readers: [
new PeriodicExportingMetricReader({
exporter: new OTLPMetricExporter({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I may be missing whether inheriting the caller's OTEL configuration is intentional, but OTLPMetricExporter reads OTEL_EXPORTER_OTLP_HEADERS and OTEL_EXPORTER_OTLP_METRICS_HEADERS when headers is omitted. I set dummy authorization and x-api-key values and confirmed both were forwarded to the configured collector. A user running the CLI in an environment configured for another OTLP backend could therefore send those credentials to telemetry.agentcore.aws.dev. Would passing headers: {} here make sense?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nice catch, was not aware of these environment variables. Was going to add this as a follow-up, but let me inject the right header now to avoid this behavior.

.child({ metricName, metricValue: value, metricAttributes: attributes })
.info(`sending telemetry metric to collector`);

this.getHistogram(metricName).record(value, attributes);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the resource attributes are still ending up in both places. InMemoryMetricEvent.emit() merges resourceAttributes into the sink attributes, then this line records that merged map after the provider already registered the same fields with resourceFromAttributes. In a local OTLP capture, all eight resource keys appeared under both resource.attributes and histogram.dataPoints[].attributes. Since the goal here is to separate resource and metric attributes, would it make sense for the sink contract to receive only metric attributes, with FileSystemSink composing its JSONL entry from its configured resource attributes? The test could also assert that resource keys are absent from the datapoint. This is really a low finding and could always be changed in the future.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good catch, I think removing is the right call to simplify.

aidandaly24
aidandaly24 previously approved these changes Aug 12, 2026
@Hweinstock
Hweinstock dismissed stale reviews from aidandaly24 and notgitika via 7e9a97eAugust 13, 2026 13:05
@notgitika

Copy link
Copy Markdown
Contributor

looks good to me!

@Hweinstock
Hweinstock merged commit 01eb60b into aws:refactorAug 14, 2026
8 of 11 checks passed
@Hweinstock
Hweinstock deleted the feat/otel-sink branch August 14, 2026 19:09
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@Hweinstock@codecov-commenter@notgitika@AlexanderRichey@aidandaly24
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(tel): implement otel sink - #1888

Merged
Hweinstock merged 14 commits into
aws:refactorfrom
Hweinstock:feat/otel-sink
Aug 14, 2026
Merged

feat(tel): implement otel sink#1888
Hweinstock merged 14 commits into
aws:refactorfrom
Hweinstock:feat/otel-sink

Conversation

@Hweinstock

@HweinstockHweinstock commented Jul 31, 2026

Copy link
Copy Markdown
Contributor

Problem

The AgentCore CLI is not currently publishing telemetry to our collector.

Solution

  • otel sdk requires the resource attributes separated from the metric attributes (even though they are merged on the request) so we explicitly inject those into each sink.
  • implement an otel histogram metric sink. We use histogram because our values represent duration, so histogram gives us count, and distributions over the values in a single instrument.

Testing

  • run our collector locally, and override local endpoint to point to it. Then verified metrics hit cloudwatch in dev account. (Requires one small backend change that is already in the pipeline).
  • added a unit test that spins up a local server, and verifies it only receives request when telemetry is enabled, and requests match otel format.

@github-actionsgithub-actionsBot added the agentcore-harness-reviewing AgentCore Harness review in progress label Jul 31, 2026
@github-actionsgithub-actionsBot removed the agentcore-harness-reviewing AgentCore Harness review in progress label Jul 31, 2026
@codecov-commenter

codecov-commenter commented Jul 31, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.48485% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 96.70%. Comparing base (7e29eda) to head (6806436).
⚠️ Report is 12 commits behind head on refactor.

Files with missing linesPatch %Lines
src/telemetry/otelSink.tsx98.07%1 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #1888 +/- ##
=========================================
Coverage 96.69% 96.70% =========================================
Files 291 292 +1 Lines 16012 16068 +56 =========================================
+ Hits 15483 15538 +55 - Misses 529 530 +1 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@Hweinstock

Copy link
Copy Markdown
ContributorAuthor

^ the line missing coverage is getName() on the collector sink.

@HweinstockHweinstock changed the title 0feat(tel): implement otel sinkfeat(tel): implement otel sinkJul 31, 2026
@Hweinstock
Hweinstock marked this pull request as ready for review July 31, 2026 19:22

async shutdown(): Promise<void> {
try {
await this.meterProvider.forceFlush({ timeoutMillis: this.flushTimeoutMs });

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

outside this try block, there is meterProvider.shutdown() which also flushes the metric again.

https://opentelemetry.io/docs/specs/otel/metrics/sdk/#shutdown

can we ensure one export somehow?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This method provides a way for provider to do any cleanup required.
Shutdown MUST be called only once for each MeterProvider instance. After the call to Shutdown, subsequent attempts to get a Meter are not allowed. SDKs SHOULD return a valid no-op Meter for these calls, if possible.
Shutdown SHOULD provide a way to let the caller know whether it succeeded, failed or timed out.
Shutdown SHOULD complete or abort within some timeout. Shutdown MAY be implemented as a blocking API or an asynchronous API which notifies the caller via a callback or an event. [OpenTelemetry SDK](https://opentelemetry.io/docs/specs/otel/overview/#sdk) authors MAY decide if they want to make the shutdown timeout configurable.
Shutdown MUST be implemented at least by invoking Shutdown on all registered [MetricReader](https://opentelemetry.io/docs/specs/otel/metrics/sdk/#metricreader) and [MetricExporter](https://opentelemetry.io/docs/specs/otel/metrics/sdk/#metricexporter) instances.

from https://opentelemetry.io/docs/specs/otel/metrics/sdk/#shutdown.

I don't see any explicit lines in the protocol linked for shutdown that it also flushes. Based on some testing, I think it does internally, but I think its safer to make that behavior explicit. If we flush twice, its a no-op anyway.


if (globalConfig.telemetry.enabled)
metricSinks.push(
new OtelHistogramSink({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like a wrong/malformed endpoint makes getMetricSinks() reject, and shutdown() propagates that rejection, erroring out in the CLI command. Is that understanding correct? can we make it best-effort?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we (the dev team) should be the only ones modifying the endpoint for testing purposes. In which case, I think the ideal behavior is that we reject early.

If a user decides to go into the global config and add an invalid override, I think rejecting is reasonable.

notgitika
notgitika previously approved these changes Aug 1, 2026

@notgitikanotgitika left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGMT

Comment threadsrc/telemetry/otelSink.tsx Outdated
try {
await this.meterProvider.forceFlush({ timeoutMillis: this.flushTimeoutMs });
} catch (e) {
const error = e instanceof Error ? e : new Error(String(e));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to use our AgentCoreError.fromError() here instead?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah I think that makes sense to ensure its consistent.

@Hweinstock

Copy link
Copy Markdown
ContributorAuthor

slack failure appears unrelated.

Configure AWS credentials (central secrets reader)
1m 0s
Run aws-actions/configure-aws-credentials@v6
Retry validateCredentials: attempt 1 of 12 failed: Credentials could not be loaded, please check your action inputs: Could not load credentials from any providers. Retrying after 23ms.

notgitika
notgitika previously approved these changes Aug 12, 2026
resource: resourceFromAttributes(config.resourceAttributes),
readers: [
new PeriodicExportingMetricReader({
exporter: new OTLPMetricExporter({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I may be missing whether inheriting the caller's OTEL configuration is intentional, but OTLPMetricExporter reads OTEL_EXPORTER_OTLP_HEADERS and OTEL_EXPORTER_OTLP_METRICS_HEADERS when headers is omitted. I set dummy authorization and x-api-key values and confirmed both were forwarded to the configured collector. A user running the CLI in an environment configured for another OTLP backend could therefore send those credentials to telemetry.agentcore.aws.dev. Would passing headers: {} here make sense?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nice catch, was not aware of these environment variables. Was going to add this as a follow-up, but let me inject the right header now to avoid this behavior.

.child({ metricName, metricValue: value, metricAttributes: attributes })
.info(`sending telemetry metric to collector`);

this.getHistogram(metricName).record(value, attributes);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the resource attributes are still ending up in both places. InMemoryMetricEvent.emit() merges resourceAttributes into the sink attributes, then this line records that merged map after the provider already registered the same fields with resourceFromAttributes. In a local OTLP capture, all eight resource keys appeared under both resource.attributes and histogram.dataPoints[].attributes. Since the goal here is to separate resource and metric attributes, would it make sense for the sink contract to receive only metric attributes, with FileSystemSink composing its JSONL entry from its configured resource attributes? The test could also assert that resource keys are absent from the datapoint. This is really a low finding and could always be changed in the future.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good catch, I think removing is the right call to simplify.

aidandaly24
aidandaly24 previously approved these changes Aug 12, 2026
@Hweinstock
Hweinstock dismissed stale reviews from aidandaly24 and notgitika via 7e9a97eAugust 13, 2026 13:05
@notgitika

Copy link
Copy Markdown
Contributor

looks good to me!

@Hweinstock
Hweinstock merged commit 01eb60b into aws:refactorAug 14, 2026
8 of 11 checks passed
@Hweinstock
Hweinstock deleted the feat/otel-sink branch August 14, 2026 19:09
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@Hweinstock@codecov-commenter@notgitika@AlexanderRichey@aidandaly24
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(tel): implement otel sink - #1888

Merged
Hweinstock merged 14 commits into
aws:refactorfrom
Hweinstock:feat/otel-sink
Aug 14, 2026
Merged

feat(tel): implement otel sink#1888
Hweinstock merged 14 commits into
aws:refactorfrom
Hweinstock:feat/otel-sink

Conversation

@Hweinstock

@HweinstockHweinstock commented Jul 31, 2026

Copy link
Copy Markdown
Contributor

Problem

The AgentCore CLI is not currently publishing telemetry to our collector.

Solution

  • otel sdk requires the resource attributes separated from the metric attributes (even though they are merged on the request) so we explicitly inject those into each sink.
  • implement an otel histogram metric sink. We use histogram because our values represent duration, so histogram gives us count, and distributions over the values in a single instrument.

Testing

  • run our collector locally, and override local endpoint to point to it. Then verified metrics hit cloudwatch in dev account. (Requires one small backend change that is already in the pipeline).
  • added a unit test that spins up a local server, and verifies it only receives request when telemetry is enabled, and requests match otel format.

@github-actionsgithub-actionsBot added the agentcore-harness-reviewing AgentCore Harness review in progress label Jul 31, 2026
@github-actionsgithub-actionsBot removed the agentcore-harness-reviewing AgentCore Harness review in progress label Jul 31, 2026
@codecov-commenter

codecov-commenter commented Jul 31, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.48485% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 96.70%. Comparing base (7e29eda) to head (6806436).
⚠️ Report is 12 commits behind head on refactor.

Files with missing linesPatch %Lines
src/telemetry/otelSink.tsx98.07%1 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## refactor #1888 +/- ##
=========================================
Coverage 96.69% 96.70% =========================================
Files 291 292 +1 Lines 16012 16068 +56 =========================================
+ Hits 15483 15538 +55 - Misses 529 530 +1 

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@Hweinstock

Copy link
Copy Markdown
ContributorAuthor

^ the line missing coverage is getName() on the collector sink.

@HweinstockHweinstock changed the title 0feat(tel): implement otel sinkfeat(tel): implement otel sinkJul 31, 2026
@Hweinstock
Hweinstock marked this pull request as ready for review July 31, 2026 19:22

async shutdown(): Promise<void> {
try {
await this.meterProvider.forceFlush({ timeoutMillis: this.flushTimeoutMs });

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

outside this try block, there is meterProvider.shutdown() which also flushes the metric again.

https://opentelemetry.io/docs/specs/otel/metrics/sdk/#shutdown

can we ensure one export somehow?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This method provides a way for provider to do any cleanup required.
Shutdown MUST be called only once for each MeterProvider instance. After the call to Shutdown, subsequent attempts to get a Meter are not allowed. SDKs SHOULD return a valid no-op Meter for these calls, if possible.
Shutdown SHOULD provide a way to let the caller know whether it succeeded, failed or timed out.
Shutdown SHOULD complete or abort within some timeout. Shutdown MAY be implemented as a blocking API or an asynchronous API which notifies the caller via a callback or an event. [OpenTelemetry SDK](https://opentelemetry.io/docs/specs/otel/overview/#sdk) authors MAY decide if they want to make the shutdown timeout configurable.
Shutdown MUST be implemented at least by invoking Shutdown on all registered [MetricReader](https://opentelemetry.io/docs/specs/otel/metrics/sdk/#metricreader) and [MetricExporter](https://opentelemetry.io/docs/specs/otel/metrics/sdk/#metricexporter) instances.

from https://opentelemetry.io/docs/specs/otel/metrics/sdk/#shutdown.

I don't see any explicit lines in the protocol linked for shutdown that it also flushes. Based on some testing, I think it does internally, but I think its safer to make that behavior explicit. If we flush twice, its a no-op anyway.


if (globalConfig.telemetry.enabled)
metricSinks.push(
new OtelHistogramSink({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it seems like a wrong/malformed endpoint makes getMetricSinks() reject, and shutdown() propagates that rejection, erroring out in the CLI command. Is that understanding correct? can we make it best-effort?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we (the dev team) should be the only ones modifying the endpoint for testing purposes. In which case, I think the ideal behavior is that we reject early.

If a user decides to go into the global config and add an invalid override, I think rejecting is reasonable.

notgitika
notgitika previously approved these changes Aug 1, 2026

@notgitikanotgitika left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGMT

Comment threadsrc/telemetry/otelSink.tsx Outdated
try {
await this.meterProvider.forceFlush({ timeoutMillis: this.flushTimeoutMs });
} catch (e) {
const error = e instanceof Error ? e : new Error(String(e));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to use our AgentCoreError.fromError() here instead?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah I think that makes sense to ensure its consistent.

@Hweinstock

Copy link
Copy Markdown
ContributorAuthor

slack failure appears unrelated.

Configure AWS credentials (central secrets reader)
1m 0s
Run aws-actions/configure-aws-credentials@v6
Retry validateCredentials: attempt 1 of 12 failed: Credentials could not be loaded, please check your action inputs: Could not load credentials from any providers. Retrying after 23ms.

notgitika
notgitika previously approved these changes Aug 12, 2026
resource: resourceFromAttributes(config.resourceAttributes),
readers: [
new PeriodicExportingMetricReader({
exporter: new OTLPMetricExporter({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I may be missing whether inheriting the caller's OTEL configuration is intentional, but OTLPMetricExporter reads OTEL_EXPORTER_OTLP_HEADERS and OTEL_EXPORTER_OTLP_METRICS_HEADERS when headers is omitted. I set dummy authorization and x-api-key values and confirmed both were forwarded to the configured collector. A user running the CLI in an environment configured for another OTLP backend could therefore send those credentials to telemetry.agentcore.aws.dev. Would passing headers: {} here make sense?

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nice catch, was not aware of these environment variables. Was going to add this as a follow-up, but let me inject the right header now to avoid this behavior.

.child({ metricName, metricValue: value, metricAttributes: attributes })
.info(`sending telemetry metric to collector`);

this.getHistogram(metricName).record(value, attributes);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the resource attributes are still ending up in both places. InMemoryMetricEvent.emit() merges resourceAttributes into the sink attributes, then this line records that merged map after the provider already registered the same fields with resourceFromAttributes. In a local OTLP capture, all eight resource keys appeared under both resource.attributes and histogram.dataPoints[].attributes. Since the goal here is to separate resource and metric attributes, would it make sense for the sink contract to receive only metric attributes, with FileSystemSink composing its JSONL entry from its configured resource attributes? The test could also assert that resource keys are absent from the datapoint. This is really a low finding and could always be changed in the future.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good catch, I think removing is the right call to simplify.

aidandaly24
aidandaly24 previously approved these changes Aug 12, 2026
@Hweinstock
Hweinstock dismissed stale reviews from aidandaly24 and notgitika via 7e9a97eAugust 13, 2026 13:05
@notgitika

Copy link
Copy Markdown
Contributor

looks good to me!

@Hweinstock
Hweinstock merged commit 01eb60b into aws:refactorAug 14, 2026
8 of 11 checks passed
@Hweinstock
Hweinstock deleted the feat/otel-sink branch August 14, 2026 19:09
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants

@Hweinstock@codecov-commenter@notgitika@AlexanderRichey@aidandaly24