feat: streaming-event telemetry collector + task.intent directive (0.5.6) - #7

Merged
drewstone merged 2 commits into
mainfrom
feat/stream-event-collector
May 10, 2026
Merged

feat: streaming-event telemetry collector + task.intent directive (0.5.6)#7
drewstone merged 2 commits into
mainfrom
feat/stream-event-collector

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Why

createRuntimeEventCollector only accepts AgentRuntimeEvent (the
sink-style events emitted by runAgentTask). It does not handle
RuntimeStreamEvent (the events yielded by runAgentTaskStream).
Every product agent that needs sanitized telemetry on top of streaming
has the same workaround in front of it: call
sanitizeRuntimeStreamEvent(event, options) event-by-event inside the
for await loop, then re-implement summary/aggregation by hand.

The gtm-agent reference migration (@tangle-network/agent-runtime is
already a dep at ^0.5.3, but no streaming integration has been wired
yet — it would hit this exact gap when it does) and any future product
agent moving from runAgentTask to runAgentTaskStream would all
re-derive the same boilerplate. Closing the gap in core now keeps the
opt-in redaction story consistent across both entry points.

API change

Adds createRuntimeStreamEventCollector(options?: RuntimeTelemetryOptions)
as a sibling factory to createRuntimeEventCollector. It returns:

{
events: Array<Record<string,unknown>>onEvent(event: RuntimeStreamEvent): voidsummary(): RuntimeStreamEventSummary}

Honors the same RuntimeTelemetryOptions redaction flags
(includeInputs, includeUserAnswers, includeControlPayloads,
includeEvidenceIds, includeMetadata,
includeRequirementDescriptions, includeEvalDetails). The
summary() rollup gives eventCount, eventCountsByType,
firstSessionId, finalStatus, finalReason, and concatenated
finalText from text_delta events.

Sibling factory vs unified union

I considered three shapes:

  1. Unified factory with a discriminated union over both event
    types.
    Rejected: the stream and non-stream events share type
    literals (task_start, readiness_end, task_end) but with
    different field shapes (timestamp and session on the stream
    side, knowledge decision on the stream side, etc.). A unified
    dispatcher would have to discriminate on the presence of optional
    fields, which is brittle and silently misroutes events. Consumers
    would also lose precise types at the callsite.
  2. Refactor the two event types into one shape and migrate both
    callers.
    Rejected: a much larger blast radius for a packaging-
    level improvement. Backward-incompatible for every existing
    runAgentTask consumer.
  3. Sibling factory. Picked. Clear typing, zero impact on existing
    createRuntimeEventCollector consumers, identical opt-in semantics.

README directive on task.intent

task.intent flows through sanitized telemetry by default (it's the
stable operation label, not redactable by includeInputs). The README
"Sanitized telemetry" section now states explicitly:

Never set task.intent to user input — use a fixed string
describing the operation kind (e.g. "Run a chat turn", "Score a tax return"). If you need to log user-visible intent, route it
through inputs (which are redacted by default) instead.

Same directive is repeated in the new example's README.

New example

examples/sanitized-telemetry-streaming/ mirrors
examples/sanitized-telemetry/ for streaming. Uses
createIterableBackend to yield a synthetic script (text_delta,
tool_call with sensitive args, tool_result with a secret token,
artifact with an internal s3 uri), runs cleanly with no creds, prints
both the default-redacted and verbose opt-in views plus the
summary() rollup. Wired into examples/README.md.

Test plan

Added three vitest cases in tests/runtime.test.ts:

  • Redaction contract (load-bearing). Runs a stream through the
    collector with sensitive tool_call.args, tool_result.result,
    artifact.uri, artifact.metadata, task.inputs, and
    task.metadata. Asserts the serialized events contain none of:
    rm -rf, sk-leaked, cat /etc/secret.txt, secret-bucket,
    cust-99, redact@example.com. This is the test that fails if we
    ever leak user input.
  • Opt-in surfaces specific fields. Same setup, with
    includeInputs/includeControlPayloads/includeEvidenceIds/includeMetadata
    on. Asserts the previously-redacted fields are now present.
  • Summary rollup. Asserts firstSessionId, finalStatus,
    finalReason, finalText, and reconciles
    eventCount === sum(eventCountsByType).

All 19 tests pass (16 existing + 3 new):

Test Files 1 passed (1)
Tests 19 passed (19)

pnpm typecheck and pnpm build also clean.

The new example runs cleanly:

pnpm dlx tsx examples/sanitized-telemetry-streaming/sanitized-telemetry-streaming.ts

Default view: "inputs":"[redacted]", "metadata":"[redacted]",
tool_call events have no args field, tool_result events have no
result field, artifact events omit uri/metadata. Verbose view:
all fields visible. (Note: existing examples all use the same
pnpm tsx invocation — tsx isn't a local bin, so pnpm dlx tsx
matches the pattern of every other example in the repo.)

Versioning note

This PR ships 0.5.4 → 0.5.6, intentionally skipping 0.5.5. PR #6
currently holds the 0.5.5 bump for the agent-eval / agent-knowledge
dep tree unification. If PR #6 lands first, the version delta becomes
0.5.5 → 0.5.6 (the file already reads 0.5.6 so no rebase action
needed). If this PR lands first, PR #6's 0.5.4 → 0.5.5 bump becomes
a no-op vs. main and that PR's author should rebase / collapse the
version bump. Coordinated with the PR #6 author via this note.

…5.6)
Add createRuntimeStreamEventCollector — a sibling of
createRuntimeEventCollector typed for RuntimeStreamEvent. Honors the
same RuntimeTelemetryOptions redaction flags (includeInputs,
includeUserAnswers, includeControlPayloads, includeEvidenceIds,
includeMetadata, includeRequirementDescriptions, includeEvalDetails)
and returns the same {events, onEvent} interface plus a summary()
function that rolls up event counts, session id, final status, and
concatenated text_delta.text.
Sibling factory rather than overload because stream and non-stream
events have different field shapes (timestamps, sessions, text/tool
deltas) and overlapping type literals (task_start, readiness_end, …) —
a unified dispatcher would silently misroute events.
Adds the streaming-collector example mirror at
examples/sanitized-telemetry-streaming/. Documents in README that
task.intent flows through sanitized telemetry by default and must
never carry user input; route user-visible intent through inputs
(redacted by default) instead.
Bumps 0.5.4 → 0.5.6 (intentionally skipping 0.5.5; PR #6 currently
holds 0.5.5 and is expected to land in series).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@drewstone
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat: streaming-event telemetry collector + task.intent directive (0.5.6) - #7

Merged
drewstone merged 2 commits into
mainfrom
feat/stream-event-collector
May 10, 2026
Merged

feat: streaming-event telemetry collector + task.intent directive (0.5.6)#7
drewstone merged 2 commits into
mainfrom
feat/stream-event-collector

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Why

createRuntimeEventCollector only accepts AgentRuntimeEvent (the
sink-style events emitted by runAgentTask). It does not handle
RuntimeStreamEvent (the events yielded by runAgentTaskStream).
Every product agent that needs sanitized telemetry on top of streaming
has the same workaround in front of it: call
sanitizeRuntimeStreamEvent(event, options) event-by-event inside the
for await loop, then re-implement summary/aggregation by hand.

The gtm-agent reference migration (@tangle-network/agent-runtime is
already a dep at ^0.5.3, but no streaming integration has been wired
yet — it would hit this exact gap when it does) and any future product
agent moving from runAgentTask to runAgentTaskStream would all
re-derive the same boilerplate. Closing the gap in core now keeps the
opt-in redaction story consistent across both entry points.

API change

Adds createRuntimeStreamEventCollector(options?: RuntimeTelemetryOptions)
as a sibling factory to createRuntimeEventCollector. It returns:

{
events: Array<Record<string,unknown>>onEvent(event: RuntimeStreamEvent): voidsummary(): RuntimeStreamEventSummary}

Honors the same RuntimeTelemetryOptions redaction flags
(includeInputs, includeUserAnswers, includeControlPayloads,
includeEvidenceIds, includeMetadata,
includeRequirementDescriptions, includeEvalDetails). The
summary() rollup gives eventCount, eventCountsByType,
firstSessionId, finalStatus, finalReason, and concatenated
finalText from text_delta events.

Sibling factory vs unified union

I considered three shapes:

  1. Unified factory with a discriminated union over both event
    types.
    Rejected: the stream and non-stream events share type
    literals (task_start, readiness_end, task_end) but with
    different field shapes (timestamp and session on the stream
    side, knowledge decision on the stream side, etc.). A unified
    dispatcher would have to discriminate on the presence of optional
    fields, which is brittle and silently misroutes events. Consumers
    would also lose precise types at the callsite.
  2. Refactor the two event types into one shape and migrate both
    callers.
    Rejected: a much larger blast radius for a packaging-
    level improvement. Backward-incompatible for every existing
    runAgentTask consumer.
  3. Sibling factory. Picked. Clear typing, zero impact on existing
    createRuntimeEventCollector consumers, identical opt-in semantics.

README directive on task.intent

task.intent flows through sanitized telemetry by default (it's the
stable operation label, not redactable by includeInputs). The README
"Sanitized telemetry" section now states explicitly:

Never set task.intent to user input — use a fixed string
describing the operation kind (e.g. "Run a chat turn", "Score a tax return"). If you need to log user-visible intent, route it
through inputs (which are redacted by default) instead.

Same directive is repeated in the new example's README.

New example

examples/sanitized-telemetry-streaming/ mirrors
examples/sanitized-telemetry/ for streaming. Uses
createIterableBackend to yield a synthetic script (text_delta,
tool_call with sensitive args, tool_result with a secret token,
artifact with an internal s3 uri), runs cleanly with no creds, prints
both the default-redacted and verbose opt-in views plus the
summary() rollup. Wired into examples/README.md.

Test plan

Added three vitest cases in tests/runtime.test.ts:

  • Redaction contract (load-bearing). Runs a stream through the
    collector with sensitive tool_call.args, tool_result.result,
    artifact.uri, artifact.metadata, task.inputs, and
    task.metadata. Asserts the serialized events contain none of:
    rm -rf, sk-leaked, cat /etc/secret.txt, secret-bucket,
    cust-99, redact@example.com. This is the test that fails if we
    ever leak user input.
  • Opt-in surfaces specific fields. Same setup, with
    includeInputs/includeControlPayloads/includeEvidenceIds/includeMetadata
    on. Asserts the previously-redacted fields are now present.
  • Summary rollup. Asserts firstSessionId, finalStatus,
    finalReason, finalText, and reconciles
    eventCount === sum(eventCountsByType).

All 19 tests pass (16 existing + 3 new):

Test Files 1 passed (1)
Tests 19 passed (19)

pnpm typecheck and pnpm build also clean.

The new example runs cleanly:

pnpm dlx tsx examples/sanitized-telemetry-streaming/sanitized-telemetry-streaming.ts

Default view: "inputs":"[redacted]", "metadata":"[redacted]",
tool_call events have no args field, tool_result events have no
result field, artifact events omit uri/metadata. Verbose view:
all fields visible. (Note: existing examples all use the same
pnpm tsx invocation — tsx isn't a local bin, so pnpm dlx tsx
matches the pattern of every other example in the repo.)

Versioning note

This PR ships 0.5.4 → 0.5.6, intentionally skipping 0.5.5. PR #6
currently holds the 0.5.5 bump for the agent-eval / agent-knowledge
dep tree unification. If PR #6 lands first, the version delta becomes
0.5.5 → 0.5.6 (the file already reads 0.5.6 so no rebase action
needed). If this PR lands first, PR #6's 0.5.4 → 0.5.5 bump becomes
a no-op vs. main and that PR's author should rebase / collapse the
version bump. Coordinated with the PR #6 author via this note.

…5.6)
Add createRuntimeStreamEventCollector — a sibling of
createRuntimeEventCollector typed for RuntimeStreamEvent. Honors the
same RuntimeTelemetryOptions redaction flags (includeInputs,
includeUserAnswers, includeControlPayloads, includeEvidenceIds,
includeMetadata, includeRequirementDescriptions, includeEvalDetails)
and returns the same {events, onEvent} interface plus a summary()
function that rolls up event counts, session id, final status, and
concatenated text_delta.text.
Sibling factory rather than overload because stream and non-stream
events have different field shapes (timestamps, sessions, text/tool
deltas) and overlapping type literals (task_start, readiness_end, …) —
a unified dispatcher would silently misroute events.
Adds the streaming-collector example mirror at
examples/sanitized-telemetry-streaming/. Documents in README that
task.intent flows through sanitized telemetry by default and must
never carry user input; route user-visible intent through inputs
(redacted by default) instead.
Bumps 0.5.4 → 0.5.6 (intentionally skipping 0.5.5; PR #6 currently
holds 0.5.5 and is expected to land in series).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@drewstone
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: streaming-event telemetry collector + task.intent directive (0.5.6) - #7

Merged
drewstone merged 2 commits into
mainfrom
feat/stream-event-collector
May 10, 2026
Merged

feat: streaming-event telemetry collector + task.intent directive (0.5.6)#7
drewstone merged 2 commits into
mainfrom
feat/stream-event-collector

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Why

createRuntimeEventCollector only accepts AgentRuntimeEvent (the
sink-style events emitted by runAgentTask). It does not handle
RuntimeStreamEvent (the events yielded by runAgentTaskStream).
Every product agent that needs sanitized telemetry on top of streaming
has the same workaround in front of it: call
sanitizeRuntimeStreamEvent(event, options) event-by-event inside the
for await loop, then re-implement summary/aggregation by hand.

The gtm-agent reference migration (@tangle-network/agent-runtime is
already a dep at ^0.5.3, but no streaming integration has been wired
yet — it would hit this exact gap when it does) and any future product
agent moving from runAgentTask to runAgentTaskStream would all
re-derive the same boilerplate. Closing the gap in core now keeps the
opt-in redaction story consistent across both entry points.

API change

Adds createRuntimeStreamEventCollector(options?: RuntimeTelemetryOptions)
as a sibling factory to createRuntimeEventCollector. It returns:

{
events: Array<Record<string,unknown>>onEvent(event: RuntimeStreamEvent): voidsummary(): RuntimeStreamEventSummary}

Honors the same RuntimeTelemetryOptions redaction flags
(includeInputs, includeUserAnswers, includeControlPayloads,
includeEvidenceIds, includeMetadata,
includeRequirementDescriptions, includeEvalDetails). The
summary() rollup gives eventCount, eventCountsByType,
firstSessionId, finalStatus, finalReason, and concatenated
finalText from text_delta events.

Sibling factory vs unified union

I considered three shapes:

  1. Unified factory with a discriminated union over both event
    types.
    Rejected: the stream and non-stream events share type
    literals (task_start, readiness_end, task_end) but with
    different field shapes (timestamp and session on the stream
    side, knowledge decision on the stream side, etc.). A unified
    dispatcher would have to discriminate on the presence of optional
    fields, which is brittle and silently misroutes events. Consumers
    would also lose precise types at the callsite.
  2. Refactor the two event types into one shape and migrate both
    callers.
    Rejected: a much larger blast radius for a packaging-
    level improvement. Backward-incompatible for every existing
    runAgentTask consumer.
  3. Sibling factory. Picked. Clear typing, zero impact on existing
    createRuntimeEventCollector consumers, identical opt-in semantics.

README directive on task.intent

task.intent flows through sanitized telemetry by default (it's the
stable operation label, not redactable by includeInputs). The README
"Sanitized telemetry" section now states explicitly:

Never set task.intent to user input — use a fixed string
describing the operation kind (e.g. "Run a chat turn", "Score a tax return"). If you need to log user-visible intent, route it
through inputs (which are redacted by default) instead.

Same directive is repeated in the new example's README.

New example

examples/sanitized-telemetry-streaming/ mirrors
examples/sanitized-telemetry/ for streaming. Uses
createIterableBackend to yield a synthetic script (text_delta,
tool_call with sensitive args, tool_result with a secret token,
artifact with an internal s3 uri), runs cleanly with no creds, prints
both the default-redacted and verbose opt-in views plus the
summary() rollup. Wired into examples/README.md.

Test plan

Added three vitest cases in tests/runtime.test.ts:

  • Redaction contract (load-bearing). Runs a stream through the
    collector with sensitive tool_call.args, tool_result.result,
    artifact.uri, artifact.metadata, task.inputs, and
    task.metadata. Asserts the serialized events contain none of:
    rm -rf, sk-leaked, cat /etc/secret.txt, secret-bucket,
    cust-99, redact@example.com. This is the test that fails if we
    ever leak user input.
  • Opt-in surfaces specific fields. Same setup, with
    includeInputs/includeControlPayloads/includeEvidenceIds/includeMetadata
    on. Asserts the previously-redacted fields are now present.
  • Summary rollup. Asserts firstSessionId, finalStatus,
    finalReason, finalText, and reconciles
    eventCount === sum(eventCountsByType).

All 19 tests pass (16 existing + 3 new):

Test Files 1 passed (1)
Tests 19 passed (19)

pnpm typecheck and pnpm build also clean.

The new example runs cleanly:

pnpm dlx tsx examples/sanitized-telemetry-streaming/sanitized-telemetry-streaming.ts

Default view: "inputs":"[redacted]", "metadata":"[redacted]",
tool_call events have no args field, tool_result events have no
result field, artifact events omit uri/metadata. Verbose view:
all fields visible. (Note: existing examples all use the same
pnpm tsx invocation — tsx isn't a local bin, so pnpm dlx tsx
matches the pattern of every other example in the repo.)

Versioning note

This PR ships 0.5.4 → 0.5.6, intentionally skipping 0.5.5. PR #6
currently holds the 0.5.5 bump for the agent-eval / agent-knowledge
dep tree unification. If PR #6 lands first, the version delta becomes
0.5.5 → 0.5.6 (the file already reads 0.5.6 so no rebase action
needed). If this PR lands first, PR #6's 0.5.4 → 0.5.5 bump becomes
a no-op vs. main and that PR's author should rebase / collapse the
version bump. Coordinated with the PR #6 author via this note.

…5.6)
Add createRuntimeStreamEventCollector — a sibling of
createRuntimeEventCollector typed for RuntimeStreamEvent. Honors the
same RuntimeTelemetryOptions redaction flags (includeInputs,
includeUserAnswers, includeControlPayloads, includeEvidenceIds,
includeMetadata, includeRequirementDescriptions, includeEvalDetails)
and returns the same {events, onEvent} interface plus a summary()
function that rolls up event counts, session id, final status, and
concatenated text_delta.text.
Sibling factory rather than overload because stream and non-stream
events have different field shapes (timestamps, sessions, text/tool
deltas) and overlapping type literals (task_start, readiness_end, …) —
a unified dispatcher would silently misroute events.
Adds the streaming-collector example mirror at
examples/sanitized-telemetry-streaming/. Documents in README that
task.intent flows through sanitized telemetry by default and must
never carry user input; route user-visible intent through inputs
(redacted by default) instead.
Bumps 0.5.4 → 0.5.6 (intentionally skipping 0.5.5; PR #6 currently
holds 0.5.5 and is expected to land in series).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@drewstone
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: streaming-event telemetry collector + task.intent directive (0.5.6) - #7

Merged
drewstone merged 2 commits into
mainfrom
feat/stream-event-collector
May 10, 2026
Merged

feat: streaming-event telemetry collector + task.intent directive (0.5.6)#7
drewstone merged 2 commits into
mainfrom
feat/stream-event-collector

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Why

createRuntimeEventCollector only accepts AgentRuntimeEvent (the
sink-style events emitted by runAgentTask). It does not handle
RuntimeStreamEvent (the events yielded by runAgentTaskStream).
Every product agent that needs sanitized telemetry on top of streaming
has the same workaround in front of it: call
sanitizeRuntimeStreamEvent(event, options) event-by-event inside the
for await loop, then re-implement summary/aggregation by hand.

The gtm-agent reference migration (@tangle-network/agent-runtime is
already a dep at ^0.5.3, but no streaming integration has been wired
yet — it would hit this exact gap when it does) and any future product
agent moving from runAgentTask to runAgentTaskStream would all
re-derive the same boilerplate. Closing the gap in core now keeps the
opt-in redaction story consistent across both entry points.

API change

Adds createRuntimeStreamEventCollector(options?: RuntimeTelemetryOptions)
as a sibling factory to createRuntimeEventCollector. It returns:

{
events: Array<Record<string,unknown>>onEvent(event: RuntimeStreamEvent): voidsummary(): RuntimeStreamEventSummary}

Honors the same RuntimeTelemetryOptions redaction flags
(includeInputs, includeUserAnswers, includeControlPayloads,
includeEvidenceIds, includeMetadata,
includeRequirementDescriptions, includeEvalDetails). The
summary() rollup gives eventCount, eventCountsByType,
firstSessionId, finalStatus, finalReason, and concatenated
finalText from text_delta events.

Sibling factory vs unified union

I considered three shapes:

  1. Unified factory with a discriminated union over both event
    types.
    Rejected: the stream and non-stream events share type
    literals (task_start, readiness_end, task_end) but with
    different field shapes (timestamp and session on the stream
    side, knowledge decision on the stream side, etc.). A unified
    dispatcher would have to discriminate on the presence of optional
    fields, which is brittle and silently misroutes events. Consumers
    would also lose precise types at the callsite.
  2. Refactor the two event types into one shape and migrate both
    callers.
    Rejected: a much larger blast radius for a packaging-
    level improvement. Backward-incompatible for every existing
    runAgentTask consumer.
  3. Sibling factory. Picked. Clear typing, zero impact on existing
    createRuntimeEventCollector consumers, identical opt-in semantics.

README directive on task.intent

task.intent flows through sanitized telemetry by default (it's the
stable operation label, not redactable by includeInputs). The README
"Sanitized telemetry" section now states explicitly:

Never set task.intent to user input — use a fixed string
describing the operation kind (e.g. "Run a chat turn", "Score a tax return"). If you need to log user-visible intent, route it
through inputs (which are redacted by default) instead.

Same directive is repeated in the new example's README.

New example

examples/sanitized-telemetry-streaming/ mirrors
examples/sanitized-telemetry/ for streaming. Uses
createIterableBackend to yield a synthetic script (text_delta,
tool_call with sensitive args, tool_result with a secret token,
artifact with an internal s3 uri), runs cleanly with no creds, prints
both the default-redacted and verbose opt-in views plus the
summary() rollup. Wired into examples/README.md.

Test plan

Added three vitest cases in tests/runtime.test.ts:

  • Redaction contract (load-bearing). Runs a stream through the
    collector with sensitive tool_call.args, tool_result.result,
    artifact.uri, artifact.metadata, task.inputs, and
    task.metadata. Asserts the serialized events contain none of:
    rm -rf, sk-leaked, cat /etc/secret.txt, secret-bucket,
    cust-99, redact@example.com. This is the test that fails if we
    ever leak user input.
  • Opt-in surfaces specific fields. Same setup, with
    includeInputs/includeControlPayloads/includeEvidenceIds/includeMetadata
    on. Asserts the previously-redacted fields are now present.
  • Summary rollup. Asserts firstSessionId, finalStatus,
    finalReason, finalText, and reconciles
    eventCount === sum(eventCountsByType).

All 19 tests pass (16 existing + 3 new):

Test Files 1 passed (1)
Tests 19 passed (19)

pnpm typecheck and pnpm build also clean.

The new example runs cleanly:

pnpm dlx tsx examples/sanitized-telemetry-streaming/sanitized-telemetry-streaming.ts

Default view: "inputs":"[redacted]", "metadata":"[redacted]",
tool_call events have no args field, tool_result events have no
result field, artifact events omit uri/metadata. Verbose view:
all fields visible. (Note: existing examples all use the same
pnpm tsx invocation — tsx isn't a local bin, so pnpm dlx tsx
matches the pattern of every other example in the repo.)

Versioning note

This PR ships 0.5.4 → 0.5.6, intentionally skipping 0.5.5. PR #6
currently holds the 0.5.5 bump for the agent-eval / agent-knowledge
dep tree unification. If PR #6 lands first, the version delta becomes
0.5.5 → 0.5.6 (the file already reads 0.5.6 so no rebase action
needed). If this PR lands first, PR #6's 0.5.4 → 0.5.5 bump becomes
a no-op vs. main and that PR's author should rebase / collapse the
version bump. Coordinated with the PR #6 author via this note.

…5.6)
Add createRuntimeStreamEventCollector — a sibling of
createRuntimeEventCollector typed for RuntimeStreamEvent. Honors the
same RuntimeTelemetryOptions redaction flags (includeInputs,
includeUserAnswers, includeControlPayloads, includeEvidenceIds,
includeMetadata, includeRequirementDescriptions, includeEvalDetails)
and returns the same {events, onEvent} interface plus a summary()
function that rolls up event counts, session id, final status, and
concatenated text_delta.text.
Sibling factory rather than overload because stream and non-stream
events have different field shapes (timestamps, sessions, text/tool
deltas) and overlapping type literals (task_start, readiness_end, …) —
a unified dispatcher would silently misroute events.
Adds the streaming-collector example mirror at
examples/sanitized-telemetry-streaming/. Documents in README that
task.intent flows through sanitized telemetry by default and must
never carry user input; route user-visible intent through inputs
(redacted by default) instead.
Bumps 0.5.4 → 0.5.6 (intentionally skipping 0.5.5; PR #6 currently
holds 0.5.5 and is expected to land in series).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@drewstone
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat: streaming-event telemetry collector + task.intent directive (0.5.6) - #7

Merged
drewstone merged 2 commits into
mainfrom
feat/stream-event-collector
May 10, 2026
Merged

feat: streaming-event telemetry collector + task.intent directive (0.5.6)#7
drewstone merged 2 commits into
mainfrom
feat/stream-event-collector

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Why

createRuntimeEventCollector only accepts AgentRuntimeEvent (the
sink-style events emitted by runAgentTask). It does not handle
RuntimeStreamEvent (the events yielded by runAgentTaskStream).
Every product agent that needs sanitized telemetry on top of streaming
has the same workaround in front of it: call
sanitizeRuntimeStreamEvent(event, options) event-by-event inside the
for await loop, then re-implement summary/aggregation by hand.

The gtm-agent reference migration (@tangle-network/agent-runtime is
already a dep at ^0.5.3, but no streaming integration has been wired
yet — it would hit this exact gap when it does) and any future product
agent moving from runAgentTask to runAgentTaskStream would all
re-derive the same boilerplate. Closing the gap in core now keeps the
opt-in redaction story consistent across both entry points.

API change

Adds createRuntimeStreamEventCollector(options?: RuntimeTelemetryOptions)
as a sibling factory to createRuntimeEventCollector. It returns:

{
events: Array<Record<string,unknown>>onEvent(event: RuntimeStreamEvent): voidsummary(): RuntimeStreamEventSummary}

Honors the same RuntimeTelemetryOptions redaction flags
(includeInputs, includeUserAnswers, includeControlPayloads,
includeEvidenceIds, includeMetadata,
includeRequirementDescriptions, includeEvalDetails). The
summary() rollup gives eventCount, eventCountsByType,
firstSessionId, finalStatus, finalReason, and concatenated
finalText from text_delta events.

Sibling factory vs unified union

I considered three shapes:

  1. Unified factory with a discriminated union over both event
    types.
    Rejected: the stream and non-stream events share type
    literals (task_start, readiness_end, task_end) but with
    different field shapes (timestamp and session on the stream
    side, knowledge decision on the stream side, etc.). A unified
    dispatcher would have to discriminate on the presence of optional
    fields, which is brittle and silently misroutes events. Consumers
    would also lose precise types at the callsite.
  2. Refactor the two event types into one shape and migrate both
    callers.
    Rejected: a much larger blast radius for a packaging-
    level improvement. Backward-incompatible for every existing
    runAgentTask consumer.
  3. Sibling factory. Picked. Clear typing, zero impact on existing
    createRuntimeEventCollector consumers, identical opt-in semantics.

README directive on task.intent

task.intent flows through sanitized telemetry by default (it's the
stable operation label, not redactable by includeInputs). The README
"Sanitized telemetry" section now states explicitly:

Never set task.intent to user input — use a fixed string
describing the operation kind (e.g. "Run a chat turn", "Score a tax return"). If you need to log user-visible intent, route it
through inputs (which are redacted by default) instead.

Same directive is repeated in the new example's README.

New example

examples/sanitized-telemetry-streaming/ mirrors
examples/sanitized-telemetry/ for streaming. Uses
createIterableBackend to yield a synthetic script (text_delta,
tool_call with sensitive args, tool_result with a secret token,
artifact with an internal s3 uri), runs cleanly with no creds, prints
both the default-redacted and verbose opt-in views plus the
summary() rollup. Wired into examples/README.md.

Test plan

Added three vitest cases in tests/runtime.test.ts:

  • Redaction contract (load-bearing). Runs a stream through the
    collector with sensitive tool_call.args, tool_result.result,
    artifact.uri, artifact.metadata, task.inputs, and
    task.metadata. Asserts the serialized events contain none of:
    rm -rf, sk-leaked, cat /etc/secret.txt, secret-bucket,
    cust-99, redact@example.com. This is the test that fails if we
    ever leak user input.
  • Opt-in surfaces specific fields. Same setup, with
    includeInputs/includeControlPayloads/includeEvidenceIds/includeMetadata
    on. Asserts the previously-redacted fields are now present.
  • Summary rollup. Asserts firstSessionId, finalStatus,
    finalReason, finalText, and reconciles
    eventCount === sum(eventCountsByType).

All 19 tests pass (16 existing + 3 new):

Test Files 1 passed (1)
Tests 19 passed (19)

pnpm typecheck and pnpm build also clean.

The new example runs cleanly:

pnpm dlx tsx examples/sanitized-telemetry-streaming/sanitized-telemetry-streaming.ts

Default view: "inputs":"[redacted]", "metadata":"[redacted]",
tool_call events have no args field, tool_result events have no
result field, artifact events omit uri/metadata. Verbose view:
all fields visible. (Note: existing examples all use the same
pnpm tsx invocation — tsx isn't a local bin, so pnpm dlx tsx
matches the pattern of every other example in the repo.)

Versioning note

This PR ships 0.5.4 → 0.5.6, intentionally skipping 0.5.5. PR #6
currently holds the 0.5.5 bump for the agent-eval / agent-knowledge
dep tree unification. If PR #6 lands first, the version delta becomes
0.5.5 → 0.5.6 (the file already reads 0.5.6 so no rebase action
needed). If this PR lands first, PR #6's 0.5.4 → 0.5.5 bump becomes
a no-op vs. main and that PR's author should rebase / collapse the
version bump. Coordinated with the PR #6 author via this note.

…5.6)
Add createRuntimeStreamEventCollector — a sibling of
createRuntimeEventCollector typed for RuntimeStreamEvent. Honors the
same RuntimeTelemetryOptions redaction flags (includeInputs,
includeUserAnswers, includeControlPayloads, includeEvidenceIds,
includeMetadata, includeRequirementDescriptions, includeEvalDetails)
and returns the same {events, onEvent} interface plus a summary()
function that rolls up event counts, session id, final status, and
concatenated text_delta.text.
Sibling factory rather than overload because stream and non-stream
events have different field shapes (timestamps, sessions, text/tool
deltas) and overlapping type literals (task_start, readiness_end, …) —
a unified dispatcher would silently misroute events.
Adds the streaming-collector example mirror at
examples/sanitized-telemetry-streaming/. Documents in README that
task.intent flows through sanitized telemetry by default and must
never carry user input; route user-visible intent through inputs
(redacted by default) instead.
Bumps 0.5.4 → 0.5.6 (intentionally skipping 0.5.5; PR #6 currently
holds 0.5.5 and is expected to land in series).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@drewstone
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: streaming-event telemetry collector + task.intent directive (0.5.6) - #7

Merged
drewstone merged 2 commits into
mainfrom
feat/stream-event-collector
May 10, 2026
Merged

feat: streaming-event telemetry collector + task.intent directive (0.5.6)#7
drewstone merged 2 commits into
mainfrom
feat/stream-event-collector

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Why

createRuntimeEventCollector only accepts AgentRuntimeEvent (the
sink-style events emitted by runAgentTask). It does not handle
RuntimeStreamEvent (the events yielded by runAgentTaskStream).
Every product agent that needs sanitized telemetry on top of streaming
has the same workaround in front of it: call
sanitizeRuntimeStreamEvent(event, options) event-by-event inside the
for await loop, then re-implement summary/aggregation by hand.

The gtm-agent reference migration (@tangle-network/agent-runtime is
already a dep at ^0.5.3, but no streaming integration has been wired
yet — it would hit this exact gap when it does) and any future product
agent moving from runAgentTask to runAgentTaskStream would all
re-derive the same boilerplate. Closing the gap in core now keeps the
opt-in redaction story consistent across both entry points.

API change

Adds createRuntimeStreamEventCollector(options?: RuntimeTelemetryOptions)
as a sibling factory to createRuntimeEventCollector. It returns:

{
events: Array<Record<string,unknown>>onEvent(event: RuntimeStreamEvent): voidsummary(): RuntimeStreamEventSummary}

Honors the same RuntimeTelemetryOptions redaction flags
(includeInputs, includeUserAnswers, includeControlPayloads,
includeEvidenceIds, includeMetadata,
includeRequirementDescriptions, includeEvalDetails). The
summary() rollup gives eventCount, eventCountsByType,
firstSessionId, finalStatus, finalReason, and concatenated
finalText from text_delta events.

Sibling factory vs unified union

I considered three shapes:

  1. Unified factory with a discriminated union over both event
    types.
    Rejected: the stream and non-stream events share type
    literals (task_start, readiness_end, task_end) but with
    different field shapes (timestamp and session on the stream
    side, knowledge decision on the stream side, etc.). A unified
    dispatcher would have to discriminate on the presence of optional
    fields, which is brittle and silently misroutes events. Consumers
    would also lose precise types at the callsite.
  2. Refactor the two event types into one shape and migrate both
    callers.
    Rejected: a much larger blast radius for a packaging-
    level improvement. Backward-incompatible for every existing
    runAgentTask consumer.
  3. Sibling factory. Picked. Clear typing, zero impact on existing
    createRuntimeEventCollector consumers, identical opt-in semantics.

README directive on task.intent

task.intent flows through sanitized telemetry by default (it's the
stable operation label, not redactable by includeInputs). The README
"Sanitized telemetry" section now states explicitly:

Never set task.intent to user input — use a fixed string
describing the operation kind (e.g. "Run a chat turn", "Score a tax return"). If you need to log user-visible intent, route it
through inputs (which are redacted by default) instead.

Same directive is repeated in the new example's README.

New example

examples/sanitized-telemetry-streaming/ mirrors
examples/sanitized-telemetry/ for streaming. Uses
createIterableBackend to yield a synthetic script (text_delta,
tool_call with sensitive args, tool_result with a secret token,
artifact with an internal s3 uri), runs cleanly with no creds, prints
both the default-redacted and verbose opt-in views plus the
summary() rollup. Wired into examples/README.md.

Test plan

Added three vitest cases in tests/runtime.test.ts:

  • Redaction contract (load-bearing). Runs a stream through the
    collector with sensitive tool_call.args, tool_result.result,
    artifact.uri, artifact.metadata, task.inputs, and
    task.metadata. Asserts the serialized events contain none of:
    rm -rf, sk-leaked, cat /etc/secret.txt, secret-bucket,
    cust-99, redact@example.com. This is the test that fails if we
    ever leak user input.
  • Opt-in surfaces specific fields. Same setup, with
    includeInputs/includeControlPayloads/includeEvidenceIds/includeMetadata
    on. Asserts the previously-redacted fields are now present.
  • Summary rollup. Asserts firstSessionId, finalStatus,
    finalReason, finalText, and reconciles
    eventCount === sum(eventCountsByType).

All 19 tests pass (16 existing + 3 new):

Test Files 1 passed (1)
Tests 19 passed (19)

pnpm typecheck and pnpm build also clean.

The new example runs cleanly:

pnpm dlx tsx examples/sanitized-telemetry-streaming/sanitized-telemetry-streaming.ts

Default view: "inputs":"[redacted]", "metadata":"[redacted]",
tool_call events have no args field, tool_result events have no
result field, artifact events omit uri/metadata. Verbose view:
all fields visible. (Note: existing examples all use the same
pnpm tsx invocation — tsx isn't a local bin, so pnpm dlx tsx
matches the pattern of every other example in the repo.)

Versioning note

This PR ships 0.5.4 → 0.5.6, intentionally skipping 0.5.5. PR #6
currently holds the 0.5.5 bump for the agent-eval / agent-knowledge
dep tree unification. If PR #6 lands first, the version delta becomes
0.5.5 → 0.5.6 (the file already reads 0.5.6 so no rebase action
needed). If this PR lands first, PR #6's 0.5.4 → 0.5.5 bump becomes
a no-op vs. main and that PR's author should rebase / collapse the
version bump. Coordinated with the PR #6 author via this note.

…5.6)
Add createRuntimeStreamEventCollector — a sibling of
createRuntimeEventCollector typed for RuntimeStreamEvent. Honors the
same RuntimeTelemetryOptions redaction flags (includeInputs,
includeUserAnswers, includeControlPayloads, includeEvidenceIds,
includeMetadata, includeRequirementDescriptions, includeEvalDetails)
and returns the same {events, onEvent} interface plus a summary()
function that rolls up event counts, session id, final status, and
concatenated text_delta.text.
Sibling factory rather than overload because stream and non-stream
events have different field shapes (timestamps, sessions, text/tool
deltas) and overlapping type literals (task_start, readiness_end, …) —
a unified dispatcher would silently misroute events.
Adds the streaming-collector example mirror at
examples/sanitized-telemetry-streaming/. Documents in README that
task.intent flows through sanitized telemetry by default and must
never carry user input; route user-visible intent through inputs
(redacted by default) instead.
Bumps 0.5.4 → 0.5.6 (intentionally skipping 0.5.5; PR #6 currently
holds 0.5.5 and is expected to land in series).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@drewstone
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: streaming-event telemetry collector + task.intent directive (0.5.6) - #7

Merged
drewstone merged 2 commits into
mainfrom
feat/stream-event-collector
May 10, 2026
Merged

feat: streaming-event telemetry collector + task.intent directive (0.5.6)#7
drewstone merged 2 commits into
mainfrom
feat/stream-event-collector

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Why

createRuntimeEventCollector only accepts AgentRuntimeEvent (the
sink-style events emitted by runAgentTask). It does not handle
RuntimeStreamEvent (the events yielded by runAgentTaskStream).
Every product agent that needs sanitized telemetry on top of streaming
has the same workaround in front of it: call
sanitizeRuntimeStreamEvent(event, options) event-by-event inside the
for await loop, then re-implement summary/aggregation by hand.

The gtm-agent reference migration (@tangle-network/agent-runtime is
already a dep at ^0.5.3, but no streaming integration has been wired
yet — it would hit this exact gap when it does) and any future product
agent moving from runAgentTask to runAgentTaskStream would all
re-derive the same boilerplate. Closing the gap in core now keeps the
opt-in redaction story consistent across both entry points.

API change

Adds createRuntimeStreamEventCollector(options?: RuntimeTelemetryOptions)
as a sibling factory to createRuntimeEventCollector. It returns:

{
events: Array<Record<string,unknown>>onEvent(event: RuntimeStreamEvent): voidsummary(): RuntimeStreamEventSummary}

Honors the same RuntimeTelemetryOptions redaction flags
(includeInputs, includeUserAnswers, includeControlPayloads,
includeEvidenceIds, includeMetadata,
includeRequirementDescriptions, includeEvalDetails). The
summary() rollup gives eventCount, eventCountsByType,
firstSessionId, finalStatus, finalReason, and concatenated
finalText from text_delta events.

Sibling factory vs unified union

I considered three shapes:

  1. Unified factory with a discriminated union over both event
    types.
    Rejected: the stream and non-stream events share type
    literals (task_start, readiness_end, task_end) but with
    different field shapes (timestamp and session on the stream
    side, knowledge decision on the stream side, etc.). A unified
    dispatcher would have to discriminate on the presence of optional
    fields, which is brittle and silently misroutes events. Consumers
    would also lose precise types at the callsite.
  2. Refactor the two event types into one shape and migrate both
    callers.
    Rejected: a much larger blast radius for a packaging-
    level improvement. Backward-incompatible for every existing
    runAgentTask consumer.
  3. Sibling factory. Picked. Clear typing, zero impact on existing
    createRuntimeEventCollector consumers, identical opt-in semantics.

README directive on task.intent

task.intent flows through sanitized telemetry by default (it's the
stable operation label, not redactable by includeInputs). The README
"Sanitized telemetry" section now states explicitly:

Never set task.intent to user input — use a fixed string
describing the operation kind (e.g. "Run a chat turn", "Score a tax return"). If you need to log user-visible intent, route it
through inputs (which are redacted by default) instead.

Same directive is repeated in the new example's README.

New example

examples/sanitized-telemetry-streaming/ mirrors
examples/sanitized-telemetry/ for streaming. Uses
createIterableBackend to yield a synthetic script (text_delta,
tool_call with sensitive args, tool_result with a secret token,
artifact with an internal s3 uri), runs cleanly with no creds, prints
both the default-redacted and verbose opt-in views plus the
summary() rollup. Wired into examples/README.md.

Test plan

Added three vitest cases in tests/runtime.test.ts:

  • Redaction contract (load-bearing). Runs a stream through the
    collector with sensitive tool_call.args, tool_result.result,
    artifact.uri, artifact.metadata, task.inputs, and
    task.metadata. Asserts the serialized events contain none of:
    rm -rf, sk-leaked, cat /etc/secret.txt, secret-bucket,
    cust-99, redact@example.com. This is the test that fails if we
    ever leak user input.
  • Opt-in surfaces specific fields. Same setup, with
    includeInputs/includeControlPayloads/includeEvidenceIds/includeMetadata
    on. Asserts the previously-redacted fields are now present.
  • Summary rollup. Asserts firstSessionId, finalStatus,
    finalReason, finalText, and reconciles
    eventCount === sum(eventCountsByType).

All 19 tests pass (16 existing + 3 new):

Test Files 1 passed (1)
Tests 19 passed (19)

pnpm typecheck and pnpm build also clean.

The new example runs cleanly:

pnpm dlx tsx examples/sanitized-telemetry-streaming/sanitized-telemetry-streaming.ts

Default view: "inputs":"[redacted]", "metadata":"[redacted]",
tool_call events have no args field, tool_result events have no
result field, artifact events omit uri/metadata. Verbose view:
all fields visible. (Note: existing examples all use the same
pnpm tsx invocation — tsx isn't a local bin, so pnpm dlx tsx
matches the pattern of every other example in the repo.)

Versioning note

This PR ships 0.5.4 → 0.5.6, intentionally skipping 0.5.5. PR #6
currently holds the 0.5.5 bump for the agent-eval / agent-knowledge
dep tree unification. If PR #6 lands first, the version delta becomes
0.5.5 → 0.5.6 (the file already reads 0.5.6 so no rebase action
needed). If this PR lands first, PR #6's 0.5.4 → 0.5.5 bump becomes
a no-op vs. main and that PR's author should rebase / collapse the
version bump. Coordinated with the PR #6 author via this note.

…5.6)
Add createRuntimeStreamEventCollector — a sibling of
createRuntimeEventCollector typed for RuntimeStreamEvent. Honors the
same RuntimeTelemetryOptions redaction flags (includeInputs,
includeUserAnswers, includeControlPayloads, includeEvidenceIds,
includeMetadata, includeRequirementDescriptions, includeEvalDetails)
and returns the same {events, onEvent} interface plus a summary()
function that rolls up event counts, session id, final status, and
concatenated text_delta.text.
Sibling factory rather than overload because stream and non-stream
events have different field shapes (timestamps, sessions, text/tool
deltas) and overlapping type literals (task_start, readiness_end, …) —
a unified dispatcher would silently misroute events.
Adds the streaming-collector example mirror at
examples/sanitized-telemetry-streaming/. Documents in README that
task.intent flows through sanitized telemetry by default and must
never carry user input; route user-visible intent through inputs
(redacted by default) instead.
Bumps 0.5.4 → 0.5.6 (intentionally skipping 0.5.5; PR #6 currently
holds 0.5.5 and is expected to land in series).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@drewstone
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat: streaming-event telemetry collector + task.intent directive (0.5.6) - #7

Merged
drewstone merged 2 commits into
mainfrom
feat/stream-event-collector
May 10, 2026
Merged

feat: streaming-event telemetry collector + task.intent directive (0.5.6)#7
drewstone merged 2 commits into
mainfrom
feat/stream-event-collector

Conversation

@drewstone

Copy link
Copy Markdown
Contributor

Why

createRuntimeEventCollector only accepts AgentRuntimeEvent (the
sink-style events emitted by runAgentTask). It does not handle
RuntimeStreamEvent (the events yielded by runAgentTaskStream).
Every product agent that needs sanitized telemetry on top of streaming
has the same workaround in front of it: call
sanitizeRuntimeStreamEvent(event, options) event-by-event inside the
for await loop, then re-implement summary/aggregation by hand.

The gtm-agent reference migration (@tangle-network/agent-runtime is
already a dep at ^0.5.3, but no streaming integration has been wired
yet — it would hit this exact gap when it does) and any future product
agent moving from runAgentTask to runAgentTaskStream would all
re-derive the same boilerplate. Closing the gap in core now keeps the
opt-in redaction story consistent across both entry points.

API change

Adds createRuntimeStreamEventCollector(options?: RuntimeTelemetryOptions)
as a sibling factory to createRuntimeEventCollector. It returns:

{
events: Array<Record<string,unknown>>onEvent(event: RuntimeStreamEvent): voidsummary(): RuntimeStreamEventSummary}

Honors the same RuntimeTelemetryOptions redaction flags
(includeInputs, includeUserAnswers, includeControlPayloads,
includeEvidenceIds, includeMetadata,
includeRequirementDescriptions, includeEvalDetails). The
summary() rollup gives eventCount, eventCountsByType,
firstSessionId, finalStatus, finalReason, and concatenated
finalText from text_delta events.

Sibling factory vs unified union

I considered three shapes:

  1. Unified factory with a discriminated union over both event
    types.
    Rejected: the stream and non-stream events share type
    literals (task_start, readiness_end, task_end) but with
    different field shapes (timestamp and session on the stream
    side, knowledge decision on the stream side, etc.). A unified
    dispatcher would have to discriminate on the presence of optional
    fields, which is brittle and silently misroutes events. Consumers
    would also lose precise types at the callsite.
  2. Refactor the two event types into one shape and migrate both
    callers.
    Rejected: a much larger blast radius for a packaging-
    level improvement. Backward-incompatible for every existing
    runAgentTask consumer.
  3. Sibling factory. Picked. Clear typing, zero impact on existing
    createRuntimeEventCollector consumers, identical opt-in semantics.

README directive on task.intent

task.intent flows through sanitized telemetry by default (it's the
stable operation label, not redactable by includeInputs). The README
"Sanitized telemetry" section now states explicitly:

Never set task.intent to user input — use a fixed string
describing the operation kind (e.g. "Run a chat turn", "Score a tax return"). If you need to log user-visible intent, route it
through inputs (which are redacted by default) instead.

Same directive is repeated in the new example's README.

New example

examples/sanitized-telemetry-streaming/ mirrors
examples/sanitized-telemetry/ for streaming. Uses
createIterableBackend to yield a synthetic script (text_delta,
tool_call with sensitive args, tool_result with a secret token,
artifact with an internal s3 uri), runs cleanly with no creds, prints
both the default-redacted and verbose opt-in views plus the
summary() rollup. Wired into examples/README.md.

Test plan

Added three vitest cases in tests/runtime.test.ts:

  • Redaction contract (load-bearing). Runs a stream through the
    collector with sensitive tool_call.args, tool_result.result,
    artifact.uri, artifact.metadata, task.inputs, and
    task.metadata. Asserts the serialized events contain none of:
    rm -rf, sk-leaked, cat /etc/secret.txt, secret-bucket,
    cust-99, redact@example.com. This is the test that fails if we
    ever leak user input.
  • Opt-in surfaces specific fields. Same setup, with
    includeInputs/includeControlPayloads/includeEvidenceIds/includeMetadata
    on. Asserts the previously-redacted fields are now present.
  • Summary rollup. Asserts firstSessionId, finalStatus,
    finalReason, finalText, and reconciles
    eventCount === sum(eventCountsByType).

All 19 tests pass (16 existing + 3 new):

Test Files 1 passed (1)
Tests 19 passed (19)

pnpm typecheck and pnpm build also clean.

The new example runs cleanly:

pnpm dlx tsx examples/sanitized-telemetry-streaming/sanitized-telemetry-streaming.ts

Default view: "inputs":"[redacted]", "metadata":"[redacted]",
tool_call events have no args field, tool_result events have no
result field, artifact events omit uri/metadata. Verbose view:
all fields visible. (Note: existing examples all use the same
pnpm tsx invocation — tsx isn't a local bin, so pnpm dlx tsx
matches the pattern of every other example in the repo.)

Versioning note

This PR ships 0.5.4 → 0.5.6, intentionally skipping 0.5.5. PR #6
currently holds the 0.5.5 bump for the agent-eval / agent-knowledge
dep tree unification. If PR #6 lands first, the version delta becomes
0.5.5 → 0.5.6 (the file already reads 0.5.6 so no rebase action
needed). If this PR lands first, PR #6's 0.5.4 → 0.5.5 bump becomes
a no-op vs. main and that PR's author should rebase / collapse the
version bump. Coordinated with the PR #6 author via this note.

…5.6)
Add createRuntimeStreamEventCollector — a sibling of
createRuntimeEventCollector typed for RuntimeStreamEvent. Honors the
same RuntimeTelemetryOptions redaction flags (includeInputs,
includeUserAnswers, includeControlPayloads, includeEvidenceIds,
includeMetadata, includeRequirementDescriptions, includeEvalDetails)
and returns the same {events, onEvent} interface plus a summary()
function that rolls up event counts, session id, final status, and
concatenated text_delta.text.
Sibling factory rather than overload because stream and non-stream
events have different field shapes (timestamps, sessions, text/tool
deltas) and overlapping type literals (task_start, readiness_end, …) —
a unified dispatcher would silently misroute events.
Adds the streaming-collector example mirror at
examples/sanitized-telemetry-streaming/. Documents in README that
task.intent flows through sanitized telemetry by default and must
never carry user input; route user-visible intent through inputs
(redacted by default) instead.
Bumps 0.5.4 → 0.5.6 (intentionally skipping 0.5.5; PR #6 currently
holds 0.5.5 and is expected to land in series).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@drewstone