Feat/mcp structured tool results - #135

Open
Asaf-prog wants to merge 5 commits into
mainfrom
feat/mcp-structured-tool-results
Open

Feat/mcp structured tool results#135
Asaf-prog wants to merge 5 commits into
mainfrom
feat/mcp-structured-tool-results

Conversation

@Asaf-prog

Copy link
Copy Markdown
Collaborator

Summary

Preserve structured MCP tool results throughout Extra's execution pipeline instead of collapsing every successful tool result into plain text.

This builds on #133, which fixed MCP text extraction, and adds a provider-agnostic normalized tool-result model that keeps:

  • model-facing text;
  • machine-readable structured data;
  • bounded artifact metadata.

The structured result now survives hooks, persistence, idempotent replay, and HITL resume without leaking provider-specific LangChain objects into the core runtime.

What changed

  • Added an immutable normalized tool-result abstraction owned by Extra.
  • Preserve MCP structuredContent instead of dropping it.
  • Preserve local structured dict / list tool results where applicable.
  • Use deterministic canonical JSON for structured-only model-facing fallback.
  • Keep only text in the model conversation; structured/artifact data is not automatically exposed through ToolMessage.
  • Added versioned serialization with backwards compatibility for legacy text-only execution records.
  • Preserve structured results through idempotent replay and HITL resume.
  • Added atomic claim/wait semantics so concurrent duplicate calls do not execute the same tool side effect more than once.
  • Keep existing text-based transform_tool_result hooks compatible while preserving structured data.
  • Added bounded artifact normalization and safer handling of malformed/non-serializable provider values.
  • Avoid logging raw structured tool results or provider-supplied values.

Runtime flow

Provider / LangChain result

Tool-result normalization

NormalizedToolResult
├── text
├── structured
└── bounded artifact metadata

Hooks

Execution ledger

Replay / HITL resume

text only

Model conversation

Backwards compatibility

Tests

Added/updated coverage for:

  • MCP text-only and multi-block results;
  • MCP structuredContent;
  • structured-only results;
  • deterministic JSON fallback;
  • local structured and string tool results;
  • malformed/non-serializable provider values;
  • artifact size and nesting limits;
  • sequential and concurrent idempotent replay;
  • terminal execution-state protection;
  • legacy ledger compatibility;
  • structured result preservation through hooks;
  • HITL suspend/resume;
  • prevention of structured data leaking into logs/callback metadata.

Validation

  • Ruff format/lint passed.
  • Mypy passed.
  • Pytest: 965 passed.
  • git diff --check passed.

Out of scope

  • Full multimodal MCP support.
  • Binary artifact persistence.
  • Widget rendering for structured/artifact results.
  • Universal output-schema validation.
  • Provider-specific model serialization strategies.

Add an immutable, versioned tool-result contract and make execution claims
safe for concurrent replay. Preserve legacy text records while preventing
terminal ledger overwrites.
Normalize MCP and local tool outputs once at the provider boundary, expose
structured values to trusted hooks, and keep model messages and logs text-only.
Cover structured-only, malformed, replay, concurrency, and HITL behavior.
Document normalization ownership, model visibility, artifact bounds, hook
compatibility, and concurrency-safe persisted replay in ADR 0004 and runtime
guides.
@Asaf-progAsaf-prog self-assigned this Sep 1, 2026

@AmitAvital1AmitAvital1 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes for correctness and contract issues found in the structured-result execution path; details are inline.

await self._publish_processing_failure(call, "Tool error processing failed")
raise
result = NormalizedToolResult.text_only(model_text or f"Tool error: {exc}")
await self._execution_manager.finish_execution(call.exec_id, status="failed", result=result)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cancellation while awaiting this write leaves the claim at started: _record_error() is outside the recovery wrapper used by _record_success().
Please shield/finalize it and add a cancellation test; otherwise duplicate callers can wait forever.

try:
if isinstance(result, ToolMessage):
model_text, content_structured = _tool_message_content(result.content)
artifact_structured, artifact = _split_artifact(result.artifact)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This branch relies on the artifact contract introduced by langchain-mcp-adapters 0.2.0, but pyproject.toml still permits 0.1.x.
Could we bump the minimum version? Otherwise valid installations silently lose structuredContent, and older artifact shapes can fail normalization.

object.__setattr__(
self,
"_structured_json",
_canonical_json(structured, field_name="structured result")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

structured is canonicalized and later decoded, copied, and persisted with only a depth check; unlike artifact, there is no item or byte budget.
A large MCP response can multiply memory use across normalization, hooks, and replay—please reuse a bounded-JSON policy here.

Comment threadsrc/agent_engine/approvals/models.py Outdated
from agent_engine.runtime.tool_models import ToolProviderName
from agent_engine.runtime.tool_results import PersistedToolResult

ToolExecutionStatus = Literal["started", "succeeded", "failed"]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this be a StrEnum, consistent with RunStatus and ApprovalStatus below?
The raw values are repeated in the repository, manager, and invoker; enum members would remove these magic state strings.

Make terminal execution writes cancellation-safe, model execution states with
a StrEnum, bound structured JSON payloads, and require the MCP adapter version
that provides the structured-content artifact contract.
Add regressions for cancellation with duplicate waiters and oversized MCP
structured results.
Record the structured JSON budgets, execution status enum, cancellation-safe
ledger expectations, and langchain-mcp-adapters 0.2 minimum.
@Asaf-prog

Copy link
Copy Markdown
CollaboratorAuthor

Thanks for the detailed review — I addressed all four findings:

  • Made terminal execution writes cancellation-safe. Failed results are persisted before secondary usage/hook processing, and cancellation waits for the shielded ledger write to finish. Added a regression proving concurrent duplicate callers do not hang or re-execute the provider.
  • Raised the langchain-mcp-adapters minimum version to >=0.2, matching the structured-content artifact contract used by the normalizer.
  • Added bounded JSON validation for structured results: depth, total values, string/key bytes, integer size, and final encoded size are limited. Oversized MCP structured results now produce controlled tool errors.
  • Replaced repeated execution-state strings with the wire-compatible ToolExecutionStatus StrEnum.
    Documentation and ADR 0004 were updated accordingly. The full repository gate passes: 969 tests, Ruff, formatting, and mypy

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Asaf-prog@AmitAvital1
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Feat/mcp structured tool results - #135

Open
Asaf-prog wants to merge 5 commits into
mainfrom
feat/mcp-structured-tool-results
Open

Feat/mcp structured tool results#135
Asaf-prog wants to merge 5 commits into
mainfrom
feat/mcp-structured-tool-results

Conversation

@Asaf-prog

Copy link
Copy Markdown
Collaborator

Summary

Preserve structured MCP tool results throughout Extra's execution pipeline instead of collapsing every successful tool result into plain text.

This builds on #133, which fixed MCP text extraction, and adds a provider-agnostic normalized tool-result model that keeps:

  • model-facing text;
  • machine-readable structured data;
  • bounded artifact metadata.

The structured result now survives hooks, persistence, idempotent replay, and HITL resume without leaking provider-specific LangChain objects into the core runtime.

What changed

  • Added an immutable normalized tool-result abstraction owned by Extra.
  • Preserve MCP structuredContent instead of dropping it.
  • Preserve local structured dict / list tool results where applicable.
  • Use deterministic canonical JSON for structured-only model-facing fallback.
  • Keep only text in the model conversation; structured/artifact data is not automatically exposed through ToolMessage.
  • Added versioned serialization with backwards compatibility for legacy text-only execution records.
  • Preserve structured results through idempotent replay and HITL resume.
  • Added atomic claim/wait semantics so concurrent duplicate calls do not execute the same tool side effect more than once.
  • Keep existing text-based transform_tool_result hooks compatible while preserving structured data.
  • Added bounded artifact normalization and safer handling of malformed/non-serializable provider values.
  • Avoid logging raw structured tool results or provider-supplied values.

Runtime flow

Provider / LangChain result

Tool-result normalization

NormalizedToolResult
├── text
├── structured
└── bounded artifact metadata

Hooks

Execution ledger

Replay / HITL resume

text only

Model conversation

Backwards compatibility

Tests

Added/updated coverage for:

  • MCP text-only and multi-block results;
  • MCP structuredContent;
  • structured-only results;
  • deterministic JSON fallback;
  • local structured and string tool results;
  • malformed/non-serializable provider values;
  • artifact size and nesting limits;
  • sequential and concurrent idempotent replay;
  • terminal execution-state protection;
  • legacy ledger compatibility;
  • structured result preservation through hooks;
  • HITL suspend/resume;
  • prevention of structured data leaking into logs/callback metadata.

Validation

  • Ruff format/lint passed.
  • Mypy passed.
  • Pytest: 965 passed.
  • git diff --check passed.

Out of scope

  • Full multimodal MCP support.
  • Binary artifact persistence.
  • Widget rendering for structured/artifact results.
  • Universal output-schema validation.
  • Provider-specific model serialization strategies.

Add an immutable, versioned tool-result contract and make execution claims
safe for concurrent replay. Preserve legacy text records while preventing
terminal ledger overwrites.
Normalize MCP and local tool outputs once at the provider boundary, expose
structured values to trusted hooks, and keep model messages and logs text-only.
Cover structured-only, malformed, replay, concurrency, and HITL behavior.
Document normalization ownership, model visibility, artifact bounds, hook
compatibility, and concurrency-safe persisted replay in ADR 0004 and runtime
guides.
@Asaf-progAsaf-prog self-assigned this Sep 1, 2026

@AmitAvital1AmitAvital1 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes for correctness and contract issues found in the structured-result execution path; details are inline.

await self._publish_processing_failure(call, "Tool error processing failed")
raise
result = NormalizedToolResult.text_only(model_text or f"Tool error: {exc}")
await self._execution_manager.finish_execution(call.exec_id, status="failed", result=result)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cancellation while awaiting this write leaves the claim at started: _record_error() is outside the recovery wrapper used by _record_success().
Please shield/finalize it and add a cancellation test; otherwise duplicate callers can wait forever.

try:
if isinstance(result, ToolMessage):
model_text, content_structured = _tool_message_content(result.content)
artifact_structured, artifact = _split_artifact(result.artifact)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This branch relies on the artifact contract introduced by langchain-mcp-adapters 0.2.0, but pyproject.toml still permits 0.1.x.
Could we bump the minimum version? Otherwise valid installations silently lose structuredContent, and older artifact shapes can fail normalization.

object.__setattr__(
self,
"_structured_json",
_canonical_json(structured, field_name="structured result")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

structured is canonicalized and later decoded, copied, and persisted with only a depth check; unlike artifact, there is no item or byte budget.
A large MCP response can multiply memory use across normalization, hooks, and replay—please reuse a bounded-JSON policy here.

Comment threadsrc/agent_engine/approvals/models.py Outdated
from agent_engine.runtime.tool_models import ToolProviderName
from agent_engine.runtime.tool_results import PersistedToolResult

ToolExecutionStatus = Literal["started", "succeeded", "failed"]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this be a StrEnum, consistent with RunStatus and ApprovalStatus below?
The raw values are repeated in the repository, manager, and invoker; enum members would remove these magic state strings.

Make terminal execution writes cancellation-safe, model execution states with
a StrEnum, bound structured JSON payloads, and require the MCP adapter version
that provides the structured-content artifact contract.
Add regressions for cancellation with duplicate waiters and oversized MCP
structured results.
Record the structured JSON budgets, execution status enum, cancellation-safe
ledger expectations, and langchain-mcp-adapters 0.2 minimum.
@Asaf-prog

Copy link
Copy Markdown
CollaboratorAuthor

Thanks for the detailed review — I addressed all four findings:

  • Made terminal execution writes cancellation-safe. Failed results are persisted before secondary usage/hook processing, and cancellation waits for the shielded ledger write to finish. Added a regression proving concurrent duplicate callers do not hang or re-execute the provider.
  • Raised the langchain-mcp-adapters minimum version to >=0.2, matching the structured-content artifact contract used by the normalizer.
  • Added bounded JSON validation for structured results: depth, total values, string/key bytes, integer size, and final encoded size are limited. Oversized MCP structured results now produce controlled tool errors.
  • Replaced repeated execution-state strings with the wire-compatible ToolExecutionStatus StrEnum.
    Documentation and ADR 0004 were updated accordingly. The full repository gate passes: 969 tests, Ruff, formatting, and mypy

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Asaf-prog@AmitAvital1
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Feat/mcp structured tool results - #135

Open
Asaf-prog wants to merge 5 commits into
mainfrom
feat/mcp-structured-tool-results
Open

Feat/mcp structured tool results#135
Asaf-prog wants to merge 5 commits into
mainfrom
feat/mcp-structured-tool-results

Conversation

@Asaf-prog

Copy link
Copy Markdown
Collaborator

Summary

Preserve structured MCP tool results throughout Extra's execution pipeline instead of collapsing every successful tool result into plain text.

This builds on #133, which fixed MCP text extraction, and adds a provider-agnostic normalized tool-result model that keeps:

  • model-facing text;
  • machine-readable structured data;
  • bounded artifact metadata.

The structured result now survives hooks, persistence, idempotent replay, and HITL resume without leaking provider-specific LangChain objects into the core runtime.

What changed

  • Added an immutable normalized tool-result abstraction owned by Extra.
  • Preserve MCP structuredContent instead of dropping it.
  • Preserve local structured dict / list tool results where applicable.
  • Use deterministic canonical JSON for structured-only model-facing fallback.
  • Keep only text in the model conversation; structured/artifact data is not automatically exposed through ToolMessage.
  • Added versioned serialization with backwards compatibility for legacy text-only execution records.
  • Preserve structured results through idempotent replay and HITL resume.
  • Added atomic claim/wait semantics so concurrent duplicate calls do not execute the same tool side effect more than once.
  • Keep existing text-based transform_tool_result hooks compatible while preserving structured data.
  • Added bounded artifact normalization and safer handling of malformed/non-serializable provider values.
  • Avoid logging raw structured tool results or provider-supplied values.

Runtime flow

Provider / LangChain result

Tool-result normalization

NormalizedToolResult
├── text
├── structured
└── bounded artifact metadata

Hooks

Execution ledger

Replay / HITL resume

text only

Model conversation

Backwards compatibility

Tests

Added/updated coverage for:

  • MCP text-only and multi-block results;
  • MCP structuredContent;
  • structured-only results;
  • deterministic JSON fallback;
  • local structured and string tool results;
  • malformed/non-serializable provider values;
  • artifact size and nesting limits;
  • sequential and concurrent idempotent replay;
  • terminal execution-state protection;
  • legacy ledger compatibility;
  • structured result preservation through hooks;
  • HITL suspend/resume;
  • prevention of structured data leaking into logs/callback metadata.

Validation

  • Ruff format/lint passed.
  • Mypy passed.
  • Pytest: 965 passed.
  • git diff --check passed.

Out of scope

  • Full multimodal MCP support.
  • Binary artifact persistence.
  • Widget rendering for structured/artifact results.
  • Universal output-schema validation.
  • Provider-specific model serialization strategies.

Add an immutable, versioned tool-result contract and make execution claims
safe for concurrent replay. Preserve legacy text records while preventing
terminal ledger overwrites.
Normalize MCP and local tool outputs once at the provider boundary, expose
structured values to trusted hooks, and keep model messages and logs text-only.
Cover structured-only, malformed, replay, concurrency, and HITL behavior.
Document normalization ownership, model visibility, artifact bounds, hook
compatibility, and concurrency-safe persisted replay in ADR 0004 and runtime
guides.
@Asaf-progAsaf-prog self-assigned this Sep 1, 2026

@AmitAvital1AmitAvital1 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes for correctness and contract issues found in the structured-result execution path; details are inline.

await self._publish_processing_failure(call, "Tool error processing failed")
raise
result = NormalizedToolResult.text_only(model_text or f"Tool error: {exc}")
await self._execution_manager.finish_execution(call.exec_id, status="failed", result=result)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cancellation while awaiting this write leaves the claim at started: _record_error() is outside the recovery wrapper used by _record_success().
Please shield/finalize it and add a cancellation test; otherwise duplicate callers can wait forever.

try:
if isinstance(result, ToolMessage):
model_text, content_structured = _tool_message_content(result.content)
artifact_structured, artifact = _split_artifact(result.artifact)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This branch relies on the artifact contract introduced by langchain-mcp-adapters 0.2.0, but pyproject.toml still permits 0.1.x.
Could we bump the minimum version? Otherwise valid installations silently lose structuredContent, and older artifact shapes can fail normalization.

object.__setattr__(
self,
"_structured_json",
_canonical_json(structured, field_name="structured result")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

structured is canonicalized and later decoded, copied, and persisted with only a depth check; unlike artifact, there is no item or byte budget.
A large MCP response can multiply memory use across normalization, hooks, and replay—please reuse a bounded-JSON policy here.

Comment threadsrc/agent_engine/approvals/models.py Outdated
from agent_engine.runtime.tool_models import ToolProviderName
from agent_engine.runtime.tool_results import PersistedToolResult

ToolExecutionStatus = Literal["started", "succeeded", "failed"]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this be a StrEnum, consistent with RunStatus and ApprovalStatus below?
The raw values are repeated in the repository, manager, and invoker; enum members would remove these magic state strings.

Make terminal execution writes cancellation-safe, model execution states with
a StrEnum, bound structured JSON payloads, and require the MCP adapter version
that provides the structured-content artifact contract.
Add regressions for cancellation with duplicate waiters and oversized MCP
structured results.
Record the structured JSON budgets, execution status enum, cancellation-safe
ledger expectations, and langchain-mcp-adapters 0.2 minimum.
@Asaf-prog

Copy link
Copy Markdown
CollaboratorAuthor

Thanks for the detailed review — I addressed all four findings:

  • Made terminal execution writes cancellation-safe. Failed results are persisted before secondary usage/hook processing, and cancellation waits for the shielded ledger write to finish. Added a regression proving concurrent duplicate callers do not hang or re-execute the provider.
  • Raised the langchain-mcp-adapters minimum version to >=0.2, matching the structured-content artifact contract used by the normalizer.
  • Added bounded JSON validation for structured results: depth, total values, string/key bytes, integer size, and final encoded size are limited. Oversized MCP structured results now produce controlled tool errors.
  • Replaced repeated execution-state strings with the wire-compatible ToolExecutionStatus StrEnum.
    Documentation and ADR 0004 were updated accordingly. The full repository gate passes: 969 tests, Ruff, formatting, and mypy

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Asaf-prog@AmitAvital1
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Feat/mcp structured tool results - #135

Open
Asaf-prog wants to merge 5 commits into
mainfrom
feat/mcp-structured-tool-results
Open

Feat/mcp structured tool results#135
Asaf-prog wants to merge 5 commits into
mainfrom
feat/mcp-structured-tool-results

Conversation

@Asaf-prog

Copy link
Copy Markdown
Collaborator

Summary

Preserve structured MCP tool results throughout Extra's execution pipeline instead of collapsing every successful tool result into plain text.

This builds on #133, which fixed MCP text extraction, and adds a provider-agnostic normalized tool-result model that keeps:

  • model-facing text;
  • machine-readable structured data;
  • bounded artifact metadata.

The structured result now survives hooks, persistence, idempotent replay, and HITL resume without leaking provider-specific LangChain objects into the core runtime.

What changed

  • Added an immutable normalized tool-result abstraction owned by Extra.
  • Preserve MCP structuredContent instead of dropping it.
  • Preserve local structured dict / list tool results where applicable.
  • Use deterministic canonical JSON for structured-only model-facing fallback.
  • Keep only text in the model conversation; structured/artifact data is not automatically exposed through ToolMessage.
  • Added versioned serialization with backwards compatibility for legacy text-only execution records.
  • Preserve structured results through idempotent replay and HITL resume.
  • Added atomic claim/wait semantics so concurrent duplicate calls do not execute the same tool side effect more than once.
  • Keep existing text-based transform_tool_result hooks compatible while preserving structured data.
  • Added bounded artifact normalization and safer handling of malformed/non-serializable provider values.
  • Avoid logging raw structured tool results or provider-supplied values.

Runtime flow

Provider / LangChain result

Tool-result normalization

NormalizedToolResult
├── text
├── structured
└── bounded artifact metadata

Hooks

Execution ledger

Replay / HITL resume

text only

Model conversation

Backwards compatibility

Tests

Added/updated coverage for:

  • MCP text-only and multi-block results;
  • MCP structuredContent;
  • structured-only results;
  • deterministic JSON fallback;
  • local structured and string tool results;
  • malformed/non-serializable provider values;
  • artifact size and nesting limits;
  • sequential and concurrent idempotent replay;
  • terminal execution-state protection;
  • legacy ledger compatibility;
  • structured result preservation through hooks;
  • HITL suspend/resume;
  • prevention of structured data leaking into logs/callback metadata.

Validation

  • Ruff format/lint passed.
  • Mypy passed.
  • Pytest: 965 passed.
  • git diff --check passed.

Out of scope

  • Full multimodal MCP support.
  • Binary artifact persistence.
  • Widget rendering for structured/artifact results.
  • Universal output-schema validation.
  • Provider-specific model serialization strategies.

Add an immutable, versioned tool-result contract and make execution claims
safe for concurrent replay. Preserve legacy text records while preventing
terminal ledger overwrites.
Normalize MCP and local tool outputs once at the provider boundary, expose
structured values to trusted hooks, and keep model messages and logs text-only.
Cover structured-only, malformed, replay, concurrency, and HITL behavior.
Document normalization ownership, model visibility, artifact bounds, hook
compatibility, and concurrency-safe persisted replay in ADR 0004 and runtime
guides.
@Asaf-progAsaf-prog self-assigned this Sep 1, 2026

@AmitAvital1AmitAvital1 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes for correctness and contract issues found in the structured-result execution path; details are inline.

await self._publish_processing_failure(call, "Tool error processing failed")
raise
result = NormalizedToolResult.text_only(model_text or f"Tool error: {exc}")
await self._execution_manager.finish_execution(call.exec_id, status="failed", result=result)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cancellation while awaiting this write leaves the claim at started: _record_error() is outside the recovery wrapper used by _record_success().
Please shield/finalize it and add a cancellation test; otherwise duplicate callers can wait forever.

try:
if isinstance(result, ToolMessage):
model_text, content_structured = _tool_message_content(result.content)
artifact_structured, artifact = _split_artifact(result.artifact)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This branch relies on the artifact contract introduced by langchain-mcp-adapters 0.2.0, but pyproject.toml still permits 0.1.x.
Could we bump the minimum version? Otherwise valid installations silently lose structuredContent, and older artifact shapes can fail normalization.

object.__setattr__(
self,
"_structured_json",
_canonical_json(structured, field_name="structured result")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

structured is canonicalized and later decoded, copied, and persisted with only a depth check; unlike artifact, there is no item or byte budget.
A large MCP response can multiply memory use across normalization, hooks, and replay—please reuse a bounded-JSON policy here.

Comment threadsrc/agent_engine/approvals/models.py Outdated
from agent_engine.runtime.tool_models import ToolProviderName
from agent_engine.runtime.tool_results import PersistedToolResult

ToolExecutionStatus = Literal["started", "succeeded", "failed"]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this be a StrEnum, consistent with RunStatus and ApprovalStatus below?
The raw values are repeated in the repository, manager, and invoker; enum members would remove these magic state strings.

Make terminal execution writes cancellation-safe, model execution states with
a StrEnum, bound structured JSON payloads, and require the MCP adapter version
that provides the structured-content artifact contract.
Add regressions for cancellation with duplicate waiters and oversized MCP
structured results.
Record the structured JSON budgets, execution status enum, cancellation-safe
ledger expectations, and langchain-mcp-adapters 0.2 minimum.
@Asaf-prog

Copy link
Copy Markdown
CollaboratorAuthor

Thanks for the detailed review — I addressed all four findings:

  • Made terminal execution writes cancellation-safe. Failed results are persisted before secondary usage/hook processing, and cancellation waits for the shielded ledger write to finish. Added a regression proving concurrent duplicate callers do not hang or re-execute the provider.
  • Raised the langchain-mcp-adapters minimum version to >=0.2, matching the structured-content artifact contract used by the normalizer.
  • Added bounded JSON validation for structured results: depth, total values, string/key bytes, integer size, and final encoded size are limited. Oversized MCP structured results now produce controlled tool errors.
  • Replaced repeated execution-state strings with the wire-compatible ToolExecutionStatus StrEnum.
    Documentation and ADR 0004 were updated accordingly. The full repository gate passes: 969 tests, Ruff, formatting, and mypy

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Asaf-prog@AmitAvital1
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Feat/mcp structured tool results - #135

Open
Asaf-prog wants to merge 5 commits into
mainfrom
feat/mcp-structured-tool-results
Open

Feat/mcp structured tool results#135
Asaf-prog wants to merge 5 commits into
mainfrom
feat/mcp-structured-tool-results

Conversation

@Asaf-prog

Copy link
Copy Markdown
Collaborator

Summary

Preserve structured MCP tool results throughout Extra's execution pipeline instead of collapsing every successful tool result into plain text.

This builds on #133, which fixed MCP text extraction, and adds a provider-agnostic normalized tool-result model that keeps:

  • model-facing text;
  • machine-readable structured data;
  • bounded artifact metadata.

The structured result now survives hooks, persistence, idempotent replay, and HITL resume without leaking provider-specific LangChain objects into the core runtime.

What changed

  • Added an immutable normalized tool-result abstraction owned by Extra.
  • Preserve MCP structuredContent instead of dropping it.
  • Preserve local structured dict / list tool results where applicable.
  • Use deterministic canonical JSON for structured-only model-facing fallback.
  • Keep only text in the model conversation; structured/artifact data is not automatically exposed through ToolMessage.
  • Added versioned serialization with backwards compatibility for legacy text-only execution records.
  • Preserve structured results through idempotent replay and HITL resume.
  • Added atomic claim/wait semantics so concurrent duplicate calls do not execute the same tool side effect more than once.
  • Keep existing text-based transform_tool_result hooks compatible while preserving structured data.
  • Added bounded artifact normalization and safer handling of malformed/non-serializable provider values.
  • Avoid logging raw structured tool results or provider-supplied values.

Runtime flow

Provider / LangChain result

Tool-result normalization

NormalizedToolResult
├── text
├── structured
└── bounded artifact metadata

Hooks

Execution ledger

Replay / HITL resume

text only

Model conversation

Backwards compatibility

Tests

Added/updated coverage for:

  • MCP text-only and multi-block results;
  • MCP structuredContent;
  • structured-only results;
  • deterministic JSON fallback;
  • local structured and string tool results;
  • malformed/non-serializable provider values;
  • artifact size and nesting limits;
  • sequential and concurrent idempotent replay;
  • terminal execution-state protection;
  • legacy ledger compatibility;
  • structured result preservation through hooks;
  • HITL suspend/resume;
  • prevention of structured data leaking into logs/callback metadata.

Validation

  • Ruff format/lint passed.
  • Mypy passed.
  • Pytest: 965 passed.
  • git diff --check passed.

Out of scope

  • Full multimodal MCP support.
  • Binary artifact persistence.
  • Widget rendering for structured/artifact results.
  • Universal output-schema validation.
  • Provider-specific model serialization strategies.

Add an immutable, versioned tool-result contract and make execution claims
safe for concurrent replay. Preserve legacy text records while preventing
terminal ledger overwrites.
Normalize MCP and local tool outputs once at the provider boundary, expose
structured values to trusted hooks, and keep model messages and logs text-only.
Cover structured-only, malformed, replay, concurrency, and HITL behavior.
Document normalization ownership, model visibility, artifact bounds, hook
compatibility, and concurrency-safe persisted replay in ADR 0004 and runtime
guides.
@Asaf-progAsaf-prog self-assigned this Sep 1, 2026

@AmitAvital1AmitAvital1 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes for correctness and contract issues found in the structured-result execution path; details are inline.

await self._publish_processing_failure(call, "Tool error processing failed")
raise
result = NormalizedToolResult.text_only(model_text or f"Tool error: {exc}")
await self._execution_manager.finish_execution(call.exec_id, status="failed", result=result)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cancellation while awaiting this write leaves the claim at started: _record_error() is outside the recovery wrapper used by _record_success().
Please shield/finalize it and add a cancellation test; otherwise duplicate callers can wait forever.

try:
if isinstance(result, ToolMessage):
model_text, content_structured = _tool_message_content(result.content)
artifact_structured, artifact = _split_artifact(result.artifact)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This branch relies on the artifact contract introduced by langchain-mcp-adapters 0.2.0, but pyproject.toml still permits 0.1.x.
Could we bump the minimum version? Otherwise valid installations silently lose structuredContent, and older artifact shapes can fail normalization.

object.__setattr__(
self,
"_structured_json",
_canonical_json(structured, field_name="structured result")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

structured is canonicalized and later decoded, copied, and persisted with only a depth check; unlike artifact, there is no item or byte budget.
A large MCP response can multiply memory use across normalization, hooks, and replay—please reuse a bounded-JSON policy here.

Comment threadsrc/agent_engine/approvals/models.py Outdated
from agent_engine.runtime.tool_models import ToolProviderName
from agent_engine.runtime.tool_results import PersistedToolResult

ToolExecutionStatus = Literal["started", "succeeded", "failed"]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this be a StrEnum, consistent with RunStatus and ApprovalStatus below?
The raw values are repeated in the repository, manager, and invoker; enum members would remove these magic state strings.

Make terminal execution writes cancellation-safe, model execution states with
a StrEnum, bound structured JSON payloads, and require the MCP adapter version
that provides the structured-content artifact contract.
Add regressions for cancellation with duplicate waiters and oversized MCP
structured results.
Record the structured JSON budgets, execution status enum, cancellation-safe
ledger expectations, and langchain-mcp-adapters 0.2 minimum.
@Asaf-prog

Copy link
Copy Markdown
CollaboratorAuthor

Thanks for the detailed review — I addressed all four findings:

  • Made terminal execution writes cancellation-safe. Failed results are persisted before secondary usage/hook processing, and cancellation waits for the shielded ledger write to finish. Added a regression proving concurrent duplicate callers do not hang or re-execute the provider.
  • Raised the langchain-mcp-adapters minimum version to >=0.2, matching the structured-content artifact contract used by the normalizer.
  • Added bounded JSON validation for structured results: depth, total values, string/key bytes, integer size, and final encoded size are limited. Oversized MCP structured results now produce controlled tool errors.
  • Replaced repeated execution-state strings with the wire-compatible ToolExecutionStatus StrEnum.
    Documentation and ADR 0004 were updated accordingly. The full repository gate passes: 969 tests, Ruff, formatting, and mypy

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Asaf-prog@AmitAvital1
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Feat/mcp structured tool results - #135

Open
Asaf-prog wants to merge 5 commits into
mainfrom
feat/mcp-structured-tool-results
Open

Feat/mcp structured tool results#135
Asaf-prog wants to merge 5 commits into
mainfrom
feat/mcp-structured-tool-results

Conversation

@Asaf-prog

Copy link
Copy Markdown
Collaborator

Summary

Preserve structured MCP tool results throughout Extra's execution pipeline instead of collapsing every successful tool result into plain text.

This builds on #133, which fixed MCP text extraction, and adds a provider-agnostic normalized tool-result model that keeps:

  • model-facing text;
  • machine-readable structured data;
  • bounded artifact metadata.

The structured result now survives hooks, persistence, idempotent replay, and HITL resume without leaking provider-specific LangChain objects into the core runtime.

What changed

  • Added an immutable normalized tool-result abstraction owned by Extra.
  • Preserve MCP structuredContent instead of dropping it.
  • Preserve local structured dict / list tool results where applicable.
  • Use deterministic canonical JSON for structured-only model-facing fallback.
  • Keep only text in the model conversation; structured/artifact data is not automatically exposed through ToolMessage.
  • Added versioned serialization with backwards compatibility for legacy text-only execution records.
  • Preserve structured results through idempotent replay and HITL resume.
  • Added atomic claim/wait semantics so concurrent duplicate calls do not execute the same tool side effect more than once.
  • Keep existing text-based transform_tool_result hooks compatible while preserving structured data.
  • Added bounded artifact normalization and safer handling of malformed/non-serializable provider values.
  • Avoid logging raw structured tool results or provider-supplied values.

Runtime flow

Provider / LangChain result

Tool-result normalization

NormalizedToolResult
├── text
├── structured
└── bounded artifact metadata

Hooks

Execution ledger

Replay / HITL resume

text only

Model conversation

Backwards compatibility

Tests

Added/updated coverage for:

  • MCP text-only and multi-block results;
  • MCP structuredContent;
  • structured-only results;
  • deterministic JSON fallback;
  • local structured and string tool results;
  • malformed/non-serializable provider values;
  • artifact size and nesting limits;
  • sequential and concurrent idempotent replay;
  • terminal execution-state protection;
  • legacy ledger compatibility;
  • structured result preservation through hooks;
  • HITL suspend/resume;
  • prevention of structured data leaking into logs/callback metadata.

Validation

  • Ruff format/lint passed.
  • Mypy passed.
  • Pytest: 965 passed.
  • git diff --check passed.

Out of scope

  • Full multimodal MCP support.
  • Binary artifact persistence.
  • Widget rendering for structured/artifact results.
  • Universal output-schema validation.
  • Provider-specific model serialization strategies.

Add an immutable, versioned tool-result contract and make execution claims
safe for concurrent replay. Preserve legacy text records while preventing
terminal ledger overwrites.
Normalize MCP and local tool outputs once at the provider boundary, expose
structured values to trusted hooks, and keep model messages and logs text-only.
Cover structured-only, malformed, replay, concurrency, and HITL behavior.
Document normalization ownership, model visibility, artifact bounds, hook
compatibility, and concurrency-safe persisted replay in ADR 0004 and runtime
guides.
@Asaf-progAsaf-prog self-assigned this Sep 1, 2026

@AmitAvital1AmitAvital1 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes for correctness and contract issues found in the structured-result execution path; details are inline.

await self._publish_processing_failure(call, "Tool error processing failed")
raise
result = NormalizedToolResult.text_only(model_text or f"Tool error: {exc}")
await self._execution_manager.finish_execution(call.exec_id, status="failed", result=result)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cancellation while awaiting this write leaves the claim at started: _record_error() is outside the recovery wrapper used by _record_success().
Please shield/finalize it and add a cancellation test; otherwise duplicate callers can wait forever.

try:
if isinstance(result, ToolMessage):
model_text, content_structured = _tool_message_content(result.content)
artifact_structured, artifact = _split_artifact(result.artifact)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This branch relies on the artifact contract introduced by langchain-mcp-adapters 0.2.0, but pyproject.toml still permits 0.1.x.
Could we bump the minimum version? Otherwise valid installations silently lose structuredContent, and older artifact shapes can fail normalization.

object.__setattr__(
self,
"_structured_json",
_canonical_json(structured, field_name="structured result")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

structured is canonicalized and later decoded, copied, and persisted with only a depth check; unlike artifact, there is no item or byte budget.
A large MCP response can multiply memory use across normalization, hooks, and replay—please reuse a bounded-JSON policy here.

Comment threadsrc/agent_engine/approvals/models.py Outdated
from agent_engine.runtime.tool_models import ToolProviderName
from agent_engine.runtime.tool_results import PersistedToolResult

ToolExecutionStatus = Literal["started", "succeeded", "failed"]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this be a StrEnum, consistent with RunStatus and ApprovalStatus below?
The raw values are repeated in the repository, manager, and invoker; enum members would remove these magic state strings.

Make terminal execution writes cancellation-safe, model execution states with
a StrEnum, bound structured JSON payloads, and require the MCP adapter version
that provides the structured-content artifact contract.
Add regressions for cancellation with duplicate waiters and oversized MCP
structured results.
Record the structured JSON budgets, execution status enum, cancellation-safe
ledger expectations, and langchain-mcp-adapters 0.2 minimum.
@Asaf-prog

Copy link
Copy Markdown
CollaboratorAuthor

Thanks for the detailed review — I addressed all four findings:

  • Made terminal execution writes cancellation-safe. Failed results are persisted before secondary usage/hook processing, and cancellation waits for the shielded ledger write to finish. Added a regression proving concurrent duplicate callers do not hang or re-execute the provider.
  • Raised the langchain-mcp-adapters minimum version to >=0.2, matching the structured-content artifact contract used by the normalizer.
  • Added bounded JSON validation for structured results: depth, total values, string/key bytes, integer size, and final encoded size are limited. Oversized MCP structured results now produce controlled tool errors.
  • Replaced repeated execution-state strings with the wire-compatible ToolExecutionStatus StrEnum.
    Documentation and ADR 0004 were updated accordingly. The full repository gate passes: 969 tests, Ruff, formatting, and mypy

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Asaf-prog@AmitAvital1
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Feat/mcp structured tool results - #135

Open
Asaf-prog wants to merge 5 commits into
mainfrom
feat/mcp-structured-tool-results
Open

Feat/mcp structured tool results#135
Asaf-prog wants to merge 5 commits into
mainfrom
feat/mcp-structured-tool-results

Conversation

@Asaf-prog

Copy link
Copy Markdown
Collaborator

Summary

Preserve structured MCP tool results throughout Extra's execution pipeline instead of collapsing every successful tool result into plain text.

This builds on #133, which fixed MCP text extraction, and adds a provider-agnostic normalized tool-result model that keeps:

  • model-facing text;
  • machine-readable structured data;
  • bounded artifact metadata.

The structured result now survives hooks, persistence, idempotent replay, and HITL resume without leaking provider-specific LangChain objects into the core runtime.

What changed

  • Added an immutable normalized tool-result abstraction owned by Extra.
  • Preserve MCP structuredContent instead of dropping it.
  • Preserve local structured dict / list tool results where applicable.
  • Use deterministic canonical JSON for structured-only model-facing fallback.
  • Keep only text in the model conversation; structured/artifact data is not automatically exposed through ToolMessage.
  • Added versioned serialization with backwards compatibility for legacy text-only execution records.
  • Preserve structured results through idempotent replay and HITL resume.
  • Added atomic claim/wait semantics so concurrent duplicate calls do not execute the same tool side effect more than once.
  • Keep existing text-based transform_tool_result hooks compatible while preserving structured data.
  • Added bounded artifact normalization and safer handling of malformed/non-serializable provider values.
  • Avoid logging raw structured tool results or provider-supplied values.

Runtime flow

Provider / LangChain result

Tool-result normalization

NormalizedToolResult
├── text
├── structured
└── bounded artifact metadata

Hooks

Execution ledger

Replay / HITL resume

text only

Model conversation

Backwards compatibility

Tests

Added/updated coverage for:

  • MCP text-only and multi-block results;
  • MCP structuredContent;
  • structured-only results;
  • deterministic JSON fallback;
  • local structured and string tool results;
  • malformed/non-serializable provider values;
  • artifact size and nesting limits;
  • sequential and concurrent idempotent replay;
  • terminal execution-state protection;
  • legacy ledger compatibility;
  • structured result preservation through hooks;
  • HITL suspend/resume;
  • prevention of structured data leaking into logs/callback metadata.

Validation

  • Ruff format/lint passed.
  • Mypy passed.
  • Pytest: 965 passed.
  • git diff --check passed.

Out of scope

  • Full multimodal MCP support.
  • Binary artifact persistence.
  • Widget rendering for structured/artifact results.
  • Universal output-schema validation.
  • Provider-specific model serialization strategies.

Add an immutable, versioned tool-result contract and make execution claims
safe for concurrent replay. Preserve legacy text records while preventing
terminal ledger overwrites.
Normalize MCP and local tool outputs once at the provider boundary, expose
structured values to trusted hooks, and keep model messages and logs text-only.
Cover structured-only, malformed, replay, concurrency, and HITL behavior.
Document normalization ownership, model visibility, artifact bounds, hook
compatibility, and concurrency-safe persisted replay in ADR 0004 and runtime
guides.
@Asaf-progAsaf-prog self-assigned this Sep 1, 2026

@AmitAvital1AmitAvital1 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes for correctness and contract issues found in the structured-result execution path; details are inline.

await self._publish_processing_failure(call, "Tool error processing failed")
raise
result = NormalizedToolResult.text_only(model_text or f"Tool error: {exc}")
await self._execution_manager.finish_execution(call.exec_id, status="failed", result=result)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cancellation while awaiting this write leaves the claim at started: _record_error() is outside the recovery wrapper used by _record_success().
Please shield/finalize it and add a cancellation test; otherwise duplicate callers can wait forever.

try:
if isinstance(result, ToolMessage):
model_text, content_structured = _tool_message_content(result.content)
artifact_structured, artifact = _split_artifact(result.artifact)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This branch relies on the artifact contract introduced by langchain-mcp-adapters 0.2.0, but pyproject.toml still permits 0.1.x.
Could we bump the minimum version? Otherwise valid installations silently lose structuredContent, and older artifact shapes can fail normalization.

object.__setattr__(
self,
"_structured_json",
_canonical_json(structured, field_name="structured result")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

structured is canonicalized and later decoded, copied, and persisted with only a depth check; unlike artifact, there is no item or byte budget.
A large MCP response can multiply memory use across normalization, hooks, and replay—please reuse a bounded-JSON policy here.

Comment threadsrc/agent_engine/approvals/models.py Outdated
from agent_engine.runtime.tool_models import ToolProviderName
from agent_engine.runtime.tool_results import PersistedToolResult

ToolExecutionStatus = Literal["started", "succeeded", "failed"]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this be a StrEnum, consistent with RunStatus and ApprovalStatus below?
The raw values are repeated in the repository, manager, and invoker; enum members would remove these magic state strings.

Make terminal execution writes cancellation-safe, model execution states with
a StrEnum, bound structured JSON payloads, and require the MCP adapter version
that provides the structured-content artifact contract.
Add regressions for cancellation with duplicate waiters and oversized MCP
structured results.
Record the structured JSON budgets, execution status enum, cancellation-safe
ledger expectations, and langchain-mcp-adapters 0.2 minimum.
@Asaf-prog

Copy link
Copy Markdown
CollaboratorAuthor

Thanks for the detailed review — I addressed all four findings:

  • Made terminal execution writes cancellation-safe. Failed results are persisted before secondary usage/hook processing, and cancellation waits for the shielded ledger write to finish. Added a regression proving concurrent duplicate callers do not hang or re-execute the provider.
  • Raised the langchain-mcp-adapters minimum version to >=0.2, matching the structured-content artifact contract used by the normalizer.
  • Added bounded JSON validation for structured results: depth, total values, string/key bytes, integer size, and final encoded size are limited. Oversized MCP structured results now produce controlled tool errors.
  • Replaced repeated execution-state strings with the wire-compatible ToolExecutionStatus StrEnum.
    Documentation and ADR 0004 were updated accordingly. The full repository gate passes: 969 tests, Ruff, formatting, and mypy

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Asaf-prog@AmitAvital1
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Feat/mcp structured tool results - #135

Open
Asaf-prog wants to merge 5 commits into
mainfrom
feat/mcp-structured-tool-results
Open

Feat/mcp structured tool results#135
Asaf-prog wants to merge 5 commits into
mainfrom
feat/mcp-structured-tool-results

Conversation

@Asaf-prog

Copy link
Copy Markdown
Collaborator

Summary

Preserve structured MCP tool results throughout Extra's execution pipeline instead of collapsing every successful tool result into plain text.

This builds on #133, which fixed MCP text extraction, and adds a provider-agnostic normalized tool-result model that keeps:

  • model-facing text;
  • machine-readable structured data;
  • bounded artifact metadata.

The structured result now survives hooks, persistence, idempotent replay, and HITL resume without leaking provider-specific LangChain objects into the core runtime.

What changed

  • Added an immutable normalized tool-result abstraction owned by Extra.
  • Preserve MCP structuredContent instead of dropping it.
  • Preserve local structured dict / list tool results where applicable.
  • Use deterministic canonical JSON for structured-only model-facing fallback.
  • Keep only text in the model conversation; structured/artifact data is not automatically exposed through ToolMessage.
  • Added versioned serialization with backwards compatibility for legacy text-only execution records.
  • Preserve structured results through idempotent replay and HITL resume.
  • Added atomic claim/wait semantics so concurrent duplicate calls do not execute the same tool side effect more than once.
  • Keep existing text-based transform_tool_result hooks compatible while preserving structured data.
  • Added bounded artifact normalization and safer handling of malformed/non-serializable provider values.
  • Avoid logging raw structured tool results or provider-supplied values.

Runtime flow

Provider / LangChain result

Tool-result normalization

NormalizedToolResult
├── text
├── structured
└── bounded artifact metadata

Hooks

Execution ledger

Replay / HITL resume

text only

Model conversation

Backwards compatibility

Tests

Added/updated coverage for:

  • MCP text-only and multi-block results;
  • MCP structuredContent;
  • structured-only results;
  • deterministic JSON fallback;
  • local structured and string tool results;
  • malformed/non-serializable provider values;
  • artifact size and nesting limits;
  • sequential and concurrent idempotent replay;
  • terminal execution-state protection;
  • legacy ledger compatibility;
  • structured result preservation through hooks;
  • HITL suspend/resume;
  • prevention of structured data leaking into logs/callback metadata.

Validation

  • Ruff format/lint passed.
  • Mypy passed.
  • Pytest: 965 passed.
  • git diff --check passed.

Out of scope

  • Full multimodal MCP support.
  • Binary artifact persistence.
  • Widget rendering for structured/artifact results.
  • Universal output-schema validation.
  • Provider-specific model serialization strategies.

Add an immutable, versioned tool-result contract and make execution claims
safe for concurrent replay. Preserve legacy text records while preventing
terminal ledger overwrites.
Normalize MCP and local tool outputs once at the provider boundary, expose
structured values to trusted hooks, and keep model messages and logs text-only.
Cover structured-only, malformed, replay, concurrency, and HITL behavior.
Document normalization ownership, model visibility, artifact bounds, hook
compatibility, and concurrency-safe persisted replay in ADR 0004 and runtime
guides.
@Asaf-progAsaf-prog self-assigned this Sep 1, 2026

@AmitAvital1AmitAvital1 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes for correctness and contract issues found in the structured-result execution path; details are inline.

await self._publish_processing_failure(call, "Tool error processing failed")
raise
result = NormalizedToolResult.text_only(model_text or f"Tool error: {exc}")
await self._execution_manager.finish_execution(call.exec_id, status="failed", result=result)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cancellation while awaiting this write leaves the claim at started: _record_error() is outside the recovery wrapper used by _record_success().
Please shield/finalize it and add a cancellation test; otherwise duplicate callers can wait forever.

try:
if isinstance(result, ToolMessage):
model_text, content_structured = _tool_message_content(result.content)
artifact_structured, artifact = _split_artifact(result.artifact)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This branch relies on the artifact contract introduced by langchain-mcp-adapters 0.2.0, but pyproject.toml still permits 0.1.x.
Could we bump the minimum version? Otherwise valid installations silently lose structuredContent, and older artifact shapes can fail normalization.

object.__setattr__(
self,
"_structured_json",
_canonical_json(structured, field_name="structured result")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

structured is canonicalized and later decoded, copied, and persisted with only a depth check; unlike artifact, there is no item or byte budget.
A large MCP response can multiply memory use across normalization, hooks, and replay—please reuse a bounded-JSON policy here.

Comment threadsrc/agent_engine/approvals/models.py Outdated
from agent_engine.runtime.tool_models import ToolProviderName
from agent_engine.runtime.tool_results import PersistedToolResult

ToolExecutionStatus = Literal["started", "succeeded", "failed"]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this be a StrEnum, consistent with RunStatus and ApprovalStatus below?
The raw values are repeated in the repository, manager, and invoker; enum members would remove these magic state strings.

Make terminal execution writes cancellation-safe, model execution states with
a StrEnum, bound structured JSON payloads, and require the MCP adapter version
that provides the structured-content artifact contract.
Add regressions for cancellation with duplicate waiters and oversized MCP
structured results.
Record the structured JSON budgets, execution status enum, cancellation-safe
ledger expectations, and langchain-mcp-adapters 0.2 minimum.
@Asaf-prog

Copy link
Copy Markdown
CollaboratorAuthor

Thanks for the detailed review — I addressed all four findings:

  • Made terminal execution writes cancellation-safe. Failed results are persisted before secondary usage/hook processing, and cancellation waits for the shielded ledger write to finish. Added a regression proving concurrent duplicate callers do not hang or re-execute the provider.
  • Raised the langchain-mcp-adapters minimum version to >=0.2, matching the structured-content artifact contract used by the normalizer.
  • Added bounded JSON validation for structured results: depth, total values, string/key bytes, integer size, and final encoded size are limited. Oversized MCP structured results now produce controlled tool errors.
  • Replaced repeated execution-state strings with the wire-compatible ToolExecutionStatus StrEnum.
    Documentation and ADR 0004 were updated accordingly. The full repository gate passes: 969 tests, Ruff, formatting, and mypy

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@Asaf-prog@AmitAvital1