Runtime robustness: shell execution, safety policy, permission asks, MCP startup, and tool-call handling - #951

Merged
alcholiclg merged 55 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness
Aug 27, 2026
Merged

Runtime robustness: shell execution, safety policy, permission asks, MCP startup, and tool-call handling#951
alcholiclg merged 55 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness

Conversation

@alcholiclg

Copy link
Copy Markdown
Collaborator

Change Summary

  • Shell safety policy: quote-aware redirect parsing (fixes the false Create path outside allowed directories: /dev/null) block), heredoc bodies treated as code rather than paths, subshell/brace-group unwrapping, OS temp dir writable by default; real violations (writes outside allowed dirs, glob creates, variable-expansion targets) still refused
  • New ask categories: interpreter_exec for inline interpreter code (python3 -c, heredocs) — rememberable per project via "always allow"; sensitive_read for credential paths (~/.ssh/*, *.pem, …) — denied in auto mode, never rememberable, kept separate from the write-protection list so e.g. .git/config stays readable
  • Ask handling: interactive mode waits indefinitely; full-access times out after 25 min with feedback that marks it a timeout, not a refusal; public pending-ask APIs (is_awaiting / awaiting_request_ids / resolve_matching / cancel_pending) so front-ends stop reading private state
  • Shell executor: commands run verbatim in a non-login shell — python3 -V and python3 -V ; true no longer resolve different interpreters; agent-friendly environment (PAGER=cat, GIT_TERMINAL_PROMPT=0, NO_COLOR=1, …) injected unconditionally; tool description states no state persists between calls
  • MCP: parallel connects under per-server owner tasks (anyio-safe teardown), stdio startup timeout, per-server failure isolation; Pydantic union errors deduplicated with actionable guidance prepended
  • Tool calls: schema-driven coercion of string-typed numerics ("19.5" → 19.5), unique-prefix tool-name resolution with candidate suggestions on miss, read_file re-reads return content with an "unchanged" note instead of a stub
  • Prompting: the system prompt documents framework-managed directories (sessions/, .ms_agent/) so transcript matches are not mistaken for user content
  • Tests: 6 new test files + 2 extended; suite baseline unchanged (25 pre-existing env/network failures); verified end-to-end through the WebUI (approval flows, policy blocks, MCP coercion, PATH/env)

Related issue number

Checklist

  • The pull request title is a good summary of the changes - it will be used in the changelog
  • Unit tests for the changes exist
  • Run pre-commit install and pre-commit run --all-files before git commit, and passed lint check.
  • Documentation reflects the changes where applicable

The orchestrator now owns the write discipline around a backend:
- schedule_add() runs the extraction-LLM + embedding cost (seconds) in a
background task; flush_pending() is the teardown barrier so the last
write is never dropped, and an inline fallback keeps writes when no
loop is running.
- retrieval/ingestion/flush serialize on one per-store asyncio lock
(embedded qdrant underneath is lock-free single-client code).
- a content-hash delta ledger (<base_dir>/ingest_state.json) makes each
ingest send only messages the store has not seen; hashes are recorded
only after a confirmed write, so a failed ingest retries naturally.
- ingest_status reports the last outcome (state/count/error/pending) so
a UI can show memory working instead of silence.
Mem0Backend: per-turn retrieval cache (rounds 2..N of a tool-calling
turn reuse round 1's search instead of paying an embedding round-trip
each), on_messages returns the event count and propagates failures --
the orchestrator is the swallow-and-report layer now and needs the
exception to keep failed messages un-marked for retry.
…-side close
- add_memory(add_after_step) now fires only when a round closes the turn
(assistant reply with no tool calls) and dispatches through the
backend's schedule_add when available: tool rounds are intermediate
state, and ingesting every round cost O(rounds x history) extraction
calls where the closing ingest covers the whole turn.
- an interrupted round advances the ingest ledger WITHOUT ingesting
(mark_ingested): a half-finished answer is not durable conversational
truth and must not be swept into the next turn's delta.
- cleanup_tools drains scheduled ingestion (flush only -- memory
instances are shared across agents of one store, so closing here would
yank the store from a sibling agent); the new
SharedMemoryManager.close_matching(base_dir) is the owner-of-last-
resort that actually closes instances and releases the embedded
store's exclusive file lock.
The number of recalled memories injected per turn was hardcoded twice
(search default 20, then a [:10] formatting slice). MemoryConfig gains
recall_top_k (default 10, read from the unified_memory node) and the
mem0 adapter threads it through search and formatting — consumers can
now size recall to their context budget.
…E) instead of scattered config fields, gated by personalization.enabled.
When those files change mid-conversation the next user turn carries a durable <system-reminder> naming them, so the model can tell a changed file from its own faulty memory.
…end's MEMORY.md snapshot in step with edits made outside the agent.
Also translates the memory tool descriptions and prompt headings to English.
# Conflicts:
#	.gitignore
#	ms_agent/memory/unified/backends/mem0_adapter.py
#	ms_agent/memory/unified/orchestrator.py
#	setup.py
…ts last entry deleted, instead of leaving the previous round's block in place.
- an interrupt marks only its own round, and never messages a scheduled
ingest still owns (both lost the write silently)
- one shared instance per store, not per model, reconfigured in place
- close() is terminal: a straggler can no longer reopen a released store
- the store lock is per (loop, path), and search() takes it too
- search() honours its limit; injected memories carry their date
- memories are written in the language the user used
# Conflicts:
#	ms_agent/memory/unified/backends/mem0_adapter.py
…ters
Thinking support is per-model with no naming rule, and an unsupported model may
reject the whole request (DashScope returns 400) instead of ignoring the flag.
So we ask, and on a refusal retry once with it off, remembering the model.
…orwards, and read OpenRouter's reasoning field
…o a vision-disabled model neither claims nor disowns them
…x/runtime-robustness
# Conflicts:
#	ms_agent/agent/llm_agent.py
#	ms_agent/memory/unified/backends/mem0_adapter.py
#	ms_agent/memory/unified/orchestrator.py
…hat arrive mid-stream
- images go out only when the model's own switch says so; a provider's declared
vision capability no longer implies it
- a 400 delivered on the first streamed chunk is repaired like an eager one
- thinking refusals are repaired on the Anthropic and Responses paths too
- a tool call the model is still writing is reported instead of nothing at all
- an unreadable managed MCP config is logged instead of silently yielding none
… tell the agent why a search failed instead of reporting no results
…ps cutting structured results into invalid JSON
… and gate interpreter execution and credential reads behind mode-aware asks
… unique tool-name prefixes, and return content instead of a stub on unchanged re-reads
…io startup timeouts and per-server failure isolation
…ranscript matches are not mistaken for user content
@alcholiclg
alcholiclg merged commit a8ad587 into modelscope:mainAug 27, 2026
1 check failed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@alcholiclg@wangxingjun778
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Runtime robustness: shell execution, safety policy, permission asks, MCP startup, and tool-call handling - #951

Merged
alcholiclg merged 55 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness
Aug 27, 2026
Merged

Runtime robustness: shell execution, safety policy, permission asks, MCP startup, and tool-call handling#951
alcholiclg merged 55 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness

Conversation

@alcholiclg

Copy link
Copy Markdown
Collaborator

Change Summary

  • Shell safety policy: quote-aware redirect parsing (fixes the false Create path outside allowed directories: /dev/null) block), heredoc bodies treated as code rather than paths, subshell/brace-group unwrapping, OS temp dir writable by default; real violations (writes outside allowed dirs, glob creates, variable-expansion targets) still refused
  • New ask categories: interpreter_exec for inline interpreter code (python3 -c, heredocs) — rememberable per project via "always allow"; sensitive_read for credential paths (~/.ssh/*, *.pem, …) — denied in auto mode, never rememberable, kept separate from the write-protection list so e.g. .git/config stays readable
  • Ask handling: interactive mode waits indefinitely; full-access times out after 25 min with feedback that marks it a timeout, not a refusal; public pending-ask APIs (is_awaiting / awaiting_request_ids / resolve_matching / cancel_pending) so front-ends stop reading private state
  • Shell executor: commands run verbatim in a non-login shell — python3 -V and python3 -V ; true no longer resolve different interpreters; agent-friendly environment (PAGER=cat, GIT_TERMINAL_PROMPT=0, NO_COLOR=1, …) injected unconditionally; tool description states no state persists between calls
  • MCP: parallel connects under per-server owner tasks (anyio-safe teardown), stdio startup timeout, per-server failure isolation; Pydantic union errors deduplicated with actionable guidance prepended
  • Tool calls: schema-driven coercion of string-typed numerics ("19.5" → 19.5), unique-prefix tool-name resolution with candidate suggestions on miss, read_file re-reads return content with an "unchanged" note instead of a stub
  • Prompting: the system prompt documents framework-managed directories (sessions/, .ms_agent/) so transcript matches are not mistaken for user content
  • Tests: 6 new test files + 2 extended; suite baseline unchanged (25 pre-existing env/network failures); verified end-to-end through the WebUI (approval flows, policy blocks, MCP coercion, PATH/env)

Related issue number

Checklist

  • The pull request title is a good summary of the changes - it will be used in the changelog
  • Unit tests for the changes exist
  • Run pre-commit install and pre-commit run --all-files before git commit, and passed lint check.
  • Documentation reflects the changes where applicable

The orchestrator now owns the write discipline around a backend:
- schedule_add() runs the extraction-LLM + embedding cost (seconds) in a
background task; flush_pending() is the teardown barrier so the last
write is never dropped, and an inline fallback keeps writes when no
loop is running.
- retrieval/ingestion/flush serialize on one per-store asyncio lock
(embedded qdrant underneath is lock-free single-client code).
- a content-hash delta ledger (<base_dir>/ingest_state.json) makes each
ingest send only messages the store has not seen; hashes are recorded
only after a confirmed write, so a failed ingest retries naturally.
- ingest_status reports the last outcome (state/count/error/pending) so
a UI can show memory working instead of silence.
Mem0Backend: per-turn retrieval cache (rounds 2..N of a tool-calling
turn reuse round 1's search instead of paying an embedding round-trip
each), on_messages returns the event count and propagates failures --
the orchestrator is the swallow-and-report layer now and needs the
exception to keep failed messages un-marked for retry.
…-side close
- add_memory(add_after_step) now fires only when a round closes the turn
(assistant reply with no tool calls) and dispatches through the
backend's schedule_add when available: tool rounds are intermediate
state, and ingesting every round cost O(rounds x history) extraction
calls where the closing ingest covers the whole turn.
- an interrupted round advances the ingest ledger WITHOUT ingesting
(mark_ingested): a half-finished answer is not durable conversational
truth and must not be swept into the next turn's delta.
- cleanup_tools drains scheduled ingestion (flush only -- memory
instances are shared across agents of one store, so closing here would
yank the store from a sibling agent); the new
SharedMemoryManager.close_matching(base_dir) is the owner-of-last-
resort that actually closes instances and releases the embedded
store's exclusive file lock.
The number of recalled memories injected per turn was hardcoded twice
(search default 20, then a [:10] formatting slice). MemoryConfig gains
recall_top_k (default 10, read from the unified_memory node) and the
mem0 adapter threads it through search and formatting — consumers can
now size recall to their context budget.
…E) instead of scattered config fields, gated by personalization.enabled.
When those files change mid-conversation the next user turn carries a durable <system-reminder> naming them, so the model can tell a changed file from its own faulty memory.
…end's MEMORY.md snapshot in step with edits made outside the agent.
Also translates the memory tool descriptions and prompt headings to English.
# Conflicts:
#	.gitignore
#	ms_agent/memory/unified/backends/mem0_adapter.py
#	ms_agent/memory/unified/orchestrator.py
#	setup.py
…ts last entry deleted, instead of leaving the previous round's block in place.
- an interrupt marks only its own round, and never messages a scheduled
ingest still owns (both lost the write silently)
- one shared instance per store, not per model, reconfigured in place
- close() is terminal: a straggler can no longer reopen a released store
- the store lock is per (loop, path), and search() takes it too
- search() honours its limit; injected memories carry their date
- memories are written in the language the user used
# Conflicts:
#	ms_agent/memory/unified/backends/mem0_adapter.py
…ters
Thinking support is per-model with no naming rule, and an unsupported model may
reject the whole request (DashScope returns 400) instead of ignoring the flag.
So we ask, and on a refusal retry once with it off, remembering the model.
…orwards, and read OpenRouter's reasoning field
…o a vision-disabled model neither claims nor disowns them
…x/runtime-robustness
# Conflicts:
#	ms_agent/agent/llm_agent.py
#	ms_agent/memory/unified/backends/mem0_adapter.py
#	ms_agent/memory/unified/orchestrator.py
…hat arrive mid-stream
- images go out only when the model's own switch says so; a provider's declared
vision capability no longer implies it
- a 400 delivered on the first streamed chunk is repaired like an eager one
- thinking refusals are repaired on the Anthropic and Responses paths too
- a tool call the model is still writing is reported instead of nothing at all
- an unreadable managed MCP config is logged instead of silently yielding none
… tell the agent why a search failed instead of reporting no results
…ps cutting structured results into invalid JSON
… and gate interpreter execution and credential reads behind mode-aware asks
… unique tool-name prefixes, and return content instead of a stub on unchanged re-reads
…io startup timeouts and per-server failure isolation
…ranscript matches are not mistaken for user content
@alcholiclg
alcholiclg merged commit a8ad587 into modelscope:mainAug 27, 2026
1 check failed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@alcholiclg@wangxingjun778
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Runtime robustness: shell execution, safety policy, permission asks, MCP startup, and tool-call handling - #951

Merged
alcholiclg merged 55 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness
Aug 27, 2026
Merged

Runtime robustness: shell execution, safety policy, permission asks, MCP startup, and tool-call handling#951
alcholiclg merged 55 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness

Conversation

@alcholiclg

Copy link
Copy Markdown
Collaborator

Change Summary

  • Shell safety policy: quote-aware redirect parsing (fixes the false Create path outside allowed directories: /dev/null) block), heredoc bodies treated as code rather than paths, subshell/brace-group unwrapping, OS temp dir writable by default; real violations (writes outside allowed dirs, glob creates, variable-expansion targets) still refused
  • New ask categories: interpreter_exec for inline interpreter code (python3 -c, heredocs) — rememberable per project via "always allow"; sensitive_read for credential paths (~/.ssh/*, *.pem, …) — denied in auto mode, never rememberable, kept separate from the write-protection list so e.g. .git/config stays readable
  • Ask handling: interactive mode waits indefinitely; full-access times out after 25 min with feedback that marks it a timeout, not a refusal; public pending-ask APIs (is_awaiting / awaiting_request_ids / resolve_matching / cancel_pending) so front-ends stop reading private state
  • Shell executor: commands run verbatim in a non-login shell — python3 -V and python3 -V ; true no longer resolve different interpreters; agent-friendly environment (PAGER=cat, GIT_TERMINAL_PROMPT=0, NO_COLOR=1, …) injected unconditionally; tool description states no state persists between calls
  • MCP: parallel connects under per-server owner tasks (anyio-safe teardown), stdio startup timeout, per-server failure isolation; Pydantic union errors deduplicated with actionable guidance prepended
  • Tool calls: schema-driven coercion of string-typed numerics ("19.5" → 19.5), unique-prefix tool-name resolution with candidate suggestions on miss, read_file re-reads return content with an "unchanged" note instead of a stub
  • Prompting: the system prompt documents framework-managed directories (sessions/, .ms_agent/) so transcript matches are not mistaken for user content
  • Tests: 6 new test files + 2 extended; suite baseline unchanged (25 pre-existing env/network failures); verified end-to-end through the WebUI (approval flows, policy blocks, MCP coercion, PATH/env)

Related issue number

Checklist

  • The pull request title is a good summary of the changes - it will be used in the changelog
  • Unit tests for the changes exist
  • Run pre-commit install and pre-commit run --all-files before git commit, and passed lint check.
  • Documentation reflects the changes where applicable

The orchestrator now owns the write discipline around a backend:
- schedule_add() runs the extraction-LLM + embedding cost (seconds) in a
background task; flush_pending() is the teardown barrier so the last
write is never dropped, and an inline fallback keeps writes when no
loop is running.
- retrieval/ingestion/flush serialize on one per-store asyncio lock
(embedded qdrant underneath is lock-free single-client code).
- a content-hash delta ledger (<base_dir>/ingest_state.json) makes each
ingest send only messages the store has not seen; hashes are recorded
only after a confirmed write, so a failed ingest retries naturally.
- ingest_status reports the last outcome (state/count/error/pending) so
a UI can show memory working instead of silence.
Mem0Backend: per-turn retrieval cache (rounds 2..N of a tool-calling
turn reuse round 1's search instead of paying an embedding round-trip
each), on_messages returns the event count and propagates failures --
the orchestrator is the swallow-and-report layer now and needs the
exception to keep failed messages un-marked for retry.
…-side close
- add_memory(add_after_step) now fires only when a round closes the turn
(assistant reply with no tool calls) and dispatches through the
backend's schedule_add when available: tool rounds are intermediate
state, and ingesting every round cost O(rounds x history) extraction
calls where the closing ingest covers the whole turn.
- an interrupted round advances the ingest ledger WITHOUT ingesting
(mark_ingested): a half-finished answer is not durable conversational
truth and must not be swept into the next turn's delta.
- cleanup_tools drains scheduled ingestion (flush only -- memory
instances are shared across agents of one store, so closing here would
yank the store from a sibling agent); the new
SharedMemoryManager.close_matching(base_dir) is the owner-of-last-
resort that actually closes instances and releases the embedded
store's exclusive file lock.
The number of recalled memories injected per turn was hardcoded twice
(search default 20, then a [:10] formatting slice). MemoryConfig gains
recall_top_k (default 10, read from the unified_memory node) and the
mem0 adapter threads it through search and formatting — consumers can
now size recall to their context budget.
…E) instead of scattered config fields, gated by personalization.enabled.
When those files change mid-conversation the next user turn carries a durable <system-reminder> naming them, so the model can tell a changed file from its own faulty memory.
…end's MEMORY.md snapshot in step with edits made outside the agent.
Also translates the memory tool descriptions and prompt headings to English.
# Conflicts:
#	.gitignore
#	ms_agent/memory/unified/backends/mem0_adapter.py
#	ms_agent/memory/unified/orchestrator.py
#	setup.py
…ts last entry deleted, instead of leaving the previous round's block in place.
- an interrupt marks only its own round, and never messages a scheduled
ingest still owns (both lost the write silently)
- one shared instance per store, not per model, reconfigured in place
- close() is terminal: a straggler can no longer reopen a released store
- the store lock is per (loop, path), and search() takes it too
- search() honours its limit; injected memories carry their date
- memories are written in the language the user used
# Conflicts:
#	ms_agent/memory/unified/backends/mem0_adapter.py
…ters
Thinking support is per-model with no naming rule, and an unsupported model may
reject the whole request (DashScope returns 400) instead of ignoring the flag.
So we ask, and on a refusal retry once with it off, remembering the model.
…orwards, and read OpenRouter's reasoning field
…o a vision-disabled model neither claims nor disowns them
…x/runtime-robustness
# Conflicts:
#	ms_agent/agent/llm_agent.py
#	ms_agent/memory/unified/backends/mem0_adapter.py
#	ms_agent/memory/unified/orchestrator.py
…hat arrive mid-stream
- images go out only when the model's own switch says so; a provider's declared
vision capability no longer implies it
- a 400 delivered on the first streamed chunk is repaired like an eager one
- thinking refusals are repaired on the Anthropic and Responses paths too
- a tool call the model is still writing is reported instead of nothing at all
- an unreadable managed MCP config is logged instead of silently yielding none
… tell the agent why a search failed instead of reporting no results
…ps cutting structured results into invalid JSON
… and gate interpreter execution and credential reads behind mode-aware asks
… unique tool-name prefixes, and return content instead of a stub on unchanged re-reads
…io startup timeouts and per-server failure isolation
…ranscript matches are not mistaken for user content
@alcholiclg
alcholiclg merged commit a8ad587 into modelscope:mainAug 27, 2026
1 check failed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@alcholiclg@wangxingjun778
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Runtime robustness: shell execution, safety policy, permission asks, MCP startup, and tool-call handling - #951

Merged
alcholiclg merged 55 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness
Aug 27, 2026
Merged

Runtime robustness: shell execution, safety policy, permission asks, MCP startup, and tool-call handling#951
alcholiclg merged 55 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness

Conversation

@alcholiclg

Copy link
Copy Markdown
Collaborator

Change Summary

  • Shell safety policy: quote-aware redirect parsing (fixes the false Create path outside allowed directories: /dev/null) block), heredoc bodies treated as code rather than paths, subshell/brace-group unwrapping, OS temp dir writable by default; real violations (writes outside allowed dirs, glob creates, variable-expansion targets) still refused
  • New ask categories: interpreter_exec for inline interpreter code (python3 -c, heredocs) — rememberable per project via "always allow"; sensitive_read for credential paths (~/.ssh/*, *.pem, …) — denied in auto mode, never rememberable, kept separate from the write-protection list so e.g. .git/config stays readable
  • Ask handling: interactive mode waits indefinitely; full-access times out after 25 min with feedback that marks it a timeout, not a refusal; public pending-ask APIs (is_awaiting / awaiting_request_ids / resolve_matching / cancel_pending) so front-ends stop reading private state
  • Shell executor: commands run verbatim in a non-login shell — python3 -V and python3 -V ; true no longer resolve different interpreters; agent-friendly environment (PAGER=cat, GIT_TERMINAL_PROMPT=0, NO_COLOR=1, …) injected unconditionally; tool description states no state persists between calls
  • MCP: parallel connects under per-server owner tasks (anyio-safe teardown), stdio startup timeout, per-server failure isolation; Pydantic union errors deduplicated with actionable guidance prepended
  • Tool calls: schema-driven coercion of string-typed numerics ("19.5" → 19.5), unique-prefix tool-name resolution with candidate suggestions on miss, read_file re-reads return content with an "unchanged" note instead of a stub
  • Prompting: the system prompt documents framework-managed directories (sessions/, .ms_agent/) so transcript matches are not mistaken for user content
  • Tests: 6 new test files + 2 extended; suite baseline unchanged (25 pre-existing env/network failures); verified end-to-end through the WebUI (approval flows, policy blocks, MCP coercion, PATH/env)

Related issue number

Checklist

  • The pull request title is a good summary of the changes - it will be used in the changelog
  • Unit tests for the changes exist
  • Run pre-commit install and pre-commit run --all-files before git commit, and passed lint check.
  • Documentation reflects the changes where applicable

The orchestrator now owns the write discipline around a backend:
- schedule_add() runs the extraction-LLM + embedding cost (seconds) in a
background task; flush_pending() is the teardown barrier so the last
write is never dropped, and an inline fallback keeps writes when no
loop is running.
- retrieval/ingestion/flush serialize on one per-store asyncio lock
(embedded qdrant underneath is lock-free single-client code).
- a content-hash delta ledger (<base_dir>/ingest_state.json) makes each
ingest send only messages the store has not seen; hashes are recorded
only after a confirmed write, so a failed ingest retries naturally.
- ingest_status reports the last outcome (state/count/error/pending) so
a UI can show memory working instead of silence.
Mem0Backend: per-turn retrieval cache (rounds 2..N of a tool-calling
turn reuse round 1's search instead of paying an embedding round-trip
each), on_messages returns the event count and propagates failures --
the orchestrator is the swallow-and-report layer now and needs the
exception to keep failed messages un-marked for retry.
…-side close
- add_memory(add_after_step) now fires only when a round closes the turn
(assistant reply with no tool calls) and dispatches through the
backend's schedule_add when available: tool rounds are intermediate
state, and ingesting every round cost O(rounds x history) extraction
calls where the closing ingest covers the whole turn.
- an interrupted round advances the ingest ledger WITHOUT ingesting
(mark_ingested): a half-finished answer is not durable conversational
truth and must not be swept into the next turn's delta.
- cleanup_tools drains scheduled ingestion (flush only -- memory
instances are shared across agents of one store, so closing here would
yank the store from a sibling agent); the new
SharedMemoryManager.close_matching(base_dir) is the owner-of-last-
resort that actually closes instances and releases the embedded
store's exclusive file lock.
The number of recalled memories injected per turn was hardcoded twice
(search default 20, then a [:10] formatting slice). MemoryConfig gains
recall_top_k (default 10, read from the unified_memory node) and the
mem0 adapter threads it through search and formatting — consumers can
now size recall to their context budget.
…E) instead of scattered config fields, gated by personalization.enabled.
When those files change mid-conversation the next user turn carries a durable <system-reminder> naming them, so the model can tell a changed file from its own faulty memory.
…end's MEMORY.md snapshot in step with edits made outside the agent.
Also translates the memory tool descriptions and prompt headings to English.
# Conflicts:
#	.gitignore
#	ms_agent/memory/unified/backends/mem0_adapter.py
#	ms_agent/memory/unified/orchestrator.py
#	setup.py
…ts last entry deleted, instead of leaving the previous round's block in place.
- an interrupt marks only its own round, and never messages a scheduled
ingest still owns (both lost the write silently)
- one shared instance per store, not per model, reconfigured in place
- close() is terminal: a straggler can no longer reopen a released store
- the store lock is per (loop, path), and search() takes it too
- search() honours its limit; injected memories carry their date
- memories are written in the language the user used
# Conflicts:
#	ms_agent/memory/unified/backends/mem0_adapter.py
…ters
Thinking support is per-model with no naming rule, and an unsupported model may
reject the whole request (DashScope returns 400) instead of ignoring the flag.
So we ask, and on a refusal retry once with it off, remembering the model.
…orwards, and read OpenRouter's reasoning field
…o a vision-disabled model neither claims nor disowns them
…x/runtime-robustness
# Conflicts:
#	ms_agent/agent/llm_agent.py
#	ms_agent/memory/unified/backends/mem0_adapter.py
#	ms_agent/memory/unified/orchestrator.py
…hat arrive mid-stream
- images go out only when the model's own switch says so; a provider's declared
vision capability no longer implies it
- a 400 delivered on the first streamed chunk is repaired like an eager one
- thinking refusals are repaired on the Anthropic and Responses paths too
- a tool call the model is still writing is reported instead of nothing at all
- an unreadable managed MCP config is logged instead of silently yielding none
… tell the agent why a search failed instead of reporting no results
…ps cutting structured results into invalid JSON
… and gate interpreter execution and credential reads behind mode-aware asks
… unique tool-name prefixes, and return content instead of a stub on unchanged re-reads
…io startup timeouts and per-server failure isolation
…ranscript matches are not mistaken for user content
@alcholiclg
alcholiclg merged commit a8ad587 into modelscope:mainAug 27, 2026
1 check failed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@alcholiclg@wangxingjun778
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Runtime robustness: shell execution, safety policy, permission asks, MCP startup, and tool-call handling - #951

Merged
alcholiclg merged 55 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness
Aug 27, 2026
Merged

Runtime robustness: shell execution, safety policy, permission asks, MCP startup, and tool-call handling#951
alcholiclg merged 55 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness

Conversation

@alcholiclg

Copy link
Copy Markdown
Collaborator

Change Summary

  • Shell safety policy: quote-aware redirect parsing (fixes the false Create path outside allowed directories: /dev/null) block), heredoc bodies treated as code rather than paths, subshell/brace-group unwrapping, OS temp dir writable by default; real violations (writes outside allowed dirs, glob creates, variable-expansion targets) still refused
  • New ask categories: interpreter_exec for inline interpreter code (python3 -c, heredocs) — rememberable per project via "always allow"; sensitive_read for credential paths (~/.ssh/*, *.pem, …) — denied in auto mode, never rememberable, kept separate from the write-protection list so e.g. .git/config stays readable
  • Ask handling: interactive mode waits indefinitely; full-access times out after 25 min with feedback that marks it a timeout, not a refusal; public pending-ask APIs (is_awaiting / awaiting_request_ids / resolve_matching / cancel_pending) so front-ends stop reading private state
  • Shell executor: commands run verbatim in a non-login shell — python3 -V and python3 -V ; true no longer resolve different interpreters; agent-friendly environment (PAGER=cat, GIT_TERMINAL_PROMPT=0, NO_COLOR=1, …) injected unconditionally; tool description states no state persists between calls
  • MCP: parallel connects under per-server owner tasks (anyio-safe teardown), stdio startup timeout, per-server failure isolation; Pydantic union errors deduplicated with actionable guidance prepended
  • Tool calls: schema-driven coercion of string-typed numerics ("19.5" → 19.5), unique-prefix tool-name resolution with candidate suggestions on miss, read_file re-reads return content with an "unchanged" note instead of a stub
  • Prompting: the system prompt documents framework-managed directories (sessions/, .ms_agent/) so transcript matches are not mistaken for user content
  • Tests: 6 new test files + 2 extended; suite baseline unchanged (25 pre-existing env/network failures); verified end-to-end through the WebUI (approval flows, policy blocks, MCP coercion, PATH/env)

Related issue number

Checklist

  • The pull request title is a good summary of the changes - it will be used in the changelog
  • Unit tests for the changes exist
  • Run pre-commit install and pre-commit run --all-files before git commit, and passed lint check.
  • Documentation reflects the changes where applicable

The orchestrator now owns the write discipline around a backend:
- schedule_add() runs the extraction-LLM + embedding cost (seconds) in a
background task; flush_pending() is the teardown barrier so the last
write is never dropped, and an inline fallback keeps writes when no
loop is running.
- retrieval/ingestion/flush serialize on one per-store asyncio lock
(embedded qdrant underneath is lock-free single-client code).
- a content-hash delta ledger (<base_dir>/ingest_state.json) makes each
ingest send only messages the store has not seen; hashes are recorded
only after a confirmed write, so a failed ingest retries naturally.
- ingest_status reports the last outcome (state/count/error/pending) so
a UI can show memory working instead of silence.
Mem0Backend: per-turn retrieval cache (rounds 2..N of a tool-calling
turn reuse round 1's search instead of paying an embedding round-trip
each), on_messages returns the event count and propagates failures --
the orchestrator is the swallow-and-report layer now and needs the
exception to keep failed messages un-marked for retry.
…-side close
- add_memory(add_after_step) now fires only when a round closes the turn
(assistant reply with no tool calls) and dispatches through the
backend's schedule_add when available: tool rounds are intermediate
state, and ingesting every round cost O(rounds x history) extraction
calls where the closing ingest covers the whole turn.
- an interrupted round advances the ingest ledger WITHOUT ingesting
(mark_ingested): a half-finished answer is not durable conversational
truth and must not be swept into the next turn's delta.
- cleanup_tools drains scheduled ingestion (flush only -- memory
instances are shared across agents of one store, so closing here would
yank the store from a sibling agent); the new
SharedMemoryManager.close_matching(base_dir) is the owner-of-last-
resort that actually closes instances and releases the embedded
store's exclusive file lock.
The number of recalled memories injected per turn was hardcoded twice
(search default 20, then a [:10] formatting slice). MemoryConfig gains
recall_top_k (default 10, read from the unified_memory node) and the
mem0 adapter threads it through search and formatting — consumers can
now size recall to their context budget.
…E) instead of scattered config fields, gated by personalization.enabled.
When those files change mid-conversation the next user turn carries a durable <system-reminder> naming them, so the model can tell a changed file from its own faulty memory.
…end's MEMORY.md snapshot in step with edits made outside the agent.
Also translates the memory tool descriptions and prompt headings to English.
# Conflicts:
#	.gitignore
#	ms_agent/memory/unified/backends/mem0_adapter.py
#	ms_agent/memory/unified/orchestrator.py
#	setup.py
…ts last entry deleted, instead of leaving the previous round's block in place.
- an interrupt marks only its own round, and never messages a scheduled
ingest still owns (both lost the write silently)
- one shared instance per store, not per model, reconfigured in place
- close() is terminal: a straggler can no longer reopen a released store
- the store lock is per (loop, path), and search() takes it too
- search() honours its limit; injected memories carry their date
- memories are written in the language the user used
# Conflicts:
#	ms_agent/memory/unified/backends/mem0_adapter.py
…ters
Thinking support is per-model with no naming rule, and an unsupported model may
reject the whole request (DashScope returns 400) instead of ignoring the flag.
So we ask, and on a refusal retry once with it off, remembering the model.
…orwards, and read OpenRouter's reasoning field
…o a vision-disabled model neither claims nor disowns them
…x/runtime-robustness
# Conflicts:
#	ms_agent/agent/llm_agent.py
#	ms_agent/memory/unified/backends/mem0_adapter.py
#	ms_agent/memory/unified/orchestrator.py
…hat arrive mid-stream
- images go out only when the model's own switch says so; a provider's declared
vision capability no longer implies it
- a 400 delivered on the first streamed chunk is repaired like an eager one
- thinking refusals are repaired on the Anthropic and Responses paths too
- a tool call the model is still writing is reported instead of nothing at all
- an unreadable managed MCP config is logged instead of silently yielding none
… tell the agent why a search failed instead of reporting no results
…ps cutting structured results into invalid JSON
… and gate interpreter execution and credential reads behind mode-aware asks
… unique tool-name prefixes, and return content instead of a stub on unchanged re-reads
…io startup timeouts and per-server failure isolation
…ranscript matches are not mistaken for user content
@alcholiclg
alcholiclg merged commit a8ad587 into modelscope:mainAug 27, 2026
1 check failed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@alcholiclg@wangxingjun778
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Runtime robustness: shell execution, safety policy, permission asks, MCP startup, and tool-call handling - #951

Merged
alcholiclg merged 55 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness
Aug 27, 2026
Merged

Runtime robustness: shell execution, safety policy, permission asks, MCP startup, and tool-call handling#951
alcholiclg merged 55 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness

Conversation

@alcholiclg

Copy link
Copy Markdown
Collaborator

Change Summary

  • Shell safety policy: quote-aware redirect parsing (fixes the false Create path outside allowed directories: /dev/null) block), heredoc bodies treated as code rather than paths, subshell/brace-group unwrapping, OS temp dir writable by default; real violations (writes outside allowed dirs, glob creates, variable-expansion targets) still refused
  • New ask categories: interpreter_exec for inline interpreter code (python3 -c, heredocs) — rememberable per project via "always allow"; sensitive_read for credential paths (~/.ssh/*, *.pem, …) — denied in auto mode, never rememberable, kept separate from the write-protection list so e.g. .git/config stays readable
  • Ask handling: interactive mode waits indefinitely; full-access times out after 25 min with feedback that marks it a timeout, not a refusal; public pending-ask APIs (is_awaiting / awaiting_request_ids / resolve_matching / cancel_pending) so front-ends stop reading private state
  • Shell executor: commands run verbatim in a non-login shell — python3 -V and python3 -V ; true no longer resolve different interpreters; agent-friendly environment (PAGER=cat, GIT_TERMINAL_PROMPT=0, NO_COLOR=1, …) injected unconditionally; tool description states no state persists between calls
  • MCP: parallel connects under per-server owner tasks (anyio-safe teardown), stdio startup timeout, per-server failure isolation; Pydantic union errors deduplicated with actionable guidance prepended
  • Tool calls: schema-driven coercion of string-typed numerics ("19.5" → 19.5), unique-prefix tool-name resolution with candidate suggestions on miss, read_file re-reads return content with an "unchanged" note instead of a stub
  • Prompting: the system prompt documents framework-managed directories (sessions/, .ms_agent/) so transcript matches are not mistaken for user content
  • Tests: 6 new test files + 2 extended; suite baseline unchanged (25 pre-existing env/network failures); verified end-to-end through the WebUI (approval flows, policy blocks, MCP coercion, PATH/env)

Related issue number

Checklist

  • The pull request title is a good summary of the changes - it will be used in the changelog
  • Unit tests for the changes exist
  • Run pre-commit install and pre-commit run --all-files before git commit, and passed lint check.
  • Documentation reflects the changes where applicable

The orchestrator now owns the write discipline around a backend:
- schedule_add() runs the extraction-LLM + embedding cost (seconds) in a
background task; flush_pending() is the teardown barrier so the last
write is never dropped, and an inline fallback keeps writes when no
loop is running.
- retrieval/ingestion/flush serialize on one per-store asyncio lock
(embedded qdrant underneath is lock-free single-client code).
- a content-hash delta ledger (<base_dir>/ingest_state.json) makes each
ingest send only messages the store has not seen; hashes are recorded
only after a confirmed write, so a failed ingest retries naturally.
- ingest_status reports the last outcome (state/count/error/pending) so
a UI can show memory working instead of silence.
Mem0Backend: per-turn retrieval cache (rounds 2..N of a tool-calling
turn reuse round 1's search instead of paying an embedding round-trip
each), on_messages returns the event count and propagates failures --
the orchestrator is the swallow-and-report layer now and needs the
exception to keep failed messages un-marked for retry.
…-side close
- add_memory(add_after_step) now fires only when a round closes the turn
(assistant reply with no tool calls) and dispatches through the
backend's schedule_add when available: tool rounds are intermediate
state, and ingesting every round cost O(rounds x history) extraction
calls where the closing ingest covers the whole turn.
- an interrupted round advances the ingest ledger WITHOUT ingesting
(mark_ingested): a half-finished answer is not durable conversational
truth and must not be swept into the next turn's delta.
- cleanup_tools drains scheduled ingestion (flush only -- memory
instances are shared across agents of one store, so closing here would
yank the store from a sibling agent); the new
SharedMemoryManager.close_matching(base_dir) is the owner-of-last-
resort that actually closes instances and releases the embedded
store's exclusive file lock.
The number of recalled memories injected per turn was hardcoded twice
(search default 20, then a [:10] formatting slice). MemoryConfig gains
recall_top_k (default 10, read from the unified_memory node) and the
mem0 adapter threads it through search and formatting — consumers can
now size recall to their context budget.
…E) instead of scattered config fields, gated by personalization.enabled.
When those files change mid-conversation the next user turn carries a durable <system-reminder> naming them, so the model can tell a changed file from its own faulty memory.
…end's MEMORY.md snapshot in step with edits made outside the agent.
Also translates the memory tool descriptions and prompt headings to English.
# Conflicts:
#	.gitignore
#	ms_agent/memory/unified/backends/mem0_adapter.py
#	ms_agent/memory/unified/orchestrator.py
#	setup.py
…ts last entry deleted, instead of leaving the previous round's block in place.
- an interrupt marks only its own round, and never messages a scheduled
ingest still owns (both lost the write silently)
- one shared instance per store, not per model, reconfigured in place
- close() is terminal: a straggler can no longer reopen a released store
- the store lock is per (loop, path), and search() takes it too
- search() honours its limit; injected memories carry their date
- memories are written in the language the user used
# Conflicts:
#	ms_agent/memory/unified/backends/mem0_adapter.py
…ters
Thinking support is per-model with no naming rule, and an unsupported model may
reject the whole request (DashScope returns 400) instead of ignoring the flag.
So we ask, and on a refusal retry once with it off, remembering the model.
…orwards, and read OpenRouter's reasoning field
…o a vision-disabled model neither claims nor disowns them
…x/runtime-robustness
# Conflicts:
#	ms_agent/agent/llm_agent.py
#	ms_agent/memory/unified/backends/mem0_adapter.py
#	ms_agent/memory/unified/orchestrator.py
…hat arrive mid-stream
- images go out only when the model's own switch says so; a provider's declared
vision capability no longer implies it
- a 400 delivered on the first streamed chunk is repaired like an eager one
- thinking refusals are repaired on the Anthropic and Responses paths too
- a tool call the model is still writing is reported instead of nothing at all
- an unreadable managed MCP config is logged instead of silently yielding none
… tell the agent why a search failed instead of reporting no results
…ps cutting structured results into invalid JSON
… and gate interpreter execution and credential reads behind mode-aware asks
… unique tool-name prefixes, and return content instead of a stub on unchanged re-reads
…io startup timeouts and per-server failure isolation
…ranscript matches are not mistaken for user content
@alcholiclg
alcholiclg merged commit a8ad587 into modelscope:mainAug 27, 2026
1 check failed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@alcholiclg@wangxingjun778
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Runtime robustness: shell execution, safety policy, permission asks, MCP startup, and tool-call handling - #951

Merged
alcholiclg merged 55 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness
Aug 27, 2026
Merged

Runtime robustness: shell execution, safety policy, permission asks, MCP startup, and tool-call handling#951
alcholiclg merged 55 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness

Conversation

@alcholiclg

Copy link
Copy Markdown
Collaborator

Change Summary

  • Shell safety policy: quote-aware redirect parsing (fixes the false Create path outside allowed directories: /dev/null) block), heredoc bodies treated as code rather than paths, subshell/brace-group unwrapping, OS temp dir writable by default; real violations (writes outside allowed dirs, glob creates, variable-expansion targets) still refused
  • New ask categories: interpreter_exec for inline interpreter code (python3 -c, heredocs) — rememberable per project via "always allow"; sensitive_read for credential paths (~/.ssh/*, *.pem, …) — denied in auto mode, never rememberable, kept separate from the write-protection list so e.g. .git/config stays readable
  • Ask handling: interactive mode waits indefinitely; full-access times out after 25 min with feedback that marks it a timeout, not a refusal; public pending-ask APIs (is_awaiting / awaiting_request_ids / resolve_matching / cancel_pending) so front-ends stop reading private state
  • Shell executor: commands run verbatim in a non-login shell — python3 -V and python3 -V ; true no longer resolve different interpreters; agent-friendly environment (PAGER=cat, GIT_TERMINAL_PROMPT=0, NO_COLOR=1, …) injected unconditionally; tool description states no state persists between calls
  • MCP: parallel connects under per-server owner tasks (anyio-safe teardown), stdio startup timeout, per-server failure isolation; Pydantic union errors deduplicated with actionable guidance prepended
  • Tool calls: schema-driven coercion of string-typed numerics ("19.5" → 19.5), unique-prefix tool-name resolution with candidate suggestions on miss, read_file re-reads return content with an "unchanged" note instead of a stub
  • Prompting: the system prompt documents framework-managed directories (sessions/, .ms_agent/) so transcript matches are not mistaken for user content
  • Tests: 6 new test files + 2 extended; suite baseline unchanged (25 pre-existing env/network failures); verified end-to-end through the WebUI (approval flows, policy blocks, MCP coercion, PATH/env)

Related issue number

Checklist

  • The pull request title is a good summary of the changes - it will be used in the changelog
  • Unit tests for the changes exist
  • Run pre-commit install and pre-commit run --all-files before git commit, and passed lint check.
  • Documentation reflects the changes where applicable

The orchestrator now owns the write discipline around a backend:
- schedule_add() runs the extraction-LLM + embedding cost (seconds) in a
background task; flush_pending() is the teardown barrier so the last
write is never dropped, and an inline fallback keeps writes when no
loop is running.
- retrieval/ingestion/flush serialize on one per-store asyncio lock
(embedded qdrant underneath is lock-free single-client code).
- a content-hash delta ledger (<base_dir>/ingest_state.json) makes each
ingest send only messages the store has not seen; hashes are recorded
only after a confirmed write, so a failed ingest retries naturally.
- ingest_status reports the last outcome (state/count/error/pending) so
a UI can show memory working instead of silence.
Mem0Backend: per-turn retrieval cache (rounds 2..N of a tool-calling
turn reuse round 1's search instead of paying an embedding round-trip
each), on_messages returns the event count and propagates failures --
the orchestrator is the swallow-and-report layer now and needs the
exception to keep failed messages un-marked for retry.
…-side close
- add_memory(add_after_step) now fires only when a round closes the turn
(assistant reply with no tool calls) and dispatches through the
backend's schedule_add when available: tool rounds are intermediate
state, and ingesting every round cost O(rounds x history) extraction
calls where the closing ingest covers the whole turn.
- an interrupted round advances the ingest ledger WITHOUT ingesting
(mark_ingested): a half-finished answer is not durable conversational
truth and must not be swept into the next turn's delta.
- cleanup_tools drains scheduled ingestion (flush only -- memory
instances are shared across agents of one store, so closing here would
yank the store from a sibling agent); the new
SharedMemoryManager.close_matching(base_dir) is the owner-of-last-
resort that actually closes instances and releases the embedded
store's exclusive file lock.
The number of recalled memories injected per turn was hardcoded twice
(search default 20, then a [:10] formatting slice). MemoryConfig gains
recall_top_k (default 10, read from the unified_memory node) and the
mem0 adapter threads it through search and formatting — consumers can
now size recall to their context budget.
…E) instead of scattered config fields, gated by personalization.enabled.
When those files change mid-conversation the next user turn carries a durable <system-reminder> naming them, so the model can tell a changed file from its own faulty memory.
…end's MEMORY.md snapshot in step with edits made outside the agent.
Also translates the memory tool descriptions and prompt headings to English.
# Conflicts:
#	.gitignore
#	ms_agent/memory/unified/backends/mem0_adapter.py
#	ms_agent/memory/unified/orchestrator.py
#	setup.py
…ts last entry deleted, instead of leaving the previous round's block in place.
- an interrupt marks only its own round, and never messages a scheduled
ingest still owns (both lost the write silently)
- one shared instance per store, not per model, reconfigured in place
- close() is terminal: a straggler can no longer reopen a released store
- the store lock is per (loop, path), and search() takes it too
- search() honours its limit; injected memories carry their date
- memories are written in the language the user used
# Conflicts:
#	ms_agent/memory/unified/backends/mem0_adapter.py
…ters
Thinking support is per-model with no naming rule, and an unsupported model may
reject the whole request (DashScope returns 400) instead of ignoring the flag.
So we ask, and on a refusal retry once with it off, remembering the model.
…orwards, and read OpenRouter's reasoning field
…o a vision-disabled model neither claims nor disowns them
…x/runtime-robustness
# Conflicts:
#	ms_agent/agent/llm_agent.py
#	ms_agent/memory/unified/backends/mem0_adapter.py
#	ms_agent/memory/unified/orchestrator.py
…hat arrive mid-stream
- images go out only when the model's own switch says so; a provider's declared
vision capability no longer implies it
- a 400 delivered on the first streamed chunk is repaired like an eager one
- thinking refusals are repaired on the Anthropic and Responses paths too
- a tool call the model is still writing is reported instead of nothing at all
- an unreadable managed MCP config is logged instead of silently yielding none
… tell the agent why a search failed instead of reporting no results
…ps cutting structured results into invalid JSON
… and gate interpreter execution and credential reads behind mode-aware asks
… unique tool-name prefixes, and return content instead of a stub on unchanged re-reads
…io startup timeouts and per-server failure isolation
…ranscript matches are not mistaken for user content
@alcholiclg
alcholiclg merged commit a8ad587 into modelscope:mainAug 27, 2026
1 check failed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@alcholiclg@wangxingjun778
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Runtime robustness: shell execution, safety policy, permission asks, MCP startup, and tool-call handling - #951

Merged
alcholiclg merged 55 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness
Aug 27, 2026
Merged

Runtime robustness: shell execution, safety policy, permission asks, MCP startup, and tool-call handling#951
alcholiclg merged 55 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness

Conversation

@alcholiclg

Copy link
Copy Markdown
Collaborator

Change Summary

  • Shell safety policy: quote-aware redirect parsing (fixes the false Create path outside allowed directories: /dev/null) block), heredoc bodies treated as code rather than paths, subshell/brace-group unwrapping, OS temp dir writable by default; real violations (writes outside allowed dirs, glob creates, variable-expansion targets) still refused
  • New ask categories: interpreter_exec for inline interpreter code (python3 -c, heredocs) — rememberable per project via "always allow"; sensitive_read for credential paths (~/.ssh/*, *.pem, …) — denied in auto mode, never rememberable, kept separate from the write-protection list so e.g. .git/config stays readable
  • Ask handling: interactive mode waits indefinitely; full-access times out after 25 min with feedback that marks it a timeout, not a refusal; public pending-ask APIs (is_awaiting / awaiting_request_ids / resolve_matching / cancel_pending) so front-ends stop reading private state
  • Shell executor: commands run verbatim in a non-login shell — python3 -V and python3 -V ; true no longer resolve different interpreters; agent-friendly environment (PAGER=cat, GIT_TERMINAL_PROMPT=0, NO_COLOR=1, …) injected unconditionally; tool description states no state persists between calls
  • MCP: parallel connects under per-server owner tasks (anyio-safe teardown), stdio startup timeout, per-server failure isolation; Pydantic union errors deduplicated with actionable guidance prepended
  • Tool calls: schema-driven coercion of string-typed numerics ("19.5" → 19.5), unique-prefix tool-name resolution with candidate suggestions on miss, read_file re-reads return content with an "unchanged" note instead of a stub
  • Prompting: the system prompt documents framework-managed directories (sessions/, .ms_agent/) so transcript matches are not mistaken for user content
  • Tests: 6 new test files + 2 extended; suite baseline unchanged (25 pre-existing env/network failures); verified end-to-end through the WebUI (approval flows, policy blocks, MCP coercion, PATH/env)

Related issue number

Checklist

  • The pull request title is a good summary of the changes - it will be used in the changelog
  • Unit tests for the changes exist
  • Run pre-commit install and pre-commit run --all-files before git commit, and passed lint check.
  • Documentation reflects the changes where applicable

The orchestrator now owns the write discipline around a backend:
- schedule_add() runs the extraction-LLM + embedding cost (seconds) in a
background task; flush_pending() is the teardown barrier so the last
write is never dropped, and an inline fallback keeps writes when no
loop is running.
- retrieval/ingestion/flush serialize on one per-store asyncio lock
(embedded qdrant underneath is lock-free single-client code).
- a content-hash delta ledger (<base_dir>/ingest_state.json) makes each
ingest send only messages the store has not seen; hashes are recorded
only after a confirmed write, so a failed ingest retries naturally.
- ingest_status reports the last outcome (state/count/error/pending) so
a UI can show memory working instead of silence.
Mem0Backend: per-turn retrieval cache (rounds 2..N of a tool-calling
turn reuse round 1's search instead of paying an embedding round-trip
each), on_messages returns the event count and propagates failures --
the orchestrator is the swallow-and-report layer now and needs the
exception to keep failed messages un-marked for retry.
…-side close
- add_memory(add_after_step) now fires only when a round closes the turn
(assistant reply with no tool calls) and dispatches through the
backend's schedule_add when available: tool rounds are intermediate
state, and ingesting every round cost O(rounds x history) extraction
calls where the closing ingest covers the whole turn.
- an interrupted round advances the ingest ledger WITHOUT ingesting
(mark_ingested): a half-finished answer is not durable conversational
truth and must not be swept into the next turn's delta.
- cleanup_tools drains scheduled ingestion (flush only -- memory
instances are shared across agents of one store, so closing here would
yank the store from a sibling agent); the new
SharedMemoryManager.close_matching(base_dir) is the owner-of-last-
resort that actually closes instances and releases the embedded
store's exclusive file lock.
The number of recalled memories injected per turn was hardcoded twice
(search default 20, then a [:10] formatting slice). MemoryConfig gains
recall_top_k (default 10, read from the unified_memory node) and the
mem0 adapter threads it through search and formatting — consumers can
now size recall to their context budget.
…E) instead of scattered config fields, gated by personalization.enabled.
When those files change mid-conversation the next user turn carries a durable <system-reminder> naming them, so the model can tell a changed file from its own faulty memory.
…end's MEMORY.md snapshot in step with edits made outside the agent.
Also translates the memory tool descriptions and prompt headings to English.
# Conflicts:
#	.gitignore
#	ms_agent/memory/unified/backends/mem0_adapter.py
#	ms_agent/memory/unified/orchestrator.py
#	setup.py
…ts last entry deleted, instead of leaving the previous round's block in place.
- an interrupt marks only its own round, and never messages a scheduled
ingest still owns (both lost the write silently)
- one shared instance per store, not per model, reconfigured in place
- close() is terminal: a straggler can no longer reopen a released store
- the store lock is per (loop, path), and search() takes it too
- search() honours its limit; injected memories carry their date
- memories are written in the language the user used
# Conflicts:
#	ms_agent/memory/unified/backends/mem0_adapter.py
…ters
Thinking support is per-model with no naming rule, and an unsupported model may
reject the whole request (DashScope returns 400) instead of ignoring the flag.
So we ask, and on a refusal retry once with it off, remembering the model.
…orwards, and read OpenRouter's reasoning field
…o a vision-disabled model neither claims nor disowns them
…x/runtime-robustness
# Conflicts:
#	ms_agent/agent/llm_agent.py
#	ms_agent/memory/unified/backends/mem0_adapter.py
#	ms_agent/memory/unified/orchestrator.py
…hat arrive mid-stream
- images go out only when the model's own switch says so; a provider's declared
vision capability no longer implies it
- a 400 delivered on the first streamed chunk is repaired like an eager one
- thinking refusals are repaired on the Anthropic and Responses paths too
- a tool call the model is still writing is reported instead of nothing at all
- an unreadable managed MCP config is logged instead of silently yielding none
… tell the agent why a search failed instead of reporting no results
…ps cutting structured results into invalid JSON
… and gate interpreter execution and credential reads behind mode-aware asks
… unique tool-name prefixes, and return content instead of a stub on unchanged re-reads
…io startup timeouts and per-server failure isolation
…ranscript matches are not mistaken for user content
@alcholiclg
alcholiclg merged commit a8ad587 into modelscope:mainAug 27, 2026
1 check failed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@alcholiclg@wangxingjun778