Add current local-runtime and inference wire contracts - #3

Merged
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest
Aug 30, 2026
Merged

Add current local-runtime and inference wire contracts#3
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest

Conversation

@senamakel

Copy link
Copy Markdown
Member

Summary

  • add runtime-neutral reasoning, cache, continuation, and provider metadata needed by current TinyAgents
  • preserve current local runtime probing, warm-up, model validation, and Ollama embedding behavior
  • fix strict schema fallback, tool id normalization, cache accounting, Responses request/reasoning handling, and JSON Schema union validation
  • keep tracing test-only and use httpdate instead of chrono

Verification

  • cargo test
  • cargo clippy --all-targets -- -D warnings
  • downstream: cargo test --workspace
  • downstream: cargo clippy --workspace --all-targets -- -D warnings

Co-authored-by: Medulla <medulla@tinyhumans.ai>
@chatgpt-codex-connector

chatgpt-codex-connectorBot commented Aug 30, 2026

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

ReviewStatusCommitReview trigger
📝 Code ReviewCompleted2026-08-30T17:59:12.023682Z70412cePR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

Copy link
Copy Markdown

Warning

Review limit reached

  • Run on-demand review

On-demand reviews are free for the next 21 days. After that, they cost $0.25 per reviewed file.

Or wait 8 minutes for your next included review.

View limit details

Limit details: You’ve used the included review currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 4c6cd872-9fcd-4f14-9b4d-cbf8d31df189

📥 Commits

Reviewing files that changed from the base of the PR and between 2c40a5f and 70412ce.

📒 Files selected for processing (20)
  • crates/tinyinference/src/cache/mod.rs
  • crates/tinyinference/src/cache/types.rs
  • crates/tinyinference/src/embeddings/mod.rs
  • crates/tinyinference/src/embeddings/ollama.rs
  • crates/tinyinference/src/message/mod.rs
  • crates/tinyinference/src/message/types.rs
  • crates/tinyinference/src/model/mod.rs
  • crates/tinyinference/src/model/types.rs
  • crates/tinyinference/src/providers/mock.rs
  • crates/tinyinference/src/providers/openai/convert.rs
  • crates/tinyinference/src/providers/openai/local.rs
  • crates/tinyinference/src/providers/openai/local_test.rs
  • crates/tinyinference/src/providers/openai/mod.rs
  • crates/tinyinference/src/providers/openai/responses.rs
  • crates/tinyinference/src/providers/openai/sse.rs
  • crates/tinyinference/src/providers/openai/test.rs
  • crates/tinyinference/src/providers/openai/transport.rs
  • crates/tinyinference/src/providers/openai/types.rs
  • crates/tinyinference/src/providers/types.rs
  • crates/tinyinference/src/tool.rs

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@tinysweeper

Copy link
Copy Markdown

How this change flows

0 changed behaviours across 2 relationships. 2 surrounding behaviours are shown (60 graph nodes walked). 49 further behaviours left out to keep the diagram readable.

flowchart LR
n0["translate_request"]:::impacted
n1["translate_request_with"]:::impacted
n0 -->|calls| n1
n0 -->|tests| n1
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading

Green: changed behaviour. Grey: surrounding behaviour. Arrows name the call, use, implementation, or test relationship. Orange: has findings. Red: has a finding that blocks the merge.

tinysweeper 0.1.0

@tinysweepertinysweeperBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

$0.0000 · 0 in / 0 out

@senamakel
senamakel merged commit cc8aca4 into mainAug 30, 2026
13 checks passed

@chatgpt-codex-connectorchatgpt-codex-connectorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:70412ce592

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +1415 to +1416
if degrade.native_tools {
self.native_tools_on_wire.store(false, Ordering::Relaxed);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Parse prompt-guided calls after degrading native tools

When a model profile advertises tool calling but the server returns a 400 saying tools are unsupported, this latch makes the retry use prompt-guided <tool_call> output. The response paths at invoke and stream, however, only call prompt_tools::apply_to_response when self.profile.tool_calling is false; the profile remains true here, so the retry can succeed while returning the tool-call markup as ordinary assistant text instead of a normalized call. Base response parsing on this latch as well as the static profile.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

} else {
Vec::new()
};
let tool_choice = (!tools.is_empty()).then(|| translate_tool_choice(&request.tool_choice));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Use the Responses shape for named tool choice

For ToolChoice::Tool, this reuses the Chat Completions translator and serializes {"type":"function","function":{"name":...}}. The Responses API uses the flattened {"type":"function","name":...} shape, which the removed responses_tool_choice helper previously produced, so named-tool requests on the Responses path are rejected instead of forcing the requested tool.

Useful? React with 👍 / 👎.

Comment on lines +470 to 475
ModelResponse {
message: AssistantMessage {
id: None,
content: vec![ContentBlock::Text(text)],
tool_calls,
content,
tool_calls: Vec::new(),
usage,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Decode Responses function calls

Whenever /responses returns a function_call output item, this now unconditionally emits an empty tool_calls list; the response wire struct also no longer retains the call id, name, or arguments. The previous parser normalized these items, including malformed arguments, so tool-using Responses calls now appear to consumers as empty assistant replies and the requested tool is never executed.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +312 to +313
Message::User(m) => message_text(&m.content),
Message::Assistant(m) => message_text(&m.content),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve images in Responses input

For a Responses-primary request containing ContentBlock::Image, message_text retains only text blocks, so this branch silently discards every image; an image-only user turn is skipped altogether at the subsequent empty-text check. The prior translation emitted input_image content parts, so vision requests now run against missing input and can return plausible but incorrect results rather than an error.

Useful? React with 👍 / 👎.

} else {
"stop".to_string()
}),
finish_reason: Some("stop".to_string()),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve incomplete Responses finish reasons

Every Responses result is now marked stop, including responses whose status is incomplete because max_output_tokens was reached. Consumers therefore cannot distinguish truncated output from a completed answer and may accept or cache partial text/JSON; retain the response status and map incomplete_details.reason as the previous parser did.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +186 to +191
model_info
.iter()
.filter(|(key, _)| key.ends_with(".context_length") || key.as_str() == "context_length")
.filter_map(|(_, value)| value.as_u64())
.filter(|value| *value > 0)
.min()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Select the language-model context length

On multimodal Ollama models, model_info can contain both the language architecture window (for example gemma3.context_length) and a small projector window such as clip.context_length = 77. Taking the minimum advertises 77 tokens as the chat model's input capacity, causing capability checks or context compaction to reject or discard nearly every normal prompt; identify the main architecture's context field instead of minimizing unrelated components.

Useful? React with 👍 / 👎.

Comment on lines 343 to +344
if live.is_empty() {
return Err(Error::Validation(
"Ollama embedding batches must not contain blank inputs".into(),
));
return Ok(vec![Vec::new(); texts.len()]);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reject blank embedding batches instead of returning zero-width vectors

For any nonzero-dimensional OllamaEmbeddingModel, an all-blank batch now returns one empty vector per input, violating EmbeddingModel's fixed-dimension contract. Passing such output through Retriever::index immediately fails in InMemoryVectorStore::add, while direct callers can accidentally persist invalid vectors elsewhere; return a validation error or valid vectors of dimensions() instead.

AGENTS.md reference: AGENTS.md:L53-L58

Useful? React with 👍 / 👎.

Comment on lines +452 to +457
let parsed: ResponsesResponse =
serde_json::from_value(value.clone()).unwrap_or_else(|_| ResponsesResponse {
output: Vec::new(),
output_text: None,
usage: None,
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Propagate invalid Responses payload shapes

If a successful HTTP response has an incompatible schema, such as {"output":"not-an-array"}, deserialization now falls back to an empty response and invoke_responses returns a successful blank assistant message. This hides provider incompatibilities and malformed payloads that previously surfaced as serialization errors, making failures indistinguishable from genuine empty completions; keep parsing fallible and propagate the decode error.

Useful? React with 👍 / 👎.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@senamakel
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Add current local-runtime and inference wire contracts - #3

Merged
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest
Aug 30, 2026
Merged

Add current local-runtime and inference wire contracts#3
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest

Conversation

@senamakel

Copy link
Copy Markdown
Member

Summary

  • add runtime-neutral reasoning, cache, continuation, and provider metadata needed by current TinyAgents
  • preserve current local runtime probing, warm-up, model validation, and Ollama embedding behavior
  • fix strict schema fallback, tool id normalization, cache accounting, Responses request/reasoning handling, and JSON Schema union validation
  • keep tracing test-only and use httpdate instead of chrono

Verification

  • cargo test
  • cargo clippy --all-targets -- -D warnings
  • downstream: cargo test --workspace
  • downstream: cargo clippy --workspace --all-targets -- -D warnings

Co-authored-by: Medulla <medulla@tinyhumans.ai>
@chatgpt-codex-connector

chatgpt-codex-connectorBot commented Aug 30, 2026

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

ReviewStatusCommitReview trigger
📝 Code ReviewCompleted2026-08-30T17:59:12.023682Z70412cePR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

Copy link
Copy Markdown

Warning

Review limit reached

  • Run on-demand review

On-demand reviews are free for the next 21 days. After that, they cost $0.25 per reviewed file.

Or wait 8 minutes for your next included review.

View limit details

Limit details: You’ve used the included review currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 4c6cd872-9fcd-4f14-9b4d-cbf8d31df189

📥 Commits

Reviewing files that changed from the base of the PR and between 2c40a5f and 70412ce.

📒 Files selected for processing (20)
  • crates/tinyinference/src/cache/mod.rs
  • crates/tinyinference/src/cache/types.rs
  • crates/tinyinference/src/embeddings/mod.rs
  • crates/tinyinference/src/embeddings/ollama.rs
  • crates/tinyinference/src/message/mod.rs
  • crates/tinyinference/src/message/types.rs
  • crates/tinyinference/src/model/mod.rs
  • crates/tinyinference/src/model/types.rs
  • crates/tinyinference/src/providers/mock.rs
  • crates/tinyinference/src/providers/openai/convert.rs
  • crates/tinyinference/src/providers/openai/local.rs
  • crates/tinyinference/src/providers/openai/local_test.rs
  • crates/tinyinference/src/providers/openai/mod.rs
  • crates/tinyinference/src/providers/openai/responses.rs
  • crates/tinyinference/src/providers/openai/sse.rs
  • crates/tinyinference/src/providers/openai/test.rs
  • crates/tinyinference/src/providers/openai/transport.rs
  • crates/tinyinference/src/providers/openai/types.rs
  • crates/tinyinference/src/providers/types.rs
  • crates/tinyinference/src/tool.rs

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@tinysweeper

Copy link
Copy Markdown

How this change flows

0 changed behaviours across 2 relationships. 2 surrounding behaviours are shown (60 graph nodes walked). 49 further behaviours left out to keep the diagram readable.

flowchart LR
n0["translate_request"]:::impacted
n1["translate_request_with"]:::impacted
n0 -->|calls| n1
n0 -->|tests| n1
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading

Green: changed behaviour. Grey: surrounding behaviour. Arrows name the call, use, implementation, or test relationship. Orange: has findings. Red: has a finding that blocks the merge.

tinysweeper 0.1.0

@tinysweepertinysweeperBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

$0.0000 · 0 in / 0 out

@senamakel
senamakel merged commit cc8aca4 into mainAug 30, 2026
13 checks passed

@chatgpt-codex-connectorchatgpt-codex-connectorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:70412ce592

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +1415 to +1416
if degrade.native_tools {
self.native_tools_on_wire.store(false, Ordering::Relaxed);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Parse prompt-guided calls after degrading native tools

When a model profile advertises tool calling but the server returns a 400 saying tools are unsupported, this latch makes the retry use prompt-guided <tool_call> output. The response paths at invoke and stream, however, only call prompt_tools::apply_to_response when self.profile.tool_calling is false; the profile remains true here, so the retry can succeed while returning the tool-call markup as ordinary assistant text instead of a normalized call. Base response parsing on this latch as well as the static profile.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

} else {
Vec::new()
};
let tool_choice = (!tools.is_empty()).then(|| translate_tool_choice(&request.tool_choice));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Use the Responses shape for named tool choice

For ToolChoice::Tool, this reuses the Chat Completions translator and serializes {"type":"function","function":{"name":...}}. The Responses API uses the flattened {"type":"function","name":...} shape, which the removed responses_tool_choice helper previously produced, so named-tool requests on the Responses path are rejected instead of forcing the requested tool.

Useful? React with 👍 / 👎.

Comment on lines +470 to 475
ModelResponse {
message: AssistantMessage {
id: None,
content: vec![ContentBlock::Text(text)],
tool_calls,
content,
tool_calls: Vec::new(),
usage,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Decode Responses function calls

Whenever /responses returns a function_call output item, this now unconditionally emits an empty tool_calls list; the response wire struct also no longer retains the call id, name, or arguments. The previous parser normalized these items, including malformed arguments, so tool-using Responses calls now appear to consumers as empty assistant replies and the requested tool is never executed.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +312 to +313
Message::User(m) => message_text(&m.content),
Message::Assistant(m) => message_text(&m.content),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve images in Responses input

For a Responses-primary request containing ContentBlock::Image, message_text retains only text blocks, so this branch silently discards every image; an image-only user turn is skipped altogether at the subsequent empty-text check. The prior translation emitted input_image content parts, so vision requests now run against missing input and can return plausible but incorrect results rather than an error.

Useful? React with 👍 / 👎.

} else {
"stop".to_string()
}),
finish_reason: Some("stop".to_string()),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve incomplete Responses finish reasons

Every Responses result is now marked stop, including responses whose status is incomplete because max_output_tokens was reached. Consumers therefore cannot distinguish truncated output from a completed answer and may accept or cache partial text/JSON; retain the response status and map incomplete_details.reason as the previous parser did.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +186 to +191
model_info
.iter()
.filter(|(key, _)| key.ends_with(".context_length") || key.as_str() == "context_length")
.filter_map(|(_, value)| value.as_u64())
.filter(|value| *value > 0)
.min()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Select the language-model context length

On multimodal Ollama models, model_info can contain both the language architecture window (for example gemma3.context_length) and a small projector window such as clip.context_length = 77. Taking the minimum advertises 77 tokens as the chat model's input capacity, causing capability checks or context compaction to reject or discard nearly every normal prompt; identify the main architecture's context field instead of minimizing unrelated components.

Useful? React with 👍 / 👎.

Comment on lines 343 to +344
if live.is_empty() {
return Err(Error::Validation(
"Ollama embedding batches must not contain blank inputs".into(),
));
return Ok(vec![Vec::new(); texts.len()]);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reject blank embedding batches instead of returning zero-width vectors

For any nonzero-dimensional OllamaEmbeddingModel, an all-blank batch now returns one empty vector per input, violating EmbeddingModel's fixed-dimension contract. Passing such output through Retriever::index immediately fails in InMemoryVectorStore::add, while direct callers can accidentally persist invalid vectors elsewhere; return a validation error or valid vectors of dimensions() instead.

AGENTS.md reference: AGENTS.md:L53-L58

Useful? React with 👍 / 👎.

Comment on lines +452 to +457
let parsed: ResponsesResponse =
serde_json::from_value(value.clone()).unwrap_or_else(|_| ResponsesResponse {
output: Vec::new(),
output_text: None,
usage: None,
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Propagate invalid Responses payload shapes

If a successful HTTP response has an incompatible schema, such as {"output":"not-an-array"}, deserialization now falls back to an empty response and invoke_responses returns a successful blank assistant message. This hides provider incompatibilities and malformed payloads that previously surfaced as serialization errors, making failures indistinguishable from genuine empty completions; keep parsing fallible and propagate the decode error.

Useful? React with 👍 / 👎.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@senamakel
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add current local-runtime and inference wire contracts - #3

Merged
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest
Aug 30, 2026
Merged

Add current local-runtime and inference wire contracts#3
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest

Conversation

@senamakel

Copy link
Copy Markdown
Member

Summary

  • add runtime-neutral reasoning, cache, continuation, and provider metadata needed by current TinyAgents
  • preserve current local runtime probing, warm-up, model validation, and Ollama embedding behavior
  • fix strict schema fallback, tool id normalization, cache accounting, Responses request/reasoning handling, and JSON Schema union validation
  • keep tracing test-only and use httpdate instead of chrono

Verification

  • cargo test
  • cargo clippy --all-targets -- -D warnings
  • downstream: cargo test --workspace
  • downstream: cargo clippy --workspace --all-targets -- -D warnings

Co-authored-by: Medulla <medulla@tinyhumans.ai>
@chatgpt-codex-connector

chatgpt-codex-connectorBot commented Aug 30, 2026

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

ReviewStatusCommitReview trigger
📝 Code ReviewCompleted2026-08-30T17:59:12.023682Z70412cePR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

Copy link
Copy Markdown

Warning

Review limit reached

  • Run on-demand review

On-demand reviews are free for the next 21 days. After that, they cost $0.25 per reviewed file.

Or wait 8 minutes for your next included review.

View limit details

Limit details: You’ve used the included review currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 4c6cd872-9fcd-4f14-9b4d-cbf8d31df189

📥 Commits

Reviewing files that changed from the base of the PR and between 2c40a5f and 70412ce.

📒 Files selected for processing (20)
  • crates/tinyinference/src/cache/mod.rs
  • crates/tinyinference/src/cache/types.rs
  • crates/tinyinference/src/embeddings/mod.rs
  • crates/tinyinference/src/embeddings/ollama.rs
  • crates/tinyinference/src/message/mod.rs
  • crates/tinyinference/src/message/types.rs
  • crates/tinyinference/src/model/mod.rs
  • crates/tinyinference/src/model/types.rs
  • crates/tinyinference/src/providers/mock.rs
  • crates/tinyinference/src/providers/openai/convert.rs
  • crates/tinyinference/src/providers/openai/local.rs
  • crates/tinyinference/src/providers/openai/local_test.rs
  • crates/tinyinference/src/providers/openai/mod.rs
  • crates/tinyinference/src/providers/openai/responses.rs
  • crates/tinyinference/src/providers/openai/sse.rs
  • crates/tinyinference/src/providers/openai/test.rs
  • crates/tinyinference/src/providers/openai/transport.rs
  • crates/tinyinference/src/providers/openai/types.rs
  • crates/tinyinference/src/providers/types.rs
  • crates/tinyinference/src/tool.rs

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@tinysweeper

Copy link
Copy Markdown

How this change flows

0 changed behaviours across 2 relationships. 2 surrounding behaviours are shown (60 graph nodes walked). 49 further behaviours left out to keep the diagram readable.

flowchart LR
n0["translate_request"]:::impacted
n1["translate_request_with"]:::impacted
n0 -->|calls| n1
n0 -->|tests| n1
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading

Green: changed behaviour. Grey: surrounding behaviour. Arrows name the call, use, implementation, or test relationship. Orange: has findings. Red: has a finding that blocks the merge.

tinysweeper 0.1.0

@tinysweepertinysweeperBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

$0.0000 · 0 in / 0 out

@senamakel
senamakel merged commit cc8aca4 into mainAug 30, 2026
13 checks passed

@chatgpt-codex-connectorchatgpt-codex-connectorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:70412ce592

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +1415 to +1416
if degrade.native_tools {
self.native_tools_on_wire.store(false, Ordering::Relaxed);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Parse prompt-guided calls after degrading native tools

When a model profile advertises tool calling but the server returns a 400 saying tools are unsupported, this latch makes the retry use prompt-guided <tool_call> output. The response paths at invoke and stream, however, only call prompt_tools::apply_to_response when self.profile.tool_calling is false; the profile remains true here, so the retry can succeed while returning the tool-call markup as ordinary assistant text instead of a normalized call. Base response parsing on this latch as well as the static profile.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

} else {
Vec::new()
};
let tool_choice = (!tools.is_empty()).then(|| translate_tool_choice(&request.tool_choice));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Use the Responses shape for named tool choice

For ToolChoice::Tool, this reuses the Chat Completions translator and serializes {"type":"function","function":{"name":...}}. The Responses API uses the flattened {"type":"function","name":...} shape, which the removed responses_tool_choice helper previously produced, so named-tool requests on the Responses path are rejected instead of forcing the requested tool.

Useful? React with 👍 / 👎.

Comment on lines +470 to 475
ModelResponse {
message: AssistantMessage {
id: None,
content: vec![ContentBlock::Text(text)],
tool_calls,
content,
tool_calls: Vec::new(),
usage,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Decode Responses function calls

Whenever /responses returns a function_call output item, this now unconditionally emits an empty tool_calls list; the response wire struct also no longer retains the call id, name, or arguments. The previous parser normalized these items, including malformed arguments, so tool-using Responses calls now appear to consumers as empty assistant replies and the requested tool is never executed.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +312 to +313
Message::User(m) => message_text(&m.content),
Message::Assistant(m) => message_text(&m.content),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve images in Responses input

For a Responses-primary request containing ContentBlock::Image, message_text retains only text blocks, so this branch silently discards every image; an image-only user turn is skipped altogether at the subsequent empty-text check. The prior translation emitted input_image content parts, so vision requests now run against missing input and can return plausible but incorrect results rather than an error.

Useful? React with 👍 / 👎.

} else {
"stop".to_string()
}),
finish_reason: Some("stop".to_string()),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve incomplete Responses finish reasons

Every Responses result is now marked stop, including responses whose status is incomplete because max_output_tokens was reached. Consumers therefore cannot distinguish truncated output from a completed answer and may accept or cache partial text/JSON; retain the response status and map incomplete_details.reason as the previous parser did.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +186 to +191
model_info
.iter()
.filter(|(key, _)| key.ends_with(".context_length") || key.as_str() == "context_length")
.filter_map(|(_, value)| value.as_u64())
.filter(|value| *value > 0)
.min()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Select the language-model context length

On multimodal Ollama models, model_info can contain both the language architecture window (for example gemma3.context_length) and a small projector window such as clip.context_length = 77. Taking the minimum advertises 77 tokens as the chat model's input capacity, causing capability checks or context compaction to reject or discard nearly every normal prompt; identify the main architecture's context field instead of minimizing unrelated components.

Useful? React with 👍 / 👎.

Comment on lines 343 to +344
if live.is_empty() {
return Err(Error::Validation(
"Ollama embedding batches must not contain blank inputs".into(),
));
return Ok(vec![Vec::new(); texts.len()]);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reject blank embedding batches instead of returning zero-width vectors

For any nonzero-dimensional OllamaEmbeddingModel, an all-blank batch now returns one empty vector per input, violating EmbeddingModel's fixed-dimension contract. Passing such output through Retriever::index immediately fails in InMemoryVectorStore::add, while direct callers can accidentally persist invalid vectors elsewhere; return a validation error or valid vectors of dimensions() instead.

AGENTS.md reference: AGENTS.md:L53-L58

Useful? React with 👍 / 👎.

Comment on lines +452 to +457
let parsed: ResponsesResponse =
serde_json::from_value(value.clone()).unwrap_or_else(|_| ResponsesResponse {
output: Vec::new(),
output_text: None,
usage: None,
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Propagate invalid Responses payload shapes

If a successful HTTP response has an incompatible schema, such as {"output":"not-an-array"}, deserialization now falls back to an empty response and invoke_responses returns a successful blank assistant message. This hides provider incompatibilities and malformed payloads that previously surfaced as serialization errors, making failures indistinguishable from genuine empty completions; keep parsing fallible and propagate the decode error.

Useful? React with 👍 / 👎.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@senamakel
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add current local-runtime and inference wire contracts - #3

Merged
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest
Aug 30, 2026
Merged

Add current local-runtime and inference wire contracts#3
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest

Conversation

@senamakel

Copy link
Copy Markdown
Member

Summary

  • add runtime-neutral reasoning, cache, continuation, and provider metadata needed by current TinyAgents
  • preserve current local runtime probing, warm-up, model validation, and Ollama embedding behavior
  • fix strict schema fallback, tool id normalization, cache accounting, Responses request/reasoning handling, and JSON Schema union validation
  • keep tracing test-only and use httpdate instead of chrono

Verification

  • cargo test
  • cargo clippy --all-targets -- -D warnings
  • downstream: cargo test --workspace
  • downstream: cargo clippy --workspace --all-targets -- -D warnings

Co-authored-by: Medulla <medulla@tinyhumans.ai>
@chatgpt-codex-connector

chatgpt-codex-connectorBot commented Aug 30, 2026

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

ReviewStatusCommitReview trigger
📝 Code ReviewCompleted2026-08-30T17:59:12.023682Z70412cePR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

Copy link
Copy Markdown

Warning

Review limit reached

  • Run on-demand review

On-demand reviews are free for the next 21 days. After that, they cost $0.25 per reviewed file.

Or wait 8 minutes for your next included review.

View limit details

Limit details: You’ve used the included review currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 4c6cd872-9fcd-4f14-9b4d-cbf8d31df189

📥 Commits

Reviewing files that changed from the base of the PR and between 2c40a5f and 70412ce.

📒 Files selected for processing (20)
  • crates/tinyinference/src/cache/mod.rs
  • crates/tinyinference/src/cache/types.rs
  • crates/tinyinference/src/embeddings/mod.rs
  • crates/tinyinference/src/embeddings/ollama.rs
  • crates/tinyinference/src/message/mod.rs
  • crates/tinyinference/src/message/types.rs
  • crates/tinyinference/src/model/mod.rs
  • crates/tinyinference/src/model/types.rs
  • crates/tinyinference/src/providers/mock.rs
  • crates/tinyinference/src/providers/openai/convert.rs
  • crates/tinyinference/src/providers/openai/local.rs
  • crates/tinyinference/src/providers/openai/local_test.rs
  • crates/tinyinference/src/providers/openai/mod.rs
  • crates/tinyinference/src/providers/openai/responses.rs
  • crates/tinyinference/src/providers/openai/sse.rs
  • crates/tinyinference/src/providers/openai/test.rs
  • crates/tinyinference/src/providers/openai/transport.rs
  • crates/tinyinference/src/providers/openai/types.rs
  • crates/tinyinference/src/providers/types.rs
  • crates/tinyinference/src/tool.rs

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@tinysweeper

Copy link
Copy Markdown

How this change flows

0 changed behaviours across 2 relationships. 2 surrounding behaviours are shown (60 graph nodes walked). 49 further behaviours left out to keep the diagram readable.

flowchart LR
n0["translate_request"]:::impacted
n1["translate_request_with"]:::impacted
n0 -->|calls| n1
n0 -->|tests| n1
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading

Green: changed behaviour. Grey: surrounding behaviour. Arrows name the call, use, implementation, or test relationship. Orange: has findings. Red: has a finding that blocks the merge.

tinysweeper 0.1.0

@tinysweepertinysweeperBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

$0.0000 · 0 in / 0 out

@senamakel
senamakel merged commit cc8aca4 into mainAug 30, 2026
13 checks passed

@chatgpt-codex-connectorchatgpt-codex-connectorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:70412ce592

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +1415 to +1416
if degrade.native_tools {
self.native_tools_on_wire.store(false, Ordering::Relaxed);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Parse prompt-guided calls after degrading native tools

When a model profile advertises tool calling but the server returns a 400 saying tools are unsupported, this latch makes the retry use prompt-guided <tool_call> output. The response paths at invoke and stream, however, only call prompt_tools::apply_to_response when self.profile.tool_calling is false; the profile remains true here, so the retry can succeed while returning the tool-call markup as ordinary assistant text instead of a normalized call. Base response parsing on this latch as well as the static profile.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

} else {
Vec::new()
};
let tool_choice = (!tools.is_empty()).then(|| translate_tool_choice(&request.tool_choice));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Use the Responses shape for named tool choice

For ToolChoice::Tool, this reuses the Chat Completions translator and serializes {"type":"function","function":{"name":...}}. The Responses API uses the flattened {"type":"function","name":...} shape, which the removed responses_tool_choice helper previously produced, so named-tool requests on the Responses path are rejected instead of forcing the requested tool.

Useful? React with 👍 / 👎.

Comment on lines +470 to 475
ModelResponse {
message: AssistantMessage {
id: None,
content: vec![ContentBlock::Text(text)],
tool_calls,
content,
tool_calls: Vec::new(),
usage,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Decode Responses function calls

Whenever /responses returns a function_call output item, this now unconditionally emits an empty tool_calls list; the response wire struct also no longer retains the call id, name, or arguments. The previous parser normalized these items, including malformed arguments, so tool-using Responses calls now appear to consumers as empty assistant replies and the requested tool is never executed.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +312 to +313
Message::User(m) => message_text(&m.content),
Message::Assistant(m) => message_text(&m.content),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve images in Responses input

For a Responses-primary request containing ContentBlock::Image, message_text retains only text blocks, so this branch silently discards every image; an image-only user turn is skipped altogether at the subsequent empty-text check. The prior translation emitted input_image content parts, so vision requests now run against missing input and can return plausible but incorrect results rather than an error.

Useful? React with 👍 / 👎.

} else {
"stop".to_string()
}),
finish_reason: Some("stop".to_string()),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve incomplete Responses finish reasons

Every Responses result is now marked stop, including responses whose status is incomplete because max_output_tokens was reached. Consumers therefore cannot distinguish truncated output from a completed answer and may accept or cache partial text/JSON; retain the response status and map incomplete_details.reason as the previous parser did.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +186 to +191
model_info
.iter()
.filter(|(key, _)| key.ends_with(".context_length") || key.as_str() == "context_length")
.filter_map(|(_, value)| value.as_u64())
.filter(|value| *value > 0)
.min()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Select the language-model context length

On multimodal Ollama models, model_info can contain both the language architecture window (for example gemma3.context_length) and a small projector window such as clip.context_length = 77. Taking the minimum advertises 77 tokens as the chat model's input capacity, causing capability checks or context compaction to reject or discard nearly every normal prompt; identify the main architecture's context field instead of minimizing unrelated components.

Useful? React with 👍 / 👎.

Comment on lines 343 to +344
if live.is_empty() {
return Err(Error::Validation(
"Ollama embedding batches must not contain blank inputs".into(),
));
return Ok(vec![Vec::new(); texts.len()]);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reject blank embedding batches instead of returning zero-width vectors

For any nonzero-dimensional OllamaEmbeddingModel, an all-blank batch now returns one empty vector per input, violating EmbeddingModel's fixed-dimension contract. Passing such output through Retriever::index immediately fails in InMemoryVectorStore::add, while direct callers can accidentally persist invalid vectors elsewhere; return a validation error or valid vectors of dimensions() instead.

AGENTS.md reference: AGENTS.md:L53-L58

Useful? React with 👍 / 👎.

Comment on lines +452 to +457
let parsed: ResponsesResponse =
serde_json::from_value(value.clone()).unwrap_or_else(|_| ResponsesResponse {
output: Vec::new(),
output_text: None,
usage: None,
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Propagate invalid Responses payload shapes

If a successful HTTP response has an incompatible schema, such as {"output":"not-an-array"}, deserialization now falls back to an empty response and invoke_responses returns a successful blank assistant message. This hides provider incompatibilities and malformed payloads that previously surfaced as serialization errors, making failures indistinguishable from genuine empty completions; keep parsing fallible and propagate the decode error.

Useful? React with 👍 / 👎.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@senamakel
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Add current local-runtime and inference wire contracts - #3

Merged
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest
Aug 30, 2026
Merged

Add current local-runtime and inference wire contracts#3
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest

Conversation

@senamakel

Copy link
Copy Markdown
Member

Summary

  • add runtime-neutral reasoning, cache, continuation, and provider metadata needed by current TinyAgents
  • preserve current local runtime probing, warm-up, model validation, and Ollama embedding behavior
  • fix strict schema fallback, tool id normalization, cache accounting, Responses request/reasoning handling, and JSON Schema union validation
  • keep tracing test-only and use httpdate instead of chrono

Verification

  • cargo test
  • cargo clippy --all-targets -- -D warnings
  • downstream: cargo test --workspace
  • downstream: cargo clippy --workspace --all-targets -- -D warnings

Co-authored-by: Medulla <medulla@tinyhumans.ai>
@chatgpt-codex-connector

chatgpt-codex-connectorBot commented Aug 30, 2026

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

ReviewStatusCommitReview trigger
📝 Code ReviewCompleted2026-08-30T17:59:12.023682Z70412cePR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

Copy link
Copy Markdown

Warning

Review limit reached

  • Run on-demand review

On-demand reviews are free for the next 21 days. After that, they cost $0.25 per reviewed file.

Or wait 8 minutes for your next included review.

View limit details

Limit details: You’ve used the included review currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 4c6cd872-9fcd-4f14-9b4d-cbf8d31df189

📥 Commits

Reviewing files that changed from the base of the PR and between 2c40a5f and 70412ce.

📒 Files selected for processing (20)
  • crates/tinyinference/src/cache/mod.rs
  • crates/tinyinference/src/cache/types.rs
  • crates/tinyinference/src/embeddings/mod.rs
  • crates/tinyinference/src/embeddings/ollama.rs
  • crates/tinyinference/src/message/mod.rs
  • crates/tinyinference/src/message/types.rs
  • crates/tinyinference/src/model/mod.rs
  • crates/tinyinference/src/model/types.rs
  • crates/tinyinference/src/providers/mock.rs
  • crates/tinyinference/src/providers/openai/convert.rs
  • crates/tinyinference/src/providers/openai/local.rs
  • crates/tinyinference/src/providers/openai/local_test.rs
  • crates/tinyinference/src/providers/openai/mod.rs
  • crates/tinyinference/src/providers/openai/responses.rs
  • crates/tinyinference/src/providers/openai/sse.rs
  • crates/tinyinference/src/providers/openai/test.rs
  • crates/tinyinference/src/providers/openai/transport.rs
  • crates/tinyinference/src/providers/openai/types.rs
  • crates/tinyinference/src/providers/types.rs
  • crates/tinyinference/src/tool.rs

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@tinysweeper

Copy link
Copy Markdown

How this change flows

0 changed behaviours across 2 relationships. 2 surrounding behaviours are shown (60 graph nodes walked). 49 further behaviours left out to keep the diagram readable.

flowchart LR
n0["translate_request"]:::impacted
n1["translate_request_with"]:::impacted
n0 -->|calls| n1
n0 -->|tests| n1
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading

Green: changed behaviour. Grey: surrounding behaviour. Arrows name the call, use, implementation, or test relationship. Orange: has findings. Red: has a finding that blocks the merge.

tinysweeper 0.1.0

@tinysweepertinysweeperBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

$0.0000 · 0 in / 0 out

@senamakel
senamakel merged commit cc8aca4 into mainAug 30, 2026
13 checks passed

@chatgpt-codex-connectorchatgpt-codex-connectorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:70412ce592

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +1415 to +1416
if degrade.native_tools {
self.native_tools_on_wire.store(false, Ordering::Relaxed);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Parse prompt-guided calls after degrading native tools

When a model profile advertises tool calling but the server returns a 400 saying tools are unsupported, this latch makes the retry use prompt-guided <tool_call> output. The response paths at invoke and stream, however, only call prompt_tools::apply_to_response when self.profile.tool_calling is false; the profile remains true here, so the retry can succeed while returning the tool-call markup as ordinary assistant text instead of a normalized call. Base response parsing on this latch as well as the static profile.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

} else {
Vec::new()
};
let tool_choice = (!tools.is_empty()).then(|| translate_tool_choice(&request.tool_choice));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Use the Responses shape for named tool choice

For ToolChoice::Tool, this reuses the Chat Completions translator and serializes {"type":"function","function":{"name":...}}. The Responses API uses the flattened {"type":"function","name":...} shape, which the removed responses_tool_choice helper previously produced, so named-tool requests on the Responses path are rejected instead of forcing the requested tool.

Useful? React with 👍 / 👎.

Comment on lines +470 to 475
ModelResponse {
message: AssistantMessage {
id: None,
content: vec![ContentBlock::Text(text)],
tool_calls,
content,
tool_calls: Vec::new(),
usage,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Decode Responses function calls

Whenever /responses returns a function_call output item, this now unconditionally emits an empty tool_calls list; the response wire struct also no longer retains the call id, name, or arguments. The previous parser normalized these items, including malformed arguments, so tool-using Responses calls now appear to consumers as empty assistant replies and the requested tool is never executed.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +312 to +313
Message::User(m) => message_text(&m.content),
Message::Assistant(m) => message_text(&m.content),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve images in Responses input

For a Responses-primary request containing ContentBlock::Image, message_text retains only text blocks, so this branch silently discards every image; an image-only user turn is skipped altogether at the subsequent empty-text check. The prior translation emitted input_image content parts, so vision requests now run against missing input and can return plausible but incorrect results rather than an error.

Useful? React with 👍 / 👎.

} else {
"stop".to_string()
}),
finish_reason: Some("stop".to_string()),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve incomplete Responses finish reasons

Every Responses result is now marked stop, including responses whose status is incomplete because max_output_tokens was reached. Consumers therefore cannot distinguish truncated output from a completed answer and may accept or cache partial text/JSON; retain the response status and map incomplete_details.reason as the previous parser did.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +186 to +191
model_info
.iter()
.filter(|(key, _)| key.ends_with(".context_length") || key.as_str() == "context_length")
.filter_map(|(_, value)| value.as_u64())
.filter(|value| *value > 0)
.min()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Select the language-model context length

On multimodal Ollama models, model_info can contain both the language architecture window (for example gemma3.context_length) and a small projector window such as clip.context_length = 77. Taking the minimum advertises 77 tokens as the chat model's input capacity, causing capability checks or context compaction to reject or discard nearly every normal prompt; identify the main architecture's context field instead of minimizing unrelated components.

Useful? React with 👍 / 👎.

Comment on lines 343 to +344
if live.is_empty() {
return Err(Error::Validation(
"Ollama embedding batches must not contain blank inputs".into(),
));
return Ok(vec![Vec::new(); texts.len()]);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reject blank embedding batches instead of returning zero-width vectors

For any nonzero-dimensional OllamaEmbeddingModel, an all-blank batch now returns one empty vector per input, violating EmbeddingModel's fixed-dimension contract. Passing such output through Retriever::index immediately fails in InMemoryVectorStore::add, while direct callers can accidentally persist invalid vectors elsewhere; return a validation error or valid vectors of dimensions() instead.

AGENTS.md reference: AGENTS.md:L53-L58

Useful? React with 👍 / 👎.

Comment on lines +452 to +457
let parsed: ResponsesResponse =
serde_json::from_value(value.clone()).unwrap_or_else(|_| ResponsesResponse {
output: Vec::new(),
output_text: None,
usage: None,
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Propagate invalid Responses payload shapes

If a successful HTTP response has an incompatible schema, such as {"output":"not-an-array"}, deserialization now falls back to an empty response and invoke_responses returns a successful blank assistant message. This hides provider incompatibilities and malformed payloads that previously surfaced as serialization errors, making failures indistinguishable from genuine empty completions; keep parsing fallible and propagate the decode error.

Useful? React with 👍 / 👎.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@senamakel
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add current local-runtime and inference wire contracts - #3

Merged
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest
Aug 30, 2026
Merged

Add current local-runtime and inference wire contracts#3
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest

Conversation

@senamakel

Copy link
Copy Markdown
Member

Summary

  • add runtime-neutral reasoning, cache, continuation, and provider metadata needed by current TinyAgents
  • preserve current local runtime probing, warm-up, model validation, and Ollama embedding behavior
  • fix strict schema fallback, tool id normalization, cache accounting, Responses request/reasoning handling, and JSON Schema union validation
  • keep tracing test-only and use httpdate instead of chrono

Verification

  • cargo test
  • cargo clippy --all-targets -- -D warnings
  • downstream: cargo test --workspace
  • downstream: cargo clippy --workspace --all-targets -- -D warnings

Co-authored-by: Medulla <medulla@tinyhumans.ai>
@chatgpt-codex-connector

chatgpt-codex-connectorBot commented Aug 30, 2026

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

ReviewStatusCommitReview trigger
📝 Code ReviewCompleted2026-08-30T17:59:12.023682Z70412cePR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

Copy link
Copy Markdown

Warning

Review limit reached

  • Run on-demand review

On-demand reviews are free for the next 21 days. After that, they cost $0.25 per reviewed file.

Or wait 8 minutes for your next included review.

View limit details

Limit details: You’ve used the included review currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 4c6cd872-9fcd-4f14-9b4d-cbf8d31df189

📥 Commits

Reviewing files that changed from the base of the PR and between 2c40a5f and 70412ce.

📒 Files selected for processing (20)
  • crates/tinyinference/src/cache/mod.rs
  • crates/tinyinference/src/cache/types.rs
  • crates/tinyinference/src/embeddings/mod.rs
  • crates/tinyinference/src/embeddings/ollama.rs
  • crates/tinyinference/src/message/mod.rs
  • crates/tinyinference/src/message/types.rs
  • crates/tinyinference/src/model/mod.rs
  • crates/tinyinference/src/model/types.rs
  • crates/tinyinference/src/providers/mock.rs
  • crates/tinyinference/src/providers/openai/convert.rs
  • crates/tinyinference/src/providers/openai/local.rs
  • crates/tinyinference/src/providers/openai/local_test.rs
  • crates/tinyinference/src/providers/openai/mod.rs
  • crates/tinyinference/src/providers/openai/responses.rs
  • crates/tinyinference/src/providers/openai/sse.rs
  • crates/tinyinference/src/providers/openai/test.rs
  • crates/tinyinference/src/providers/openai/transport.rs
  • crates/tinyinference/src/providers/openai/types.rs
  • crates/tinyinference/src/providers/types.rs
  • crates/tinyinference/src/tool.rs

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@tinysweeper

Copy link
Copy Markdown

How this change flows

0 changed behaviours across 2 relationships. 2 surrounding behaviours are shown (60 graph nodes walked). 49 further behaviours left out to keep the diagram readable.

flowchart LR
n0["translate_request"]:::impacted
n1["translate_request_with"]:::impacted
n0 -->|calls| n1
n0 -->|tests| n1
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading

Green: changed behaviour. Grey: surrounding behaviour. Arrows name the call, use, implementation, or test relationship. Orange: has findings. Red: has a finding that blocks the merge.

tinysweeper 0.1.0

@tinysweepertinysweeperBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

$0.0000 · 0 in / 0 out

@senamakel
senamakel merged commit cc8aca4 into mainAug 30, 2026
13 checks passed

@chatgpt-codex-connectorchatgpt-codex-connectorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:70412ce592

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +1415 to +1416
if degrade.native_tools {
self.native_tools_on_wire.store(false, Ordering::Relaxed);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Parse prompt-guided calls after degrading native tools

When a model profile advertises tool calling but the server returns a 400 saying tools are unsupported, this latch makes the retry use prompt-guided <tool_call> output. The response paths at invoke and stream, however, only call prompt_tools::apply_to_response when self.profile.tool_calling is false; the profile remains true here, so the retry can succeed while returning the tool-call markup as ordinary assistant text instead of a normalized call. Base response parsing on this latch as well as the static profile.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

} else {
Vec::new()
};
let tool_choice = (!tools.is_empty()).then(|| translate_tool_choice(&request.tool_choice));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Use the Responses shape for named tool choice

For ToolChoice::Tool, this reuses the Chat Completions translator and serializes {"type":"function","function":{"name":...}}. The Responses API uses the flattened {"type":"function","name":...} shape, which the removed responses_tool_choice helper previously produced, so named-tool requests on the Responses path are rejected instead of forcing the requested tool.

Useful? React with 👍 / 👎.

Comment on lines +470 to 475
ModelResponse {
message: AssistantMessage {
id: None,
content: vec![ContentBlock::Text(text)],
tool_calls,
content,
tool_calls: Vec::new(),
usage,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Decode Responses function calls

Whenever /responses returns a function_call output item, this now unconditionally emits an empty tool_calls list; the response wire struct also no longer retains the call id, name, or arguments. The previous parser normalized these items, including malformed arguments, so tool-using Responses calls now appear to consumers as empty assistant replies and the requested tool is never executed.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +312 to +313
Message::User(m) => message_text(&m.content),
Message::Assistant(m) => message_text(&m.content),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve images in Responses input

For a Responses-primary request containing ContentBlock::Image, message_text retains only text blocks, so this branch silently discards every image; an image-only user turn is skipped altogether at the subsequent empty-text check. The prior translation emitted input_image content parts, so vision requests now run against missing input and can return plausible but incorrect results rather than an error.

Useful? React with 👍 / 👎.

} else {
"stop".to_string()
}),
finish_reason: Some("stop".to_string()),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve incomplete Responses finish reasons

Every Responses result is now marked stop, including responses whose status is incomplete because max_output_tokens was reached. Consumers therefore cannot distinguish truncated output from a completed answer and may accept or cache partial text/JSON; retain the response status and map incomplete_details.reason as the previous parser did.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +186 to +191
model_info
.iter()
.filter(|(key, _)| key.ends_with(".context_length") || key.as_str() == "context_length")
.filter_map(|(_, value)| value.as_u64())
.filter(|value| *value > 0)
.min()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Select the language-model context length

On multimodal Ollama models, model_info can contain both the language architecture window (for example gemma3.context_length) and a small projector window such as clip.context_length = 77. Taking the minimum advertises 77 tokens as the chat model's input capacity, causing capability checks or context compaction to reject or discard nearly every normal prompt; identify the main architecture's context field instead of minimizing unrelated components.

Useful? React with 👍 / 👎.

Comment on lines 343 to +344
if live.is_empty() {
return Err(Error::Validation(
"Ollama embedding batches must not contain blank inputs".into(),
));
return Ok(vec![Vec::new(); texts.len()]);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reject blank embedding batches instead of returning zero-width vectors

For any nonzero-dimensional OllamaEmbeddingModel, an all-blank batch now returns one empty vector per input, violating EmbeddingModel's fixed-dimension contract. Passing such output through Retriever::index immediately fails in InMemoryVectorStore::add, while direct callers can accidentally persist invalid vectors elsewhere; return a validation error or valid vectors of dimensions() instead.

AGENTS.md reference: AGENTS.md:L53-L58

Useful? React with 👍 / 👎.

Comment on lines +452 to +457
let parsed: ResponsesResponse =
serde_json::from_value(value.clone()).unwrap_or_else(|_| ResponsesResponse {
output: Vec::new(),
output_text: None,
usage: None,
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Propagate invalid Responses payload shapes

If a successful HTTP response has an incompatible schema, such as {"output":"not-an-array"}, deserialization now falls back to an empty response and invoke_responses returns a successful blank assistant message. This hides provider incompatibilities and malformed payloads that previously surfaced as serialization errors, making failures indistinguishable from genuine empty completions; keep parsing fallible and propagate the decode error.

Useful? React with 👍 / 👎.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@senamakel
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add current local-runtime and inference wire contracts - #3

Merged
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest
Aug 30, 2026
Merged

Add current local-runtime and inference wire contracts#3
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest

Conversation

@senamakel

Copy link
Copy Markdown
Member

Summary

  • add runtime-neutral reasoning, cache, continuation, and provider metadata needed by current TinyAgents
  • preserve current local runtime probing, warm-up, model validation, and Ollama embedding behavior
  • fix strict schema fallback, tool id normalization, cache accounting, Responses request/reasoning handling, and JSON Schema union validation
  • keep tracing test-only and use httpdate instead of chrono

Verification

  • cargo test
  • cargo clippy --all-targets -- -D warnings
  • downstream: cargo test --workspace
  • downstream: cargo clippy --workspace --all-targets -- -D warnings

Co-authored-by: Medulla <medulla@tinyhumans.ai>
@chatgpt-codex-connector

chatgpt-codex-connectorBot commented Aug 30, 2026

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

ReviewStatusCommitReview trigger
📝 Code ReviewCompleted2026-08-30T17:59:12.023682Z70412cePR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

Copy link
Copy Markdown

Warning

Review limit reached

  • Run on-demand review

On-demand reviews are free for the next 21 days. After that, they cost $0.25 per reviewed file.

Or wait 8 minutes for your next included review.

View limit details

Limit details: You’ve used the included review currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 4c6cd872-9fcd-4f14-9b4d-cbf8d31df189

📥 Commits

Reviewing files that changed from the base of the PR and between 2c40a5f and 70412ce.

📒 Files selected for processing (20)
  • crates/tinyinference/src/cache/mod.rs
  • crates/tinyinference/src/cache/types.rs
  • crates/tinyinference/src/embeddings/mod.rs
  • crates/tinyinference/src/embeddings/ollama.rs
  • crates/tinyinference/src/message/mod.rs
  • crates/tinyinference/src/message/types.rs
  • crates/tinyinference/src/model/mod.rs
  • crates/tinyinference/src/model/types.rs
  • crates/tinyinference/src/providers/mock.rs
  • crates/tinyinference/src/providers/openai/convert.rs
  • crates/tinyinference/src/providers/openai/local.rs
  • crates/tinyinference/src/providers/openai/local_test.rs
  • crates/tinyinference/src/providers/openai/mod.rs
  • crates/tinyinference/src/providers/openai/responses.rs
  • crates/tinyinference/src/providers/openai/sse.rs
  • crates/tinyinference/src/providers/openai/test.rs
  • crates/tinyinference/src/providers/openai/transport.rs
  • crates/tinyinference/src/providers/openai/types.rs
  • crates/tinyinference/src/providers/types.rs
  • crates/tinyinference/src/tool.rs

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@tinysweeper

Copy link
Copy Markdown

How this change flows

0 changed behaviours across 2 relationships. 2 surrounding behaviours are shown (60 graph nodes walked). 49 further behaviours left out to keep the diagram readable.

flowchart LR
n0["translate_request"]:::impacted
n1["translate_request_with"]:::impacted
n0 -->|calls| n1
n0 -->|tests| n1
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading

Green: changed behaviour. Grey: surrounding behaviour. Arrows name the call, use, implementation, or test relationship. Orange: has findings. Red: has a finding that blocks the merge.

tinysweeper 0.1.0

@tinysweepertinysweeperBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

$0.0000 · 0 in / 0 out

@senamakel
senamakel merged commit cc8aca4 into mainAug 30, 2026
13 checks passed

@chatgpt-codex-connectorchatgpt-codex-connectorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:70412ce592

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +1415 to +1416
if degrade.native_tools {
self.native_tools_on_wire.store(false, Ordering::Relaxed);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Parse prompt-guided calls after degrading native tools

When a model profile advertises tool calling but the server returns a 400 saying tools are unsupported, this latch makes the retry use prompt-guided <tool_call> output. The response paths at invoke and stream, however, only call prompt_tools::apply_to_response when self.profile.tool_calling is false; the profile remains true here, so the retry can succeed while returning the tool-call markup as ordinary assistant text instead of a normalized call. Base response parsing on this latch as well as the static profile.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

} else {
Vec::new()
};
let tool_choice = (!tools.is_empty()).then(|| translate_tool_choice(&request.tool_choice));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Use the Responses shape for named tool choice

For ToolChoice::Tool, this reuses the Chat Completions translator and serializes {"type":"function","function":{"name":...}}. The Responses API uses the flattened {"type":"function","name":...} shape, which the removed responses_tool_choice helper previously produced, so named-tool requests on the Responses path are rejected instead of forcing the requested tool.

Useful? React with 👍 / 👎.

Comment on lines +470 to 475
ModelResponse {
message: AssistantMessage {
id: None,
content: vec![ContentBlock::Text(text)],
tool_calls,
content,
tool_calls: Vec::new(),
usage,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Decode Responses function calls

Whenever /responses returns a function_call output item, this now unconditionally emits an empty tool_calls list; the response wire struct also no longer retains the call id, name, or arguments. The previous parser normalized these items, including malformed arguments, so tool-using Responses calls now appear to consumers as empty assistant replies and the requested tool is never executed.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +312 to +313
Message::User(m) => message_text(&m.content),
Message::Assistant(m) => message_text(&m.content),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve images in Responses input

For a Responses-primary request containing ContentBlock::Image, message_text retains only text blocks, so this branch silently discards every image; an image-only user turn is skipped altogether at the subsequent empty-text check. The prior translation emitted input_image content parts, so vision requests now run against missing input and can return plausible but incorrect results rather than an error.

Useful? React with 👍 / 👎.

} else {
"stop".to_string()
}),
finish_reason: Some("stop".to_string()),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve incomplete Responses finish reasons

Every Responses result is now marked stop, including responses whose status is incomplete because max_output_tokens was reached. Consumers therefore cannot distinguish truncated output from a completed answer and may accept or cache partial text/JSON; retain the response status and map incomplete_details.reason as the previous parser did.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +186 to +191
model_info
.iter()
.filter(|(key, _)| key.ends_with(".context_length") || key.as_str() == "context_length")
.filter_map(|(_, value)| value.as_u64())
.filter(|value| *value > 0)
.min()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Select the language-model context length

On multimodal Ollama models, model_info can contain both the language architecture window (for example gemma3.context_length) and a small projector window such as clip.context_length = 77. Taking the minimum advertises 77 tokens as the chat model's input capacity, causing capability checks or context compaction to reject or discard nearly every normal prompt; identify the main architecture's context field instead of minimizing unrelated components.

Useful? React with 👍 / 👎.

Comment on lines 343 to +344
if live.is_empty() {
return Err(Error::Validation(
"Ollama embedding batches must not contain blank inputs".into(),
));
return Ok(vec![Vec::new(); texts.len()]);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reject blank embedding batches instead of returning zero-width vectors

For any nonzero-dimensional OllamaEmbeddingModel, an all-blank batch now returns one empty vector per input, violating EmbeddingModel's fixed-dimension contract. Passing such output through Retriever::index immediately fails in InMemoryVectorStore::add, while direct callers can accidentally persist invalid vectors elsewhere; return a validation error or valid vectors of dimensions() instead.

AGENTS.md reference: AGENTS.md:L53-L58

Useful? React with 👍 / 👎.

Comment on lines +452 to +457
let parsed: ResponsesResponse =
serde_json::from_value(value.clone()).unwrap_or_else(|_| ResponsesResponse {
output: Vec::new(),
output_text: None,
usage: None,
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Propagate invalid Responses payload shapes

If a successful HTTP response has an incompatible schema, such as {"output":"not-an-array"}, deserialization now falls back to an empty response and invoke_responses returns a successful blank assistant message. This hides provider incompatibilities and malformed payloads that previously surfaced as serialization errors, making failures indistinguishable from genuine empty completions; keep parsing fallible and propagate the decode error.

Useful? React with 👍 / 👎.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@senamakel
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Add current local-runtime and inference wire contracts - #3

Merged
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest
Aug 30, 2026
Merged

Add current local-runtime and inference wire contracts#3
senamakel merged 1 commit into
mainfrom
runtime-metadata-and-latest

Conversation

@senamakel

Copy link
Copy Markdown
Member

Summary

  • add runtime-neutral reasoning, cache, continuation, and provider metadata needed by current TinyAgents
  • preserve current local runtime probing, warm-up, model validation, and Ollama embedding behavior
  • fix strict schema fallback, tool id normalization, cache accounting, Responses request/reasoning handling, and JSON Schema union validation
  • keep tracing test-only and use httpdate instead of chrono

Verification

  • cargo test
  • cargo clippy --all-targets -- -D warnings
  • downstream: cargo test --workspace
  • downstream: cargo clippy --workspace --all-targets -- -D warnings

Co-authored-by: Medulla <medulla@tinyhumans.ai>
@chatgpt-codex-connector

chatgpt-codex-connectorBot commented Aug 30, 2026

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

ReviewStatusCommitReview trigger
📝 Code ReviewCompleted2026-08-30T17:59:12.023682Z70412cePR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

Copy link
Copy Markdown

Warning

Review limit reached

  • Run on-demand review

On-demand reviews are free for the next 21 days. After that, they cost $0.25 per reviewed file.

Or wait 8 minutes for your next included review.

View limit details

Limit details: You’ve used the included review currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 4c6cd872-9fcd-4f14-9b4d-cbf8d31df189

📥 Commits

Reviewing files that changed from the base of the PR and between 2c40a5f and 70412ce.

📒 Files selected for processing (20)
  • crates/tinyinference/src/cache/mod.rs
  • crates/tinyinference/src/cache/types.rs
  • crates/tinyinference/src/embeddings/mod.rs
  • crates/tinyinference/src/embeddings/ollama.rs
  • crates/tinyinference/src/message/mod.rs
  • crates/tinyinference/src/message/types.rs
  • crates/tinyinference/src/model/mod.rs
  • crates/tinyinference/src/model/types.rs
  • crates/tinyinference/src/providers/mock.rs
  • crates/tinyinference/src/providers/openai/convert.rs
  • crates/tinyinference/src/providers/openai/local.rs
  • crates/tinyinference/src/providers/openai/local_test.rs
  • crates/tinyinference/src/providers/openai/mod.rs
  • crates/tinyinference/src/providers/openai/responses.rs
  • crates/tinyinference/src/providers/openai/sse.rs
  • crates/tinyinference/src/providers/openai/test.rs
  • crates/tinyinference/src/providers/openai/transport.rs
  • crates/tinyinference/src/providers/openai/types.rs
  • crates/tinyinference/src/providers/types.rs
  • crates/tinyinference/src/tool.rs

Warning

Your free Security trial is over. An organization admin can activate Security or dismiss this notice.


Comment @coderabbitai help to get the list of available commands.

@tinysweeper

Copy link
Copy Markdown

How this change flows

0 changed behaviours across 2 relationships. 2 surrounding behaviours are shown (60 graph nodes walked). 49 further behaviours left out to keep the diagram readable.

flowchart LR
n0["translate_request"]:::impacted
n1["translate_request_with"]:::impacted
n0 -->|calls| n1
n0 -->|tests| n1
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading

Green: changed behaviour. Grey: surrounding behaviour. Arrows name the call, use, implementation, or test relationship. Orange: has findings. Red: has a finding that blocks the merge.

tinysweeper 0.1.0

@tinysweepertinysweeperBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

$0.0000 · 0 in / 0 out

@senamakel
senamakel merged commit cc8aca4 into mainAug 30, 2026
13 checks passed

@chatgpt-codex-connectorchatgpt-codex-connectorBot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit:70412ce592

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +1415 to +1416
if degrade.native_tools {
self.native_tools_on_wire.store(false, Ordering::Relaxed);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Parse prompt-guided calls after degrading native tools

When a model profile advertises tool calling but the server returns a 400 saying tools are unsupported, this latch makes the retry use prompt-guided <tool_call> output. The response paths at invoke and stream, however, only call prompt_tools::apply_to_response when self.profile.tool_calling is false; the profile remains true here, so the retry can succeed while returning the tool-call markup as ordinary assistant text instead of a normalized call. Base response parsing on this latch as well as the static profile.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

} else {
Vec::new()
};
let tool_choice = (!tools.is_empty()).then(|| translate_tool_choice(&request.tool_choice));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Use the Responses shape for named tool choice

For ToolChoice::Tool, this reuses the Chat Completions translator and serializes {"type":"function","function":{"name":...}}. The Responses API uses the flattened {"type":"function","name":...} shape, which the removed responses_tool_choice helper previously produced, so named-tool requests on the Responses path are rejected instead of forcing the requested tool.

Useful? React with 👍 / 👎.

Comment on lines +470 to 475
ModelResponse {
message: AssistantMessage {
id: None,
content: vec![ContentBlock::Text(text)],
tool_calls,
content,
tool_calls: Vec::new(),
usage,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Decode Responses function calls

Whenever /responses returns a function_call output item, this now unconditionally emits an empty tool_calls list; the response wire struct also no longer retains the call id, name, or arguments. The previous parser normalized these items, including malformed arguments, so tool-using Responses calls now appear to consumers as empty assistant replies and the requested tool is never executed.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +312 to +313
Message::User(m) => message_text(&m.content),
Message::Assistant(m) => message_text(&m.content),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve images in Responses input

For a Responses-primary request containing ContentBlock::Image, message_text retains only text blocks, so this branch silently discards every image; an image-only user turn is skipped altogether at the subsequent empty-text check. The prior translation emitted input_image content parts, so vision requests now run against missing input and can return plausible but incorrect results rather than an error.

Useful? React with 👍 / 👎.

} else {
"stop".to_string()
}),
finish_reason: Some("stop".to_string()),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve incomplete Responses finish reasons

Every Responses result is now marked stop, including responses whose status is incomplete because max_output_tokens was reached. Consumers therefore cannot distinguish truncated output from a completed answer and may accept or cache partial text/JSON; retain the response status and map incomplete_details.reason as the previous parser did.

AGENTS.md reference: AGENTS.md:L43-L45

Useful? React with 👍 / 👎.

Comment on lines +186 to +191
model_info
.iter()
.filter(|(key, _)| key.ends_with(".context_length") || key.as_str() == "context_length")
.filter_map(|(_, value)| value.as_u64())
.filter(|value| *value > 0)
.min()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Select the language-model context length

On multimodal Ollama models, model_info can contain both the language architecture window (for example gemma3.context_length) and a small projector window such as clip.context_length = 77. Taking the minimum advertises 77 tokens as the chat model's input capacity, causing capability checks or context compaction to reject or discard nearly every normal prompt; identify the main architecture's context field instead of minimizing unrelated components.

Useful? React with 👍 / 👎.

Comment on lines 343 to +344
if live.is_empty() {
return Err(Error::Validation(
"Ollama embedding batches must not contain blank inputs".into(),
));
return Ok(vec![Vec::new(); texts.len()]);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reject blank embedding batches instead of returning zero-width vectors

For any nonzero-dimensional OllamaEmbeddingModel, an all-blank batch now returns one empty vector per input, violating EmbeddingModel's fixed-dimension contract. Passing such output through Retriever::index immediately fails in InMemoryVectorStore::add, while direct callers can accidentally persist invalid vectors elsewhere; return a validation error or valid vectors of dimensions() instead.

AGENTS.md reference: AGENTS.md:L53-L58

Useful? React with 👍 / 👎.

Comment on lines +452 to +457
let parsed: ResponsesResponse =
serde_json::from_value(value.clone()).unwrap_or_else(|_| ResponsesResponse {
output: Vec::new(),
output_text: None,
usage: None,
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Propagate invalid Responses payload shapes

If a successful HTTP response has an incompatible schema, such as {"output":"not-an-array"}, deserialization now falls back to an empty response and invoke_responses returns a successful blank assistant message. This hides provider incompatibilities and malformed payloads that previously surfaced as serialization errors, making failures indistinguishable from genuine empty completions; keep parsing fallible and propagate the decode error.

Useful? React with 👍 / 👎.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@senamakel