Skip to content

Ollama tool-calling loop does not complete for query/chat on v0.4.5 and v0.5.0-rc1 #205

Description

@mike-011

Summary

openkb add works with local Ollama models, but openkb query does not reliably complete the tool-calling loop. This remains reproducible on OpenKB v0.5.0-rc1, even with the litellm.drop_params configuration introduced for Ollama compatibility.

The models can emit tool calls, but the CLI either:

  • prints raw tool-call JSON instead of executing the tool and continuing;
  • returns {};
  • returns an incomplete/raw response;
  • or times out without a final answer.

Environment

  • Linux x86_64
  • OpenKB v0.5.0-rc1
  • openai-agents==0.17.3
  • Ollama running as a separate local service
  • Ollama accessed through its OpenAI-compatible/LiteLLM route
  • Knowledge base containing indexed Markdown summaries
  • Query executed through the normal CLI:
openkb --kb-dir <kb-dir> query "What topics are discussed in these messages?"

No cloud API, private endpoint, API key, token, or other credential is required to reproduce the local-backend behavior.

Configuration

The KB configuration was equivalent to:

model: ollama/llama3.1:8b-instruct-q8_0api_base: http://ollama-host:11434language: enpageindex_threshold: 20parallel_tool_calls: falselitellm:
drop_params: true

The hostname above is intentionally a placeholder. The real endpoint is private and is not part of this report.

Reproduction

  1. Create or use a KB with several indexed Markdown documents.
  2. Configure the KB for a local Ollama model as shown above.
  3. Run:
openkb --kb-dir <kb-dir> query "What topics are discussed in these messages?"
  1. Observe that the query does not produce a normal final answer grounded in the KB.

Observed result with llama3.1:8b-instruct-q8_0

The CLI returned raw JSON similar to:

{
"reqId": "<redacted>",
"message": "<a requested wiki path was not found>",
"toolCalls": [
{
"id": "<redacted>",
"type": "function",
"name": "read_file",
"arguments": {"path": "summaries/index.md"}
}
]
}

The tool call was not followed by a normal final answer.

The same model was also tested through the non-streaming Runner.run path. In that path it attempted to call a tool named search_strategy, which was not one of the tools registered by OpenKB:

ModelBehaviorError: Tool search_strategy not found in agent wiki-query

Additional local-model results

Using the same KB, question, and OpenKB v0.5.0-rc1 CLI:

ModelResult
llama3.1:8b-instruct-q8_0raw tool-call JSON; no final answer
qwen3:14b{}
qwen3.5:9bempty output
gemma4:12btimed out
deepseek-r1:14b{}
qwen2.5-coder:14braw/incomplete response claiming that index.md was unavailable
llama3.2:1braw function-call JSON

openkb add had previously completed successfully on the same type of KB, including document summaries. The failure is specific to the query/chat agent path and tool-loop completion, not basic LiteLLM connectivity.

Expected behavior

For an Ollama-compatible model that returns a valid structured tool call, OpenKB should:

  1. parse the tool call;
  2. execute the registered OpenKB tool;
  3. append the tool result to the agent conversation using the correct Chat Completions/LiteLLM format;
  4. continue the agent loop;
  5. return a final natural-language answer.

If a model returns an unsupported or malformed tool call, OpenKB should report a clear actionable error rather than emitting raw JSON or returning {}.

Investigation notes

Request

Could the maintainers please confirm:

  1. which Ollama/LiteLLM model and version combinations are officially supported for query/chat;
  2. whether OpenKB expects OpenAI Responses API semantics or Chat Completions semantics for the LiteLLM provider;
  3. whether the streaming path is expected to support Ollama tool calls;
  4. whether a non-streaming fallback should be used for LiteLLM/Ollama providers;
  5. whether malformed model-emitted tool names should be handled with a clearer validation error.

A small regression test covering one structured Ollama tool call, tool execution, and a final answer would help prevent this from recurring.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
       blocks
      (function() {
      function addCopyButtons() {
      document.querySelectorAll('pre code').forEach(function(codeBlock) {
      if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
      codeBlock.parentElement.setAttribute('data-copy-added', 'true');
      var btn = document.createElement('button');
      btn.textContent = 'Copy';
      btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
      btn.onmouseover = function() { this.style.opacity = '1'; };
      btn.onmouseout = function() { this.style.opacity = '0.7'; };
      btn.onclick = function() {
      navigator.clipboard.writeText(codeBlock.textContent).then(function() {
      btn.textContent = 'Copied!';
      setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
      });
      };
      codeBlock.parentElement.style.position = 'relative';
      codeBlock.parentElement.appendChild(btn);
      });
      }
      addCopyButtons();
      // Re-run on dynamic content
      var observer = new MutationObserver(addCopyButtons);
      observer.observe(document.body, { childList: true, subtree: true });
      })();
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      Ollama tool-calling loop does not complete for query/chat on v0.4.5 and v0.5.0-rc1 · Issue #205 · VectifyAI/OpenKB · GitHub
      Skip to content

      Ollama tool-calling loop does not complete for query/chat on v0.4.5 and v0.5.0-rc1 #205

      Description

      @mike-011

      Summary

      openkb add works with local Ollama models, but openkb query does not reliably complete the tool-calling loop. This remains reproducible on OpenKB v0.5.0-rc1, even with the litellm.drop_params configuration introduced for Ollama compatibility.

      The models can emit tool calls, but the CLI either:

      • prints raw tool-call JSON instead of executing the tool and continuing;
      • returns {};
      • returns an incomplete/raw response;
      • or times out without a final answer.

      Environment

      • Linux x86_64
      • OpenKB v0.5.0-rc1
      • openai-agents==0.17.3
      • Ollama running as a separate local service
      • Ollama accessed through its OpenAI-compatible/LiteLLM route
      • Knowledge base containing indexed Markdown summaries
      • Query executed through the normal CLI:
      openkb --kb-dir <kb-dir> query "What topics are discussed in these messages?"
      

      No cloud API, private endpoint, API key, token, or other credential is required to reproduce the local-backend behavior.

      Configuration

      The KB configuration was equivalent to:

      model: ollama/llama3.1:8b-instruct-q8_0api_base: http://ollama-host:11434language: enpageindex_threshold: 20parallel_tool_calls: falselitellm:
      drop_params: true

      The hostname above is intentionally a placeholder. The real endpoint is private and is not part of this report.

      Reproduction

      1. Create or use a KB with several indexed Markdown documents.
      2. Configure the KB for a local Ollama model as shown above.
      3. Run:
      openkb --kb-dir <kb-dir> query "What topics are discussed in these messages?"
      1. Observe that the query does not produce a normal final answer grounded in the KB.

      Observed result with llama3.1:8b-instruct-q8_0

      The CLI returned raw JSON similar to:

      {
      "reqId": "<redacted>",
      "message": "<a requested wiki path was not found>",
      "toolCalls": [
      {
      "id": "<redacted>",
      "type": "function",
      "name": "read_file",
      "arguments": {"path": "summaries/index.md"}
      }
      ]
      }

      The tool call was not followed by a normal final answer.

      The same model was also tested through the non-streaming Runner.run path. In that path it attempted to call a tool named search_strategy, which was not one of the tools registered by OpenKB:

      ModelBehaviorError: Tool search_strategy not found in agent wiki-query
      

      Additional local-model results

      Using the same KB, question, and OpenKB v0.5.0-rc1 CLI:

      ModelResult
      llama3.1:8b-instruct-q8_0raw tool-call JSON; no final answer
      qwen3:14b{}
      qwen3.5:9bempty output
      gemma4:12btimed out
      deepseek-r1:14b{}
      qwen2.5-coder:14braw/incomplete response claiming that index.md was unavailable
      llama3.2:1braw function-call JSON

      openkb add had previously completed successfully on the same type of KB, including document summaries. The failure is specific to the query/chat agent path and tool-loop completion, not basic LiteLLM connectivity.

      Expected behavior

      For an Ollama-compatible model that returns a valid structured tool call, OpenKB should:

      1. parse the tool call;
      2. execute the registered OpenKB tool;
      3. append the tool result to the agent conversation using the correct Chat Completions/LiteLLM format;
      4. continue the agent loop;
      5. return a final natural-language answer.

      If a model returns an unsupported or malformed tool call, OpenKB should report a clear actionable error rather than emitting raw JSON or returning {}.

      Investigation notes

      Request

      Could the maintainers please confirm:

      1. which Ollama/LiteLLM model and version combinations are officially supported for query/chat;
      2. whether OpenKB expects OpenAI Responses API semantics or Chat Completions semantics for the LiteLLM provider;
      3. whether the streaming path is expected to support Ollama tool calls;
      4. whether a non-streaming fallback should be used for LiteLLM/Ollama providers;
      5. whether malformed model-emitted tool names should be handled with a clearer validation error.

      A small regression test covering one structured Ollama tool call, tool execution, and a final answer would help prevent this from recurring.

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        No labels
        No labels

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Ollama tool-calling loop does not complete for query/chat on v0.4.5 and v0.5.0-rc1 · Issue #205 · VectifyAI/OpenKB · GitHub
          Skip to content

          Ollama tool-calling loop does not complete for query/chat on v0.4.5 and v0.5.0-rc1 #205

          Description

          @mike-011

          Summary

          openkb add works with local Ollama models, but openkb query does not reliably complete the tool-calling loop. This remains reproducible on OpenKB v0.5.0-rc1, even with the litellm.drop_params configuration introduced for Ollama compatibility.

          The models can emit tool calls, but the CLI either:

          • prints raw tool-call JSON instead of executing the tool and continuing;
          • returns {};
          • returns an incomplete/raw response;
          • or times out without a final answer.

          Environment

          • Linux x86_64
          • OpenKB v0.5.0-rc1
          • openai-agents==0.17.3
          • Ollama running as a separate local service
          • Ollama accessed through its OpenAI-compatible/LiteLLM route
          • Knowledge base containing indexed Markdown summaries
          • Query executed through the normal CLI:
          openkb --kb-dir <kb-dir> query "What topics are discussed in these messages?"
          

          No cloud API, private endpoint, API key, token, or other credential is required to reproduce the local-backend behavior.

          Configuration

          The KB configuration was equivalent to:

          model: ollama/llama3.1:8b-instruct-q8_0api_base: http://ollama-host:11434language: enpageindex_threshold: 20parallel_tool_calls: falselitellm:
          drop_params: true

          The hostname above is intentionally a placeholder. The real endpoint is private and is not part of this report.

          Reproduction

          1. Create or use a KB with several indexed Markdown documents.
          2. Configure the KB for a local Ollama model as shown above.
          3. Run:
          openkb --kb-dir <kb-dir> query "What topics are discussed in these messages?"
          1. Observe that the query does not produce a normal final answer grounded in the KB.

          Observed result with llama3.1:8b-instruct-q8_0

          The CLI returned raw JSON similar to:

          {
          "reqId": "<redacted>",
          "message": "<a requested wiki path was not found>",
          "toolCalls": [
          {
          "id": "<redacted>",
          "type": "function",
          "name": "read_file",
          "arguments": {"path": "summaries/index.md"}
          }
          ]
          }

          The tool call was not followed by a normal final answer.

          The same model was also tested through the non-streaming Runner.run path. In that path it attempted to call a tool named search_strategy, which was not one of the tools registered by OpenKB:

          ModelBehaviorError: Tool search_strategy not found in agent wiki-query
          

          Additional local-model results

          Using the same KB, question, and OpenKB v0.5.0-rc1 CLI:

          ModelResult
          llama3.1:8b-instruct-q8_0raw tool-call JSON; no final answer
          qwen3:14b{}
          qwen3.5:9bempty output
          gemma4:12btimed out
          deepseek-r1:14b{}
          qwen2.5-coder:14braw/incomplete response claiming that index.md was unavailable
          llama3.2:1braw function-call JSON

          openkb add had previously completed successfully on the same type of KB, including document summaries. The failure is specific to the query/chat agent path and tool-loop completion, not basic LiteLLM connectivity.

          Expected behavior

          For an Ollama-compatible model that returns a valid structured tool call, OpenKB should:

          1. parse the tool call;
          2. execute the registered OpenKB tool;
          3. append the tool result to the agent conversation using the correct Chat Completions/LiteLLM format;
          4. continue the agent loop;
          5. return a final natural-language answer.

          If a model returns an unsupported or malformed tool call, OpenKB should report a clear actionable error rather than emitting raw JSON or returning {}.

          Investigation notes

          Request

          Could the maintainers please confirm:

          1. which Ollama/LiteLLM model and version combinations are officially supported for query/chat;
          2. whether OpenKB expects OpenAI Responses API semantics or Chat Completions semantics for the LiteLLM provider;
          3. whether the streaming path is expected to support Ollama tool calls;
          4. whether a non-streaming fallback should be used for LiteLLM/Ollama providers;
          5. whether malformed model-emitted tool names should be handled with a clearer validation error.

          A small regression test covering one structured Ollama tool call, tool execution, and a final answer would help prevent this from recurring.

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            No labels
            No labels

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Ollama tool-calling loop does not complete for query/chat on v0.4.5 and v0.5.0-rc1 · Issue #205 · VectifyAI/OpenKB · GitHub
              Skip to content

              Ollama tool-calling loop does not complete for query/chat on v0.4.5 and v0.5.0-rc1 #205

              Description

              @mike-011

              Summary

              openkb add works with local Ollama models, but openkb query does not reliably complete the tool-calling loop. This remains reproducible on OpenKB v0.5.0-rc1, even with the litellm.drop_params configuration introduced for Ollama compatibility.

              The models can emit tool calls, but the CLI either:

              • prints raw tool-call JSON instead of executing the tool and continuing;
              • returns {};
              • returns an incomplete/raw response;
              • or times out without a final answer.

              Environment

              • Linux x86_64
              • OpenKB v0.5.0-rc1
              • openai-agents==0.17.3
              • Ollama running as a separate local service
              • Ollama accessed through its OpenAI-compatible/LiteLLM route
              • Knowledge base containing indexed Markdown summaries
              • Query executed through the normal CLI:
              openkb --kb-dir <kb-dir> query "What topics are discussed in these messages?"
              

              No cloud API, private endpoint, API key, token, or other credential is required to reproduce the local-backend behavior.

              Configuration

              The KB configuration was equivalent to:

              model: ollama/llama3.1:8b-instruct-q8_0api_base: http://ollama-host:11434language: enpageindex_threshold: 20parallel_tool_calls: falselitellm:
              drop_params: true

              The hostname above is intentionally a placeholder. The real endpoint is private and is not part of this report.

              Reproduction

              1. Create or use a KB with several indexed Markdown documents.
              2. Configure the KB for a local Ollama model as shown above.
              3. Run:
              openkb --kb-dir <kb-dir> query "What topics are discussed in these messages?"
              1. Observe that the query does not produce a normal final answer grounded in the KB.

              Observed result with llama3.1:8b-instruct-q8_0

              The CLI returned raw JSON similar to:

              {
              "reqId": "<redacted>",
              "message": "<a requested wiki path was not found>",
              "toolCalls": [
              {
              "id": "<redacted>",
              "type": "function",
              "name": "read_file",
              "arguments": {"path": "summaries/index.md"}
              }
              ]
              }

              The tool call was not followed by a normal final answer.

              The same model was also tested through the non-streaming Runner.run path. In that path it attempted to call a tool named search_strategy, which was not one of the tools registered by OpenKB:

              ModelBehaviorError: Tool search_strategy not found in agent wiki-query
              

              Additional local-model results

              Using the same KB, question, and OpenKB v0.5.0-rc1 CLI:

              ModelResult
              llama3.1:8b-instruct-q8_0raw tool-call JSON; no final answer
              qwen3:14b{}
              qwen3.5:9bempty output
              gemma4:12btimed out
              deepseek-r1:14b{}
              qwen2.5-coder:14braw/incomplete response claiming that index.md was unavailable
              llama3.2:1braw function-call JSON

              openkb add had previously completed successfully on the same type of KB, including document summaries. The failure is specific to the query/chat agent path and tool-loop completion, not basic LiteLLM connectivity.

              Expected behavior

              For an Ollama-compatible model that returns a valid structured tool call, OpenKB should:

              1. parse the tool call;
              2. execute the registered OpenKB tool;
              3. append the tool result to the agent conversation using the correct Chat Completions/LiteLLM format;
              4. continue the agent loop;
              5. return a final natural-language answer.

              If a model returns an unsupported or malformed tool call, OpenKB should report a clear actionable error rather than emitting raw JSON or returning {}.

              Investigation notes

              Request

              Could the maintainers please confirm:

              1. which Ollama/LiteLLM model and version combinations are officially supported for query/chat;
              2. whether OpenKB expects OpenAI Responses API semantics or Chat Completions semantics for the LiteLLM provider;
              3. whether the streaming path is expected to support Ollama tool calls;
              4. whether a non-streaming fallback should be used for LiteLLM/Ollama providers;
              5. whether malformed model-emitted tool names should be handled with a clearer validation error.

              A small regression test covering one structured Ollama tool call, tool execution, and a final answer would help prevent this from recurring.

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                No labels
                No labels

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Ollama tool-calling loop does not complete for query/chat on v0.4.5 and v0.5.0-rc1 · Issue #205 · VectifyAI/OpenKB · GitHub
                  Skip to content

                  Ollama tool-calling loop does not complete for query/chat on v0.4.5 and v0.5.0-rc1 #205

                  Description

                  @mike-011

                  Summary

                  openkb add works with local Ollama models, but openkb query does not reliably complete the tool-calling loop. This remains reproducible on OpenKB v0.5.0-rc1, even with the litellm.drop_params configuration introduced for Ollama compatibility.

                  The models can emit tool calls, but the CLI either:

                  • prints raw tool-call JSON instead of executing the tool and continuing;
                  • returns {};
                  • returns an incomplete/raw response;
                  • or times out without a final answer.

                  Environment

                  • Linux x86_64
                  • OpenKB v0.5.0-rc1
                  • openai-agents==0.17.3
                  • Ollama running as a separate local service
                  • Ollama accessed through its OpenAI-compatible/LiteLLM route
                  • Knowledge base containing indexed Markdown summaries
                  • Query executed through the normal CLI:
                  openkb --kb-dir <kb-dir> query "What topics are discussed in these messages?"
                  

                  No cloud API, private endpoint, API key, token, or other credential is required to reproduce the local-backend behavior.

                  Configuration

                  The KB configuration was equivalent to:

                  model: ollama/llama3.1:8b-instruct-q8_0api_base: http://ollama-host:11434language: enpageindex_threshold: 20parallel_tool_calls: falselitellm:
                  drop_params: true

                  The hostname above is intentionally a placeholder. The real endpoint is private and is not part of this report.

                  Reproduction

                  1. Create or use a KB with several indexed Markdown documents.
                  2. Configure the KB for a local Ollama model as shown above.
                  3. Run:
                  openkb --kb-dir <kb-dir> query "What topics are discussed in these messages?"
                  1. Observe that the query does not produce a normal final answer grounded in the KB.

                  Observed result with llama3.1:8b-instruct-q8_0

                  The CLI returned raw JSON similar to:

                  {
                  "reqId": "<redacted>",
                  "message": "<a requested wiki path was not found>",
                  "toolCalls": [
                  {
                  "id": "<redacted>",
                  "type": "function",
                  "name": "read_file",
                  "arguments": {"path": "summaries/index.md"}
                  }
                  ]
                  }

                  The tool call was not followed by a normal final answer.

                  The same model was also tested through the non-streaming Runner.run path. In that path it attempted to call a tool named search_strategy, which was not one of the tools registered by OpenKB:

                  ModelBehaviorError: Tool search_strategy not found in agent wiki-query
                  

                  Additional local-model results

                  Using the same KB, question, and OpenKB v0.5.0-rc1 CLI:

                  ModelResult
                  llama3.1:8b-instruct-q8_0raw tool-call JSON; no final answer
                  qwen3:14b{}
                  qwen3.5:9bempty output
                  gemma4:12btimed out
                  deepseek-r1:14b{}
                  qwen2.5-coder:14braw/incomplete response claiming that index.md was unavailable
                  llama3.2:1braw function-call JSON

                  openkb add had previously completed successfully on the same type of KB, including document summaries. The failure is specific to the query/chat agent path and tool-loop completion, not basic LiteLLM connectivity.

                  Expected behavior

                  For an Ollama-compatible model that returns a valid structured tool call, OpenKB should:

                  1. parse the tool call;
                  2. execute the registered OpenKB tool;
                  3. append the tool result to the agent conversation using the correct Chat Completions/LiteLLM format;
                  4. continue the agent loop;
                  5. return a final natural-language answer.

                  If a model returns an unsupported or malformed tool call, OpenKB should report a clear actionable error rather than emitting raw JSON or returning {}.

                  Investigation notes

                  Request

                  Could the maintainers please confirm:

                  1. which Ollama/LiteLLM model and version combinations are officially supported for query/chat;
                  2. whether OpenKB expects OpenAI Responses API semantics or Chat Completions semantics for the LiteLLM provider;
                  3. whether the streaming path is expected to support Ollama tool calls;
                  4. whether a non-streaming fallback should be used for LiteLLM/Ollama providers;
                  5. whether malformed model-emitted tool names should be handled with a clearer validation error.

                  A small regression test covering one structured Ollama tool call, tool execution, and a final answer would help prevent this from recurring.

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    No labels
                    No labels

                    Type

                    No type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Ollama tool-calling loop does not complete for query/chat on v0.4.5 and v0.5.0-rc1 · Issue #205 · VectifyAI/OpenKB · GitHub
                      Skip to content

                      Ollama tool-calling loop does not complete for query/chat on v0.4.5 and v0.5.0-rc1 #205

                      Description

                      @mike-011

                      Summary

                      openkb add works with local Ollama models, but openkb query does not reliably complete the tool-calling loop. This remains reproducible on OpenKB v0.5.0-rc1, even with the litellm.drop_params configuration introduced for Ollama compatibility.

                      The models can emit tool calls, but the CLI either:

                      • prints raw tool-call JSON instead of executing the tool and continuing;
                      • returns {};
                      • returns an incomplete/raw response;
                      • or times out without a final answer.

                      Environment

                      • Linux x86_64
                      • OpenKB v0.5.0-rc1
                      • openai-agents==0.17.3
                      • Ollama running as a separate local service
                      • Ollama accessed through its OpenAI-compatible/LiteLLM route
                      • Knowledge base containing indexed Markdown summaries
                      • Query executed through the normal CLI:
                      openkb --kb-dir <kb-dir> query "What topics are discussed in these messages?"
                      

                      No cloud API, private endpoint, API key, token, or other credential is required to reproduce the local-backend behavior.

                      Configuration

                      The KB configuration was equivalent to:

                      model: ollama/llama3.1:8b-instruct-q8_0api_base: http://ollama-host:11434language: enpageindex_threshold: 20parallel_tool_calls: falselitellm:
                      drop_params: true

                      The hostname above is intentionally a placeholder. The real endpoint is private and is not part of this report.

                      Reproduction

                      1. Create or use a KB with several indexed Markdown documents.
                      2. Configure the KB for a local Ollama model as shown above.
                      3. Run:
                      openkb --kb-dir <kb-dir> query "What topics are discussed in these messages?"
                      1. Observe that the query does not produce a normal final answer grounded in the KB.

                      Observed result with llama3.1:8b-instruct-q8_0

                      The CLI returned raw JSON similar to:

                      {
                      "reqId": "<redacted>",
                      "message": "<a requested wiki path was not found>",
                      "toolCalls": [
                      {
                      "id": "<redacted>",
                      "type": "function",
                      "name": "read_file",
                      "arguments": {"path": "summaries/index.md"}
                      }
                      ]
                      }

                      The tool call was not followed by a normal final answer.

                      The same model was also tested through the non-streaming Runner.run path. In that path it attempted to call a tool named search_strategy, which was not one of the tools registered by OpenKB:

                      ModelBehaviorError: Tool search_strategy not found in agent wiki-query
                      

                      Additional local-model results

                      Using the same KB, question, and OpenKB v0.5.0-rc1 CLI:

                      ModelResult
                      llama3.1:8b-instruct-q8_0raw tool-call JSON; no final answer
                      qwen3:14b{}
                      qwen3.5:9bempty output
                      gemma4:12btimed out
                      deepseek-r1:14b{}
                      qwen2.5-coder:14braw/incomplete response claiming that index.md was unavailable
                      llama3.2:1braw function-call JSON

                      openkb add had previously completed successfully on the same type of KB, including document summaries. The failure is specific to the query/chat agent path and tool-loop completion, not basic LiteLLM connectivity.

                      Expected behavior

                      For an Ollama-compatible model that returns a valid structured tool call, OpenKB should:

                      1. parse the tool call;
                      2. execute the registered OpenKB tool;
                      3. append the tool result to the agent conversation using the correct Chat Completions/LiteLLM format;
                      4. continue the agent loop;
                      5. return a final natural-language answer.

                      If a model returns an unsupported or malformed tool call, OpenKB should report a clear actionable error rather than emitting raw JSON or returning {}.

                      Investigation notes

                      Request

                      Could the maintainers please confirm:

                      1. which Ollama/LiteLLM model and version combinations are officially supported for query/chat;
                      2. whether OpenKB expects OpenAI Responses API semantics or Chat Completions semantics for the LiteLLM provider;
                      3. whether the streaming path is expected to support Ollama tool calls;
                      4. whether a non-streaming fallback should be used for LiteLLM/Ollama providers;
                      5. whether malformed model-emitted tool names should be handled with a clearer validation error.

                      A small regression test covering one structured Ollama tool call, tool execution, and a final answer would help prevent this from recurring.

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        No labels
                        No labels

                        Type

                        No type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Ollama tool-calling loop does not complete for query/chat on v0.4.5 and v0.5.0-rc1 · Issue #205 · VectifyAI/OpenKB · GitHub
                          Skip to content

                          Ollama tool-calling loop does not complete for query/chat on v0.4.5 and v0.5.0-rc1 #205

                          Description

                          @mike-011

                          Summary

                          openkb add works with local Ollama models, but openkb query does not reliably complete the tool-calling loop. This remains reproducible on OpenKB v0.5.0-rc1, even with the litellm.drop_params configuration introduced for Ollama compatibility.

                          The models can emit tool calls, but the CLI either:

                          • prints raw tool-call JSON instead of executing the tool and continuing;
                          • returns {};
                          • returns an incomplete/raw response;
                          • or times out without a final answer.

                          Environment

                          • Linux x86_64
                          • OpenKB v0.5.0-rc1
                          • openai-agents==0.17.3
                          • Ollama running as a separate local service
                          • Ollama accessed through its OpenAI-compatible/LiteLLM route
                          • Knowledge base containing indexed Markdown summaries
                          • Query executed through the normal CLI:
                          openkb --kb-dir <kb-dir> query "What topics are discussed in these messages?"
                          

                          No cloud API, private endpoint, API key, token, or other credential is required to reproduce the local-backend behavior.

                          Configuration

                          The KB configuration was equivalent to:

                          model: ollama/llama3.1:8b-instruct-q8_0api_base: http://ollama-host:11434language: enpageindex_threshold: 20parallel_tool_calls: falselitellm:
                          drop_params: true

                          The hostname above is intentionally a placeholder. The real endpoint is private and is not part of this report.

                          Reproduction

                          1. Create or use a KB with several indexed Markdown documents.
                          2. Configure the KB for a local Ollama model as shown above.
                          3. Run:
                          openkb --kb-dir <kb-dir> query "What topics are discussed in these messages?"
                          1. Observe that the query does not produce a normal final answer grounded in the KB.

                          Observed result with llama3.1:8b-instruct-q8_0

                          The CLI returned raw JSON similar to:

                          {
                          "reqId": "<redacted>",
                          "message": "<a requested wiki path was not found>",
                          "toolCalls": [
                          {
                          "id": "<redacted>",
                          "type": "function",
                          "name": "read_file",
                          "arguments": {"path": "summaries/index.md"}
                          }
                          ]
                          }

                          The tool call was not followed by a normal final answer.

                          The same model was also tested through the non-streaming Runner.run path. In that path it attempted to call a tool named search_strategy, which was not one of the tools registered by OpenKB:

                          ModelBehaviorError: Tool search_strategy not found in agent wiki-query
                          

                          Additional local-model results

                          Using the same KB, question, and OpenKB v0.5.0-rc1 CLI:

                          ModelResult
                          llama3.1:8b-instruct-q8_0raw tool-call JSON; no final answer
                          qwen3:14b{}
                          qwen3.5:9bempty output
                          gemma4:12btimed out
                          deepseek-r1:14b{}
                          qwen2.5-coder:14braw/incomplete response claiming that index.md was unavailable
                          llama3.2:1braw function-call JSON

                          openkb add had previously completed successfully on the same type of KB, including document summaries. The failure is specific to the query/chat agent path and tool-loop completion, not basic LiteLLM connectivity.

                          Expected behavior

                          For an Ollama-compatible model that returns a valid structured tool call, OpenKB should:

                          1. parse the tool call;
                          2. execute the registered OpenKB tool;
                          3. append the tool result to the agent conversation using the correct Chat Completions/LiteLLM format;
                          4. continue the agent loop;
                          5. return a final natural-language answer.

                          If a model returns an unsupported or malformed tool call, OpenKB should report a clear actionable error rather than emitting raw JSON or returning {}.

                          Investigation notes

                          Request

                          Could the maintainers please confirm:

                          1. which Ollama/LiteLLM model and version combinations are officially supported for query/chat;
                          2. whether OpenKB expects OpenAI Responses API semantics or Chat Completions semantics for the LiteLLM provider;
                          3. whether the streaming path is expected to support Ollama tool calls;
                          4. whether a non-streaming fallback should be used for LiteLLM/Ollama providers;
                          5. whether malformed model-emitted tool names should be handled with a clearer validation error.

                          A small regression test covering one structured Ollama tool call, tool execution, and a final answer would help prevent this from recurring.

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            No labels
                            No labels

                            Type

                            No type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Ollama tool-calling loop does not complete for query/chat on v0.4.5 and v0.5.0-rc1 · Issue #205 · VectifyAI/OpenKB · GitHub
                              Skip to content

                              Ollama tool-calling loop does not complete for query/chat on v0.4.5 and v0.5.0-rc1 #205

                              Description

                              @mike-011

                              Summary

                              openkb add works with local Ollama models, but openkb query does not reliably complete the tool-calling loop. This remains reproducible on OpenKB v0.5.0-rc1, even with the litellm.drop_params configuration introduced for Ollama compatibility.

                              The models can emit tool calls, but the CLI either:

                              • prints raw tool-call JSON instead of executing the tool and continuing;
                              • returns {};
                              • returns an incomplete/raw response;
                              • or times out without a final answer.

                              Environment

                              • Linux x86_64
                              • OpenKB v0.5.0-rc1
                              • openai-agents==0.17.3
                              • Ollama running as a separate local service
                              • Ollama accessed through its OpenAI-compatible/LiteLLM route
                              • Knowledge base containing indexed Markdown summaries
                              • Query executed through the normal CLI:
                              openkb --kb-dir <kb-dir> query "What topics are discussed in these messages?"
                              

                              No cloud API, private endpoint, API key, token, or other credential is required to reproduce the local-backend behavior.

                              Configuration

                              The KB configuration was equivalent to:

                              model: ollama/llama3.1:8b-instruct-q8_0api_base: http://ollama-host:11434language: enpageindex_threshold: 20parallel_tool_calls: falselitellm:
                              drop_params: true

                              The hostname above is intentionally a placeholder. The real endpoint is private and is not part of this report.

                              Reproduction

                              1. Create or use a KB with several indexed Markdown documents.
                              2. Configure the KB for a local Ollama model as shown above.
                              3. Run:
                              openkb --kb-dir <kb-dir> query "What topics are discussed in these messages?"
                              1. Observe that the query does not produce a normal final answer grounded in the KB.

                              Observed result with llama3.1:8b-instruct-q8_0

                              The CLI returned raw JSON similar to:

                              {
                              "reqId": "<redacted>",
                              "message": "<a requested wiki path was not found>",
                              "toolCalls": [
                              {
                              "id": "<redacted>",
                              "type": "function",
                              "name": "read_file",
                              "arguments": {"path": "summaries/index.md"}
                              }
                              ]
                              }

                              The tool call was not followed by a normal final answer.

                              The same model was also tested through the non-streaming Runner.run path. In that path it attempted to call a tool named search_strategy, which was not one of the tools registered by OpenKB:

                              ModelBehaviorError: Tool search_strategy not found in agent wiki-query
                              

                              Additional local-model results

                              Using the same KB, question, and OpenKB v0.5.0-rc1 CLI:

                              ModelResult
                              llama3.1:8b-instruct-q8_0raw tool-call JSON; no final answer
                              qwen3:14b{}
                              qwen3.5:9bempty output
                              gemma4:12btimed out
                              deepseek-r1:14b{}
                              qwen2.5-coder:14braw/incomplete response claiming that index.md was unavailable
                              llama3.2:1braw function-call JSON

                              openkb add had previously completed successfully on the same type of KB, including document summaries. The failure is specific to the query/chat agent path and tool-loop completion, not basic LiteLLM connectivity.

                              Expected behavior

                              For an Ollama-compatible model that returns a valid structured tool call, OpenKB should:

                              1. parse the tool call;
                              2. execute the registered OpenKB tool;
                              3. append the tool result to the agent conversation using the correct Chat Completions/LiteLLM format;
                              4. continue the agent loop;
                              5. return a final natural-language answer.

                              If a model returns an unsupported or malformed tool call, OpenKB should report a clear actionable error rather than emitting raw JSON or returning {}.

                              Investigation notes

                              Request

                              Could the maintainers please confirm:

                              1. which Ollama/LiteLLM model and version combinations are officially supported for query/chat;
                              2. whether OpenKB expects OpenAI Responses API semantics or Chat Completions semantics for the LiteLLM provider;
                              3. whether the streaming path is expected to support Ollama tool calls;
                              4. whether a non-streaming fallback should be used for LiteLLM/Ollama providers;
                              5. whether malformed model-emitted tool names should be handled with a clearer validation error.

                              A small regression test covering one structured Ollama tool call, tool execution, and a final answer would help prevent this from recurring.

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                No labels
                                No labels

                                Type

                                No type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions