Skip to content

feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls - #40

Closed
pitimon wants to merge 2 commits into
VectifyAI:mainfrom
pitimon:feat/39-response-cache-opt-in
Closed

feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls#40
pitimon wants to merge 2 commits into
VectifyAI:mainfrom
pitimon:feat/39-response-cache-opt-in

Conversation

@pitimon

Copy link
Copy Markdown
Contributor

Summary

  • New per-KB config flags response_cache: bool = false and response_cache_ttl: int | None = None.
  • When enabled and the active model starts with openrouter/, compiler forwards extra_headers containing X-OpenRouter-Cache: true (and optionally X-OpenRouter-Cache-TTL: <seconds>) on every LiteLLM call.
  • OpenRouter returns identical-payload responses in 80–300 ms with zero token billing (docs) — direct win on the compile-retry path (failed compile re-run) and dev iteration.

Why

openkb add only registers a doc's hash after compilation succeeds. When compilation fails partway, the retry runs every LLM call again (summary → plan → N+M concept pages) with identical prompts. Without Response Caching, every retry rebills full token cost. Same situation for repeated openkb lint and developer iteration loops.

Behaviour

  • Default OFF. Response caching stores responses on OpenRouter — incompatible with strict zero-data-retention postures (e.g. KBs holding regulated/classified content). Users opt in deliberately.
  • Headers are emitted only when model.startswith("openrouter/"). Direct Anthropic/OpenAI/etc. requests remain byte-identical to today.
  • TTL is cast to int() before stringifying, so YAML quoting quirks ("600" vs 600) don't reach the header value.
  • Complementary to feat(compiler): add cache_control breakpoints for Anthropic prompt caching #38: prompt caching reduces input cost on the cached prefix per call; response caching skips the model entirely on identical-payload re-runs. They compose.

Scope

  • compile_short_doc, compile_long_doc, _compile_concepts — the only direct LiteLLM callers in the project.
  • Out of scope:query, chat, linter — those use the OpenAI Agents SDK; routing custom headers through the SDK requires a separate, larger change.

Config example

# .openkb/config.yamlmodel: openrouter/anthropic/claude-sonnet-4.5response_cache: trueresponse_cache_ttl: 600# optional, 1..86400 seconds, OpenRouter default 300

Test plan

  • TestResponseCacheHeaders — 7 unit tests covering disabled, missing key, non-OpenRouter model, OpenRouter+enabled, TTL emit/omit, _build_llm_kwargs packaging.
  • TestResponseCacheIntegration — 2 end-to-end tests: flag-on forwards extra_headers on every sync LLM call; flag-off (default) emits no extra_headers (regression guard).
  • Full suite: 244 passed (10 new).
  • Manual smoke against a real OpenRouter key — flip the flag, run the same openkb add twice, observe X-OpenRouter-Cache-Status: HIT on the second run (left for reviewer).

Depends on

#38 — uses the **kwargs symmetry fix on _llm_call_async. Either merge order works after the simple rebase if #38 lands first.

Refs #39

itarun.p added 2 commits May 4, 2026 10:07
…ching
Compiler reuses base context A (system + doc) across N+M+2 LLM calls per
document. Without cache_control markers, every call rebills the full
document content as input tokens.
Adds two breakpoints:
- end of doc_msg: caches (system + doc) for summary, plan, every concept
- end of assistant summary: caches (system + doc + summary) for plan and
every concept generation call
For non-Anthropic providers, the list-of-blocks payload is a valid
OpenAI-compatible shape; LiteLLM normalizes cache_control away.
Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call
(memory observation #82886).
Refs VectifyAI#37
…alls
When response_cache is enabled in the per-KB config and the active model
is routed via openrouter/, compile_short_doc and compile_long_doc forward
extra_headers={"X-OpenRouter-Cache": "true", optional X-OpenRouter-Cache-TTL}
on every LiteLLM call. OpenRouter then returns a cached response in
80-300ms with zero token billing on identical follow-up requests, which
benefits the compile-retry path and repeated lint runs.
Default OFF — opt-in only. Response caching stores responses on
OpenRouter, which conflicts with strict zero-data-retention postures.
Skips header emission when the model is not openrouter/-routed, so direct
Anthropic/OpenAI/etc. calls remain byte-identical to before.
Scope is intentionally limited to compiler.py (the only direct LiteLLM
caller). query/chat/linter route through the OpenAI Agents SDK; threading
custom headers there is a separate change.
Refs VectifyAI#39
Depends on VectifyAI#38
@pitimon

Copy link
Copy Markdown
ContributorAuthor

Smoke test — two-layer verification

L1 — openkb wires the headers through

Enabled in ~/kb-isms-test/.openkb/config.yaml:

response_cache: trueresponse_cache_ttl: 3600

Ran openkb -v add <doc.pdf> (PageIndex/long-doc path, model openrouter/anthropic/claude-sonnet-4.5). Every LLM call's debug log shows the headers attached:

openkb.agent.compiler DEBUG: LLM kwargs [overview]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concepts-plan]: {'max_tokens': 1024, 'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: information-security-risk-assessment]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: risk-management-methodology]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: threat-vulnerability-assessment]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [update: information-asset-classification]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [update: isms-implementation]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}

Every call from compile_long_doc and _compile_concepts carries the cache headers. When the flag is left at its default (response_cache: false), no extra_headers appears in any call — verified by the regression test in this PR (TestResponseCacheIntegration::test_flag_off_no_extra_headers).

L2 — OpenRouter actually serves a cached response

Bypassed openkb to isolate the OpenRouter side. Two identical curl POSTs to /api/v1/chat/completions against anthropic/claude-haiku-4.5 with the same headers and payload, ~1 second apart:

Call 1 (cold)Call 2 (warm)
X-OpenRouter-Cache-StatusMISSHIT
X-OpenRouter-Cache-Age0
prompt_tokens220
completion_tokens240
total_tokens460
cost$0.000142$0
Response body1 stable completionbyte-identical to call 1

Both calls returned the same message.content. Call 2 paid zero tokens and returned in well under one second, exactly matching the OpenRouter response-caching docs.

Combined effect

Both layers compose. The retry path of openkb add (failed compile → re-run with identical prompts) and re-run lint will now hit OpenRouter's cache on every step instead of paying full token cost. Privacy default is preserved — opt-in only, and only OpenRouter-routed models receive the headers.

@KylinMountain

Copy link
Copy Markdown
Collaborator

@pitimon thanks for your contribution! Could you help resolve the conflicts?

@KylinMountain

Copy link
Copy Markdown
Collaborator

As it has been implemented by a litemllm extra header in config, so I will close this PR, thanks for your contributing. you can refer the example/configuration.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@pitimon@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls by pitimon · Pull Request #40 · VectifyAI/OpenKB · GitHub
Skip to content

feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls - #40

Closed
pitimon wants to merge 2 commits into
VectifyAI:mainfrom
pitimon:feat/39-response-cache-opt-in
Closed

feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls#40
pitimon wants to merge 2 commits into
VectifyAI:mainfrom
pitimon:feat/39-response-cache-opt-in

Conversation

@pitimon

Copy link
Copy Markdown
Contributor

Summary

  • New per-KB config flags response_cache: bool = false and response_cache_ttl: int | None = None.
  • When enabled and the active model starts with openrouter/, compiler forwards extra_headers containing X-OpenRouter-Cache: true (and optionally X-OpenRouter-Cache-TTL: <seconds>) on every LiteLLM call.
  • OpenRouter returns identical-payload responses in 80–300 ms with zero token billing (docs) — direct win on the compile-retry path (failed compile re-run) and dev iteration.

Why

openkb add only registers a doc's hash after compilation succeeds. When compilation fails partway, the retry runs every LLM call again (summary → plan → N+M concept pages) with identical prompts. Without Response Caching, every retry rebills full token cost. Same situation for repeated openkb lint and developer iteration loops.

Behaviour

  • Default OFF. Response caching stores responses on OpenRouter — incompatible with strict zero-data-retention postures (e.g. KBs holding regulated/classified content). Users opt in deliberately.
  • Headers are emitted only when model.startswith("openrouter/"). Direct Anthropic/OpenAI/etc. requests remain byte-identical to today.
  • TTL is cast to int() before stringifying, so YAML quoting quirks ("600" vs 600) don't reach the header value.
  • Complementary to feat(compiler): add cache_control breakpoints for Anthropic prompt caching #38: prompt caching reduces input cost on the cached prefix per call; response caching skips the model entirely on identical-payload re-runs. They compose.

Scope

  • compile_short_doc, compile_long_doc, _compile_concepts — the only direct LiteLLM callers in the project.
  • Out of scope:query, chat, linter — those use the OpenAI Agents SDK; routing custom headers through the SDK requires a separate, larger change.

Config example

# .openkb/config.yamlmodel: openrouter/anthropic/claude-sonnet-4.5response_cache: trueresponse_cache_ttl: 600# optional, 1..86400 seconds, OpenRouter default 300

Test plan

  • TestResponseCacheHeaders — 7 unit tests covering disabled, missing key, non-OpenRouter model, OpenRouter+enabled, TTL emit/omit, _build_llm_kwargs packaging.
  • TestResponseCacheIntegration — 2 end-to-end tests: flag-on forwards extra_headers on every sync LLM call; flag-off (default) emits no extra_headers (regression guard).
  • Full suite: 244 passed (10 new).
  • Manual smoke against a real OpenRouter key — flip the flag, run the same openkb add twice, observe X-OpenRouter-Cache-Status: HIT on the second run (left for reviewer).

Depends on

#38 — uses the **kwargs symmetry fix on _llm_call_async. Either merge order works after the simple rebase if #38 lands first.

Refs #39

itarun.p added 2 commits May 4, 2026 10:07
…ching
Compiler reuses base context A (system + doc) across N+M+2 LLM calls per
document. Without cache_control markers, every call rebills the full
document content as input tokens.
Adds two breakpoints:
- end of doc_msg: caches (system + doc) for summary, plan, every concept
- end of assistant summary: caches (system + doc + summary) for plan and
every concept generation call
For non-Anthropic providers, the list-of-blocks payload is a valid
OpenAI-compatible shape; LiteLLM normalizes cache_control away.
Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call
(memory observation #82886).
Refs VectifyAI#37
…alls
When response_cache is enabled in the per-KB config and the active model
is routed via openrouter/, compile_short_doc and compile_long_doc forward
extra_headers={"X-OpenRouter-Cache": "true", optional X-OpenRouter-Cache-TTL}
on every LiteLLM call. OpenRouter then returns a cached response in
80-300ms with zero token billing on identical follow-up requests, which
benefits the compile-retry path and repeated lint runs.
Default OFF — opt-in only. Response caching stores responses on
OpenRouter, which conflicts with strict zero-data-retention postures.
Skips header emission when the model is not openrouter/-routed, so direct
Anthropic/OpenAI/etc. calls remain byte-identical to before.
Scope is intentionally limited to compiler.py (the only direct LiteLLM
caller). query/chat/linter route through the OpenAI Agents SDK; threading
custom headers there is a separate change.
Refs VectifyAI#39
Depends on VectifyAI#38
@pitimon

Copy link
Copy Markdown
ContributorAuthor

Smoke test — two-layer verification

L1 — openkb wires the headers through

Enabled in ~/kb-isms-test/.openkb/config.yaml:

response_cache: trueresponse_cache_ttl: 3600

Ran openkb -v add <doc.pdf> (PageIndex/long-doc path, model openrouter/anthropic/claude-sonnet-4.5). Every LLM call's debug log shows the headers attached:

openkb.agent.compiler DEBUG: LLM kwargs [overview]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concepts-plan]: {'max_tokens': 1024, 'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: information-security-risk-assessment]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: risk-management-methodology]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: threat-vulnerability-assessment]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [update: information-asset-classification]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [update: isms-implementation]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}

Every call from compile_long_doc and _compile_concepts carries the cache headers. When the flag is left at its default (response_cache: false), no extra_headers appears in any call — verified by the regression test in this PR (TestResponseCacheIntegration::test_flag_off_no_extra_headers).

L2 — OpenRouter actually serves a cached response

Bypassed openkb to isolate the OpenRouter side. Two identical curl POSTs to /api/v1/chat/completions against anthropic/claude-haiku-4.5 with the same headers and payload, ~1 second apart:

Call 1 (cold)Call 2 (warm)
X-OpenRouter-Cache-StatusMISSHIT
X-OpenRouter-Cache-Age0
prompt_tokens220
completion_tokens240
total_tokens460
cost$0.000142$0
Response body1 stable completionbyte-identical to call 1

Both calls returned the same message.content. Call 2 paid zero tokens and returned in well under one second, exactly matching the OpenRouter response-caching docs.

Combined effect

Both layers compose. The retry path of openkb add (failed compile → re-run with identical prompts) and re-run lint will now hit OpenRouter's cache on every step instead of paying full token cost. Privacy default is preserved — opt-in only, and only OpenRouter-routed models receive the headers.

@KylinMountain

Copy link
Copy Markdown
Collaborator

@pitimon thanks for your contribution! Could you help resolve the conflicts?

@KylinMountain

Copy link
Copy Markdown
Collaborator

As it has been implemented by a litemllm extra header in config, so I will close this PR, thanks for your contributing. you can refer the example/configuration.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@pitimon@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls by pitimon · Pull Request #40 · VectifyAI/OpenKB · GitHub
Skip to content

feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls - #40

Closed
pitimon wants to merge 2 commits into
VectifyAI:mainfrom
pitimon:feat/39-response-cache-opt-in
Closed

feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls#40
pitimon wants to merge 2 commits into
VectifyAI:mainfrom
pitimon:feat/39-response-cache-opt-in

Conversation

@pitimon

Copy link
Copy Markdown
Contributor

Summary

  • New per-KB config flags response_cache: bool = false and response_cache_ttl: int | None = None.
  • When enabled and the active model starts with openrouter/, compiler forwards extra_headers containing X-OpenRouter-Cache: true (and optionally X-OpenRouter-Cache-TTL: <seconds>) on every LiteLLM call.
  • OpenRouter returns identical-payload responses in 80–300 ms with zero token billing (docs) — direct win on the compile-retry path (failed compile re-run) and dev iteration.

Why

openkb add only registers a doc's hash after compilation succeeds. When compilation fails partway, the retry runs every LLM call again (summary → plan → N+M concept pages) with identical prompts. Without Response Caching, every retry rebills full token cost. Same situation for repeated openkb lint and developer iteration loops.

Behaviour

  • Default OFF. Response caching stores responses on OpenRouter — incompatible with strict zero-data-retention postures (e.g. KBs holding regulated/classified content). Users opt in deliberately.
  • Headers are emitted only when model.startswith("openrouter/"). Direct Anthropic/OpenAI/etc. requests remain byte-identical to today.
  • TTL is cast to int() before stringifying, so YAML quoting quirks ("600" vs 600) don't reach the header value.
  • Complementary to feat(compiler): add cache_control breakpoints for Anthropic prompt caching #38: prompt caching reduces input cost on the cached prefix per call; response caching skips the model entirely on identical-payload re-runs. They compose.

Scope

  • compile_short_doc, compile_long_doc, _compile_concepts — the only direct LiteLLM callers in the project.
  • Out of scope:query, chat, linter — those use the OpenAI Agents SDK; routing custom headers through the SDK requires a separate, larger change.

Config example

# .openkb/config.yamlmodel: openrouter/anthropic/claude-sonnet-4.5response_cache: trueresponse_cache_ttl: 600# optional, 1..86400 seconds, OpenRouter default 300

Test plan

  • TestResponseCacheHeaders — 7 unit tests covering disabled, missing key, non-OpenRouter model, OpenRouter+enabled, TTL emit/omit, _build_llm_kwargs packaging.
  • TestResponseCacheIntegration — 2 end-to-end tests: flag-on forwards extra_headers on every sync LLM call; flag-off (default) emits no extra_headers (regression guard).
  • Full suite: 244 passed (10 new).
  • Manual smoke against a real OpenRouter key — flip the flag, run the same openkb add twice, observe X-OpenRouter-Cache-Status: HIT on the second run (left for reviewer).

Depends on

#38 — uses the **kwargs symmetry fix on _llm_call_async. Either merge order works after the simple rebase if #38 lands first.

Refs #39

itarun.p added 2 commits May 4, 2026 10:07
…ching
Compiler reuses base context A (system + doc) across N+M+2 LLM calls per
document. Without cache_control markers, every call rebills the full
document content as input tokens.
Adds two breakpoints:
- end of doc_msg: caches (system + doc) for summary, plan, every concept
- end of assistant summary: caches (system + doc + summary) for plan and
every concept generation call
For non-Anthropic providers, the list-of-blocks payload is a valid
OpenAI-compatible shape; LiteLLM normalizes cache_control away.
Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call
(memory observation #82886).
Refs VectifyAI#37
…alls
When response_cache is enabled in the per-KB config and the active model
is routed via openrouter/, compile_short_doc and compile_long_doc forward
extra_headers={"X-OpenRouter-Cache": "true", optional X-OpenRouter-Cache-TTL}
on every LiteLLM call. OpenRouter then returns a cached response in
80-300ms with zero token billing on identical follow-up requests, which
benefits the compile-retry path and repeated lint runs.
Default OFF — opt-in only. Response caching stores responses on
OpenRouter, which conflicts with strict zero-data-retention postures.
Skips header emission when the model is not openrouter/-routed, so direct
Anthropic/OpenAI/etc. calls remain byte-identical to before.
Scope is intentionally limited to compiler.py (the only direct LiteLLM
caller). query/chat/linter route through the OpenAI Agents SDK; threading
custom headers there is a separate change.
Refs VectifyAI#39
Depends on VectifyAI#38
@pitimon

Copy link
Copy Markdown
ContributorAuthor

Smoke test — two-layer verification

L1 — openkb wires the headers through

Enabled in ~/kb-isms-test/.openkb/config.yaml:

response_cache: trueresponse_cache_ttl: 3600

Ran openkb -v add <doc.pdf> (PageIndex/long-doc path, model openrouter/anthropic/claude-sonnet-4.5). Every LLM call's debug log shows the headers attached:

openkb.agent.compiler DEBUG: LLM kwargs [overview]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concepts-plan]: {'max_tokens': 1024, 'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: information-security-risk-assessment]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: risk-management-methodology]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: threat-vulnerability-assessment]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [update: information-asset-classification]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [update: isms-implementation]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}

Every call from compile_long_doc and _compile_concepts carries the cache headers. When the flag is left at its default (response_cache: false), no extra_headers appears in any call — verified by the regression test in this PR (TestResponseCacheIntegration::test_flag_off_no_extra_headers).

L2 — OpenRouter actually serves a cached response

Bypassed openkb to isolate the OpenRouter side. Two identical curl POSTs to /api/v1/chat/completions against anthropic/claude-haiku-4.5 with the same headers and payload, ~1 second apart:

Call 1 (cold)Call 2 (warm)
X-OpenRouter-Cache-StatusMISSHIT
X-OpenRouter-Cache-Age0
prompt_tokens220
completion_tokens240
total_tokens460
cost$0.000142$0
Response body1 stable completionbyte-identical to call 1

Both calls returned the same message.content. Call 2 paid zero tokens and returned in well under one second, exactly matching the OpenRouter response-caching docs.

Combined effect

Both layers compose. The retry path of openkb add (failed compile → re-run with identical prompts) and re-run lint will now hit OpenRouter's cache on every step instead of paying full token cost. Privacy default is preserved — opt-in only, and only OpenRouter-routed models receive the headers.

@KylinMountain

Copy link
Copy Markdown
Collaborator

@pitimon thanks for your contribution! Could you help resolve the conflicts?

@KylinMountain

Copy link
Copy Markdown
Collaborator

As it has been implemented by a litemllm extra header in config, so I will close this PR, thanks for your contributing. you can refer the example/configuration.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@pitimon@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls by pitimon · Pull Request #40 · VectifyAI/OpenKB · GitHub
Skip to content

feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls - #40

Closed
pitimon wants to merge 2 commits into
VectifyAI:mainfrom
pitimon:feat/39-response-cache-opt-in
Closed

feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls#40
pitimon wants to merge 2 commits into
VectifyAI:mainfrom
pitimon:feat/39-response-cache-opt-in

Conversation

@pitimon

Copy link
Copy Markdown
Contributor

Summary

  • New per-KB config flags response_cache: bool = false and response_cache_ttl: int | None = None.
  • When enabled and the active model starts with openrouter/, compiler forwards extra_headers containing X-OpenRouter-Cache: true (and optionally X-OpenRouter-Cache-TTL: <seconds>) on every LiteLLM call.
  • OpenRouter returns identical-payload responses in 80–300 ms with zero token billing (docs) — direct win on the compile-retry path (failed compile re-run) and dev iteration.

Why

openkb add only registers a doc's hash after compilation succeeds. When compilation fails partway, the retry runs every LLM call again (summary → plan → N+M concept pages) with identical prompts. Without Response Caching, every retry rebills full token cost. Same situation for repeated openkb lint and developer iteration loops.

Behaviour

  • Default OFF. Response caching stores responses on OpenRouter — incompatible with strict zero-data-retention postures (e.g. KBs holding regulated/classified content). Users opt in deliberately.
  • Headers are emitted only when model.startswith("openrouter/"). Direct Anthropic/OpenAI/etc. requests remain byte-identical to today.
  • TTL is cast to int() before stringifying, so YAML quoting quirks ("600" vs 600) don't reach the header value.
  • Complementary to feat(compiler): add cache_control breakpoints for Anthropic prompt caching #38: prompt caching reduces input cost on the cached prefix per call; response caching skips the model entirely on identical-payload re-runs. They compose.

Scope

  • compile_short_doc, compile_long_doc, _compile_concepts — the only direct LiteLLM callers in the project.
  • Out of scope:query, chat, linter — those use the OpenAI Agents SDK; routing custom headers through the SDK requires a separate, larger change.

Config example

# .openkb/config.yamlmodel: openrouter/anthropic/claude-sonnet-4.5response_cache: trueresponse_cache_ttl: 600# optional, 1..86400 seconds, OpenRouter default 300

Test plan

  • TestResponseCacheHeaders — 7 unit tests covering disabled, missing key, non-OpenRouter model, OpenRouter+enabled, TTL emit/omit, _build_llm_kwargs packaging.
  • TestResponseCacheIntegration — 2 end-to-end tests: flag-on forwards extra_headers on every sync LLM call; flag-off (default) emits no extra_headers (regression guard).
  • Full suite: 244 passed (10 new).
  • Manual smoke against a real OpenRouter key — flip the flag, run the same openkb add twice, observe X-OpenRouter-Cache-Status: HIT on the second run (left for reviewer).

Depends on

#38 — uses the **kwargs symmetry fix on _llm_call_async. Either merge order works after the simple rebase if #38 lands first.

Refs #39

itarun.p added 2 commits May 4, 2026 10:07
…ching
Compiler reuses base context A (system + doc) across N+M+2 LLM calls per
document. Without cache_control markers, every call rebills the full
document content as input tokens.
Adds two breakpoints:
- end of doc_msg: caches (system + doc) for summary, plan, every concept
- end of assistant summary: caches (system + doc + summary) for plan and
every concept generation call
For non-Anthropic providers, the list-of-blocks payload is a valid
OpenAI-compatible shape; LiteLLM normalizes cache_control away.
Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call
(memory observation #82886).
Refs VectifyAI#37
…alls
When response_cache is enabled in the per-KB config and the active model
is routed via openrouter/, compile_short_doc and compile_long_doc forward
extra_headers={"X-OpenRouter-Cache": "true", optional X-OpenRouter-Cache-TTL}
on every LiteLLM call. OpenRouter then returns a cached response in
80-300ms with zero token billing on identical follow-up requests, which
benefits the compile-retry path and repeated lint runs.
Default OFF — opt-in only. Response caching stores responses on
OpenRouter, which conflicts with strict zero-data-retention postures.
Skips header emission when the model is not openrouter/-routed, so direct
Anthropic/OpenAI/etc. calls remain byte-identical to before.
Scope is intentionally limited to compiler.py (the only direct LiteLLM
caller). query/chat/linter route through the OpenAI Agents SDK; threading
custom headers there is a separate change.
Refs VectifyAI#39
Depends on VectifyAI#38
@pitimon

Copy link
Copy Markdown
ContributorAuthor

Smoke test — two-layer verification

L1 — openkb wires the headers through

Enabled in ~/kb-isms-test/.openkb/config.yaml:

response_cache: trueresponse_cache_ttl: 3600

Ran openkb -v add <doc.pdf> (PageIndex/long-doc path, model openrouter/anthropic/claude-sonnet-4.5). Every LLM call's debug log shows the headers attached:

openkb.agent.compiler DEBUG: LLM kwargs [overview]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concepts-plan]: {'max_tokens': 1024, 'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: information-security-risk-assessment]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: risk-management-methodology]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: threat-vulnerability-assessment]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [update: information-asset-classification]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [update: isms-implementation]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}

Every call from compile_long_doc and _compile_concepts carries the cache headers. When the flag is left at its default (response_cache: false), no extra_headers appears in any call — verified by the regression test in this PR (TestResponseCacheIntegration::test_flag_off_no_extra_headers).

L2 — OpenRouter actually serves a cached response

Bypassed openkb to isolate the OpenRouter side. Two identical curl POSTs to /api/v1/chat/completions against anthropic/claude-haiku-4.5 with the same headers and payload, ~1 second apart:

Call 1 (cold)Call 2 (warm)
X-OpenRouter-Cache-StatusMISSHIT
X-OpenRouter-Cache-Age0
prompt_tokens220
completion_tokens240
total_tokens460
cost$0.000142$0
Response body1 stable completionbyte-identical to call 1

Both calls returned the same message.content. Call 2 paid zero tokens and returned in well under one second, exactly matching the OpenRouter response-caching docs.

Combined effect

Both layers compose. The retry path of openkb add (failed compile → re-run with identical prompts) and re-run lint will now hit OpenRouter's cache on every step instead of paying full token cost. Privacy default is preserved — opt-in only, and only OpenRouter-routed models receive the headers.

@KylinMountain

Copy link
Copy Markdown
Collaborator

@pitimon thanks for your contribution! Could you help resolve the conflicts?

@KylinMountain

Copy link
Copy Markdown
Collaborator

As it has been implemented by a litemllm extra header in config, so I will close this PR, thanks for your contributing. you can refer the example/configuration.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@pitimon@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls by pitimon · Pull Request #40 · VectifyAI/OpenKB · GitHub
Skip to content

feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls - #40

Closed
pitimon wants to merge 2 commits into
VectifyAI:mainfrom
pitimon:feat/39-response-cache-opt-in
Closed

feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls#40
pitimon wants to merge 2 commits into
VectifyAI:mainfrom
pitimon:feat/39-response-cache-opt-in

Conversation

@pitimon

Copy link
Copy Markdown
Contributor

Summary

  • New per-KB config flags response_cache: bool = false and response_cache_ttl: int | None = None.
  • When enabled and the active model starts with openrouter/, compiler forwards extra_headers containing X-OpenRouter-Cache: true (and optionally X-OpenRouter-Cache-TTL: <seconds>) on every LiteLLM call.
  • OpenRouter returns identical-payload responses in 80–300 ms with zero token billing (docs) — direct win on the compile-retry path (failed compile re-run) and dev iteration.

Why

openkb add only registers a doc's hash after compilation succeeds. When compilation fails partway, the retry runs every LLM call again (summary → plan → N+M concept pages) with identical prompts. Without Response Caching, every retry rebills full token cost. Same situation for repeated openkb lint and developer iteration loops.

Behaviour

  • Default OFF. Response caching stores responses on OpenRouter — incompatible with strict zero-data-retention postures (e.g. KBs holding regulated/classified content). Users opt in deliberately.
  • Headers are emitted only when model.startswith("openrouter/"). Direct Anthropic/OpenAI/etc. requests remain byte-identical to today.
  • TTL is cast to int() before stringifying, so YAML quoting quirks ("600" vs 600) don't reach the header value.
  • Complementary to feat(compiler): add cache_control breakpoints for Anthropic prompt caching #38: prompt caching reduces input cost on the cached prefix per call; response caching skips the model entirely on identical-payload re-runs. They compose.

Scope

  • compile_short_doc, compile_long_doc, _compile_concepts — the only direct LiteLLM callers in the project.
  • Out of scope:query, chat, linter — those use the OpenAI Agents SDK; routing custom headers through the SDK requires a separate, larger change.

Config example

# .openkb/config.yamlmodel: openrouter/anthropic/claude-sonnet-4.5response_cache: trueresponse_cache_ttl: 600# optional, 1..86400 seconds, OpenRouter default 300

Test plan

  • TestResponseCacheHeaders — 7 unit tests covering disabled, missing key, non-OpenRouter model, OpenRouter+enabled, TTL emit/omit, _build_llm_kwargs packaging.
  • TestResponseCacheIntegration — 2 end-to-end tests: flag-on forwards extra_headers on every sync LLM call; flag-off (default) emits no extra_headers (regression guard).
  • Full suite: 244 passed (10 new).
  • Manual smoke against a real OpenRouter key — flip the flag, run the same openkb add twice, observe X-OpenRouter-Cache-Status: HIT on the second run (left for reviewer).

Depends on

#38 — uses the **kwargs symmetry fix on _llm_call_async. Either merge order works after the simple rebase if #38 lands first.

Refs #39

itarun.p added 2 commits May 4, 2026 10:07
…ching
Compiler reuses base context A (system + doc) across N+M+2 LLM calls per
document. Without cache_control markers, every call rebills the full
document content as input tokens.
Adds two breakpoints:
- end of doc_msg: caches (system + doc) for summary, plan, every concept
- end of assistant summary: caches (system + doc + summary) for plan and
every concept generation call
For non-Anthropic providers, the list-of-blocks payload is a valid
OpenAI-compatible shape; LiteLLM normalizes cache_control away.
Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call
(memory observation #82886).
Refs VectifyAI#37
…alls
When response_cache is enabled in the per-KB config and the active model
is routed via openrouter/, compile_short_doc and compile_long_doc forward
extra_headers={"X-OpenRouter-Cache": "true", optional X-OpenRouter-Cache-TTL}
on every LiteLLM call. OpenRouter then returns a cached response in
80-300ms with zero token billing on identical follow-up requests, which
benefits the compile-retry path and repeated lint runs.
Default OFF — opt-in only. Response caching stores responses on
OpenRouter, which conflicts with strict zero-data-retention postures.
Skips header emission when the model is not openrouter/-routed, so direct
Anthropic/OpenAI/etc. calls remain byte-identical to before.
Scope is intentionally limited to compiler.py (the only direct LiteLLM
caller). query/chat/linter route through the OpenAI Agents SDK; threading
custom headers there is a separate change.
Refs VectifyAI#39
Depends on VectifyAI#38
@pitimon

Copy link
Copy Markdown
ContributorAuthor

Smoke test — two-layer verification

L1 — openkb wires the headers through

Enabled in ~/kb-isms-test/.openkb/config.yaml:

response_cache: trueresponse_cache_ttl: 3600

Ran openkb -v add <doc.pdf> (PageIndex/long-doc path, model openrouter/anthropic/claude-sonnet-4.5). Every LLM call's debug log shows the headers attached:

openkb.agent.compiler DEBUG: LLM kwargs [overview]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concepts-plan]: {'max_tokens': 1024, 'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: information-security-risk-assessment]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: risk-management-methodology]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: threat-vulnerability-assessment]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [update: information-asset-classification]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [update: isms-implementation]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}

Every call from compile_long_doc and _compile_concepts carries the cache headers. When the flag is left at its default (response_cache: false), no extra_headers appears in any call — verified by the regression test in this PR (TestResponseCacheIntegration::test_flag_off_no_extra_headers).

L2 — OpenRouter actually serves a cached response

Bypassed openkb to isolate the OpenRouter side. Two identical curl POSTs to /api/v1/chat/completions against anthropic/claude-haiku-4.5 with the same headers and payload, ~1 second apart:

Call 1 (cold)Call 2 (warm)
X-OpenRouter-Cache-StatusMISSHIT
X-OpenRouter-Cache-Age0
prompt_tokens220
completion_tokens240
total_tokens460
cost$0.000142$0
Response body1 stable completionbyte-identical to call 1

Both calls returned the same message.content. Call 2 paid zero tokens and returned in well under one second, exactly matching the OpenRouter response-caching docs.

Combined effect

Both layers compose. The retry path of openkb add (failed compile → re-run with identical prompts) and re-run lint will now hit OpenRouter's cache on every step instead of paying full token cost. Privacy default is preserved — opt-in only, and only OpenRouter-routed models receive the headers.

@KylinMountain

Copy link
Copy Markdown
Collaborator

@pitimon thanks for your contribution! Could you help resolve the conflicts?

@KylinMountain

Copy link
Copy Markdown
Collaborator

As it has been implemented by a litemllm extra header in config, so I will close this PR, thanks for your contributing. you can refer the example/configuration.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@pitimon@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls by pitimon · Pull Request #40 · VectifyAI/OpenKB · GitHub
Skip to content

feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls - #40

Closed
pitimon wants to merge 2 commits into
VectifyAI:mainfrom
pitimon:feat/39-response-cache-opt-in
Closed

feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls#40
pitimon wants to merge 2 commits into
VectifyAI:mainfrom
pitimon:feat/39-response-cache-opt-in

Conversation

@pitimon

Copy link
Copy Markdown
Contributor

Summary

  • New per-KB config flags response_cache: bool = false and response_cache_ttl: int | None = None.
  • When enabled and the active model starts with openrouter/, compiler forwards extra_headers containing X-OpenRouter-Cache: true (and optionally X-OpenRouter-Cache-TTL: <seconds>) on every LiteLLM call.
  • OpenRouter returns identical-payload responses in 80–300 ms with zero token billing (docs) — direct win on the compile-retry path (failed compile re-run) and dev iteration.

Why

openkb add only registers a doc's hash after compilation succeeds. When compilation fails partway, the retry runs every LLM call again (summary → plan → N+M concept pages) with identical prompts. Without Response Caching, every retry rebills full token cost. Same situation for repeated openkb lint and developer iteration loops.

Behaviour

  • Default OFF. Response caching stores responses on OpenRouter — incompatible with strict zero-data-retention postures (e.g. KBs holding regulated/classified content). Users opt in deliberately.
  • Headers are emitted only when model.startswith("openrouter/"). Direct Anthropic/OpenAI/etc. requests remain byte-identical to today.
  • TTL is cast to int() before stringifying, so YAML quoting quirks ("600" vs 600) don't reach the header value.
  • Complementary to feat(compiler): add cache_control breakpoints for Anthropic prompt caching #38: prompt caching reduces input cost on the cached prefix per call; response caching skips the model entirely on identical-payload re-runs. They compose.

Scope

  • compile_short_doc, compile_long_doc, _compile_concepts — the only direct LiteLLM callers in the project.
  • Out of scope:query, chat, linter — those use the OpenAI Agents SDK; routing custom headers through the SDK requires a separate, larger change.

Config example

# .openkb/config.yamlmodel: openrouter/anthropic/claude-sonnet-4.5response_cache: trueresponse_cache_ttl: 600# optional, 1..86400 seconds, OpenRouter default 300

Test plan

  • TestResponseCacheHeaders — 7 unit tests covering disabled, missing key, non-OpenRouter model, OpenRouter+enabled, TTL emit/omit, _build_llm_kwargs packaging.
  • TestResponseCacheIntegration — 2 end-to-end tests: flag-on forwards extra_headers on every sync LLM call; flag-off (default) emits no extra_headers (regression guard).
  • Full suite: 244 passed (10 new).
  • Manual smoke against a real OpenRouter key — flip the flag, run the same openkb add twice, observe X-OpenRouter-Cache-Status: HIT on the second run (left for reviewer).

Depends on

#38 — uses the **kwargs symmetry fix on _llm_call_async. Either merge order works after the simple rebase if #38 lands first.

Refs #39

itarun.p added 2 commits May 4, 2026 10:07
…ching
Compiler reuses base context A (system + doc) across N+M+2 LLM calls per
document. Without cache_control markers, every call rebills the full
document content as input tokens.
Adds two breakpoints:
- end of doc_msg: caches (system + doc) for summary, plan, every concept
- end of assistant summary: caches (system + doc + summary) for plan and
every concept generation call
For non-Anthropic providers, the list-of-blocks payload is a valid
OpenAI-compatible shape; LiteLLM normalizes cache_control away.
Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call
(memory observation #82886).
Refs VectifyAI#37
…alls
When response_cache is enabled in the per-KB config and the active model
is routed via openrouter/, compile_short_doc and compile_long_doc forward
extra_headers={"X-OpenRouter-Cache": "true", optional X-OpenRouter-Cache-TTL}
on every LiteLLM call. OpenRouter then returns a cached response in
80-300ms with zero token billing on identical follow-up requests, which
benefits the compile-retry path and repeated lint runs.
Default OFF — opt-in only. Response caching stores responses on
OpenRouter, which conflicts with strict zero-data-retention postures.
Skips header emission when the model is not openrouter/-routed, so direct
Anthropic/OpenAI/etc. calls remain byte-identical to before.
Scope is intentionally limited to compiler.py (the only direct LiteLLM
caller). query/chat/linter route through the OpenAI Agents SDK; threading
custom headers there is a separate change.
Refs VectifyAI#39
Depends on VectifyAI#38
@pitimon

Copy link
Copy Markdown
ContributorAuthor

Smoke test — two-layer verification

L1 — openkb wires the headers through

Enabled in ~/kb-isms-test/.openkb/config.yaml:

response_cache: trueresponse_cache_ttl: 3600

Ran openkb -v add <doc.pdf> (PageIndex/long-doc path, model openrouter/anthropic/claude-sonnet-4.5). Every LLM call's debug log shows the headers attached:

openkb.agent.compiler DEBUG: LLM kwargs [overview]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concepts-plan]: {'max_tokens': 1024, 'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: information-security-risk-assessment]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: risk-management-methodology]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: threat-vulnerability-assessment]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [update: information-asset-classification]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [update: isms-implementation]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}

Every call from compile_long_doc and _compile_concepts carries the cache headers. When the flag is left at its default (response_cache: false), no extra_headers appears in any call — verified by the regression test in this PR (TestResponseCacheIntegration::test_flag_off_no_extra_headers).

L2 — OpenRouter actually serves a cached response

Bypassed openkb to isolate the OpenRouter side. Two identical curl POSTs to /api/v1/chat/completions against anthropic/claude-haiku-4.5 with the same headers and payload, ~1 second apart:

Call 1 (cold)Call 2 (warm)
X-OpenRouter-Cache-StatusMISSHIT
X-OpenRouter-Cache-Age0
prompt_tokens220
completion_tokens240
total_tokens460
cost$0.000142$0
Response body1 stable completionbyte-identical to call 1

Both calls returned the same message.content. Call 2 paid zero tokens and returned in well under one second, exactly matching the OpenRouter response-caching docs.

Combined effect

Both layers compose. The retry path of openkb add (failed compile → re-run with identical prompts) and re-run lint will now hit OpenRouter's cache on every step instead of paying full token cost. Privacy default is preserved — opt-in only, and only OpenRouter-routed models receive the headers.

@KylinMountain

Copy link
Copy Markdown
Collaborator

@pitimon thanks for your contribution! Could you help resolve the conflicts?

@KylinMountain

Copy link
Copy Markdown
Collaborator

As it has been implemented by a litemllm extra header in config, so I will close this PR, thanks for your contributing. you can refer the example/configuration.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@pitimon@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls by pitimon · Pull Request #40 · VectifyAI/OpenKB · GitHub
Skip to content

feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls - #40

Closed
pitimon wants to merge 2 commits into
VectifyAI:mainfrom
pitimon:feat/39-response-cache-opt-in
Closed

feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls#40
pitimon wants to merge 2 commits into
VectifyAI:mainfrom
pitimon:feat/39-response-cache-opt-in

Conversation

@pitimon

Copy link
Copy Markdown
Contributor

Summary

  • New per-KB config flags response_cache: bool = false and response_cache_ttl: int | None = None.
  • When enabled and the active model starts with openrouter/, compiler forwards extra_headers containing X-OpenRouter-Cache: true (and optionally X-OpenRouter-Cache-TTL: <seconds>) on every LiteLLM call.
  • OpenRouter returns identical-payload responses in 80–300 ms with zero token billing (docs) — direct win on the compile-retry path (failed compile re-run) and dev iteration.

Why

openkb add only registers a doc's hash after compilation succeeds. When compilation fails partway, the retry runs every LLM call again (summary → plan → N+M concept pages) with identical prompts. Without Response Caching, every retry rebills full token cost. Same situation for repeated openkb lint and developer iteration loops.

Behaviour

  • Default OFF. Response caching stores responses on OpenRouter — incompatible with strict zero-data-retention postures (e.g. KBs holding regulated/classified content). Users opt in deliberately.
  • Headers are emitted only when model.startswith("openrouter/"). Direct Anthropic/OpenAI/etc. requests remain byte-identical to today.
  • TTL is cast to int() before stringifying, so YAML quoting quirks ("600" vs 600) don't reach the header value.
  • Complementary to feat(compiler): add cache_control breakpoints for Anthropic prompt caching #38: prompt caching reduces input cost on the cached prefix per call; response caching skips the model entirely on identical-payload re-runs. They compose.

Scope

  • compile_short_doc, compile_long_doc, _compile_concepts — the only direct LiteLLM callers in the project.
  • Out of scope:query, chat, linter — those use the OpenAI Agents SDK; routing custom headers through the SDK requires a separate, larger change.

Config example

# .openkb/config.yamlmodel: openrouter/anthropic/claude-sonnet-4.5response_cache: trueresponse_cache_ttl: 600# optional, 1..86400 seconds, OpenRouter default 300

Test plan

  • TestResponseCacheHeaders — 7 unit tests covering disabled, missing key, non-OpenRouter model, OpenRouter+enabled, TTL emit/omit, _build_llm_kwargs packaging.
  • TestResponseCacheIntegration — 2 end-to-end tests: flag-on forwards extra_headers on every sync LLM call; flag-off (default) emits no extra_headers (regression guard).
  • Full suite: 244 passed (10 new).
  • Manual smoke against a real OpenRouter key — flip the flag, run the same openkb add twice, observe X-OpenRouter-Cache-Status: HIT on the second run (left for reviewer).

Depends on

#38 — uses the **kwargs symmetry fix on _llm_call_async. Either merge order works after the simple rebase if #38 lands first.

Refs #39

itarun.p added 2 commits May 4, 2026 10:07
…ching
Compiler reuses base context A (system + doc) across N+M+2 LLM calls per
document. Without cache_control markers, every call rebills the full
document content as input tokens.
Adds two breakpoints:
- end of doc_msg: caches (system + doc) for summary, plan, every concept
- end of assistant summary: caches (system + doc + summary) for plan and
every concept generation call
For non-Anthropic providers, the list-of-blocks payload is a valid
OpenAI-compatible shape; LiteLLM normalizes cache_control away.
Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call
(memory observation #82886).
Refs VectifyAI#37
…alls
When response_cache is enabled in the per-KB config and the active model
is routed via openrouter/, compile_short_doc and compile_long_doc forward
extra_headers={"X-OpenRouter-Cache": "true", optional X-OpenRouter-Cache-TTL}
on every LiteLLM call. OpenRouter then returns a cached response in
80-300ms with zero token billing on identical follow-up requests, which
benefits the compile-retry path and repeated lint runs.
Default OFF — opt-in only. Response caching stores responses on
OpenRouter, which conflicts with strict zero-data-retention postures.
Skips header emission when the model is not openrouter/-routed, so direct
Anthropic/OpenAI/etc. calls remain byte-identical to before.
Scope is intentionally limited to compiler.py (the only direct LiteLLM
caller). query/chat/linter route through the OpenAI Agents SDK; threading
custom headers there is a separate change.
Refs VectifyAI#39
Depends on VectifyAI#38
@pitimon

Copy link
Copy Markdown
ContributorAuthor

Smoke test — two-layer verification

L1 — openkb wires the headers through

Enabled in ~/kb-isms-test/.openkb/config.yaml:

response_cache: trueresponse_cache_ttl: 3600

Ran openkb -v add <doc.pdf> (PageIndex/long-doc path, model openrouter/anthropic/claude-sonnet-4.5). Every LLM call's debug log shows the headers attached:

openkb.agent.compiler DEBUG: LLM kwargs [overview]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concepts-plan]: {'max_tokens': 1024, 'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: information-security-risk-assessment]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: risk-management-methodology]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: threat-vulnerability-assessment]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [update: information-asset-classification]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [update: isms-implementation]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}

Every call from compile_long_doc and _compile_concepts carries the cache headers. When the flag is left at its default (response_cache: false), no extra_headers appears in any call — verified by the regression test in this PR (TestResponseCacheIntegration::test_flag_off_no_extra_headers).

L2 — OpenRouter actually serves a cached response

Bypassed openkb to isolate the OpenRouter side. Two identical curl POSTs to /api/v1/chat/completions against anthropic/claude-haiku-4.5 with the same headers and payload, ~1 second apart:

Call 1 (cold)Call 2 (warm)
X-OpenRouter-Cache-StatusMISSHIT
X-OpenRouter-Cache-Age0
prompt_tokens220
completion_tokens240
total_tokens460
cost$0.000142$0
Response body1 stable completionbyte-identical to call 1

Both calls returned the same message.content. Call 2 paid zero tokens and returned in well under one second, exactly matching the OpenRouter response-caching docs.

Combined effect

Both layers compose. The retry path of openkb add (failed compile → re-run with identical prompts) and re-run lint will now hit OpenRouter's cache on every step instead of paying full token cost. Privacy default is preserved — opt-in only, and only OpenRouter-routed models receive the headers.

@KylinMountain

Copy link
Copy Markdown
Collaborator

@pitimon thanks for your contribution! Could you help resolve the conflicts?

@KylinMountain

Copy link
Copy Markdown
Collaborator

As it has been implemented by a litemllm extra header in config, so I will close this PR, thanks for your contributing. you can refer the example/configuration.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@pitimon@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls by pitimon · Pull Request #40 · VectifyAI/OpenKB · GitHub
Skip to content

feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls - #40

Closed
pitimon wants to merge 2 commits into
VectifyAI:mainfrom
pitimon:feat/39-response-cache-opt-in
Closed

feat(compiler): opt-in OpenRouter Response Caching for compiler LLM calls#40
pitimon wants to merge 2 commits into
VectifyAI:mainfrom
pitimon:feat/39-response-cache-opt-in

Conversation

@pitimon

Copy link
Copy Markdown
Contributor

Summary

  • New per-KB config flags response_cache: bool = false and response_cache_ttl: int | None = None.
  • When enabled and the active model starts with openrouter/, compiler forwards extra_headers containing X-OpenRouter-Cache: true (and optionally X-OpenRouter-Cache-TTL: <seconds>) on every LiteLLM call.
  • OpenRouter returns identical-payload responses in 80–300 ms with zero token billing (docs) — direct win on the compile-retry path (failed compile re-run) and dev iteration.

Why

openkb add only registers a doc's hash after compilation succeeds. When compilation fails partway, the retry runs every LLM call again (summary → plan → N+M concept pages) with identical prompts. Without Response Caching, every retry rebills full token cost. Same situation for repeated openkb lint and developer iteration loops.

Behaviour

  • Default OFF. Response caching stores responses on OpenRouter — incompatible with strict zero-data-retention postures (e.g. KBs holding regulated/classified content). Users opt in deliberately.
  • Headers are emitted only when model.startswith("openrouter/"). Direct Anthropic/OpenAI/etc. requests remain byte-identical to today.
  • TTL is cast to int() before stringifying, so YAML quoting quirks ("600" vs 600) don't reach the header value.
  • Complementary to feat(compiler): add cache_control breakpoints for Anthropic prompt caching #38: prompt caching reduces input cost on the cached prefix per call; response caching skips the model entirely on identical-payload re-runs. They compose.

Scope

  • compile_short_doc, compile_long_doc, _compile_concepts — the only direct LiteLLM callers in the project.
  • Out of scope:query, chat, linter — those use the OpenAI Agents SDK; routing custom headers through the SDK requires a separate, larger change.

Config example

# .openkb/config.yamlmodel: openrouter/anthropic/claude-sonnet-4.5response_cache: trueresponse_cache_ttl: 600# optional, 1..86400 seconds, OpenRouter default 300

Test plan

  • TestResponseCacheHeaders — 7 unit tests covering disabled, missing key, non-OpenRouter model, OpenRouter+enabled, TTL emit/omit, _build_llm_kwargs packaging.
  • TestResponseCacheIntegration — 2 end-to-end tests: flag-on forwards extra_headers on every sync LLM call; flag-off (default) emits no extra_headers (regression guard).
  • Full suite: 244 passed (10 new).
  • Manual smoke against a real OpenRouter key — flip the flag, run the same openkb add twice, observe X-OpenRouter-Cache-Status: HIT on the second run (left for reviewer).

Depends on

#38 — uses the **kwargs symmetry fix on _llm_call_async. Either merge order works after the simple rebase if #38 lands first.

Refs #39

itarun.p added 2 commits May 4, 2026 10:07
…ching
Compiler reuses base context A (system + doc) across N+M+2 LLM calls per
document. Without cache_control markers, every call rebills the full
document content as input tokens.
Adds two breakpoints:
- end of doc_msg: caches (system + doc) for summary, plan, every concept
- end of assistant summary: caches (system + doc + summary) for plan and
every concept generation call
For non-Anthropic providers, the list-of-blocks payload is a valid
OpenAI-compatible shape; LiteLLM normalizes cache_control away.
Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call
(memory observation #82886).
Refs VectifyAI#37
…alls
When response_cache is enabled in the per-KB config and the active model
is routed via openrouter/, compile_short_doc and compile_long_doc forward
extra_headers={"X-OpenRouter-Cache": "true", optional X-OpenRouter-Cache-TTL}
on every LiteLLM call. OpenRouter then returns a cached response in
80-300ms with zero token billing on identical follow-up requests, which
benefits the compile-retry path and repeated lint runs.
Default OFF — opt-in only. Response caching stores responses on
OpenRouter, which conflicts with strict zero-data-retention postures.
Skips header emission when the model is not openrouter/-routed, so direct
Anthropic/OpenAI/etc. calls remain byte-identical to before.
Scope is intentionally limited to compiler.py (the only direct LiteLLM
caller). query/chat/linter route through the OpenAI Agents SDK; threading
custom headers there is a separate change.
Refs VectifyAI#39
Depends on VectifyAI#38
@pitimon

Copy link
Copy Markdown
ContributorAuthor

Smoke test — two-layer verification

L1 — openkb wires the headers through

Enabled in ~/kb-isms-test/.openkb/config.yaml:

response_cache: trueresponse_cache_ttl: 3600

Ran openkb -v add <doc.pdf> (PageIndex/long-doc path, model openrouter/anthropic/claude-sonnet-4.5). Every LLM call's debug log shows the headers attached:

openkb.agent.compiler DEBUG: LLM kwargs [overview]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concepts-plan]: {'max_tokens': 1024, 'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: information-security-risk-assessment]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: risk-management-methodology]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [concept: threat-vulnerability-assessment]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [update: information-asset-classification]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}
openkb.agent.compiler DEBUG: LLM kwargs [update: isms-implementation]: {'extra_headers': {'X-OpenRouter-Cache': 'true', 'X-OpenRouter-Cache-TTL': '3600'}}

Every call from compile_long_doc and _compile_concepts carries the cache headers. When the flag is left at its default (response_cache: false), no extra_headers appears in any call — verified by the regression test in this PR (TestResponseCacheIntegration::test_flag_off_no_extra_headers).

L2 — OpenRouter actually serves a cached response

Bypassed openkb to isolate the OpenRouter side. Two identical curl POSTs to /api/v1/chat/completions against anthropic/claude-haiku-4.5 with the same headers and payload, ~1 second apart:

Call 1 (cold)Call 2 (warm)
X-OpenRouter-Cache-StatusMISSHIT
X-OpenRouter-Cache-Age0
prompt_tokens220
completion_tokens240
total_tokens460
cost$0.000142$0
Response body1 stable completionbyte-identical to call 1

Both calls returned the same message.content. Call 2 paid zero tokens and returned in well under one second, exactly matching the OpenRouter response-caching docs.

Combined effect

Both layers compose. The retry path of openkb add (failed compile → re-run with identical prompts) and re-run lint will now hit OpenRouter's cache on every step instead of paying full token cost. Privacy default is preserved — opt-in only, and only OpenRouter-routed models receive the headers.

@KylinMountain

Copy link
Copy Markdown
Collaborator

@pitimon thanks for your contribution! Could you help resolve the conflicts?

@KylinMountain

Copy link
Copy Markdown
Collaborator

As it has been implemented by a litemllm extra header in config, so I will close this PR, thanks for your contributing. you can refer the example/configuration.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@pitimon@KylinMountain