Skip to content

feat(compiler): add cache_control breakpoints for Anthropic prompt caching - #38

Merged
KylinMountain merged 1 commit into
VectifyAI:mainfrom
pitimon:feat/37-cache-control-prompt-caching
May 11, 2026
Merged

feat(compiler): add cache_control breakpoints for Anthropic prompt caching#38
KylinMountain merged 1 commit into
VectifyAI:mainfrom
pitimon:feat/37-cache-control-prompt-caching

Conversation

@pitimon

Copy link
Copy Markdown
Contributor

Summary

  • Adds two cache_control: {"type": "ephemeral"} breakpoints to the compiler so Anthropic prompt caching can hit on every reuse of the base context.
  • Breakpoint 1 — end of doc_msg: caches (system + doc) across summary, concepts plan, and every concept page call (1 + 1 + N + M calls per document).
  • Breakpoint 2 — end of assistant summary: caches (system + doc + summary) across the plan call and every concept generation call.
  • Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call.

Why

Compiler architecture is explicitly designed around base-context reuse (per CLAUDE.md: "Designed around prompt-cache reuse: a single base context A reused across summary → concept-plan → concept-page calls") but no cache markers were emitted, so every call rebilled the full document content. For Anthropic Sonnet 4.5 via OpenRouter or direct, this typically reduces input cost by ~90% on the cached prefix and lowers TTFT after the first call.

Compatibility

  • Anthropic / OpenRouter→Anthropic: cache_control honored.
  • OpenAI: list-of-blocks content is a valid OpenAI-compatible shape; cache_control field is ignored harmlessly.
  • Other providers via LiteLLM: LiteLLM normalizes / strips unknown fields.

_fmt_messages was extended to handle both string and list-of-blocks content shapes for debug output.

Test plan

  • Existing 232 tests pass unchanged (mocks accept *args, **kwargs).
  • New TestCacheControl class — 2 tests:
    • test_short_doc_marks_doc_and_summary captures messages from sync + async LiteLLM calls and asserts cache_control markers are present on doc_msg (all 3 sync calls, all async calls) and on the assistant summary (plan + concept calls).
    • test_long_doc_marks_doc_message asserts the breakpoint on doc_msg for the long-doc path.
  • Full suite: 234 passed (1.05s).
  • Manual smoke against a real Anthropic API key: observe cached_tokens in prompt_tokens_details on calls 2..N (left for reviewer).

Out of scope

  • OpenRouter Response Caching (X-OpenRouter-Cache: true header) — different mechanism, evaluated separately.
  • Splitting compiler.py (~847 lines, pre-existing >800 condition deepened by ~35 lines) — recommend follow-up extracting compiler/messages.py.

Refs #37

…ching
Compiler reuses base context A (system + doc) across N+M+2 LLM calls per
document. Without cache_control markers, every call rebills the full
document content as input tokens.
Adds two breakpoints:
- end of doc_msg: caches (system + doc) for summary, plan, every concept
- end of assistant summary: caches (system + doc + summary) for plan and
every concept generation call
For non-Anthropic providers, the list-of-blocks payload is a valid
OpenAI-compatible shape; LiteLLM normalizes cache_control away.
Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call
(memory observation #82886).
Refs VectifyAI#37
@pitimon

Copy link
Copy Markdown
ContributorAuthor

Smoke test against live OpenRouter → Anthropic Sonnet 4.5

Ingested a fresh OCR'd 21-page Thai procedure document into a short-doc-pipeline KB. Token usage from openkb -v add (verbose mode prints per-call usage):

Stepinput tokenscached tokenshit rate
summary (call 1)6,9900miss (cache being written)
concepts-plan9,5026,98773.5% — breakpoint 1 (system + doc)
concept: preventive-maintenance8,7458,54297.7% — breakpoint 2 (+ summary)
concept: network-management8,7418,54297.7%
concept: network-security-zones8,7488,54297.7%
update: change-management12,6038,54267.8% (existing concept body extends prompt)
update: configuration-management13,7368,54262.2%

Both breakpoints work as designed:

  1. End of doc_msg — every call from concepts-plan onward shows cached >= 6,987, confirming the (system + doc) prefix is being reused.
  2. End of assistant summary — every concept-generation call shows cached = 8,542 (≈ 1,555 tokens beyond breakpoint 1), confirming the (system + doc + summary) prefix is reused across the N+M concept-page calls.

Total tokens cached on this single ingest: ~50,200 across 6 reuse calls. For a typical document with 5+ concept calls, the cached prefix is paid once and reused at the discounted rate for all subsequent calls — matching the ~90% savings expected for Anthropic prompt caching.

Routed via openrouter/anthropic/claude-sonnet-4.5; LiteLLM passes the cache_control markers through to Anthropic transparently.

@KylinMountainKylinMountain left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@pitimon@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
feat(compiler): add cache_control breakpoints for Anthropic prompt caching by pitimon · Pull Request #38 · VectifyAI/OpenKB · GitHub
Skip to content

feat(compiler): add cache_control breakpoints for Anthropic prompt caching - #38

Merged
KylinMountain merged 1 commit into
VectifyAI:mainfrom
pitimon:feat/37-cache-control-prompt-caching
May 11, 2026
Merged

feat(compiler): add cache_control breakpoints for Anthropic prompt caching#38
KylinMountain merged 1 commit into
VectifyAI:mainfrom
pitimon:feat/37-cache-control-prompt-caching

Conversation

@pitimon

Copy link
Copy Markdown
Contributor

Summary

  • Adds two cache_control: {"type": "ephemeral"} breakpoints to the compiler so Anthropic prompt caching can hit on every reuse of the base context.
  • Breakpoint 1 — end of doc_msg: caches (system + doc) across summary, concepts plan, and every concept page call (1 + 1 + N + M calls per document).
  • Breakpoint 2 — end of assistant summary: caches (system + doc + summary) across the plan call and every concept generation call.
  • Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call.

Why

Compiler architecture is explicitly designed around base-context reuse (per CLAUDE.md: "Designed around prompt-cache reuse: a single base context A reused across summary → concept-plan → concept-page calls") but no cache markers were emitted, so every call rebilled the full document content. For Anthropic Sonnet 4.5 via OpenRouter or direct, this typically reduces input cost by ~90% on the cached prefix and lowers TTFT after the first call.

Compatibility

  • Anthropic / OpenRouter→Anthropic: cache_control honored.
  • OpenAI: list-of-blocks content is a valid OpenAI-compatible shape; cache_control field is ignored harmlessly.
  • Other providers via LiteLLM: LiteLLM normalizes / strips unknown fields.

_fmt_messages was extended to handle both string and list-of-blocks content shapes for debug output.

Test plan

  • Existing 232 tests pass unchanged (mocks accept *args, **kwargs).
  • New TestCacheControl class — 2 tests:
    • test_short_doc_marks_doc_and_summary captures messages from sync + async LiteLLM calls and asserts cache_control markers are present on doc_msg (all 3 sync calls, all async calls) and on the assistant summary (plan + concept calls).
    • test_long_doc_marks_doc_message asserts the breakpoint on doc_msg for the long-doc path.
  • Full suite: 234 passed (1.05s).
  • Manual smoke against a real Anthropic API key: observe cached_tokens in prompt_tokens_details on calls 2..N (left for reviewer).

Out of scope

  • OpenRouter Response Caching (X-OpenRouter-Cache: true header) — different mechanism, evaluated separately.
  • Splitting compiler.py (~847 lines, pre-existing >800 condition deepened by ~35 lines) — recommend follow-up extracting compiler/messages.py.

Refs #37

…ching
Compiler reuses base context A (system + doc) across N+M+2 LLM calls per
document. Without cache_control markers, every call rebills the full
document content as input tokens.
Adds two breakpoints:
- end of doc_msg: caches (system + doc) for summary, plan, every concept
- end of assistant summary: caches (system + doc + summary) for plan and
every concept generation call
For non-Anthropic providers, the list-of-blocks payload is a valid
OpenAI-compatible shape; LiteLLM normalizes cache_control away.
Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call
(memory observation #82886).
Refs VectifyAI#37
@pitimon

Copy link
Copy Markdown
ContributorAuthor

Smoke test against live OpenRouter → Anthropic Sonnet 4.5

Ingested a fresh OCR'd 21-page Thai procedure document into a short-doc-pipeline KB. Token usage from openkb -v add (verbose mode prints per-call usage):

Stepinput tokenscached tokenshit rate
summary (call 1)6,9900miss (cache being written)
concepts-plan9,5026,98773.5% — breakpoint 1 (system + doc)
concept: preventive-maintenance8,7458,54297.7% — breakpoint 2 (+ summary)
concept: network-management8,7418,54297.7%
concept: network-security-zones8,7488,54297.7%
update: change-management12,6038,54267.8% (existing concept body extends prompt)
update: configuration-management13,7368,54262.2%

Both breakpoints work as designed:

  1. End of doc_msg — every call from concepts-plan onward shows cached >= 6,987, confirming the (system + doc) prefix is being reused.
  2. End of assistant summary — every concept-generation call shows cached = 8,542 (≈ 1,555 tokens beyond breakpoint 1), confirming the (system + doc + summary) prefix is reused across the N+M concept-page calls.

Total tokens cached on this single ingest: ~50,200 across 6 reuse calls. For a typical document with 5+ concept calls, the cached prefix is paid once and reused at the discounted rate for all subsequent calls — matching the ~90% savings expected for Anthropic prompt caching.

Routed via openrouter/anthropic/claude-sonnet-4.5; LiteLLM passes the cache_control markers through to Anthropic transparently.

@KylinMountainKylinMountain left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@pitimon@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(compiler): add cache_control breakpoints for Anthropic prompt caching by pitimon · Pull Request #38 · VectifyAI/OpenKB · GitHub
Skip to content

feat(compiler): add cache_control breakpoints for Anthropic prompt caching - #38

Merged
KylinMountain merged 1 commit into
VectifyAI:mainfrom
pitimon:feat/37-cache-control-prompt-caching
May 11, 2026
Merged

feat(compiler): add cache_control breakpoints for Anthropic prompt caching#38
KylinMountain merged 1 commit into
VectifyAI:mainfrom
pitimon:feat/37-cache-control-prompt-caching

Conversation

@pitimon

Copy link
Copy Markdown
Contributor

Summary

  • Adds two cache_control: {"type": "ephemeral"} breakpoints to the compiler so Anthropic prompt caching can hit on every reuse of the base context.
  • Breakpoint 1 — end of doc_msg: caches (system + doc) across summary, concepts plan, and every concept page call (1 + 1 + N + M calls per document).
  • Breakpoint 2 — end of assistant summary: caches (system + doc + summary) across the plan call and every concept generation call.
  • Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call.

Why

Compiler architecture is explicitly designed around base-context reuse (per CLAUDE.md: "Designed around prompt-cache reuse: a single base context A reused across summary → concept-plan → concept-page calls") but no cache markers were emitted, so every call rebilled the full document content. For Anthropic Sonnet 4.5 via OpenRouter or direct, this typically reduces input cost by ~90% on the cached prefix and lowers TTFT after the first call.

Compatibility

  • Anthropic / OpenRouter→Anthropic: cache_control honored.
  • OpenAI: list-of-blocks content is a valid OpenAI-compatible shape; cache_control field is ignored harmlessly.
  • Other providers via LiteLLM: LiteLLM normalizes / strips unknown fields.

_fmt_messages was extended to handle both string and list-of-blocks content shapes for debug output.

Test plan

  • Existing 232 tests pass unchanged (mocks accept *args, **kwargs).
  • New TestCacheControl class — 2 tests:
    • test_short_doc_marks_doc_and_summary captures messages from sync + async LiteLLM calls and asserts cache_control markers are present on doc_msg (all 3 sync calls, all async calls) and on the assistant summary (plan + concept calls).
    • test_long_doc_marks_doc_message asserts the breakpoint on doc_msg for the long-doc path.
  • Full suite: 234 passed (1.05s).
  • Manual smoke against a real Anthropic API key: observe cached_tokens in prompt_tokens_details on calls 2..N (left for reviewer).

Out of scope

  • OpenRouter Response Caching (X-OpenRouter-Cache: true header) — different mechanism, evaluated separately.
  • Splitting compiler.py (~847 lines, pre-existing >800 condition deepened by ~35 lines) — recommend follow-up extracting compiler/messages.py.

Refs #37

…ching
Compiler reuses base context A (system + doc) across N+M+2 LLM calls per
document. Without cache_control markers, every call rebills the full
document content as input tokens.
Adds two breakpoints:
- end of doc_msg: caches (system + doc) for summary, plan, every concept
- end of assistant summary: caches (system + doc + summary) for plan and
every concept generation call
For non-Anthropic providers, the list-of-blocks payload is a valid
OpenAI-compatible shape; LiteLLM normalizes cache_control away.
Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call
(memory observation #82886).
Refs VectifyAI#37
@pitimon

Copy link
Copy Markdown
ContributorAuthor

Smoke test against live OpenRouter → Anthropic Sonnet 4.5

Ingested a fresh OCR'd 21-page Thai procedure document into a short-doc-pipeline KB. Token usage from openkb -v add (verbose mode prints per-call usage):

Stepinput tokenscached tokenshit rate
summary (call 1)6,9900miss (cache being written)
concepts-plan9,5026,98773.5% — breakpoint 1 (system + doc)
concept: preventive-maintenance8,7458,54297.7% — breakpoint 2 (+ summary)
concept: network-management8,7418,54297.7%
concept: network-security-zones8,7488,54297.7%
update: change-management12,6038,54267.8% (existing concept body extends prompt)
update: configuration-management13,7368,54262.2%

Both breakpoints work as designed:

  1. End of doc_msg — every call from concepts-plan onward shows cached >= 6,987, confirming the (system + doc) prefix is being reused.
  2. End of assistant summary — every concept-generation call shows cached = 8,542 (≈ 1,555 tokens beyond breakpoint 1), confirming the (system + doc + summary) prefix is reused across the N+M concept-page calls.

Total tokens cached on this single ingest: ~50,200 across 6 reuse calls. For a typical document with 5+ concept calls, the cached prefix is paid once and reused at the discounted rate for all subsequent calls — matching the ~90% savings expected for Anthropic prompt caching.

Routed via openrouter/anthropic/claude-sonnet-4.5; LiteLLM passes the cache_control markers through to Anthropic transparently.

@KylinMountainKylinMountain left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@pitimon@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(compiler): add cache_control breakpoints for Anthropic prompt caching by pitimon · Pull Request #38 · VectifyAI/OpenKB · GitHub
Skip to content

feat(compiler): add cache_control breakpoints for Anthropic prompt caching - #38

Merged
KylinMountain merged 1 commit into
VectifyAI:mainfrom
pitimon:feat/37-cache-control-prompt-caching
May 11, 2026
Merged

feat(compiler): add cache_control breakpoints for Anthropic prompt caching#38
KylinMountain merged 1 commit into
VectifyAI:mainfrom
pitimon:feat/37-cache-control-prompt-caching

Conversation

@pitimon

Copy link
Copy Markdown
Contributor

Summary

  • Adds two cache_control: {"type": "ephemeral"} breakpoints to the compiler so Anthropic prompt caching can hit on every reuse of the base context.
  • Breakpoint 1 — end of doc_msg: caches (system + doc) across summary, concepts plan, and every concept page call (1 + 1 + N + M calls per document).
  • Breakpoint 2 — end of assistant summary: caches (system + doc + summary) across the plan call and every concept generation call.
  • Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call.

Why

Compiler architecture is explicitly designed around base-context reuse (per CLAUDE.md: "Designed around prompt-cache reuse: a single base context A reused across summary → concept-plan → concept-page calls") but no cache markers were emitted, so every call rebilled the full document content. For Anthropic Sonnet 4.5 via OpenRouter or direct, this typically reduces input cost by ~90% on the cached prefix and lowers TTFT after the first call.

Compatibility

  • Anthropic / OpenRouter→Anthropic: cache_control honored.
  • OpenAI: list-of-blocks content is a valid OpenAI-compatible shape; cache_control field is ignored harmlessly.
  • Other providers via LiteLLM: LiteLLM normalizes / strips unknown fields.

_fmt_messages was extended to handle both string and list-of-blocks content shapes for debug output.

Test plan

  • Existing 232 tests pass unchanged (mocks accept *args, **kwargs).
  • New TestCacheControl class — 2 tests:
    • test_short_doc_marks_doc_and_summary captures messages from sync + async LiteLLM calls and asserts cache_control markers are present on doc_msg (all 3 sync calls, all async calls) and on the assistant summary (plan + concept calls).
    • test_long_doc_marks_doc_message asserts the breakpoint on doc_msg for the long-doc path.
  • Full suite: 234 passed (1.05s).
  • Manual smoke against a real Anthropic API key: observe cached_tokens in prompt_tokens_details on calls 2..N (left for reviewer).

Out of scope

  • OpenRouter Response Caching (X-OpenRouter-Cache: true header) — different mechanism, evaluated separately.
  • Splitting compiler.py (~847 lines, pre-existing >800 condition deepened by ~35 lines) — recommend follow-up extracting compiler/messages.py.

Refs #37

…ching
Compiler reuses base context A (system + doc) across N+M+2 LLM calls per
document. Without cache_control markers, every call rebills the full
document content as input tokens.
Adds two breakpoints:
- end of doc_msg: caches (system + doc) for summary, plan, every concept
- end of assistant summary: caches (system + doc + summary) for plan and
every concept generation call
For non-Anthropic providers, the list-of-blocks payload is a valid
OpenAI-compatible shape; LiteLLM normalizes cache_control away.
Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call
(memory observation #82886).
Refs VectifyAI#37
@pitimon

Copy link
Copy Markdown
ContributorAuthor

Smoke test against live OpenRouter → Anthropic Sonnet 4.5

Ingested a fresh OCR'd 21-page Thai procedure document into a short-doc-pipeline KB. Token usage from openkb -v add (verbose mode prints per-call usage):

Stepinput tokenscached tokenshit rate
summary (call 1)6,9900miss (cache being written)
concepts-plan9,5026,98773.5% — breakpoint 1 (system + doc)
concept: preventive-maintenance8,7458,54297.7% — breakpoint 2 (+ summary)
concept: network-management8,7418,54297.7%
concept: network-security-zones8,7488,54297.7%
update: change-management12,6038,54267.8% (existing concept body extends prompt)
update: configuration-management13,7368,54262.2%

Both breakpoints work as designed:

  1. End of doc_msg — every call from concepts-plan onward shows cached >= 6,987, confirming the (system + doc) prefix is being reused.
  2. End of assistant summary — every concept-generation call shows cached = 8,542 (≈ 1,555 tokens beyond breakpoint 1), confirming the (system + doc + summary) prefix is reused across the N+M concept-page calls.

Total tokens cached on this single ingest: ~50,200 across 6 reuse calls. For a typical document with 5+ concept calls, the cached prefix is paid once and reused at the discounted rate for all subsequent calls — matching the ~90% savings expected for Anthropic prompt caching.

Routed via openrouter/anthropic/claude-sonnet-4.5; LiteLLM passes the cache_control markers through to Anthropic transparently.

@KylinMountainKylinMountain left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@pitimon@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' feat(compiler): add cache_control breakpoints for Anthropic prompt caching by pitimon · Pull Request #38 · VectifyAI/OpenKB · GitHub
Skip to content

feat(compiler): add cache_control breakpoints for Anthropic prompt caching - #38

Merged
KylinMountain merged 1 commit into
VectifyAI:mainfrom
pitimon:feat/37-cache-control-prompt-caching
May 11, 2026
Merged

feat(compiler): add cache_control breakpoints for Anthropic prompt caching#38
KylinMountain merged 1 commit into
VectifyAI:mainfrom
pitimon:feat/37-cache-control-prompt-caching

Conversation

@pitimon

Copy link
Copy Markdown
Contributor

Summary

  • Adds two cache_control: {"type": "ephemeral"} breakpoints to the compiler so Anthropic prompt caching can hit on every reuse of the base context.
  • Breakpoint 1 — end of doc_msg: caches (system + doc) across summary, concepts plan, and every concept page call (1 + 1 + N + M calls per document).
  • Breakpoint 2 — end of assistant summary: caches (system + doc + summary) across the plan call and every concept generation call.
  • Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call.

Why

Compiler architecture is explicitly designed around base-context reuse (per CLAUDE.md: "Designed around prompt-cache reuse: a single base context A reused across summary → concept-plan → concept-page calls") but no cache markers were emitted, so every call rebilled the full document content. For Anthropic Sonnet 4.5 via OpenRouter or direct, this typically reduces input cost by ~90% on the cached prefix and lowers TTFT after the first call.

Compatibility

  • Anthropic / OpenRouter→Anthropic: cache_control honored.
  • OpenAI: list-of-blocks content is a valid OpenAI-compatible shape; cache_control field is ignored harmlessly.
  • Other providers via LiteLLM: LiteLLM normalizes / strips unknown fields.

_fmt_messages was extended to handle both string and list-of-blocks content shapes for debug output.

Test plan

  • Existing 232 tests pass unchanged (mocks accept *args, **kwargs).
  • New TestCacheControl class — 2 tests:
    • test_short_doc_marks_doc_and_summary captures messages from sync + async LiteLLM calls and asserts cache_control markers are present on doc_msg (all 3 sync calls, all async calls) and on the assistant summary (plan + concept calls).
    • test_long_doc_marks_doc_message asserts the breakpoint on doc_msg for the long-doc path.
  • Full suite: 234 passed (1.05s).
  • Manual smoke against a real Anthropic API key: observe cached_tokens in prompt_tokens_details on calls 2..N (left for reviewer).

Out of scope

  • OpenRouter Response Caching (X-OpenRouter-Cache: true header) — different mechanism, evaluated separately.
  • Splitting compiler.py (~847 lines, pre-existing >800 condition deepened by ~35 lines) — recommend follow-up extracting compiler/messages.py.

Refs #37

…ching
Compiler reuses base context A (system + doc) across N+M+2 LLM calls per
document. Without cache_control markers, every call rebills the full
document content as input tokens.
Adds two breakpoints:
- end of doc_msg: caches (system + doc) for summary, plan, every concept
- end of assistant summary: caches (system + doc + summary) for plan and
every concept generation call
For non-Anthropic providers, the list-of-blocks payload is a valid
OpenAI-compatible shape; LiteLLM normalizes cache_control away.
Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call
(memory observation #82886).
Refs VectifyAI#37
@pitimon

Copy link
Copy Markdown
ContributorAuthor

Smoke test against live OpenRouter → Anthropic Sonnet 4.5

Ingested a fresh OCR'd 21-page Thai procedure document into a short-doc-pipeline KB. Token usage from openkb -v add (verbose mode prints per-call usage):

Stepinput tokenscached tokenshit rate
summary (call 1)6,9900miss (cache being written)
concepts-plan9,5026,98773.5% — breakpoint 1 (system + doc)
concept: preventive-maintenance8,7458,54297.7% — breakpoint 2 (+ summary)
concept: network-management8,7418,54297.7%
concept: network-security-zones8,7488,54297.7%
update: change-management12,6038,54267.8% (existing concept body extends prompt)
update: configuration-management13,7368,54262.2%

Both breakpoints work as designed:

  1. End of doc_msg — every call from concepts-plan onward shows cached >= 6,987, confirming the (system + doc) prefix is being reused.
  2. End of assistant summary — every concept-generation call shows cached = 8,542 (≈ 1,555 tokens beyond breakpoint 1), confirming the (system + doc + summary) prefix is reused across the N+M concept-page calls.

Total tokens cached on this single ingest: ~50,200 across 6 reuse calls. For a typical document with 5+ concept calls, the cached prefix is paid once and reused at the discounted rate for all subsequent calls — matching the ~90% savings expected for Anthropic prompt caching.

Routed via openrouter/anthropic/claude-sonnet-4.5; LiteLLM passes the cache_control markers through to Anthropic transparently.

@KylinMountainKylinMountain left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@pitimon@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(compiler): add cache_control breakpoints for Anthropic prompt caching by pitimon · Pull Request #38 · VectifyAI/OpenKB · GitHub
Skip to content

feat(compiler): add cache_control breakpoints for Anthropic prompt caching - #38

Merged
KylinMountain merged 1 commit into
VectifyAI:mainfrom
pitimon:feat/37-cache-control-prompt-caching
May 11, 2026
Merged

feat(compiler): add cache_control breakpoints for Anthropic prompt caching#38
KylinMountain merged 1 commit into
VectifyAI:mainfrom
pitimon:feat/37-cache-control-prompt-caching

Conversation

@pitimon

Copy link
Copy Markdown
Contributor

Summary

  • Adds two cache_control: {"type": "ephemeral"} breakpoints to the compiler so Anthropic prompt caching can hit on every reuse of the base context.
  • Breakpoint 1 — end of doc_msg: caches (system + doc) across summary, concepts plan, and every concept page call (1 + 1 + N + M calls per document).
  • Breakpoint 2 — end of assistant summary: caches (system + doc + summary) across the plan call and every concept generation call.
  • Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call.

Why

Compiler architecture is explicitly designed around base-context reuse (per CLAUDE.md: "Designed around prompt-cache reuse: a single base context A reused across summary → concept-plan → concept-page calls") but no cache markers were emitted, so every call rebilled the full document content. For Anthropic Sonnet 4.5 via OpenRouter or direct, this typically reduces input cost by ~90% on the cached prefix and lowers TTFT after the first call.

Compatibility

  • Anthropic / OpenRouter→Anthropic: cache_control honored.
  • OpenAI: list-of-blocks content is a valid OpenAI-compatible shape; cache_control field is ignored harmlessly.
  • Other providers via LiteLLM: LiteLLM normalizes / strips unknown fields.

_fmt_messages was extended to handle both string and list-of-blocks content shapes for debug output.

Test plan

  • Existing 232 tests pass unchanged (mocks accept *args, **kwargs).
  • New TestCacheControl class — 2 tests:
    • test_short_doc_marks_doc_and_summary captures messages from sync + async LiteLLM calls and asserts cache_control markers are present on doc_msg (all 3 sync calls, all async calls) and on the assistant summary (plan + concept calls).
    • test_long_doc_marks_doc_message asserts the breakpoint on doc_msg for the long-doc path.
  • Full suite: 234 passed (1.05s).
  • Manual smoke against a real Anthropic API key: observe cached_tokens in prompt_tokens_details on calls 2..N (left for reviewer).

Out of scope

  • OpenRouter Response Caching (X-OpenRouter-Cache: true header) — different mechanism, evaluated separately.
  • Splitting compiler.py (~847 lines, pre-existing >800 condition deepened by ~35 lines) — recommend follow-up extracting compiler/messages.py.

Refs #37

…ching
Compiler reuses base context A (system + doc) across N+M+2 LLM calls per
document. Without cache_control markers, every call rebills the full
document content as input tokens.
Adds two breakpoints:
- end of doc_msg: caches (system + doc) for summary, plan, every concept
- end of assistant summary: caches (system + doc + summary) for plan and
every concept generation call
For non-Anthropic providers, the list-of-blocks payload is a valid
OpenAI-compatible shape; LiteLLM normalizes cache_control away.
Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call
(memory observation #82886).
Refs VectifyAI#37
@pitimon

Copy link
Copy Markdown
ContributorAuthor

Smoke test against live OpenRouter → Anthropic Sonnet 4.5

Ingested a fresh OCR'd 21-page Thai procedure document into a short-doc-pipeline KB. Token usage from openkb -v add (verbose mode prints per-call usage):

Stepinput tokenscached tokenshit rate
summary (call 1)6,9900miss (cache being written)
concepts-plan9,5026,98773.5% — breakpoint 1 (system + doc)
concept: preventive-maintenance8,7458,54297.7% — breakpoint 2 (+ summary)
concept: network-management8,7418,54297.7%
concept: network-security-zones8,7488,54297.7%
update: change-management12,6038,54267.8% (existing concept body extends prompt)
update: configuration-management13,7368,54262.2%

Both breakpoints work as designed:

  1. End of doc_msg — every call from concepts-plan onward shows cached >= 6,987, confirming the (system + doc) prefix is being reused.
  2. End of assistant summary — every concept-generation call shows cached = 8,542 (≈ 1,555 tokens beyond breakpoint 1), confirming the (system + doc + summary) prefix is reused across the N+M concept-page calls.

Total tokens cached on this single ingest: ~50,200 across 6 reuse calls. For a typical document with 5+ concept calls, the cached prefix is paid once and reused at the discounted rate for all subsequent calls — matching the ~90% savings expected for Anthropic prompt caching.

Routed via openrouter/anthropic/claude-sonnet-4.5; LiteLLM passes the cache_control markers through to Anthropic transparently.

@KylinMountainKylinMountain left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@pitimon@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' feat(compiler): add cache_control breakpoints for Anthropic prompt caching by pitimon · Pull Request #38 · VectifyAI/OpenKB · GitHub
Skip to content

feat(compiler): add cache_control breakpoints for Anthropic prompt caching - #38

Merged
KylinMountain merged 1 commit into
VectifyAI:mainfrom
pitimon:feat/37-cache-control-prompt-caching
May 11, 2026
Merged

feat(compiler): add cache_control breakpoints for Anthropic prompt caching#38
KylinMountain merged 1 commit into
VectifyAI:mainfrom
pitimon:feat/37-cache-control-prompt-caching

Conversation

@pitimon

Copy link
Copy Markdown
Contributor

Summary

  • Adds two cache_control: {"type": "ephemeral"} breakpoints to the compiler so Anthropic prompt caching can hit on every reuse of the base context.
  • Breakpoint 1 — end of doc_msg: caches (system + doc) across summary, concepts plan, and every concept page call (1 + 1 + N + M calls per document).
  • Breakpoint 2 — end of assistant summary: caches (system + doc + summary) across the plan call and every concept generation call.
  • Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call.

Why

Compiler architecture is explicitly designed around base-context reuse (per CLAUDE.md: "Designed around prompt-cache reuse: a single base context A reused across summary → concept-plan → concept-page calls") but no cache markers were emitted, so every call rebilled the full document content. For Anthropic Sonnet 4.5 via OpenRouter or direct, this typically reduces input cost by ~90% on the cached prefix and lowers TTFT after the first call.

Compatibility

  • Anthropic / OpenRouter→Anthropic: cache_control honored.
  • OpenAI: list-of-blocks content is a valid OpenAI-compatible shape; cache_control field is ignored harmlessly.
  • Other providers via LiteLLM: LiteLLM normalizes / strips unknown fields.

_fmt_messages was extended to handle both string and list-of-blocks content shapes for debug output.

Test plan

  • Existing 232 tests pass unchanged (mocks accept *args, **kwargs).
  • New TestCacheControl class — 2 tests:
    • test_short_doc_marks_doc_and_summary captures messages from sync + async LiteLLM calls and asserts cache_control markers are present on doc_msg (all 3 sync calls, all async calls) and on the assistant summary (plan + concept calls).
    • test_long_doc_marks_doc_message asserts the breakpoint on doc_msg for the long-doc path.
  • Full suite: 234 passed (1.05s).
  • Manual smoke against a real Anthropic API key: observe cached_tokens in prompt_tokens_details on calls 2..N (left for reviewer).

Out of scope

  • OpenRouter Response Caching (X-OpenRouter-Cache: true header) — different mechanism, evaluated separately.
  • Splitting compiler.py (~847 lines, pre-existing >800 condition deepened by ~35 lines) — recommend follow-up extracting compiler/messages.py.

Refs #37

…ching
Compiler reuses base context A (system + doc) across N+M+2 LLM calls per
document. Without cache_control markers, every call rebills the full
document content as input tokens.
Adds two breakpoints:
- end of doc_msg: caches (system + doc) for summary, plan, every concept
- end of assistant summary: caches (system + doc + summary) for plan and
every concept generation call
For non-Anthropic providers, the list-of-blocks payload is a valid
OpenAI-compatible shape; LiteLLM normalizes cache_control away.
Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call
(memory observation #82886).
Refs VectifyAI#37
@pitimon

Copy link
Copy Markdown
ContributorAuthor

Smoke test against live OpenRouter → Anthropic Sonnet 4.5

Ingested a fresh OCR'd 21-page Thai procedure document into a short-doc-pipeline KB. Token usage from openkb -v add (verbose mode prints per-call usage):

Stepinput tokenscached tokenshit rate
summary (call 1)6,9900miss (cache being written)
concepts-plan9,5026,98773.5% — breakpoint 1 (system + doc)
concept: preventive-maintenance8,7458,54297.7% — breakpoint 2 (+ summary)
concept: network-management8,7418,54297.7%
concept: network-security-zones8,7488,54297.7%
update: change-management12,6038,54267.8% (existing concept body extends prompt)
update: configuration-management13,7368,54262.2%

Both breakpoints work as designed:

  1. End of doc_msg — every call from concepts-plan onward shows cached >= 6,987, confirming the (system + doc) prefix is being reused.
  2. End of assistant summary — every concept-generation call shows cached = 8,542 (≈ 1,555 tokens beyond breakpoint 1), confirming the (system + doc + summary) prefix is reused across the N+M concept-page calls.

Total tokens cached on this single ingest: ~50,200 across 6 reuse calls. For a typical document with 5+ concept calls, the cached prefix is paid once and reused at the discounted rate for all subsequent calls — matching the ~90% savings expected for Anthropic prompt caching.

Routed via openrouter/anthropic/claude-sonnet-4.5; LiteLLM passes the cache_control markers through to Anthropic transparently.

@KylinMountainKylinMountain left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@pitimon@KylinMountain
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); feat(compiler): add cache_control breakpoints for Anthropic prompt caching by pitimon · Pull Request #38 · VectifyAI/OpenKB · GitHub
Skip to content

feat(compiler): add cache_control breakpoints for Anthropic prompt caching - #38

Merged
KylinMountain merged 1 commit into
VectifyAI:mainfrom
pitimon:feat/37-cache-control-prompt-caching
May 11, 2026
Merged

feat(compiler): add cache_control breakpoints for Anthropic prompt caching#38
KylinMountain merged 1 commit into
VectifyAI:mainfrom
pitimon:feat/37-cache-control-prompt-caching

Conversation

@pitimon

Copy link
Copy Markdown
Contributor

Summary

  • Adds two cache_control: {"type": "ephemeral"} breakpoints to the compiler so Anthropic prompt caching can hit on every reuse of the base context.
  • Breakpoint 1 — end of doc_msg: caches (system + doc) across summary, concepts plan, and every concept page call (1 + 1 + N + M calls per document).
  • Breakpoint 2 — end of assistant summary: caches (system + doc + summary) across the plan call and every concept generation call.
  • Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call.

Why

Compiler architecture is explicitly designed around base-context reuse (per CLAUDE.md: "Designed around prompt-cache reuse: a single base context A reused across summary → concept-plan → concept-page calls") but no cache markers were emitted, so every call rebilled the full document content. For Anthropic Sonnet 4.5 via OpenRouter or direct, this typically reduces input cost by ~90% on the cached prefix and lowers TTFT after the first call.

Compatibility

  • Anthropic / OpenRouter→Anthropic: cache_control honored.
  • OpenAI: list-of-blocks content is a valid OpenAI-compatible shape; cache_control field is ignored harmlessly.
  • Other providers via LiteLLM: LiteLLM normalizes / strips unknown fields.

_fmt_messages was extended to handle both string and list-of-blocks content shapes for debug output.

Test plan

  • Existing 232 tests pass unchanged (mocks accept *args, **kwargs).
  • New TestCacheControl class — 2 tests:
    • test_short_doc_marks_doc_and_summary captures messages from sync + async LiteLLM calls and asserts cache_control markers are present on doc_msg (all 3 sync calls, all async calls) and on the assistant summary (plan + concept calls).
    • test_long_doc_marks_doc_message asserts the breakpoint on doc_msg for the long-doc path.
  • Full suite: 234 passed (1.05s).
  • Manual smoke against a real Anthropic API key: observe cached_tokens in prompt_tokens_details on calls 2..N (left for reviewer).

Out of scope

  • OpenRouter Response Caching (X-OpenRouter-Cache: true header) — different mechanism, evaluated separately.
  • Splitting compiler.py (~847 lines, pre-existing >800 condition deepened by ~35 lines) — recommend follow-up extracting compiler/messages.py.

Refs #37

…ching
Compiler reuses base context A (system + doc) across N+M+2 LLM calls per
document. Without cache_control markers, every call rebills the full
document content as input tokens.
Adds two breakpoints:
- end of doc_msg: caches (system + doc) for summary, plan, every concept
- end of assistant summary: caches (system + doc + summary) for plan and
every concept generation call
For non-Anthropic providers, the list-of-blocks payload is a valid
OpenAI-compatible shape; LiteLLM normalizes cache_control away.
Side fix: _llm_call_async now forwards **kwargs for parity with _llm_call
(memory observation #82886).
Refs VectifyAI#37
@pitimon

Copy link
Copy Markdown
ContributorAuthor

Smoke test against live OpenRouter → Anthropic Sonnet 4.5

Ingested a fresh OCR'd 21-page Thai procedure document into a short-doc-pipeline KB. Token usage from openkb -v add (verbose mode prints per-call usage):

Stepinput tokenscached tokenshit rate
summary (call 1)6,9900miss (cache being written)
concepts-plan9,5026,98773.5% — breakpoint 1 (system + doc)
concept: preventive-maintenance8,7458,54297.7% — breakpoint 2 (+ summary)
concept: network-management8,7418,54297.7%
concept: network-security-zones8,7488,54297.7%
update: change-management12,6038,54267.8% (existing concept body extends prompt)
update: configuration-management13,7368,54262.2%

Both breakpoints work as designed:

  1. End of doc_msg — every call from concepts-plan onward shows cached >= 6,987, confirming the (system + doc) prefix is being reused.
  2. End of assistant summary — every concept-generation call shows cached = 8,542 (≈ 1,555 tokens beyond breakpoint 1), confirming the (system + doc + summary) prefix is reused across the N+M concept-page calls.

Total tokens cached on this single ingest: ~50,200 across 6 reuse calls. For a typical document with 5+ concept calls, the cached prefix is paid once and reused at the discounted rate for all subsequent calls — matching the ~90% savings expected for Anthropic prompt caching.

Routed via openrouter/anthropic/claude-sonnet-4.5; LiteLLM passes the cache_control markers through to Anthropic transparently.

@KylinMountainKylinMountain left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@pitimon@KylinMountain