Skip to content

fix: isolate embedding cache by model and prefix - #213

Open
Daniel Peng (original4422) wants to merge 1 commit into
braintrustdata:mainfrom
original4422:fix/embedding-cache-configuration
Open

fix: isolate embedding cache by model and prefix#213
Daniel Peng (original4422) wants to merge 1 commit into
braintrustdata:mainfrom
original4422:fix/embedding-cache-configuration

Conversation

@original4422

Copy link
Copy Markdown

Summary

  • key the class-level embedding cache by model, prefix, and normalized input
  • share the same cache-key logic between sync and async embedding paths
  • add deterministic sync and async regressions for configuration isolation and same-configuration cache hits

Root cause

EmbeddingSimilarity used only the normalized input as its class-level cache key even though the embedding request also depends on model and prefix. Instances with different configurations could therefore reuse vectors produced by another model or prefixed input.

Validation

  • pytest py/autoevals/test_embeddings.py -k 'cache_isolated' -q (2 passed)
  • pytest py/autoevals/test_values.py py/autoevals/test_json.py -q (8 passed)
  • pre-commit run --all-files
  • python -m build
  • python -m twine check dist/*

The full local suite completed with 72 passed and 17 integration failures because OpenAI/Braintrust API credentials are unavailable locally; those failures stop at client initialization and do not exercise this cache change.

Fixes#211

CopilotAI lite review requested due to automatic review settings August 19, 2026 19:07

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes a correctness issue in the Python EmbeddingSimilarity scorer where its class-level embedding cache could be inadvertently shared across different configurations, leading to cross-model / cross-prefix cache poisoning. It aligns with the library’s goal of providing reliable, deterministic evaluators by ensuring cached embedding results are only reused when the embedding request configuration matches.

Changes:

  • Introduces a shared cache-key function for EmbeddingSimilarity and keys the cache by (model, prefix, normalized_input).
  • Applies the same cache-keying behavior to both sync and async embedding paths.
  • Adds regression tests to confirm cache isolation across models/prefixes and cache hits within the same configuration.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

FileDescription
py/autoevals/string.pyFixes EmbeddingSimilarity cache keying to prevent cross-configuration reuse and shares key logic between sync/async paths.
py/autoevals/test_embeddings.pyAdds deterministic regression tests (sync + async) and a per-test cache reset fixture to validate isolation and same-config cache hits.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@original4422

Copy link
Copy Markdown
Author

Hi David Elner (@delner), when you have a chance, would you mind taking a look at this cache-key fix and letting me know if any changes are needed? Thank you!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

EmbeddingSimilarity class-level embedding cache ignores model and prefix — cross-model cache poisoning

2 participants

@original4422
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
fix: isolate embedding cache by model and prefix by original4422 · Pull Request #213 · braintrustdata/autoevals · GitHub
Skip to content

fix: isolate embedding cache by model and prefix - #213

Open
Daniel Peng (original4422) wants to merge 1 commit into
braintrustdata:mainfrom
original4422:fix/embedding-cache-configuration
Open

fix: isolate embedding cache by model and prefix#213
Daniel Peng (original4422) wants to merge 1 commit into
braintrustdata:mainfrom
original4422:fix/embedding-cache-configuration

Conversation

@original4422

Copy link
Copy Markdown

Summary

  • key the class-level embedding cache by model, prefix, and normalized input
  • share the same cache-key logic between sync and async embedding paths
  • add deterministic sync and async regressions for configuration isolation and same-configuration cache hits

Root cause

EmbeddingSimilarity used only the normalized input as its class-level cache key even though the embedding request also depends on model and prefix. Instances with different configurations could therefore reuse vectors produced by another model or prefixed input.

Validation

  • pytest py/autoevals/test_embeddings.py -k 'cache_isolated' -q (2 passed)
  • pytest py/autoevals/test_values.py py/autoevals/test_json.py -q (8 passed)
  • pre-commit run --all-files
  • python -m build
  • python -m twine check dist/*

The full local suite completed with 72 passed and 17 integration failures because OpenAI/Braintrust API credentials are unavailable locally; those failures stop at client initialization and do not exercise this cache change.

Fixes#211

CopilotAI lite review requested due to automatic review settings August 19, 2026 19:07

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes a correctness issue in the Python EmbeddingSimilarity scorer where its class-level embedding cache could be inadvertently shared across different configurations, leading to cross-model / cross-prefix cache poisoning. It aligns with the library’s goal of providing reliable, deterministic evaluators by ensuring cached embedding results are only reused when the embedding request configuration matches.

Changes:

  • Introduces a shared cache-key function for EmbeddingSimilarity and keys the cache by (model, prefix, normalized_input).
  • Applies the same cache-keying behavior to both sync and async embedding paths.
  • Adds regression tests to confirm cache isolation across models/prefixes and cache hits within the same configuration.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

FileDescription
py/autoevals/string.pyFixes EmbeddingSimilarity cache keying to prevent cross-configuration reuse and shares key logic between sync/async paths.
py/autoevals/test_embeddings.pyAdds deterministic regression tests (sync + async) and a per-test cache reset fixture to validate isolation and same-config cache hits.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@original4422

Copy link
Copy Markdown
Author

Hi David Elner (@delner), when you have a chance, would you mind taking a look at this cache-key fix and letting me know if any changes are needed? Thank you!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

EmbeddingSimilarity class-level embedding cache ignores model and prefix — cross-model cache poisoning

2 participants

@original4422
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix: isolate embedding cache by model and prefix by original4422 · Pull Request #213 · braintrustdata/autoevals · GitHub
Skip to content

fix: isolate embedding cache by model and prefix - #213

Open
Daniel Peng (original4422) wants to merge 1 commit into
braintrustdata:mainfrom
original4422:fix/embedding-cache-configuration
Open

fix: isolate embedding cache by model and prefix#213
Daniel Peng (original4422) wants to merge 1 commit into
braintrustdata:mainfrom
original4422:fix/embedding-cache-configuration

Conversation

@original4422

Copy link
Copy Markdown

Summary

  • key the class-level embedding cache by model, prefix, and normalized input
  • share the same cache-key logic between sync and async embedding paths
  • add deterministic sync and async regressions for configuration isolation and same-configuration cache hits

Root cause

EmbeddingSimilarity used only the normalized input as its class-level cache key even though the embedding request also depends on model and prefix. Instances with different configurations could therefore reuse vectors produced by another model or prefixed input.

Validation

  • pytest py/autoevals/test_embeddings.py -k 'cache_isolated' -q (2 passed)
  • pytest py/autoevals/test_values.py py/autoevals/test_json.py -q (8 passed)
  • pre-commit run --all-files
  • python -m build
  • python -m twine check dist/*

The full local suite completed with 72 passed and 17 integration failures because OpenAI/Braintrust API credentials are unavailable locally; those failures stop at client initialization and do not exercise this cache change.

Fixes#211

CopilotAI lite review requested due to automatic review settings August 19, 2026 19:07

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes a correctness issue in the Python EmbeddingSimilarity scorer where its class-level embedding cache could be inadvertently shared across different configurations, leading to cross-model / cross-prefix cache poisoning. It aligns with the library’s goal of providing reliable, deterministic evaluators by ensuring cached embedding results are only reused when the embedding request configuration matches.

Changes:

  • Introduces a shared cache-key function for EmbeddingSimilarity and keys the cache by (model, prefix, normalized_input).
  • Applies the same cache-keying behavior to both sync and async embedding paths.
  • Adds regression tests to confirm cache isolation across models/prefixes and cache hits within the same configuration.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

FileDescription
py/autoevals/string.pyFixes EmbeddingSimilarity cache keying to prevent cross-configuration reuse and shares key logic between sync/async paths.
py/autoevals/test_embeddings.pyAdds deterministic regression tests (sync + async) and a per-test cache reset fixture to validate isolation and same-config cache hits.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@original4422

Copy link
Copy Markdown
Author

Hi David Elner (@delner), when you have a chance, would you mind taking a look at this cache-key fix and letting me know if any changes are needed? Thank you!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

EmbeddingSimilarity class-level embedding cache ignores model and prefix — cross-model cache poisoning

2 participants

@original4422
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix: isolate embedding cache by model and prefix by original4422 · Pull Request #213 · braintrustdata/autoevals · GitHub
Skip to content

fix: isolate embedding cache by model and prefix - #213

Open
Daniel Peng (original4422) wants to merge 1 commit into
braintrustdata:mainfrom
original4422:fix/embedding-cache-configuration
Open

fix: isolate embedding cache by model and prefix#213
Daniel Peng (original4422) wants to merge 1 commit into
braintrustdata:mainfrom
original4422:fix/embedding-cache-configuration

Conversation

@original4422

Copy link
Copy Markdown

Summary

  • key the class-level embedding cache by model, prefix, and normalized input
  • share the same cache-key logic between sync and async embedding paths
  • add deterministic sync and async regressions for configuration isolation and same-configuration cache hits

Root cause

EmbeddingSimilarity used only the normalized input as its class-level cache key even though the embedding request also depends on model and prefix. Instances with different configurations could therefore reuse vectors produced by another model or prefixed input.

Validation

  • pytest py/autoevals/test_embeddings.py -k 'cache_isolated' -q (2 passed)
  • pytest py/autoevals/test_values.py py/autoevals/test_json.py -q (8 passed)
  • pre-commit run --all-files
  • python -m build
  • python -m twine check dist/*

The full local suite completed with 72 passed and 17 integration failures because OpenAI/Braintrust API credentials are unavailable locally; those failures stop at client initialization and do not exercise this cache change.

Fixes#211

CopilotAI lite review requested due to automatic review settings August 19, 2026 19:07

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes a correctness issue in the Python EmbeddingSimilarity scorer where its class-level embedding cache could be inadvertently shared across different configurations, leading to cross-model / cross-prefix cache poisoning. It aligns with the library’s goal of providing reliable, deterministic evaluators by ensuring cached embedding results are only reused when the embedding request configuration matches.

Changes:

  • Introduces a shared cache-key function for EmbeddingSimilarity and keys the cache by (model, prefix, normalized_input).
  • Applies the same cache-keying behavior to both sync and async embedding paths.
  • Adds regression tests to confirm cache isolation across models/prefixes and cache hits within the same configuration.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

FileDescription
py/autoevals/string.pyFixes EmbeddingSimilarity cache keying to prevent cross-configuration reuse and shares key logic between sync/async paths.
py/autoevals/test_embeddings.pyAdds deterministic regression tests (sync + async) and a per-test cache reset fixture to validate isolation and same-config cache hits.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@original4422

Copy link
Copy Markdown
Author

Hi David Elner (@delner), when you have a chance, would you mind taking a look at this cache-key fix and letting me know if any changes are needed? Thank you!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

EmbeddingSimilarity class-level embedding cache ignores model and prefix — cross-model cache poisoning

2 participants

@original4422
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' fix: isolate embedding cache by model and prefix by original4422 · Pull Request #213 · braintrustdata/autoevals · GitHub
Skip to content

fix: isolate embedding cache by model and prefix - #213

Open
Daniel Peng (original4422) wants to merge 1 commit into
braintrustdata:mainfrom
original4422:fix/embedding-cache-configuration
Open

fix: isolate embedding cache by model and prefix#213
Daniel Peng (original4422) wants to merge 1 commit into
braintrustdata:mainfrom
original4422:fix/embedding-cache-configuration

Conversation

@original4422

Copy link
Copy Markdown

Summary

  • key the class-level embedding cache by model, prefix, and normalized input
  • share the same cache-key logic between sync and async embedding paths
  • add deterministic sync and async regressions for configuration isolation and same-configuration cache hits

Root cause

EmbeddingSimilarity used only the normalized input as its class-level cache key even though the embedding request also depends on model and prefix. Instances with different configurations could therefore reuse vectors produced by another model or prefixed input.

Validation

  • pytest py/autoevals/test_embeddings.py -k 'cache_isolated' -q (2 passed)
  • pytest py/autoevals/test_values.py py/autoevals/test_json.py -q (8 passed)
  • pre-commit run --all-files
  • python -m build
  • python -m twine check dist/*

The full local suite completed with 72 passed and 17 integration failures because OpenAI/Braintrust API credentials are unavailable locally; those failures stop at client initialization and do not exercise this cache change.

Fixes#211

CopilotAI lite review requested due to automatic review settings August 19, 2026 19:07

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes a correctness issue in the Python EmbeddingSimilarity scorer where its class-level embedding cache could be inadvertently shared across different configurations, leading to cross-model / cross-prefix cache poisoning. It aligns with the library’s goal of providing reliable, deterministic evaluators by ensuring cached embedding results are only reused when the embedding request configuration matches.

Changes:

  • Introduces a shared cache-key function for EmbeddingSimilarity and keys the cache by (model, prefix, normalized_input).
  • Applies the same cache-keying behavior to both sync and async embedding paths.
  • Adds regression tests to confirm cache isolation across models/prefixes and cache hits within the same configuration.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

FileDescription
py/autoevals/string.pyFixes EmbeddingSimilarity cache keying to prevent cross-configuration reuse and shares key logic between sync/async paths.
py/autoevals/test_embeddings.pyAdds deterministic regression tests (sync + async) and a per-test cache reset fixture to validate isolation and same-config cache hits.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@original4422

Copy link
Copy Markdown
Author

Hi David Elner (@delner), when you have a chance, would you mind taking a look at this cache-key fix and letting me know if any changes are needed? Thank you!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

EmbeddingSimilarity class-level embedding cache ignores model and prefix — cross-model cache poisoning

2 participants

@original4422
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix: isolate embedding cache by model and prefix by original4422 · Pull Request #213 · braintrustdata/autoevals · GitHub
Skip to content

fix: isolate embedding cache by model and prefix - #213

Open
Daniel Peng (original4422) wants to merge 1 commit into
braintrustdata:mainfrom
original4422:fix/embedding-cache-configuration
Open

fix: isolate embedding cache by model and prefix#213
Daniel Peng (original4422) wants to merge 1 commit into
braintrustdata:mainfrom
original4422:fix/embedding-cache-configuration

Conversation

@original4422

Copy link
Copy Markdown

Summary

  • key the class-level embedding cache by model, prefix, and normalized input
  • share the same cache-key logic between sync and async embedding paths
  • add deterministic sync and async regressions for configuration isolation and same-configuration cache hits

Root cause

EmbeddingSimilarity used only the normalized input as its class-level cache key even though the embedding request also depends on model and prefix. Instances with different configurations could therefore reuse vectors produced by another model or prefixed input.

Validation

  • pytest py/autoevals/test_embeddings.py -k 'cache_isolated' -q (2 passed)
  • pytest py/autoevals/test_values.py py/autoevals/test_json.py -q (8 passed)
  • pre-commit run --all-files
  • python -m build
  • python -m twine check dist/*

The full local suite completed with 72 passed and 17 integration failures because OpenAI/Braintrust API credentials are unavailable locally; those failures stop at client initialization and do not exercise this cache change.

Fixes#211

CopilotAI lite review requested due to automatic review settings August 19, 2026 19:07

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes a correctness issue in the Python EmbeddingSimilarity scorer where its class-level embedding cache could be inadvertently shared across different configurations, leading to cross-model / cross-prefix cache poisoning. It aligns with the library’s goal of providing reliable, deterministic evaluators by ensuring cached embedding results are only reused when the embedding request configuration matches.

Changes:

  • Introduces a shared cache-key function for EmbeddingSimilarity and keys the cache by (model, prefix, normalized_input).
  • Applies the same cache-keying behavior to both sync and async embedding paths.
  • Adds regression tests to confirm cache isolation across models/prefixes and cache hits within the same configuration.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

FileDescription
py/autoevals/string.pyFixes EmbeddingSimilarity cache keying to prevent cross-configuration reuse and shares key logic between sync/async paths.
py/autoevals/test_embeddings.pyAdds deterministic regression tests (sync + async) and a per-test cache reset fixture to validate isolation and same-config cache hits.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@original4422

Copy link
Copy Markdown
Author

Hi David Elner (@delner), when you have a chance, would you mind taking a look at this cache-key fix and letting me know if any changes are needed? Thank you!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

EmbeddingSimilarity class-level embedding cache ignores model and prefix — cross-model cache poisoning

2 participants

@original4422
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); })(); fix: isolate embedding cache by model and prefix by original4422 · Pull Request #213 · braintrustdata/autoevals · GitHub
Skip to content

fix: isolate embedding cache by model and prefix - #213

Open
Daniel Peng (original4422) wants to merge 1 commit into
braintrustdata:mainfrom
original4422:fix/embedding-cache-configuration
Open

fix: isolate embedding cache by model and prefix#213
Daniel Peng (original4422) wants to merge 1 commit into
braintrustdata:mainfrom
original4422:fix/embedding-cache-configuration

Conversation

@original4422

Copy link
Copy Markdown

Summary

  • key the class-level embedding cache by model, prefix, and normalized input
  • share the same cache-key logic between sync and async embedding paths
  • add deterministic sync and async regressions for configuration isolation and same-configuration cache hits

Root cause

EmbeddingSimilarity used only the normalized input as its class-level cache key even though the embedding request also depends on model and prefix. Instances with different configurations could therefore reuse vectors produced by another model or prefixed input.

Validation

  • pytest py/autoevals/test_embeddings.py -k 'cache_isolated' -q (2 passed)
  • pytest py/autoevals/test_values.py py/autoevals/test_json.py -q (8 passed)
  • pre-commit run --all-files
  • python -m build
  • python -m twine check dist/*

The full local suite completed with 72 passed and 17 integration failures because OpenAI/Braintrust API credentials are unavailable locally; those failures stop at client initialization and do not exercise this cache change.

Fixes#211

CopilotAI lite review requested due to automatic review settings August 19, 2026 19:07

CopilotAI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes a correctness issue in the Python EmbeddingSimilarity scorer where its class-level embedding cache could be inadvertently shared across different configurations, leading to cross-model / cross-prefix cache poisoning. It aligns with the library’s goal of providing reliable, deterministic evaluators by ensuring cached embedding results are only reused when the embedding request configuration matches.

Changes:

  • Introduces a shared cache-key function for EmbeddingSimilarity and keys the cache by (model, prefix, normalized_input).
  • Applies the same cache-keying behavior to both sync and async embedding paths.
  • Adds regression tests to confirm cache isolation across models/prefixes and cache hits within the same configuration.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

FileDescription
py/autoevals/string.pyFixes EmbeddingSimilarity cache keying to prevent cross-configuration reuse and shares key logic between sync/async paths.
py/autoevals/test_embeddings.pyAdds deterministic regression tests (sync + async) and a per-test cache reset fixture to validate isolation and same-config cache hits.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@original4422

Copy link
Copy Markdown
Author

Hi David Elner (@delner), when you have a chance, would you mind taking a look at this cache-key fix and letting me know if any changes are needed? Thank you!

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

EmbeddingSimilarity class-level embedding cache ignores model and prefix — cross-model cache poisoning

2 participants

@original4422