Skip to content

fix(json): stop a skipped sub-score being averaged as a mismatch - #210

Open
Arthi Arumugam (arthi-arumugam-git) wants to merge 1 commit into
braintrustdata:mainfrom
arthi-arumugam-git:fix/json-diff-skipped-scores
Open

fix(json): stop a skipped sub-score being averaged as a mismatch#210
Arthi Arumugam (arthi-arumugam-git) wants to merge 1 commit into
braintrustdata:mainfrom
arthi-arumugam-git:fix/json-diff-skipped-scores

Conversation

@arthi-arumugam-git

Copy link
Copy Markdown

The problem

Score documents the contract:

If the score is None, the evaluation is considered to be skipped.

JSONDiff.json_diff filters those out before averaging, but the two branches then divide by
different things:

# dictbase_scores= [sforsinbase_scoresifsisnotNone]
returnsum(base_scores) /len(base_scores) # count AFTER filtering# listbase_scores= [sforsinbase_scoresifsisnotNone]
returnsum(base_scores) /max(len(o1), len(o2)) # count BEFORE filtering

In an object a skipped key is removed from the numerator and the denominator, so it is
ignored, which matches the documented meaning. In an array it is removed from the numerator
only, so it is averaged in as a zero.

The same skip therefore lands differently depending on the container:

inputscorer skips one valuescore on main
{"a": "skip", "b": "same"}yes1.0
["skip", "same"]yes0.5

Same values, same scorer, same skip.

This is reachable with any scorer that can abstain, which is the case the None score exists
for: an LLM judge that declines to answer, a scorer that cannot parse a field. The array
result is silently lower and nothing raises.

Second, smaller issue

The dict branch divides by len(base_scores) with no guard. An object whose comparisons are
all skipped leaves that list empty and raises ZeroDivisionError. The empty-object case is
handled above it, so this needs a non-empty object rather than an edge case in the inputs.

The change

Both branches now drop skipped comparisons from the denominator, and return None when
nothing is left to average, which propagates the skip upward instead of inventing a number or
raising.

The list denominator stays max(len(o1), len(o2)) minus the skips. That distinction is
deliberate: an element with no counterpart is a genuine difference and must stay counted,
which is what max(...) contributes over the zip. Only the skips come out.

Tests

py/autoevals/test_json_skipped_scores.py, eight tests, four of which fail on main:

FAILED test_skipped_element_in_a_list_is_not_scored_as_zero
FAILED test_list_and_dict_treat_an_identical_skip_the_same_way
FAILED test_a_fully_skipped_object_reports_a_skip_rather_than_raising
FAILED test_a_fully_skipped_list_reports_a_skip_rather_than_raising

The other four pin behaviour that must not move: missing elements are still penalised
(["a"] vs ["a", "b"] is still 0.5), and unskipped lists, dicts and empty containers score
exactly as before.

py/autoevals/test_json.py: 4 passed. Across py/autoevals the results are identical with
and without this diff (21 failed, 39 passed, 8 errors, all needing API keys or litellm).
black and isort clean at line-length 119.

Score documents that "If the score is None, the evaluation is considered to be
skipped", and json_diff filters those out before averaging. The two branches then
disagreed about what to divide by.
The dict branch divided by len(base_scores), the count after filtering, so a
skipped key was excluded from both the numerator and the denominator and
correctly ignored. The list branch divided by max(len(o1), len(o2)), which still
counts the skipped element, so the same skip was averaged in as a zero.
The result is that one skipped comparison scores 1.0 inside an object and 0.5
inside a two-element array, for identical values and an identical scorer.
The dict branch also divided by len(base_scores) with no guard. An object whose
every comparison is skipped leaves that list empty and raises ZeroDivisionError
rather than reporting a skip.
Both branches now drop skipped comparisons from the denominator and return None,
propagating the skip, when nothing is left to average. The list denominator stays
max(len(o1), len(o2)) minus the skips, so elements with no counterpart are still
counted as a real difference; only the skips come out.
Eight tests, four of which fail on main. The other four pin what must not move:
missing elements are still penalised, and unskipped lists, dicts and empty
containers score exactly as before.
@arthi-arumugam-git

Copy link
Copy Markdown
Author

Flagging in case it is not visible from your side: the workflow run on this PR needs maintainer approval before anything executes, so the checks here are empty rather than failing.

Locally the eight added cases pass, four of which fail on main, and the rest of the py/autoevals suite is byte-identical with and without the change (the pre-existing failures there need API keys). Happy for CI to run whenever someone can enable it, and to fix anything it turns up.

@arthi-arumugam-git

Copy link
Copy Markdown
Author

The red checks that showed up on this PR today are not test failures. The six runs were created on 2 Aug, sat waiting for maintainer approval, and GitHub expired them at the 30 day mark (created 2026-08-02 02:33 UTC, marked failed 2026-09-01 02:36 UTC). All six have zero jobs and no logs, so nothing actually ran. The same thing has happened to the other fork PRs on this repo, so it does not look specific to this branch.

I re-checked the change locally today on Python 3.13:

  • the 8 new tests plus the existing test_json.py pass, 12 passed
  • with json.py put back to main and the tests kept, 4 fail, the same 4 listed in the description
  • the rest of py/autoevals gives an identical 28 failures with and without this change (all of them missing API keys or missing optional modules)
  • black and ruff are clean on both changed files

The branch is still level with main, so there is nothing to rebase. Since these runs have already expired, a fresh one would need to be queued before it can be approved. I can push an empty commit to trigger that if it is useful, just say the word.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@arthi-arumugam-git
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
fix(json): stop a skipped sub-score being averaged as a mismatch by arthi-arumugam-git · Pull Request #210 · braintrustdata/autoevals · GitHub
Skip to content

fix(json): stop a skipped sub-score being averaged as a mismatch - #210

Open
Arthi Arumugam (arthi-arumugam-git) wants to merge 1 commit into
braintrustdata:mainfrom
arthi-arumugam-git:fix/json-diff-skipped-scores
Open

fix(json): stop a skipped sub-score being averaged as a mismatch#210
Arthi Arumugam (arthi-arumugam-git) wants to merge 1 commit into
braintrustdata:mainfrom
arthi-arumugam-git:fix/json-diff-skipped-scores

Conversation

@arthi-arumugam-git

Copy link
Copy Markdown

The problem

Score documents the contract:

If the score is None, the evaluation is considered to be skipped.

JSONDiff.json_diff filters those out before averaging, but the two branches then divide by
different things:

# dictbase_scores= [sforsinbase_scoresifsisnotNone]
returnsum(base_scores) /len(base_scores) # count AFTER filtering# listbase_scores= [sforsinbase_scoresifsisnotNone]
returnsum(base_scores) /max(len(o1), len(o2)) # count BEFORE filtering

In an object a skipped key is removed from the numerator and the denominator, so it is
ignored, which matches the documented meaning. In an array it is removed from the numerator
only, so it is averaged in as a zero.

The same skip therefore lands differently depending on the container:

inputscorer skips one valuescore on main
{"a": "skip", "b": "same"}yes1.0
["skip", "same"]yes0.5

Same values, same scorer, same skip.

This is reachable with any scorer that can abstain, which is the case the None score exists
for: an LLM judge that declines to answer, a scorer that cannot parse a field. The array
result is silently lower and nothing raises.

Second, smaller issue

The dict branch divides by len(base_scores) with no guard. An object whose comparisons are
all skipped leaves that list empty and raises ZeroDivisionError. The empty-object case is
handled above it, so this needs a non-empty object rather than an edge case in the inputs.

The change

Both branches now drop skipped comparisons from the denominator, and return None when
nothing is left to average, which propagates the skip upward instead of inventing a number or
raising.

The list denominator stays max(len(o1), len(o2)) minus the skips. That distinction is
deliberate: an element with no counterpart is a genuine difference and must stay counted,
which is what max(...) contributes over the zip. Only the skips come out.

Tests

py/autoevals/test_json_skipped_scores.py, eight tests, four of which fail on main:

FAILED test_skipped_element_in_a_list_is_not_scored_as_zero
FAILED test_list_and_dict_treat_an_identical_skip_the_same_way
FAILED test_a_fully_skipped_object_reports_a_skip_rather_than_raising
FAILED test_a_fully_skipped_list_reports_a_skip_rather_than_raising

The other four pin behaviour that must not move: missing elements are still penalised
(["a"] vs ["a", "b"] is still 0.5), and unskipped lists, dicts and empty containers score
exactly as before.

py/autoevals/test_json.py: 4 passed. Across py/autoevals the results are identical with
and without this diff (21 failed, 39 passed, 8 errors, all needing API keys or litellm).
black and isort clean at line-length 119.

Score documents that "If the score is None, the evaluation is considered to be
skipped", and json_diff filters those out before averaging. The two branches then
disagreed about what to divide by.
The dict branch divided by len(base_scores), the count after filtering, so a
skipped key was excluded from both the numerator and the denominator and
correctly ignored. The list branch divided by max(len(o1), len(o2)), which still
counts the skipped element, so the same skip was averaged in as a zero.
The result is that one skipped comparison scores 1.0 inside an object and 0.5
inside a two-element array, for identical values and an identical scorer.
The dict branch also divided by len(base_scores) with no guard. An object whose
every comparison is skipped leaves that list empty and raises ZeroDivisionError
rather than reporting a skip.
Both branches now drop skipped comparisons from the denominator and return None,
propagating the skip, when nothing is left to average. The list denominator stays
max(len(o1), len(o2)) minus the skips, so elements with no counterpart are still
counted as a real difference; only the skips come out.
Eight tests, four of which fail on main. The other four pin what must not move:
missing elements are still penalised, and unskipped lists, dicts and empty
containers score exactly as before.
@arthi-arumugam-git

Copy link
Copy Markdown
Author

Flagging in case it is not visible from your side: the workflow run on this PR needs maintainer approval before anything executes, so the checks here are empty rather than failing.

Locally the eight added cases pass, four of which fail on main, and the rest of the py/autoevals suite is byte-identical with and without the change (the pre-existing failures there need API keys). Happy for CI to run whenever someone can enable it, and to fix anything it turns up.

@arthi-arumugam-git

Copy link
Copy Markdown
Author

The red checks that showed up on this PR today are not test failures. The six runs were created on 2 Aug, sat waiting for maintainer approval, and GitHub expired them at the 30 day mark (created 2026-08-02 02:33 UTC, marked failed 2026-09-01 02:36 UTC). All six have zero jobs and no logs, so nothing actually ran. The same thing has happened to the other fork PRs on this repo, so it does not look specific to this branch.

I re-checked the change locally today on Python 3.13:

  • the 8 new tests plus the existing test_json.py pass, 12 passed
  • with json.py put back to main and the tests kept, 4 fail, the same 4 listed in the description
  • the rest of py/autoevals gives an identical 28 failures with and without this change (all of them missing API keys or missing optional modules)
  • black and ruff are clean on both changed files

The branch is still level with main, so there is nothing to rebase. Since these runs have already expired, a fresh one would need to be queued before it can be approved. I can push an empty commit to trigger that if it is useful, just say the word.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@arthi-arumugam-git
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(json): stop a skipped sub-score being averaged as a mismatch by arthi-arumugam-git · Pull Request #210 · braintrustdata/autoevals · GitHub
Skip to content

fix(json): stop a skipped sub-score being averaged as a mismatch - #210

Open
Arthi Arumugam (arthi-arumugam-git) wants to merge 1 commit into
braintrustdata:mainfrom
arthi-arumugam-git:fix/json-diff-skipped-scores
Open

fix(json): stop a skipped sub-score being averaged as a mismatch#210
Arthi Arumugam (arthi-arumugam-git) wants to merge 1 commit into
braintrustdata:mainfrom
arthi-arumugam-git:fix/json-diff-skipped-scores

Conversation

@arthi-arumugam-git

Copy link
Copy Markdown

The problem

Score documents the contract:

If the score is None, the evaluation is considered to be skipped.

JSONDiff.json_diff filters those out before averaging, but the two branches then divide by
different things:

# dictbase_scores= [sforsinbase_scoresifsisnotNone]
returnsum(base_scores) /len(base_scores) # count AFTER filtering# listbase_scores= [sforsinbase_scoresifsisnotNone]
returnsum(base_scores) /max(len(o1), len(o2)) # count BEFORE filtering

In an object a skipped key is removed from the numerator and the denominator, so it is
ignored, which matches the documented meaning. In an array it is removed from the numerator
only, so it is averaged in as a zero.

The same skip therefore lands differently depending on the container:

inputscorer skips one valuescore on main
{"a": "skip", "b": "same"}yes1.0
["skip", "same"]yes0.5

Same values, same scorer, same skip.

This is reachable with any scorer that can abstain, which is the case the None score exists
for: an LLM judge that declines to answer, a scorer that cannot parse a field. The array
result is silently lower and nothing raises.

Second, smaller issue

The dict branch divides by len(base_scores) with no guard. An object whose comparisons are
all skipped leaves that list empty and raises ZeroDivisionError. The empty-object case is
handled above it, so this needs a non-empty object rather than an edge case in the inputs.

The change

Both branches now drop skipped comparisons from the denominator, and return None when
nothing is left to average, which propagates the skip upward instead of inventing a number or
raising.

The list denominator stays max(len(o1), len(o2)) minus the skips. That distinction is
deliberate: an element with no counterpart is a genuine difference and must stay counted,
which is what max(...) contributes over the zip. Only the skips come out.

Tests

py/autoevals/test_json_skipped_scores.py, eight tests, four of which fail on main:

FAILED test_skipped_element_in_a_list_is_not_scored_as_zero
FAILED test_list_and_dict_treat_an_identical_skip_the_same_way
FAILED test_a_fully_skipped_object_reports_a_skip_rather_than_raising
FAILED test_a_fully_skipped_list_reports_a_skip_rather_than_raising

The other four pin behaviour that must not move: missing elements are still penalised
(["a"] vs ["a", "b"] is still 0.5), and unskipped lists, dicts and empty containers score
exactly as before.

py/autoevals/test_json.py: 4 passed. Across py/autoevals the results are identical with
and without this diff (21 failed, 39 passed, 8 errors, all needing API keys or litellm).
black and isort clean at line-length 119.

Score documents that "If the score is None, the evaluation is considered to be
skipped", and json_diff filters those out before averaging. The two branches then
disagreed about what to divide by.
The dict branch divided by len(base_scores), the count after filtering, so a
skipped key was excluded from both the numerator and the denominator and
correctly ignored. The list branch divided by max(len(o1), len(o2)), which still
counts the skipped element, so the same skip was averaged in as a zero.
The result is that one skipped comparison scores 1.0 inside an object and 0.5
inside a two-element array, for identical values and an identical scorer.
The dict branch also divided by len(base_scores) with no guard. An object whose
every comparison is skipped leaves that list empty and raises ZeroDivisionError
rather than reporting a skip.
Both branches now drop skipped comparisons from the denominator and return None,
propagating the skip, when nothing is left to average. The list denominator stays
max(len(o1), len(o2)) minus the skips, so elements with no counterpart are still
counted as a real difference; only the skips come out.
Eight tests, four of which fail on main. The other four pin what must not move:
missing elements are still penalised, and unskipped lists, dicts and empty
containers score exactly as before.
@arthi-arumugam-git

Copy link
Copy Markdown
Author

Flagging in case it is not visible from your side: the workflow run on this PR needs maintainer approval before anything executes, so the checks here are empty rather than failing.

Locally the eight added cases pass, four of which fail on main, and the rest of the py/autoevals suite is byte-identical with and without the change (the pre-existing failures there need API keys). Happy for CI to run whenever someone can enable it, and to fix anything it turns up.

@arthi-arumugam-git

Copy link
Copy Markdown
Author

The red checks that showed up on this PR today are not test failures. The six runs were created on 2 Aug, sat waiting for maintainer approval, and GitHub expired them at the 30 day mark (created 2026-08-02 02:33 UTC, marked failed 2026-09-01 02:36 UTC). All six have zero jobs and no logs, so nothing actually ran. The same thing has happened to the other fork PRs on this repo, so it does not look specific to this branch.

I re-checked the change locally today on Python 3.13:

  • the 8 new tests plus the existing test_json.py pass, 12 passed
  • with json.py put back to main and the tests kept, 4 fail, the same 4 listed in the description
  • the rest of py/autoevals gives an identical 28 failures with and without this change (all of them missing API keys or missing optional modules)
  • black and ruff are clean on both changed files

The branch is still level with main, so there is nothing to rebase. Since these runs have already expired, a fresh one would need to be queued before it can be approved. I can push an empty commit to trigger that if it is useful, just say the word.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@arthi-arumugam-git
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(json): stop a skipped sub-score being averaged as a mismatch by arthi-arumugam-git · Pull Request #210 · braintrustdata/autoevals · GitHub
Skip to content

fix(json): stop a skipped sub-score being averaged as a mismatch - #210

Open
Arthi Arumugam (arthi-arumugam-git) wants to merge 1 commit into
braintrustdata:mainfrom
arthi-arumugam-git:fix/json-diff-skipped-scores
Open

fix(json): stop a skipped sub-score being averaged as a mismatch#210
Arthi Arumugam (arthi-arumugam-git) wants to merge 1 commit into
braintrustdata:mainfrom
arthi-arumugam-git:fix/json-diff-skipped-scores

Conversation

@arthi-arumugam-git

Copy link
Copy Markdown

The problem

Score documents the contract:

If the score is None, the evaluation is considered to be skipped.

JSONDiff.json_diff filters those out before averaging, but the two branches then divide by
different things:

# dictbase_scores= [sforsinbase_scoresifsisnotNone]
returnsum(base_scores) /len(base_scores) # count AFTER filtering# listbase_scores= [sforsinbase_scoresifsisnotNone]
returnsum(base_scores) /max(len(o1), len(o2)) # count BEFORE filtering

In an object a skipped key is removed from the numerator and the denominator, so it is
ignored, which matches the documented meaning. In an array it is removed from the numerator
only, so it is averaged in as a zero.

The same skip therefore lands differently depending on the container:

inputscorer skips one valuescore on main
{"a": "skip", "b": "same"}yes1.0
["skip", "same"]yes0.5

Same values, same scorer, same skip.

This is reachable with any scorer that can abstain, which is the case the None score exists
for: an LLM judge that declines to answer, a scorer that cannot parse a field. The array
result is silently lower and nothing raises.

Second, smaller issue

The dict branch divides by len(base_scores) with no guard. An object whose comparisons are
all skipped leaves that list empty and raises ZeroDivisionError. The empty-object case is
handled above it, so this needs a non-empty object rather than an edge case in the inputs.

The change

Both branches now drop skipped comparisons from the denominator, and return None when
nothing is left to average, which propagates the skip upward instead of inventing a number or
raising.

The list denominator stays max(len(o1), len(o2)) minus the skips. That distinction is
deliberate: an element with no counterpart is a genuine difference and must stay counted,
which is what max(...) contributes over the zip. Only the skips come out.

Tests

py/autoevals/test_json_skipped_scores.py, eight tests, four of which fail on main:

FAILED test_skipped_element_in_a_list_is_not_scored_as_zero
FAILED test_list_and_dict_treat_an_identical_skip_the_same_way
FAILED test_a_fully_skipped_object_reports_a_skip_rather_than_raising
FAILED test_a_fully_skipped_list_reports_a_skip_rather_than_raising

The other four pin behaviour that must not move: missing elements are still penalised
(["a"] vs ["a", "b"] is still 0.5), and unskipped lists, dicts and empty containers score
exactly as before.

py/autoevals/test_json.py: 4 passed. Across py/autoevals the results are identical with
and without this diff (21 failed, 39 passed, 8 errors, all needing API keys or litellm).
black and isort clean at line-length 119.

Score documents that "If the score is None, the evaluation is considered to be
skipped", and json_diff filters those out before averaging. The two branches then
disagreed about what to divide by.
The dict branch divided by len(base_scores), the count after filtering, so a
skipped key was excluded from both the numerator and the denominator and
correctly ignored. The list branch divided by max(len(o1), len(o2)), which still
counts the skipped element, so the same skip was averaged in as a zero.
The result is that one skipped comparison scores 1.0 inside an object and 0.5
inside a two-element array, for identical values and an identical scorer.
The dict branch also divided by len(base_scores) with no guard. An object whose
every comparison is skipped leaves that list empty and raises ZeroDivisionError
rather than reporting a skip.
Both branches now drop skipped comparisons from the denominator and return None,
propagating the skip, when nothing is left to average. The list denominator stays
max(len(o1), len(o2)) minus the skips, so elements with no counterpart are still
counted as a real difference; only the skips come out.
Eight tests, four of which fail on main. The other four pin what must not move:
missing elements are still penalised, and unskipped lists, dicts and empty
containers score exactly as before.
@arthi-arumugam-git

Copy link
Copy Markdown
Author

Flagging in case it is not visible from your side: the workflow run on this PR needs maintainer approval before anything executes, so the checks here are empty rather than failing.

Locally the eight added cases pass, four of which fail on main, and the rest of the py/autoevals suite is byte-identical with and without the change (the pre-existing failures there need API keys). Happy for CI to run whenever someone can enable it, and to fix anything it turns up.

@arthi-arumugam-git

Copy link
Copy Markdown
Author

The red checks that showed up on this PR today are not test failures. The six runs were created on 2 Aug, sat waiting for maintainer approval, and GitHub expired them at the 30 day mark (created 2026-08-02 02:33 UTC, marked failed 2026-09-01 02:36 UTC). All six have zero jobs and no logs, so nothing actually ran. The same thing has happened to the other fork PRs on this repo, so it does not look specific to this branch.

I re-checked the change locally today on Python 3.13:

  • the 8 new tests plus the existing test_json.py pass, 12 passed
  • with json.py put back to main and the tests kept, 4 fail, the same 4 listed in the description
  • the rest of py/autoevals gives an identical 28 failures with and without this change (all of them missing API keys or missing optional modules)
  • black and ruff are clean on both changed files

The branch is still level with main, so there is nothing to rebase. Since these runs have already expired, a fresh one would need to be queued before it can be approved. I can push an empty commit to trigger that if it is useful, just say the word.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@arthi-arumugam-git
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' fix(json): stop a skipped sub-score being averaged as a mismatch by arthi-arumugam-git · Pull Request #210 · braintrustdata/autoevals · GitHub
Skip to content

fix(json): stop a skipped sub-score being averaged as a mismatch - #210

Open
Arthi Arumugam (arthi-arumugam-git) wants to merge 1 commit into
braintrustdata:mainfrom
arthi-arumugam-git:fix/json-diff-skipped-scores
Open

fix(json): stop a skipped sub-score being averaged as a mismatch#210
Arthi Arumugam (arthi-arumugam-git) wants to merge 1 commit into
braintrustdata:mainfrom
arthi-arumugam-git:fix/json-diff-skipped-scores

Conversation

@arthi-arumugam-git

Copy link
Copy Markdown

The problem

Score documents the contract:

If the score is None, the evaluation is considered to be skipped.

JSONDiff.json_diff filters those out before averaging, but the two branches then divide by
different things:

# dictbase_scores= [sforsinbase_scoresifsisnotNone]
returnsum(base_scores) /len(base_scores) # count AFTER filtering# listbase_scores= [sforsinbase_scoresifsisnotNone]
returnsum(base_scores) /max(len(o1), len(o2)) # count BEFORE filtering

In an object a skipped key is removed from the numerator and the denominator, so it is
ignored, which matches the documented meaning. In an array it is removed from the numerator
only, so it is averaged in as a zero.

The same skip therefore lands differently depending on the container:

inputscorer skips one valuescore on main
{"a": "skip", "b": "same"}yes1.0
["skip", "same"]yes0.5

Same values, same scorer, same skip.

This is reachable with any scorer that can abstain, which is the case the None score exists
for: an LLM judge that declines to answer, a scorer that cannot parse a field. The array
result is silently lower and nothing raises.

Second, smaller issue

The dict branch divides by len(base_scores) with no guard. An object whose comparisons are
all skipped leaves that list empty and raises ZeroDivisionError. The empty-object case is
handled above it, so this needs a non-empty object rather than an edge case in the inputs.

The change

Both branches now drop skipped comparisons from the denominator, and return None when
nothing is left to average, which propagates the skip upward instead of inventing a number or
raising.

The list denominator stays max(len(o1), len(o2)) minus the skips. That distinction is
deliberate: an element with no counterpart is a genuine difference and must stay counted,
which is what max(...) contributes over the zip. Only the skips come out.

Tests

py/autoevals/test_json_skipped_scores.py, eight tests, four of which fail on main:

FAILED test_skipped_element_in_a_list_is_not_scored_as_zero
FAILED test_list_and_dict_treat_an_identical_skip_the_same_way
FAILED test_a_fully_skipped_object_reports_a_skip_rather_than_raising
FAILED test_a_fully_skipped_list_reports_a_skip_rather_than_raising

The other four pin behaviour that must not move: missing elements are still penalised
(["a"] vs ["a", "b"] is still 0.5), and unskipped lists, dicts and empty containers score
exactly as before.

py/autoevals/test_json.py: 4 passed. Across py/autoevals the results are identical with
and without this diff (21 failed, 39 passed, 8 errors, all needing API keys or litellm).
black and isort clean at line-length 119.

Score documents that "If the score is None, the evaluation is considered to be
skipped", and json_diff filters those out before averaging. The two branches then
disagreed about what to divide by.
The dict branch divided by len(base_scores), the count after filtering, so a
skipped key was excluded from both the numerator and the denominator and
correctly ignored. The list branch divided by max(len(o1), len(o2)), which still
counts the skipped element, so the same skip was averaged in as a zero.
The result is that one skipped comparison scores 1.0 inside an object and 0.5
inside a two-element array, for identical values and an identical scorer.
The dict branch also divided by len(base_scores) with no guard. An object whose
every comparison is skipped leaves that list empty and raises ZeroDivisionError
rather than reporting a skip.
Both branches now drop skipped comparisons from the denominator and return None,
propagating the skip, when nothing is left to average. The list denominator stays
max(len(o1), len(o2)) minus the skips, so elements with no counterpart are still
counted as a real difference; only the skips come out.
Eight tests, four of which fail on main. The other four pin what must not move:
missing elements are still penalised, and unskipped lists, dicts and empty
containers score exactly as before.
@arthi-arumugam-git

Copy link
Copy Markdown
Author

Flagging in case it is not visible from your side: the workflow run on this PR needs maintainer approval before anything executes, so the checks here are empty rather than failing.

Locally the eight added cases pass, four of which fail on main, and the rest of the py/autoevals suite is byte-identical with and without the change (the pre-existing failures there need API keys). Happy for CI to run whenever someone can enable it, and to fix anything it turns up.

@arthi-arumugam-git

Copy link
Copy Markdown
Author

The red checks that showed up on this PR today are not test failures. The six runs were created on 2 Aug, sat waiting for maintainer approval, and GitHub expired them at the 30 day mark (created 2026-08-02 02:33 UTC, marked failed 2026-09-01 02:36 UTC). All six have zero jobs and no logs, so nothing actually ran. The same thing has happened to the other fork PRs on this repo, so it does not look specific to this branch.

I re-checked the change locally today on Python 3.13:

  • the 8 new tests plus the existing test_json.py pass, 12 passed
  • with json.py put back to main and the tests kept, 4 fail, the same 4 listed in the description
  • the rest of py/autoevals gives an identical 28 failures with and without this change (all of them missing API keys or missing optional modules)
  • black and ruff are clean on both changed files

The branch is still level with main, so there is nothing to rebase. Since these runs have already expired, a fresh one would need to be queued before it can be approved. I can push an empty commit to trigger that if it is useful, just say the word.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@arthi-arumugam-git
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(json): stop a skipped sub-score being averaged as a mismatch by arthi-arumugam-git · Pull Request #210 · braintrustdata/autoevals · GitHub
Skip to content

fix(json): stop a skipped sub-score being averaged as a mismatch - #210

Open
Arthi Arumugam (arthi-arumugam-git) wants to merge 1 commit into
braintrustdata:mainfrom
arthi-arumugam-git:fix/json-diff-skipped-scores
Open

fix(json): stop a skipped sub-score being averaged as a mismatch#210
Arthi Arumugam (arthi-arumugam-git) wants to merge 1 commit into
braintrustdata:mainfrom
arthi-arumugam-git:fix/json-diff-skipped-scores

Conversation

@arthi-arumugam-git

Copy link
Copy Markdown

The problem

Score documents the contract:

If the score is None, the evaluation is considered to be skipped.

JSONDiff.json_diff filters those out before averaging, but the two branches then divide by
different things:

# dictbase_scores= [sforsinbase_scoresifsisnotNone]
returnsum(base_scores) /len(base_scores) # count AFTER filtering# listbase_scores= [sforsinbase_scoresifsisnotNone]
returnsum(base_scores) /max(len(o1), len(o2)) # count BEFORE filtering

In an object a skipped key is removed from the numerator and the denominator, so it is
ignored, which matches the documented meaning. In an array it is removed from the numerator
only, so it is averaged in as a zero.

The same skip therefore lands differently depending on the container:

inputscorer skips one valuescore on main
{"a": "skip", "b": "same"}yes1.0
["skip", "same"]yes0.5

Same values, same scorer, same skip.

This is reachable with any scorer that can abstain, which is the case the None score exists
for: an LLM judge that declines to answer, a scorer that cannot parse a field. The array
result is silently lower and nothing raises.

Second, smaller issue

The dict branch divides by len(base_scores) with no guard. An object whose comparisons are
all skipped leaves that list empty and raises ZeroDivisionError. The empty-object case is
handled above it, so this needs a non-empty object rather than an edge case in the inputs.

The change

Both branches now drop skipped comparisons from the denominator, and return None when
nothing is left to average, which propagates the skip upward instead of inventing a number or
raising.

The list denominator stays max(len(o1), len(o2)) minus the skips. That distinction is
deliberate: an element with no counterpart is a genuine difference and must stay counted,
which is what max(...) contributes over the zip. Only the skips come out.

Tests

py/autoevals/test_json_skipped_scores.py, eight tests, four of which fail on main:

FAILED test_skipped_element_in_a_list_is_not_scored_as_zero
FAILED test_list_and_dict_treat_an_identical_skip_the_same_way
FAILED test_a_fully_skipped_object_reports_a_skip_rather_than_raising
FAILED test_a_fully_skipped_list_reports_a_skip_rather_than_raising

The other four pin behaviour that must not move: missing elements are still penalised
(["a"] vs ["a", "b"] is still 0.5), and unskipped lists, dicts and empty containers score
exactly as before.

py/autoevals/test_json.py: 4 passed. Across py/autoevals the results are identical with
and without this diff (21 failed, 39 passed, 8 errors, all needing API keys or litellm).
black and isort clean at line-length 119.

Score documents that "If the score is None, the evaluation is considered to be
skipped", and json_diff filters those out before averaging. The two branches then
disagreed about what to divide by.
The dict branch divided by len(base_scores), the count after filtering, so a
skipped key was excluded from both the numerator and the denominator and
correctly ignored. The list branch divided by max(len(o1), len(o2)), which still
counts the skipped element, so the same skip was averaged in as a zero.
The result is that one skipped comparison scores 1.0 inside an object and 0.5
inside a two-element array, for identical values and an identical scorer.
The dict branch also divided by len(base_scores) with no guard. An object whose
every comparison is skipped leaves that list empty and raises ZeroDivisionError
rather than reporting a skip.
Both branches now drop skipped comparisons from the denominator and return None,
propagating the skip, when nothing is left to average. The list denominator stays
max(len(o1), len(o2)) minus the skips, so elements with no counterpart are still
counted as a real difference; only the skips come out.
Eight tests, four of which fail on main. The other four pin what must not move:
missing elements are still penalised, and unskipped lists, dicts and empty
containers score exactly as before.
@arthi-arumugam-git

Copy link
Copy Markdown
Author

Flagging in case it is not visible from your side: the workflow run on this PR needs maintainer approval before anything executes, so the checks here are empty rather than failing.

Locally the eight added cases pass, four of which fail on main, and the rest of the py/autoevals suite is byte-identical with and without the change (the pre-existing failures there need API keys). Happy for CI to run whenever someone can enable it, and to fix anything it turns up.

@arthi-arumugam-git

Copy link
Copy Markdown
Author

The red checks that showed up on this PR today are not test failures. The six runs were created on 2 Aug, sat waiting for maintainer approval, and GitHub expired them at the 30 day mark (created 2026-08-02 02:33 UTC, marked failed 2026-09-01 02:36 UTC). All six have zero jobs and no logs, so nothing actually ran. The same thing has happened to the other fork PRs on this repo, so it does not look specific to this branch.

I re-checked the change locally today on Python 3.13:

  • the 8 new tests plus the existing test_json.py pass, 12 passed
  • with json.py put back to main and the tests kept, 4 fail, the same 4 listed in the description
  • the rest of py/autoevals gives an identical 28 failures with and without this change (all of them missing API keys or missing optional modules)
  • black and ruff are clean on both changed files

The branch is still level with main, so there is nothing to rebase. Since these runs have already expired, a fresh one would need to be queued before it can be approved. I can push an empty commit to trigger that if it is useful, just say the word.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@arthi-arumugam-git
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' fix(json): stop a skipped sub-score being averaged as a mismatch by arthi-arumugam-git · Pull Request #210 · braintrustdata/autoevals · GitHub
Skip to content

fix(json): stop a skipped sub-score being averaged as a mismatch - #210

Open
Arthi Arumugam (arthi-arumugam-git) wants to merge 1 commit into
braintrustdata:mainfrom
arthi-arumugam-git:fix/json-diff-skipped-scores
Open

fix(json): stop a skipped sub-score being averaged as a mismatch#210
Arthi Arumugam (arthi-arumugam-git) wants to merge 1 commit into
braintrustdata:mainfrom
arthi-arumugam-git:fix/json-diff-skipped-scores

Conversation

@arthi-arumugam-git

Copy link
Copy Markdown

The problem

Score documents the contract:

If the score is None, the evaluation is considered to be skipped.

JSONDiff.json_diff filters those out before averaging, but the two branches then divide by
different things:

# dictbase_scores= [sforsinbase_scoresifsisnotNone]
returnsum(base_scores) /len(base_scores) # count AFTER filtering# listbase_scores= [sforsinbase_scoresifsisnotNone]
returnsum(base_scores) /max(len(o1), len(o2)) # count BEFORE filtering

In an object a skipped key is removed from the numerator and the denominator, so it is
ignored, which matches the documented meaning. In an array it is removed from the numerator
only, so it is averaged in as a zero.

The same skip therefore lands differently depending on the container:

inputscorer skips one valuescore on main
{"a": "skip", "b": "same"}yes1.0
["skip", "same"]yes0.5

Same values, same scorer, same skip.

This is reachable with any scorer that can abstain, which is the case the None score exists
for: an LLM judge that declines to answer, a scorer that cannot parse a field. The array
result is silently lower and nothing raises.

Second, smaller issue

The dict branch divides by len(base_scores) with no guard. An object whose comparisons are
all skipped leaves that list empty and raises ZeroDivisionError. The empty-object case is
handled above it, so this needs a non-empty object rather than an edge case in the inputs.

The change

Both branches now drop skipped comparisons from the denominator, and return None when
nothing is left to average, which propagates the skip upward instead of inventing a number or
raising.

The list denominator stays max(len(o1), len(o2)) minus the skips. That distinction is
deliberate: an element with no counterpart is a genuine difference and must stay counted,
which is what max(...) contributes over the zip. Only the skips come out.

Tests

py/autoevals/test_json_skipped_scores.py, eight tests, four of which fail on main:

FAILED test_skipped_element_in_a_list_is_not_scored_as_zero
FAILED test_list_and_dict_treat_an_identical_skip_the_same_way
FAILED test_a_fully_skipped_object_reports_a_skip_rather_than_raising
FAILED test_a_fully_skipped_list_reports_a_skip_rather_than_raising

The other four pin behaviour that must not move: missing elements are still penalised
(["a"] vs ["a", "b"] is still 0.5), and unskipped lists, dicts and empty containers score
exactly as before.

py/autoevals/test_json.py: 4 passed. Across py/autoevals the results are identical with
and without this diff (21 failed, 39 passed, 8 errors, all needing API keys or litellm).
black and isort clean at line-length 119.

Score documents that "If the score is None, the evaluation is considered to be
skipped", and json_diff filters those out before averaging. The two branches then
disagreed about what to divide by.
The dict branch divided by len(base_scores), the count after filtering, so a
skipped key was excluded from both the numerator and the denominator and
correctly ignored. The list branch divided by max(len(o1), len(o2)), which still
counts the skipped element, so the same skip was averaged in as a zero.
The result is that one skipped comparison scores 1.0 inside an object and 0.5
inside a two-element array, for identical values and an identical scorer.
The dict branch also divided by len(base_scores) with no guard. An object whose
every comparison is skipped leaves that list empty and raises ZeroDivisionError
rather than reporting a skip.
Both branches now drop skipped comparisons from the denominator and return None,
propagating the skip, when nothing is left to average. The list denominator stays
max(len(o1), len(o2)) minus the skips, so elements with no counterpart are still
counted as a real difference; only the skips come out.
Eight tests, four of which fail on main. The other four pin what must not move:
missing elements are still penalised, and unskipped lists, dicts and empty
containers score exactly as before.
@arthi-arumugam-git

Copy link
Copy Markdown
Author

Flagging in case it is not visible from your side: the workflow run on this PR needs maintainer approval before anything executes, so the checks here are empty rather than failing.

Locally the eight added cases pass, four of which fail on main, and the rest of the py/autoevals suite is byte-identical with and without the change (the pre-existing failures there need API keys). Happy for CI to run whenever someone can enable it, and to fix anything it turns up.

@arthi-arumugam-git

Copy link
Copy Markdown
Author

The red checks that showed up on this PR today are not test failures. The six runs were created on 2 Aug, sat waiting for maintainer approval, and GitHub expired them at the 30 day mark (created 2026-08-02 02:33 UTC, marked failed 2026-09-01 02:36 UTC). All six have zero jobs and no logs, so nothing actually ran. The same thing has happened to the other fork PRs on this repo, so it does not look specific to this branch.

I re-checked the change locally today on Python 3.13:

  • the 8 new tests plus the existing test_json.py pass, 12 passed
  • with json.py put back to main and the tests kept, 4 fail, the same 4 listed in the description
  • the rest of py/autoevals gives an identical 28 failures with and without this change (all of them missing API keys or missing optional modules)
  • black and ruff are clean on both changed files

The branch is still level with main, so there is nothing to rebase. Since these runs have already expired, a fresh one would need to be queued before it can be approved. I can push an empty commit to trigger that if it is useful, just say the word.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@arthi-arumugam-git
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); fix(json): stop a skipped sub-score being averaged as a mismatch by arthi-arumugam-git · Pull Request #210 · braintrustdata/autoevals · GitHub
Skip to content

fix(json): stop a skipped sub-score being averaged as a mismatch - #210

Open
Arthi Arumugam (arthi-arumugam-git) wants to merge 1 commit into
braintrustdata:mainfrom
arthi-arumugam-git:fix/json-diff-skipped-scores
Open

fix(json): stop a skipped sub-score being averaged as a mismatch#210
Arthi Arumugam (arthi-arumugam-git) wants to merge 1 commit into
braintrustdata:mainfrom
arthi-arumugam-git:fix/json-diff-skipped-scores

Conversation

@arthi-arumugam-git

Copy link
Copy Markdown

The problem

Score documents the contract:

If the score is None, the evaluation is considered to be skipped.

JSONDiff.json_diff filters those out before averaging, but the two branches then divide by
different things:

# dictbase_scores= [sforsinbase_scoresifsisnotNone]
returnsum(base_scores) /len(base_scores) # count AFTER filtering# listbase_scores= [sforsinbase_scoresifsisnotNone]
returnsum(base_scores) /max(len(o1), len(o2)) # count BEFORE filtering

In an object a skipped key is removed from the numerator and the denominator, so it is
ignored, which matches the documented meaning. In an array it is removed from the numerator
only, so it is averaged in as a zero.

The same skip therefore lands differently depending on the container:

inputscorer skips one valuescore on main
{"a": "skip", "b": "same"}yes1.0
["skip", "same"]yes0.5

Same values, same scorer, same skip.

This is reachable with any scorer that can abstain, which is the case the None score exists
for: an LLM judge that declines to answer, a scorer that cannot parse a field. The array
result is silently lower and nothing raises.

Second, smaller issue

The dict branch divides by len(base_scores) with no guard. An object whose comparisons are
all skipped leaves that list empty and raises ZeroDivisionError. The empty-object case is
handled above it, so this needs a non-empty object rather than an edge case in the inputs.

The change

Both branches now drop skipped comparisons from the denominator, and return None when
nothing is left to average, which propagates the skip upward instead of inventing a number or
raising.

The list denominator stays max(len(o1), len(o2)) minus the skips. That distinction is
deliberate: an element with no counterpart is a genuine difference and must stay counted,
which is what max(...) contributes over the zip. Only the skips come out.

Tests

py/autoevals/test_json_skipped_scores.py, eight tests, four of which fail on main:

FAILED test_skipped_element_in_a_list_is_not_scored_as_zero
FAILED test_list_and_dict_treat_an_identical_skip_the_same_way
FAILED test_a_fully_skipped_object_reports_a_skip_rather_than_raising
FAILED test_a_fully_skipped_list_reports_a_skip_rather_than_raising

The other four pin behaviour that must not move: missing elements are still penalised
(["a"] vs ["a", "b"] is still 0.5), and unskipped lists, dicts and empty containers score
exactly as before.

py/autoevals/test_json.py: 4 passed. Across py/autoevals the results are identical with
and without this diff (21 failed, 39 passed, 8 errors, all needing API keys or litellm).
black and isort clean at line-length 119.

Score documents that "If the score is None, the evaluation is considered to be
skipped", and json_diff filters those out before averaging. The two branches then
disagreed about what to divide by.
The dict branch divided by len(base_scores), the count after filtering, so a
skipped key was excluded from both the numerator and the denominator and
correctly ignored. The list branch divided by max(len(o1), len(o2)), which still
counts the skipped element, so the same skip was averaged in as a zero.
The result is that one skipped comparison scores 1.0 inside an object and 0.5
inside a two-element array, for identical values and an identical scorer.
The dict branch also divided by len(base_scores) with no guard. An object whose
every comparison is skipped leaves that list empty and raises ZeroDivisionError
rather than reporting a skip.
Both branches now drop skipped comparisons from the denominator and return None,
propagating the skip, when nothing is left to average. The list denominator stays
max(len(o1), len(o2)) minus the skips, so elements with no counterpart are still
counted as a real difference; only the skips come out.
Eight tests, four of which fail on main. The other four pin what must not move:
missing elements are still penalised, and unskipped lists, dicts and empty
containers score exactly as before.
@arthi-arumugam-git

Copy link
Copy Markdown
Author

Flagging in case it is not visible from your side: the workflow run on this PR needs maintainer approval before anything executes, so the checks here are empty rather than failing.

Locally the eight added cases pass, four of which fail on main, and the rest of the py/autoevals suite is byte-identical with and without the change (the pre-existing failures there need API keys). Happy for CI to run whenever someone can enable it, and to fix anything it turns up.

@arthi-arumugam-git

Copy link
Copy Markdown
Author

The red checks that showed up on this PR today are not test failures. The six runs were created on 2 Aug, sat waiting for maintainer approval, and GitHub expired them at the 30 day mark (created 2026-08-02 02:33 UTC, marked failed 2026-09-01 02:36 UTC). All six have zero jobs and no logs, so nothing actually ran. The same thing has happened to the other fork PRs on this repo, so it does not look specific to this branch.

I re-checked the change locally today on Python 3.13:

  • the 8 new tests plus the existing test_json.py pass, 12 passed
  • with json.py put back to main and the tests kept, 4 fail, the same 4 listed in the description
  • the rest of py/autoevals gives an identical 28 failures with and without this change (all of them missing API keys or missing optional modules)
  • black and ruff are clean on both changed files

The branch is still level with main, so there is nothing to rebase. Since these runs have already expired, a fresh one would need to be queued before it can be approved. I can push an empty commit to trigger that if it is useful, just say the word.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@arthi-arumugam-git