Uh oh!
There was an error while loading. Please reload this page.
[codex] docs(eval): publish nine-arm Terminal-Bench 2.1 leaderboard - #3004
Merged
Conversation
4 tasks
M4n5terforce-pushed
the
codex/tb21-nine-arm-report
branch
from
August 26, 2026 08:46
f5af587 to
40045ffCompareM4n5terforce-pushed
the
codex/tb21-nine-arm-report
branch
from
August 26, 2026 10:03
40045ff to
63a861cCompareAstro-Han
marked this pull request as ready for review
September 1, 2026 02:08
Astro-Han
approved these changes
Sep 1, 2026
Astro-Han
left a comment
Contributor
There was a problem hiding this comment.
Pushed one commit to your branch: the test failure was just the missing ASF header on nine-arm.md. The audit passes now.
Also marked the PR ready for review — these numbers deserve to be published, and I want to cite them in the website positioning discussion (#4307).
简体中文
往你的分支推了一个提交:test 挂掉的唯一原因是 nine-arm.md 缺 ASF header,现在审计通过了。
顺手转成了正式 PR —— 这份数据值得发布,我想在 #4307 的网站定位讨论里引用它。
Uh oh!
There was an error while loading. Please reload this page.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for freeto join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Docs-only: publish the final nine-harness Terminal-Bench 2.1 comparison around DeepSeek V4 Flash at max reasoning effort, plus an 89-task accepted-outcome CSV.
This extends the four-arm report in #2208 with OpenCode, Kimi Code, ZCode, Pi, and DeepSeek Harness (DSH).
Full report:
docs/eval/terminal-bench-2.1-deepseek-v4-flash-nine-arm.mdLeaderboard
All nine harnesses pass the same 28/89 tasks. Five tasks fail on all nine.
Scope and accounting
d49e28f1e4ddd13d289e85a5f312a66750951932deepseek-v4-flash, thinking enabled, reasoning effortmaxDSH recovery disclosure
DSH initially recorded 61/89. Real-machine diagnosis found three Eval-specific defects: a five-minute Bash timeout that interrupted
dpkg, interactivetzdatasetup, and background-service teardown before verifier execution. Commit28ebe0949fixed those paths.Seven cells were rerun. Four changed from fail to pass (
hf-model-inference,merge-diff-arc-agi-task,polyglot-c-py,regex-log), producing the final 65/89 result. Three remained genuine verifier failures.Verification
73, 69, 66, 65, 63, 60, 58, 53, 49c27c3bcbfc3ebe8e21cc250dc409f02f49ae055032eaf2f19fd6349986e94e6agit diff --checkpassedReview focus