Uh oh!
There was an error while loading. Please reload this page.
docs: Terminal-Bench 2.1 Maka vs Kimi Code v11 results - #1217
Merged
Conversation
Astro-Hanforce-pushed
the
docs/tb21-maka-vs-kimi-code-v11-report
branch
from
July 18, 2026 17:49
df50c87 to
8a267c0CompareAstro-Hanforce-pushed
the
docs/tb21-maka-vs-kimi-code-v11-report
branch
2 times, most recently
from
July 18, 2026 17:54
e6c3442 to
785fb2bCompareAstro-Hanforce-pushed
the
docs/tb21-maka-vs-kimi-code-v11-report
branch
from
July 18, 2026 17:58
785fb2b to
df50c87CompareAstro-Hanforce-pushed
the
docs/tb21-maka-vs-kimi-code-v11-report
branch
4 times, most recently
from
July 18, 2026 19:23
d20cb7b to
d28d7f0CompareCapture the paired pass@1 outcome, aligned thinking settings, known harness asymmetries, and interpretation boundaries without committing eval artifacts.
Astro-Hanforce-pushed
the
docs/tb21-maka-vs-kimi-code-v11-report
branch
from
July 18, 2026 19:24
d28d7f0 to
f105202CompareUh oh!
There was an error while loading. Please reload this page.
This was referenced Jul 20, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for freeto join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Docs-only: English write-up of Terminal-Bench 2.1 harness A/B run
k3-maka-vs-kimi-code-tbench-2.1-full-v11(Maka vs Kimi Code, both k3 / max thinking), plus a redaction-minimal accepted-outcome CSV.Full report:
docs/eval/terminal-bench-2.1-maka-vs-kimi-code-v11.mdResults
Official reference: 88.3% for Kimi K3 + KimiCode harness (tech blog, accessed 2026-07-19).
Kimi Code's accepted 53/89 is operator-adjusted (mechanical first verifier = 50/89). Three score-changing recoveries:
prove-plus-comm,hf-model-inference,pypi-server. Run admissions: 203 physical → 178 accepted (25 recovery). Single-repetition; descriptive, not statistical.Absolute scores are throughput-limited
Decode was ~34–37 output tok/s with ~14 s median first-token latency against task-native agent budgets. About a third of tasks per arm hit the deadline mid-generation and rarely passed (14–15%).
Finished-in-time pass rates (95.1% / 85.7%; 95.9% / 89.8% on the 49 tasks both arms finished) sit around the official 88.3%. Under healthy throughput, Maka would plausibly reach or exceed that reference and Kimi Code land near it — subsets are self-selected, so a re-run is decisive.
Official frontier is already compressed (~88.8% / 88.3%). Unconstrained subsets pass at 85.7–95.9%. TB 2.1 is close to saturated at the frontier; harder benchmarks will carry more signal.
Relative gap is harness asymmetry
Thinking effort, thinking retention, and wall-clock are aligned (same degraded network path). The +10.1 pp (Maka) persists among finished tasks (95.1% vs 85.7%). Differences below are listed, not ranked as score causes.
Background Bash is not the exclusive-win story (2/12 Maka-only tasks used
run_in_backgroundon Kimi). Heavy-task self-check is in the build but was off (zero tool/gate events). Recorded cost is $0 (account-plan); list-price uses official K3 API rates.Verification
controller/results.jsonl{,.attempts.jsonl}, trialresult.json, Kimiprovider-request-telemetry.json, Makamaka-cell-output.json/ runtime streams,harness-ab-report.json, tasktask.tomlbudgetsgit diff --checkclean