Uh oh!
There was an error while loading. Please reload this page.
Actions: loop-eng/loop-bench
Actions
Showing runs from all workflows
22 workflow runs
22 workflow runs
Auto-format Python with ruff — all CI checks should pass now
CI
#16:
Commit 184dc43
pushed
by
rajfirke
Fix ruff: remove unused math import from test_analysis.py
CI
#14:
Commit a58637d
pushed
by
rajfirke
Fix CI: remove unused tempfile import, replace require() with top-lev…
CI
#13:
Commit a304406
pushed
by
rajfirke
Add bench init, fix 15 bugs from multi-pass audit (Phases 10+11)
CI
#12:
Commit aaafcfb
pushed
by
rajfirke
Add leaderboard static site with sortable table and filters
CI
#10:
Commit b1d8d30
pushed
by
rajfirke
Implement Python analysis tools: compare, statistics, visualize, lead…
CI
#9:
Commit 4e3f4ee
pushed
by
rajfirke
Add 20 feature/refactoring/multi-step tasks — 30 total benchmark suite
CI
#8:
Commit d3d9f05
pushed
by
rajfirke
Add 3 baseline loop adapters: minimal, reflexion, plan-first
CI
#7:
Commit a055972
pushed
by
rajfirke
Add metrics engine: 11 calculators, composite scoring, bootstrap CI
CI
#6:
Commit 11c9b73
pushed
by
rajfirke
Wire up harness core: CLI, loader, evaluator, LTF collector, runner p…
CI
#5:
Commit 5ca9522
pushed
by
rajfirke
Add 10 bug-fix benchmark tasks with synthetic repos and hidden tests
CI
#3:
Commit 124d546
pushed
by
rajfirke
Add spec + schemas: BENCHMARK.md, metrics.md, JSON Schemas, Ajv valid…
CI
#2:
Commit 72c9d41
pushed
by
rajfirke