Skip to content

🦀 PinchBench

Real-world benchmarks for AI coding agents

PinchBench measures how well LLM models perform as the brain of an OpenClaw agent. Instead of synthetic tests, we throw real tasks at agents: scheduling meetings, writing code, triaging email, researching topics, and managing files.


Repositories

RepoDescription
skillBenchmark runner and task definitions — run it yourself
leaderboardThe pinchbench.com leaderboard frontend
apiThe public PinchBench API at api.pinchbench.com
scriptsThe offical PinchBench run automation with default_models.yml

Run the Benchmark

git clone https://github.com/pinchbench/skill.git
cd skill
./scripts/run.sh --model anthropic/claude-sonnet-4

Results upload to the public leaderboard. Get started →


Claw-some AI agent testing. Made with 🦀 by the humans at https://kilo.ai 🦞

Popular repositories Loading

  1. skill skillPublic

    PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai

    Python 1.3k 151

  2. leaderboard leaderboardPublic

    PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai

    TypeScript 41 15

  3. api apiPublic

    Public API for pinchbench.com to display results. Also contains admin interface for managing PinchBench

    TypeScript 7 8

  4. scripts scriptsPublic

    Shell 2 8

  5. .github .githubPublic

    PinchBench organization profile and community health files

Repositories

Showing 5 of 5 repositories

Top languages

Loading…

Most used topics

Loading…