We build software engineering agents, the benchmarks to evaluate them, and the models behind them — aiming for AI-Powered Software 3.0.
Prometheus — our SWE agent system
https://github.com/EuniAI/PrometheusContextBench — evaluate context retrieval vs. actual usage in coding agents
https://github.com/EuniAI/ContextBenchEnvAgent — generate reproducible Docker environments from environment specs
https://github.com/EuniAI/EnvAgentawesome-code-agents — a curated map of code agents, benchmarks, and tooling
https://github.com/EuniAI/awesome-code-agents
Issues and PRs are welcome. If you’re not sure where to start, open an issue with what you want to build or evaluate.