Evaluation-driven autonomous development: a harness for agent loops whose acceptance criterion is a metric, not a test suite. Held-out metrics, sealed scoring, integrity checks.
verificationautonomous-researchllm-agentsreward-hackingagent-harnessevaluation-driven-developmentloop-engineeringheld-out-validation
-
Updated
Jul 31, 2026 - Python