You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A Multi-Agent System (MAS) evaluation framework using PydanticAI that generates and evaluates scientific paper reviews through a three-tiered assessment approach: traditional metrics, LLM-as-a-Judge, and graph-based complexity analysis.
Fault-injecting OpenEnv training environment for vibe-coded SaaS incidents. 30 scenarios grounded in 2025-26 production failures. Drop-in OpenClaw-RL pool server. Claude Code skill included.
Winner of the tau2-bench competition on AgentBeats (Berkeley AgentX) at the competition close: a customer-service purple agent reaching 82.5% in the telecom domain. A plain A2A service with no agent framework; the edge is distilled per-domain policy playbooks injected ahead of the task policy.