diff --git a/README.md b/README.md index c61d6e7..a5a0a1c 100644 --- a/README.md +++ b/README.md @@ -2,194 +2,222 @@ [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](LICENSE) [![NIST AI RMF](https://img.shields.io/badge/NIST%20AI%20RMF-Informed-0055A4?style=flat-square)](https://airc.nist.gov/home) -[![Discussions](https://img.shields.io/badge/Discussions-Join-7289da?style=flat-square&logo=github)](https://github.com/simaba/agent-eval/discussions) -A structured framework for evaluating AI agents across task performance, safety, reliability, and governance dimensions, designed for regulated-industry deployment. +A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score. ---- +The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds. -## Choose this repo when +## The central idea -Use this repository when you need **measurement and pass/fail logic** for AI agents: +An evaluation result is only useful when five things are explicit: -- what to measure -- which scenarios to test -- how to score performance, safety, and reliability -- what should appear in an evaluation report -- how to structure evaluation evidence for human review and future automation +1. **the decision it supports** — prototype, compare, pilot, expand, hold, roll back, or retire; +2. **the operating population it represents** — users, tasks, languages, tools, permissions, and exclusions; +3. **the measurement instrument** — executable rule, expert review, user outcome, or model judge; +4. **the uncertainty and coverage limits** — sample size, evaluator error, slice gaps, version drift, and scenario bias; +5. **the decision semantics** — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk. -Do **not** start here if you need the oversight model. Use [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance). +Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures. -Do **not** start here if you need orchestration patterns. Use [`agent-orchestration`](https://github.com/simaba/agent-orchestration). +## Start here -Do **not** start here if you want runnable behavior. Use [`agent-simulator`](https://github.com/simaba/agent-simulator). +| Artifact | Use it for | +|---|---| +| [`docs/evaluation-validity.md`](docs/evaluation-validity.md) | Designing evidence that can support a stated decision | +| [`docs/decision-semantics.md`](docs/decision-semantics.md) | Keeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct | +| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human review and sign-off discussion | +| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable evaluation records | +| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Fictional structured example | +| [`tools/validate_evaluation_report.py`](tools/validate_evaluation_report.py) | Schema and decision-coherence validation | ---- +## What the framework evaluates -## Maturity +The dimensions below are categories for designing measurements. They are not automatically valid metrics. -This is a **practitioner evaluation framework**, not a certified benchmark suite. The metrics and targets are intended as starting points for structured evaluation. Real deployments should adapt thresholds to the domain, risk tier, user population, applicable regulation, and operational failure tolerance. +### Task and outcome performance ---- +- end-to-end task completion; +- correctness against a defined reference or acceptance rule; +- user or workflow outcome; +- tool selection and parameter accuracy; +- correction and recovery behavior; +- unnecessary step or interaction cost. + +### Safety, privacy, and policy behavior + +- prohibited action or tool invocation; +- sensitive-data exposure or inappropriate retention; +- refusal, deferral, and escalation at defined boundaries; +- prompt-injection and untrusted-content handling; +- behavior under conflicting instructions; +- containment after a control or dependency failure. + +### Operational reliability -## Practical artifacts +- latency and timeout behavior under stated load; +- unhandled and handled failure rates; +- retry, fallback, and circuit-breaker behavior; +- dependency and tool degradation; +- cost and resource consumption; +- recovery time and state consistency. -| Artifact | Use for | +### Governance and auditability + +- decision and tool-call provenance; +- version and permission traceability; +- human override and stop mechanisms; +- evidence retention with privacy controls; +- owner, review, and residual-risk records; +- regression triggers after material change. + +## Metrics need operational definitions + +Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining. + +For each metric, record: + +| Field | Example question | |---|---| -| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human-readable agent evaluation report | -| [`examples/sample-evaluation-report.md`](examples/sample-evaluation-report.md) | Filled generic Markdown example | -| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable report schema for structured evaluation evidence | -| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Filled JSON example aligned to the schema | +| Construct | What property are we trying to understand? | +| Observable rule | What exactly counts as success or failure? | +| Unit | Response, task, conversation, tool call, user, or time window? | +| Denominator | Which attempts are included, excluded, or missing? | +| Evaluator | Rule, test, expert, user, or model judge? | +| Instrument error | How can the evaluator itself be wrong? | +| Decision use | Which action changes if the metric moves? | ---- +Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable. -## Evaluation dimensions +## Evaluation suite design -### 1. Task performance +A credible suite usually separates several purposes: -| Metric | Description | Measurement Method | -|---|---|---| -| **Task Completion Rate** | % of tasks completed without error | automated test suite against benchmark scenarios | -| **Goal Achievement Rate** | % of tasks where the final goal was met | human or LLM-judge evaluation | -| **Step Efficiency** | average steps taken vs. optimal path | comparison to expert baseline | -| **Tool Use Accuracy** | correct tool selection and parameter passing | automated verification | -| **Output Quality** | correctness and usefulness of agent outputs | rubric-based judging | +| Suite | Primary purpose | +|---|---| +| Operating-distribution sample | Estimate routine performance under stated workload assumptions | +| Critical-risk suite | Oversample rare, high-consequence conditions and hard gates | +| Boundary suite | Test ambiguity, refusal, escalation, and policy transitions | +| Fault-injection suite | Exercise tool, retrieval, network, permission, and state failures | +| Regression suite | Preserve previously discovered failures and important invariants | +| Exploratory suite | Discover new slices and failure modes; not used alone for release claims | -### 2. Safety and reliability +A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately. -| Metric | Description | Example target | -|---|---|---| -| **Hallucination Rate** | factually incorrect statements in outputs | < 2% on factual tasks, adjusted by risk tier | -| **Harmful Output Rate** | outputs violating safety guidelines | 0% hard limit for critical harms | -| **Refusal Appropriateness** | correct refusal of inappropriate requests | > 95%, adjusted by policy scope | -| **Prompt Injection Resistance** | resistance to adversarial inputs | tested via red-team suite | -| **Consistency** | same input produces the same output class | > 90% across repeated runs, adjusted by task type | +## Evaluator hierarchy -### 3. Operational reliability +Use the simplest evaluator capable of measuring the property. -| Metric | Description | Example target | -|---|---|---| -| **Availability** | uptime under normal and peak load | ≥ 99.5%, adjusted by service criticality | -| **P99 Latency** | 99th-percentile response time | ≤ configured SLA | -| **Error Rate** | % of requests ending in unhandled errors | < 0.1%, adjusted by workload | -| **Cost per Task** | token cost per successful task | tracked against budget | -| **Graceful Degradation** | behavior when tools or dependencies fail | tested via fault injection | +1. **Executable checks** for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior. +2. **Qualified human review** for domain judgment, consequential ambiguity, and disputed failures. +3. **Model judges** for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects. -### 4. Governance and auditability +A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process. -| Metric | Description | Pass Criterion | -|---|---|---| -| **Decision Traceability** | can every output be traced to inputs and decision evidence? | full trace available | -| **Tool Call Logging** | all external tool calls logged with parameters | 100% logged, with sensitive data controls | -| **Human Escalation Rate** | % of tasks escalated to human review | tracked with threshold defined | -| **Sensitive Data Handling** | PII or sensitive data not leaked in outputs or logs | 0 critical violations | -| **Override Capability** | operator can halt or override agent behavior | verified in testing | +## Repeated runs and uncertainty ---- +For nondeterministic systems, report: -## Evaluation protocol +- attempts per scenario; +- run and seed policy where applicable; +- model, prompt, retrieval, tool, permission, and environment versions; +- missing or failed runs; +- scenario-level recurrence and variability; +- uncertainty intervals for estimated rates or means; +- coverage gaps and evaluator limitations that statistical intervals do not capture. -The target numbers above should not be treated as universal pass/fail thresholds. Use this protocol to make evaluations more defensible. +“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk. -| Protocol element | Recommended practice | -|---|---| -| Scenario set | Include normal, ambiguous, adversarial, tool-failure, escalation, and out-of-distribution tasks | -| Sample size | Define the minimum number of tasks per scenario class before testing starts | -| Risk tiering | Use stricter thresholds for safety-critical, regulated, or irreversible decisions | -| Repeated runs | Run the same scenario multiple times when model nondeterminism matters | -| Human adjudication | Use expert review for high-impact failures, disputed LLM-judge results, and policy-boundary cases | -| Severity levels | Classify findings as minor, major, critical, or blocker | -| Confidence | Report uncertainty, small sample caveats, and known benchmark limitations | -| Regression testing | Re-run the suite after model, prompt, tool, retrieval, or policy changes | -| Evidence retention | Store prompts, inputs, outputs, tool traces, evaluator notes, and sign-off decisions with appropriate data controls | - -### Suggested release decision logic - -| Decision | Use when | -|---|---| -| **Approve** | No blockers, critical controls pass, residual risks accepted by named owners | -| **Approve with conditions** | Minor or moderate issues exist but mitigations and monitoring are defined | -| **Hold** | Major issues require fixes before release | -| **Reject** | Critical safety, privacy, compliance, or operational failures are unresolved | +## Slice analysis ---- +Predefine decision-relevant slices such as: -## Evaluation scenarios +- language or locale; +- task and user type; +- ambiguity level; +- tool authority and permission scope; +- data sensitivity; +- retrieval quality; +- accessibility need; +- dependency failure mode; +- safety or policy boundary. -### Standard benchmark suite +Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing. -| Scenario | Description | Difficulty | -|---|---|---| -| Single-step task | straightforward, well-defined task | Basic | -| Multi-step task | sequence of actions with intermediate dependencies | Intermediate | -| Ambiguous instruction | underspecified request requiring clarification | Intermediate | -| Tool failure | dependent tool returns an error | Intermediate | -| Adversarial input | prompt injection or manipulation attempt | Advanced | -| Novel situation | task outside the usual distribution | Advanced | -| Cascading failure | multiple dependencies fail in sequence | Advanced | -| Human escalation | agent should recognize the need to escalate | Advanced | +## Threshold and gate design -### Regulated-industry scenarios +This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective. -| Industry | Scenario | Key Evaluation Criteria | -|---|---|---| -| Healthcare | clinical documentation assistant | factual accuracy, refusal quality, low hallucination | -| Finance | transaction-analysis agent | compliance, auditability, no PII leakage | -| Insurance | claims-assessment agent | fairness, traceability, escalation quality | -| Legal | contract-review agent | citation accuracy, scope awareness, human escalation | +A threshold should record: ---- +- decision owner and intended decision; +- baseline or comparator; +- affected population and harm model; +- measurement method and uncertainty; +- minimum sample and slice coverage; +- non-compensable hard gates; +- quality or operational targets that may be optimized; +- exception and residual-risk process; +- expiry or review trigger. -## Human-readable report template +A weighted aggregate must never hide a failed hard gate or an unresolved critical finding. -Use [`templates/evaluation-report.md`](templates/evaluation-report.md) for review meetings, sign-off packs, and stakeholder alignment. +## Decision semantics -## Machine-readable report schema +The schema and validator keep these concepts separate: -Use [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) when you want evaluation results to be structured for later validation, CI checks, dashboards, or trend analysis. +| Concept | Meaning | +|---|---| +| System risk tier | impact context for the evaluated system and use case | +| Scenario / finding severity | consequence of a specific failure | +| Blocker | unresolved condition that prevents the current decision | +| Required action | follow-up work accepted under a bounded conditional decision | +| Condition | explicit constraint attached to that decision | +| Residual risk | remaining uncertainty or harm accepted by a named authority | -The schema captures: +See [`docs/decision-semantics.md`](docs/decision-semantics.md) for the validator rules. -- metadata -- recommendation and blockers -- scope -- scenarios -- scorecard results -- findings -- oversight triggers -- release decision and sign-off +## Validation ---- +The validator checks schema conformance and selected decision coherence: -## NIST AI RMF mapping +```bash +python tools/validate_evaluation_report.py \ + schemas/evaluation-report.schema.json \ + examples/sample-evaluation-report.json +``` -Full mapping: [docs/nist-rmf-mapping.md](docs/nist-rmf-mapping.md) +It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable. -Key practitioner mapping: +## Failure review -- **MS.1**: AI risk identification through safety and adversarial scenarios -- **MS.3**: structured evaluation techniques and benchmark design -- **MS.5**: subgroup and fairness measurement -- **GV.4**: human oversight via escalation and override checks +Do not stop at the scorecard. For material failures, preserve: ---- +- scenario and run identifiers; +- relevant input, context, and tool trace; +- evaluator result and disagreement; +- severity rationale; +- containment and remediation; +- root-cause hypothesis versus confirmed cause; +- regression test created; +- owner and disposition. -## Scope and disclaimer +The failure corpus often becomes more valuable than the original aggregate score. -This repository is shared in a personal capacity. It is not legal advice, compliance certification, regulatory approval, safety certification, benchmark certification, or official guidance from NIST, the EU, ISO, or any employer. +## Maturity and scope -References to NIST AI RMF, EU AI Act, safety controls, or industry obligations are practitioner mappings and examples. Always verify against official sources before using this framework for compliance, safety, or release decisions. +This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority. ---- +References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance. ## Related repositories -| Repository | What it adds | +| Repository | Distinct role | |---|---| -| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | oversight model, trust structure, accountability controls | -| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns and orchestration logic | -| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable execution and failure-mode examples | -| [`nist-rmf-guide`](https://github.com/simaba/nist-rmf-guide) | practitioner guide for broader RMF implementation | -| [`governance-playbook`](https://github.com/simaba/governance-playbook) | enterprise operating model beyond agent evaluation | +| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable bounded-agent behavior and failure paths | +| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns | +| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | authority, oversight, containment, and accountability | +| [`automotive-llm-eval-harness`](https://github.com/simaba/automotive-llm-eval-harness) | compact scorer for synthetic automotive evaluation artifacts | + +--- -*Maintained by [Sima Bagheri](https://github.com/simaba) · Connect on [LinkedIn](https://www.linkedin.com/in/simaba/)* +*Maintained by [Sima Bagheri](https://github.com/simaba).* diff --git a/docs/evaluation-validity.md b/docs/evaluation-validity.md new file mode 100644 index 0000000..f800dec --- /dev/null +++ b/docs/evaluation-validity.md @@ -0,0 +1,208 @@ +# Evaluation Validity and Evidence Design + +An evaluation can be internally consistent and still support the wrong conclusion. This guide describes the evidence questions that should be answered before a score is treated as a product, safety, or release signal. + +## Start from the decision + +Define the decision the evaluation is intended to support: + +- continue prototyping; +- compare two configurations; +- enter a bounded pilot; +- expand a pilot population; +- approve a particular tool permission; +- hold or roll back a release; +- retire a model, prompt, tool, or workflow. + +A benchmark without a decision context often accumulates convenient metrics rather than decision-relevant evidence. + +## Define the unit of analysis + +State exactly what one observation represents: + +- one model response; +- one end-to-end task attempt; +- one user journey; +- one tool invocation; +- one conversation; +- one incident window; +- one repeated run of a fixed scenario. + +Do not mix units in one denominator. For example, “failure rate” is ambiguous when some records are messages, others are tasks, and others are conversations. + +## Specify the target population + +A scenario set is a sample from an intended operating environment. Document: + +- user groups and languages; +- task categories and frequency assumptions; +- tool and data-access conditions; +- normal, ambiguous, adversarial, failure, and out-of-distribution cases; +- exclusions; +- known differences between the evaluation sample and expected use. + +A balanced benchmark can be useful for finding failures but may not estimate production frequency. A production-shaped sample can estimate prevalence but may underrepresent rare critical cases. Use separate suites when those purposes conflict. + +## Separate coverage from prevalence + +Report both: + +1. **coverage:** which risk, task, and failure classes were exercised; +2. **prevalence-weighted performance:** results under an explicit operating-distribution assumption. + +Do not average a rare catastrophic scenario with routine tasks and call the result “overall safety.” Hard gates, scenario-level reporting, and residual-risk review are more appropriate for non-compensable harms. + +## Operationalize each metric + +For every metric, record: + +| Field | Question | +|---|---| +| Construct | What property are we trying to understand? | +| Operational definition | What observable event counts as success or failure? | +| Unit | Response, task, conversation, user, run, or time window? | +| Denominator | Which attempted cases are included or excluded? | +| Evaluator | Rule, test, expert, user, or model judge? | +| Error modes | How can the measurement itself be wrong? | +| Decision use | Which decision changes if the metric moves? | + +Labels such as “hallucination,” “quality,” “helpfulness,” and “safety” are not measurements until their decision rules are specified. + +## Evaluator validity + +### Deterministic checks + +Use executable checks for properties that can be observed directly: schema validity, prohibited tool calls, citation presence, exact constraints, timeout behavior, or unauthorized state changes. + +A deterministic check can be precise and still be incomplete. Passing it does not imply semantic correctness. + +### Human review + +Define reviewer qualifications, instructions, adjudication, and blinding where relevant. Measure agreement on a representative sample before interpreting reviewer labels as stable ground truth. + +Raw agreement alone can be misleading when one label dominates. Report the confusion pattern and adjudication rate, not only a single coefficient. + +### Model judges + +A model judge is a measurement instrument, not an authority. Before using it for consequential conclusions: + +- define the rubric independently of the judged outputs; +- test order, verbosity, style, and self-preference effects; +- compare against qualified human review on representative and difficult cases; +- record judge model, prompt, configuration, and version; +- preserve disagreement and adjudication rather than silently averaging it away; +- avoid using the same model family as generator and sole judge when correlated errors matter. + +Do not describe an LLM-judge score as objective or ground truth. + +## Repeated runs and dependence + +Agent behavior can vary across repeated attempts, but repeated outputs from the same scenario are not automatically independent observations. + +Record: + +- run count per scenario; +- sampling and seed policy where applicable; +- model, prompt, retrieval, tools, permissions, and environment versions; +- cache and state behavior; +- whether failures cluster by scenario, user journey, or dependency. + +Report scenario-level variability and worst-case recurrence, not only the mean across all attempts. + +## Uncertainty + +Use confidence intervals or other uncertainty summaries when estimating a rate or mean from a sample. Also report sources of uncertainty that statistical intervals do not capture: + +- scenario-selection bias; +- incomplete coverage; +- evaluator error; +- distribution shift; +- version drift; +- hidden tool or retrieval state; +- small counts for critical slices. + +A narrow interval around a biased benchmark is still misleading. + +For a zero-observed-failure result, report the sample count and an upper bound or plain-language limitation. “0 observed in 50 cases” is not evidence of zero risk. + +## Threshold design + +Thresholds should come from the decision context, not from a generic table. + +A threshold record should include: + +- decision and owner; +- affected population and harm model; +- baseline or comparator; +- measurement method and uncertainty; +- hard gates that cannot be traded off; +- quality or operational thresholds that can be optimized; +- minimum sample and slice requirements; +- exception and residual-risk process; +- expiry or review trigger. + +Avoid universal targets such as “hallucination below 2%” without defining the task distribution, severity, evaluator, and confidence required. + +## Slice analysis + +Predefine slices that can reveal concentrated failure: + +- language or locale; +- user or task type; +- ambiguity level; +- tool permission; +- data sensitivity; +- retrieval quality; +- input length; +- dependency failure mode; +- safety or policy boundary; +- accessibility need. + +Do not search hundreds of slices after the fact and report only the worst or best without disclosing the search. Exploratory slice discovery should lead to a versioned confirmatory test set. + +## Comparison design + +For model, prompt, or harness comparisons: + +- keep task, tools, permissions, and evaluator conditions equivalent; +- use paired analysis when the same scenarios are attempted by both systems; +- randomize presentation order for subjective judging; +- report missing and failed runs rather than dropping them; +- distinguish statistical difference from operational significance; +- account for multiple comparisons when many variants are tested; +- preserve version and cost context. + +A leaderboard number is not meaningful when the systems received different authority, context, or recovery paths. + +## Failure review + +Aggregate scores are navigation aids. Review failures as evidence. + +For material failures, preserve: + +- scenario and run identifiers; +- input and relevant context; +- tool trace and system state; +- evaluator result and disagreement; +- failure taxonomy and severity rationale; +- containment or remediation; +- regression test created; +- owner and disposition. + +Separate root-cause hypotheses from confirmed causes. + +## Release evidence package + +A defensible evaluation package should make it possible for another reviewer to answer: + +1. What system and version were tested? +2. What decision was being supported? +3. What operating population did the scenarios represent? +4. Which critical conditions were intentionally oversampled? +5. How were outcomes measured and how reliable were the evaluators? +6. What uncertainty and exclusions remain? +7. Which hard gates failed or passed? +8. Which residual risks were accepted, by whom, and under what conditions? +9. What changes invalidate the result and trigger regression testing? + +The package supports accountable judgment. It does not automate release authority or prove safety.