Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
296 changes: 162 additions & 134 deletions README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -2,194 +2,222 @@

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](LICENSE)
[![NIST AI RMF](https://img.shields.io/badge/NIST%20AI%20RMF-Informed-0055A4?style=flat-square)](https://airc.nist.gov/home)
[![Discussions](https://img.shields.io/badge/Discussions-Join-7289da?style=flat-square&logo=github)](https://github.com/simaba/agent-eval/discussions)

A structured framework for evaluating AI agents across task performance, safety, reliability, and governance dimensions, designed for regulated-industry deployment.
A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score.

---
The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds.

## Choose this repo when
## The central idea

Use this repository when you need **measurement and pass/fail logic** for AI agents:
An evaluation result is only useful when five things are explicit:

- what to measure
- which scenarios to test
- how to score performance, safety, and reliability
- what should appear in an evaluation report
- how to structure evaluation evidence for human review and future automation
1. **the decision it supports** — prototype, compare, pilot, expand, hold, roll back, or retire;
2. **the operating population it represents** — users, tasks, languages, tools, permissions, and exclusions;
3. **the measurement instrument** — executable rule, expert review, user outcome, or model judge;
4. **the uncertainty and coverage limits** — sample size, evaluator error, slice gaps, version drift, and scenario bias;
5. **the decision semantics** — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk.

Do **not** start here if you need the oversight model. Use [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance).
Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures.

Do **not** start here if you need orchestration patterns. Use [`agent-orchestration`](https://github.com/simaba/agent-orchestration).
## Start here

Do **not** start here if you want runnable behavior. Use [`agent-simulator`](https://github.com/simaba/agent-simulator).
| Artifact | Use it for |
|---|---|
| [`docs/evaluation-validity.md`](docs/evaluation-validity.md) | Designing evidence that can support a stated decision |
| [`docs/decision-semantics.md`](docs/decision-semantics.md) | Keeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct |
| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human review and sign-off discussion |
| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable evaluation records |
| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Fictional structured example |
| [`tools/validate_evaluation_report.py`](tools/validate_evaluation_report.py) | Schema and decision-coherence validation |

---
## What the framework evaluates

## Maturity
The dimensions below are categories for designing measurements. They are not automatically valid metrics.

This is a **practitioner evaluation framework**, not a certified benchmark suite. The metrics and targets are intended as starting points for structured evaluation. Real deployments should adapt thresholds to the domain, risk tier, user population, applicable regulation, and operational failure tolerance.
### Task and outcome performance

---
- end-to-end task completion;
- correctness against a defined reference or acceptance rule;
- user or workflow outcome;
- tool selection and parameter accuracy;
- correction and recovery behavior;
- unnecessary step or interaction cost.

### Safety, privacy, and policy behavior

- prohibited action or tool invocation;
- sensitive-data exposure or inappropriate retention;
- refusal, deferral, and escalation at defined boundaries;
- prompt-injection and untrusted-content handling;
- behavior under conflicting instructions;
- containment after a control or dependency failure.

### Operational reliability

## Practical artifacts
- latency and timeout behavior under stated load;
- unhandled and handled failure rates;
- retry, fallback, and circuit-breaker behavior;
- dependency and tool degradation;
- cost and resource consumption;
- recovery time and state consistency.

| Artifact | Use for |
### Governance and auditability

- decision and tool-call provenance;
- version and permission traceability;
- human override and stop mechanisms;
- evidence retention with privacy controls;
- owner, review, and residual-risk records;
- regression triggers after material change.

## Metrics need operational definitions

Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining.

For each metric, record:

| Field | Example question |
|---|---|
| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human-readable agent evaluation report |
| [`examples/sample-evaluation-report.md`](examples/sample-evaluation-report.md) | Filled generic Markdown example |
| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable report schema for structured evaluation evidence |
| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Filled JSON example aligned to the schema |
| Construct | What property are we trying to understand? |
| Observable rule | What exactly counts as success or failure? |
| Unit | Response, task, conversation, tool call, user, or time window? |
| Denominator | Which attempts are included, excluded, or missing? |
| Evaluator | Rule, test, expert, user, or model judge? |
| Instrument error | How can the evaluator itself be wrong? |
| Decision use | Which action changes if the metric moves? |

---
Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable.

## Evaluation dimensions
## Evaluation suite design

### 1. Task performance
A credible suite usually separates several purposes:

| Metric | Description | Measurement Method |
|---|---|---|
| **Task Completion Rate** | % of tasks completed without error | automated test suite against benchmark scenarios |
| **Goal Achievement Rate** | % of tasks where the final goal was met | human or LLM-judge evaluation |
| **Step Efficiency** | average steps taken vs. optimal path | comparison to expert baseline |
| **Tool Use Accuracy** | correct tool selection and parameter passing | automated verification |
| **Output Quality** | correctness and usefulness of agent outputs | rubric-based judging |
| Suite | Primary purpose |
|---|---|
| Operating-distribution sample | Estimate routine performance under stated workload assumptions |
| Critical-risk suite | Oversample rare, high-consequence conditions and hard gates |
| Boundary suite | Test ambiguity, refusal, escalation, and policy transitions |
| Fault-injection suite | Exercise tool, retrieval, network, permission, and state failures |
| Regression suite | Preserve previously discovered failures and important invariants |
| Exploratory suite | Discover new slices and failure modes; not used alone for release claims |

### 2. Safety and reliability
A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately.

| Metric | Description | Example target |
|---|---|---|
| **Hallucination Rate** | factually incorrect statements in outputs | < 2% on factual tasks, adjusted by risk tier |
| **Harmful Output Rate** | outputs violating safety guidelines | 0% hard limit for critical harms |
| **Refusal Appropriateness** | correct refusal of inappropriate requests | > 95%, adjusted by policy scope |
| **Prompt Injection Resistance** | resistance to adversarial inputs | tested via red-team suite |
| **Consistency** | same input produces the same output class | > 90% across repeated runs, adjusted by task type |
## Evaluator hierarchy

### 3. Operational reliability
Use the simplest evaluator capable of measuring the property.

| Metric | Description | Example target |
|---|---|---|
| **Availability** | uptime under normal and peak load | ≥ 99.5%, adjusted by service criticality |
| **P99 Latency** | 99th-percentile response time | ≤ configured SLA |
| **Error Rate** | % of requests ending in unhandled errors | < 0.1%, adjusted by workload |
| **Cost per Task** | token cost per successful task | tracked against budget |
| **Graceful Degradation** | behavior when tools or dependencies fail | tested via fault injection |
1. **Executable checks** for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior.
2. **Qualified human review** for domain judgment, consequential ambiguity, and disputed failures.
3. **Model judges** for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects.

### 4. Governance and auditability
A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process.

| Metric | Description | Pass Criterion |
|---|---|---|
| **Decision Traceability** | can every output be traced to inputs and decision evidence? | full trace available |
| **Tool Call Logging** | all external tool calls logged with parameters | 100% logged, with sensitive data controls |
| **Human Escalation Rate** | % of tasks escalated to human review | tracked with threshold defined |
| **Sensitive Data Handling** | PII or sensitive data not leaked in outputs or logs | 0 critical violations |
| **Override Capability** | operator can halt or override agent behavior | verified in testing |
## Repeated runs and uncertainty

---
For nondeterministic systems, report:

## Evaluation protocol
- attempts per scenario;
- run and seed policy where applicable;
- model, prompt, retrieval, tool, permission, and environment versions;
- missing or failed runs;
- scenario-level recurrence and variability;
- uncertainty intervals for estimated rates or means;
- coverage gaps and evaluator limitations that statistical intervals do not capture.

The target numbers above should not be treated as universal pass/fail thresholds. Use this protocol to make evaluations more defensible.
“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk.

| Protocol element | Recommended practice |
|---|---|
| Scenario set | Include normal, ambiguous, adversarial, tool-failure, escalation, and out-of-distribution tasks |
| Sample size | Define the minimum number of tasks per scenario class before testing starts |
| Risk tiering | Use stricter thresholds for safety-critical, regulated, or irreversible decisions |
| Repeated runs | Run the same scenario multiple times when model nondeterminism matters |
| Human adjudication | Use expert review for high-impact failures, disputed LLM-judge results, and policy-boundary cases |
| Severity levels | Classify findings as minor, major, critical, or blocker |
| Confidence | Report uncertainty, small sample caveats, and known benchmark limitations |
| Regression testing | Re-run the suite after model, prompt, tool, retrieval, or policy changes |
| Evidence retention | Store prompts, inputs, outputs, tool traces, evaluator notes, and sign-off decisions with appropriate data controls |

### Suggested release decision logic

| Decision | Use when |
|---|---|
| **Approve** | No blockers, critical controls pass, residual risks accepted by named owners |
| **Approve with conditions** | Minor or moderate issues exist but mitigations and monitoring are defined |
| **Hold** | Major issues require fixes before release |
| **Reject** | Critical safety, privacy, compliance, or operational failures are unresolved |
## Slice analysis

---
Predefine decision-relevant slices such as:

## Evaluation scenarios
- language or locale;
- task and user type;
- ambiguity level;
- tool authority and permission scope;
- data sensitivity;
- retrieval quality;
- accessibility need;
- dependency failure mode;
- safety or policy boundary.

### Standard benchmark suite
Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing.

| Scenario | Description | Difficulty |
|---|---|---|
| Single-step task | straightforward, well-defined task | Basic |
| Multi-step task | sequence of actions with intermediate dependencies | Intermediate |
| Ambiguous instruction | underspecified request requiring clarification | Intermediate |
| Tool failure | dependent tool returns an error | Intermediate |
| Adversarial input | prompt injection or manipulation attempt | Advanced |
| Novel situation | task outside the usual distribution | Advanced |
| Cascading failure | multiple dependencies fail in sequence | Advanced |
| Human escalation | agent should recognize the need to escalate | Advanced |
## Threshold and gate design

### Regulated-industry scenarios
This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective.

| Industry | Scenario | Key Evaluation Criteria |
|---|---|---|
| Healthcare | clinical documentation assistant | factual accuracy, refusal quality, low hallucination |
| Finance | transaction-analysis agent | compliance, auditability, no PII leakage |
| Insurance | claims-assessment agent | fairness, traceability, escalation quality |
| Legal | contract-review agent | citation accuracy, scope awareness, human escalation |
A threshold should record:

---
- decision owner and intended decision;
- baseline or comparator;
- affected population and harm model;
- measurement method and uncertainty;
- minimum sample and slice coverage;
- non-compensable hard gates;
- quality or operational targets that may be optimized;
- exception and residual-risk process;
- expiry or review trigger.

## Human-readable report template
A weighted aggregate must never hide a failed hard gate or an unresolved critical finding.

Use [`templates/evaluation-report.md`](templates/evaluation-report.md) for review meetings, sign-off packs, and stakeholder alignment.
## Decision semantics

## Machine-readable report schema
The schema and validator keep these concepts separate:

Use [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) when you want evaluation results to be structured for later validation, CI checks, dashboards, or trend analysis.
| Concept | Meaning |
|---|---|
| System risk tier | impact context for the evaluated system and use case |
| Scenario / finding severity | consequence of a specific failure |
| Blocker | unresolved condition that prevents the current decision |
| Required action | follow-up work accepted under a bounded conditional decision |
| Condition | explicit constraint attached to that decision |
| Residual risk | remaining uncertainty or harm accepted by a named authority |

The schema captures:
See [`docs/decision-semantics.md`](docs/decision-semantics.md) for the validator rules.

- metadata
- recommendation and blockers
- scope
- scenarios
- scorecard results
- findings
- oversight triggers
- release decision and sign-off
## Validation

---
The validator checks schema conformance and selected decision coherence:

## NIST AI RMF mapping
```bash
python tools/validate_evaluation_report.py \
schemas/evaluation-report.schema.json \
examples/sample-evaluation-report.json
```

Full mapping: [docs/nist-rmf-mapping.md](docs/nist-rmf-mapping.md)
It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable.

Key practitioner mapping:
## Failure review

- **MS.1**: AI risk identification through safety and adversarial scenarios
- **MS.3**: structured evaluation techniques and benchmark design
- **MS.5**: subgroup and fairness measurement
- **GV.4**: human oversight via escalation and override checks
Do not stop at the scorecard. For material failures, preserve:

---
- scenario and run identifiers;
- relevant input, context, and tool trace;
- evaluator result and disagreement;
- severity rationale;
- containment and remediation;
- root-cause hypothesis versus confirmed cause;
- regression test created;
- owner and disposition.

## Scope and disclaimer
The failure corpus often becomes more valuable than the original aggregate score.

This repository is shared in a personal capacity. It is not legal advice, compliance certification, regulatory approval, safety certification, benchmark certification, or official guidance from NIST, the EU, ISO, or any employer.
## Maturity and scope

References to NIST AI RMF, EU AI Act, safety controls, or industry obligations are practitioner mappings and examples. Always verify against official sources before using this framework for compliance, safety, or release decisions.
This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority.

---
References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance.

## Related repositories

| Repository | What it adds |
| Repository | Distinct role |
|---|---|
| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | oversight model, trust structure, accountability controls |
| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns and orchestration logic |
| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable execution and failure-mode examples |
| [`nist-rmf-guide`](https://github.com/simaba/nist-rmf-guide) | practitioner guide for broader RMF implementation |
| [`governance-playbook`](https://github.com/simaba/governance-playbook) | enterprise operating model beyond agent evaluation |
| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable bounded-agent behavior and failure paths |
| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns |
| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | authority, oversight, containment, and accountability |
| [`automotive-llm-eval-harness`](https://github.com/simaba/automotive-llm-eval-harness) | compact scorer for synthetic automotive evaluation artifacts |

---

*Maintained by [Sima Bagheri](https://github.com/simaba) · Connect on [LinkedIn](https://www.linkedin.com/in/simaba/)*
*Maintained by [Sima Bagheri](https://github.com/simaba).*
Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
296 changes: 162 additions & 134 deletions README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -2,194 +2,222 @@

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](LICENSE)
[![NIST AI RMF](https://img.shields.io/badge/NIST%20AI%20RMF-Informed-0055A4?style=flat-square)](https://airc.nist.gov/home)
[![Discussions](https://img.shields.io/badge/Discussions-Join-7289da?style=flat-square&logo=github)](https://github.com/simaba/agent-eval/discussions)

A structured framework for evaluating AI agents across task performance, safety, reliability, and governance dimensions, designed for regulated-industry deployment.
A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score.

---
The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds.

## Choose this repo when
## The central idea

Use this repository when you need **measurement and pass/fail logic** for AI agents:
An evaluation result is only useful when five things are explicit:

- what to measure
- which scenarios to test
- how to score performance, safety, and reliability
- what should appear in an evaluation report
- how to structure evaluation evidence for human review and future automation
1. **the decision it supports** — prototype, compare, pilot, expand, hold, roll back, or retire;
2. **the operating population it represents** — users, tasks, languages, tools, permissions, and exclusions;
3. **the measurement instrument** — executable rule, expert review, user outcome, or model judge;
4. **the uncertainty and coverage limits** — sample size, evaluator error, slice gaps, version drift, and scenario bias;
5. **the decision semantics** — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk.

Do **not** start here if you need the oversight model. Use [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance).
Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures.

Do **not** start here if you need orchestration patterns. Use [`agent-orchestration`](https://github.com/simaba/agent-orchestration).
## Start here

Do **not** start here if you want runnable behavior. Use [`agent-simulator`](https://github.com/simaba/agent-simulator).
| Artifact | Use it for |
|---|---|
| [`docs/evaluation-validity.md`](docs/evaluation-validity.md) | Designing evidence that can support a stated decision |
| [`docs/decision-semantics.md`](docs/decision-semantics.md) | Keeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct |
| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human review and sign-off discussion |
| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable evaluation records |
| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Fictional structured example |
| [`tools/validate_evaluation_report.py`](tools/validate_evaluation_report.py) | Schema and decision-coherence validation |

---
## What the framework evaluates

## Maturity
The dimensions below are categories for designing measurements. They are not automatically valid metrics.

This is a **practitioner evaluation framework**, not a certified benchmark suite. The metrics and targets are intended as starting points for structured evaluation. Real deployments should adapt thresholds to the domain, risk tier, user population, applicable regulation, and operational failure tolerance.
### Task and outcome performance

---
- end-to-end task completion;
- correctness against a defined reference or acceptance rule;
- user or workflow outcome;
- tool selection and parameter accuracy;
- correction and recovery behavior;
- unnecessary step or interaction cost.

### Safety, privacy, and policy behavior

- prohibited action or tool invocation;
- sensitive-data exposure or inappropriate retention;
- refusal, deferral, and escalation at defined boundaries;
- prompt-injection and untrusted-content handling;
- behavior under conflicting instructions;
- containment after a control or dependency failure.

### Operational reliability

## Practical artifacts
- latency and timeout behavior under stated load;
- unhandled and handled failure rates;
- retry, fallback, and circuit-breaker behavior;
- dependency and tool degradation;
- cost and resource consumption;
- recovery time and state consistency.

| Artifact | Use for |
### Governance and auditability

- decision and tool-call provenance;
- version and permission traceability;
- human override and stop mechanisms;
- evidence retention with privacy controls;
- owner, review, and residual-risk records;
- regression triggers after material change.

## Metrics need operational definitions

Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining.

For each metric, record:

| Field | Example question |
|---|---|
| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human-readable agent evaluation report |
| [`examples/sample-evaluation-report.md`](examples/sample-evaluation-report.md) | Filled generic Markdown example |
| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable report schema for structured evaluation evidence |
| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Filled JSON example aligned to the schema |
| Construct | What property are we trying to understand? |
| Observable rule | What exactly counts as success or failure? |
| Unit | Response, task, conversation, tool call, user, or time window? |
| Denominator | Which attempts are included, excluded, or missing? |
| Evaluator | Rule, test, expert, user, or model judge? |
| Instrument error | How can the evaluator itself be wrong? |
| Decision use | Which action changes if the metric moves? |

---
Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable.

## Evaluation dimensions
## Evaluation suite design

### 1. Task performance
A credible suite usually separates several purposes:

| Metric | Description | Measurement Method |
|---|---|---|
| **Task Completion Rate** | % of tasks completed without error | automated test suite against benchmark scenarios |
| **Goal Achievement Rate** | % of tasks where the final goal was met | human or LLM-judge evaluation |
| **Step Efficiency** | average steps taken vs. optimal path | comparison to expert baseline |
| **Tool Use Accuracy** | correct tool selection and parameter passing | automated verification |
| **Output Quality** | correctness and usefulness of agent outputs | rubric-based judging |
| Suite | Primary purpose |
|---|---|
| Operating-distribution sample | Estimate routine performance under stated workload assumptions |
| Critical-risk suite | Oversample rare, high-consequence conditions and hard gates |
| Boundary suite | Test ambiguity, refusal, escalation, and policy transitions |
| Fault-injection suite | Exercise tool, retrieval, network, permission, and state failures |
| Regression suite | Preserve previously discovered failures and important invariants |
| Exploratory suite | Discover new slices and failure modes; not used alone for release claims |

### 2. Safety and reliability
A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately.

| Metric | Description | Example target |
|---|---|---|
| **Hallucination Rate** | factually incorrect statements in outputs | < 2% on factual tasks, adjusted by risk tier |
| **Harmful Output Rate** | outputs violating safety guidelines | 0% hard limit for critical harms |
| **Refusal Appropriateness** | correct refusal of inappropriate requests | > 95%, adjusted by policy scope |
| **Prompt Injection Resistance** | resistance to adversarial inputs | tested via red-team suite |
| **Consistency** | same input produces the same output class | > 90% across repeated runs, adjusted by task type |
## Evaluator hierarchy

### 3. Operational reliability
Use the simplest evaluator capable of measuring the property.

| Metric | Description | Example target |
|---|---|---|
| **Availability** | uptime under normal and peak load | ≥ 99.5%, adjusted by service criticality |
| **P99 Latency** | 99th-percentile response time | ≤ configured SLA |
| **Error Rate** | % of requests ending in unhandled errors | < 0.1%, adjusted by workload |
| **Cost per Task** | token cost per successful task | tracked against budget |
| **Graceful Degradation** | behavior when tools or dependencies fail | tested via fault injection |
1. **Executable checks** for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior.
2. **Qualified human review** for domain judgment, consequential ambiguity, and disputed failures.
3. **Model judges** for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects.

### 4. Governance and auditability
A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process.

| Metric | Description | Pass Criterion |
|---|---|---|
| **Decision Traceability** | can every output be traced to inputs and decision evidence? | full trace available |
| **Tool Call Logging** | all external tool calls logged with parameters | 100% logged, with sensitive data controls |
| **Human Escalation Rate** | % of tasks escalated to human review | tracked with threshold defined |
| **Sensitive Data Handling** | PII or sensitive data not leaked in outputs or logs | 0 critical violations |
| **Override Capability** | operator can halt or override agent behavior | verified in testing |
## Repeated runs and uncertainty

---
For nondeterministic systems, report:

## Evaluation protocol
- attempts per scenario;
- run and seed policy where applicable;
- model, prompt, retrieval, tool, permission, and environment versions;
- missing or failed runs;
- scenario-level recurrence and variability;
- uncertainty intervals for estimated rates or means;
- coverage gaps and evaluator limitations that statistical intervals do not capture.

The target numbers above should not be treated as universal pass/fail thresholds. Use this protocol to make evaluations more defensible.
“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk.

| Protocol element | Recommended practice |
|---|---|
| Scenario set | Include normal, ambiguous, adversarial, tool-failure, escalation, and out-of-distribution tasks |
| Sample size | Define the minimum number of tasks per scenario class before testing starts |
| Risk tiering | Use stricter thresholds for safety-critical, regulated, or irreversible decisions |
| Repeated runs | Run the same scenario multiple times when model nondeterminism matters |
| Human adjudication | Use expert review for high-impact failures, disputed LLM-judge results, and policy-boundary cases |
| Severity levels | Classify findings as minor, major, critical, or blocker |
| Confidence | Report uncertainty, small sample caveats, and known benchmark limitations |
| Regression testing | Re-run the suite after model, prompt, tool, retrieval, or policy changes |
| Evidence retention | Store prompts, inputs, outputs, tool traces, evaluator notes, and sign-off decisions with appropriate data controls |

### Suggested release decision logic

| Decision | Use when |
|---|---|
| **Approve** | No blockers, critical controls pass, residual risks accepted by named owners |
| **Approve with conditions** | Minor or moderate issues exist but mitigations and monitoring are defined |
| **Hold** | Major issues require fixes before release |
| **Reject** | Critical safety, privacy, compliance, or operational failures are unresolved |
## Slice analysis

---
Predefine decision-relevant slices such as:

## Evaluation scenarios
- language or locale;
- task and user type;
- ambiguity level;
- tool authority and permission scope;
- data sensitivity;
- retrieval quality;
- accessibility need;
- dependency failure mode;
- safety or policy boundary.

### Standard benchmark suite
Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing.

| Scenario | Description | Difficulty |
|---|---|---|
| Single-step task | straightforward, well-defined task | Basic |
| Multi-step task | sequence of actions with intermediate dependencies | Intermediate |
| Ambiguous instruction | underspecified request requiring clarification | Intermediate |
| Tool failure | dependent tool returns an error | Intermediate |
| Adversarial input | prompt injection or manipulation attempt | Advanced |
| Novel situation | task outside the usual distribution | Advanced |
| Cascading failure | multiple dependencies fail in sequence | Advanced |
| Human escalation | agent should recognize the need to escalate | Advanced |
## Threshold and gate design

### Regulated-industry scenarios
This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective.

| Industry | Scenario | Key Evaluation Criteria |
|---|---|---|
| Healthcare | clinical documentation assistant | factual accuracy, refusal quality, low hallucination |
| Finance | transaction-analysis agent | compliance, auditability, no PII leakage |
| Insurance | claims-assessment agent | fairness, traceability, escalation quality |
| Legal | contract-review agent | citation accuracy, scope awareness, human escalation |
A threshold should record:

---
- decision owner and intended decision;
- baseline or comparator;
- affected population and harm model;
- measurement method and uncertainty;
- minimum sample and slice coverage;
- non-compensable hard gates;
- quality or operational targets that may be optimized;
- exception and residual-risk process;
- expiry or review trigger.

## Human-readable report template
A weighted aggregate must never hide a failed hard gate or an unresolved critical finding.

Use [`templates/evaluation-report.md`](templates/evaluation-report.md) for review meetings, sign-off packs, and stakeholder alignment.
## Decision semantics

## Machine-readable report schema
The schema and validator keep these concepts separate:

Use [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) when you want evaluation results to be structured for later validation, CI checks, dashboards, or trend analysis.
| Concept | Meaning |
|---|---|
| System risk tier | impact context for the evaluated system and use case |
| Scenario / finding severity | consequence of a specific failure |
| Blocker | unresolved condition that prevents the current decision |
| Required action | follow-up work accepted under a bounded conditional decision |
| Condition | explicit constraint attached to that decision |
| Residual risk | remaining uncertainty or harm accepted by a named authority |

The schema captures:
See [`docs/decision-semantics.md`](docs/decision-semantics.md) for the validator rules.

- metadata
- recommendation and blockers
- scope
- scenarios
- scorecard results
- findings
- oversight triggers
- release decision and sign-off
## Validation

---
The validator checks schema conformance and selected decision coherence:

## NIST AI RMF mapping
```bash
python tools/validate_evaluation_report.py \
schemas/evaluation-report.schema.json \
examples/sample-evaluation-report.json
```

Full mapping: [docs/nist-rmf-mapping.md](docs/nist-rmf-mapping.md)
It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable.

Key practitioner mapping:
## Failure review

- **MS.1**: AI risk identification through safety and adversarial scenarios
- **MS.3**: structured evaluation techniques and benchmark design
- **MS.5**: subgroup and fairness measurement
- **GV.4**: human oversight via escalation and override checks
Do not stop at the scorecard. For material failures, preserve:

---
- scenario and run identifiers;
- relevant input, context, and tool trace;
- evaluator result and disagreement;
- severity rationale;
- containment and remediation;
- root-cause hypothesis versus confirmed cause;
- regression test created;
- owner and disposition.

## Scope and disclaimer
The failure corpus often becomes more valuable than the original aggregate score.

This repository is shared in a personal capacity. It is not legal advice, compliance certification, regulatory approval, safety certification, benchmark certification, or official guidance from NIST, the EU, ISO, or any employer.
## Maturity and scope

References to NIST AI RMF, EU AI Act, safety controls, or industry obligations are practitioner mappings and examples. Always verify against official sources before using this framework for compliance, safety, or release decisions.
This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority.

---
References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance.

## Related repositories

| Repository | What it adds |
| Repository | Distinct role |
|---|---|
| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | oversight model, trust structure, accountability controls |
| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns and orchestration logic |
| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable execution and failure-mode examples |
| [`nist-rmf-guide`](https://github.com/simaba/nist-rmf-guide) | practitioner guide for broader RMF implementation |
| [`governance-playbook`](https://github.com/simaba/governance-playbook) | enterprise operating model beyond agent evaluation |
| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable bounded-agent behavior and failure paths |
| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns |
| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | authority, oversight, containment, and accountability |
| [`automotive-llm-eval-harness`](https://github.com/simaba/automotive-llm-eval-harness) | compact scorer for synthetic automotive evaluation artifacts |

---

*Maintained by [Sima Bagheri](https://github.com/simaba) · Connect on [LinkedIn](https://www.linkedin.com/in/simaba/)*
*Maintained by [Sima Bagheri](https://github.com/simaba).*
Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
296 changes: 162 additions & 134 deletions README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -2,194 +2,222 @@

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](LICENSE)
[![NIST AI RMF](https://img.shields.io/badge/NIST%20AI%20RMF-Informed-0055A4?style=flat-square)](https://airc.nist.gov/home)
[![Discussions](https://img.shields.io/badge/Discussions-Join-7289da?style=flat-square&logo=github)](https://github.com/simaba/agent-eval/discussions)

A structured framework for evaluating AI agents across task performance, safety, reliability, and governance dimensions, designed for regulated-industry deployment.
A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score.

---
The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds.

## Choose this repo when
## The central idea

Use this repository when you need **measurement and pass/fail logic** for AI agents:
An evaluation result is only useful when five things are explicit:

- what to measure
- which scenarios to test
- how to score performance, safety, and reliability
- what should appear in an evaluation report
- how to structure evaluation evidence for human review and future automation
1. **the decision it supports** — prototype, compare, pilot, expand, hold, roll back, or retire;
2. **the operating population it represents** — users, tasks, languages, tools, permissions, and exclusions;
3. **the measurement instrument** — executable rule, expert review, user outcome, or model judge;
4. **the uncertainty and coverage limits** — sample size, evaluator error, slice gaps, version drift, and scenario bias;
5. **the decision semantics** — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk.

Do **not** start here if you need the oversight model. Use [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance).
Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures.

Do **not** start here if you need orchestration patterns. Use [`agent-orchestration`](https://github.com/simaba/agent-orchestration).
## Start here

Do **not** start here if you want runnable behavior. Use [`agent-simulator`](https://github.com/simaba/agent-simulator).
| Artifact | Use it for |
|---|---|
| [`docs/evaluation-validity.md`](docs/evaluation-validity.md) | Designing evidence that can support a stated decision |
| [`docs/decision-semantics.md`](docs/decision-semantics.md) | Keeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct |
| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human review and sign-off discussion |
| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable evaluation records |
| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Fictional structured example |
| [`tools/validate_evaluation_report.py`](tools/validate_evaluation_report.py) | Schema and decision-coherence validation |

---
## What the framework evaluates

## Maturity
The dimensions below are categories for designing measurements. They are not automatically valid metrics.

This is a **practitioner evaluation framework**, not a certified benchmark suite. The metrics and targets are intended as starting points for structured evaluation. Real deployments should adapt thresholds to the domain, risk tier, user population, applicable regulation, and operational failure tolerance.
### Task and outcome performance

---
- end-to-end task completion;
- correctness against a defined reference or acceptance rule;
- user or workflow outcome;
- tool selection and parameter accuracy;
- correction and recovery behavior;
- unnecessary step or interaction cost.

### Safety, privacy, and policy behavior

- prohibited action or tool invocation;
- sensitive-data exposure or inappropriate retention;
- refusal, deferral, and escalation at defined boundaries;
- prompt-injection and untrusted-content handling;
- behavior under conflicting instructions;
- containment after a control or dependency failure.

### Operational reliability

## Practical artifacts
- latency and timeout behavior under stated load;
- unhandled and handled failure rates;
- retry, fallback, and circuit-breaker behavior;
- dependency and tool degradation;
- cost and resource consumption;
- recovery time and state consistency.

| Artifact | Use for |
### Governance and auditability

- decision and tool-call provenance;
- version and permission traceability;
- human override and stop mechanisms;
- evidence retention with privacy controls;
- owner, review, and residual-risk records;
- regression triggers after material change.

## Metrics need operational definitions

Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining.

For each metric, record:

| Field | Example question |
|---|---|
| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human-readable agent evaluation report |
| [`examples/sample-evaluation-report.md`](examples/sample-evaluation-report.md) | Filled generic Markdown example |
| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable report schema for structured evaluation evidence |
| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Filled JSON example aligned to the schema |
| Construct | What property are we trying to understand? |
| Observable rule | What exactly counts as success or failure? |
| Unit | Response, task, conversation, tool call, user, or time window? |
| Denominator | Which attempts are included, excluded, or missing? |
| Evaluator | Rule, test, expert, user, or model judge? |
| Instrument error | How can the evaluator itself be wrong? |
| Decision use | Which action changes if the metric moves? |

---
Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable.

## Evaluation dimensions
## Evaluation suite design

### 1. Task performance
A credible suite usually separates several purposes:

| Metric | Description | Measurement Method |
|---|---|---|
| **Task Completion Rate** | % of tasks completed without error | automated test suite against benchmark scenarios |
| **Goal Achievement Rate** | % of tasks where the final goal was met | human or LLM-judge evaluation |
| **Step Efficiency** | average steps taken vs. optimal path | comparison to expert baseline |
| **Tool Use Accuracy** | correct tool selection and parameter passing | automated verification |
| **Output Quality** | correctness and usefulness of agent outputs | rubric-based judging |
| Suite | Primary purpose |
|---|---|
| Operating-distribution sample | Estimate routine performance under stated workload assumptions |
| Critical-risk suite | Oversample rare, high-consequence conditions and hard gates |
| Boundary suite | Test ambiguity, refusal, escalation, and policy transitions |
| Fault-injection suite | Exercise tool, retrieval, network, permission, and state failures |
| Regression suite | Preserve previously discovered failures and important invariants |
| Exploratory suite | Discover new slices and failure modes; not used alone for release claims |

### 2. Safety and reliability
A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately.

| Metric | Description | Example target |
|---|---|---|
| **Hallucination Rate** | factually incorrect statements in outputs | < 2% on factual tasks, adjusted by risk tier |
| **Harmful Output Rate** | outputs violating safety guidelines | 0% hard limit for critical harms |
| **Refusal Appropriateness** | correct refusal of inappropriate requests | > 95%, adjusted by policy scope |
| **Prompt Injection Resistance** | resistance to adversarial inputs | tested via red-team suite |
| **Consistency** | same input produces the same output class | > 90% across repeated runs, adjusted by task type |
## Evaluator hierarchy

### 3. Operational reliability
Use the simplest evaluator capable of measuring the property.

| Metric | Description | Example target |
|---|---|---|
| **Availability** | uptime under normal and peak load | ≥ 99.5%, adjusted by service criticality |
| **P99 Latency** | 99th-percentile response time | ≤ configured SLA |
| **Error Rate** | % of requests ending in unhandled errors | < 0.1%, adjusted by workload |
| **Cost per Task** | token cost per successful task | tracked against budget |
| **Graceful Degradation** | behavior when tools or dependencies fail | tested via fault injection |
1. **Executable checks** for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior.
2. **Qualified human review** for domain judgment, consequential ambiguity, and disputed failures.
3. **Model judges** for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects.

### 4. Governance and auditability
A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process.

| Metric | Description | Pass Criterion |
|---|---|---|
| **Decision Traceability** | can every output be traced to inputs and decision evidence? | full trace available |
| **Tool Call Logging** | all external tool calls logged with parameters | 100% logged, with sensitive data controls |
| **Human Escalation Rate** | % of tasks escalated to human review | tracked with threshold defined |
| **Sensitive Data Handling** | PII or sensitive data not leaked in outputs or logs | 0 critical violations |
| **Override Capability** | operator can halt or override agent behavior | verified in testing |
## Repeated runs and uncertainty

---
For nondeterministic systems, report:

## Evaluation protocol
- attempts per scenario;
- run and seed policy where applicable;
- model, prompt, retrieval, tool, permission, and environment versions;
- missing or failed runs;
- scenario-level recurrence and variability;
- uncertainty intervals for estimated rates or means;
- coverage gaps and evaluator limitations that statistical intervals do not capture.

The target numbers above should not be treated as universal pass/fail thresholds. Use this protocol to make evaluations more defensible.
“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk.

| Protocol element | Recommended practice |
|---|---|
| Scenario set | Include normal, ambiguous, adversarial, tool-failure, escalation, and out-of-distribution tasks |
| Sample size | Define the minimum number of tasks per scenario class before testing starts |
| Risk tiering | Use stricter thresholds for safety-critical, regulated, or irreversible decisions |
| Repeated runs | Run the same scenario multiple times when model nondeterminism matters |
| Human adjudication | Use expert review for high-impact failures, disputed LLM-judge results, and policy-boundary cases |
| Severity levels | Classify findings as minor, major, critical, or blocker |
| Confidence | Report uncertainty, small sample caveats, and known benchmark limitations |
| Regression testing | Re-run the suite after model, prompt, tool, retrieval, or policy changes |
| Evidence retention | Store prompts, inputs, outputs, tool traces, evaluator notes, and sign-off decisions with appropriate data controls |

### Suggested release decision logic

| Decision | Use when |
|---|---|
| **Approve** | No blockers, critical controls pass, residual risks accepted by named owners |
| **Approve with conditions** | Minor or moderate issues exist but mitigations and monitoring are defined |
| **Hold** | Major issues require fixes before release |
| **Reject** | Critical safety, privacy, compliance, or operational failures are unresolved |
## Slice analysis

---
Predefine decision-relevant slices such as:

## Evaluation scenarios
- language or locale;
- task and user type;
- ambiguity level;
- tool authority and permission scope;
- data sensitivity;
- retrieval quality;
- accessibility need;
- dependency failure mode;
- safety or policy boundary.

### Standard benchmark suite
Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing.

| Scenario | Description | Difficulty |
|---|---|---|
| Single-step task | straightforward, well-defined task | Basic |
| Multi-step task | sequence of actions with intermediate dependencies | Intermediate |
| Ambiguous instruction | underspecified request requiring clarification | Intermediate |
| Tool failure | dependent tool returns an error | Intermediate |
| Adversarial input | prompt injection or manipulation attempt | Advanced |
| Novel situation | task outside the usual distribution | Advanced |
| Cascading failure | multiple dependencies fail in sequence | Advanced |
| Human escalation | agent should recognize the need to escalate | Advanced |
## Threshold and gate design

### Regulated-industry scenarios
This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective.

| Industry | Scenario | Key Evaluation Criteria |
|---|---|---|
| Healthcare | clinical documentation assistant | factual accuracy, refusal quality, low hallucination |
| Finance | transaction-analysis agent | compliance, auditability, no PII leakage |
| Insurance | claims-assessment agent | fairness, traceability, escalation quality |
| Legal | contract-review agent | citation accuracy, scope awareness, human escalation |
A threshold should record:

---
- decision owner and intended decision;
- baseline or comparator;
- affected population and harm model;
- measurement method and uncertainty;
- minimum sample and slice coverage;
- non-compensable hard gates;
- quality or operational targets that may be optimized;
- exception and residual-risk process;
- expiry or review trigger.

## Human-readable report template
A weighted aggregate must never hide a failed hard gate or an unresolved critical finding.

Use [`templates/evaluation-report.md`](templates/evaluation-report.md) for review meetings, sign-off packs, and stakeholder alignment.
## Decision semantics

## Machine-readable report schema
The schema and validator keep these concepts separate:

Use [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) when you want evaluation results to be structured for later validation, CI checks, dashboards, or trend analysis.
| Concept | Meaning |
|---|---|
| System risk tier | impact context for the evaluated system and use case |
| Scenario / finding severity | consequence of a specific failure |
| Blocker | unresolved condition that prevents the current decision |
| Required action | follow-up work accepted under a bounded conditional decision |
| Condition | explicit constraint attached to that decision |
| Residual risk | remaining uncertainty or harm accepted by a named authority |

The schema captures:
See [`docs/decision-semantics.md`](docs/decision-semantics.md) for the validator rules.

- metadata
- recommendation and blockers
- scope
- scenarios
- scorecard results
- findings
- oversight triggers
- release decision and sign-off
## Validation

---
The validator checks schema conformance and selected decision coherence:

## NIST AI RMF mapping
```bash
python tools/validate_evaluation_report.py \
schemas/evaluation-report.schema.json \
examples/sample-evaluation-report.json
```

Full mapping: [docs/nist-rmf-mapping.md](docs/nist-rmf-mapping.md)
It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable.

Key practitioner mapping:
## Failure review

- **MS.1**: AI risk identification through safety and adversarial scenarios
- **MS.3**: structured evaluation techniques and benchmark design
- **MS.5**: subgroup and fairness measurement
- **GV.4**: human oversight via escalation and override checks
Do not stop at the scorecard. For material failures, preserve:

---
- scenario and run identifiers;
- relevant input, context, and tool trace;
- evaluator result and disagreement;
- severity rationale;
- containment and remediation;
- root-cause hypothesis versus confirmed cause;
- regression test created;
- owner and disposition.

## Scope and disclaimer
The failure corpus often becomes more valuable than the original aggregate score.

This repository is shared in a personal capacity. It is not legal advice, compliance certification, regulatory approval, safety certification, benchmark certification, or official guidance from NIST, the EU, ISO, or any employer.
## Maturity and scope

References to NIST AI RMF, EU AI Act, safety controls, or industry obligations are practitioner mappings and examples. Always verify against official sources before using this framework for compliance, safety, or release decisions.
This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority.

---
References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance.

## Related repositories

| Repository | What it adds |
| Repository | Distinct role |
|---|---|
| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | oversight model, trust structure, accountability controls |
| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns and orchestration logic |
| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable execution and failure-mode examples |
| [`nist-rmf-guide`](https://github.com/simaba/nist-rmf-guide) | practitioner guide for broader RMF implementation |
| [`governance-playbook`](https://github.com/simaba/governance-playbook) | enterprise operating model beyond agent evaluation |
| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable bounded-agent behavior and failure paths |
| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns |
| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | authority, oversight, containment, and accountability |
| [`automotive-llm-eval-harness`](https://github.com/simaba/automotive-llm-eval-harness) | compact scorer for synthetic automotive evaluation artifacts |

---

*Maintained by [Sima Bagheri](https://github.com/simaba) · Connect on [LinkedIn](https://www.linkedin.com/in/simaba/)*
*Maintained by [Sima Bagheri](https://github.com/simaba).*
Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
296 changes: 162 additions & 134 deletions README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -2,194 +2,222 @@

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](LICENSE)
[![NIST AI RMF](https://img.shields.io/badge/NIST%20AI%20RMF-Informed-0055A4?style=flat-square)](https://airc.nist.gov/home)
[![Discussions](https://img.shields.io/badge/Discussions-Join-7289da?style=flat-square&logo=github)](https://github.com/simaba/agent-eval/discussions)

A structured framework for evaluating AI agents across task performance, safety, reliability, and governance dimensions, designed for regulated-industry deployment.
A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score.

---
The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds.

## Choose this repo when
## The central idea

Use this repository when you need **measurement and pass/fail logic** for AI agents:
An evaluation result is only useful when five things are explicit:

- what to measure
- which scenarios to test
- how to score performance, safety, and reliability
- what should appear in an evaluation report
- how to structure evaluation evidence for human review and future automation
1. **the decision it supports** — prototype, compare, pilot, expand, hold, roll back, or retire;
2. **the operating population it represents** — users, tasks, languages, tools, permissions, and exclusions;
3. **the measurement instrument** — executable rule, expert review, user outcome, or model judge;
4. **the uncertainty and coverage limits** — sample size, evaluator error, slice gaps, version drift, and scenario bias;
5. **the decision semantics** — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk.

Do **not** start here if you need the oversight model. Use [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance).
Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures.

Do **not** start here if you need orchestration patterns. Use [`agent-orchestration`](https://github.com/simaba/agent-orchestration).
## Start here

Do **not** start here if you want runnable behavior. Use [`agent-simulator`](https://github.com/simaba/agent-simulator).
| Artifact | Use it for |
|---|---|
| [`docs/evaluation-validity.md`](docs/evaluation-validity.md) | Designing evidence that can support a stated decision |
| [`docs/decision-semantics.md`](docs/decision-semantics.md) | Keeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct |
| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human review and sign-off discussion |
| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable evaluation records |
| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Fictional structured example |
| [`tools/validate_evaluation_report.py`](tools/validate_evaluation_report.py) | Schema and decision-coherence validation |

---
## What the framework evaluates

## Maturity
The dimensions below are categories for designing measurements. They are not automatically valid metrics.

This is a **practitioner evaluation framework**, not a certified benchmark suite. The metrics and targets are intended as starting points for structured evaluation. Real deployments should adapt thresholds to the domain, risk tier, user population, applicable regulation, and operational failure tolerance.
### Task and outcome performance

---
- end-to-end task completion;
- correctness against a defined reference or acceptance rule;
- user or workflow outcome;
- tool selection and parameter accuracy;
- correction and recovery behavior;
- unnecessary step or interaction cost.

### Safety, privacy, and policy behavior

- prohibited action or tool invocation;
- sensitive-data exposure or inappropriate retention;
- refusal, deferral, and escalation at defined boundaries;
- prompt-injection and untrusted-content handling;
- behavior under conflicting instructions;
- containment after a control or dependency failure.

### Operational reliability

## Practical artifacts
- latency and timeout behavior under stated load;
- unhandled and handled failure rates;
- retry, fallback, and circuit-breaker behavior;
- dependency and tool degradation;
- cost and resource consumption;
- recovery time and state consistency.

| Artifact | Use for |
### Governance and auditability

- decision and tool-call provenance;
- version and permission traceability;
- human override and stop mechanisms;
- evidence retention with privacy controls;
- owner, review, and residual-risk records;
- regression triggers after material change.

## Metrics need operational definitions

Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining.

For each metric, record:

| Field | Example question |
|---|---|
| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human-readable agent evaluation report |
| [`examples/sample-evaluation-report.md`](examples/sample-evaluation-report.md) | Filled generic Markdown example |
| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable report schema for structured evaluation evidence |
| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Filled JSON example aligned to the schema |
| Construct | What property are we trying to understand? |
| Observable rule | What exactly counts as success or failure? |
| Unit | Response, task, conversation, tool call, user, or time window? |
| Denominator | Which attempts are included, excluded, or missing? |
| Evaluator | Rule, test, expert, user, or model judge? |
| Instrument error | How can the evaluator itself be wrong? |
| Decision use | Which action changes if the metric moves? |

---
Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable.

## Evaluation dimensions
## Evaluation suite design

### 1. Task performance
A credible suite usually separates several purposes:

| Metric | Description | Measurement Method |
|---|---|---|
| **Task Completion Rate** | % of tasks completed without error | automated test suite against benchmark scenarios |
| **Goal Achievement Rate** | % of tasks where the final goal was met | human or LLM-judge evaluation |
| **Step Efficiency** | average steps taken vs. optimal path | comparison to expert baseline |
| **Tool Use Accuracy** | correct tool selection and parameter passing | automated verification |
| **Output Quality** | correctness and usefulness of agent outputs | rubric-based judging |
| Suite | Primary purpose |
|---|---|
| Operating-distribution sample | Estimate routine performance under stated workload assumptions |
| Critical-risk suite | Oversample rare, high-consequence conditions and hard gates |
| Boundary suite | Test ambiguity, refusal, escalation, and policy transitions |
| Fault-injection suite | Exercise tool, retrieval, network, permission, and state failures |
| Regression suite | Preserve previously discovered failures and important invariants |
| Exploratory suite | Discover new slices and failure modes; not used alone for release claims |

### 2. Safety and reliability
A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately.

| Metric | Description | Example target |
|---|---|---|
| **Hallucination Rate** | factually incorrect statements in outputs | < 2% on factual tasks, adjusted by risk tier |
| **Harmful Output Rate** | outputs violating safety guidelines | 0% hard limit for critical harms |
| **Refusal Appropriateness** | correct refusal of inappropriate requests | > 95%, adjusted by policy scope |
| **Prompt Injection Resistance** | resistance to adversarial inputs | tested via red-team suite |
| **Consistency** | same input produces the same output class | > 90% across repeated runs, adjusted by task type |
## Evaluator hierarchy

### 3. Operational reliability
Use the simplest evaluator capable of measuring the property.

| Metric | Description | Example target |
|---|---|---|
| **Availability** | uptime under normal and peak load | ≥ 99.5%, adjusted by service criticality |
| **P99 Latency** | 99th-percentile response time | ≤ configured SLA |
| **Error Rate** | % of requests ending in unhandled errors | < 0.1%, adjusted by workload |
| **Cost per Task** | token cost per successful task | tracked against budget |
| **Graceful Degradation** | behavior when tools or dependencies fail | tested via fault injection |
1. **Executable checks** for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior.
2. **Qualified human review** for domain judgment, consequential ambiguity, and disputed failures.
3. **Model judges** for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects.

### 4. Governance and auditability
A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process.

| Metric | Description | Pass Criterion |
|---|---|---|
| **Decision Traceability** | can every output be traced to inputs and decision evidence? | full trace available |
| **Tool Call Logging** | all external tool calls logged with parameters | 100% logged, with sensitive data controls |
| **Human Escalation Rate** | % of tasks escalated to human review | tracked with threshold defined |
| **Sensitive Data Handling** | PII or sensitive data not leaked in outputs or logs | 0 critical violations |
| **Override Capability** | operator can halt or override agent behavior | verified in testing |
## Repeated runs and uncertainty

---
For nondeterministic systems, report:

## Evaluation protocol
- attempts per scenario;
- run and seed policy where applicable;
- model, prompt, retrieval, tool, permission, and environment versions;
- missing or failed runs;
- scenario-level recurrence and variability;
- uncertainty intervals for estimated rates or means;
- coverage gaps and evaluator limitations that statistical intervals do not capture.

The target numbers above should not be treated as universal pass/fail thresholds. Use this protocol to make evaluations more defensible.
“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk.

| Protocol element | Recommended practice |
|---|---|
| Scenario set | Include normal, ambiguous, adversarial, tool-failure, escalation, and out-of-distribution tasks |
| Sample size | Define the minimum number of tasks per scenario class before testing starts |
| Risk tiering | Use stricter thresholds for safety-critical, regulated, or irreversible decisions |
| Repeated runs | Run the same scenario multiple times when model nondeterminism matters |
| Human adjudication | Use expert review for high-impact failures, disputed LLM-judge results, and policy-boundary cases |
| Severity levels | Classify findings as minor, major, critical, or blocker |
| Confidence | Report uncertainty, small sample caveats, and known benchmark limitations |
| Regression testing | Re-run the suite after model, prompt, tool, retrieval, or policy changes |
| Evidence retention | Store prompts, inputs, outputs, tool traces, evaluator notes, and sign-off decisions with appropriate data controls |

### Suggested release decision logic

| Decision | Use when |
|---|---|
| **Approve** | No blockers, critical controls pass, residual risks accepted by named owners |
| **Approve with conditions** | Minor or moderate issues exist but mitigations and monitoring are defined |
| **Hold** | Major issues require fixes before release |
| **Reject** | Critical safety, privacy, compliance, or operational failures are unresolved |
## Slice analysis

---
Predefine decision-relevant slices such as:

## Evaluation scenarios
- language or locale;
- task and user type;
- ambiguity level;
- tool authority and permission scope;
- data sensitivity;
- retrieval quality;
- accessibility need;
- dependency failure mode;
- safety or policy boundary.

### Standard benchmark suite
Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing.

| Scenario | Description | Difficulty |
|---|---|---|
| Single-step task | straightforward, well-defined task | Basic |
| Multi-step task | sequence of actions with intermediate dependencies | Intermediate |
| Ambiguous instruction | underspecified request requiring clarification | Intermediate |
| Tool failure | dependent tool returns an error | Intermediate |
| Adversarial input | prompt injection or manipulation attempt | Advanced |
| Novel situation | task outside the usual distribution | Advanced |
| Cascading failure | multiple dependencies fail in sequence | Advanced |
| Human escalation | agent should recognize the need to escalate | Advanced |
## Threshold and gate design

### Regulated-industry scenarios
This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective.

| Industry | Scenario | Key Evaluation Criteria |
|---|---|---|
| Healthcare | clinical documentation assistant | factual accuracy, refusal quality, low hallucination |
| Finance | transaction-analysis agent | compliance, auditability, no PII leakage |
| Insurance | claims-assessment agent | fairness, traceability, escalation quality |
| Legal | contract-review agent | citation accuracy, scope awareness, human escalation |
A threshold should record:

---
- decision owner and intended decision;
- baseline or comparator;
- affected population and harm model;
- measurement method and uncertainty;
- minimum sample and slice coverage;
- non-compensable hard gates;
- quality or operational targets that may be optimized;
- exception and residual-risk process;
- expiry or review trigger.

## Human-readable report template
A weighted aggregate must never hide a failed hard gate or an unresolved critical finding.

Use [`templates/evaluation-report.md`](templates/evaluation-report.md) for review meetings, sign-off packs, and stakeholder alignment.
## Decision semantics

## Machine-readable report schema
The schema and validator keep these concepts separate:

Use [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) when you want evaluation results to be structured for later validation, CI checks, dashboards, or trend analysis.
| Concept | Meaning |
|---|---|
| System risk tier | impact context for the evaluated system and use case |
| Scenario / finding severity | consequence of a specific failure |
| Blocker | unresolved condition that prevents the current decision |
| Required action | follow-up work accepted under a bounded conditional decision |
| Condition | explicit constraint attached to that decision |
| Residual risk | remaining uncertainty or harm accepted by a named authority |

The schema captures:
See [`docs/decision-semantics.md`](docs/decision-semantics.md) for the validator rules.

- metadata
- recommendation and blockers
- scope
- scenarios
- scorecard results
- findings
- oversight triggers
- release decision and sign-off
## Validation

---
The validator checks schema conformance and selected decision coherence:

## NIST AI RMF mapping
```bash
python tools/validate_evaluation_report.py \
schemas/evaluation-report.schema.json \
examples/sample-evaluation-report.json
```

Full mapping: [docs/nist-rmf-mapping.md](docs/nist-rmf-mapping.md)
It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable.

Key practitioner mapping:
## Failure review

- **MS.1**: AI risk identification through safety and adversarial scenarios
- **MS.3**: structured evaluation techniques and benchmark design
- **MS.5**: subgroup and fairness measurement
- **GV.4**: human oversight via escalation and override checks
Do not stop at the scorecard. For material failures, preserve:

---
- scenario and run identifiers;
- relevant input, context, and tool trace;
- evaluator result and disagreement;
- severity rationale;
- containment and remediation;
- root-cause hypothesis versus confirmed cause;
- regression test created;
- owner and disposition.

## Scope and disclaimer
The failure corpus often becomes more valuable than the original aggregate score.

This repository is shared in a personal capacity. It is not legal advice, compliance certification, regulatory approval, safety certification, benchmark certification, or official guidance from NIST, the EU, ISO, or any employer.
## Maturity and scope

References to NIST AI RMF, EU AI Act, safety controls, or industry obligations are practitioner mappings and examples. Always verify against official sources before using this framework for compliance, safety, or release decisions.
This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority.

---
References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance.

## Related repositories

| Repository | What it adds |
| Repository | Distinct role |
|---|---|
| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | oversight model, trust structure, accountability controls |
| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns and orchestration logic |
| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable execution and failure-mode examples |
| [`nist-rmf-guide`](https://github.com/simaba/nist-rmf-guide) | practitioner guide for broader RMF implementation |
| [`governance-playbook`](https://github.com/simaba/governance-playbook) | enterprise operating model beyond agent evaluation |
| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable bounded-agent behavior and failure paths |
| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns |
| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | authority, oversight, containment, and accountability |
| [`automotive-llm-eval-harness`](https://github.com/simaba/automotive-llm-eval-harness) | compact scorer for synthetic automotive evaluation artifacts |

---

*Maintained by [Sima Bagheri](https://github.com/simaba) · Connect on [LinkedIn](https://www.linkedin.com/in/simaba/)*
*Maintained by [Sima Bagheri](https://github.com/simaba).*
Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
296 changes: 162 additions & 134 deletions README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -2,194 +2,222 @@

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](LICENSE)
[![NIST AI RMF](https://img.shields.io/badge/NIST%20AI%20RMF-Informed-0055A4?style=flat-square)](https://airc.nist.gov/home)
[![Discussions](https://img.shields.io/badge/Discussions-Join-7289da?style=flat-square&logo=github)](https://github.com/simaba/agent-eval/discussions)

A structured framework for evaluating AI agents across task performance, safety, reliability, and governance dimensions, designed for regulated-industry deployment.
A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score.

---
The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds.

## Choose this repo when
## The central idea

Use this repository when you need **measurement and pass/fail logic** for AI agents:
An evaluation result is only useful when five things are explicit:

- what to measure
- which scenarios to test
- how to score performance, safety, and reliability
- what should appear in an evaluation report
- how to structure evaluation evidence for human review and future automation
1. **the decision it supports** — prototype, compare, pilot, expand, hold, roll back, or retire;
2. **the operating population it represents** — users, tasks, languages, tools, permissions, and exclusions;
3. **the measurement instrument** — executable rule, expert review, user outcome, or model judge;
4. **the uncertainty and coverage limits** — sample size, evaluator error, slice gaps, version drift, and scenario bias;
5. **the decision semantics** — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk.

Do **not** start here if you need the oversight model. Use [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance).
Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures.

Do **not** start here if you need orchestration patterns. Use [`agent-orchestration`](https://github.com/simaba/agent-orchestration).
## Start here

Do **not** start here if you want runnable behavior. Use [`agent-simulator`](https://github.com/simaba/agent-simulator).
| Artifact | Use it for |
|---|---|
| [`docs/evaluation-validity.md`](docs/evaluation-validity.md) | Designing evidence that can support a stated decision |
| [`docs/decision-semantics.md`](docs/decision-semantics.md) | Keeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct |
| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human review and sign-off discussion |
| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable evaluation records |
| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Fictional structured example |
| [`tools/validate_evaluation_report.py`](tools/validate_evaluation_report.py) | Schema and decision-coherence validation |

---
## What the framework evaluates

## Maturity
The dimensions below are categories for designing measurements. They are not automatically valid metrics.

This is a **practitioner evaluation framework**, not a certified benchmark suite. The metrics and targets are intended as starting points for structured evaluation. Real deployments should adapt thresholds to the domain, risk tier, user population, applicable regulation, and operational failure tolerance.
### Task and outcome performance

---
- end-to-end task completion;
- correctness against a defined reference or acceptance rule;
- user or workflow outcome;
- tool selection and parameter accuracy;
- correction and recovery behavior;
- unnecessary step or interaction cost.

### Safety, privacy, and policy behavior

- prohibited action or tool invocation;
- sensitive-data exposure or inappropriate retention;
- refusal, deferral, and escalation at defined boundaries;
- prompt-injection and untrusted-content handling;
- behavior under conflicting instructions;
- containment after a control or dependency failure.

### Operational reliability

## Practical artifacts
- latency and timeout behavior under stated load;
- unhandled and handled failure rates;
- retry, fallback, and circuit-breaker behavior;
- dependency and tool degradation;
- cost and resource consumption;
- recovery time and state consistency.

| Artifact | Use for |
### Governance and auditability

- decision and tool-call provenance;
- version and permission traceability;
- human override and stop mechanisms;
- evidence retention with privacy controls;
- owner, review, and residual-risk records;
- regression triggers after material change.

## Metrics need operational definitions

Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining.

For each metric, record:

| Field | Example question |
|---|---|
| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human-readable agent evaluation report |
| [`examples/sample-evaluation-report.md`](examples/sample-evaluation-report.md) | Filled generic Markdown example |
| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable report schema for structured evaluation evidence |
| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Filled JSON example aligned to the schema |
| Construct | What property are we trying to understand? |
| Observable rule | What exactly counts as success or failure? |
| Unit | Response, task, conversation, tool call, user, or time window? |
| Denominator | Which attempts are included, excluded, or missing? |
| Evaluator | Rule, test, expert, user, or model judge? |
| Instrument error | How can the evaluator itself be wrong? |
| Decision use | Which action changes if the metric moves? |

---
Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable.

## Evaluation dimensions
## Evaluation suite design

### 1. Task performance
A credible suite usually separates several purposes:

| Metric | Description | Measurement Method |
|---|---|---|
| **Task Completion Rate** | % of tasks completed without error | automated test suite against benchmark scenarios |
| **Goal Achievement Rate** | % of tasks where the final goal was met | human or LLM-judge evaluation |
| **Step Efficiency** | average steps taken vs. optimal path | comparison to expert baseline |
| **Tool Use Accuracy** | correct tool selection and parameter passing | automated verification |
| **Output Quality** | correctness and usefulness of agent outputs | rubric-based judging |
| Suite | Primary purpose |
|---|---|
| Operating-distribution sample | Estimate routine performance under stated workload assumptions |
| Critical-risk suite | Oversample rare, high-consequence conditions and hard gates |
| Boundary suite | Test ambiguity, refusal, escalation, and policy transitions |
| Fault-injection suite | Exercise tool, retrieval, network, permission, and state failures |
| Regression suite | Preserve previously discovered failures and important invariants |
| Exploratory suite | Discover new slices and failure modes; not used alone for release claims |

### 2. Safety and reliability
A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately.

| Metric | Description | Example target |
|---|---|---|
| **Hallucination Rate** | factually incorrect statements in outputs | < 2% on factual tasks, adjusted by risk tier |
| **Harmful Output Rate** | outputs violating safety guidelines | 0% hard limit for critical harms |
| **Refusal Appropriateness** | correct refusal of inappropriate requests | > 95%, adjusted by policy scope |
| **Prompt Injection Resistance** | resistance to adversarial inputs | tested via red-team suite |
| **Consistency** | same input produces the same output class | > 90% across repeated runs, adjusted by task type |
## Evaluator hierarchy

### 3. Operational reliability
Use the simplest evaluator capable of measuring the property.

| Metric | Description | Example target |
|---|---|---|
| **Availability** | uptime under normal and peak load | ≥ 99.5%, adjusted by service criticality |
| **P99 Latency** | 99th-percentile response time | ≤ configured SLA |
| **Error Rate** | % of requests ending in unhandled errors | < 0.1%, adjusted by workload |
| **Cost per Task** | token cost per successful task | tracked against budget |
| **Graceful Degradation** | behavior when tools or dependencies fail | tested via fault injection |
1. **Executable checks** for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior.
2. **Qualified human review** for domain judgment, consequential ambiguity, and disputed failures.
3. **Model judges** for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects.

### 4. Governance and auditability
A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process.

| Metric | Description | Pass Criterion |
|---|---|---|
| **Decision Traceability** | can every output be traced to inputs and decision evidence? | full trace available |
| **Tool Call Logging** | all external tool calls logged with parameters | 100% logged, with sensitive data controls |
| **Human Escalation Rate** | % of tasks escalated to human review | tracked with threshold defined |
| **Sensitive Data Handling** | PII or sensitive data not leaked in outputs or logs | 0 critical violations |
| **Override Capability** | operator can halt or override agent behavior | verified in testing |
## Repeated runs and uncertainty

---
For nondeterministic systems, report:

## Evaluation protocol
- attempts per scenario;
- run and seed policy where applicable;
- model, prompt, retrieval, tool, permission, and environment versions;
- missing or failed runs;
- scenario-level recurrence and variability;
- uncertainty intervals for estimated rates or means;
- coverage gaps and evaluator limitations that statistical intervals do not capture.

The target numbers above should not be treated as universal pass/fail thresholds. Use this protocol to make evaluations more defensible.
“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk.

| Protocol element | Recommended practice |
|---|---|
| Scenario set | Include normal, ambiguous, adversarial, tool-failure, escalation, and out-of-distribution tasks |
| Sample size | Define the minimum number of tasks per scenario class before testing starts |
| Risk tiering | Use stricter thresholds for safety-critical, regulated, or irreversible decisions |
| Repeated runs | Run the same scenario multiple times when model nondeterminism matters |
| Human adjudication | Use expert review for high-impact failures, disputed LLM-judge results, and policy-boundary cases |
| Severity levels | Classify findings as minor, major, critical, or blocker |
| Confidence | Report uncertainty, small sample caveats, and known benchmark limitations |
| Regression testing | Re-run the suite after model, prompt, tool, retrieval, or policy changes |
| Evidence retention | Store prompts, inputs, outputs, tool traces, evaluator notes, and sign-off decisions with appropriate data controls |

### Suggested release decision logic

| Decision | Use when |
|---|---|
| **Approve** | No blockers, critical controls pass, residual risks accepted by named owners |
| **Approve with conditions** | Minor or moderate issues exist but mitigations and monitoring are defined |
| **Hold** | Major issues require fixes before release |
| **Reject** | Critical safety, privacy, compliance, or operational failures are unresolved |
## Slice analysis

---
Predefine decision-relevant slices such as:

## Evaluation scenarios
- language or locale;
- task and user type;
- ambiguity level;
- tool authority and permission scope;
- data sensitivity;
- retrieval quality;
- accessibility need;
- dependency failure mode;
- safety or policy boundary.

### Standard benchmark suite
Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing.

| Scenario | Description | Difficulty |
|---|---|---|
| Single-step task | straightforward, well-defined task | Basic |
| Multi-step task | sequence of actions with intermediate dependencies | Intermediate |
| Ambiguous instruction | underspecified request requiring clarification | Intermediate |
| Tool failure | dependent tool returns an error | Intermediate |
| Adversarial input | prompt injection or manipulation attempt | Advanced |
| Novel situation | task outside the usual distribution | Advanced |
| Cascading failure | multiple dependencies fail in sequence | Advanced |
| Human escalation | agent should recognize the need to escalate | Advanced |
## Threshold and gate design

### Regulated-industry scenarios
This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective.

| Industry | Scenario | Key Evaluation Criteria |
|---|---|---|
| Healthcare | clinical documentation assistant | factual accuracy, refusal quality, low hallucination |
| Finance | transaction-analysis agent | compliance, auditability, no PII leakage |
| Insurance | claims-assessment agent | fairness, traceability, escalation quality |
| Legal | contract-review agent | citation accuracy, scope awareness, human escalation |
A threshold should record:

---
- decision owner and intended decision;
- baseline or comparator;
- affected population and harm model;
- measurement method and uncertainty;
- minimum sample and slice coverage;
- non-compensable hard gates;
- quality or operational targets that may be optimized;
- exception and residual-risk process;
- expiry or review trigger.

## Human-readable report template
A weighted aggregate must never hide a failed hard gate or an unresolved critical finding.

Use [`templates/evaluation-report.md`](templates/evaluation-report.md) for review meetings, sign-off packs, and stakeholder alignment.
## Decision semantics

## Machine-readable report schema
The schema and validator keep these concepts separate:

Use [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) when you want evaluation results to be structured for later validation, CI checks, dashboards, or trend analysis.
| Concept | Meaning |
|---|---|
| System risk tier | impact context for the evaluated system and use case |
| Scenario / finding severity | consequence of a specific failure |
| Blocker | unresolved condition that prevents the current decision |
| Required action | follow-up work accepted under a bounded conditional decision |
| Condition | explicit constraint attached to that decision |
| Residual risk | remaining uncertainty or harm accepted by a named authority |

The schema captures:
See [`docs/decision-semantics.md`](docs/decision-semantics.md) for the validator rules.

- metadata
- recommendation and blockers
- scope
- scenarios
- scorecard results
- findings
- oversight triggers
- release decision and sign-off
## Validation

---
The validator checks schema conformance and selected decision coherence:

## NIST AI RMF mapping
```bash
python tools/validate_evaluation_report.py \
schemas/evaluation-report.schema.json \
examples/sample-evaluation-report.json
```

Full mapping: [docs/nist-rmf-mapping.md](docs/nist-rmf-mapping.md)
It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable.

Key practitioner mapping:
## Failure review

- **MS.1**: AI risk identification through safety and adversarial scenarios
- **MS.3**: structured evaluation techniques and benchmark design
- **MS.5**: subgroup and fairness measurement
- **GV.4**: human oversight via escalation and override checks
Do not stop at the scorecard. For material failures, preserve:

---
- scenario and run identifiers;
- relevant input, context, and tool trace;
- evaluator result and disagreement;
- severity rationale;
- containment and remediation;
- root-cause hypothesis versus confirmed cause;
- regression test created;
- owner and disposition.

## Scope and disclaimer
The failure corpus often becomes more valuable than the original aggregate score.

This repository is shared in a personal capacity. It is not legal advice, compliance certification, regulatory approval, safety certification, benchmark certification, or official guidance from NIST, the EU, ISO, or any employer.
## Maturity and scope

References to NIST AI RMF, EU AI Act, safety controls, or industry obligations are practitioner mappings and examples. Always verify against official sources before using this framework for compliance, safety, or release decisions.
This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority.

---
References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance.

## Related repositories

| Repository | What it adds |
| Repository | Distinct role |
|---|---|
| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | oversight model, trust structure, accountability controls |
| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns and orchestration logic |
| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable execution and failure-mode examples |
| [`nist-rmf-guide`](https://github.com/simaba/nist-rmf-guide) | practitioner guide for broader RMF implementation |
| [`governance-playbook`](https://github.com/simaba/governance-playbook) | enterprise operating model beyond agent evaluation |
| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable bounded-agent behavior and failure paths |
| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns |
| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | authority, oversight, containment, and accountability |
| [`automotive-llm-eval-harness`](https://github.com/simaba/automotive-llm-eval-harness) | compact scorer for synthetic automotive evaluation artifacts |

---

*Maintained by [Sima Bagheri](https://github.com/simaba) · Connect on [LinkedIn](https://www.linkedin.com/in/simaba/)*
*Maintained by [Sima Bagheri](https://github.com/simaba).*
Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
296 changes: 162 additions & 134 deletions README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -2,194 +2,222 @@

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](LICENSE)
[![NIST AI RMF](https://img.shields.io/badge/NIST%20AI%20RMF-Informed-0055A4?style=flat-square)](https://airc.nist.gov/home)
[![Discussions](https://img.shields.io/badge/Discussions-Join-7289da?style=flat-square&logo=github)](https://github.com/simaba/agent-eval/discussions)

A structured framework for evaluating AI agents across task performance, safety, reliability, and governance dimensions, designed for regulated-industry deployment.
A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score.

---
The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds.

## Choose this repo when
## The central idea

Use this repository when you need **measurement and pass/fail logic** for AI agents:
An evaluation result is only useful when five things are explicit:

- what to measure
- which scenarios to test
- how to score performance, safety, and reliability
- what should appear in an evaluation report
- how to structure evaluation evidence for human review and future automation
1. **the decision it supports** — prototype, compare, pilot, expand, hold, roll back, or retire;
2. **the operating population it represents** — users, tasks, languages, tools, permissions, and exclusions;
3. **the measurement instrument** — executable rule, expert review, user outcome, or model judge;
4. **the uncertainty and coverage limits** — sample size, evaluator error, slice gaps, version drift, and scenario bias;
5. **the decision semantics** — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk.

Do **not** start here if you need the oversight model. Use [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance).
Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures.

Do **not** start here if you need orchestration patterns. Use [`agent-orchestration`](https://github.com/simaba/agent-orchestration).
## Start here

Do **not** start here if you want runnable behavior. Use [`agent-simulator`](https://github.com/simaba/agent-simulator).
| Artifact | Use it for |
|---|---|
| [`docs/evaluation-validity.md`](docs/evaluation-validity.md) | Designing evidence that can support a stated decision |
| [`docs/decision-semantics.md`](docs/decision-semantics.md) | Keeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct |
| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human review and sign-off discussion |
| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable evaluation records |
| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Fictional structured example |
| [`tools/validate_evaluation_report.py`](tools/validate_evaluation_report.py) | Schema and decision-coherence validation |

---
## What the framework evaluates

## Maturity
The dimensions below are categories for designing measurements. They are not automatically valid metrics.

This is a **practitioner evaluation framework**, not a certified benchmark suite. The metrics and targets are intended as starting points for structured evaluation. Real deployments should adapt thresholds to the domain, risk tier, user population, applicable regulation, and operational failure tolerance.
### Task and outcome performance

---
- end-to-end task completion;
- correctness against a defined reference or acceptance rule;
- user or workflow outcome;
- tool selection and parameter accuracy;
- correction and recovery behavior;
- unnecessary step or interaction cost.

### Safety, privacy, and policy behavior

- prohibited action or tool invocation;
- sensitive-data exposure or inappropriate retention;
- refusal, deferral, and escalation at defined boundaries;
- prompt-injection and untrusted-content handling;
- behavior under conflicting instructions;
- containment after a control or dependency failure.

### Operational reliability

## Practical artifacts
- latency and timeout behavior under stated load;
- unhandled and handled failure rates;
- retry, fallback, and circuit-breaker behavior;
- dependency and tool degradation;
- cost and resource consumption;
- recovery time and state consistency.

| Artifact | Use for |
### Governance and auditability

- decision and tool-call provenance;
- version and permission traceability;
- human override and stop mechanisms;
- evidence retention with privacy controls;
- owner, review, and residual-risk records;
- regression triggers after material change.

## Metrics need operational definitions

Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining.

For each metric, record:

| Field | Example question |
|---|---|
| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human-readable agent evaluation report |
| [`examples/sample-evaluation-report.md`](examples/sample-evaluation-report.md) | Filled generic Markdown example |
| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable report schema for structured evaluation evidence |
| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Filled JSON example aligned to the schema |
| Construct | What property are we trying to understand? |
| Observable rule | What exactly counts as success or failure? |
| Unit | Response, task, conversation, tool call, user, or time window? |
| Denominator | Which attempts are included, excluded, or missing? |
| Evaluator | Rule, test, expert, user, or model judge? |
| Instrument error | How can the evaluator itself be wrong? |
| Decision use | Which action changes if the metric moves? |

---
Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable.

## Evaluation dimensions
## Evaluation suite design

### 1. Task performance
A credible suite usually separates several purposes:

| Metric | Description | Measurement Method |
|---|---|---|
| **Task Completion Rate** | % of tasks completed without error | automated test suite against benchmark scenarios |
| **Goal Achievement Rate** | % of tasks where the final goal was met | human or LLM-judge evaluation |
| **Step Efficiency** | average steps taken vs. optimal path | comparison to expert baseline |
| **Tool Use Accuracy** | correct tool selection and parameter passing | automated verification |
| **Output Quality** | correctness and usefulness of agent outputs | rubric-based judging |
| Suite | Primary purpose |
|---|---|
| Operating-distribution sample | Estimate routine performance under stated workload assumptions |
| Critical-risk suite | Oversample rare, high-consequence conditions and hard gates |
| Boundary suite | Test ambiguity, refusal, escalation, and policy transitions |
| Fault-injection suite | Exercise tool, retrieval, network, permission, and state failures |
| Regression suite | Preserve previously discovered failures and important invariants |
| Exploratory suite | Discover new slices and failure modes; not used alone for release claims |

### 2. Safety and reliability
A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately.

| Metric | Description | Example target |
|---|---|---|
| **Hallucination Rate** | factually incorrect statements in outputs | < 2% on factual tasks, adjusted by risk tier |
| **Harmful Output Rate** | outputs violating safety guidelines | 0% hard limit for critical harms |
| **Refusal Appropriateness** | correct refusal of inappropriate requests | > 95%, adjusted by policy scope |
| **Prompt Injection Resistance** | resistance to adversarial inputs | tested via red-team suite |
| **Consistency** | same input produces the same output class | > 90% across repeated runs, adjusted by task type |
## Evaluator hierarchy

### 3. Operational reliability
Use the simplest evaluator capable of measuring the property.

| Metric | Description | Example target |
|---|---|---|
| **Availability** | uptime under normal and peak load | ≥ 99.5%, adjusted by service criticality |
| **P99 Latency** | 99th-percentile response time | ≤ configured SLA |
| **Error Rate** | % of requests ending in unhandled errors | < 0.1%, adjusted by workload |
| **Cost per Task** | token cost per successful task | tracked against budget |
| **Graceful Degradation** | behavior when tools or dependencies fail | tested via fault injection |
1. **Executable checks** for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior.
2. **Qualified human review** for domain judgment, consequential ambiguity, and disputed failures.
3. **Model judges** for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects.

### 4. Governance and auditability
A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process.

| Metric | Description | Pass Criterion |
|---|---|---|
| **Decision Traceability** | can every output be traced to inputs and decision evidence? | full trace available |
| **Tool Call Logging** | all external tool calls logged with parameters | 100% logged, with sensitive data controls |
| **Human Escalation Rate** | % of tasks escalated to human review | tracked with threshold defined |
| **Sensitive Data Handling** | PII or sensitive data not leaked in outputs or logs | 0 critical violations |
| **Override Capability** | operator can halt or override agent behavior | verified in testing |
## Repeated runs and uncertainty

---
For nondeterministic systems, report:

## Evaluation protocol
- attempts per scenario;
- run and seed policy where applicable;
- model, prompt, retrieval, tool, permission, and environment versions;
- missing or failed runs;
- scenario-level recurrence and variability;
- uncertainty intervals for estimated rates or means;
- coverage gaps and evaluator limitations that statistical intervals do not capture.

The target numbers above should not be treated as universal pass/fail thresholds. Use this protocol to make evaluations more defensible.
“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk.

| Protocol element | Recommended practice |
|---|---|
| Scenario set | Include normal, ambiguous, adversarial, tool-failure, escalation, and out-of-distribution tasks |
| Sample size | Define the minimum number of tasks per scenario class before testing starts |
| Risk tiering | Use stricter thresholds for safety-critical, regulated, or irreversible decisions |
| Repeated runs | Run the same scenario multiple times when model nondeterminism matters |
| Human adjudication | Use expert review for high-impact failures, disputed LLM-judge results, and policy-boundary cases |
| Severity levels | Classify findings as minor, major, critical, or blocker |
| Confidence | Report uncertainty, small sample caveats, and known benchmark limitations |
| Regression testing | Re-run the suite after model, prompt, tool, retrieval, or policy changes |
| Evidence retention | Store prompts, inputs, outputs, tool traces, evaluator notes, and sign-off decisions with appropriate data controls |

### Suggested release decision logic

| Decision | Use when |
|---|---|
| **Approve** | No blockers, critical controls pass, residual risks accepted by named owners |
| **Approve with conditions** | Minor or moderate issues exist but mitigations and monitoring are defined |
| **Hold** | Major issues require fixes before release |
| **Reject** | Critical safety, privacy, compliance, or operational failures are unresolved |
## Slice analysis

---
Predefine decision-relevant slices such as:

## Evaluation scenarios
- language or locale;
- task and user type;
- ambiguity level;
- tool authority and permission scope;
- data sensitivity;
- retrieval quality;
- accessibility need;
- dependency failure mode;
- safety or policy boundary.

### Standard benchmark suite
Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing.

| Scenario | Description | Difficulty |
|---|---|---|
| Single-step task | straightforward, well-defined task | Basic |
| Multi-step task | sequence of actions with intermediate dependencies | Intermediate |
| Ambiguous instruction | underspecified request requiring clarification | Intermediate |
| Tool failure | dependent tool returns an error | Intermediate |
| Adversarial input | prompt injection or manipulation attempt | Advanced |
| Novel situation | task outside the usual distribution | Advanced |
| Cascading failure | multiple dependencies fail in sequence | Advanced |
| Human escalation | agent should recognize the need to escalate | Advanced |
## Threshold and gate design

### Regulated-industry scenarios
This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective.

| Industry | Scenario | Key Evaluation Criteria |
|---|---|---|
| Healthcare | clinical documentation assistant | factual accuracy, refusal quality, low hallucination |
| Finance | transaction-analysis agent | compliance, auditability, no PII leakage |
| Insurance | claims-assessment agent | fairness, traceability, escalation quality |
| Legal | contract-review agent | citation accuracy, scope awareness, human escalation |
A threshold should record:

---
- decision owner and intended decision;
- baseline or comparator;
- affected population and harm model;
- measurement method and uncertainty;
- minimum sample and slice coverage;
- non-compensable hard gates;
- quality or operational targets that may be optimized;
- exception and residual-risk process;
- expiry or review trigger.

## Human-readable report template
A weighted aggregate must never hide a failed hard gate or an unresolved critical finding.

Use [`templates/evaluation-report.md`](templates/evaluation-report.md) for review meetings, sign-off packs, and stakeholder alignment.
## Decision semantics

## Machine-readable report schema
The schema and validator keep these concepts separate:

Use [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) when you want evaluation results to be structured for later validation, CI checks, dashboards, or trend analysis.
| Concept | Meaning |
|---|---|
| System risk tier | impact context for the evaluated system and use case |
| Scenario / finding severity | consequence of a specific failure |
| Blocker | unresolved condition that prevents the current decision |
| Required action | follow-up work accepted under a bounded conditional decision |
| Condition | explicit constraint attached to that decision |
| Residual risk | remaining uncertainty or harm accepted by a named authority |

The schema captures:
See [`docs/decision-semantics.md`](docs/decision-semantics.md) for the validator rules.

- metadata
- recommendation and blockers
- scope
- scenarios
- scorecard results
- findings
- oversight triggers
- release decision and sign-off
## Validation

---
The validator checks schema conformance and selected decision coherence:

## NIST AI RMF mapping
```bash
python tools/validate_evaluation_report.py \
schemas/evaluation-report.schema.json \
examples/sample-evaluation-report.json
```

Full mapping: [docs/nist-rmf-mapping.md](docs/nist-rmf-mapping.md)
It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable.

Key practitioner mapping:
## Failure review

- **MS.1**: AI risk identification through safety and adversarial scenarios
- **MS.3**: structured evaluation techniques and benchmark design
- **MS.5**: subgroup and fairness measurement
- **GV.4**: human oversight via escalation and override checks
Do not stop at the scorecard. For material failures, preserve:

---
- scenario and run identifiers;
- relevant input, context, and tool trace;
- evaluator result and disagreement;
- severity rationale;
- containment and remediation;
- root-cause hypothesis versus confirmed cause;
- regression test created;
- owner and disposition.

## Scope and disclaimer
The failure corpus often becomes more valuable than the original aggregate score.

This repository is shared in a personal capacity. It is not legal advice, compliance certification, regulatory approval, safety certification, benchmark certification, or official guidance from NIST, the EU, ISO, or any employer.
## Maturity and scope

References to NIST AI RMF, EU AI Act, safety controls, or industry obligations are practitioner mappings and examples. Always verify against official sources before using this framework for compliance, safety, or release decisions.
This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority.

---
References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance.

## Related repositories

| Repository | What it adds |
| Repository | Distinct role |
|---|---|
| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | oversight model, trust structure, accountability controls |
| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns and orchestration logic |
| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable execution and failure-mode examples |
| [`nist-rmf-guide`](https://github.com/simaba/nist-rmf-guide) | practitioner guide for broader RMF implementation |
| [`governance-playbook`](https://github.com/simaba/governance-playbook) | enterprise operating model beyond agent evaluation |
| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable bounded-agent behavior and failure paths |
| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns |
| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | authority, oversight, containment, and accountability |
| [`automotive-llm-eval-harness`](https://github.com/simaba/automotive-llm-eval-harness) | compact scorer for synthetic automotive evaluation artifacts |

---

*Maintained by [Sima Bagheri](https://github.com/simaba) · Connect on [LinkedIn](https://www.linkedin.com/in/simaba/)*
*Maintained by [Sima Bagheri](https://github.com/simaba).*
Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
296 changes: 162 additions & 134 deletions README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -2,194 +2,222 @@

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](LICENSE)
[![NIST AI RMF](https://img.shields.io/badge/NIST%20AI%20RMF-Informed-0055A4?style=flat-square)](https://airc.nist.gov/home)
[![Discussions](https://img.shields.io/badge/Discussions-Join-7289da?style=flat-square&logo=github)](https://github.com/simaba/agent-eval/discussions)

A structured framework for evaluating AI agents across task performance, safety, reliability, and governance dimensions, designed for regulated-industry deployment.
A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score.

---
The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds.

## Choose this repo when
## The central idea

Use this repository when you need **measurement and pass/fail logic** for AI agents:
An evaluation result is only useful when five things are explicit:

- what to measure
- which scenarios to test
- how to score performance, safety, and reliability
- what should appear in an evaluation report
- how to structure evaluation evidence for human review and future automation
1. **the decision it supports** — prototype, compare, pilot, expand, hold, roll back, or retire;
2. **the operating population it represents** — users, tasks, languages, tools, permissions, and exclusions;
3. **the measurement instrument** — executable rule, expert review, user outcome, or model judge;
4. **the uncertainty and coverage limits** — sample size, evaluator error, slice gaps, version drift, and scenario bias;
5. **the decision semantics** — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk.

Do **not** start here if you need the oversight model. Use [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance).
Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures.

Do **not** start here if you need orchestration patterns. Use [`agent-orchestration`](https://github.com/simaba/agent-orchestration).
## Start here

Do **not** start here if you want runnable behavior. Use [`agent-simulator`](https://github.com/simaba/agent-simulator).
| Artifact | Use it for |
|---|---|
| [`docs/evaluation-validity.md`](docs/evaluation-validity.md) | Designing evidence that can support a stated decision |
| [`docs/decision-semantics.md`](docs/decision-semantics.md) | Keeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct |
| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human review and sign-off discussion |
| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable evaluation records |
| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Fictional structured example |
| [`tools/validate_evaluation_report.py`](tools/validate_evaluation_report.py) | Schema and decision-coherence validation |

---
## What the framework evaluates

## Maturity
The dimensions below are categories for designing measurements. They are not automatically valid metrics.

This is a **practitioner evaluation framework**, not a certified benchmark suite. The metrics and targets are intended as starting points for structured evaluation. Real deployments should adapt thresholds to the domain, risk tier, user population, applicable regulation, and operational failure tolerance.
### Task and outcome performance

---
- end-to-end task completion;
- correctness against a defined reference or acceptance rule;
- user or workflow outcome;
- tool selection and parameter accuracy;
- correction and recovery behavior;
- unnecessary step or interaction cost.

### Safety, privacy, and policy behavior

- prohibited action or tool invocation;
- sensitive-data exposure or inappropriate retention;
- refusal, deferral, and escalation at defined boundaries;
- prompt-injection and untrusted-content handling;
- behavior under conflicting instructions;
- containment after a control or dependency failure.

### Operational reliability

## Practical artifacts
- latency and timeout behavior under stated load;
- unhandled and handled failure rates;
- retry, fallback, and circuit-breaker behavior;
- dependency and tool degradation;
- cost and resource consumption;
- recovery time and state consistency.

| Artifact | Use for |
### Governance and auditability

- decision and tool-call provenance;
- version and permission traceability;
- human override and stop mechanisms;
- evidence retention with privacy controls;
- owner, review, and residual-risk records;
- regression triggers after material change.

## Metrics need operational definitions

Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining.

For each metric, record:

| Field | Example question |
|---|---|
| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human-readable agent evaluation report |
| [`examples/sample-evaluation-report.md`](examples/sample-evaluation-report.md) | Filled generic Markdown example |
| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable report schema for structured evaluation evidence |
| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Filled JSON example aligned to the schema |
| Construct | What property are we trying to understand? |
| Observable rule | What exactly counts as success or failure? |
| Unit | Response, task, conversation, tool call, user, or time window? |
| Denominator | Which attempts are included, excluded, or missing? |
| Evaluator | Rule, test, expert, user, or model judge? |
| Instrument error | How can the evaluator itself be wrong? |
| Decision use | Which action changes if the metric moves? |

---
Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable.

## Evaluation dimensions
## Evaluation suite design

### 1. Task performance
A credible suite usually separates several purposes:

| Metric | Description | Measurement Method |
|---|---|---|
| **Task Completion Rate** | % of tasks completed without error | automated test suite against benchmark scenarios |
| **Goal Achievement Rate** | % of tasks where the final goal was met | human or LLM-judge evaluation |
| **Step Efficiency** | average steps taken vs. optimal path | comparison to expert baseline |
| **Tool Use Accuracy** | correct tool selection and parameter passing | automated verification |
| **Output Quality** | correctness and usefulness of agent outputs | rubric-based judging |
| Suite | Primary purpose |
|---|---|
| Operating-distribution sample | Estimate routine performance under stated workload assumptions |
| Critical-risk suite | Oversample rare, high-consequence conditions and hard gates |
| Boundary suite | Test ambiguity, refusal, escalation, and policy transitions |
| Fault-injection suite | Exercise tool, retrieval, network, permission, and state failures |
| Regression suite | Preserve previously discovered failures and important invariants |
| Exploratory suite | Discover new slices and failure modes; not used alone for release claims |

### 2. Safety and reliability
A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately.

| Metric | Description | Example target |
|---|---|---|
| **Hallucination Rate** | factually incorrect statements in outputs | < 2% on factual tasks, adjusted by risk tier |
| **Harmful Output Rate** | outputs violating safety guidelines | 0% hard limit for critical harms |
| **Refusal Appropriateness** | correct refusal of inappropriate requests | > 95%, adjusted by policy scope |
| **Prompt Injection Resistance** | resistance to adversarial inputs | tested via red-team suite |
| **Consistency** | same input produces the same output class | > 90% across repeated runs, adjusted by task type |
## Evaluator hierarchy

### 3. Operational reliability
Use the simplest evaluator capable of measuring the property.

| Metric | Description | Example target |
|---|---|---|
| **Availability** | uptime under normal and peak load | ≥ 99.5%, adjusted by service criticality |
| **P99 Latency** | 99th-percentile response time | ≤ configured SLA |
| **Error Rate** | % of requests ending in unhandled errors | < 0.1%, adjusted by workload |
| **Cost per Task** | token cost per successful task | tracked against budget |
| **Graceful Degradation** | behavior when tools or dependencies fail | tested via fault injection |
1. **Executable checks** for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior.
2. **Qualified human review** for domain judgment, consequential ambiguity, and disputed failures.
3. **Model judges** for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects.

### 4. Governance and auditability
A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process.

| Metric | Description | Pass Criterion |
|---|---|---|
| **Decision Traceability** | can every output be traced to inputs and decision evidence? | full trace available |
| **Tool Call Logging** | all external tool calls logged with parameters | 100% logged, with sensitive data controls |
| **Human Escalation Rate** | % of tasks escalated to human review | tracked with threshold defined |
| **Sensitive Data Handling** | PII or sensitive data not leaked in outputs or logs | 0 critical violations |
| **Override Capability** | operator can halt or override agent behavior | verified in testing |
## Repeated runs and uncertainty

---
For nondeterministic systems, report:

## Evaluation protocol
- attempts per scenario;
- run and seed policy where applicable;
- model, prompt, retrieval, tool, permission, and environment versions;
- missing or failed runs;
- scenario-level recurrence and variability;
- uncertainty intervals for estimated rates or means;
- coverage gaps and evaluator limitations that statistical intervals do not capture.

The target numbers above should not be treated as universal pass/fail thresholds. Use this protocol to make evaluations more defensible.
“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk.

| Protocol element | Recommended practice |
|---|---|
| Scenario set | Include normal, ambiguous, adversarial, tool-failure, escalation, and out-of-distribution tasks |
| Sample size | Define the minimum number of tasks per scenario class before testing starts |
| Risk tiering | Use stricter thresholds for safety-critical, regulated, or irreversible decisions |
| Repeated runs | Run the same scenario multiple times when model nondeterminism matters |
| Human adjudication | Use expert review for high-impact failures, disputed LLM-judge results, and policy-boundary cases |
| Severity levels | Classify findings as minor, major, critical, or blocker |
| Confidence | Report uncertainty, small sample caveats, and known benchmark limitations |
| Regression testing | Re-run the suite after model, prompt, tool, retrieval, or policy changes |
| Evidence retention | Store prompts, inputs, outputs, tool traces, evaluator notes, and sign-off decisions with appropriate data controls |

### Suggested release decision logic

| Decision | Use when |
|---|---|
| **Approve** | No blockers, critical controls pass, residual risks accepted by named owners |
| **Approve with conditions** | Minor or moderate issues exist but mitigations and monitoring are defined |
| **Hold** | Major issues require fixes before release |
| **Reject** | Critical safety, privacy, compliance, or operational failures are unresolved |
## Slice analysis

---
Predefine decision-relevant slices such as:

## Evaluation scenarios
- language or locale;
- task and user type;
- ambiguity level;
- tool authority and permission scope;
- data sensitivity;
- retrieval quality;
- accessibility need;
- dependency failure mode;
- safety or policy boundary.

### Standard benchmark suite
Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing.

| Scenario | Description | Difficulty |
|---|---|---|
| Single-step task | straightforward, well-defined task | Basic |
| Multi-step task | sequence of actions with intermediate dependencies | Intermediate |
| Ambiguous instruction | underspecified request requiring clarification | Intermediate |
| Tool failure | dependent tool returns an error | Intermediate |
| Adversarial input | prompt injection or manipulation attempt | Advanced |
| Novel situation | task outside the usual distribution | Advanced |
| Cascading failure | multiple dependencies fail in sequence | Advanced |
| Human escalation | agent should recognize the need to escalate | Advanced |
## Threshold and gate design

### Regulated-industry scenarios
This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective.

| Industry | Scenario | Key Evaluation Criteria |
|---|---|---|
| Healthcare | clinical documentation assistant | factual accuracy, refusal quality, low hallucination |
| Finance | transaction-analysis agent | compliance, auditability, no PII leakage |
| Insurance | claims-assessment agent | fairness, traceability, escalation quality |
| Legal | contract-review agent | citation accuracy, scope awareness, human escalation |
A threshold should record:

---
- decision owner and intended decision;
- baseline or comparator;
- affected population and harm model;
- measurement method and uncertainty;
- minimum sample and slice coverage;
- non-compensable hard gates;
- quality or operational targets that may be optimized;
- exception and residual-risk process;
- expiry or review trigger.

## Human-readable report template
A weighted aggregate must never hide a failed hard gate or an unresolved critical finding.

Use [`templates/evaluation-report.md`](templates/evaluation-report.md) for review meetings, sign-off packs, and stakeholder alignment.
## Decision semantics

## Machine-readable report schema
The schema and validator keep these concepts separate:

Use [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) when you want evaluation results to be structured for later validation, CI checks, dashboards, or trend analysis.
| Concept | Meaning |
|---|---|
| System risk tier | impact context for the evaluated system and use case |
| Scenario / finding severity | consequence of a specific failure |
| Blocker | unresolved condition that prevents the current decision |
| Required action | follow-up work accepted under a bounded conditional decision |
| Condition | explicit constraint attached to that decision |
| Residual risk | remaining uncertainty or harm accepted by a named authority |

The schema captures:
See [`docs/decision-semantics.md`](docs/decision-semantics.md) for the validator rules.

- metadata
- recommendation and blockers
- scope
- scenarios
- scorecard results
- findings
- oversight triggers
- release decision and sign-off
## Validation

---
The validator checks schema conformance and selected decision coherence:

## NIST AI RMF mapping
```bash
python tools/validate_evaluation_report.py \
schemas/evaluation-report.schema.json \
examples/sample-evaluation-report.json
```

Full mapping: [docs/nist-rmf-mapping.md](docs/nist-rmf-mapping.md)
It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable.

Key practitioner mapping:
## Failure review

- **MS.1**: AI risk identification through safety and adversarial scenarios
- **MS.3**: structured evaluation techniques and benchmark design
- **MS.5**: subgroup and fairness measurement
- **GV.4**: human oversight via escalation and override checks
Do not stop at the scorecard. For material failures, preserve:

---
- scenario and run identifiers;
- relevant input, context, and tool trace;
- evaluator result and disagreement;
- severity rationale;
- containment and remediation;
- root-cause hypothesis versus confirmed cause;
- regression test created;
- owner and disposition.

## Scope and disclaimer
The failure corpus often becomes more valuable than the original aggregate score.

This repository is shared in a personal capacity. It is not legal advice, compliance certification, regulatory approval, safety certification, benchmark certification, or official guidance from NIST, the EU, ISO, or any employer.
## Maturity and scope

References to NIST AI RMF, EU AI Act, safety controls, or industry obligations are practitioner mappings and examples. Always verify against official sources before using this framework for compliance, safety, or release decisions.
This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority.

---
References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance.

## Related repositories

| Repository | What it adds |
| Repository | Distinct role |
|---|---|
| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | oversight model, trust structure, accountability controls |
| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns and orchestration logic |
| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable execution and failure-mode examples |
| [`nist-rmf-guide`](https://github.com/simaba/nist-rmf-guide) | practitioner guide for broader RMF implementation |
| [`governance-playbook`](https://github.com/simaba/governance-playbook) | enterprise operating model beyond agent evaluation |
| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable bounded-agent behavior and failure paths |
| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns |
| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | authority, oversight, containment, and accountability |
| [`automotive-llm-eval-harness`](https://github.com/simaba/automotive-llm-eval-harness) | compact scorer for synthetic automotive evaluation artifacts |

---

*Maintained by [Sima Bagheri](https://github.com/simaba) · Connect on [LinkedIn](https://www.linkedin.com/in/simaba/)*
*Maintained by [Sima Bagheri](https://github.com/simaba).*
Loading
Loading
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
296 changes: 162 additions & 134 deletions README.md
Original file line numberDiff line numberDiff line change
Expand Up@@ -2,194 +2,222 @@

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](LICENSE)
[![NIST AI RMF](https://img.shields.io/badge/NIST%20AI%20RMF-Informed-0055A4?style=flat-square)](https://airc.nist.gov/home)
[![Discussions](https://img.shields.io/badge/Discussions-Join-7289da?style=flat-square&logo=github)](https://github.com/simaba/agent-eval/discussions)

A structured framework for evaluating AI agents across task performance, safety, reliability, and governance dimensions, designed for regulated-industry deployment.
A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score.

---
The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds.

## Choose this repo when
## The central idea

Use this repository when you need **measurement and pass/fail logic** for AI agents:
An evaluation result is only useful when five things are explicit:

- what to measure
- which scenarios to test
- how to score performance, safety, and reliability
- what should appear in an evaluation report
- how to structure evaluation evidence for human review and future automation
1. **the decision it supports** — prototype, compare, pilot, expand, hold, roll back, or retire;
2. **the operating population it represents** — users, tasks, languages, tools, permissions, and exclusions;
3. **the measurement instrument** — executable rule, expert review, user outcome, or model judge;
4. **the uncertainty and coverage limits** — sample size, evaluator error, slice gaps, version drift, and scenario bias;
5. **the decision semantics** — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk.

Do **not** start here if you need the oversight model. Use [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance).
Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures.

Do **not** start here if you need orchestration patterns. Use [`agent-orchestration`](https://github.com/simaba/agent-orchestration).
## Start here

Do **not** start here if you want runnable behavior. Use [`agent-simulator`](https://github.com/simaba/agent-simulator).
| Artifact | Use it for |
|---|---|
| [`docs/evaluation-validity.md`](docs/evaluation-validity.md) | Designing evidence that can support a stated decision |
| [`docs/decision-semantics.md`](docs/decision-semantics.md) | Keeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct |
| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human review and sign-off discussion |
| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable evaluation records |
| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Fictional structured example |
| [`tools/validate_evaluation_report.py`](tools/validate_evaluation_report.py) | Schema and decision-coherence validation |

---
## What the framework evaluates

## Maturity
The dimensions below are categories for designing measurements. They are not automatically valid metrics.

This is a **practitioner evaluation framework**, not a certified benchmark suite. The metrics and targets are intended as starting points for structured evaluation. Real deployments should adapt thresholds to the domain, risk tier, user population, applicable regulation, and operational failure tolerance.
### Task and outcome performance

---
- end-to-end task completion;
- correctness against a defined reference or acceptance rule;
- user or workflow outcome;
- tool selection and parameter accuracy;
- correction and recovery behavior;
- unnecessary step or interaction cost.

### Safety, privacy, and policy behavior

- prohibited action or tool invocation;
- sensitive-data exposure or inappropriate retention;
- refusal, deferral, and escalation at defined boundaries;
- prompt-injection and untrusted-content handling;
- behavior under conflicting instructions;
- containment after a control or dependency failure.

### Operational reliability

## Practical artifacts
- latency and timeout behavior under stated load;
- unhandled and handled failure rates;
- retry, fallback, and circuit-breaker behavior;
- dependency and tool degradation;
- cost and resource consumption;
- recovery time and state consistency.

| Artifact | Use for |
### Governance and auditability

- decision and tool-call provenance;
- version and permission traceability;
- human override and stop mechanisms;
- evidence retention with privacy controls;
- owner, review, and residual-risk records;
- regression triggers after material change.

## Metrics need operational definitions

Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining.

For each metric, record:

| Field | Example question |
|---|---|
| [`templates/evaluation-report.md`](templates/evaluation-report.md) | Human-readable agent evaluation report |
| [`examples/sample-evaluation-report.md`](examples/sample-evaluation-report.md) | Filled generic Markdown example |
| [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) | Machine-readable report schema for structured evaluation evidence |
| [`examples/sample-evaluation-report.json`](examples/sample-evaluation-report.json) | Filled JSON example aligned to the schema |
| Construct | What property are we trying to understand? |
| Observable rule | What exactly counts as success or failure? |
| Unit | Response, task, conversation, tool call, user, or time window? |
| Denominator | Which attempts are included, excluded, or missing? |
| Evaluator | Rule, test, expert, user, or model judge? |
| Instrument error | How can the evaluator itself be wrong? |
| Decision use | Which action changes if the metric moves? |

---
Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable.

## Evaluation dimensions
## Evaluation suite design

### 1. Task performance
A credible suite usually separates several purposes:

| Metric | Description | Measurement Method |
|---|---|---|
| **Task Completion Rate** | % of tasks completed without error | automated test suite against benchmark scenarios |
| **Goal Achievement Rate** | % of tasks where the final goal was met | human or LLM-judge evaluation |
| **Step Efficiency** | average steps taken vs. optimal path | comparison to expert baseline |
| **Tool Use Accuracy** | correct tool selection and parameter passing | automated verification |
| **Output Quality** | correctness and usefulness of agent outputs | rubric-based judging |
| Suite | Primary purpose |
|---|---|
| Operating-distribution sample | Estimate routine performance under stated workload assumptions |
| Critical-risk suite | Oversample rare, high-consequence conditions and hard gates |
| Boundary suite | Test ambiguity, refusal, escalation, and policy transitions |
| Fault-injection suite | Exercise tool, retrieval, network, permission, and state failures |
| Regression suite | Preserve previously discovered failures and important invariants |
| Exploratory suite | Discover new slices and failure modes; not used alone for release claims |

### 2. Safety and reliability
A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately.

| Metric | Description | Example target |
|---|---|---|
| **Hallucination Rate** | factually incorrect statements in outputs | < 2% on factual tasks, adjusted by risk tier |
| **Harmful Output Rate** | outputs violating safety guidelines | 0% hard limit for critical harms |
| **Refusal Appropriateness** | correct refusal of inappropriate requests | > 95%, adjusted by policy scope |
| **Prompt Injection Resistance** | resistance to adversarial inputs | tested via red-team suite |
| **Consistency** | same input produces the same output class | > 90% across repeated runs, adjusted by task type |
## Evaluator hierarchy

### 3. Operational reliability
Use the simplest evaluator capable of measuring the property.

| Metric | Description | Example target |
|---|---|---|
| **Availability** | uptime under normal and peak load | ≥ 99.5%, adjusted by service criticality |
| **P99 Latency** | 99th-percentile response time | ≤ configured SLA |
| **Error Rate** | % of requests ending in unhandled errors | < 0.1%, adjusted by workload |
| **Cost per Task** | token cost per successful task | tracked against budget |
| **Graceful Degradation** | behavior when tools or dependencies fail | tested via fault injection |
1. **Executable checks** for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior.
2. **Qualified human review** for domain judgment, consequential ambiguity, and disputed failures.
3. **Model judges** for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects.

### 4. Governance and auditability
A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process.

| Metric | Description | Pass Criterion |
|---|---|---|
| **Decision Traceability** | can every output be traced to inputs and decision evidence? | full trace available |
| **Tool Call Logging** | all external tool calls logged with parameters | 100% logged, with sensitive data controls |
| **Human Escalation Rate** | % of tasks escalated to human review | tracked with threshold defined |
| **Sensitive Data Handling** | PII or sensitive data not leaked in outputs or logs | 0 critical violations |
| **Override Capability** | operator can halt or override agent behavior | verified in testing |
## Repeated runs and uncertainty

---
For nondeterministic systems, report:

## Evaluation protocol
- attempts per scenario;
- run and seed policy where applicable;
- model, prompt, retrieval, tool, permission, and environment versions;
- missing or failed runs;
- scenario-level recurrence and variability;
- uncertainty intervals for estimated rates or means;
- coverage gaps and evaluator limitations that statistical intervals do not capture.

The target numbers above should not be treated as universal pass/fail thresholds. Use this protocol to make evaluations more defensible.
“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk.

| Protocol element | Recommended practice |
|---|---|
| Scenario set | Include normal, ambiguous, adversarial, tool-failure, escalation, and out-of-distribution tasks |
| Sample size | Define the minimum number of tasks per scenario class before testing starts |
| Risk tiering | Use stricter thresholds for safety-critical, regulated, or irreversible decisions |
| Repeated runs | Run the same scenario multiple times when model nondeterminism matters |
| Human adjudication | Use expert review for high-impact failures, disputed LLM-judge results, and policy-boundary cases |
| Severity levels | Classify findings as minor, major, critical, or blocker |
| Confidence | Report uncertainty, small sample caveats, and known benchmark limitations |
| Regression testing | Re-run the suite after model, prompt, tool, retrieval, or policy changes |
| Evidence retention | Store prompts, inputs, outputs, tool traces, evaluator notes, and sign-off decisions with appropriate data controls |

### Suggested release decision logic

| Decision | Use when |
|---|---|
| **Approve** | No blockers, critical controls pass, residual risks accepted by named owners |
| **Approve with conditions** | Minor or moderate issues exist but mitigations and monitoring are defined |
| **Hold** | Major issues require fixes before release |
| **Reject** | Critical safety, privacy, compliance, or operational failures are unresolved |
## Slice analysis

---
Predefine decision-relevant slices such as:

## Evaluation scenarios
- language or locale;
- task and user type;
- ambiguity level;
- tool authority and permission scope;
- data sensitivity;
- retrieval quality;
- accessibility need;
- dependency failure mode;
- safety or policy boundary.

### Standard benchmark suite
Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing.

| Scenario | Description | Difficulty |
|---|---|---|
| Single-step task | straightforward, well-defined task | Basic |
| Multi-step task | sequence of actions with intermediate dependencies | Intermediate |
| Ambiguous instruction | underspecified request requiring clarification | Intermediate |
| Tool failure | dependent tool returns an error | Intermediate |
| Adversarial input | prompt injection or manipulation attempt | Advanced |
| Novel situation | task outside the usual distribution | Advanced |
| Cascading failure | multiple dependencies fail in sequence | Advanced |
| Human escalation | agent should recognize the need to escalate | Advanced |
## Threshold and gate design

### Regulated-industry scenarios
This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective.

| Industry | Scenario | Key Evaluation Criteria |
|---|---|---|
| Healthcare | clinical documentation assistant | factual accuracy, refusal quality, low hallucination |
| Finance | transaction-analysis agent | compliance, auditability, no PII leakage |
| Insurance | claims-assessment agent | fairness, traceability, escalation quality |
| Legal | contract-review agent | citation accuracy, scope awareness, human escalation |
A threshold should record:

---
- decision owner and intended decision;
- baseline or comparator;
- affected population and harm model;
- measurement method and uncertainty;
- minimum sample and slice coverage;
- non-compensable hard gates;
- quality or operational targets that may be optimized;
- exception and residual-risk process;
- expiry or review trigger.

## Human-readable report template
A weighted aggregate must never hide a failed hard gate or an unresolved critical finding.

Use [`templates/evaluation-report.md`](templates/evaluation-report.md) for review meetings, sign-off packs, and stakeholder alignment.
## Decision semantics

## Machine-readable report schema
The schema and validator keep these concepts separate:

Use [`schemas/evaluation-report.schema.json`](schemas/evaluation-report.schema.json) when you want evaluation results to be structured for later validation, CI checks, dashboards, or trend analysis.
| Concept | Meaning |
|---|---|
| System risk tier | impact context for the evaluated system and use case |
| Scenario / finding severity | consequence of a specific failure |
| Blocker | unresolved condition that prevents the current decision |
| Required action | follow-up work accepted under a bounded conditional decision |
| Condition | explicit constraint attached to that decision |
| Residual risk | remaining uncertainty or harm accepted by a named authority |

The schema captures:
See [`docs/decision-semantics.md`](docs/decision-semantics.md) for the validator rules.

- metadata
- recommendation and blockers
- scope
- scenarios
- scorecard results
- findings
- oversight triggers
- release decision and sign-off
## Validation

---
The validator checks schema conformance and selected decision coherence:

## NIST AI RMF mapping
```bash
python tools/validate_evaluation_report.py \
schemas/evaluation-report.schema.json \
examples/sample-evaluation-report.json
```

Full mapping: [docs/nist-rmf-mapping.md](docs/nist-rmf-mapping.md)
It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable.

Key practitioner mapping:
## Failure review

- **MS.1**: AI risk identification through safety and adversarial scenarios
- **MS.3**: structured evaluation techniques and benchmark design
- **MS.5**: subgroup and fairness measurement
- **GV.4**: human oversight via escalation and override checks
Do not stop at the scorecard. For material failures, preserve:

---
- scenario and run identifiers;
- relevant input, context, and tool trace;
- evaluator result and disagreement;
- severity rationale;
- containment and remediation;
- root-cause hypothesis versus confirmed cause;
- regression test created;
- owner and disposition.

## Scope and disclaimer
The failure corpus often becomes more valuable than the original aggregate score.

This repository is shared in a personal capacity. It is not legal advice, compliance certification, regulatory approval, safety certification, benchmark certification, or official guidance from NIST, the EU, ISO, or any employer.
## Maturity and scope

References to NIST AI RMF, EU AI Act, safety controls, or industry obligations are practitioner mappings and examples. Always verify against official sources before using this framework for compliance, safety, or release decisions.
This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority.

---
References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance.

## Related repositories

| Repository | What it adds |
| Repository | Distinct role |
|---|---|
| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | oversight model, trust structure, accountability controls |
| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns and orchestration logic |
| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable execution and failure-mode examples |
| [`nist-rmf-guide`](https://github.com/simaba/nist-rmf-guide) | practitioner guide for broader RMF implementation |
| [`governance-playbook`](https://github.com/simaba/governance-playbook) | enterprise operating model beyond agent evaluation |
| [`agent-simulator`](https://github.com/simaba/agent-simulator) | runnable bounded-agent behavior and failure paths |
| [`agent-orchestration`](https://github.com/simaba/agent-orchestration) | control-flow patterns |
| [`multi-agent-governance`](https://github.com/simaba/multi-agent-governance) | authority, oversight, containment, and accountability |
| [`automotive-llm-eval-harness`](https://github.com/simaba/automotive-llm-eval-harness) | compact scorer for synthetic automotive evaluation artifacts |

---

*Maintained by [Sima Bagheri](https://github.com/simaba) · Connect on [LinkedIn](https://www.linkedin.com/in/simaba/)*
*Maintained by [Sima Bagheri](https://github.com/simaba).*
Loading
Loading