Repository files navigation

AI Agent Evaluation Framework

License: MITNIST AI RMF

A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score.

The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds.

The central idea

An evaluation result is only useful when five things are explicit:

  1. the decision it supports — prototype, compare, pilot, expand, hold, roll back, or retire;
  2. the operating population it represents — users, tasks, languages, tools, permissions, and exclusions;
  3. the measurement instrument — executable rule, expert review, user outcome, or model judge;
  4. the uncertainty and coverage limits — sample size, evaluator error, slice gaps, version drift, and scenario bias;
  5. the decision semantics — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk.

Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures.

Start here

ArtifactUse it for
docs/evaluation-validity.mdDesigning evidence that can support a stated decision
docs/decision-semantics.mdKeeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct
templates/evaluation-report.mdHuman review and sign-off discussion
schemas/evaluation-report.schema.jsonMachine-readable evaluation records
examples/sample-evaluation-report.jsonFictional structured example
tools/validate_evaluation_report.pySchema and decision-coherence validation

What the framework evaluates

The dimensions below are categories for designing measurements. They are not automatically valid metrics.

Task and outcome performance

  • end-to-end task completion;
  • correctness against a defined reference or acceptance rule;
  • user or workflow outcome;
  • tool selection and parameter accuracy;
  • correction and recovery behavior;
  • unnecessary step or interaction cost.

Safety, privacy, and policy behavior

  • prohibited action or tool invocation;
  • sensitive-data exposure or inappropriate retention;
  • refusal, deferral, and escalation at defined boundaries;
  • prompt-injection and untrusted-content handling;
  • behavior under conflicting instructions;
  • containment after a control or dependency failure.

Operational reliability

  • latency and timeout behavior under stated load;
  • unhandled and handled failure rates;
  • retry, fallback, and circuit-breaker behavior;
  • dependency and tool degradation;
  • cost and resource consumption;
  • recovery time and state consistency.

Governance and auditability

  • decision and tool-call provenance;
  • version and permission traceability;
  • human override and stop mechanisms;
  • evidence retention with privacy controls;
  • owner, review, and residual-risk records;
  • regression triggers after material change.

Metrics need operational definitions

Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining.

For each metric, record:

FieldExample question
ConstructWhat property are we trying to understand?
Observable ruleWhat exactly counts as success or failure?
UnitResponse, task, conversation, tool call, user, or time window?
DenominatorWhich attempts are included, excluded, or missing?
EvaluatorRule, test, expert, user, or model judge?
Instrument errorHow can the evaluator itself be wrong?
Decision useWhich action changes if the metric moves?

Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable.

Evaluation suite design

A credible suite usually separates several purposes:

SuitePrimary purpose
Operating-distribution sampleEstimate routine performance under stated workload assumptions
Critical-risk suiteOversample rare, high-consequence conditions and hard gates
Boundary suiteTest ambiguity, refusal, escalation, and policy transitions
Fault-injection suiteExercise tool, retrieval, network, permission, and state failures
Regression suitePreserve previously discovered failures and important invariants
Exploratory suiteDiscover new slices and failure modes; not used alone for release claims

A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately.

Evaluator hierarchy

Use the simplest evaluator capable of measuring the property.

  1. Executable checks for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior.
  2. Qualified human review for domain judgment, consequential ambiguity, and disputed failures.
  3. Model judges for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects.

A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process.

Repeated runs and uncertainty

For nondeterministic systems, report:

  • attempts per scenario;
  • run and seed policy where applicable;
  • model, prompt, retrieval, tool, permission, and environment versions;
  • missing or failed runs;
  • scenario-level recurrence and variability;
  • uncertainty intervals for estimated rates or means;
  • coverage gaps and evaluator limitations that statistical intervals do not capture.

“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk.

Slice analysis

Predefine decision-relevant slices such as:

  • language or locale;
  • task and user type;
  • ambiguity level;
  • tool authority and permission scope;
  • data sensitivity;
  • retrieval quality;
  • accessibility need;
  • dependency failure mode;
  • safety or policy boundary.

Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing.

Threshold and gate design

This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective.

A threshold should record:

  • decision owner and intended decision;
  • baseline or comparator;
  • affected population and harm model;
  • measurement method and uncertainty;
  • minimum sample and slice coverage;
  • non-compensable hard gates;
  • quality or operational targets that may be optimized;
  • exception and residual-risk process;
  • expiry or review trigger.

A weighted aggregate must never hide a failed hard gate or an unresolved critical finding.

Decision semantics

The schema and validator keep these concepts separate:

ConceptMeaning
System risk tierimpact context for the evaluated system and use case
Scenario / finding severityconsequence of a specific failure
Blockerunresolved condition that prevents the current decision
Required actionfollow-up work accepted under a bounded conditional decision
Conditionexplicit constraint attached to that decision
Residual riskremaining uncertainty or harm accepted by a named authority

See docs/decision-semantics.md for the validator rules.

Validation

The validator checks schema conformance and selected decision coherence:

python tools/validate_evaluation_report.py \
schemas/evaluation-report.schema.json \
examples/sample-evaluation-report.json

It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable.

Failure review

Do not stop at the scorecard. For material failures, preserve:

  • scenario and run identifiers;
  • relevant input, context, and tool trace;
  • evaluator result and disagreement;
  • severity rationale;
  • containment and remediation;
  • root-cause hypothesis versus confirmed cause;
  • regression test created;
  • owner and disposition.

The failure corpus often becomes more valuable than the original aggregate score.

Maturity and scope

This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority.

References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance.

Related repositories

RepositoryDistinct role
agent-simulatorrunnable bounded-agent behavior and failure paths
agent-orchestrationcontrol-flow patterns
multi-agent-governanceauthority, oversight, containment, and accountability
automotive-llm-eval-harnesscompact scorer for synthetic automotive evaluation artifacts

Maintained by Sima Bagheri.

About

A practical framework for measuring AI-agent performance, safety, reliability, and governance.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

AI Agent Evaluation Framework

License: MITNIST AI RMF

A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score.

The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds.

The central idea

An evaluation result is only useful when five things are explicit:

  1. the decision it supports — prototype, compare, pilot, expand, hold, roll back, or retire;
  2. the operating population it represents — users, tasks, languages, tools, permissions, and exclusions;
  3. the measurement instrument — executable rule, expert review, user outcome, or model judge;
  4. the uncertainty and coverage limits — sample size, evaluator error, slice gaps, version drift, and scenario bias;
  5. the decision semantics — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk.

Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures.

Start here

ArtifactUse it for
docs/evaluation-validity.mdDesigning evidence that can support a stated decision
docs/decision-semantics.mdKeeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct
templates/evaluation-report.mdHuman review and sign-off discussion
schemas/evaluation-report.schema.jsonMachine-readable evaluation records
examples/sample-evaluation-report.jsonFictional structured example
tools/validate_evaluation_report.pySchema and decision-coherence validation

What the framework evaluates

The dimensions below are categories for designing measurements. They are not automatically valid metrics.

Task and outcome performance

  • end-to-end task completion;
  • correctness against a defined reference or acceptance rule;
  • user or workflow outcome;
  • tool selection and parameter accuracy;
  • correction and recovery behavior;
  • unnecessary step or interaction cost.

Safety, privacy, and policy behavior

  • prohibited action or tool invocation;
  • sensitive-data exposure or inappropriate retention;
  • refusal, deferral, and escalation at defined boundaries;
  • prompt-injection and untrusted-content handling;
  • behavior under conflicting instructions;
  • containment after a control or dependency failure.

Operational reliability

  • latency and timeout behavior under stated load;
  • unhandled and handled failure rates;
  • retry, fallback, and circuit-breaker behavior;
  • dependency and tool degradation;
  • cost and resource consumption;
  • recovery time and state consistency.

Governance and auditability

  • decision and tool-call provenance;
  • version and permission traceability;
  • human override and stop mechanisms;
  • evidence retention with privacy controls;
  • owner, review, and residual-risk records;
  • regression triggers after material change.

Metrics need operational definitions

Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining.

For each metric, record:

FieldExample question
ConstructWhat property are we trying to understand?
Observable ruleWhat exactly counts as success or failure?
UnitResponse, task, conversation, tool call, user, or time window?
DenominatorWhich attempts are included, excluded, or missing?
EvaluatorRule, test, expert, user, or model judge?
Instrument errorHow can the evaluator itself be wrong?
Decision useWhich action changes if the metric moves?

Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable.

Evaluation suite design

A credible suite usually separates several purposes:

SuitePrimary purpose
Operating-distribution sampleEstimate routine performance under stated workload assumptions
Critical-risk suiteOversample rare, high-consequence conditions and hard gates
Boundary suiteTest ambiguity, refusal, escalation, and policy transitions
Fault-injection suiteExercise tool, retrieval, network, permission, and state failures
Regression suitePreserve previously discovered failures and important invariants
Exploratory suiteDiscover new slices and failure modes; not used alone for release claims

A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately.

Evaluator hierarchy

Use the simplest evaluator capable of measuring the property.

  1. Executable checks for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior.
  2. Qualified human review for domain judgment, consequential ambiguity, and disputed failures.
  3. Model judges for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects.

A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process.

Repeated runs and uncertainty

For nondeterministic systems, report:

  • attempts per scenario;
  • run and seed policy where applicable;
  • model, prompt, retrieval, tool, permission, and environment versions;
  • missing or failed runs;
  • scenario-level recurrence and variability;
  • uncertainty intervals for estimated rates or means;
  • coverage gaps and evaluator limitations that statistical intervals do not capture.

“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk.

Slice analysis

Predefine decision-relevant slices such as:

  • language or locale;
  • task and user type;
  • ambiguity level;
  • tool authority and permission scope;
  • data sensitivity;
  • retrieval quality;
  • accessibility need;
  • dependency failure mode;
  • safety or policy boundary.

Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing.

Threshold and gate design

This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective.

A threshold should record:

  • decision owner and intended decision;
  • baseline or comparator;
  • affected population and harm model;
  • measurement method and uncertainty;
  • minimum sample and slice coverage;
  • non-compensable hard gates;
  • quality or operational targets that may be optimized;
  • exception and residual-risk process;
  • expiry or review trigger.

A weighted aggregate must never hide a failed hard gate or an unresolved critical finding.

Decision semantics

The schema and validator keep these concepts separate:

ConceptMeaning
System risk tierimpact context for the evaluated system and use case
Scenario / finding severityconsequence of a specific failure
Blockerunresolved condition that prevents the current decision
Required actionfollow-up work accepted under a bounded conditional decision
Conditionexplicit constraint attached to that decision
Residual riskremaining uncertainty or harm accepted by a named authority

See docs/decision-semantics.md for the validator rules.

Validation

The validator checks schema conformance and selected decision coherence:

python tools/validate_evaluation_report.py \
schemas/evaluation-report.schema.json \
examples/sample-evaluation-report.json

It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable.

Failure review

Do not stop at the scorecard. For material failures, preserve:

  • scenario and run identifiers;
  • relevant input, context, and tool trace;
  • evaluator result and disagreement;
  • severity rationale;
  • containment and remediation;
  • root-cause hypothesis versus confirmed cause;
  • regression test created;
  • owner and disposition.

The failure corpus often becomes more valuable than the original aggregate score.

Maturity and scope

This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority.

References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance.

Related repositories

RepositoryDistinct role
agent-simulatorrunnable bounded-agent behavior and failure paths
agent-orchestrationcontrol-flow patterns
multi-agent-governanceauthority, oversight, containment, and accountability
automotive-llm-eval-harnesscompact scorer for synthetic automotive evaluation artifacts

Maintained by Sima Bagheri.

About

A practical framework for measuring AI-agent performance, safety, reliability, and governance.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

AI Agent Evaluation Framework

License: MITNIST AI RMF

A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score.

The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds.

The central idea

An evaluation result is only useful when five things are explicit:

  1. the decision it supports — prototype, compare, pilot, expand, hold, roll back, or retire;
  2. the operating population it represents — users, tasks, languages, tools, permissions, and exclusions;
  3. the measurement instrument — executable rule, expert review, user outcome, or model judge;
  4. the uncertainty and coverage limits — sample size, evaluator error, slice gaps, version drift, and scenario bias;
  5. the decision semantics — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk.

Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures.

Start here

ArtifactUse it for
docs/evaluation-validity.mdDesigning evidence that can support a stated decision
docs/decision-semantics.mdKeeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct
templates/evaluation-report.mdHuman review and sign-off discussion
schemas/evaluation-report.schema.jsonMachine-readable evaluation records
examples/sample-evaluation-report.jsonFictional structured example
tools/validate_evaluation_report.pySchema and decision-coherence validation

What the framework evaluates

The dimensions below are categories for designing measurements. They are not automatically valid metrics.

Task and outcome performance

  • end-to-end task completion;
  • correctness against a defined reference or acceptance rule;
  • user or workflow outcome;
  • tool selection and parameter accuracy;
  • correction and recovery behavior;
  • unnecessary step or interaction cost.

Safety, privacy, and policy behavior

  • prohibited action or tool invocation;
  • sensitive-data exposure or inappropriate retention;
  • refusal, deferral, and escalation at defined boundaries;
  • prompt-injection and untrusted-content handling;
  • behavior under conflicting instructions;
  • containment after a control or dependency failure.

Operational reliability

  • latency and timeout behavior under stated load;
  • unhandled and handled failure rates;
  • retry, fallback, and circuit-breaker behavior;
  • dependency and tool degradation;
  • cost and resource consumption;
  • recovery time and state consistency.

Governance and auditability

  • decision and tool-call provenance;
  • version and permission traceability;
  • human override and stop mechanisms;
  • evidence retention with privacy controls;
  • owner, review, and residual-risk records;
  • regression triggers after material change.

Metrics need operational definitions

Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining.

For each metric, record:

FieldExample question
ConstructWhat property are we trying to understand?
Observable ruleWhat exactly counts as success or failure?
UnitResponse, task, conversation, tool call, user, or time window?
DenominatorWhich attempts are included, excluded, or missing?
EvaluatorRule, test, expert, user, or model judge?
Instrument errorHow can the evaluator itself be wrong?
Decision useWhich action changes if the metric moves?

Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable.

Evaluation suite design

A credible suite usually separates several purposes:

SuitePrimary purpose
Operating-distribution sampleEstimate routine performance under stated workload assumptions
Critical-risk suiteOversample rare, high-consequence conditions and hard gates
Boundary suiteTest ambiguity, refusal, escalation, and policy transitions
Fault-injection suiteExercise tool, retrieval, network, permission, and state failures
Regression suitePreserve previously discovered failures and important invariants
Exploratory suiteDiscover new slices and failure modes; not used alone for release claims

A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately.

Evaluator hierarchy

Use the simplest evaluator capable of measuring the property.

  1. Executable checks for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior.
  2. Qualified human review for domain judgment, consequential ambiguity, and disputed failures.
  3. Model judges for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects.

A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process.

Repeated runs and uncertainty

For nondeterministic systems, report:

  • attempts per scenario;
  • run and seed policy where applicable;
  • model, prompt, retrieval, tool, permission, and environment versions;
  • missing or failed runs;
  • scenario-level recurrence and variability;
  • uncertainty intervals for estimated rates or means;
  • coverage gaps and evaluator limitations that statistical intervals do not capture.

“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk.

Slice analysis

Predefine decision-relevant slices such as:

  • language or locale;
  • task and user type;
  • ambiguity level;
  • tool authority and permission scope;
  • data sensitivity;
  • retrieval quality;
  • accessibility need;
  • dependency failure mode;
  • safety or policy boundary.

Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing.

Threshold and gate design

This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective.

A threshold should record:

  • decision owner and intended decision;
  • baseline or comparator;
  • affected population and harm model;
  • measurement method and uncertainty;
  • minimum sample and slice coverage;
  • non-compensable hard gates;
  • quality or operational targets that may be optimized;
  • exception and residual-risk process;
  • expiry or review trigger.

A weighted aggregate must never hide a failed hard gate or an unresolved critical finding.

Decision semantics

The schema and validator keep these concepts separate:

ConceptMeaning
System risk tierimpact context for the evaluated system and use case
Scenario / finding severityconsequence of a specific failure
Blockerunresolved condition that prevents the current decision
Required actionfollow-up work accepted under a bounded conditional decision
Conditionexplicit constraint attached to that decision
Residual riskremaining uncertainty or harm accepted by a named authority

See docs/decision-semantics.md for the validator rules.

Validation

The validator checks schema conformance and selected decision coherence:

python tools/validate_evaluation_report.py \
schemas/evaluation-report.schema.json \
examples/sample-evaluation-report.json

It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable.

Failure review

Do not stop at the scorecard. For material failures, preserve:

  • scenario and run identifiers;
  • relevant input, context, and tool trace;
  • evaluator result and disagreement;
  • severity rationale;
  • containment and remediation;
  • root-cause hypothesis versus confirmed cause;
  • regression test created;
  • owner and disposition.

The failure corpus often becomes more valuable than the original aggregate score.

Maturity and scope

This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority.

References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance.

Related repositories

RepositoryDistinct role
agent-simulatorrunnable bounded-agent behavior and failure paths
agent-orchestrationcontrol-flow patterns
multi-agent-governanceauthority, oversight, containment, and accountability
automotive-llm-eval-harnesscompact scorer for synthetic automotive evaluation artifacts

Maintained by Sima Bagheri.

About

A practical framework for measuring AI-agent performance, safety, reliability, and governance.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

AI Agent Evaluation Framework

License: MITNIST AI RMF

A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score.

The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds.

The central idea

An evaluation result is only useful when five things are explicit:

  1. the decision it supports — prototype, compare, pilot, expand, hold, roll back, or retire;
  2. the operating population it represents — users, tasks, languages, tools, permissions, and exclusions;
  3. the measurement instrument — executable rule, expert review, user outcome, or model judge;
  4. the uncertainty and coverage limits — sample size, evaluator error, slice gaps, version drift, and scenario bias;
  5. the decision semantics — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk.

Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures.

Start here

ArtifactUse it for
docs/evaluation-validity.mdDesigning evidence that can support a stated decision
docs/decision-semantics.mdKeeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct
templates/evaluation-report.mdHuman review and sign-off discussion
schemas/evaluation-report.schema.jsonMachine-readable evaluation records
examples/sample-evaluation-report.jsonFictional structured example
tools/validate_evaluation_report.pySchema and decision-coherence validation

What the framework evaluates

The dimensions below are categories for designing measurements. They are not automatically valid metrics.

Task and outcome performance

  • end-to-end task completion;
  • correctness against a defined reference or acceptance rule;
  • user or workflow outcome;
  • tool selection and parameter accuracy;
  • correction and recovery behavior;
  • unnecessary step or interaction cost.

Safety, privacy, and policy behavior

  • prohibited action or tool invocation;
  • sensitive-data exposure or inappropriate retention;
  • refusal, deferral, and escalation at defined boundaries;
  • prompt-injection and untrusted-content handling;
  • behavior under conflicting instructions;
  • containment after a control or dependency failure.

Operational reliability

  • latency and timeout behavior under stated load;
  • unhandled and handled failure rates;
  • retry, fallback, and circuit-breaker behavior;
  • dependency and tool degradation;
  • cost and resource consumption;
  • recovery time and state consistency.

Governance and auditability

  • decision and tool-call provenance;
  • version and permission traceability;
  • human override and stop mechanisms;
  • evidence retention with privacy controls;
  • owner, review, and residual-risk records;
  • regression triggers after material change.

Metrics need operational definitions

Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining.

For each metric, record:

FieldExample question
ConstructWhat property are we trying to understand?
Observable ruleWhat exactly counts as success or failure?
UnitResponse, task, conversation, tool call, user, or time window?
DenominatorWhich attempts are included, excluded, or missing?
EvaluatorRule, test, expert, user, or model judge?
Instrument errorHow can the evaluator itself be wrong?
Decision useWhich action changes if the metric moves?

Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable.

Evaluation suite design

A credible suite usually separates several purposes:

SuitePrimary purpose
Operating-distribution sampleEstimate routine performance under stated workload assumptions
Critical-risk suiteOversample rare, high-consequence conditions and hard gates
Boundary suiteTest ambiguity, refusal, escalation, and policy transitions
Fault-injection suiteExercise tool, retrieval, network, permission, and state failures
Regression suitePreserve previously discovered failures and important invariants
Exploratory suiteDiscover new slices and failure modes; not used alone for release claims

A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately.

Evaluator hierarchy

Use the simplest evaluator capable of measuring the property.

  1. Executable checks for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior.
  2. Qualified human review for domain judgment, consequential ambiguity, and disputed failures.
  3. Model judges for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects.

A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process.

Repeated runs and uncertainty

For nondeterministic systems, report:

  • attempts per scenario;
  • run and seed policy where applicable;
  • model, prompt, retrieval, tool, permission, and environment versions;
  • missing or failed runs;
  • scenario-level recurrence and variability;
  • uncertainty intervals for estimated rates or means;
  • coverage gaps and evaluator limitations that statistical intervals do not capture.

“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk.

Slice analysis

Predefine decision-relevant slices such as:

  • language or locale;
  • task and user type;
  • ambiguity level;
  • tool authority and permission scope;
  • data sensitivity;
  • retrieval quality;
  • accessibility need;
  • dependency failure mode;
  • safety or policy boundary.

Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing.

Threshold and gate design

This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective.

A threshold should record:

  • decision owner and intended decision;
  • baseline or comparator;
  • affected population and harm model;
  • measurement method and uncertainty;
  • minimum sample and slice coverage;
  • non-compensable hard gates;
  • quality or operational targets that may be optimized;
  • exception and residual-risk process;
  • expiry or review trigger.

A weighted aggregate must never hide a failed hard gate or an unresolved critical finding.

Decision semantics

The schema and validator keep these concepts separate:

ConceptMeaning
System risk tierimpact context for the evaluated system and use case
Scenario / finding severityconsequence of a specific failure
Blockerunresolved condition that prevents the current decision
Required actionfollow-up work accepted under a bounded conditional decision
Conditionexplicit constraint attached to that decision
Residual riskremaining uncertainty or harm accepted by a named authority

See docs/decision-semantics.md for the validator rules.

Validation

The validator checks schema conformance and selected decision coherence:

python tools/validate_evaluation_report.py \
schemas/evaluation-report.schema.json \
examples/sample-evaluation-report.json

It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable.

Failure review

Do not stop at the scorecard. For material failures, preserve:

  • scenario and run identifiers;
  • relevant input, context, and tool trace;
  • evaluator result and disagreement;
  • severity rationale;
  • containment and remediation;
  • root-cause hypothesis versus confirmed cause;
  • regression test created;
  • owner and disposition.

The failure corpus often becomes more valuable than the original aggregate score.

Maturity and scope

This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority.

References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance.

Related repositories

RepositoryDistinct role
agent-simulatorrunnable bounded-agent behavior and failure paths
agent-orchestrationcontrol-flow patterns
multi-agent-governanceauthority, oversight, containment, and accountability
automotive-llm-eval-harnesscompact scorer for synthetic automotive evaluation artifacts

Maintained by Sima Bagheri.

About

A practical framework for measuring AI-agent performance, safety, reliability, and governance.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

AI Agent Evaluation Framework

License: MITNIST AI RMF

A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score.

The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds.

The central idea

An evaluation result is only useful when five things are explicit:

  1. the decision it supports — prototype, compare, pilot, expand, hold, roll back, or retire;
  2. the operating population it represents — users, tasks, languages, tools, permissions, and exclusions;
  3. the measurement instrument — executable rule, expert review, user outcome, or model judge;
  4. the uncertainty and coverage limits — sample size, evaluator error, slice gaps, version drift, and scenario bias;
  5. the decision semantics — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk.

Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures.

Start here

ArtifactUse it for
docs/evaluation-validity.mdDesigning evidence that can support a stated decision
docs/decision-semantics.mdKeeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct
templates/evaluation-report.mdHuman review and sign-off discussion
schemas/evaluation-report.schema.jsonMachine-readable evaluation records
examples/sample-evaluation-report.jsonFictional structured example
tools/validate_evaluation_report.pySchema and decision-coherence validation

What the framework evaluates

The dimensions below are categories for designing measurements. They are not automatically valid metrics.

Task and outcome performance

  • end-to-end task completion;
  • correctness against a defined reference or acceptance rule;
  • user or workflow outcome;
  • tool selection and parameter accuracy;
  • correction and recovery behavior;
  • unnecessary step or interaction cost.

Safety, privacy, and policy behavior

  • prohibited action or tool invocation;
  • sensitive-data exposure or inappropriate retention;
  • refusal, deferral, and escalation at defined boundaries;
  • prompt-injection and untrusted-content handling;
  • behavior under conflicting instructions;
  • containment after a control or dependency failure.

Operational reliability

  • latency and timeout behavior under stated load;
  • unhandled and handled failure rates;
  • retry, fallback, and circuit-breaker behavior;
  • dependency and tool degradation;
  • cost and resource consumption;
  • recovery time and state consistency.

Governance and auditability

  • decision and tool-call provenance;
  • version and permission traceability;
  • human override and stop mechanisms;
  • evidence retention with privacy controls;
  • owner, review, and residual-risk records;
  • regression triggers after material change.

Metrics need operational definitions

Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining.

For each metric, record:

FieldExample question
ConstructWhat property are we trying to understand?
Observable ruleWhat exactly counts as success or failure?
UnitResponse, task, conversation, tool call, user, or time window?
DenominatorWhich attempts are included, excluded, or missing?
EvaluatorRule, test, expert, user, or model judge?
Instrument errorHow can the evaluator itself be wrong?
Decision useWhich action changes if the metric moves?

Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable.

Evaluation suite design

A credible suite usually separates several purposes:

SuitePrimary purpose
Operating-distribution sampleEstimate routine performance under stated workload assumptions
Critical-risk suiteOversample rare, high-consequence conditions and hard gates
Boundary suiteTest ambiguity, refusal, escalation, and policy transitions
Fault-injection suiteExercise tool, retrieval, network, permission, and state failures
Regression suitePreserve previously discovered failures and important invariants
Exploratory suiteDiscover new slices and failure modes; not used alone for release claims

A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately.

Evaluator hierarchy

Use the simplest evaluator capable of measuring the property.

  1. Executable checks for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior.
  2. Qualified human review for domain judgment, consequential ambiguity, and disputed failures.
  3. Model judges for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects.

A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process.

Repeated runs and uncertainty

For nondeterministic systems, report:

  • attempts per scenario;
  • run and seed policy where applicable;
  • model, prompt, retrieval, tool, permission, and environment versions;
  • missing or failed runs;
  • scenario-level recurrence and variability;
  • uncertainty intervals for estimated rates or means;
  • coverage gaps and evaluator limitations that statistical intervals do not capture.

“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk.

Slice analysis

Predefine decision-relevant slices such as:

  • language or locale;
  • task and user type;
  • ambiguity level;
  • tool authority and permission scope;
  • data sensitivity;
  • retrieval quality;
  • accessibility need;
  • dependency failure mode;
  • safety or policy boundary.

Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing.

Threshold and gate design

This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective.

A threshold should record:

  • decision owner and intended decision;
  • baseline or comparator;
  • affected population and harm model;
  • measurement method and uncertainty;
  • minimum sample and slice coverage;
  • non-compensable hard gates;
  • quality or operational targets that may be optimized;
  • exception and residual-risk process;
  • expiry or review trigger.

A weighted aggregate must never hide a failed hard gate or an unresolved critical finding.

Decision semantics

The schema and validator keep these concepts separate:

ConceptMeaning
System risk tierimpact context for the evaluated system and use case
Scenario / finding severityconsequence of a specific failure
Blockerunresolved condition that prevents the current decision
Required actionfollow-up work accepted under a bounded conditional decision
Conditionexplicit constraint attached to that decision
Residual riskremaining uncertainty or harm accepted by a named authority

See docs/decision-semantics.md for the validator rules.

Validation

The validator checks schema conformance and selected decision coherence:

python tools/validate_evaluation_report.py \
schemas/evaluation-report.schema.json \
examples/sample-evaluation-report.json

It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable.

Failure review

Do not stop at the scorecard. For material failures, preserve:

  • scenario and run identifiers;
  • relevant input, context, and tool trace;
  • evaluator result and disagreement;
  • severity rationale;
  • containment and remediation;
  • root-cause hypothesis versus confirmed cause;
  • regression test created;
  • owner and disposition.

The failure corpus often becomes more valuable than the original aggregate score.

Maturity and scope

This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority.

References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance.

Related repositories

RepositoryDistinct role
agent-simulatorrunnable bounded-agent behavior and failure paths
agent-orchestrationcontrol-flow patterns
multi-agent-governanceauthority, oversight, containment, and accountability
automotive-llm-eval-harnesscompact scorer for synthetic automotive evaluation artifacts

Maintained by Sima Bagheri.

About

A practical framework for measuring AI-agent performance, safety, reliability, and governance.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

AI Agent Evaluation Framework

License: MITNIST AI RMF

A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score.

The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds.

The central idea

An evaluation result is only useful when five things are explicit:

  1. the decision it supports — prototype, compare, pilot, expand, hold, roll back, or retire;
  2. the operating population it represents — users, tasks, languages, tools, permissions, and exclusions;
  3. the measurement instrument — executable rule, expert review, user outcome, or model judge;
  4. the uncertainty and coverage limits — sample size, evaluator error, slice gaps, version drift, and scenario bias;
  5. the decision semantics — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk.

Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures.

Start here

ArtifactUse it for
docs/evaluation-validity.mdDesigning evidence that can support a stated decision
docs/decision-semantics.mdKeeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct
templates/evaluation-report.mdHuman review and sign-off discussion
schemas/evaluation-report.schema.jsonMachine-readable evaluation records
examples/sample-evaluation-report.jsonFictional structured example
tools/validate_evaluation_report.pySchema and decision-coherence validation

What the framework evaluates

The dimensions below are categories for designing measurements. They are not automatically valid metrics.

Task and outcome performance

  • end-to-end task completion;
  • correctness against a defined reference or acceptance rule;
  • user or workflow outcome;
  • tool selection and parameter accuracy;
  • correction and recovery behavior;
  • unnecessary step or interaction cost.

Safety, privacy, and policy behavior

  • prohibited action or tool invocation;
  • sensitive-data exposure or inappropriate retention;
  • refusal, deferral, and escalation at defined boundaries;
  • prompt-injection and untrusted-content handling;
  • behavior under conflicting instructions;
  • containment after a control or dependency failure.

Operational reliability

  • latency and timeout behavior under stated load;
  • unhandled and handled failure rates;
  • retry, fallback, and circuit-breaker behavior;
  • dependency and tool degradation;
  • cost and resource consumption;
  • recovery time and state consistency.

Governance and auditability

  • decision and tool-call provenance;
  • version and permission traceability;
  • human override and stop mechanisms;
  • evidence retention with privacy controls;
  • owner, review, and residual-risk records;
  • regression triggers after material change.

Metrics need operational definitions

Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining.

For each metric, record:

FieldExample question
ConstructWhat property are we trying to understand?
Observable ruleWhat exactly counts as success or failure?
UnitResponse, task, conversation, tool call, user, or time window?
DenominatorWhich attempts are included, excluded, or missing?
EvaluatorRule, test, expert, user, or model judge?
Instrument errorHow can the evaluator itself be wrong?
Decision useWhich action changes if the metric moves?

Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable.

Evaluation suite design

A credible suite usually separates several purposes:

SuitePrimary purpose
Operating-distribution sampleEstimate routine performance under stated workload assumptions
Critical-risk suiteOversample rare, high-consequence conditions and hard gates
Boundary suiteTest ambiguity, refusal, escalation, and policy transitions
Fault-injection suiteExercise tool, retrieval, network, permission, and state failures
Regression suitePreserve previously discovered failures and important invariants
Exploratory suiteDiscover new slices and failure modes; not used alone for release claims

A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately.

Evaluator hierarchy

Use the simplest evaluator capable of measuring the property.

  1. Executable checks for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior.
  2. Qualified human review for domain judgment, consequential ambiguity, and disputed failures.
  3. Model judges for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects.

A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process.

Repeated runs and uncertainty

For nondeterministic systems, report:

  • attempts per scenario;
  • run and seed policy where applicable;
  • model, prompt, retrieval, tool, permission, and environment versions;
  • missing or failed runs;
  • scenario-level recurrence and variability;
  • uncertainty intervals for estimated rates or means;
  • coverage gaps and evaluator limitations that statistical intervals do not capture.

“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk.

Slice analysis

Predefine decision-relevant slices such as:

  • language or locale;
  • task and user type;
  • ambiguity level;
  • tool authority and permission scope;
  • data sensitivity;
  • retrieval quality;
  • accessibility need;
  • dependency failure mode;
  • safety or policy boundary.

Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing.

Threshold and gate design

This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective.

A threshold should record:

  • decision owner and intended decision;
  • baseline or comparator;
  • affected population and harm model;
  • measurement method and uncertainty;
  • minimum sample and slice coverage;
  • non-compensable hard gates;
  • quality or operational targets that may be optimized;
  • exception and residual-risk process;
  • expiry or review trigger.

A weighted aggregate must never hide a failed hard gate or an unresolved critical finding.

Decision semantics

The schema and validator keep these concepts separate:

ConceptMeaning
System risk tierimpact context for the evaluated system and use case
Scenario / finding severityconsequence of a specific failure
Blockerunresolved condition that prevents the current decision
Required actionfollow-up work accepted under a bounded conditional decision
Conditionexplicit constraint attached to that decision
Residual riskremaining uncertainty or harm accepted by a named authority

See docs/decision-semantics.md for the validator rules.

Validation

The validator checks schema conformance and selected decision coherence:

python tools/validate_evaluation_report.py \
schemas/evaluation-report.schema.json \
examples/sample-evaluation-report.json

It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable.

Failure review

Do not stop at the scorecard. For material failures, preserve:

  • scenario and run identifiers;
  • relevant input, context, and tool trace;
  • evaluator result and disagreement;
  • severity rationale;
  • containment and remediation;
  • root-cause hypothesis versus confirmed cause;
  • regression test created;
  • owner and disposition.

The failure corpus often becomes more valuable than the original aggregate score.

Maturity and scope

This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority.

References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance.

Related repositories

RepositoryDistinct role
agent-simulatorrunnable bounded-agent behavior and failure paths
agent-orchestrationcontrol-flow patterns
multi-agent-governanceauthority, oversight, containment, and accountability
automotive-llm-eval-harnesscompact scorer for synthetic automotive evaluation artifacts

Maintained by Sima Bagheri.

About

A practical framework for measuring AI-agent performance, safety, reliability, and governance.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

AI Agent Evaluation Framework

License: MITNIST AI RMF

A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score.

The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds.

The central idea

An evaluation result is only useful when five things are explicit:

  1. the decision it supports — prototype, compare, pilot, expand, hold, roll back, or retire;
  2. the operating population it represents — users, tasks, languages, tools, permissions, and exclusions;
  3. the measurement instrument — executable rule, expert review, user outcome, or model judge;
  4. the uncertainty and coverage limits — sample size, evaluator error, slice gaps, version drift, and scenario bias;
  5. the decision semantics — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk.

Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures.

Start here

ArtifactUse it for
docs/evaluation-validity.mdDesigning evidence that can support a stated decision
docs/decision-semantics.mdKeeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct
templates/evaluation-report.mdHuman review and sign-off discussion
schemas/evaluation-report.schema.jsonMachine-readable evaluation records
examples/sample-evaluation-report.jsonFictional structured example
tools/validate_evaluation_report.pySchema and decision-coherence validation

What the framework evaluates

The dimensions below are categories for designing measurements. They are not automatically valid metrics.

Task and outcome performance

  • end-to-end task completion;
  • correctness against a defined reference or acceptance rule;
  • user or workflow outcome;
  • tool selection and parameter accuracy;
  • correction and recovery behavior;
  • unnecessary step or interaction cost.

Safety, privacy, and policy behavior

  • prohibited action or tool invocation;
  • sensitive-data exposure or inappropriate retention;
  • refusal, deferral, and escalation at defined boundaries;
  • prompt-injection and untrusted-content handling;
  • behavior under conflicting instructions;
  • containment after a control or dependency failure.

Operational reliability

  • latency and timeout behavior under stated load;
  • unhandled and handled failure rates;
  • retry, fallback, and circuit-breaker behavior;
  • dependency and tool degradation;
  • cost and resource consumption;
  • recovery time and state consistency.

Governance and auditability

  • decision and tool-call provenance;
  • version and permission traceability;
  • human override and stop mechanisms;
  • evidence retention with privacy controls;
  • owner, review, and residual-risk records;
  • regression triggers after material change.

Metrics need operational definitions

Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining.

For each metric, record:

FieldExample question
ConstructWhat property are we trying to understand?
Observable ruleWhat exactly counts as success or failure?
UnitResponse, task, conversation, tool call, user, or time window?
DenominatorWhich attempts are included, excluded, or missing?
EvaluatorRule, test, expert, user, or model judge?
Instrument errorHow can the evaluator itself be wrong?
Decision useWhich action changes if the metric moves?

Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable.

Evaluation suite design

A credible suite usually separates several purposes:

SuitePrimary purpose
Operating-distribution sampleEstimate routine performance under stated workload assumptions
Critical-risk suiteOversample rare, high-consequence conditions and hard gates
Boundary suiteTest ambiguity, refusal, escalation, and policy transitions
Fault-injection suiteExercise tool, retrieval, network, permission, and state failures
Regression suitePreserve previously discovered failures and important invariants
Exploratory suiteDiscover new slices and failure modes; not used alone for release claims

A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately.

Evaluator hierarchy

Use the simplest evaluator capable of measuring the property.

  1. Executable checks for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior.
  2. Qualified human review for domain judgment, consequential ambiguity, and disputed failures.
  3. Model judges for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects.

A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process.

Repeated runs and uncertainty

For nondeterministic systems, report:

  • attempts per scenario;
  • run and seed policy where applicable;
  • model, prompt, retrieval, tool, permission, and environment versions;
  • missing or failed runs;
  • scenario-level recurrence and variability;
  • uncertainty intervals for estimated rates or means;
  • coverage gaps and evaluator limitations that statistical intervals do not capture.

“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk.

Slice analysis

Predefine decision-relevant slices such as:

  • language or locale;
  • task and user type;
  • ambiguity level;
  • tool authority and permission scope;
  • data sensitivity;
  • retrieval quality;
  • accessibility need;
  • dependency failure mode;
  • safety or policy boundary.

Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing.

Threshold and gate design

This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective.

A threshold should record:

  • decision owner and intended decision;
  • baseline or comparator;
  • affected population and harm model;
  • measurement method and uncertainty;
  • minimum sample and slice coverage;
  • non-compensable hard gates;
  • quality or operational targets that may be optimized;
  • exception and residual-risk process;
  • expiry or review trigger.

A weighted aggregate must never hide a failed hard gate or an unresolved critical finding.

Decision semantics

The schema and validator keep these concepts separate:

ConceptMeaning
System risk tierimpact context for the evaluated system and use case
Scenario / finding severityconsequence of a specific failure
Blockerunresolved condition that prevents the current decision
Required actionfollow-up work accepted under a bounded conditional decision
Conditionexplicit constraint attached to that decision
Residual riskremaining uncertainty or harm accepted by a named authority

See docs/decision-semantics.md for the validator rules.

Validation

The validator checks schema conformance and selected decision coherence:

python tools/validate_evaluation_report.py \
schemas/evaluation-report.schema.json \
examples/sample-evaluation-report.json

It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable.

Failure review

Do not stop at the scorecard. For material failures, preserve:

  • scenario and run identifiers;
  • relevant input, context, and tool trace;
  • evaluator result and disagreement;
  • severity rationale;
  • containment and remediation;
  • root-cause hypothesis versus confirmed cause;
  • regression test created;
  • owner and disposition.

The failure corpus often becomes more valuable than the original aggregate score.

Maturity and scope

This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority.

References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance.

Related repositories

RepositoryDistinct role
agent-simulatorrunnable bounded-agent behavior and failure paths
agent-orchestrationcontrol-flow patterns
multi-agent-governanceauthority, oversight, containment, and accountability
automotive-llm-eval-harnesscompact scorer for synthetic automotive evaluation artifacts

Maintained by Sima Bagheri.

About

A practical framework for measuring AI-agent performance, safety, reliability, and governance.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

AI Agent Evaluation Framework

License: MITNIST AI RMF

A practitioner framework for designing, recording, and reviewing agent evaluations without collapsing coverage, measurement validity, hard gates, and release judgment into one score.

The repository provides a human-readable report template, a machine-readable schema, a semantic validator, fictional examples, and guidance on evaluation validity. It is not a benchmark leaderboard or a source of universal pass thresholds.

The central idea

An evaluation result is only useful when five things are explicit:

  1. the decision it supports — prototype, compare, pilot, expand, hold, roll back, or retire;
  2. the operating population it represents — users, tasks, languages, tools, permissions, and exclusions;
  3. the measurement instrument — executable rule, expert review, user outcome, or model judge;
  4. the uncertainty and coverage limits — sample size, evaluator error, slice gaps, version drift, and scenario bias;
  5. the decision semantics — hard gate, quality threshold, required action, condition, blocker, or accepted residual risk.

Without those elements, a precise-looking score can be less informative than a small set of well-investigated failures.

Start here

ArtifactUse it for
docs/evaluation-validity.mdDesigning evidence that can support a stated decision
docs/decision-semantics.mdKeeping risk tier, finding severity, blockers, required actions, conditions, and release decisions distinct
templates/evaluation-report.mdHuman review and sign-off discussion
schemas/evaluation-report.schema.jsonMachine-readable evaluation records
examples/sample-evaluation-report.jsonFictional structured example
tools/validate_evaluation_report.pySchema and decision-coherence validation

What the framework evaluates

The dimensions below are categories for designing measurements. They are not automatically valid metrics.

Task and outcome performance

  • end-to-end task completion;
  • correctness against a defined reference or acceptance rule;
  • user or workflow outcome;
  • tool selection and parameter accuracy;
  • correction and recovery behavior;
  • unnecessary step or interaction cost.

Safety, privacy, and policy behavior

  • prohibited action or tool invocation;
  • sensitive-data exposure or inappropriate retention;
  • refusal, deferral, and escalation at defined boundaries;
  • prompt-injection and untrusted-content handling;
  • behavior under conflicting instructions;
  • containment after a control or dependency failure.

Operational reliability

  • latency and timeout behavior under stated load;
  • unhandled and handled failure rates;
  • retry, fallback, and circuit-breaker behavior;
  • dependency and tool degradation;
  • cost and resource consumption;
  • recovery time and state consistency.

Governance and auditability

  • decision and tool-call provenance;
  • version and permission traceability;
  • human override and stop mechanisms;
  • evidence retention with privacy controls;
  • owner, review, and residual-risk records;
  • regression triggers after material change.

Metrics need operational definitions

Terms such as “hallucination,” “helpfulness,” “safety,” “consistency,” and “quality” are not self-defining.

For each metric, record:

FieldExample question
ConstructWhat property are we trying to understand?
Observable ruleWhat exactly counts as success or failure?
UnitResponse, task, conversation, tool call, user, or time window?
DenominatorWhich attempts are included, excluded, or missing?
EvaluatorRule, test, expert, user, or model judge?
Instrument errorHow can the evaluator itself be wrong?
Decision useWhich action changes if the metric moves?

Do not reuse a percentage target across systems unless the population, unit, evaluator, severity model, and decision context are comparable.

Evaluation suite design

A credible suite usually separates several purposes:

SuitePrimary purpose
Operating-distribution sampleEstimate routine performance under stated workload assumptions
Critical-risk suiteOversample rare, high-consequence conditions and hard gates
Boundary suiteTest ambiguity, refusal, escalation, and policy transitions
Fault-injection suiteExercise tool, retrieval, network, permission, and state failures
Regression suitePreserve previously discovered failures and important invariants
Exploratory suiteDiscover new slices and failure modes; not used alone for release claims

A balanced challenge set may be good at finding failures but unsuitable for estimating production prevalence. Report coverage and prevalence-weighted performance separately.

Evaluator hierarchy

Use the simplest evaluator capable of measuring the property.

  1. Executable checks for directly observable invariants such as schema validity, unauthorized tool use, state mutation, citation presence, or timeout behavior.
  2. Qualified human review for domain judgment, consequential ambiguity, and disputed failures.
  3. Model judges for scalable rubric application only after calibration against representative human review and explicit testing for order, style, self-preference, and verbosity effects.

A model-judge result is an instrument reading, not ground truth. Record the judge model, prompt, configuration, version, agreement, and adjudication process.

Repeated runs and uncertainty

For nondeterministic systems, report:

  • attempts per scenario;
  • run and seed policy where applicable;
  • model, prompt, retrieval, tool, permission, and environment versions;
  • missing or failed runs;
  • scenario-level recurrence and variability;
  • uncertainty intervals for estimated rates or means;
  • coverage gaps and evaluator limitations that statistical intervals do not capture.

“Zero observed failures” must include the number of opportunities and a plain-language limit. Zero failures in a small sample is not evidence of zero risk.

Slice analysis

Predefine decision-relevant slices such as:

  • language or locale;
  • task and user type;
  • ambiguity level;
  • tool authority and permission scope;
  • data sensitivity;
  • retrieval quality;
  • accessibility need;
  • dependency failure mode;
  • safety or policy boundary.

Exploratory slicing can discover problems, but a post-hoc search across many slices should not be presented as confirmatory evidence without disclosure and follow-up testing.

Threshold and gate design

This repository intentionally does not prescribe universal values such as a fixed hallucination rate, refusal rate, availability target, or latency objective.

A threshold should record:

  • decision owner and intended decision;
  • baseline or comparator;
  • affected population and harm model;
  • measurement method and uncertainty;
  • minimum sample and slice coverage;
  • non-compensable hard gates;
  • quality or operational targets that may be optimized;
  • exception and residual-risk process;
  • expiry or review trigger.

A weighted aggregate must never hide a failed hard gate or an unresolved critical finding.

Decision semantics

The schema and validator keep these concepts separate:

ConceptMeaning
System risk tierimpact context for the evaluated system and use case
Scenario / finding severityconsequence of a specific failure
Blockerunresolved condition that prevents the current decision
Required actionfollow-up work accepted under a bounded conditional decision
Conditionexplicit constraint attached to that decision
Residual riskremaining uncertainty or harm accepted by a named authority

See docs/decision-semantics.md for the validator rules.

Validation

The validator checks schema conformance and selected decision coherence:

python tools/validate_evaluation_report.py \
schemas/evaluation-report.schema.json \
examples/sample-evaluation-report.json

It can catch contradictions such as an unconditional release with unresolved blockers. It cannot determine whether the benchmark is representative, the evaluator is valid, the threshold is justified, or the residual risk is acceptable.

Failure review

Do not stop at the scorecard. For material failures, preserve:

  • scenario and run identifiers;
  • relevant input, context, and tool trace;
  • evaluator result and disagreement;
  • severity rationale;
  • containment and remediation;
  • root-cause hypothesis versus confirmed cause;
  • regression test created;
  • owner and disposition.

The failure corpus often becomes more valuable than the original aggregate score.

Maturity and scope

This is a practitioner evaluation framework with working schema and semantic validation. It is useful for structured reviews, evaluation-plan design, fictional examples, and future automation. It is not a certified benchmark, safety case, regulatory assessment, or substitute for qualified domain, security, privacy, legal, compliance, or release authority.

References to NIST AI RMF or regulated-industry concerns are practitioner mappings. Verify official sources and adapt the framework to the actual system, population, jurisdiction, and failure tolerance.

Related repositories

RepositoryDistinct role
agent-simulatorrunnable bounded-agent behavior and failure paths
agent-orchestrationcontrol-flow patterns
multi-agent-governanceauthority, oversight, containment, and accountability
automotive-llm-eval-harnesscompact scorer for synthetic automotive evaluation artifacts

Maintained by Sima Bagheri.

About

A practical framework for measuring AI-agent performance, safety, reliability, and governance.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages