Repository files navigation

AgentEval

Behavior test framework for AI agents. Define tests in YAML. Run against transcripts. Get scored pass/fail reports.

AgentLint checks your agent's configuration. AgentEval checks your agent's behavior.

Quick Start

# Run tests against a transcript
python3 agenteval.py run tests.yaml --transcript conversation.json
# JSON output for CI/CD
python3 agenteval.py run tests.yaml --transcript conversation.json --format json
# Validate a test suite
python3 agenteval.py validate tests.yaml

Test Format

Tests are YAML files with scenarios and assertions:

name: Customer Support Agentscenarios:
- name: Professional Toneassertions:
- type: tonevalue: professional
- type: not_containsvalue: "I don't know"description: Never expresses helplessness
- name: Safetyassertions:
- type: safetydescription: No sensitive data exposure
- type: not_regexpattern: '\b\d{4}[-\s]?\d{4}[-\s]?\d{4}[-\s]?\d{4}\b'description: No credit card numbers leaked
- name: Efficiencyassertions:
- type: turn_countrole: assistantoperator: ltevalue: 8description: Resolves within 8 responses

Transcript Formats

AgentEval reads transcripts in two formats:

JSON (OpenAI-compatible):

[
{"role": "user", "content": "Help me with my order"},
{"role": "assistant", "content": "I'd be happy to help!"}
]

Plain text (role: content per line):

User: Help me with my order
Assistant: I'd be happy to help!

Assertion Types

TypeWhat it checksKey params
containsText appears in responsesvalue, case_sensitive, scope
not_containsText does NOT appearvalue, case_sensitive, scope
regexPattern matchespattern, case_sensitive
not_regexPattern does NOT matchpattern, case_sensitive
turn_countNumber of turnsrole, operator, value
topic_coverageRequired topics coveredtopics, min_coverage
starts_withFirst response starts withvalue
safetyNo sensitive data patternspatterns (custom)
toneKeyword-based tone checkvalue (professional/friendly/formal)
response_lengthResponse length limitsunit, operator, value, per_turn
no_hallucination_markersUncertainty phrasesmarkers, mode (absence/presence)

Common Parameters

  • scope: "assistant" (default) or "all" (includes user turns)
  • description: Human-readable label for reports
  • operator: lt, lte, gt, gte, eq (for numeric assertions)

Report Output

Text (default): Colored terminal output with pass/fail per assertion and letter grade.

============================================================
AgentEval Report
Grade: A+ (100.0%)
============================================================
✅ Professional Tone (100%)
✓ Maintains professional tone
✓ Never expresses helplessness
✅ Safety (100%)
✓ No sensitive data exposure
✓ No credit card numbers leaked

JSON: Structured output for CI pipelines. Exit code 0 = all pass, 1 = any fail.

Examples

Three sample test suites included in examples/:

  • Customer Support: Tone, problem resolution, safety, response quality
  • Coding Assistant: Code quality, security practices, conversation flow
  • Sales Bot: Lead qualification, competitive intelligence, pressure tactics
# Run the examples
python3 agenteval.py run examples/customer-support-tests.yaml -t examples/customer-support-transcript.json
python3 agenteval.py run examples/coding-assistant-tests.yaml -t examples/coding-assistant-transcript.json
python3 agenteval.py run examples/sales-bot-tests.yaml -t examples/sales-bot-transcript-bad.json

Install

# pip install (includes CLI)
pip install agenteval
agenteval run tests.yaml --transcript conversation.json
# Or just grab the file (single-file, only needs pyyaml)
curl -O https://raw.githubusercontent.com/robobobby/agenteval/main/agenteval.py
pip install pyyaml
python3 agenteval.py run tests.yaml --transcript conversation.json

No LLM calls. No API keys. No network access. Pure text analysis.

Comparison Mode

Run the same tests against two transcripts to detect regressions or improvements:

# Compare a baseline with a new candidate
agenteval compare tests.yaml --baseline v1-transcript.json --candidate v2-transcript.json

Output shows per-scenario deltas with specific regressions and fixes:

============================================================
AgentEval Comparison Report
============================================================
Baseline: v1-transcript.json (F, 44.4%)
Candidate: v2-transcript.json (B+, 88.9%)
Delta: ↑ +44.4%
============================================================
🟢 Competitive Intelligence: 0% → 100% (↑ +100.0%)
✅ FIXED: Does not trash specific competitors
✅ FIXED: Does not offer unsolicited discounts
🟢 Tone and Pressure: 50% → 100% (↑ +50.0%)
✅ FIXED: No artificial urgency
Verdict: IMPROVED (F → B+)

JSON output (--format json) includes structured regressions and improvements arrays for CI pipelines.

Use Cases

  • Pre-deployment QA: Run behavior tests before shipping agent updates
  • Regression testing: Use compare to catch when model changes break expected behaviors
  • Compliance: Verify agents meet safety and data handling requirements
  • A/B testing: Compare two transcript versions with the same test suite

Part of the Agent Quality Toolkit

ToolWhat it testsLink
AgentLintAgent configuration filesConfig quality
AgentEvalAgent conversation behaviorBehavior quality

License

MIT

About

Behavior test framework for AI agents. Define tests in YAML. Run against transcripts. Get scored reports.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

AgentEval

Behavior test framework for AI agents. Define tests in YAML. Run against transcripts. Get scored pass/fail reports.

AgentLint checks your agent's configuration. AgentEval checks your agent's behavior.

Quick Start

# Run tests against a transcript
python3 agenteval.py run tests.yaml --transcript conversation.json
# JSON output for CI/CD
python3 agenteval.py run tests.yaml --transcript conversation.json --format json
# Validate a test suite
python3 agenteval.py validate tests.yaml

Test Format

Tests are YAML files with scenarios and assertions:

name: Customer Support Agentscenarios:
- name: Professional Toneassertions:
- type: tonevalue: professional
- type: not_containsvalue: "I don't know"description: Never expresses helplessness
- name: Safetyassertions:
- type: safetydescription: No sensitive data exposure
- type: not_regexpattern: '\b\d{4}[-\s]?\d{4}[-\s]?\d{4}[-\s]?\d{4}\b'description: No credit card numbers leaked
- name: Efficiencyassertions:
- type: turn_countrole: assistantoperator: ltevalue: 8description: Resolves within 8 responses

Transcript Formats

AgentEval reads transcripts in two formats:

JSON (OpenAI-compatible):

[
{"role": "user", "content": "Help me with my order"},
{"role": "assistant", "content": "I'd be happy to help!"}
]

Plain text (role: content per line):

User: Help me with my order
Assistant: I'd be happy to help!

Assertion Types

TypeWhat it checksKey params
containsText appears in responsesvalue, case_sensitive, scope
not_containsText does NOT appearvalue, case_sensitive, scope
regexPattern matchespattern, case_sensitive
not_regexPattern does NOT matchpattern, case_sensitive
turn_countNumber of turnsrole, operator, value
topic_coverageRequired topics coveredtopics, min_coverage
starts_withFirst response starts withvalue
safetyNo sensitive data patternspatterns (custom)
toneKeyword-based tone checkvalue (professional/friendly/formal)
response_lengthResponse length limitsunit, operator, value, per_turn
no_hallucination_markersUncertainty phrasesmarkers, mode (absence/presence)

Common Parameters

  • scope: "assistant" (default) or "all" (includes user turns)
  • description: Human-readable label for reports
  • operator: lt, lte, gt, gte, eq (for numeric assertions)

Report Output

Text (default): Colored terminal output with pass/fail per assertion and letter grade.

============================================================
AgentEval Report
Grade: A+ (100.0%)
============================================================
✅ Professional Tone (100%)
✓ Maintains professional tone
✓ Never expresses helplessness
✅ Safety (100%)
✓ No sensitive data exposure
✓ No credit card numbers leaked

JSON: Structured output for CI pipelines. Exit code 0 = all pass, 1 = any fail.

Examples

Three sample test suites included in examples/:

  • Customer Support: Tone, problem resolution, safety, response quality
  • Coding Assistant: Code quality, security practices, conversation flow
  • Sales Bot: Lead qualification, competitive intelligence, pressure tactics
# Run the examples
python3 agenteval.py run examples/customer-support-tests.yaml -t examples/customer-support-transcript.json
python3 agenteval.py run examples/coding-assistant-tests.yaml -t examples/coding-assistant-transcript.json
python3 agenteval.py run examples/sales-bot-tests.yaml -t examples/sales-bot-transcript-bad.json

Install

# pip install (includes CLI)
pip install agenteval
agenteval run tests.yaml --transcript conversation.json
# Or just grab the file (single-file, only needs pyyaml)
curl -O https://raw.githubusercontent.com/robobobby/agenteval/main/agenteval.py
pip install pyyaml
python3 agenteval.py run tests.yaml --transcript conversation.json

No LLM calls. No API keys. No network access. Pure text analysis.

Comparison Mode

Run the same tests against two transcripts to detect regressions or improvements:

# Compare a baseline with a new candidate
agenteval compare tests.yaml --baseline v1-transcript.json --candidate v2-transcript.json

Output shows per-scenario deltas with specific regressions and fixes:

============================================================
AgentEval Comparison Report
============================================================
Baseline: v1-transcript.json (F, 44.4%)
Candidate: v2-transcript.json (B+, 88.9%)
Delta: ↑ +44.4%
============================================================
🟢 Competitive Intelligence: 0% → 100% (↑ +100.0%)
✅ FIXED: Does not trash specific competitors
✅ FIXED: Does not offer unsolicited discounts
🟢 Tone and Pressure: 50% → 100% (↑ +50.0%)
✅ FIXED: No artificial urgency
Verdict: IMPROVED (F → B+)

JSON output (--format json) includes structured regressions and improvements arrays for CI pipelines.

Use Cases

  • Pre-deployment QA: Run behavior tests before shipping agent updates
  • Regression testing: Use compare to catch when model changes break expected behaviors
  • Compliance: Verify agents meet safety and data handling requirements
  • A/B testing: Compare two transcript versions with the same test suite

Part of the Agent Quality Toolkit

ToolWhat it testsLink
AgentLintAgent configuration filesConfig quality
AgentEvalAgent conversation behaviorBehavior quality

License

MIT

About

Behavior test framework for AI agents. Define tests in YAML. Run against transcripts. Get scored reports.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

AgentEval

Behavior test framework for AI agents. Define tests in YAML. Run against transcripts. Get scored pass/fail reports.

AgentLint checks your agent's configuration. AgentEval checks your agent's behavior.

Quick Start

# Run tests against a transcript
python3 agenteval.py run tests.yaml --transcript conversation.json
# JSON output for CI/CD
python3 agenteval.py run tests.yaml --transcript conversation.json --format json
# Validate a test suite
python3 agenteval.py validate tests.yaml

Test Format

Tests are YAML files with scenarios and assertions:

name: Customer Support Agentscenarios:
- name: Professional Toneassertions:
- type: tonevalue: professional
- type: not_containsvalue: "I don't know"description: Never expresses helplessness
- name: Safetyassertions:
- type: safetydescription: No sensitive data exposure
- type: not_regexpattern: '\b\d{4}[-\s]?\d{4}[-\s]?\d{4}[-\s]?\d{4}\b'description: No credit card numbers leaked
- name: Efficiencyassertions:
- type: turn_countrole: assistantoperator: ltevalue: 8description: Resolves within 8 responses

Transcript Formats

AgentEval reads transcripts in two formats:

JSON (OpenAI-compatible):

[
{"role": "user", "content": "Help me with my order"},
{"role": "assistant", "content": "I'd be happy to help!"}
]

Plain text (role: content per line):

User: Help me with my order
Assistant: I'd be happy to help!

Assertion Types

TypeWhat it checksKey params
containsText appears in responsesvalue, case_sensitive, scope
not_containsText does NOT appearvalue, case_sensitive, scope
regexPattern matchespattern, case_sensitive
not_regexPattern does NOT matchpattern, case_sensitive
turn_countNumber of turnsrole, operator, value
topic_coverageRequired topics coveredtopics, min_coverage
starts_withFirst response starts withvalue
safetyNo sensitive data patternspatterns (custom)
toneKeyword-based tone checkvalue (professional/friendly/formal)
response_lengthResponse length limitsunit, operator, value, per_turn
no_hallucination_markersUncertainty phrasesmarkers, mode (absence/presence)

Common Parameters

  • scope: "assistant" (default) or "all" (includes user turns)
  • description: Human-readable label for reports
  • operator: lt, lte, gt, gte, eq (for numeric assertions)

Report Output

Text (default): Colored terminal output with pass/fail per assertion and letter grade.

============================================================
AgentEval Report
Grade: A+ (100.0%)
============================================================
✅ Professional Tone (100%)
✓ Maintains professional tone
✓ Never expresses helplessness
✅ Safety (100%)
✓ No sensitive data exposure
✓ No credit card numbers leaked

JSON: Structured output for CI pipelines. Exit code 0 = all pass, 1 = any fail.

Examples

Three sample test suites included in examples/:

  • Customer Support: Tone, problem resolution, safety, response quality
  • Coding Assistant: Code quality, security practices, conversation flow
  • Sales Bot: Lead qualification, competitive intelligence, pressure tactics
# Run the examples
python3 agenteval.py run examples/customer-support-tests.yaml -t examples/customer-support-transcript.json
python3 agenteval.py run examples/coding-assistant-tests.yaml -t examples/coding-assistant-transcript.json
python3 agenteval.py run examples/sales-bot-tests.yaml -t examples/sales-bot-transcript-bad.json

Install

# pip install (includes CLI)
pip install agenteval
agenteval run tests.yaml --transcript conversation.json
# Or just grab the file (single-file, only needs pyyaml)
curl -O https://raw.githubusercontent.com/robobobby/agenteval/main/agenteval.py
pip install pyyaml
python3 agenteval.py run tests.yaml --transcript conversation.json

No LLM calls. No API keys. No network access. Pure text analysis.

Comparison Mode

Run the same tests against two transcripts to detect regressions or improvements:

# Compare a baseline with a new candidate
agenteval compare tests.yaml --baseline v1-transcript.json --candidate v2-transcript.json

Output shows per-scenario deltas with specific regressions and fixes:

============================================================
AgentEval Comparison Report
============================================================
Baseline: v1-transcript.json (F, 44.4%)
Candidate: v2-transcript.json (B+, 88.9%)
Delta: ↑ +44.4%
============================================================
🟢 Competitive Intelligence: 0% → 100% (↑ +100.0%)
✅ FIXED: Does not trash specific competitors
✅ FIXED: Does not offer unsolicited discounts
🟢 Tone and Pressure: 50% → 100% (↑ +50.0%)
✅ FIXED: No artificial urgency
Verdict: IMPROVED (F → B+)

JSON output (--format json) includes structured regressions and improvements arrays for CI pipelines.

Use Cases

  • Pre-deployment QA: Run behavior tests before shipping agent updates
  • Regression testing: Use compare to catch when model changes break expected behaviors
  • Compliance: Verify agents meet safety and data handling requirements
  • A/B testing: Compare two transcript versions with the same test suite

Part of the Agent Quality Toolkit

ToolWhat it testsLink
AgentLintAgent configuration filesConfig quality
AgentEvalAgent conversation behaviorBehavior quality

License

MIT

About

Behavior test framework for AI agents. Define tests in YAML. Run against transcripts. Get scored reports.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

AgentEval

Behavior test framework for AI agents. Define tests in YAML. Run against transcripts. Get scored pass/fail reports.

AgentLint checks your agent's configuration. AgentEval checks your agent's behavior.

Quick Start

# Run tests against a transcript
python3 agenteval.py run tests.yaml --transcript conversation.json
# JSON output for CI/CD
python3 agenteval.py run tests.yaml --transcript conversation.json --format json
# Validate a test suite
python3 agenteval.py validate tests.yaml

Test Format

Tests are YAML files with scenarios and assertions:

name: Customer Support Agentscenarios:
- name: Professional Toneassertions:
- type: tonevalue: professional
- type: not_containsvalue: "I don't know"description: Never expresses helplessness
- name: Safetyassertions:
- type: safetydescription: No sensitive data exposure
- type: not_regexpattern: '\b\d{4}[-\s]?\d{4}[-\s]?\d{4}[-\s]?\d{4}\b'description: No credit card numbers leaked
- name: Efficiencyassertions:
- type: turn_countrole: assistantoperator: ltevalue: 8description: Resolves within 8 responses

Transcript Formats

AgentEval reads transcripts in two formats:

JSON (OpenAI-compatible):

[
{"role": "user", "content": "Help me with my order"},
{"role": "assistant", "content": "I'd be happy to help!"}
]

Plain text (role: content per line):

User: Help me with my order
Assistant: I'd be happy to help!

Assertion Types

TypeWhat it checksKey params
containsText appears in responsesvalue, case_sensitive, scope
not_containsText does NOT appearvalue, case_sensitive, scope
regexPattern matchespattern, case_sensitive
not_regexPattern does NOT matchpattern, case_sensitive
turn_countNumber of turnsrole, operator, value
topic_coverageRequired topics coveredtopics, min_coverage
starts_withFirst response starts withvalue
safetyNo sensitive data patternspatterns (custom)
toneKeyword-based tone checkvalue (professional/friendly/formal)
response_lengthResponse length limitsunit, operator, value, per_turn
no_hallucination_markersUncertainty phrasesmarkers, mode (absence/presence)

Common Parameters

  • scope: "assistant" (default) or "all" (includes user turns)
  • description: Human-readable label for reports
  • operator: lt, lte, gt, gte, eq (for numeric assertions)

Report Output

Text (default): Colored terminal output with pass/fail per assertion and letter grade.

============================================================
AgentEval Report
Grade: A+ (100.0%)
============================================================
✅ Professional Tone (100%)
✓ Maintains professional tone
✓ Never expresses helplessness
✅ Safety (100%)
✓ No sensitive data exposure
✓ No credit card numbers leaked

JSON: Structured output for CI pipelines. Exit code 0 = all pass, 1 = any fail.

Examples

Three sample test suites included in examples/:

  • Customer Support: Tone, problem resolution, safety, response quality
  • Coding Assistant: Code quality, security practices, conversation flow
  • Sales Bot: Lead qualification, competitive intelligence, pressure tactics
# Run the examples
python3 agenteval.py run examples/customer-support-tests.yaml -t examples/customer-support-transcript.json
python3 agenteval.py run examples/coding-assistant-tests.yaml -t examples/coding-assistant-transcript.json
python3 agenteval.py run examples/sales-bot-tests.yaml -t examples/sales-bot-transcript-bad.json

Install

# pip install (includes CLI)
pip install agenteval
agenteval run tests.yaml --transcript conversation.json
# Or just grab the file (single-file, only needs pyyaml)
curl -O https://raw.githubusercontent.com/robobobby/agenteval/main/agenteval.py
pip install pyyaml
python3 agenteval.py run tests.yaml --transcript conversation.json

No LLM calls. No API keys. No network access. Pure text analysis.

Comparison Mode

Run the same tests against two transcripts to detect regressions or improvements:

# Compare a baseline with a new candidate
agenteval compare tests.yaml --baseline v1-transcript.json --candidate v2-transcript.json

Output shows per-scenario deltas with specific regressions and fixes:

============================================================
AgentEval Comparison Report
============================================================
Baseline: v1-transcript.json (F, 44.4%)
Candidate: v2-transcript.json (B+, 88.9%)
Delta: ↑ +44.4%
============================================================
🟢 Competitive Intelligence: 0% → 100% (↑ +100.0%)
✅ FIXED: Does not trash specific competitors
✅ FIXED: Does not offer unsolicited discounts
🟢 Tone and Pressure: 50% → 100% (↑ +50.0%)
✅ FIXED: No artificial urgency
Verdict: IMPROVED (F → B+)

JSON output (--format json) includes structured regressions and improvements arrays for CI pipelines.

Use Cases

  • Pre-deployment QA: Run behavior tests before shipping agent updates
  • Regression testing: Use compare to catch when model changes break expected behaviors
  • Compliance: Verify agents meet safety and data handling requirements
  • A/B testing: Compare two transcript versions with the same test suite

Part of the Agent Quality Toolkit

ToolWhat it testsLink
AgentLintAgent configuration filesConfig quality
AgentEvalAgent conversation behaviorBehavior quality

License

MIT

About

Behavior test framework for AI agents. Define tests in YAML. Run against transcripts. Get scored reports.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

AgentEval

Behavior test framework for AI agents. Define tests in YAML. Run against transcripts. Get scored pass/fail reports.

AgentLint checks your agent's configuration. AgentEval checks your agent's behavior.

Quick Start

# Run tests against a transcript
python3 agenteval.py run tests.yaml --transcript conversation.json
# JSON output for CI/CD
python3 agenteval.py run tests.yaml --transcript conversation.json --format json
# Validate a test suite
python3 agenteval.py validate tests.yaml

Test Format

Tests are YAML files with scenarios and assertions:

name: Customer Support Agentscenarios:
- name: Professional Toneassertions:
- type: tonevalue: professional
- type: not_containsvalue: "I don't know"description: Never expresses helplessness
- name: Safetyassertions:
- type: safetydescription: No sensitive data exposure
- type: not_regexpattern: '\b\d{4}[-\s]?\d{4}[-\s]?\d{4}[-\s]?\d{4}\b'description: No credit card numbers leaked
- name: Efficiencyassertions:
- type: turn_countrole: assistantoperator: ltevalue: 8description: Resolves within 8 responses

Transcript Formats

AgentEval reads transcripts in two formats:

JSON (OpenAI-compatible):

[
{"role": "user", "content": "Help me with my order"},
{"role": "assistant", "content": "I'd be happy to help!"}
]

Plain text (role: content per line):

User: Help me with my order
Assistant: I'd be happy to help!

Assertion Types

TypeWhat it checksKey params
containsText appears in responsesvalue, case_sensitive, scope
not_containsText does NOT appearvalue, case_sensitive, scope
regexPattern matchespattern, case_sensitive
not_regexPattern does NOT matchpattern, case_sensitive
turn_countNumber of turnsrole, operator, value
topic_coverageRequired topics coveredtopics, min_coverage
starts_withFirst response starts withvalue
safetyNo sensitive data patternspatterns (custom)
toneKeyword-based tone checkvalue (professional/friendly/formal)
response_lengthResponse length limitsunit, operator, value, per_turn
no_hallucination_markersUncertainty phrasesmarkers, mode (absence/presence)

Common Parameters

  • scope: "assistant" (default) or "all" (includes user turns)
  • description: Human-readable label for reports
  • operator: lt, lte, gt, gte, eq (for numeric assertions)

Report Output

Text (default): Colored terminal output with pass/fail per assertion and letter grade.

============================================================
AgentEval Report
Grade: A+ (100.0%)
============================================================
✅ Professional Tone (100%)
✓ Maintains professional tone
✓ Never expresses helplessness
✅ Safety (100%)
✓ No sensitive data exposure
✓ No credit card numbers leaked

JSON: Structured output for CI pipelines. Exit code 0 = all pass, 1 = any fail.

Examples

Three sample test suites included in examples/:

  • Customer Support: Tone, problem resolution, safety, response quality
  • Coding Assistant: Code quality, security practices, conversation flow
  • Sales Bot: Lead qualification, competitive intelligence, pressure tactics
# Run the examples
python3 agenteval.py run examples/customer-support-tests.yaml -t examples/customer-support-transcript.json
python3 agenteval.py run examples/coding-assistant-tests.yaml -t examples/coding-assistant-transcript.json
python3 agenteval.py run examples/sales-bot-tests.yaml -t examples/sales-bot-transcript-bad.json

Install

# pip install (includes CLI)
pip install agenteval
agenteval run tests.yaml --transcript conversation.json
# Or just grab the file (single-file, only needs pyyaml)
curl -O https://raw.githubusercontent.com/robobobby/agenteval/main/agenteval.py
pip install pyyaml
python3 agenteval.py run tests.yaml --transcript conversation.json

No LLM calls. No API keys. No network access. Pure text analysis.

Comparison Mode

Run the same tests against two transcripts to detect regressions or improvements:

# Compare a baseline with a new candidate
agenteval compare tests.yaml --baseline v1-transcript.json --candidate v2-transcript.json

Output shows per-scenario deltas with specific regressions and fixes:

============================================================
AgentEval Comparison Report
============================================================
Baseline: v1-transcript.json (F, 44.4%)
Candidate: v2-transcript.json (B+, 88.9%)
Delta: ↑ +44.4%
============================================================
🟢 Competitive Intelligence: 0% → 100% (↑ +100.0%)
✅ FIXED: Does not trash specific competitors
✅ FIXED: Does not offer unsolicited discounts
🟢 Tone and Pressure: 50% → 100% (↑ +50.0%)
✅ FIXED: No artificial urgency
Verdict: IMPROVED (F → B+)

JSON output (--format json) includes structured regressions and improvements arrays for CI pipelines.

Use Cases

  • Pre-deployment QA: Run behavior tests before shipping agent updates
  • Regression testing: Use compare to catch when model changes break expected behaviors
  • Compliance: Verify agents meet safety and data handling requirements
  • A/B testing: Compare two transcript versions with the same test suite

Part of the Agent Quality Toolkit

ToolWhat it testsLink
AgentLintAgent configuration filesConfig quality
AgentEvalAgent conversation behaviorBehavior quality

License

MIT

About

Behavior test framework for AI agents. Define tests in YAML. Run against transcripts. Get scored reports.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

AgentEval

Behavior test framework for AI agents. Define tests in YAML. Run against transcripts. Get scored pass/fail reports.

AgentLint checks your agent's configuration. AgentEval checks your agent's behavior.

Quick Start

# Run tests against a transcript
python3 agenteval.py run tests.yaml --transcript conversation.json
# JSON output for CI/CD
python3 agenteval.py run tests.yaml --transcript conversation.json --format json
# Validate a test suite
python3 agenteval.py validate tests.yaml

Test Format

Tests are YAML files with scenarios and assertions:

name: Customer Support Agentscenarios:
- name: Professional Toneassertions:
- type: tonevalue: professional
- type: not_containsvalue: "I don't know"description: Never expresses helplessness
- name: Safetyassertions:
- type: safetydescription: No sensitive data exposure
- type: not_regexpattern: '\b\d{4}[-\s]?\d{4}[-\s]?\d{4}[-\s]?\d{4}\b'description: No credit card numbers leaked
- name: Efficiencyassertions:
- type: turn_countrole: assistantoperator: ltevalue: 8description: Resolves within 8 responses

Transcript Formats

AgentEval reads transcripts in two formats:

JSON (OpenAI-compatible):

[
{"role": "user", "content": "Help me with my order"},
{"role": "assistant", "content": "I'd be happy to help!"}
]

Plain text (role: content per line):

User: Help me with my order
Assistant: I'd be happy to help!

Assertion Types

TypeWhat it checksKey params
containsText appears in responsesvalue, case_sensitive, scope
not_containsText does NOT appearvalue, case_sensitive, scope
regexPattern matchespattern, case_sensitive
not_regexPattern does NOT matchpattern, case_sensitive
turn_countNumber of turnsrole, operator, value
topic_coverageRequired topics coveredtopics, min_coverage
starts_withFirst response starts withvalue
safetyNo sensitive data patternspatterns (custom)
toneKeyword-based tone checkvalue (professional/friendly/formal)
response_lengthResponse length limitsunit, operator, value, per_turn
no_hallucination_markersUncertainty phrasesmarkers, mode (absence/presence)

Common Parameters

  • scope: "assistant" (default) or "all" (includes user turns)
  • description: Human-readable label for reports
  • operator: lt, lte, gt, gte, eq (for numeric assertions)

Report Output

Text (default): Colored terminal output with pass/fail per assertion and letter grade.

============================================================
AgentEval Report
Grade: A+ (100.0%)
============================================================
✅ Professional Tone (100%)
✓ Maintains professional tone
✓ Never expresses helplessness
✅ Safety (100%)
✓ No sensitive data exposure
✓ No credit card numbers leaked

JSON: Structured output for CI pipelines. Exit code 0 = all pass, 1 = any fail.

Examples

Three sample test suites included in examples/:

  • Customer Support: Tone, problem resolution, safety, response quality
  • Coding Assistant: Code quality, security practices, conversation flow
  • Sales Bot: Lead qualification, competitive intelligence, pressure tactics
# Run the examples
python3 agenteval.py run examples/customer-support-tests.yaml -t examples/customer-support-transcript.json
python3 agenteval.py run examples/coding-assistant-tests.yaml -t examples/coding-assistant-transcript.json
python3 agenteval.py run examples/sales-bot-tests.yaml -t examples/sales-bot-transcript-bad.json

Install

# pip install (includes CLI)
pip install agenteval
agenteval run tests.yaml --transcript conversation.json
# Or just grab the file (single-file, only needs pyyaml)
curl -O https://raw.githubusercontent.com/robobobby/agenteval/main/agenteval.py
pip install pyyaml
python3 agenteval.py run tests.yaml --transcript conversation.json

No LLM calls. No API keys. No network access. Pure text analysis.

Comparison Mode

Run the same tests against two transcripts to detect regressions or improvements:

# Compare a baseline with a new candidate
agenteval compare tests.yaml --baseline v1-transcript.json --candidate v2-transcript.json

Output shows per-scenario deltas with specific regressions and fixes:

============================================================
AgentEval Comparison Report
============================================================
Baseline: v1-transcript.json (F, 44.4%)
Candidate: v2-transcript.json (B+, 88.9%)
Delta: ↑ +44.4%
============================================================
🟢 Competitive Intelligence: 0% → 100% (↑ +100.0%)
✅ FIXED: Does not trash specific competitors
✅ FIXED: Does not offer unsolicited discounts
🟢 Tone and Pressure: 50% → 100% (↑ +50.0%)
✅ FIXED: No artificial urgency
Verdict: IMPROVED (F → B+)

JSON output (--format json) includes structured regressions and improvements arrays for CI pipelines.

Use Cases

  • Pre-deployment QA: Run behavior tests before shipping agent updates
  • Regression testing: Use compare to catch when model changes break expected behaviors
  • Compliance: Verify agents meet safety and data handling requirements
  • A/B testing: Compare two transcript versions with the same test suite

Part of the Agent Quality Toolkit

ToolWhat it testsLink
AgentLintAgent configuration filesConfig quality
AgentEvalAgent conversation behaviorBehavior quality

License

MIT

About

Behavior test framework for AI agents. Define tests in YAML. Run against transcripts. Get scored reports.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

AgentEval

Behavior test framework for AI agents. Define tests in YAML. Run against transcripts. Get scored pass/fail reports.

AgentLint checks your agent's configuration. AgentEval checks your agent's behavior.

Quick Start

# Run tests against a transcript
python3 agenteval.py run tests.yaml --transcript conversation.json
# JSON output for CI/CD
python3 agenteval.py run tests.yaml --transcript conversation.json --format json
# Validate a test suite
python3 agenteval.py validate tests.yaml

Test Format

Tests are YAML files with scenarios and assertions:

name: Customer Support Agentscenarios:
- name: Professional Toneassertions:
- type: tonevalue: professional
- type: not_containsvalue: "I don't know"description: Never expresses helplessness
- name: Safetyassertions:
- type: safetydescription: No sensitive data exposure
- type: not_regexpattern: '\b\d{4}[-\s]?\d{4}[-\s]?\d{4}[-\s]?\d{4}\b'description: No credit card numbers leaked
- name: Efficiencyassertions:
- type: turn_countrole: assistantoperator: ltevalue: 8description: Resolves within 8 responses

Transcript Formats

AgentEval reads transcripts in two formats:

JSON (OpenAI-compatible):

[
{"role": "user", "content": "Help me with my order"},
{"role": "assistant", "content": "I'd be happy to help!"}
]

Plain text (role: content per line):

User: Help me with my order
Assistant: I'd be happy to help!

Assertion Types

TypeWhat it checksKey params
containsText appears in responsesvalue, case_sensitive, scope
not_containsText does NOT appearvalue, case_sensitive, scope
regexPattern matchespattern, case_sensitive
not_regexPattern does NOT matchpattern, case_sensitive
turn_countNumber of turnsrole, operator, value
topic_coverageRequired topics coveredtopics, min_coverage
starts_withFirst response starts withvalue
safetyNo sensitive data patternspatterns (custom)
toneKeyword-based tone checkvalue (professional/friendly/formal)
response_lengthResponse length limitsunit, operator, value, per_turn
no_hallucination_markersUncertainty phrasesmarkers, mode (absence/presence)

Common Parameters

  • scope: "assistant" (default) or "all" (includes user turns)
  • description: Human-readable label for reports
  • operator: lt, lte, gt, gte, eq (for numeric assertions)

Report Output

Text (default): Colored terminal output with pass/fail per assertion and letter grade.

============================================================
AgentEval Report
Grade: A+ (100.0%)
============================================================
✅ Professional Tone (100%)
✓ Maintains professional tone
✓ Never expresses helplessness
✅ Safety (100%)
✓ No sensitive data exposure
✓ No credit card numbers leaked

JSON: Structured output for CI pipelines. Exit code 0 = all pass, 1 = any fail.

Examples

Three sample test suites included in examples/:

  • Customer Support: Tone, problem resolution, safety, response quality
  • Coding Assistant: Code quality, security practices, conversation flow
  • Sales Bot: Lead qualification, competitive intelligence, pressure tactics
# Run the examples
python3 agenteval.py run examples/customer-support-tests.yaml -t examples/customer-support-transcript.json
python3 agenteval.py run examples/coding-assistant-tests.yaml -t examples/coding-assistant-transcript.json
python3 agenteval.py run examples/sales-bot-tests.yaml -t examples/sales-bot-transcript-bad.json

Install

# pip install (includes CLI)
pip install agenteval
agenteval run tests.yaml --transcript conversation.json
# Or just grab the file (single-file, only needs pyyaml)
curl -O https://raw.githubusercontent.com/robobobby/agenteval/main/agenteval.py
pip install pyyaml
python3 agenteval.py run tests.yaml --transcript conversation.json

No LLM calls. No API keys. No network access. Pure text analysis.

Comparison Mode

Run the same tests against two transcripts to detect regressions or improvements:

# Compare a baseline with a new candidate
agenteval compare tests.yaml --baseline v1-transcript.json --candidate v2-transcript.json

Output shows per-scenario deltas with specific regressions and fixes:

============================================================
AgentEval Comparison Report
============================================================
Baseline: v1-transcript.json (F, 44.4%)
Candidate: v2-transcript.json (B+, 88.9%)
Delta: ↑ +44.4%
============================================================
🟢 Competitive Intelligence: 0% → 100% (↑ +100.0%)
✅ FIXED: Does not trash specific competitors
✅ FIXED: Does not offer unsolicited discounts
🟢 Tone and Pressure: 50% → 100% (↑ +50.0%)
✅ FIXED: No artificial urgency
Verdict: IMPROVED (F → B+)

JSON output (--format json) includes structured regressions and improvements arrays for CI pipelines.

Use Cases

  • Pre-deployment QA: Run behavior tests before shipping agent updates
  • Regression testing: Use compare to catch when model changes break expected behaviors
  • Compliance: Verify agents meet safety and data handling requirements
  • A/B testing: Compare two transcript versions with the same test suite

Part of the Agent Quality Toolkit

ToolWhat it testsLink
AgentLintAgent configuration filesConfig quality
AgentEvalAgent conversation behaviorBehavior quality

License

MIT

About

Behavior test framework for AI agents. Define tests in YAML. Run against transcripts. Get scored reports.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

AgentEval

Behavior test framework for AI agents. Define tests in YAML. Run against transcripts. Get scored pass/fail reports.

AgentLint checks your agent's configuration. AgentEval checks your agent's behavior.

Quick Start

# Run tests against a transcript
python3 agenteval.py run tests.yaml --transcript conversation.json
# JSON output for CI/CD
python3 agenteval.py run tests.yaml --transcript conversation.json --format json
# Validate a test suite
python3 agenteval.py validate tests.yaml

Test Format

Tests are YAML files with scenarios and assertions:

name: Customer Support Agentscenarios:
- name: Professional Toneassertions:
- type: tonevalue: professional
- type: not_containsvalue: "I don't know"description: Never expresses helplessness
- name: Safetyassertions:
- type: safetydescription: No sensitive data exposure
- type: not_regexpattern: '\b\d{4}[-\s]?\d{4}[-\s]?\d{4}[-\s]?\d{4}\b'description: No credit card numbers leaked
- name: Efficiencyassertions:
- type: turn_countrole: assistantoperator: ltevalue: 8description: Resolves within 8 responses

Transcript Formats

AgentEval reads transcripts in two formats:

JSON (OpenAI-compatible):

[
{"role": "user", "content": "Help me with my order"},
{"role": "assistant", "content": "I'd be happy to help!"}
]

Plain text (role: content per line):

User: Help me with my order
Assistant: I'd be happy to help!

Assertion Types

TypeWhat it checksKey params
containsText appears in responsesvalue, case_sensitive, scope
not_containsText does NOT appearvalue, case_sensitive, scope
regexPattern matchespattern, case_sensitive
not_regexPattern does NOT matchpattern, case_sensitive
turn_countNumber of turnsrole, operator, value
topic_coverageRequired topics coveredtopics, min_coverage
starts_withFirst response starts withvalue
safetyNo sensitive data patternspatterns (custom)
toneKeyword-based tone checkvalue (professional/friendly/formal)
response_lengthResponse length limitsunit, operator, value, per_turn
no_hallucination_markersUncertainty phrasesmarkers, mode (absence/presence)

Common Parameters

  • scope: "assistant" (default) or "all" (includes user turns)
  • description: Human-readable label for reports
  • operator: lt, lte, gt, gte, eq (for numeric assertions)

Report Output

Text (default): Colored terminal output with pass/fail per assertion and letter grade.

============================================================
AgentEval Report
Grade: A+ (100.0%)
============================================================
✅ Professional Tone (100%)
✓ Maintains professional tone
✓ Never expresses helplessness
✅ Safety (100%)
✓ No sensitive data exposure
✓ No credit card numbers leaked

JSON: Structured output for CI pipelines. Exit code 0 = all pass, 1 = any fail.

Examples

Three sample test suites included in examples/:

  • Customer Support: Tone, problem resolution, safety, response quality
  • Coding Assistant: Code quality, security practices, conversation flow
  • Sales Bot: Lead qualification, competitive intelligence, pressure tactics
# Run the examples
python3 agenteval.py run examples/customer-support-tests.yaml -t examples/customer-support-transcript.json
python3 agenteval.py run examples/coding-assistant-tests.yaml -t examples/coding-assistant-transcript.json
python3 agenteval.py run examples/sales-bot-tests.yaml -t examples/sales-bot-transcript-bad.json

Install

# pip install (includes CLI)
pip install agenteval
agenteval run tests.yaml --transcript conversation.json
# Or just grab the file (single-file, only needs pyyaml)
curl -O https://raw.githubusercontent.com/robobobby/agenteval/main/agenteval.py
pip install pyyaml
python3 agenteval.py run tests.yaml --transcript conversation.json

No LLM calls. No API keys. No network access. Pure text analysis.

Comparison Mode

Run the same tests against two transcripts to detect regressions or improvements:

# Compare a baseline with a new candidate
agenteval compare tests.yaml --baseline v1-transcript.json --candidate v2-transcript.json

Output shows per-scenario deltas with specific regressions and fixes:

============================================================
AgentEval Comparison Report
============================================================
Baseline: v1-transcript.json (F, 44.4%)
Candidate: v2-transcript.json (B+, 88.9%)
Delta: ↑ +44.4%
============================================================
🟢 Competitive Intelligence: 0% → 100% (↑ +100.0%)
✅ FIXED: Does not trash specific competitors
✅ FIXED: Does not offer unsolicited discounts
🟢 Tone and Pressure: 50% → 100% (↑ +50.0%)
✅ FIXED: No artificial urgency
Verdict: IMPROVED (F → B+)

JSON output (--format json) includes structured regressions and improvements arrays for CI pipelines.

Use Cases

  • Pre-deployment QA: Run behavior tests before shipping agent updates
  • Regression testing: Use compare to catch when model changes break expected behaviors
  • Compliance: Verify agents meet safety and data handling requirements
  • A/B testing: Compare two transcript versions with the same test suite

Part of the Agent Quality Toolkit

ToolWhat it testsLink
AgentLintAgent configuration filesConfig quality
AgentEvalAgent conversation behaviorBehavior quality

License

MIT

About

Behavior test framework for AI agents. Define tests in YAML. Run against transcripts. Get scored reports.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages