Skip to content

Repository files navigation

Agent behavior

Behavior specs are an open standard for defining and evaluating how an AI agent should behave across a whole trajectory.

Documentation · Specification · Quickstart · Examples

Why behavior specs

A modern agent can work for hours and make hundreds of decisions in a single trajectory, and you cannot reduce that behavior to one outcome metric. To build long-horizon agents you can trust, you have to supervise the process, not only the result.

A behavior spec records the recurring conduct you expect from an agent: how it gathers context, makes decisions, acts, and recovers when it does not know enough. It defines the standard before any reviewer, rubric, scorer, or eval measures it, so you can review traces, design evals, and align prompts against a single written source of truth.

What a behavior spec looks like

A behavior spec is a BEHAVIOR.md file with YAML frontmatter and a free-form Markdown body. Specs live under .agents/behaviors/, next to the agent they describe:

.agents/behaviors/
└── validate-rendered-deck/
├── BEHAVIOR.md # Required: YAML frontmatter and behavior description
└── references/ # Optional: rationale, examples, background docs
---name: validate-rendered-deckdescription: Render the current PowerPoint before delivery, inspect it for visual issues, and revalidate after fixes.---# Validate the rendered deck before returning it**Intent:** Each time the agent submits or resubmits a created or edited slide
deck, ensure the user receives a visually usable deck, not merely valid
PowerPoint code.
**Evidence:** The current saved deck and slide images rendered from that
version. A render made before the latest edit is not evidence about the deck
now being returned.
**Decision:** Determine whether the current rendered deck has formatting,
layout, readability, or other visible issues that make it unsuitable to
return.
**Execution:** Render the current deck to images and inspect those images
before returning it.
**Recovery:** If inspection finds a fixable issue, fix the deck, render the
updated version, and inspect the new render. If the agent cannot render or
inspect the deck, it should not claim the deck was visually validated.
**Failure modes:** Returning an unrendered deck; relying on a stale render;
fixing an issue without re-rendering; missing a visual problem that code-level
checks cannot reveal.

These six labels are optional. They are one way to organize a behavior; plain Markdown works too. For a domain-specific example, see primary-source-tax-research.

Get started

  1. Write a spec at .agents/behaviors/<name>/BEHAVIOR.md. The quickstart walks through your first one, and the portable writing-agent-behavior skill helps agents author and calibrate specs.

  2. Validate structure with the CLI in packages/agentbehavior. From a clone of this repository:

    pnpm install
    pnpm build
    pnpm exec agentbehavior validate .
  3. Evaluate traces against your specs. The runnable examples show a true/false/na judging convention over recorded trajectories.

Documentation

Full documentation is published at agentbehavior.dev:

Contributing

Agent behavior began as a collaboration between Basis and Braintrust. Contributions are welcome. See CONTRIBUTING.md for development setup and useful commands.

License

Licensed under the Apache License 2.0. See LICENSE.

About

Standards for defining and evaluating agent behavior

Topics

Resources

Contributing

Stars

323 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - braintrustdata/agentbehavior: Standards for defining and evaluating agent behavior · GitHub
Skip to content

Repository files navigation

Agent behavior

Behavior specs are an open standard for defining and evaluating how an AI agent should behave across a whole trajectory.

Documentation · Specification · Quickstart · Examples

Why behavior specs

A modern agent can work for hours and make hundreds of decisions in a single trajectory, and you cannot reduce that behavior to one outcome metric. To build long-horizon agents you can trust, you have to supervise the process, not only the result.

A behavior spec records the recurring conduct you expect from an agent: how it gathers context, makes decisions, acts, and recovers when it does not know enough. It defines the standard before any reviewer, rubric, scorer, or eval measures it, so you can review traces, design evals, and align prompts against a single written source of truth.

What a behavior spec looks like

A behavior spec is a BEHAVIOR.md file with YAML frontmatter and a free-form Markdown body. Specs live under .agents/behaviors/, next to the agent they describe:

.agents/behaviors/
└── validate-rendered-deck/
├── BEHAVIOR.md # Required: YAML frontmatter and behavior description
└── references/ # Optional: rationale, examples, background docs
---name: validate-rendered-deckdescription: Render the current PowerPoint before delivery, inspect it for visual issues, and revalidate after fixes.---# Validate the rendered deck before returning it**Intent:** Each time the agent submits or resubmits a created or edited slide
deck, ensure the user receives a visually usable deck, not merely valid
PowerPoint code.
**Evidence:** The current saved deck and slide images rendered from that
version. A render made before the latest edit is not evidence about the deck
now being returned.
**Decision:** Determine whether the current rendered deck has formatting,
layout, readability, or other visible issues that make it unsuitable to
return.
**Execution:** Render the current deck to images and inspect those images
before returning it.
**Recovery:** If inspection finds a fixable issue, fix the deck, render the
updated version, and inspect the new render. If the agent cannot render or
inspect the deck, it should not claim the deck was visually validated.
**Failure modes:** Returning an unrendered deck; relying on a stale render;
fixing an issue without re-rendering; missing a visual problem that code-level
checks cannot reveal.

These six labels are optional. They are one way to organize a behavior; plain Markdown works too. For a domain-specific example, see primary-source-tax-research.

Get started

  1. Write a spec at .agents/behaviors/<name>/BEHAVIOR.md. The quickstart walks through your first one, and the portable writing-agent-behavior skill helps agents author and calibrate specs.

  2. Validate structure with the CLI in packages/agentbehavior. From a clone of this repository:

    pnpm install
    pnpm build
    pnpm exec agentbehavior validate .
  3. Evaluate traces against your specs. The runnable examples show a true/false/na judging convention over recorded trajectories.

Documentation

Full documentation is published at agentbehavior.dev:

Contributing

Agent behavior began as a collaboration between Basis and Braintrust. Contributions are welcome. See CONTRIBUTING.md for development setup and useful commands.

License

Licensed under the Apache License 2.0. See LICENSE.

About

Standards for defining and evaluating agent behavior

Topics

Resources

Contributing

Stars

323 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - braintrustdata/agentbehavior: Standards for defining and evaluating agent behavior · GitHub
Skip to content

Repository files navigation

Agent behavior

Behavior specs are an open standard for defining and evaluating how an AI agent should behave across a whole trajectory.

Documentation · Specification · Quickstart · Examples

Why behavior specs

A modern agent can work for hours and make hundreds of decisions in a single trajectory, and you cannot reduce that behavior to one outcome metric. To build long-horizon agents you can trust, you have to supervise the process, not only the result.

A behavior spec records the recurring conduct you expect from an agent: how it gathers context, makes decisions, acts, and recovers when it does not know enough. It defines the standard before any reviewer, rubric, scorer, or eval measures it, so you can review traces, design evals, and align prompts against a single written source of truth.

What a behavior spec looks like

A behavior spec is a BEHAVIOR.md file with YAML frontmatter and a free-form Markdown body. Specs live under .agents/behaviors/, next to the agent they describe:

.agents/behaviors/
└── validate-rendered-deck/
├── BEHAVIOR.md # Required: YAML frontmatter and behavior description
└── references/ # Optional: rationale, examples, background docs
---name: validate-rendered-deckdescription: Render the current PowerPoint before delivery, inspect it for visual issues, and revalidate after fixes.---# Validate the rendered deck before returning it**Intent:** Each time the agent submits or resubmits a created or edited slide
deck, ensure the user receives a visually usable deck, not merely valid
PowerPoint code.
**Evidence:** The current saved deck and slide images rendered from that
version. A render made before the latest edit is not evidence about the deck
now being returned.
**Decision:** Determine whether the current rendered deck has formatting,
layout, readability, or other visible issues that make it unsuitable to
return.
**Execution:** Render the current deck to images and inspect those images
before returning it.
**Recovery:** If inspection finds a fixable issue, fix the deck, render the
updated version, and inspect the new render. If the agent cannot render or
inspect the deck, it should not claim the deck was visually validated.
**Failure modes:** Returning an unrendered deck; relying on a stale render;
fixing an issue without re-rendering; missing a visual problem that code-level
checks cannot reveal.

These six labels are optional. They are one way to organize a behavior; plain Markdown works too. For a domain-specific example, see primary-source-tax-research.

Get started

  1. Write a spec at .agents/behaviors/<name>/BEHAVIOR.md. The quickstart walks through your first one, and the portable writing-agent-behavior skill helps agents author and calibrate specs.

  2. Validate structure with the CLI in packages/agentbehavior. From a clone of this repository:

    pnpm install
    pnpm build
    pnpm exec agentbehavior validate .
  3. Evaluate traces against your specs. The runnable examples show a true/false/na judging convention over recorded trajectories.

Documentation

Full documentation is published at agentbehavior.dev:

Contributing

Agent behavior began as a collaboration between Basis and Braintrust. Contributions are welcome. See CONTRIBUTING.md for development setup and useful commands.

License

Licensed under the Apache License 2.0. See LICENSE.

About

Standards for defining and evaluating agent behavior

Topics

Resources

Contributing

Stars

323 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - braintrustdata/agentbehavior: Standards for defining and evaluating agent behavior · GitHub
Skip to content

Repository files navigation

Agent behavior

Behavior specs are an open standard for defining and evaluating how an AI agent should behave across a whole trajectory.

Documentation · Specification · Quickstart · Examples

Why behavior specs

A modern agent can work for hours and make hundreds of decisions in a single trajectory, and you cannot reduce that behavior to one outcome metric. To build long-horizon agents you can trust, you have to supervise the process, not only the result.

A behavior spec records the recurring conduct you expect from an agent: how it gathers context, makes decisions, acts, and recovers when it does not know enough. It defines the standard before any reviewer, rubric, scorer, or eval measures it, so you can review traces, design evals, and align prompts against a single written source of truth.

What a behavior spec looks like

A behavior spec is a BEHAVIOR.md file with YAML frontmatter and a free-form Markdown body. Specs live under .agents/behaviors/, next to the agent they describe:

.agents/behaviors/
└── validate-rendered-deck/
├── BEHAVIOR.md # Required: YAML frontmatter and behavior description
└── references/ # Optional: rationale, examples, background docs
---name: validate-rendered-deckdescription: Render the current PowerPoint before delivery, inspect it for visual issues, and revalidate after fixes.---# Validate the rendered deck before returning it**Intent:** Each time the agent submits or resubmits a created or edited slide
deck, ensure the user receives a visually usable deck, not merely valid
PowerPoint code.
**Evidence:** The current saved deck and slide images rendered from that
version. A render made before the latest edit is not evidence about the deck
now being returned.
**Decision:** Determine whether the current rendered deck has formatting,
layout, readability, or other visible issues that make it unsuitable to
return.
**Execution:** Render the current deck to images and inspect those images
before returning it.
**Recovery:** If inspection finds a fixable issue, fix the deck, render the
updated version, and inspect the new render. If the agent cannot render or
inspect the deck, it should not claim the deck was visually validated.
**Failure modes:** Returning an unrendered deck; relying on a stale render;
fixing an issue without re-rendering; missing a visual problem that code-level
checks cannot reveal.

These six labels are optional. They are one way to organize a behavior; plain Markdown works too. For a domain-specific example, see primary-source-tax-research.

Get started

  1. Write a spec at .agents/behaviors/<name>/BEHAVIOR.md. The quickstart walks through your first one, and the portable writing-agent-behavior skill helps agents author and calibrate specs.

  2. Validate structure with the CLI in packages/agentbehavior. From a clone of this repository:

    pnpm install
    pnpm build
    pnpm exec agentbehavior validate .
  3. Evaluate traces against your specs. The runnable examples show a true/false/na judging convention over recorded trajectories.

Documentation

Full documentation is published at agentbehavior.dev:

Contributing

Agent behavior began as a collaboration between Basis and Braintrust. Contributions are welcome. See CONTRIBUTING.md for development setup and useful commands.

License

Licensed under the Apache License 2.0. See LICENSE.

About

Standards for defining and evaluating agent behavior

Topics

Resources

Contributing

Stars

323 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - braintrustdata/agentbehavior: Standards for defining and evaluating agent behavior · GitHub
Skip to content

Repository files navigation

Agent behavior

Behavior specs are an open standard for defining and evaluating how an AI agent should behave across a whole trajectory.

Documentation · Specification · Quickstart · Examples

Why behavior specs

A modern agent can work for hours and make hundreds of decisions in a single trajectory, and you cannot reduce that behavior to one outcome metric. To build long-horizon agents you can trust, you have to supervise the process, not only the result.

A behavior spec records the recurring conduct you expect from an agent: how it gathers context, makes decisions, acts, and recovers when it does not know enough. It defines the standard before any reviewer, rubric, scorer, or eval measures it, so you can review traces, design evals, and align prompts against a single written source of truth.

What a behavior spec looks like

A behavior spec is a BEHAVIOR.md file with YAML frontmatter and a free-form Markdown body. Specs live under .agents/behaviors/, next to the agent they describe:

.agents/behaviors/
└── validate-rendered-deck/
├── BEHAVIOR.md # Required: YAML frontmatter and behavior description
└── references/ # Optional: rationale, examples, background docs
---name: validate-rendered-deckdescription: Render the current PowerPoint before delivery, inspect it for visual issues, and revalidate after fixes.---# Validate the rendered deck before returning it**Intent:** Each time the agent submits or resubmits a created or edited slide
deck, ensure the user receives a visually usable deck, not merely valid
PowerPoint code.
**Evidence:** The current saved deck and slide images rendered from that
version. A render made before the latest edit is not evidence about the deck
now being returned.
**Decision:** Determine whether the current rendered deck has formatting,
layout, readability, or other visible issues that make it unsuitable to
return.
**Execution:** Render the current deck to images and inspect those images
before returning it.
**Recovery:** If inspection finds a fixable issue, fix the deck, render the
updated version, and inspect the new render. If the agent cannot render or
inspect the deck, it should not claim the deck was visually validated.
**Failure modes:** Returning an unrendered deck; relying on a stale render;
fixing an issue without re-rendering; missing a visual problem that code-level
checks cannot reveal.

These six labels are optional. They are one way to organize a behavior; plain Markdown works too. For a domain-specific example, see primary-source-tax-research.

Get started

  1. Write a spec at .agents/behaviors/<name>/BEHAVIOR.md. The quickstart walks through your first one, and the portable writing-agent-behavior skill helps agents author and calibrate specs.

  2. Validate structure with the CLI in packages/agentbehavior. From a clone of this repository:

    pnpm install
    pnpm build
    pnpm exec agentbehavior validate .
  3. Evaluate traces against your specs. The runnable examples show a true/false/na judging convention over recorded trajectories.

Documentation

Full documentation is published at agentbehavior.dev:

Contributing

Agent behavior began as a collaboration between Basis and Braintrust. Contributions are welcome. See CONTRIBUTING.md for development setup and useful commands.

License

Licensed under the Apache License 2.0. See LICENSE.

About

Standards for defining and evaluating agent behavior

Topics

Resources

Contributing

Stars

323 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - braintrustdata/agentbehavior: Standards for defining and evaluating agent behavior · GitHub
Skip to content

Repository files navigation

Agent behavior

Behavior specs are an open standard for defining and evaluating how an AI agent should behave across a whole trajectory.

Documentation · Specification · Quickstart · Examples

Why behavior specs

A modern agent can work for hours and make hundreds of decisions in a single trajectory, and you cannot reduce that behavior to one outcome metric. To build long-horizon agents you can trust, you have to supervise the process, not only the result.

A behavior spec records the recurring conduct you expect from an agent: how it gathers context, makes decisions, acts, and recovers when it does not know enough. It defines the standard before any reviewer, rubric, scorer, or eval measures it, so you can review traces, design evals, and align prompts against a single written source of truth.

What a behavior spec looks like

A behavior spec is a BEHAVIOR.md file with YAML frontmatter and a free-form Markdown body. Specs live under .agents/behaviors/, next to the agent they describe:

.agents/behaviors/
└── validate-rendered-deck/
├── BEHAVIOR.md # Required: YAML frontmatter and behavior description
└── references/ # Optional: rationale, examples, background docs
---name: validate-rendered-deckdescription: Render the current PowerPoint before delivery, inspect it for visual issues, and revalidate after fixes.---# Validate the rendered deck before returning it**Intent:** Each time the agent submits or resubmits a created or edited slide
deck, ensure the user receives a visually usable deck, not merely valid
PowerPoint code.
**Evidence:** The current saved deck and slide images rendered from that
version. A render made before the latest edit is not evidence about the deck
now being returned.
**Decision:** Determine whether the current rendered deck has formatting,
layout, readability, or other visible issues that make it unsuitable to
return.
**Execution:** Render the current deck to images and inspect those images
before returning it.
**Recovery:** If inspection finds a fixable issue, fix the deck, render the
updated version, and inspect the new render. If the agent cannot render or
inspect the deck, it should not claim the deck was visually validated.
**Failure modes:** Returning an unrendered deck; relying on a stale render;
fixing an issue without re-rendering; missing a visual problem that code-level
checks cannot reveal.

These six labels are optional. They are one way to organize a behavior; plain Markdown works too. For a domain-specific example, see primary-source-tax-research.

Get started

  1. Write a spec at .agents/behaviors/<name>/BEHAVIOR.md. The quickstart walks through your first one, and the portable writing-agent-behavior skill helps agents author and calibrate specs.

  2. Validate structure with the CLI in packages/agentbehavior. From a clone of this repository:

    pnpm install
    pnpm build
    pnpm exec agentbehavior validate .
  3. Evaluate traces against your specs. The runnable examples show a true/false/na judging convention over recorded trajectories.

Documentation

Full documentation is published at agentbehavior.dev:

Contributing

Agent behavior began as a collaboration between Basis and Braintrust. Contributions are welcome. See CONTRIBUTING.md for development setup and useful commands.

License

Licensed under the Apache License 2.0. See LICENSE.

About

Standards for defining and evaluating agent behavior

Topics

Resources

Contributing

Stars

323 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); })(); GitHub - braintrustdata/agentbehavior: Standards for defining and evaluating agent behavior · GitHub
Skip to content

Repository files navigation

Agent behavior

Behavior specs are an open standard for defining and evaluating how an AI agent should behave across a whole trajectory.

Documentation · Specification · Quickstart · Examples

Why behavior specs

A modern agent can work for hours and make hundreds of decisions in a single trajectory, and you cannot reduce that behavior to one outcome metric. To build long-horizon agents you can trust, you have to supervise the process, not only the result.

A behavior spec records the recurring conduct you expect from an agent: how it gathers context, makes decisions, acts, and recovers when it does not know enough. It defines the standard before any reviewer, rubric, scorer, or eval measures it, so you can review traces, design evals, and align prompts against a single written source of truth.

What a behavior spec looks like

A behavior spec is a BEHAVIOR.md file with YAML frontmatter and a free-form Markdown body. Specs live under .agents/behaviors/, next to the agent they describe:

.agents/behaviors/
└── validate-rendered-deck/
├── BEHAVIOR.md # Required: YAML frontmatter and behavior description
└── references/ # Optional: rationale, examples, background docs
---name: validate-rendered-deckdescription: Render the current PowerPoint before delivery, inspect it for visual issues, and revalidate after fixes.---# Validate the rendered deck before returning it**Intent:** Each time the agent submits or resubmits a created or edited slide
deck, ensure the user receives a visually usable deck, not merely valid
PowerPoint code.
**Evidence:** The current saved deck and slide images rendered from that
version. A render made before the latest edit is not evidence about the deck
now being returned.
**Decision:** Determine whether the current rendered deck has formatting,
layout, readability, or other visible issues that make it unsuitable to
return.
**Execution:** Render the current deck to images and inspect those images
before returning it.
**Recovery:** If inspection finds a fixable issue, fix the deck, render the
updated version, and inspect the new render. If the agent cannot render or
inspect the deck, it should not claim the deck was visually validated.
**Failure modes:** Returning an unrendered deck; relying on a stale render;
fixing an issue without re-rendering; missing a visual problem that code-level
checks cannot reveal.

These six labels are optional. They are one way to organize a behavior; plain Markdown works too. For a domain-specific example, see primary-source-tax-research.

Get started

  1. Write a spec at .agents/behaviors/<name>/BEHAVIOR.md. The quickstart walks through your first one, and the portable writing-agent-behavior skill helps agents author and calibrate specs.

  2. Validate structure with the CLI in packages/agentbehavior. From a clone of this repository:

    pnpm install
    pnpm build
    pnpm exec agentbehavior validate .
  3. Evaluate traces against your specs. The runnable examples show a true/false/na judging convention over recorded trajectories.

Documentation

Full documentation is published at agentbehavior.dev:

Contributing

Agent behavior began as a collaboration between Basis and Braintrust. Contributions are welcome. See CONTRIBUTING.md for development setup and useful commands.

License

Licensed under the Apache License 2.0. See LICENSE.

About

Standards for defining and evaluating agent behavior

Topics

Resources

Contributing

Stars

323 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages