Repository files navigation

AI Release Governance Framework

NIST AI RMFLicense: MIT

A practitioner framework for turning evaluation, control, operational, and risk evidence into an accountable AI release decision.

The repository is deliberately broader than a checklist and narrower than an enterprise governance operating model. It focuses on the release boundary: what proposition is being approved, which evidence supports it, what remains uncertain, and who owns the decision, conditions, exceptions, residual risks, and rollback.

Start here

ArtifactUse it for
docs/release-decision-record.mddecision outcomes, evidence fields, finding dispositions, conditions, exceptions, and residual-risk acceptance
docs/nist-rmf-mapping.mdcautious practitioner cross-reference to NIST AI RMF functions
release-checklistexecutable YAML validation for a narrower set of release-readiness fields
governance-playbookbroader organizational operating model and recurring governance forums

A release decision is scoped

“Release” should always identify:

  • system and version;
  • model, prompt, retrieval, data, tools, and permissions in scope;
  • environment and infrastructure configuration;
  • user population and geography;
  • enabled actions and external effects;
  • rollout stage and exposure limit;
  • evidence cutoff date;
  • decision owner and authority;
  • conditions, exceptions, and expiry;
  • changes that invalidate the decision.

Approval of one configuration should not be interpreted as approval of future model, prompt, data, tool, permission, or population changes.

Lifecycle

purpose and risk context
↓
release proposition and scope
↓
evidence plan and hard gates
↓
evaluation and control verification
↓
operational / containment readiness
↓
release decision and dispositions
↓
staged rollout and verification
↓
monitoring, incident, change, and retirement

Release governance begins before the final review. The evidence plan, owners, hard gates, and invalidation triggers should be agreed before teams optimize against a convenient metric set.

Gate design

A gate should ask a decision-relevant question, identify the evidence needed, and define the possible dispositions.

1. Purpose, scope, and risk context

Evidence should establish:

  • intended use, users, and prohibited uses;
  • affected people and foreseeable misuse;
  • authority, data, and tool boundaries;
  • reversibility and worst credible consequences;
  • applicable internal and external obligations;
  • owner and release authority.

2. Evaluation validity

Evidence should establish:

  • task and risk coverage;
  • operating-population assumptions and exclusions;
  • hard-gate and quality-measure definitions;
  • evaluator validity and adjudication;
  • repeated-run policy and uncertainty;
  • slice results and material failures;
  • regression coverage.

Do not rely on an aggregate score without reviewing critical failures and non-compensable conditions.

3. Security, privacy, safety, and policy controls

Evidence should establish that relevant controls are:

  • defined against the actual system scope;
  • enforced in the deployment path;
  • tested under normal, adversarial, and failure conditions;
  • attributable to a current version;
  • supported by exception and incident processes;
  • subject to named ownership and change control.

Avoid prescribing a particular tool or explanation technique as universally required. The evidence method should fit the system, decision, and risk.

4. Operational and containment readiness

Evidence should cover:

  • service objectives and capacity assumptions;
  • observability at the agent, tool, workflow, and outcome levels;
  • rollback, disable, and credential-revocation paths;
  • partial and uncertain external actions;
  • data and state recovery;
  • on-call and incident decision rights;
  • user support and redress where applicable;
  • evidence retention with privacy controls.

A rollback test should verify the resulting state, not only demonstrate that an earlier software version can be redeployed.

5. Decision and follow-through

The review should distinguish:

TermMeaning
Hard gatenon-compensable requirement
Blockerunresolved condition preventing the current decision
Evidence gapmissing or unreliable support for a proposition
Required actionfollow-up work accepted under a bounded conditional release
Conditionenforceable limit on scope or operation
Exceptionauthorized deviation with compensating controls and expiry
Residual riskremaining risk accepted by an authorized owner
Observationimprovement that does not currently change the decision

See docs/release-decision-record.md for a template and review questions.

Decision outcomes

OutcomeUse when
Releaserequired hard gates pass and no conditions remain
Release with conditionsno blocker remains, but enforceable constraints or required actions limit the release
Holdevidence, remediation, or control readiness is insufficient
Do not releasecritical failure, prohibited condition, or unacceptable residual risk remains
Defer decisionthe owner postpones judgment until specified evidence or dependency is available

A conditional release must not relabel an unresolved blocker as a future action. Conditions should state scope, owner, measurement, stop trigger, and expiry.

Evidence quality

For each material conclusion, record:

  • proposition or control being assessed;
  • system scope and version;
  • evidence source, author, date, and location;
  • method and reviewer;
  • result and uncertainty;
  • coverage and exclusions;
  • freshness and invalidation trigger;
  • owner and disposition.

A current document can contain stale evidence. Freshness depends on whether the reviewed system and operating conditions have changed, not only the file date.

Staged release

A staged rollout is a control only when it has:

  • a defined population and exposure limit;
  • monitoring linked to plausible failures;
  • named go/no-go decision points;
  • stop and rollback authority;
  • minimum evidence for expansion;
  • user support and incident response;
  • conditions that prevent silent scope growth.

“Pilot” should not become an indefinite production state without renewed evidence and an accountable decision.

Change control

Re-evaluate when material changes occur to:

  • model or provider;
  • prompt, policy, routing, or orchestration;
  • retrieval corpus or data distribution;
  • tools, permissions, identities, or action authority;
  • infrastructure or deployment region;
  • user population or use case;
  • evaluator, rubric, threshold, or test set;
  • relevant law, policy, or internal control;
  • known failure or incident evidence.

The prior decision record should state which changes invalidate it.

Maturity and scope

This is a practitioner framework for planning and reviewing AI releases. It is not a certified release process, safety case, regulatory approval, legal determination, or substitute for qualified security, privacy, safety, compliance, operational, and domain review.

References to NIST AI RMF and other governance concepts are practitioner mappings. Verify official sources and adapt the framework to the actual authority, population, risk, and jurisdiction.

Related repositories

RepositoryDistinct role
release-checklistworking config validator
agent-evalevaluation validity and decision semantics
accountability-patternsdecision rights, human review, provenance, and redress
regulated-aistarter repository structure and templates

Maintained by Sima Bagheri.

About

A practical framework for AI release readiness, Lifecycle gates, decision rights, and reusable artifacts for release-stage governance.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

AI Release Governance Framework

NIST AI RMFLicense: MIT

A practitioner framework for turning evaluation, control, operational, and risk evidence into an accountable AI release decision.

The repository is deliberately broader than a checklist and narrower than an enterprise governance operating model. It focuses on the release boundary: what proposition is being approved, which evidence supports it, what remains uncertain, and who owns the decision, conditions, exceptions, residual risks, and rollback.

Start here

ArtifactUse it for
docs/release-decision-record.mddecision outcomes, evidence fields, finding dispositions, conditions, exceptions, and residual-risk acceptance
docs/nist-rmf-mapping.mdcautious practitioner cross-reference to NIST AI RMF functions
release-checklistexecutable YAML validation for a narrower set of release-readiness fields
governance-playbookbroader organizational operating model and recurring governance forums

A release decision is scoped

“Release” should always identify:

  • system and version;
  • model, prompt, retrieval, data, tools, and permissions in scope;
  • environment and infrastructure configuration;
  • user population and geography;
  • enabled actions and external effects;
  • rollout stage and exposure limit;
  • evidence cutoff date;
  • decision owner and authority;
  • conditions, exceptions, and expiry;
  • changes that invalidate the decision.

Approval of one configuration should not be interpreted as approval of future model, prompt, data, tool, permission, or population changes.

Lifecycle

purpose and risk context
↓
release proposition and scope
↓
evidence plan and hard gates
↓
evaluation and control verification
↓
operational / containment readiness
↓
release decision and dispositions
↓
staged rollout and verification
↓
monitoring, incident, change, and retirement

Release governance begins before the final review. The evidence plan, owners, hard gates, and invalidation triggers should be agreed before teams optimize against a convenient metric set.

Gate design

A gate should ask a decision-relevant question, identify the evidence needed, and define the possible dispositions.

1. Purpose, scope, and risk context

Evidence should establish:

  • intended use, users, and prohibited uses;
  • affected people and foreseeable misuse;
  • authority, data, and tool boundaries;
  • reversibility and worst credible consequences;
  • applicable internal and external obligations;
  • owner and release authority.

2. Evaluation validity

Evidence should establish:

  • task and risk coverage;
  • operating-population assumptions and exclusions;
  • hard-gate and quality-measure definitions;
  • evaluator validity and adjudication;
  • repeated-run policy and uncertainty;
  • slice results and material failures;
  • regression coverage.

Do not rely on an aggregate score without reviewing critical failures and non-compensable conditions.

3. Security, privacy, safety, and policy controls

Evidence should establish that relevant controls are:

  • defined against the actual system scope;
  • enforced in the deployment path;
  • tested under normal, adversarial, and failure conditions;
  • attributable to a current version;
  • supported by exception and incident processes;
  • subject to named ownership and change control.

Avoid prescribing a particular tool or explanation technique as universally required. The evidence method should fit the system, decision, and risk.

4. Operational and containment readiness

Evidence should cover:

  • service objectives and capacity assumptions;
  • observability at the agent, tool, workflow, and outcome levels;
  • rollback, disable, and credential-revocation paths;
  • partial and uncertain external actions;
  • data and state recovery;
  • on-call and incident decision rights;
  • user support and redress where applicable;
  • evidence retention with privacy controls.

A rollback test should verify the resulting state, not only demonstrate that an earlier software version can be redeployed.

5. Decision and follow-through

The review should distinguish:

TermMeaning
Hard gatenon-compensable requirement
Blockerunresolved condition preventing the current decision
Evidence gapmissing or unreliable support for a proposition
Required actionfollow-up work accepted under a bounded conditional release
Conditionenforceable limit on scope or operation
Exceptionauthorized deviation with compensating controls and expiry
Residual riskremaining risk accepted by an authorized owner
Observationimprovement that does not currently change the decision

See docs/release-decision-record.md for a template and review questions.

Decision outcomes

OutcomeUse when
Releaserequired hard gates pass and no conditions remain
Release with conditionsno blocker remains, but enforceable constraints or required actions limit the release
Holdevidence, remediation, or control readiness is insufficient
Do not releasecritical failure, prohibited condition, or unacceptable residual risk remains
Defer decisionthe owner postpones judgment until specified evidence or dependency is available

A conditional release must not relabel an unresolved blocker as a future action. Conditions should state scope, owner, measurement, stop trigger, and expiry.

Evidence quality

For each material conclusion, record:

  • proposition or control being assessed;
  • system scope and version;
  • evidence source, author, date, and location;
  • method and reviewer;
  • result and uncertainty;
  • coverage and exclusions;
  • freshness and invalidation trigger;
  • owner and disposition.

A current document can contain stale evidence. Freshness depends on whether the reviewed system and operating conditions have changed, not only the file date.

Staged release

A staged rollout is a control only when it has:

  • a defined population and exposure limit;
  • monitoring linked to plausible failures;
  • named go/no-go decision points;
  • stop and rollback authority;
  • minimum evidence for expansion;
  • user support and incident response;
  • conditions that prevent silent scope growth.

“Pilot” should not become an indefinite production state without renewed evidence and an accountable decision.

Change control

Re-evaluate when material changes occur to:

  • model or provider;
  • prompt, policy, routing, or orchestration;
  • retrieval corpus or data distribution;
  • tools, permissions, identities, or action authority;
  • infrastructure or deployment region;
  • user population or use case;
  • evaluator, rubric, threshold, or test set;
  • relevant law, policy, or internal control;
  • known failure or incident evidence.

The prior decision record should state which changes invalidate it.

Maturity and scope

This is a practitioner framework for planning and reviewing AI releases. It is not a certified release process, safety case, regulatory approval, legal determination, or substitute for qualified security, privacy, safety, compliance, operational, and domain review.

References to NIST AI RMF and other governance concepts are practitioner mappings. Verify official sources and adapt the framework to the actual authority, population, risk, and jurisdiction.

Related repositories

RepositoryDistinct role
release-checklistworking config validator
agent-evalevaluation validity and decision semantics
accountability-patternsdecision rights, human review, provenance, and redress
regulated-aistarter repository structure and templates

Maintained by Sima Bagheri.

About

A practical framework for AI release readiness, Lifecycle gates, decision rights, and reusable artifacts for release-stage governance.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

AI Release Governance Framework

NIST AI RMFLicense: MIT

A practitioner framework for turning evaluation, control, operational, and risk evidence into an accountable AI release decision.

The repository is deliberately broader than a checklist and narrower than an enterprise governance operating model. It focuses on the release boundary: what proposition is being approved, which evidence supports it, what remains uncertain, and who owns the decision, conditions, exceptions, residual risks, and rollback.

Start here

ArtifactUse it for
docs/release-decision-record.mddecision outcomes, evidence fields, finding dispositions, conditions, exceptions, and residual-risk acceptance
docs/nist-rmf-mapping.mdcautious practitioner cross-reference to NIST AI RMF functions
release-checklistexecutable YAML validation for a narrower set of release-readiness fields
governance-playbookbroader organizational operating model and recurring governance forums

A release decision is scoped

“Release” should always identify:

  • system and version;
  • model, prompt, retrieval, data, tools, and permissions in scope;
  • environment and infrastructure configuration;
  • user population and geography;
  • enabled actions and external effects;
  • rollout stage and exposure limit;
  • evidence cutoff date;
  • decision owner and authority;
  • conditions, exceptions, and expiry;
  • changes that invalidate the decision.

Approval of one configuration should not be interpreted as approval of future model, prompt, data, tool, permission, or population changes.

Lifecycle

purpose and risk context
↓
release proposition and scope
↓
evidence plan and hard gates
↓
evaluation and control verification
↓
operational / containment readiness
↓
release decision and dispositions
↓
staged rollout and verification
↓
monitoring, incident, change, and retirement

Release governance begins before the final review. The evidence plan, owners, hard gates, and invalidation triggers should be agreed before teams optimize against a convenient metric set.

Gate design

A gate should ask a decision-relevant question, identify the evidence needed, and define the possible dispositions.

1. Purpose, scope, and risk context

Evidence should establish:

  • intended use, users, and prohibited uses;
  • affected people and foreseeable misuse;
  • authority, data, and tool boundaries;
  • reversibility and worst credible consequences;
  • applicable internal and external obligations;
  • owner and release authority.

2. Evaluation validity

Evidence should establish:

  • task and risk coverage;
  • operating-population assumptions and exclusions;
  • hard-gate and quality-measure definitions;
  • evaluator validity and adjudication;
  • repeated-run policy and uncertainty;
  • slice results and material failures;
  • regression coverage.

Do not rely on an aggregate score without reviewing critical failures and non-compensable conditions.

3. Security, privacy, safety, and policy controls

Evidence should establish that relevant controls are:

  • defined against the actual system scope;
  • enforced in the deployment path;
  • tested under normal, adversarial, and failure conditions;
  • attributable to a current version;
  • supported by exception and incident processes;
  • subject to named ownership and change control.

Avoid prescribing a particular tool or explanation technique as universally required. The evidence method should fit the system, decision, and risk.

4. Operational and containment readiness

Evidence should cover:

  • service objectives and capacity assumptions;
  • observability at the agent, tool, workflow, and outcome levels;
  • rollback, disable, and credential-revocation paths;
  • partial and uncertain external actions;
  • data and state recovery;
  • on-call and incident decision rights;
  • user support and redress where applicable;
  • evidence retention with privacy controls.

A rollback test should verify the resulting state, not only demonstrate that an earlier software version can be redeployed.

5. Decision and follow-through

The review should distinguish:

TermMeaning
Hard gatenon-compensable requirement
Blockerunresolved condition preventing the current decision
Evidence gapmissing or unreliable support for a proposition
Required actionfollow-up work accepted under a bounded conditional release
Conditionenforceable limit on scope or operation
Exceptionauthorized deviation with compensating controls and expiry
Residual riskremaining risk accepted by an authorized owner
Observationimprovement that does not currently change the decision

See docs/release-decision-record.md for a template and review questions.

Decision outcomes

OutcomeUse when
Releaserequired hard gates pass and no conditions remain
Release with conditionsno blocker remains, but enforceable constraints or required actions limit the release
Holdevidence, remediation, or control readiness is insufficient
Do not releasecritical failure, prohibited condition, or unacceptable residual risk remains
Defer decisionthe owner postpones judgment until specified evidence or dependency is available

A conditional release must not relabel an unresolved blocker as a future action. Conditions should state scope, owner, measurement, stop trigger, and expiry.

Evidence quality

For each material conclusion, record:

  • proposition or control being assessed;
  • system scope and version;
  • evidence source, author, date, and location;
  • method and reviewer;
  • result and uncertainty;
  • coverage and exclusions;
  • freshness and invalidation trigger;
  • owner and disposition.

A current document can contain stale evidence. Freshness depends on whether the reviewed system and operating conditions have changed, not only the file date.

Staged release

A staged rollout is a control only when it has:

  • a defined population and exposure limit;
  • monitoring linked to plausible failures;
  • named go/no-go decision points;
  • stop and rollback authority;
  • minimum evidence for expansion;
  • user support and incident response;
  • conditions that prevent silent scope growth.

“Pilot” should not become an indefinite production state without renewed evidence and an accountable decision.

Change control

Re-evaluate when material changes occur to:

  • model or provider;
  • prompt, policy, routing, or orchestration;
  • retrieval corpus or data distribution;
  • tools, permissions, identities, or action authority;
  • infrastructure or deployment region;
  • user population or use case;
  • evaluator, rubric, threshold, or test set;
  • relevant law, policy, or internal control;
  • known failure or incident evidence.

The prior decision record should state which changes invalidate it.

Maturity and scope

This is a practitioner framework for planning and reviewing AI releases. It is not a certified release process, safety case, regulatory approval, legal determination, or substitute for qualified security, privacy, safety, compliance, operational, and domain review.

References to NIST AI RMF and other governance concepts are practitioner mappings. Verify official sources and adapt the framework to the actual authority, population, risk, and jurisdiction.

Related repositories

RepositoryDistinct role
release-checklistworking config validator
agent-evalevaluation validity and decision semantics
accountability-patternsdecision rights, human review, provenance, and redress
regulated-aistarter repository structure and templates

Maintained by Sima Bagheri.

About

A practical framework for AI release readiness, Lifecycle gates, decision rights, and reusable artifacts for release-stage governance.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

AI Release Governance Framework

NIST AI RMFLicense: MIT

A practitioner framework for turning evaluation, control, operational, and risk evidence into an accountable AI release decision.

The repository is deliberately broader than a checklist and narrower than an enterprise governance operating model. It focuses on the release boundary: what proposition is being approved, which evidence supports it, what remains uncertain, and who owns the decision, conditions, exceptions, residual risks, and rollback.

Start here

ArtifactUse it for
docs/release-decision-record.mddecision outcomes, evidence fields, finding dispositions, conditions, exceptions, and residual-risk acceptance
docs/nist-rmf-mapping.mdcautious practitioner cross-reference to NIST AI RMF functions
release-checklistexecutable YAML validation for a narrower set of release-readiness fields
governance-playbookbroader organizational operating model and recurring governance forums

A release decision is scoped

“Release” should always identify:

  • system and version;
  • model, prompt, retrieval, data, tools, and permissions in scope;
  • environment and infrastructure configuration;
  • user population and geography;
  • enabled actions and external effects;
  • rollout stage and exposure limit;
  • evidence cutoff date;
  • decision owner and authority;
  • conditions, exceptions, and expiry;
  • changes that invalidate the decision.

Approval of one configuration should not be interpreted as approval of future model, prompt, data, tool, permission, or population changes.

Lifecycle

purpose and risk context
↓
release proposition and scope
↓
evidence plan and hard gates
↓
evaluation and control verification
↓
operational / containment readiness
↓
release decision and dispositions
↓
staged rollout and verification
↓
monitoring, incident, change, and retirement

Release governance begins before the final review. The evidence plan, owners, hard gates, and invalidation triggers should be agreed before teams optimize against a convenient metric set.

Gate design

A gate should ask a decision-relevant question, identify the evidence needed, and define the possible dispositions.

1. Purpose, scope, and risk context

Evidence should establish:

  • intended use, users, and prohibited uses;
  • affected people and foreseeable misuse;
  • authority, data, and tool boundaries;
  • reversibility and worst credible consequences;
  • applicable internal and external obligations;
  • owner and release authority.

2. Evaluation validity

Evidence should establish:

  • task and risk coverage;
  • operating-population assumptions and exclusions;
  • hard-gate and quality-measure definitions;
  • evaluator validity and adjudication;
  • repeated-run policy and uncertainty;
  • slice results and material failures;
  • regression coverage.

Do not rely on an aggregate score without reviewing critical failures and non-compensable conditions.

3. Security, privacy, safety, and policy controls

Evidence should establish that relevant controls are:

  • defined against the actual system scope;
  • enforced in the deployment path;
  • tested under normal, adversarial, and failure conditions;
  • attributable to a current version;
  • supported by exception and incident processes;
  • subject to named ownership and change control.

Avoid prescribing a particular tool or explanation technique as universally required. The evidence method should fit the system, decision, and risk.

4. Operational and containment readiness

Evidence should cover:

  • service objectives and capacity assumptions;
  • observability at the agent, tool, workflow, and outcome levels;
  • rollback, disable, and credential-revocation paths;
  • partial and uncertain external actions;
  • data and state recovery;
  • on-call and incident decision rights;
  • user support and redress where applicable;
  • evidence retention with privacy controls.

A rollback test should verify the resulting state, not only demonstrate that an earlier software version can be redeployed.

5. Decision and follow-through

The review should distinguish:

TermMeaning
Hard gatenon-compensable requirement
Blockerunresolved condition preventing the current decision
Evidence gapmissing or unreliable support for a proposition
Required actionfollow-up work accepted under a bounded conditional release
Conditionenforceable limit on scope or operation
Exceptionauthorized deviation with compensating controls and expiry
Residual riskremaining risk accepted by an authorized owner
Observationimprovement that does not currently change the decision

See docs/release-decision-record.md for a template and review questions.

Decision outcomes

OutcomeUse when
Releaserequired hard gates pass and no conditions remain
Release with conditionsno blocker remains, but enforceable constraints or required actions limit the release
Holdevidence, remediation, or control readiness is insufficient
Do not releasecritical failure, prohibited condition, or unacceptable residual risk remains
Defer decisionthe owner postpones judgment until specified evidence or dependency is available

A conditional release must not relabel an unresolved blocker as a future action. Conditions should state scope, owner, measurement, stop trigger, and expiry.

Evidence quality

For each material conclusion, record:

  • proposition or control being assessed;
  • system scope and version;
  • evidence source, author, date, and location;
  • method and reviewer;
  • result and uncertainty;
  • coverage and exclusions;
  • freshness and invalidation trigger;
  • owner and disposition.

A current document can contain stale evidence. Freshness depends on whether the reviewed system and operating conditions have changed, not only the file date.

Staged release

A staged rollout is a control only when it has:

  • a defined population and exposure limit;
  • monitoring linked to plausible failures;
  • named go/no-go decision points;
  • stop and rollback authority;
  • minimum evidence for expansion;
  • user support and incident response;
  • conditions that prevent silent scope growth.

“Pilot” should not become an indefinite production state without renewed evidence and an accountable decision.

Change control

Re-evaluate when material changes occur to:

  • model or provider;
  • prompt, policy, routing, or orchestration;
  • retrieval corpus or data distribution;
  • tools, permissions, identities, or action authority;
  • infrastructure or deployment region;
  • user population or use case;
  • evaluator, rubric, threshold, or test set;
  • relevant law, policy, or internal control;
  • known failure or incident evidence.

The prior decision record should state which changes invalidate it.

Maturity and scope

This is a practitioner framework for planning and reviewing AI releases. It is not a certified release process, safety case, regulatory approval, legal determination, or substitute for qualified security, privacy, safety, compliance, operational, and domain review.

References to NIST AI RMF and other governance concepts are practitioner mappings. Verify official sources and adapt the framework to the actual authority, population, risk, and jurisdiction.

Related repositories

RepositoryDistinct role
release-checklistworking config validator
agent-evalevaluation validity and decision semantics
accountability-patternsdecision rights, human review, provenance, and redress
regulated-aistarter repository structure and templates

Maintained by Sima Bagheri.

About

A practical framework for AI release readiness, Lifecycle gates, decision rights, and reusable artifacts for release-stage governance.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

AI Release Governance Framework

NIST AI RMFLicense: MIT

A practitioner framework for turning evaluation, control, operational, and risk evidence into an accountable AI release decision.

The repository is deliberately broader than a checklist and narrower than an enterprise governance operating model. It focuses on the release boundary: what proposition is being approved, which evidence supports it, what remains uncertain, and who owns the decision, conditions, exceptions, residual risks, and rollback.

Start here

ArtifactUse it for
docs/release-decision-record.mddecision outcomes, evidence fields, finding dispositions, conditions, exceptions, and residual-risk acceptance
docs/nist-rmf-mapping.mdcautious practitioner cross-reference to NIST AI RMF functions
release-checklistexecutable YAML validation for a narrower set of release-readiness fields
governance-playbookbroader organizational operating model and recurring governance forums

A release decision is scoped

“Release” should always identify:

  • system and version;
  • model, prompt, retrieval, data, tools, and permissions in scope;
  • environment and infrastructure configuration;
  • user population and geography;
  • enabled actions and external effects;
  • rollout stage and exposure limit;
  • evidence cutoff date;
  • decision owner and authority;
  • conditions, exceptions, and expiry;
  • changes that invalidate the decision.

Approval of one configuration should not be interpreted as approval of future model, prompt, data, tool, permission, or population changes.

Lifecycle

purpose and risk context
↓
release proposition and scope
↓
evidence plan and hard gates
↓
evaluation and control verification
↓
operational / containment readiness
↓
release decision and dispositions
↓
staged rollout and verification
↓
monitoring, incident, change, and retirement

Release governance begins before the final review. The evidence plan, owners, hard gates, and invalidation triggers should be agreed before teams optimize against a convenient metric set.

Gate design

A gate should ask a decision-relevant question, identify the evidence needed, and define the possible dispositions.

1. Purpose, scope, and risk context

Evidence should establish:

  • intended use, users, and prohibited uses;
  • affected people and foreseeable misuse;
  • authority, data, and tool boundaries;
  • reversibility and worst credible consequences;
  • applicable internal and external obligations;
  • owner and release authority.

2. Evaluation validity

Evidence should establish:

  • task and risk coverage;
  • operating-population assumptions and exclusions;
  • hard-gate and quality-measure definitions;
  • evaluator validity and adjudication;
  • repeated-run policy and uncertainty;
  • slice results and material failures;
  • regression coverage.

Do not rely on an aggregate score without reviewing critical failures and non-compensable conditions.

3. Security, privacy, safety, and policy controls

Evidence should establish that relevant controls are:

  • defined against the actual system scope;
  • enforced in the deployment path;
  • tested under normal, adversarial, and failure conditions;
  • attributable to a current version;
  • supported by exception and incident processes;
  • subject to named ownership and change control.

Avoid prescribing a particular tool or explanation technique as universally required. The evidence method should fit the system, decision, and risk.

4. Operational and containment readiness

Evidence should cover:

  • service objectives and capacity assumptions;
  • observability at the agent, tool, workflow, and outcome levels;
  • rollback, disable, and credential-revocation paths;
  • partial and uncertain external actions;
  • data and state recovery;
  • on-call and incident decision rights;
  • user support and redress where applicable;
  • evidence retention with privacy controls.

A rollback test should verify the resulting state, not only demonstrate that an earlier software version can be redeployed.

5. Decision and follow-through

The review should distinguish:

TermMeaning
Hard gatenon-compensable requirement
Blockerunresolved condition preventing the current decision
Evidence gapmissing or unreliable support for a proposition
Required actionfollow-up work accepted under a bounded conditional release
Conditionenforceable limit on scope or operation
Exceptionauthorized deviation with compensating controls and expiry
Residual riskremaining risk accepted by an authorized owner
Observationimprovement that does not currently change the decision

See docs/release-decision-record.md for a template and review questions.

Decision outcomes

OutcomeUse when
Releaserequired hard gates pass and no conditions remain
Release with conditionsno blocker remains, but enforceable constraints or required actions limit the release
Holdevidence, remediation, or control readiness is insufficient
Do not releasecritical failure, prohibited condition, or unacceptable residual risk remains
Defer decisionthe owner postpones judgment until specified evidence or dependency is available

A conditional release must not relabel an unresolved blocker as a future action. Conditions should state scope, owner, measurement, stop trigger, and expiry.

Evidence quality

For each material conclusion, record:

  • proposition or control being assessed;
  • system scope and version;
  • evidence source, author, date, and location;
  • method and reviewer;
  • result and uncertainty;
  • coverage and exclusions;
  • freshness and invalidation trigger;
  • owner and disposition.

A current document can contain stale evidence. Freshness depends on whether the reviewed system and operating conditions have changed, not only the file date.

Staged release

A staged rollout is a control only when it has:

  • a defined population and exposure limit;
  • monitoring linked to plausible failures;
  • named go/no-go decision points;
  • stop and rollback authority;
  • minimum evidence for expansion;
  • user support and incident response;
  • conditions that prevent silent scope growth.

“Pilot” should not become an indefinite production state without renewed evidence and an accountable decision.

Change control

Re-evaluate when material changes occur to:

  • model or provider;
  • prompt, policy, routing, or orchestration;
  • retrieval corpus or data distribution;
  • tools, permissions, identities, or action authority;
  • infrastructure or deployment region;
  • user population or use case;
  • evaluator, rubric, threshold, or test set;
  • relevant law, policy, or internal control;
  • known failure or incident evidence.

The prior decision record should state which changes invalidate it.

Maturity and scope

This is a practitioner framework for planning and reviewing AI releases. It is not a certified release process, safety case, regulatory approval, legal determination, or substitute for qualified security, privacy, safety, compliance, operational, and domain review.

References to NIST AI RMF and other governance concepts are practitioner mappings. Verify official sources and adapt the framework to the actual authority, population, risk, and jurisdiction.

Related repositories

RepositoryDistinct role
release-checklistworking config validator
agent-evalevaluation validity and decision semantics
accountability-patternsdecision rights, human review, provenance, and redress
regulated-aistarter repository structure and templates

Maintained by Sima Bagheri.

About

A practical framework for AI release readiness, Lifecycle gates, decision rights, and reusable artifacts for release-stage governance.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

AI Release Governance Framework

NIST AI RMFLicense: MIT

A practitioner framework for turning evaluation, control, operational, and risk evidence into an accountable AI release decision.

The repository is deliberately broader than a checklist and narrower than an enterprise governance operating model. It focuses on the release boundary: what proposition is being approved, which evidence supports it, what remains uncertain, and who owns the decision, conditions, exceptions, residual risks, and rollback.

Start here

ArtifactUse it for
docs/release-decision-record.mddecision outcomes, evidence fields, finding dispositions, conditions, exceptions, and residual-risk acceptance
docs/nist-rmf-mapping.mdcautious practitioner cross-reference to NIST AI RMF functions
release-checklistexecutable YAML validation for a narrower set of release-readiness fields
governance-playbookbroader organizational operating model and recurring governance forums

A release decision is scoped

“Release” should always identify:

  • system and version;
  • model, prompt, retrieval, data, tools, and permissions in scope;
  • environment and infrastructure configuration;
  • user population and geography;
  • enabled actions and external effects;
  • rollout stage and exposure limit;
  • evidence cutoff date;
  • decision owner and authority;
  • conditions, exceptions, and expiry;
  • changes that invalidate the decision.

Approval of one configuration should not be interpreted as approval of future model, prompt, data, tool, permission, or population changes.

Lifecycle

purpose and risk context
↓
release proposition and scope
↓
evidence plan and hard gates
↓
evaluation and control verification
↓
operational / containment readiness
↓
release decision and dispositions
↓
staged rollout and verification
↓
monitoring, incident, change, and retirement

Release governance begins before the final review. The evidence plan, owners, hard gates, and invalidation triggers should be agreed before teams optimize against a convenient metric set.

Gate design

A gate should ask a decision-relevant question, identify the evidence needed, and define the possible dispositions.

1. Purpose, scope, and risk context

Evidence should establish:

  • intended use, users, and prohibited uses;
  • affected people and foreseeable misuse;
  • authority, data, and tool boundaries;
  • reversibility and worst credible consequences;
  • applicable internal and external obligations;
  • owner and release authority.

2. Evaluation validity

Evidence should establish:

  • task and risk coverage;
  • operating-population assumptions and exclusions;
  • hard-gate and quality-measure definitions;
  • evaluator validity and adjudication;
  • repeated-run policy and uncertainty;
  • slice results and material failures;
  • regression coverage.

Do not rely on an aggregate score without reviewing critical failures and non-compensable conditions.

3. Security, privacy, safety, and policy controls

Evidence should establish that relevant controls are:

  • defined against the actual system scope;
  • enforced in the deployment path;
  • tested under normal, adversarial, and failure conditions;
  • attributable to a current version;
  • supported by exception and incident processes;
  • subject to named ownership and change control.

Avoid prescribing a particular tool or explanation technique as universally required. The evidence method should fit the system, decision, and risk.

4. Operational and containment readiness

Evidence should cover:

  • service objectives and capacity assumptions;
  • observability at the agent, tool, workflow, and outcome levels;
  • rollback, disable, and credential-revocation paths;
  • partial and uncertain external actions;
  • data and state recovery;
  • on-call and incident decision rights;
  • user support and redress where applicable;
  • evidence retention with privacy controls.

A rollback test should verify the resulting state, not only demonstrate that an earlier software version can be redeployed.

5. Decision and follow-through

The review should distinguish:

TermMeaning
Hard gatenon-compensable requirement
Blockerunresolved condition preventing the current decision
Evidence gapmissing or unreliable support for a proposition
Required actionfollow-up work accepted under a bounded conditional release
Conditionenforceable limit on scope or operation
Exceptionauthorized deviation with compensating controls and expiry
Residual riskremaining risk accepted by an authorized owner
Observationimprovement that does not currently change the decision

See docs/release-decision-record.md for a template and review questions.

Decision outcomes

OutcomeUse when
Releaserequired hard gates pass and no conditions remain
Release with conditionsno blocker remains, but enforceable constraints or required actions limit the release
Holdevidence, remediation, or control readiness is insufficient
Do not releasecritical failure, prohibited condition, or unacceptable residual risk remains
Defer decisionthe owner postpones judgment until specified evidence or dependency is available

A conditional release must not relabel an unresolved blocker as a future action. Conditions should state scope, owner, measurement, stop trigger, and expiry.

Evidence quality

For each material conclusion, record:

  • proposition or control being assessed;
  • system scope and version;
  • evidence source, author, date, and location;
  • method and reviewer;
  • result and uncertainty;
  • coverage and exclusions;
  • freshness and invalidation trigger;
  • owner and disposition.

A current document can contain stale evidence. Freshness depends on whether the reviewed system and operating conditions have changed, not only the file date.

Staged release

A staged rollout is a control only when it has:

  • a defined population and exposure limit;
  • monitoring linked to plausible failures;
  • named go/no-go decision points;
  • stop and rollback authority;
  • minimum evidence for expansion;
  • user support and incident response;
  • conditions that prevent silent scope growth.

“Pilot” should not become an indefinite production state without renewed evidence and an accountable decision.

Change control

Re-evaluate when material changes occur to:

  • model or provider;
  • prompt, policy, routing, or orchestration;
  • retrieval corpus or data distribution;
  • tools, permissions, identities, or action authority;
  • infrastructure or deployment region;
  • user population or use case;
  • evaluator, rubric, threshold, or test set;
  • relevant law, policy, or internal control;
  • known failure or incident evidence.

The prior decision record should state which changes invalidate it.

Maturity and scope

This is a practitioner framework for planning and reviewing AI releases. It is not a certified release process, safety case, regulatory approval, legal determination, or substitute for qualified security, privacy, safety, compliance, operational, and domain review.

References to NIST AI RMF and other governance concepts are practitioner mappings. Verify official sources and adapt the framework to the actual authority, population, risk, and jurisdiction.

Related repositories

RepositoryDistinct role
release-checklistworking config validator
agent-evalevaluation validity and decision semantics
accountability-patternsdecision rights, human review, provenance, and redress
regulated-aistarter repository structure and templates

Maintained by Sima Bagheri.

About

A practical framework for AI release readiness, Lifecycle gates, decision rights, and reusable artifacts for release-stage governance.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

AI Release Governance Framework

NIST AI RMFLicense: MIT

A practitioner framework for turning evaluation, control, operational, and risk evidence into an accountable AI release decision.

The repository is deliberately broader than a checklist and narrower than an enterprise governance operating model. It focuses on the release boundary: what proposition is being approved, which evidence supports it, what remains uncertain, and who owns the decision, conditions, exceptions, residual risks, and rollback.

Start here

ArtifactUse it for
docs/release-decision-record.mddecision outcomes, evidence fields, finding dispositions, conditions, exceptions, and residual-risk acceptance
docs/nist-rmf-mapping.mdcautious practitioner cross-reference to NIST AI RMF functions
release-checklistexecutable YAML validation for a narrower set of release-readiness fields
governance-playbookbroader organizational operating model and recurring governance forums

A release decision is scoped

“Release” should always identify:

  • system and version;
  • model, prompt, retrieval, data, tools, and permissions in scope;
  • environment and infrastructure configuration;
  • user population and geography;
  • enabled actions and external effects;
  • rollout stage and exposure limit;
  • evidence cutoff date;
  • decision owner and authority;
  • conditions, exceptions, and expiry;
  • changes that invalidate the decision.

Approval of one configuration should not be interpreted as approval of future model, prompt, data, tool, permission, or population changes.

Lifecycle

purpose and risk context
↓
release proposition and scope
↓
evidence plan and hard gates
↓
evaluation and control verification
↓
operational / containment readiness
↓
release decision and dispositions
↓
staged rollout and verification
↓
monitoring, incident, change, and retirement

Release governance begins before the final review. The evidence plan, owners, hard gates, and invalidation triggers should be agreed before teams optimize against a convenient metric set.

Gate design

A gate should ask a decision-relevant question, identify the evidence needed, and define the possible dispositions.

1. Purpose, scope, and risk context

Evidence should establish:

  • intended use, users, and prohibited uses;
  • affected people and foreseeable misuse;
  • authority, data, and tool boundaries;
  • reversibility and worst credible consequences;
  • applicable internal and external obligations;
  • owner and release authority.

2. Evaluation validity

Evidence should establish:

  • task and risk coverage;
  • operating-population assumptions and exclusions;
  • hard-gate and quality-measure definitions;
  • evaluator validity and adjudication;
  • repeated-run policy and uncertainty;
  • slice results and material failures;
  • regression coverage.

Do not rely on an aggregate score without reviewing critical failures and non-compensable conditions.

3. Security, privacy, safety, and policy controls

Evidence should establish that relevant controls are:

  • defined against the actual system scope;
  • enforced in the deployment path;
  • tested under normal, adversarial, and failure conditions;
  • attributable to a current version;
  • supported by exception and incident processes;
  • subject to named ownership and change control.

Avoid prescribing a particular tool or explanation technique as universally required. The evidence method should fit the system, decision, and risk.

4. Operational and containment readiness

Evidence should cover:

  • service objectives and capacity assumptions;
  • observability at the agent, tool, workflow, and outcome levels;
  • rollback, disable, and credential-revocation paths;
  • partial and uncertain external actions;
  • data and state recovery;
  • on-call and incident decision rights;
  • user support and redress where applicable;
  • evidence retention with privacy controls.

A rollback test should verify the resulting state, not only demonstrate that an earlier software version can be redeployed.

5. Decision and follow-through

The review should distinguish:

TermMeaning
Hard gatenon-compensable requirement
Blockerunresolved condition preventing the current decision
Evidence gapmissing or unreliable support for a proposition
Required actionfollow-up work accepted under a bounded conditional release
Conditionenforceable limit on scope or operation
Exceptionauthorized deviation with compensating controls and expiry
Residual riskremaining risk accepted by an authorized owner
Observationimprovement that does not currently change the decision

See docs/release-decision-record.md for a template and review questions.

Decision outcomes

OutcomeUse when
Releaserequired hard gates pass and no conditions remain
Release with conditionsno blocker remains, but enforceable constraints or required actions limit the release
Holdevidence, remediation, or control readiness is insufficient
Do not releasecritical failure, prohibited condition, or unacceptable residual risk remains
Defer decisionthe owner postpones judgment until specified evidence or dependency is available

A conditional release must not relabel an unresolved blocker as a future action. Conditions should state scope, owner, measurement, stop trigger, and expiry.

Evidence quality

For each material conclusion, record:

  • proposition or control being assessed;
  • system scope and version;
  • evidence source, author, date, and location;
  • method and reviewer;
  • result and uncertainty;
  • coverage and exclusions;
  • freshness and invalidation trigger;
  • owner and disposition.

A current document can contain stale evidence. Freshness depends on whether the reviewed system and operating conditions have changed, not only the file date.

Staged release

A staged rollout is a control only when it has:

  • a defined population and exposure limit;
  • monitoring linked to plausible failures;
  • named go/no-go decision points;
  • stop and rollback authority;
  • minimum evidence for expansion;
  • user support and incident response;
  • conditions that prevent silent scope growth.

“Pilot” should not become an indefinite production state without renewed evidence and an accountable decision.

Change control

Re-evaluate when material changes occur to:

  • model or provider;
  • prompt, policy, routing, or orchestration;
  • retrieval corpus or data distribution;
  • tools, permissions, identities, or action authority;
  • infrastructure or deployment region;
  • user population or use case;
  • evaluator, rubric, threshold, or test set;
  • relevant law, policy, or internal control;
  • known failure or incident evidence.

The prior decision record should state which changes invalidate it.

Maturity and scope

This is a practitioner framework for planning and reviewing AI releases. It is not a certified release process, safety case, regulatory approval, legal determination, or substitute for qualified security, privacy, safety, compliance, operational, and domain review.

References to NIST AI RMF and other governance concepts are practitioner mappings. Verify official sources and adapt the framework to the actual authority, population, risk, and jurisdiction.

Related repositories

RepositoryDistinct role
release-checklistworking config validator
agent-evalevaluation validity and decision semantics
accountability-patternsdecision rights, human review, provenance, and redress
regulated-aistarter repository structure and templates

Maintained by Sima Bagheri.

About

A practical framework for AI release readiness, Lifecycle gates, decision rights, and reusable artifacts for release-stage governance.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

AI Release Governance Framework

NIST AI RMFLicense: MIT

A practitioner framework for turning evaluation, control, operational, and risk evidence into an accountable AI release decision.

The repository is deliberately broader than a checklist and narrower than an enterprise governance operating model. It focuses on the release boundary: what proposition is being approved, which evidence supports it, what remains uncertain, and who owns the decision, conditions, exceptions, residual risks, and rollback.

Start here

ArtifactUse it for
docs/release-decision-record.mddecision outcomes, evidence fields, finding dispositions, conditions, exceptions, and residual-risk acceptance
docs/nist-rmf-mapping.mdcautious practitioner cross-reference to NIST AI RMF functions
release-checklistexecutable YAML validation for a narrower set of release-readiness fields
governance-playbookbroader organizational operating model and recurring governance forums

A release decision is scoped

“Release” should always identify:

  • system and version;
  • model, prompt, retrieval, data, tools, and permissions in scope;
  • environment and infrastructure configuration;
  • user population and geography;
  • enabled actions and external effects;
  • rollout stage and exposure limit;
  • evidence cutoff date;
  • decision owner and authority;
  • conditions, exceptions, and expiry;
  • changes that invalidate the decision.

Approval of one configuration should not be interpreted as approval of future model, prompt, data, tool, permission, or population changes.

Lifecycle

purpose and risk context
↓
release proposition and scope
↓
evidence plan and hard gates
↓
evaluation and control verification
↓
operational / containment readiness
↓
release decision and dispositions
↓
staged rollout and verification
↓
monitoring, incident, change, and retirement

Release governance begins before the final review. The evidence plan, owners, hard gates, and invalidation triggers should be agreed before teams optimize against a convenient metric set.

Gate design

A gate should ask a decision-relevant question, identify the evidence needed, and define the possible dispositions.

1. Purpose, scope, and risk context

Evidence should establish:

  • intended use, users, and prohibited uses;
  • affected people and foreseeable misuse;
  • authority, data, and tool boundaries;
  • reversibility and worst credible consequences;
  • applicable internal and external obligations;
  • owner and release authority.

2. Evaluation validity

Evidence should establish:

  • task and risk coverage;
  • operating-population assumptions and exclusions;
  • hard-gate and quality-measure definitions;
  • evaluator validity and adjudication;
  • repeated-run policy and uncertainty;
  • slice results and material failures;
  • regression coverage.

Do not rely on an aggregate score without reviewing critical failures and non-compensable conditions.

3. Security, privacy, safety, and policy controls

Evidence should establish that relevant controls are:

  • defined against the actual system scope;
  • enforced in the deployment path;
  • tested under normal, adversarial, and failure conditions;
  • attributable to a current version;
  • supported by exception and incident processes;
  • subject to named ownership and change control.

Avoid prescribing a particular tool or explanation technique as universally required. The evidence method should fit the system, decision, and risk.

4. Operational and containment readiness

Evidence should cover:

  • service objectives and capacity assumptions;
  • observability at the agent, tool, workflow, and outcome levels;
  • rollback, disable, and credential-revocation paths;
  • partial and uncertain external actions;
  • data and state recovery;
  • on-call and incident decision rights;
  • user support and redress where applicable;
  • evidence retention with privacy controls.

A rollback test should verify the resulting state, not only demonstrate that an earlier software version can be redeployed.

5. Decision and follow-through

The review should distinguish:

TermMeaning
Hard gatenon-compensable requirement
Blockerunresolved condition preventing the current decision
Evidence gapmissing or unreliable support for a proposition
Required actionfollow-up work accepted under a bounded conditional release
Conditionenforceable limit on scope or operation
Exceptionauthorized deviation with compensating controls and expiry
Residual riskremaining risk accepted by an authorized owner
Observationimprovement that does not currently change the decision

See docs/release-decision-record.md for a template and review questions.

Decision outcomes

OutcomeUse when
Releaserequired hard gates pass and no conditions remain
Release with conditionsno blocker remains, but enforceable constraints or required actions limit the release
Holdevidence, remediation, or control readiness is insufficient
Do not releasecritical failure, prohibited condition, or unacceptable residual risk remains
Defer decisionthe owner postpones judgment until specified evidence or dependency is available

A conditional release must not relabel an unresolved blocker as a future action. Conditions should state scope, owner, measurement, stop trigger, and expiry.

Evidence quality

For each material conclusion, record:

  • proposition or control being assessed;
  • system scope and version;
  • evidence source, author, date, and location;
  • method and reviewer;
  • result and uncertainty;
  • coverage and exclusions;
  • freshness and invalidation trigger;
  • owner and disposition.

A current document can contain stale evidence. Freshness depends on whether the reviewed system and operating conditions have changed, not only the file date.

Staged release

A staged rollout is a control only when it has:

  • a defined population and exposure limit;
  • monitoring linked to plausible failures;
  • named go/no-go decision points;
  • stop and rollback authority;
  • minimum evidence for expansion;
  • user support and incident response;
  • conditions that prevent silent scope growth.

“Pilot” should not become an indefinite production state without renewed evidence and an accountable decision.

Change control

Re-evaluate when material changes occur to:

  • model or provider;
  • prompt, policy, routing, or orchestration;
  • retrieval corpus or data distribution;
  • tools, permissions, identities, or action authority;
  • infrastructure or deployment region;
  • user population or use case;
  • evaluator, rubric, threshold, or test set;
  • relevant law, policy, or internal control;
  • known failure or incident evidence.

The prior decision record should state which changes invalidate it.

Maturity and scope

This is a practitioner framework for planning and reviewing AI releases. It is not a certified release process, safety case, regulatory approval, legal determination, or substitute for qualified security, privacy, safety, compliance, operational, and domain review.

References to NIST AI RMF and other governance concepts are practitioner mappings. Verify official sources and adapt the framework to the actual authority, population, risk, and jurisdiction.

Related repositories

RepositoryDistinct role
release-checklistworking config validator
agent-evalevaluation validity and decision semantics
accountability-patternsdecision rights, human review, provenance, and redress
regulated-aistarter repository structure and templates

Maintained by Sima Bagheri.

About

A practical framework for AI release readiness, Lifecycle gates, decision rights, and reusable artifacts for release-stage governance.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages