Skip to content

docs: capability-behavior-vulnerability taxonomy audit (all 80 records) - #224

Merged
chaksaray merged 1 commit into
developfrom
docs/capability-vulnerability-audit
Aug 29, 2026
Merged

docs: capability-behavior-vulnerability taxonomy audit (all 80 records)#224
chaksaray merged 1 commit into
developfrom
docs/capability-vulnerability-audit

Conversation

@chaksaray

@chaksaray chaksaray commented Aug 29, 2026

Copy link
Copy Markdown
Contributor

Full audit against the adopted A/B/C/D/E design: for every record, what security property is violated, and what makes it a vulnerability rather than merely a capability or attacker behavior.

Step 0 (the inserted verification step): confirmed AVE-2026-00074, 00078, and 00080 against the live corpus before treating them as positive controls. All three exist and match what the audit design assumed about them — no correction needed.

Method: full seven-question treatment for Priority 1 (5 capability-heavy records), Priority 2 (11 technique-vs-vulnerability), Priority 3 (12 conventional-vs-agentic), and the 7 positive controls — 35 records total, all classified individually against the real record content, not from memory or the label alone. Sprint 4's remaining 45 records get a real but proportionately lighter pass (classification, boundary, decision, and the Q7 three-line artifact), per the task's own instruction not to review all 80 with equal effort.

Result: 38 VALID-AVE, 23 INSUFFICIENT-BOUNDARY (real mechanism, just needs the violated property stated more explicitly — a prose fix, not structural), 2 TECHNIQUE-CONFLATION (00009, 00032), 1 CAPABILITY-CONFLATION (00038), 16 GENERIC-VULNERABILITY (conventional CWE-territory mechanisms with a real spread of agentic-justification strength — three flagged as a genuinely open question: 00052/00053/00060 are pure MCP-implementation bugs with no behavioral component at all, and AVE's value there is a coverage argument, not a mechanism-level distinction).

Positive controls held up: all seven classified cleanly as A under the same standard applied to everything else — real, positive evidence the method discriminates rather than defaulting either direction, since Priority 1–3 produced a genuine spread across all five classifications on that same standard.

No schema change proposed. security_condition/security_boundary were the two candidates this audit was asked to evaluate. The evidence doesn't support either: every record already has an identifiable boundary and violated property — the real gap was almost always articulation in the existing description/behavioral_fingerprint prose, not a missing structured field. Recommendations (R-001 through R-005) are documentation and process fixes, not data-model changes.

Deliberately does not open the community-invitation umbrella issue. astrogilda and narko4u both have active, current threads (#218/#219, the #214 review offer, the GenAI Crosswalk manual-PR recommendation) that predate this audit. Opening a third simultaneous ask without checking that queue first is what this task's own sequencing note asked not to do by default — noted explicitly in the audit's governance section for whoever picks that up next.

All 80 records still validate, fixtures intact, 342 tests pass (docs-only change, no records touched).


Umbrella tracking/discussion issue for this audit: #225. Not a "closes" relationship deliberately — #225 stays open for challenge and disagreement after this merges, not resolved by the PR landing.

Full audit per the adopted A/B/C/D/E design: for every record, what
security property is violated, and what makes it a vulnerability
rather than merely a capability or attacker behavior.

Step 0 (inserted verification step): confirmed AVE-2026-00074,
00078, and 00080 against the live corpus before treating them as
positive controls -- all three exist and match what the audit design
assumed about them.

Full seven-question treatment for Priority 1 (5 capability-heavy
records), Priority 2 (11 technique-vs-vulnerability records),
Priority 3 (12 conventional-vs-agentic records), and the seven
positive controls -- 35 records total. Sprint 4's remaining 45
records get a real but proportionately lighter pass (classification,
boundary, decision, Q7 three-line artifact), per the task's own
instruction not to review all 80 with equal effort.

Result: 38 VALID-AVE, 23 INSUFFICIENT-BOUNDARY (real mechanism, needs
clearer description text, not a structural change), 2
TECHNIQUE-CONFLATION (00009, 00032), 1 CAPABILITY-CONFLATION (00038),
16 GENERIC-VULNERABILITY (conventional CWE-territory mechanisms with
varying strength of agentic-specific justification -- three flagged
as a genuinely open maintainer question: 00052/00053/00060 are pure
MCP-implementation bugs with no behavioral component at all).

No schema change proposed. security_condition/security_boundary were
the two candidates evaluated; the audit's own evidence is that every
record already has an identifiable boundary and violated property --
the real gap was almost always articulation in existing prose fields,
not a missing structured field.

Deliberately does not open the community-invitation umbrella issue.
astrogilda and narko4u both have active, current threads (#218/#219,
the #214 review offer, the GenAI Crosswalk manual-PR recommendation)
that predate this audit; opening a third simultaneous ask without
checking that queue first is what this task's own sequencing note
asked not to do by default.
@chaksaray chaksaray linked an issue Aug 29, 2026 that may be closed by this pull request
@chaksaray
chaksaray merged commit 824279a into develop Aug 29, 2026
6 checks passed
@chaksaray
chaksaray deleted the docs/capability-vulnerability-audit branch August 29, 2026 04:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Taxonomy Audit: Capability vs Behavior vs Vulnerability

1 participant