feat(ISSUE-263): W5-C1 golden set 标定 (#286) - #80
Merged
Conversation
新增 golden set 建设基础设施: - pipeline/adversary/golden_set.py:golden set 加载、盲化注入(剥离 id/来源元数据 同批混排)、全量重放、回归断言(不合格样本仍不合格)。 - pipeline/adversary/calibrate.py:criteria 变更时重新标定,SHA 强一致校验 (标定记录与 criteria SHA 绑定),标定记录落盘(calibrate-record.json); 反特判断言(不合格样本 max score < 阈值)。 - pipeline/adversary/fixtures/golden/:含已知不合格样本(构造独立于判定脚本, 由历史 verdict 标定,不引用 llm_verifier.py 或 evidence_check.py 逻辑)+ 盲化后的混排集。 与 W3-C3 llm_verifier.py、W3-C5 evidence_check.py 接口兼容: - 接受 verifier-report/v1 或 verified-report/v1 作为输入; - 输出 golden-run-report/v1 与 calibrate-record/v1,供下游 CI required check 消费。 Card: Cloudbird-Software/.github#286
randypanding
added a commit
that referenced
this pull request
Aug 26, 2026
对近一周(#21..#124)全部 PR 复盘后的机械债清理:仅删除 AST 级验证 「全仓零引用」的未用导入/未用名,不改任何判定逻辑、阈值、白名单或 policy 数据。逐文件出处: - pipeline/adversary/cnb_bridge.py:删未用 `from typing import Any`(#73/#74) - pipeline/adversary/golden_set.py:删未用 `from typing import Any`(#80/#82/#83) - pipeline/adversary/holdout_registry.py:删未用 `from typing import Any`(#81/#82) - pipeline/adversary/e2e/e2e-runner.py:删未用 `from typing import Any`(#89) - pipeline/adversary/llm_verifier.py:删未用 `import math`;可选库导入行去掉 未用名 extract_score(call_verifier/create_openai_client 均在用,保留)(#72/#76) - pipeline/entropy/tests/test_e2e.py:删未用 `import sys`(#56) - pipeline/selftest-c/tests/test_registry.py:删未用 `import copy`(#103) - pipeline/trust-gate/tests/test_adjudicate.py:删未用 `import copy`(#63) - pipeline/trust-gate/tests/test_cli.py:from-import 去掉未用名 PREDICATES/UNLOCK_STATE(保留 trust_gate 可导入性冒烟导入与 noqa 惯例)(#63) - scripts/dep-supply-chain-check.py:删未用 `import copy`(#36/#43) 刻意不动(已核验非死代码):各模块 `from __future__ import annotations`; fuzz/sast/symbolic 的 `_yamlmini` 双模式导入守卫(noqa F401,保证包路径); golden_set 等 try-import yaml 的环境 fail-closed 守卫;org-gate / suppression-gate / adversary-gate 等关卡 workflow 与 policy/suppressions.yaml 基线数据——门语义一概不变。 验证: - py_compile 全部 scripts/pipeline *.py 通过;bash -n 全部 *.sh 通过 - workflows/policy/pipeline 共 62 个 YAML 解析通过 - scripts/test-integrity-fixtures/run.sh、scripts/suppression-budget-selftest.sh 通过 - python -m unittest:trust-gate test_adjudicate+test_cli 17 例、 selftest-c tests.test_registry 14 例、entropy tests.test_e2e 10 例——全绿 Co-authored-by: randypanding <randypanding@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
W5-C1: golden set 标定(盲化反例 + 逐 run 重放 + 变更重标定)
Closes Cloudbird-Software/.github#286
变更内容
pipeline/adversary/golden_set.py— golden set 加载、盲化注入、全量重放、回归断言pipeline/adversary/calibrate.py— criteria 变更重新标定 + SHA 强一致校验 + 标定记录落盘pipeline/adversary/fixtures/golden/— 已知不合格样本(构造独立于判定脚本)+ 盲化混排集AC 覆盖
接口兼容
llm_verifier.py(verifier-report/v1 输入格式)evidence_check.py(verified-report/v1 输出格式)Card: Cloudbird-Software/.github#286