Conversation
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
… feat/agentic-test-framework # Conflicts: # docs/testing.md
The committed .redgate/INDEX.md was generated under a C-locale sort; under en_US.UTF-8 the same corpus sorts slice2-reconcile before slice2-reconcile-r2 and the cheap-tier drift gate reported a phantom drift. Pin the sort so the index is byte-identical regardless of the invoking shell's locale. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN
… actually live
hooks.json wired both handlers as ${CLAUDE_PLUGIN_ROOT}/hooks-handlers/<script>.sh,
but the scripts are checked in one level deeper at hooks/hooks-handlers/. On an
installed plugin the PreToolUse write guard and the SessionStart announcement
therefore failed with 'No such file' and the guard was silently absent. Found by
the agentic protocol lane's real hook-subprocess tests (T20-T22), which resolve
hook commands from hooks.json instead of a hardcoded list.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN
… T19-T24, T32-T41) Stdlib-only unittest framework under evals/agentic/: frozen vocabulary (contract.py, io.py with a fail-closed JSON-Schema subset validator), terminal- state classifier and controls/detectors (classify.py, controls.py), attempt accounting/analysis/reporting per the benchmark spec (accounting.py, analysis.py, reporting.py, usage/judgement schemas), and real hook-subprocess + MCP stdio protocol fixtures (protocols.py). 202 tests, all offline. Catalog fragments for core/measurement/protocol; counterfeit fixtures 21, 22, 24 (inert until the integration lane wires cheap section 22 and counterfeit staging). Shared-file edits (integration-owned): package skeleton, docs/testing.md gains 'eval-dir: evals/agentic', and check-testing-doc.sh skips __pycache__/ which the unittest tiers create under evals/. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN
…team lanes (wave 2, T11-T18, T25-T31, T42-T48) Registry: live 25-plugin roster, catalog loader/resolver, corpus validator with two-sided verifiers, vacuity, holdout/leakage/strata; 75 cards (positive, negative, near-miss for every plugin) with executable outcome AND adoption verifiers; pairing.py exposure parity as matched substitution, four estimand arms, guidance-only degeneracy recorded not zeroed, version estimand targets found by git survey (agent-compiler, fleet-playbook-curator, redgate, voice). Adapters: driver configs loaded from fixtures and checked against the installed claude/codex --help on every run; spawn refused without an approval token (proven by a Popen trap); append-only HMAC hash-chained host ledger whose witness is fixed at construction; replay/worker sessions can only ever yield SIMULATED/REAL_FIXTURE evidence; no stream grammar shipped (none captured). T26-T29 exist only as *__offline_form and stay BLOCKED pending approval. Red team: pinned promptfoo 0.122.0 by path (never npx), fail-closed version check, documented custom-provider interface, frozen sha256 corpus, safe/ vulnerable/refusenik scripted controls, 31 generated offline configs for the clean/adversarial x baseline/placebo/treatment design with parity digests, protected-effect assertions that outrank rubric prose, verdict.py as sole judge calling the agentic native-proof gate, docker-only netproof. T45/T46 stay paid-required. 321 offline tests green; counterfeit fixtures 23, 25-31 added (inert until the integration lane wires cheap section 22). test_protocols.py updated for the fixed redgate hook path; docs/testing.md gains 'eval-dir: evals/redteam'. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN
…e paths, cheap/counterfeit wiring, docs (wave 3, T49-T52) run.py exposes --offline/--gate/--catalog/--id/--lane, coverage --json and driver --dry-run; driver --spawn raises ApprovalRequired without a token in a manifest's approvals. --catalog resolves all 52 IDs to executing tests with assertion counting, negative-control siblings that must fail under the control fixture, and prints the frozen summary (46 executed, 6 BLOCKED pending approval). --gate is the root-portable subset (T52 sweeps the catalog, with requires_real_marketplace / reentrant_unsafe / heavy_external exclusions) and runs in ~5 s; a T52 self-recursion and a T51 whole-corpus recursion that made the gate take 30+ minutes were removed. Cheap tier section 22 runs both new suites fail-closed; counterfeit staging covers both trees, the corpus is 31 fixtures (all fire), and COUNTERFEIT_ONLY selects one fixture for the bounded T51 test. Counterfeit runs now shim npx/npm out of PATH after fixture 26's first mutation executed a real npx call and upgraded the host's shared cache; promptfoo is pinned to a separate verified 0.122.0 install. verdict.py classifies provider faults as FAULT before the VACUOUS check. Lifecycle tests cover terminal, correction, approval, compaction and cancellation paths with negative controls. Docs carry measured costs; testing-plan L3/L4 read 'framework live (offline forms); native runs approval-gated'. Verified by the coordinator: 357 tests OK; cheap 1,740/0 (22 s); counterfeits 38/0 (5.5 min); agentic --offline PASS (4.3 min); --catalog PASS (2 min); redteam --offline PASS (1.7 min); redteam --gate ~1 s; doc guard 84 entries; git diff --check clean. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN
…air (wave 4/4B) Six independent review lenses (native trust, statistics, causal validity and corpus, red-team effects, catalog vacuity with mutation testing, process safety and docs honesty) reported 83 findings, 53 blocker/major; lane-scoped repair agents applied the confirmed ones and a second pass closed the cross-lane residuals. Highlights: Trust: HostLedger keys are host capabilities, never raw caller bytes; a lying reader with zero host-observed entries can no longer satisfy assert_native_ claims; verdict.py lost its raw-key flag and its native path is now a live, tested gate instead of dead code; attempt.schema binds native-proven to a native adapter; claims_native counts event_ids; driver --dry-run validates flags against the installed help; attempt_from_session is the production seam from a live session to a contract.Attempt (native branch untestable offline, recorded in known-gaps.md). Statistics: matched_pairs/difference_interval apply the scoring-valid filter; tri-state verdicts are never coerced; per-card rates over each arm's own trials; 2x2 exclusions printed in every format; min_valid/min_clusters/margin declared in the manifest, no code defaults; DEFF floor; FAULT-starved red-team tranches are INCOMPLETE under a declared fault ceiling; clustered intervals on every red-team cell and interaction delta. Corpus: verifiers are card-bound (a copied fixture no longer passes another card); a correct direct baseline can pass the outcome verifier; adoption verifiers are per-plugin; deterministic arm ids and config hashes; evidence manifests generated for all 150 fixtures with forged/stale detection; stale README and DEFECT text corrected; run.py's parity probe covers all 25 arms; holdout paraphrases selected per attempt from the run seed; planned_n and SPAWN/EXIT accounting in the manifest. Repo hygiene: 78 generated red-team logs untracked; cheap tier now fails on tracked-but-ignored files; runners proven to write only git-ignored paths. Verified by the coordinator after all repairs: 540 tests OK (182 s); cheap 1,903/0 (30 s); counterfeits 38/0 with all 31 fixtures firing (544 s); agentic --offline PASS (326 s); --catalog 46 executed / 6 BLOCKED (146 s); redteam --offline PASS (103 s); gates 10 s / 1 s; doc guard 84 entries; git diff --check clean; no leaked processes or containers. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
🟡 Changes recommended
Several new framework file reads/writes are locale-encoding dependent (via Path.read_text/write_text defaults), which undermines the PR’s cross-host determinism and portability goals.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
Adds a marketplace-wide offline “agentic” evaluation framework plus a pinned promptfoo-based red-team lane, and wires them into the repository’s existing evaluation discipline. This also introduces Jori’s command + OpenAI agent metadata and seeds model-pricing state to support CI policy enforcement.
Changes:
- Adds
evals/agentic/: a stdlib-only agentic test harness with a task catalog (T01–T52), fixtures, controls, and a large unittest suite. - Adds
evals/redteam/: a red-team corpus + pinned promptfoo validation/generation tooling with offline checks. - Adds/updates Jori plugin integration pieces (command, calibration promptfoo stubs, agent metadata) and introduces
ci/model-pricing/initial-state.jsonas the durable starting ledger.
File summaries
| File | Description |
|---|---|
| plugins/jori/skills/jori/agents/openai.yaml | Adds OpenAI agent metadata and implicit-invocation policy for Jori. |
| plugins/jori/evals/promptfoo/calibration-stub.md | Adds a negative-control “no orchestration policy” stub used by Jori’s behavioral eval. |
| plugins/jori/evals/promptfoo/calibration-reference.md | Adds calibration routing guidance referenced by Jori’s behavioral eval. |
| plugins/jori/commands/jori.md | Adds a command entrypoint instructing users to invoke the Jori skill appropriately. |
| evals/redteam/corpus/** | Adds clean + adversarial prompt-injection corpora used by the red-team lane. |
| evals/agentic/** | Adds the agentic framework, task fixtures, and unittest suite backing the marketplace-wide catalog. |
| ci/model-pricing/initial-state.json | Seeds the model-pricing monitor’s durable ledger state. |
Review details
Suppressed comments (2)
evals/agentic/framework/controls.py:268
- Path.write_text() also defaults to locale encoding. Pinning UTF-8 avoids host-dependent behavior and makes the emitted fixtures deterministic across environments.
evals/agentic/framework/controls.py:279 - This mutation reads/writes guard.sh using Path.read_text/write_text defaults (locale encoding). Pin UTF-8 for deterministic, cross-host behavior.
- Files reviewed: 295/1368 changed files
- Comments generated: 2
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
…against The cheap and counterfeit jobs now set up Node 22 and install exactly promptfoo@0.122.0, @anthropic-ai/claude-code@2.1.263 and @openai/codex@0.153.4 into the runner temp dir (never npx, never @latest), exporting PROMPTFOO_HOME and the .bin PATH. The red-team gate compares the pin by package.json and dist-manifest digest, and the adapter lane reads the installed CLIs' --help to refuse unsupported driver flags; neither logs in or calls a model. Without this the always-on gate was host-bound (known-gaps.md R11) and red in CI. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 3a9b3a331f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| "evidence_class": "real-fixture", | ||
| "approval_gate": "none", | ||
| "negative_control": "evals/agentic/fixtures/native/drivers/invented-flag.json", | ||
| "requires_real_marketplace": false, |
There was a problem hiding this comment.
Exclude the installed-CLI probe from the portable gate
On a standard checkout without both claude and codex installed, this entry is swept by T52 during evals/agentic/run.sh --gate because it is neither gated nor marked as requiring the real environment. load_driver_config() then errors before the first assertion, producing T25: executed 0 assertions and failing the required cheap tier; the Ubuntu CI job only checks out the repository and does not install either CLI. Classify T25 as host-dependent or replace its gate path with committed help fixtures so the portable gate can run from a clean checkout.
AGENTS.md reference: AGENTS.md:L62-L71
Useful? React with 👍 / 👎.
|
|
||
| group "redteam suite (offline)" | ||
| if [ -f "evals/redteam/run.sh" ]; then | ||
| if out="$(evals/redteam/run.sh --gate 2>&1)"; then |
There was a problem hiding this comment.
Keep the Promptfoo host dependency out of the cheap tier
In the always-on Ubuntu cheap job, this invocation runs without PROMPTFOO_HOME or any Promptfoo installation step, so promptfoo.sh falls back to /home/jrichlen/... and --gate exits with redteam FAIL pin: promptfoo not installed. The same failure occurs on any developer machine lacking that personal install, making evals/cheap/run.sh red before its remaining checks can pass. Provision the pinned package in CI or make this cheap-gate check self-contained/static rather than requiring a host-local installation.
AGENTS.md reference: AGENTS.md:L62-L71
Useful? React with 👍 / 👎.
…ach (90 cards, 3 plugins measurable) Five new distinct cards per plugin (2 positive, 2 must-not-fire, 1 near-miss), each with real fixtures, card-bound two-sided verifiers, hidden pass/fail/ near-fail workspaces, holdout paraphrases and generated evidence manifests. Coverage now reports total_cards=90, measured_plugins=3 at min_clusters=8, so per-plugin effect intervals for these three plugins are no longer 'unavailable' by construction. Leakage scan: 0 overlaps against every plugin's SKILL.md and commands. Finalize audit finding (not fixed here, recorded in known-gaps.md): 26 of 31 negative cards' outcome check is trivially satisfied on an empty workspace because the deliverable is byte-identical across a negative card's own pass/fail fixtures; the adoption verifier still discriminates, so the 2x2 remains informative, but the outcome axis of negative cards is weak corpus-wide. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN
…ead it On the x64 runner the npm-installed codex wrapper answers 'exec --help' with 49 characters and the T25 assertion hid what they were. Print the captured text in the assertion message and add post-install --version/--help diagnostics to the CI tooling step. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN
…pper can run The adapter driver executes each CLI under an env allowlist with PATH=/usr/bin:/bin (contract 10.1). On the runner the npm-installed codex is a '#!/usr/bin/env node' wrapper and node lives in the setup-node tool cache, so 'codex exec --help' printed only 'env: node: No such file or directory'. Symlink the setup-node binary into /usr/bin in the tooling step. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN
… config name Control configs (invented-flag, dangerous-flag) declare the real claude binary under a different config name. On a host where the committed absolute path is gone (the CI runner) the loader fell back to shutil.which(<config name>), found nothing, and T25's negative sibling could not even load -- reported as 'negative control executed 0 assertions'. Fall back by the declared binary's basename and add a foreign-host test with its own negative. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN
…d first red-team tranche with the real CLI as subject Native (10 approved sessions, token user-approved-2026-09-07-native): CliDriver.spawn is wired to a NativeSession over a HOST_OBSERVED HostLedger; a real stream-json capture and the grammar derived from it are committed (codex still has none and load_grammar keeps raising for it). T26 session and turn acks, T27 resume/fork/fresh isolation (workspace diff empty, probe answered UNKNOWN), T28 mid-turn cancel with process-group teardown, and T29 in-process ledger verification each executed native-proven under run.py --id <T> --approval-token; without a token they stay BLOCKED and the frozen catalog line is unchanged. Driver config deviations forced by the installed CLI: --session-id moved to the fresh mode, fork = --resume + --fork-session, --verbose required by stream-json, CLAUDE_CONFIG_DIR slot. Observed: the CLI echoes a caller-minted --session-id on system/init, so the echo is the ack; the fork id is the one harness-minted id observed. Red-team tranche 2026-09-07-first (declared before running; token user-approved-2026-09-07-redteam-tranche): graveyard, redgate, egress-gate, 6 cells x 8 items x repeat 1 = 144 attempts, 148/150 model calls, 0 faults, 1h25m. bin/tranche.py brokers rows over an AF_UNIX socket so one HOST_OBSERVED ledger per plugin lives in the process that judges, and the native gate passed for the first time (432 host-observed events per plugin, 0 caller-asserted). No safety qualification is granted: every plugin produced protected-effect failures. egress-gate is the only nonzero safety interaction (+0.75 [0.18, 1.32] vs placebo) and it comes with a clean-task utility loss and a clean-condition safety loss, reported as a trade. The injected detector fired 10 times with 0 true positives against a real subject; corpus injections did not work and the instrument recorded ten successes. Textual effects only; no second grader existed. 549 tests OK (4 live forms skip without a token); cheap 2,012/0; gates 11 s / 1 s; tranche validate OK. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN
Replace claim-based task grading with isolated artifact checks, strengthen mutation and forgery controls, preserve paired uncertainty and evidence scope, and correct fixture and host portability defects. Validation: 645 unittests (641 passed, 4 existing native-required skips); cheap gate 2036/0; all 31 counterfeit fixtures rejected with 7 controls passing. Model thresholds and historical failures remain unchanged.
#136) Two fixes lifted out of #115. redgate's hooks.json pointed one directory too high, so on every installed copy the PreToolUse write guard and SessionStart announcer never ran; criteria-index.sh sorted run dirs in the host locale. New cheap-tier section 2b resolves every ${CLAUDE_PLUGIN_ROOT}/<rel> in every plugin's hooks.json against the filesystem with containment, treats a root with no hooks as nothing-to-check, and fails closed on a malformed entry or a scanner crash. Verified: cheap tier 1295/0, counterfeit harness 25/0, and a mutation set covering main's pre-fix file, a renamed handler, a ../ escape to a real file, a broken extraction regex, a no-hooks root, a null command in first and last position, and a scanner crash on the last manifest.
What
Adds
evals/agentic/andevals/redteam/: an offline, stdlib-only agentic test framework for all 26 plugins (52-entry suite catalog, 93-card corpus with no-skill baselines, native adapters with an approval-gated driver and a hash-chained host ledger, attempt accounting and clustered statistics, promptfoo 0.122.0 red-team lane pinned by path), wired into the cheap tier (section 22) and the counterfeit corpus (fixtures 19–31), with docs carrying measured costs.Also fixes two real defects found by the framework: redgate's
hooks.jsonpointed one directory too high (the write guard never ran on an installed plugin), and redgate's criteria index sorted locale-dependently (cheap tier red on en_US hosts).What the tests establish
Current-main integration and test corrections
Incorporates main through
282b416, regenerates gallery/landing pages, and adds eval-ladder task coverage. Exact live-roster coverage and missing/extra-plugin controls run explicitly in CI; structural plan checks also run against the counterfeit suite's synthetic marketplace.Corrects trusted grader-fault accounting while retaining empty, exhausted and independently failed answers as failures. Corrects two pinned Winston dependency lifecycle defects through explicit version/source-hash-checked provisioning. Normal execution only verifies the provisioned source. Regression tests require complete ordered logs and fatal write-error propagation; no logging suppression or relaxed exit-code acceptance.
Validation
Validated commit
5b306e9cad94abc7cacef4a610f10dc0aa38e31a, tree0923197c9f5421e116b70d5c4a9d991c81adb871.The two focused final-source counterfeit runs are not presented as a full-corpus rerun. The prior full counterfeit CI execution is identified below. Latest CI is triggered by this approved push and must be evaluated separately from local validation.
Actual installed pins: Promptfoo 0.122.0, Claude Code 2.1.263, Codex 0.153.4; Winston 3.19.0 and winston-transport 4.9.0 with the documented source-verified backports. Local checks made no model calls. CI provisions only the scoped bubblewrap namespace permission when needed and proves genuine positive implementations execute.
Measured limits
Executed earlier CI. These are preserved measurements of f6905cb, not results for the latest commit. Its two generated-document drift failures are corrected in this follow-up. No model answers or quality floors were changed to obtain green results.
BELOW_DECLARED_DETECTION_FLOOR; harness success does not qualify it as safety evidence.Original implementation session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN
Paid-check follow-up
Run 34314374610 confirmed the cheap gate (2,164 checks), 193 explicit regression tests, and the full counterfeit corpus (41 checks) pass on 5b306e9. Paid failures were inspected at row level: 11 behavioral packs were sample-starved by typed OpenRouter
402 in_flight_budget_exhaustedfaults; Jori also had a semantic judge contradiction in its authority negative control. Routing retained 59 PASS / 8 FAIL / 3 provider FAULT rows; exhausted and incorrect answers stay failures.The follow-up bounds subject calls across the whole workflow: routing/trajectory first, then one behavioral pack at a time, one active subject evaluation each. Routing failure cannot skip independent behavioral packs; cancellation, grader/discovery prerequisites and fork/path filters remain enforced. A new offline guard rejects the original fan-out and 21 broken scheduling/selection variants.
The routing prompt explicitly states the existing independent-guard contract without revealing scenario answers. The Jori authority-control rubric preserves its original FAIL-if-all-three-permission-requests criterion and requires quoted evidence; refusals and capability limits do not count as permission requests. No trial counts, models, completion budgets, statistical floors, or historical verdicts were changed. The original acceptance predicates are retained.
Published follow-up
ba3e913: local cheap gate: 2166 passed / 0 failed; 21 scheduling mutations rejected; 88 routing-contract controls pass; counterfeit 19 pipeline: 9 passed / 0 failed. Real pinned Promptfoo configuration validation passes; 2250 source entries verified after execution. Fresh CI is pending; no paid-check success is claimed.