Conversation
Jori demonstration — right-sizing an orchestration planThis task predates Jori. A planning-only worker read the committed Jori skill and reference, then answered the exact pre-skill input below. The wrapper allowed no child agents, external API calls, or edits; the runtime did not independently expose a model or effort setting. Before (verbatim user input):
After (verbatim planning output):
Rules shown: the output scales route to uncertainty, quality consequence, I/O burden, cost, and latency; keeps delegation single-level; bounds work before dispatch; includes whole-cycle allowance; and treats model roles as calibration hypotheses rather than proven rankings. Observed miss: The plan names a separate reviewer but leaves its model and effort, and the round and output allowances, unspecified. Those details still need a dispatch brief. Limits: A real planning worker ran; the proposed worker team was not dispatched or executed. This is not a no-skill baseline and does not establish cost savings, causal lift, or validated Sol/Luna/Terra/Astra settings. |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
🟡 Changes recommended
The new plugins/jori/evals/promptfoo/promptfooconfig.yaml diverges from established repo conventions (schema header + subject token budget), which can reduce editor validation and increase rubric flakiness.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
Adds a new jori plugin to the agent-plugins marketplace, providing a coordination layer for complex work (bounded delegation, authority control, and evidence-backed status), plus supporting eval coverage and marketplace/routing registration.
Changes:
- Adds the
plugins/jori/plugin surface (skill, command, activation fragment, and dashboard template). - Registers
joriacross marketplace + routing roster, and documents the new eval packs indocs/testing.md. - Introduces
jorievals: a cheap invariant pack, a promptfoo behavioral pack, and a counterfeit fixture to ensure invariant drift is caught.
File summaries
| File | Description |
|---|---|
| README.md | Adds jori to the top-level plugin list. |
| plugins/jori/skills/jori/SKILL.md | Defines the Jori coordination invariant and operating rules. |
| plugins/jori/skills/jori/references/orchestration.md | Adds routing/calibration reference material for coordination decisions. |
| plugins/jori/skills/jori/assets/dashboard-template.md | Provides a portable dashboard template for complex work. |
| plugins/jori/skills/jori/agents/openai.yaml | Adds OpenAI agent metadata for implicit invocation + default prompt. |
| plugins/jori/README.md | Documents install/use guidance for the jori plugin. |
| plugins/jori/evals/promptfoo/promptfooconfig.yaml | Adds behavioral promptfoo pack (including calibration negative controls). |
| plugins/jori/evals/promptfoo/prompt.txt | Defines the promptfoo single-turn/tool-less harness prompt. |
| plugins/jori/evals/promptfoo/calibration-stub.md | Adds invariant-free stub for discriminating-scenario calibration. |
| plugins/jori/evals/cheap/checks.sh | Adds cheap-tier checks for manifest wiring and load-bearing invariant phrases. |
| plugins/jori/context/AGENTS.fragment.md | Adds optional activation fragment for persistent coordination behavior. |
| plugins/jori/commands/jori.md | Adds /jori command entrypoint content. |
| plugins/jori/AGENTS.md | Adds plugin entrypoint documentation and restates the invariant. |
| plugins/jori/.codex-plugin/plugin.json | Registers the plugin for Codex with UI metadata and default prompt. |
| plugins/jori/.claude-plugin/plugin.json | Registers the plugin for Claude with marketplace-facing metadata. |
| evals/routing/roster.txt | Adds jori to the routing roster list. |
| evals/counterfeits/fixtures/18-jori-invariant/mutate.sh | Adds a counterfeit mutation to ensure invariant drift is detected. |
| evals/counterfeits/fixtures/18-jori-invariant/DEFECT.md | Documents the expected counterfeit failure substring. |
| docs/testing.md | Documents Jori’s behavioral pack scope/limits and adds jori/* packs to the inventory list. |
| .claude-plugin/marketplace.json | Registers jori in the marketplace catalog. |
Review details
- Files reviewed: 22/22 changed files
- Comments generated: 2
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
CI cost estimate from recorded tokens$0.8919 across both completed runs, calculated before rounding. This is a published-rate estimate for recorded evaluation tokens, not an invoice or the full workflow bill.
Rates verified September 5, 2026: OpenRouter Individual result usage reconciles with aggregate totals: 724,022 tokens, covering Jori, routing, trajectory, and recorded Jori grading. Cached tokens are zero; reasoning is included in output and counted once. The grader identity comes from the evaluation configuration; exports independently identify the subject model. Failed rows are included: the second run's Jori pack was 21/24, including two GitHub-case failures and one authority-control failure. Excluded: calls/retries absent from the artifacts, model-resolution pings, native Codex coordination/reviews, runner charges, and account/provider billing adjustments. Transport request counts and final billed cost remain unknown. Exported |
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Current Jori demonstration — bounded orchestration planningA real native planning worker (requested Luna, medium effort) read the Jori skill and both references at Before — exact user input:
After — exact worker response:
Rules demonstrated: separate known questions from discovery; fan out only distinct evidence paths; consider duplicated context, review and retries; keep model roles provisional; retain approval boundaries. Observed gap: calibration calls for a stronger reference but does not explicitly require an independent model-family judge for our CI comparisons. That requirement still needs to be specified in the evaluation brief. Limits: no no-skill baseline, causal improvement, or savings were measured. This output does not validate GLM dispatch, monitor strategy quality, scheduling, or unattended behavior. |
|
@claude fix failing checks and add baseline no skill tests or benchmarking to evaluate skills |
|
@copilot resolve the merge conflicts in this pull request and fix failing checks |
The branch is already clean: the latest merge commit has no unresolved index entries or conflict markers. I ran the repository cheap suite and isolated Jori checks (all passed: 1,239/0 overall and 16/0 Jori). The latest CI failure is the unrelated routing tier: 4 scenarios fell below its 0.80 floor, with missing ROUTE lines in model output; Jori, install, trajectory, and deterministic checks passed. I did not change unrelated routing behavior. |
Add Jori coordination and optional model-price monitoring
Adds Jori to the marketplace for coordinating complex work through bounded agents and an adaptable project dashboard. The package includes a command, activation context, templates, evidence-based handoffs, and sourced cross-provider routing guidance. GitHub guidance distinguishes Copilot Auto's HyDRA router from the HydraFusion research preview and preserves uncertainty about undocumented controls.
The three active subject evaluation suites now use
z-ai/glm-5.3-flashwith explicitmaxreasoning passed through to OpenRouter. Claude Sonnet remains the independent Jori judge. These configuration changes do not establish cross-model quality parity or change installed Codex models.The optional price monitor checks public OpenRouter metadata daily after activation. It watches five subject candidates plus separate Sonnet OpenRouter pricing, retaining provider, tier, quantization, context, cache, and pricing conditions. Material changes can trigger a bounded GLM strategy proposal. Results are artifacts containing evidence and an allowlisted
apply:falseconfiguration-change plan; the workflow cannot create PRs or apply model changes.Monitoring ships disabled, with strategy disabled and both spending allowances at $0. Proposed activation ceilings are $0.05 per review and $1 per month, backed by a dedicated limited key and durable reservations committed before requests. Unchanged prices require no model call. Activation, paid calibration of proposals, and model promotion require review.
Validation
Reviewed implementation commit:
2567cf0598848f1f4fd52beaaf594d069b4765d4. Current PR headb12d8942a3a307b172e12db225b5acdc5e8c6c48merges main and retains that commit and all three GLM/max configurations.reasoning.effort: maxsetting.The previous run at 1c93242 returned 119 subject HTTP 403 errors from the configured monthly key limit, with no valid model outputs or grader verdicts. Its failures are transport evidence, not a quality comparison. The historical token-cost comment covers only its two named runs; it is not a monthly bill or current-run estimate.
Recurring monitoring is not activated. No merge or automatic model promotion is included.