Skip to content

[Outcome Report] Outcome Health Report — 2026-06-10 #38245

Description

@github-actions

Workflow Health — 2026-06-10

Executive read: 28 outcomes evaluated across 8 runs; 19 items pending (67%), indicating active processing. Most volume is in Issue Monster and Smoke Claude workflows. Acceptance rate is strong at 71%, but zero-touch rate of 0% indicates all accepted items required human engagement.

WorkflowStatusLifecycle healthReferences
Smoke Claude🟨🟩⬜🟥🟨🟨🟥🟨🟨🟨🟨🟨🔴 stuck🟨 38238 · 🟩 comment · ⬜ discussion · 🟥 38197 · 🟨 review
Issue Monster🟨🟨🟨🟨🟨🟨🔴 stuck🟨 38119 · 🟨 38044 · 🟨 38010
Smoke Codex🟩🟨🟨🔴 stuck🟩 comment · 🟨 38239 · 🟨 comment
AI Moderator🟨🔴 stuck
Changeset Generator🟨🟩🟡 in flight🟨 38197 · 🟩 38197
Design Decision Gate 🏗️🟨🟩🟡 in flight🟨 38197 · 🟩 comment
Daily Sentrux Report⚪ underdefined38243
Agent Container Smoke Test🟩🟢 resolving🟩 comment

Legend:

  • Status: 🟩 accepted · 🟥 rejected · 🟨 pending · ⬜ unknown
  • Lifecycle health: 🟢 resolving · 🟡 in flight · 🟠 aging · 🔴 stuck · ⚪ underdefined

🔴 Action Items

  1. Issue Monster is underdefined: All 6 outcomes are pending with no clear acceptance/rejection signal. Review the workflow's prompt and acceptance criteria. The ~56 minute pending age with no follow-up suggests outputs may not align with actual needs.

  2. Smoke Claude needs output quality review: Waste rate of 16.7% (2 rejected items) combined with 8 pending and 1 unknown outcome. Only 8% acceptance rate indicates the workflow's output quality or targeting needs improvement.

  3. Smoke Codex pending backlog: 2 of 3 items pending (~53 minutes), with no engagement. Very small sample size (3 total), but monitor for resolution.

  4. Design Decision Gate has aging pending item: PR Enforce AI credit resolution order; set built-in defaults to 5000 (daily) and 1000 (per-run) #38197 is pending for 1.2+ hours with no review progress. This PR appears to be a focus point across multiple workflows — consider prioritizing its review.

  5. Zero-touch acceptance is 0%: All 5 accepted items required human engagement. This indicates agent output quality could be improved to require less human refinement. Consider reviewing output prompts for clarity and validating that agent outputs match user expectations.

  6. Data quality note: 1 item (3.6% of outcomes) evaluated with only existence fallback. Minor signal, but indicates one outcome type lacks dedicated evaluator logic.

Detailed metrics, evidence quality, workflow counts, and trends

Outcome Scorecard — 2026-06-10

MetricValueStatus
Acceptance rate71%🟡 60-80%
Zero-touch rate0%🔴 <25%
Waste rate7%🟢 <10%
Median time to resolution
Accepted5 / 28
— strong evidence0merged, completed, approved
— medium evidence5engaged, retained
— weak evidence0existence only
Rejected2
Ignored0no observable follow-up
Zero-touch0 / 5
Pending19
Runs checked8

Per-Workflow Breakdown

WorkflowAcceptedRejectedPendingUnknownAcceptanceWaste
Smoke Claude12818%16%
AI Moderator00100%0%
Agent Container Smoke Test1000100%0%
Changeset Generator101050%0%
Daily Sentrux Report00010%0%
Design Decision Gate 🏗️101050%0%
Issue Monster00600%0%
Smoke Codex102033%0%

Evidence Quality

⚠️1 item was evaluated using only a generic existence check (signal: target_exists_only). This contributes to weak evidence and may overstate acceptance. Dedicated evaluators for add_reviewer, submit_pull_request_review, update_issue, update_pull_request, and other types provide stronger evidence.

📊 Measured by Outcome Collector · 60.2 AIC · ⌖ 4.62 AIC

  • expires on Jun 16, 2026, 5:05 PM UTC-08:00

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions