Skip to content

RLCR: detect external blockers, allow min/max plan scope, surface stagnation via delta reviews #180

Description

@donglinz

Context

During an RLCR session, the loop ran five rounds (Round 0 → Round 4) before the stagnation circuit breaker forced exit. The plan declared an all-or-nothing scope (lower bound = upper bound = full target deliverable). An external execution-environment blocker appeared from Round 2 onward and was explicitly surfaced with resolution options in the Round 3 summary's update request. The loop continued anyway for two more rounds; the reviewer's verdict drifted from ADVANCED (Rounds 0–2) to STALLED (Rounds 3–4) as the mainline gap list barely changed. Local round contracts were well-scoped and cleanly executed, but they accepted scope reduction the plan forbade, so the reviewer never credited the work as plan-level progress. The methodology produced clean local rounds and the circuit breaker did eventually fire — but it fired after three rounds of diminishing returns.

This issue summarizes the methodology-improvement suggestions distilled from a post-loop analysis. All content is sanitized; no project-specific details below.

Suggestions

Theme 1: Detecting and reacting to external blockers early

1. Treat reported "environment-bound" deferrals as a plan-level event, not a per-round footnote.

Once the implementer reports an unresolvable external constraint that prevents touching the mainline deliverable, every subsequent round will inherit the same constraint and produce only peripheral work. In the observed session, the implementer surfaced the blocker as early as Round 2 and explicitly listed four resolution options in a Round 3 update request, but the loop continued for two more rounds anyway. When a round summary contains an explicit "Goal Tracker Update Request" presenting decision options that only the user can resolve, the methodology should pause the implement-review loop and route to a human-decision gate. Continuing the loop after a clearly-articulated blocker is a form of busy-waiting that the circuit breaker eventually catches, but a dedicated escalation gate would catch it sooner and more cleanly.

2. Distinguish "implementer cannot proceed" from "implementer chose to defer."

Round summaries listed the same set of unimplemented tasks as "deferred" for environment reasons across multiple rounds. The reviewer flagged this each time as "unjustified deferrals" because, per the plan boundaries, no deferrals were valid. Both sides were correct under their own framing, but the loop had no way to converge. The methodology should require the round summary to label each deferred item with one of three causes (scope choice / dependency on earlier in-loop work / external blocker) and require the reviewer to acknowledge that taxonomy. External-blocker items should automatically trigger the escalation gate from suggestion #1 rather than recurring as "unjustified deferral" findings.

Theme 2: Plan rigidity vs. execution reality

3. A plan whose lower bound equals its upper bound is brittle under any unexpected friction.

The plan collapsed the acceptable outcome range to a single point: full feature parity, no partial credit. Any friction (environment, missing tools, scope discovery) renders the entire loop incapable of producing a "complete" result, and the reviewer is structurally forced to return "incomplete" every time. Every review across all five rounds reported zero acceptance criteria fully addressed (or a small partial), even though substantive structural work was landed. This makes the reviewer's per-round verdict almost non-informative — it cannot distinguish "great round" from "bad round" because both report the same overall progress percentage. Plans should declare both a target scope and a minimum acceptable scope, with an explicit decision rule for what to do when execution can hit the minimum but not the target. The reviewer can then meaningfully grade rounds against the minimum boundary while still pointing toward the target.

4. Validate that the execution environment supports the plan before the loop starts.

The plan presumed a working hardware-and-software stack appropriate to the target task. The actual execution environment was missing key prerequisites, and this was discovered only after the first round of work was underway. Rounds 2 onward were entirely shaped by working around the environment instead of executing the plan. The plan-generation phase should produce an explicit "environment prerequisites" checklist, and the loop should run an environment-probe pre-flight (executing minimal smoke checks for each prerequisite) before Round 0. If a prerequisite fails the probe, the plan is revised or the user is asked before any implement-review cycles run.

Theme 3: Review effectiveness and stagnation detection

5. Reviews were specific and actionable, but the stagnation signal was buried.

The reviewer's findings were consistently concrete: precise citations, exact failure modes, and clear "next steps" lists. However, across rounds, large portions of the "required next implementation plan" text were near-verbatim copies of prior rounds. A reader skimming any single review would not realize that the mainline gap list had been almost identical for three consecutive rounds. The signal that the loop was circling lived in the diff between rounds, not in any single round. The reviewer prompt should include a directive to compute a "delta from previous round" summary — explicitly listing which mainline gaps moved, which closed, and which are unchanged. A rising "unchanged mainline gaps" count across rounds is a strong stagnation indicator that should feed into a softer circuit breaker (warning at N=2, hard stop at N=3) rather than waiting for the hard stagnation breaker to fire.

6. Reviews caught real issues with low false-positive rate, but the issues they caught shrank in importance over time.

The severity profile shifted across the five rounds: early reviews concerned the structural contract; later reviews concerned narrowing literal interpretations of one acceptance criterion and re-flagging cleanup hygiene. The methodology spent its last two rounds polishing the edges of a single acceptance criterion while the central acceptance criteria made no progress. The review machinery functioned, but it ran out of high-value work to do. Reviews should be required to bucket findings as mainline-vs-peripheral and to flag when, for two rounds in a row, all closed findings are peripheral. That signal — "we are closing only peripheral issues" — is a clean stagnation criterion that complements the unchanged-mainline-gaps signal.

Theme 4: Communication and round contracts

7. Round contracts narrowed scope appropriately, but the narrowing was not reconciled with the plan boundaries.

Each round contract correctly identified a tractable local objective. Local execution against those contracts was clean. But the contracts implicitly accepted scope reduction that the plan explicitly forbade, and the reviewer correctly refused to recognize that local progress as plan progress. Implementer and reviewer were operating on different success contracts; both were right in their own frame. A round contract should either (a) be a strict subset of the plan-level deliverables and the reviewer is told to grade only that subset, or (b) explicitly request a plan amendment if the round will not advance the plan. Mixing local-scope contracts with plan-scope reviews is what produced the verdict mismatch.

8. Per-round summaries were thorough, but they normalized non-progress.

The implementer's summaries grew progressively better at justifying why the central block was not advanced each round. By the final round, the deferral justification was a polished, internally-consistent paragraph that referenced earlier rounds and a pending user decision. This is good craftsmanship in writing, but it had the effect of making non-progress feel like a stable, defensible state and partially masked the underlying problem. Round summaries should include a quantitative "tasks closed this round" and "tasks closed cumulatively / total tasks" header at the top. Once that ratio plateaus across two rounds, the loop should trigger a planning check-in rather than a sixth round.

What worked well

  • Clean local rounds. Each round was contained, the contract was followed, the validations described were what was actually performed, and the diff between rounds was small and focused.
  • Honest blocker reporting. The implementer surfaced the external blocker as soon as it was confirmed and proposed concrete resolution paths rather than hiding it.
  • Specific, citation-rich reviews. The reviewer consistently pointed to exact citations and the precise mismatch between expected and observed behavior, which made each round's "next action" unambiguous.
  • Circuit breaker eventually fired. The stagnation breaker did catch the loop and force an exit, preventing an unbounded sequence of low-value rounds. The improvement opportunity is making that safety net trigger sooner and more gracefully.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
       blocks
      (function() {
      function addCopyButtons() {
      document.querySelectorAll('pre code').forEach(function(codeBlock) {
      if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
      codeBlock.parentElement.setAttribute('data-copy-added', 'true');
      var btn = document.createElement('button');
      btn.textContent = 'Copy';
      btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
      btn.onmouseover = function() { this.style.opacity = '1'; };
      btn.onmouseout = function() { this.style.opacity = '0.7'; };
      btn.onclick = function() {
      navigator.clipboard.writeText(codeBlock.textContent).then(function() {
      btn.textContent = 'Copied!';
      setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
      });
      };
      codeBlock.parentElement.style.position = 'relative';
      codeBlock.parentElement.appendChild(btn);
      });
      }
      addCopyButtons();
      // Re-run on dynamic content
      var observer = new MutationObserver(addCopyButtons);
      observer.observe(document.body, { childList: true, subtree: true });
      })();
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      RLCR: detect external blockers, allow min/max plan scope, surface stagnation via delta reviews · Issue #180 · PolyArch/humanize · GitHub
      Skip to content

      RLCR: detect external blockers, allow min/max plan scope, surface stagnation via delta reviews #180

      Description

      @donglinz

      Context

      During an RLCR session, the loop ran five rounds (Round 0 → Round 4) before the stagnation circuit breaker forced exit. The plan declared an all-or-nothing scope (lower bound = upper bound = full target deliverable). An external execution-environment blocker appeared from Round 2 onward and was explicitly surfaced with resolution options in the Round 3 summary's update request. The loop continued anyway for two more rounds; the reviewer's verdict drifted from ADVANCED (Rounds 0–2) to STALLED (Rounds 3–4) as the mainline gap list barely changed. Local round contracts were well-scoped and cleanly executed, but they accepted scope reduction the plan forbade, so the reviewer never credited the work as plan-level progress. The methodology produced clean local rounds and the circuit breaker did eventually fire — but it fired after three rounds of diminishing returns.

      This issue summarizes the methodology-improvement suggestions distilled from a post-loop analysis. All content is sanitized; no project-specific details below.

      Suggestions

      Theme 1: Detecting and reacting to external blockers early

      1. Treat reported "environment-bound" deferrals as a plan-level event, not a per-round footnote.

      Once the implementer reports an unresolvable external constraint that prevents touching the mainline deliverable, every subsequent round will inherit the same constraint and produce only peripheral work. In the observed session, the implementer surfaced the blocker as early as Round 2 and explicitly listed four resolution options in a Round 3 update request, but the loop continued for two more rounds anyway. When a round summary contains an explicit "Goal Tracker Update Request" presenting decision options that only the user can resolve, the methodology should pause the implement-review loop and route to a human-decision gate. Continuing the loop after a clearly-articulated blocker is a form of busy-waiting that the circuit breaker eventually catches, but a dedicated escalation gate would catch it sooner and more cleanly.

      2. Distinguish "implementer cannot proceed" from "implementer chose to defer."

      Round summaries listed the same set of unimplemented tasks as "deferred" for environment reasons across multiple rounds. The reviewer flagged this each time as "unjustified deferrals" because, per the plan boundaries, no deferrals were valid. Both sides were correct under their own framing, but the loop had no way to converge. The methodology should require the round summary to label each deferred item with one of three causes (scope choice / dependency on earlier in-loop work / external blocker) and require the reviewer to acknowledge that taxonomy. External-blocker items should automatically trigger the escalation gate from suggestion #1 rather than recurring as "unjustified deferral" findings.

      Theme 2: Plan rigidity vs. execution reality

      3. A plan whose lower bound equals its upper bound is brittle under any unexpected friction.

      The plan collapsed the acceptable outcome range to a single point: full feature parity, no partial credit. Any friction (environment, missing tools, scope discovery) renders the entire loop incapable of producing a "complete" result, and the reviewer is structurally forced to return "incomplete" every time. Every review across all five rounds reported zero acceptance criteria fully addressed (or a small partial), even though substantive structural work was landed. This makes the reviewer's per-round verdict almost non-informative — it cannot distinguish "great round" from "bad round" because both report the same overall progress percentage. Plans should declare both a target scope and a minimum acceptable scope, with an explicit decision rule for what to do when execution can hit the minimum but not the target. The reviewer can then meaningfully grade rounds against the minimum boundary while still pointing toward the target.

      4. Validate that the execution environment supports the plan before the loop starts.

      The plan presumed a working hardware-and-software stack appropriate to the target task. The actual execution environment was missing key prerequisites, and this was discovered only after the first round of work was underway. Rounds 2 onward were entirely shaped by working around the environment instead of executing the plan. The plan-generation phase should produce an explicit "environment prerequisites" checklist, and the loop should run an environment-probe pre-flight (executing minimal smoke checks for each prerequisite) before Round 0. If a prerequisite fails the probe, the plan is revised or the user is asked before any implement-review cycles run.

      Theme 3: Review effectiveness and stagnation detection

      5. Reviews were specific and actionable, but the stagnation signal was buried.

      The reviewer's findings were consistently concrete: precise citations, exact failure modes, and clear "next steps" lists. However, across rounds, large portions of the "required next implementation plan" text were near-verbatim copies of prior rounds. A reader skimming any single review would not realize that the mainline gap list had been almost identical for three consecutive rounds. The signal that the loop was circling lived in the diff between rounds, not in any single round. The reviewer prompt should include a directive to compute a "delta from previous round" summary — explicitly listing which mainline gaps moved, which closed, and which are unchanged. A rising "unchanged mainline gaps" count across rounds is a strong stagnation indicator that should feed into a softer circuit breaker (warning at N=2, hard stop at N=3) rather than waiting for the hard stagnation breaker to fire.

      6. Reviews caught real issues with low false-positive rate, but the issues they caught shrank in importance over time.

      The severity profile shifted across the five rounds: early reviews concerned the structural contract; later reviews concerned narrowing literal interpretations of one acceptance criterion and re-flagging cleanup hygiene. The methodology spent its last two rounds polishing the edges of a single acceptance criterion while the central acceptance criteria made no progress. The review machinery functioned, but it ran out of high-value work to do. Reviews should be required to bucket findings as mainline-vs-peripheral and to flag when, for two rounds in a row, all closed findings are peripheral. That signal — "we are closing only peripheral issues" — is a clean stagnation criterion that complements the unchanged-mainline-gaps signal.

      Theme 4: Communication and round contracts

      7. Round contracts narrowed scope appropriately, but the narrowing was not reconciled with the plan boundaries.

      Each round contract correctly identified a tractable local objective. Local execution against those contracts was clean. But the contracts implicitly accepted scope reduction that the plan explicitly forbade, and the reviewer correctly refused to recognize that local progress as plan progress. Implementer and reviewer were operating on different success contracts; both were right in their own frame. A round contract should either (a) be a strict subset of the plan-level deliverables and the reviewer is told to grade only that subset, or (b) explicitly request a plan amendment if the round will not advance the plan. Mixing local-scope contracts with plan-scope reviews is what produced the verdict mismatch.

      8. Per-round summaries were thorough, but they normalized non-progress.

      The implementer's summaries grew progressively better at justifying why the central block was not advanced each round. By the final round, the deferral justification was a polished, internally-consistent paragraph that referenced earlier rounds and a pending user decision. This is good craftsmanship in writing, but it had the effect of making non-progress feel like a stable, defensible state and partially masked the underlying problem. Round summaries should include a quantitative "tasks closed this round" and "tasks closed cumulatively / total tasks" header at the top. Once that ratio plateaus across two rounds, the loop should trigger a planning check-in rather than a sixth round.

      What worked well

      • Clean local rounds. Each round was contained, the contract was followed, the validations described were what was actually performed, and the diff between rounds was small and focused.
      • Honest blocker reporting. The implementer surfaced the external blocker as soon as it was confirmed and proposed concrete resolution paths rather than hiding it.
      • Specific, citation-rich reviews. The reviewer consistently pointed to exact citations and the precise mismatch between expected and observed behavior, which made each round's "next action" unambiguous.
      • Circuit breaker eventually fired. The stagnation breaker did catch the loop and force an exit, preventing an unbounded sequence of low-value rounds. The improvement opportunity is making that safety net trigger sooner and more gracefully.

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        No labels
        No labels

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' RLCR: detect external blockers, allow min/max plan scope, surface stagnation via delta reviews · Issue #180 · PolyArch/humanize · GitHub
          Skip to content

          RLCR: detect external blockers, allow min/max plan scope, surface stagnation via delta reviews #180

          Description

          @donglinz

          Context

          During an RLCR session, the loop ran five rounds (Round 0 → Round 4) before the stagnation circuit breaker forced exit. The plan declared an all-or-nothing scope (lower bound = upper bound = full target deliverable). An external execution-environment blocker appeared from Round 2 onward and was explicitly surfaced with resolution options in the Round 3 summary's update request. The loop continued anyway for two more rounds; the reviewer's verdict drifted from ADVANCED (Rounds 0–2) to STALLED (Rounds 3–4) as the mainline gap list barely changed. Local round contracts were well-scoped and cleanly executed, but they accepted scope reduction the plan forbade, so the reviewer never credited the work as plan-level progress. The methodology produced clean local rounds and the circuit breaker did eventually fire — but it fired after three rounds of diminishing returns.

          This issue summarizes the methodology-improvement suggestions distilled from a post-loop analysis. All content is sanitized; no project-specific details below.

          Suggestions

          Theme 1: Detecting and reacting to external blockers early

          1. Treat reported "environment-bound" deferrals as a plan-level event, not a per-round footnote.

          Once the implementer reports an unresolvable external constraint that prevents touching the mainline deliverable, every subsequent round will inherit the same constraint and produce only peripheral work. In the observed session, the implementer surfaced the blocker as early as Round 2 and explicitly listed four resolution options in a Round 3 update request, but the loop continued for two more rounds anyway. When a round summary contains an explicit "Goal Tracker Update Request" presenting decision options that only the user can resolve, the methodology should pause the implement-review loop and route to a human-decision gate. Continuing the loop after a clearly-articulated blocker is a form of busy-waiting that the circuit breaker eventually catches, but a dedicated escalation gate would catch it sooner and more cleanly.

          2. Distinguish "implementer cannot proceed" from "implementer chose to defer."

          Round summaries listed the same set of unimplemented tasks as "deferred" for environment reasons across multiple rounds. The reviewer flagged this each time as "unjustified deferrals" because, per the plan boundaries, no deferrals were valid. Both sides were correct under their own framing, but the loop had no way to converge. The methodology should require the round summary to label each deferred item with one of three causes (scope choice / dependency on earlier in-loop work / external blocker) and require the reviewer to acknowledge that taxonomy. External-blocker items should automatically trigger the escalation gate from suggestion #1 rather than recurring as "unjustified deferral" findings.

          Theme 2: Plan rigidity vs. execution reality

          3. A plan whose lower bound equals its upper bound is brittle under any unexpected friction.

          The plan collapsed the acceptable outcome range to a single point: full feature parity, no partial credit. Any friction (environment, missing tools, scope discovery) renders the entire loop incapable of producing a "complete" result, and the reviewer is structurally forced to return "incomplete" every time. Every review across all five rounds reported zero acceptance criteria fully addressed (or a small partial), even though substantive structural work was landed. This makes the reviewer's per-round verdict almost non-informative — it cannot distinguish "great round" from "bad round" because both report the same overall progress percentage. Plans should declare both a target scope and a minimum acceptable scope, with an explicit decision rule for what to do when execution can hit the minimum but not the target. The reviewer can then meaningfully grade rounds against the minimum boundary while still pointing toward the target.

          4. Validate that the execution environment supports the plan before the loop starts.

          The plan presumed a working hardware-and-software stack appropriate to the target task. The actual execution environment was missing key prerequisites, and this was discovered only after the first round of work was underway. Rounds 2 onward were entirely shaped by working around the environment instead of executing the plan. The plan-generation phase should produce an explicit "environment prerequisites" checklist, and the loop should run an environment-probe pre-flight (executing minimal smoke checks for each prerequisite) before Round 0. If a prerequisite fails the probe, the plan is revised or the user is asked before any implement-review cycles run.

          Theme 3: Review effectiveness and stagnation detection

          5. Reviews were specific and actionable, but the stagnation signal was buried.

          The reviewer's findings were consistently concrete: precise citations, exact failure modes, and clear "next steps" lists. However, across rounds, large portions of the "required next implementation plan" text were near-verbatim copies of prior rounds. A reader skimming any single review would not realize that the mainline gap list had been almost identical for three consecutive rounds. The signal that the loop was circling lived in the diff between rounds, not in any single round. The reviewer prompt should include a directive to compute a "delta from previous round" summary — explicitly listing which mainline gaps moved, which closed, and which are unchanged. A rising "unchanged mainline gaps" count across rounds is a strong stagnation indicator that should feed into a softer circuit breaker (warning at N=2, hard stop at N=3) rather than waiting for the hard stagnation breaker to fire.

          6. Reviews caught real issues with low false-positive rate, but the issues they caught shrank in importance over time.

          The severity profile shifted across the five rounds: early reviews concerned the structural contract; later reviews concerned narrowing literal interpretations of one acceptance criterion and re-flagging cleanup hygiene. The methodology spent its last two rounds polishing the edges of a single acceptance criterion while the central acceptance criteria made no progress. The review machinery functioned, but it ran out of high-value work to do. Reviews should be required to bucket findings as mainline-vs-peripheral and to flag when, for two rounds in a row, all closed findings are peripheral. That signal — "we are closing only peripheral issues" — is a clean stagnation criterion that complements the unchanged-mainline-gaps signal.

          Theme 4: Communication and round contracts

          7. Round contracts narrowed scope appropriately, but the narrowing was not reconciled with the plan boundaries.

          Each round contract correctly identified a tractable local objective. Local execution against those contracts was clean. But the contracts implicitly accepted scope reduction that the plan explicitly forbade, and the reviewer correctly refused to recognize that local progress as plan progress. Implementer and reviewer were operating on different success contracts; both were right in their own frame. A round contract should either (a) be a strict subset of the plan-level deliverables and the reviewer is told to grade only that subset, or (b) explicitly request a plan amendment if the round will not advance the plan. Mixing local-scope contracts with plan-scope reviews is what produced the verdict mismatch.

          8. Per-round summaries were thorough, but they normalized non-progress.

          The implementer's summaries grew progressively better at justifying why the central block was not advanced each round. By the final round, the deferral justification was a polished, internally-consistent paragraph that referenced earlier rounds and a pending user decision. This is good craftsmanship in writing, but it had the effect of making non-progress feel like a stable, defensible state and partially masked the underlying problem. Round summaries should include a quantitative "tasks closed this round" and "tasks closed cumulatively / total tasks" header at the top. Once that ratio plateaus across two rounds, the loop should trigger a planning check-in rather than a sixth round.

          What worked well

          • Clean local rounds. Each round was contained, the contract was followed, the validations described were what was actually performed, and the diff between rounds was small and focused.
          • Honest blocker reporting. The implementer surfaced the external blocker as soon as it was confirmed and proposed concrete resolution paths rather than hiding it.
          • Specific, citation-rich reviews. The reviewer consistently pointed to exact citations and the precise mismatch between expected and observed behavior, which made each round's "next action" unambiguous.
          • Circuit breaker eventually fired. The stagnation breaker did catch the loop and force an exit, preventing an unbounded sequence of low-value rounds. The improvement opportunity is making that safety net trigger sooner and more gracefully.

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            No labels
            No labels

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' RLCR: detect external blockers, allow min/max plan scope, surface stagnation via delta reviews · Issue #180 · PolyArch/humanize · GitHub
              Skip to content

              RLCR: detect external blockers, allow min/max plan scope, surface stagnation via delta reviews #180

              Description

              @donglinz

              Context

              During an RLCR session, the loop ran five rounds (Round 0 → Round 4) before the stagnation circuit breaker forced exit. The plan declared an all-or-nothing scope (lower bound = upper bound = full target deliverable). An external execution-environment blocker appeared from Round 2 onward and was explicitly surfaced with resolution options in the Round 3 summary's update request. The loop continued anyway for two more rounds; the reviewer's verdict drifted from ADVANCED (Rounds 0–2) to STALLED (Rounds 3–4) as the mainline gap list barely changed. Local round contracts were well-scoped and cleanly executed, but they accepted scope reduction the plan forbade, so the reviewer never credited the work as plan-level progress. The methodology produced clean local rounds and the circuit breaker did eventually fire — but it fired after three rounds of diminishing returns.

              This issue summarizes the methodology-improvement suggestions distilled from a post-loop analysis. All content is sanitized; no project-specific details below.

              Suggestions

              Theme 1: Detecting and reacting to external blockers early

              1. Treat reported "environment-bound" deferrals as a plan-level event, not a per-round footnote.

              Once the implementer reports an unresolvable external constraint that prevents touching the mainline deliverable, every subsequent round will inherit the same constraint and produce only peripheral work. In the observed session, the implementer surfaced the blocker as early as Round 2 and explicitly listed four resolution options in a Round 3 update request, but the loop continued for two more rounds anyway. When a round summary contains an explicit "Goal Tracker Update Request" presenting decision options that only the user can resolve, the methodology should pause the implement-review loop and route to a human-decision gate. Continuing the loop after a clearly-articulated blocker is a form of busy-waiting that the circuit breaker eventually catches, but a dedicated escalation gate would catch it sooner and more cleanly.

              2. Distinguish "implementer cannot proceed" from "implementer chose to defer."

              Round summaries listed the same set of unimplemented tasks as "deferred" for environment reasons across multiple rounds. The reviewer flagged this each time as "unjustified deferrals" because, per the plan boundaries, no deferrals were valid. Both sides were correct under their own framing, but the loop had no way to converge. The methodology should require the round summary to label each deferred item with one of three causes (scope choice / dependency on earlier in-loop work / external blocker) and require the reviewer to acknowledge that taxonomy. External-blocker items should automatically trigger the escalation gate from suggestion #1 rather than recurring as "unjustified deferral" findings.

              Theme 2: Plan rigidity vs. execution reality

              3. A plan whose lower bound equals its upper bound is brittle under any unexpected friction.

              The plan collapsed the acceptable outcome range to a single point: full feature parity, no partial credit. Any friction (environment, missing tools, scope discovery) renders the entire loop incapable of producing a "complete" result, and the reviewer is structurally forced to return "incomplete" every time. Every review across all five rounds reported zero acceptance criteria fully addressed (or a small partial), even though substantive structural work was landed. This makes the reviewer's per-round verdict almost non-informative — it cannot distinguish "great round" from "bad round" because both report the same overall progress percentage. Plans should declare both a target scope and a minimum acceptable scope, with an explicit decision rule for what to do when execution can hit the minimum but not the target. The reviewer can then meaningfully grade rounds against the minimum boundary while still pointing toward the target.

              4. Validate that the execution environment supports the plan before the loop starts.

              The plan presumed a working hardware-and-software stack appropriate to the target task. The actual execution environment was missing key prerequisites, and this was discovered only after the first round of work was underway. Rounds 2 onward were entirely shaped by working around the environment instead of executing the plan. The plan-generation phase should produce an explicit "environment prerequisites" checklist, and the loop should run an environment-probe pre-flight (executing minimal smoke checks for each prerequisite) before Round 0. If a prerequisite fails the probe, the plan is revised or the user is asked before any implement-review cycles run.

              Theme 3: Review effectiveness and stagnation detection

              5. Reviews were specific and actionable, but the stagnation signal was buried.

              The reviewer's findings were consistently concrete: precise citations, exact failure modes, and clear "next steps" lists. However, across rounds, large portions of the "required next implementation plan" text were near-verbatim copies of prior rounds. A reader skimming any single review would not realize that the mainline gap list had been almost identical for three consecutive rounds. The signal that the loop was circling lived in the diff between rounds, not in any single round. The reviewer prompt should include a directive to compute a "delta from previous round" summary — explicitly listing which mainline gaps moved, which closed, and which are unchanged. A rising "unchanged mainline gaps" count across rounds is a strong stagnation indicator that should feed into a softer circuit breaker (warning at N=2, hard stop at N=3) rather than waiting for the hard stagnation breaker to fire.

              6. Reviews caught real issues with low false-positive rate, but the issues they caught shrank in importance over time.

              The severity profile shifted across the five rounds: early reviews concerned the structural contract; later reviews concerned narrowing literal interpretations of one acceptance criterion and re-flagging cleanup hygiene. The methodology spent its last two rounds polishing the edges of a single acceptance criterion while the central acceptance criteria made no progress. The review machinery functioned, but it ran out of high-value work to do. Reviews should be required to bucket findings as mainline-vs-peripheral and to flag when, for two rounds in a row, all closed findings are peripheral. That signal — "we are closing only peripheral issues" — is a clean stagnation criterion that complements the unchanged-mainline-gaps signal.

              Theme 4: Communication and round contracts

              7. Round contracts narrowed scope appropriately, but the narrowing was not reconciled with the plan boundaries.

              Each round contract correctly identified a tractable local objective. Local execution against those contracts was clean. But the contracts implicitly accepted scope reduction that the plan explicitly forbade, and the reviewer correctly refused to recognize that local progress as plan progress. Implementer and reviewer were operating on different success contracts; both were right in their own frame. A round contract should either (a) be a strict subset of the plan-level deliverables and the reviewer is told to grade only that subset, or (b) explicitly request a plan amendment if the round will not advance the plan. Mixing local-scope contracts with plan-scope reviews is what produced the verdict mismatch.

              8. Per-round summaries were thorough, but they normalized non-progress.

              The implementer's summaries grew progressively better at justifying why the central block was not advanced each round. By the final round, the deferral justification was a polished, internally-consistent paragraph that referenced earlier rounds and a pending user decision. This is good craftsmanship in writing, but it had the effect of making non-progress feel like a stable, defensible state and partially masked the underlying problem. Round summaries should include a quantitative "tasks closed this round" and "tasks closed cumulatively / total tasks" header at the top. Once that ratio plateaus across two rounds, the loop should trigger a planning check-in rather than a sixth round.

              What worked well

              • Clean local rounds. Each round was contained, the contract was followed, the validations described were what was actually performed, and the diff between rounds was small and focused.
              • Honest blocker reporting. The implementer surfaced the external blocker as soon as it was confirmed and proposed concrete resolution paths rather than hiding it.
              • Specific, citation-rich reviews. The reviewer consistently pointed to exact citations and the precise mismatch between expected and observed behavior, which made each round's "next action" unambiguous.
              • Circuit breaker eventually fired. The stagnation breaker did catch the loop and force an exit, preventing an unbounded sequence of low-value rounds. The improvement opportunity is making that safety net trigger sooner and more gracefully.

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                No labels
                No labels

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' RLCR: detect external blockers, allow min/max plan scope, surface stagnation via delta reviews · Issue #180 · PolyArch/humanize · GitHub
                  Skip to content

                  RLCR: detect external blockers, allow min/max plan scope, surface stagnation via delta reviews #180

                  Description

                  @donglinz

                  Context

                  During an RLCR session, the loop ran five rounds (Round 0 → Round 4) before the stagnation circuit breaker forced exit. The plan declared an all-or-nothing scope (lower bound = upper bound = full target deliverable). An external execution-environment blocker appeared from Round 2 onward and was explicitly surfaced with resolution options in the Round 3 summary's update request. The loop continued anyway for two more rounds; the reviewer's verdict drifted from ADVANCED (Rounds 0–2) to STALLED (Rounds 3–4) as the mainline gap list barely changed. Local round contracts were well-scoped and cleanly executed, but they accepted scope reduction the plan forbade, so the reviewer never credited the work as plan-level progress. The methodology produced clean local rounds and the circuit breaker did eventually fire — but it fired after three rounds of diminishing returns.

                  This issue summarizes the methodology-improvement suggestions distilled from a post-loop analysis. All content is sanitized; no project-specific details below.

                  Suggestions

                  Theme 1: Detecting and reacting to external blockers early

                  1. Treat reported "environment-bound" deferrals as a plan-level event, not a per-round footnote.

                  Once the implementer reports an unresolvable external constraint that prevents touching the mainline deliverable, every subsequent round will inherit the same constraint and produce only peripheral work. In the observed session, the implementer surfaced the blocker as early as Round 2 and explicitly listed four resolution options in a Round 3 update request, but the loop continued for two more rounds anyway. When a round summary contains an explicit "Goal Tracker Update Request" presenting decision options that only the user can resolve, the methodology should pause the implement-review loop and route to a human-decision gate. Continuing the loop after a clearly-articulated blocker is a form of busy-waiting that the circuit breaker eventually catches, but a dedicated escalation gate would catch it sooner and more cleanly.

                  2. Distinguish "implementer cannot proceed" from "implementer chose to defer."

                  Round summaries listed the same set of unimplemented tasks as "deferred" for environment reasons across multiple rounds. The reviewer flagged this each time as "unjustified deferrals" because, per the plan boundaries, no deferrals were valid. Both sides were correct under their own framing, but the loop had no way to converge. The methodology should require the round summary to label each deferred item with one of three causes (scope choice / dependency on earlier in-loop work / external blocker) and require the reviewer to acknowledge that taxonomy. External-blocker items should automatically trigger the escalation gate from suggestion #1 rather than recurring as "unjustified deferral" findings.

                  Theme 2: Plan rigidity vs. execution reality

                  3. A plan whose lower bound equals its upper bound is brittle under any unexpected friction.

                  The plan collapsed the acceptable outcome range to a single point: full feature parity, no partial credit. Any friction (environment, missing tools, scope discovery) renders the entire loop incapable of producing a "complete" result, and the reviewer is structurally forced to return "incomplete" every time. Every review across all five rounds reported zero acceptance criteria fully addressed (or a small partial), even though substantive structural work was landed. This makes the reviewer's per-round verdict almost non-informative — it cannot distinguish "great round" from "bad round" because both report the same overall progress percentage. Plans should declare both a target scope and a minimum acceptable scope, with an explicit decision rule for what to do when execution can hit the minimum but not the target. The reviewer can then meaningfully grade rounds against the minimum boundary while still pointing toward the target.

                  4. Validate that the execution environment supports the plan before the loop starts.

                  The plan presumed a working hardware-and-software stack appropriate to the target task. The actual execution environment was missing key prerequisites, and this was discovered only after the first round of work was underway. Rounds 2 onward were entirely shaped by working around the environment instead of executing the plan. The plan-generation phase should produce an explicit "environment prerequisites" checklist, and the loop should run an environment-probe pre-flight (executing minimal smoke checks for each prerequisite) before Round 0. If a prerequisite fails the probe, the plan is revised or the user is asked before any implement-review cycles run.

                  Theme 3: Review effectiveness and stagnation detection

                  5. Reviews were specific and actionable, but the stagnation signal was buried.

                  The reviewer's findings were consistently concrete: precise citations, exact failure modes, and clear "next steps" lists. However, across rounds, large portions of the "required next implementation plan" text were near-verbatim copies of prior rounds. A reader skimming any single review would not realize that the mainline gap list had been almost identical for three consecutive rounds. The signal that the loop was circling lived in the diff between rounds, not in any single round. The reviewer prompt should include a directive to compute a "delta from previous round" summary — explicitly listing which mainline gaps moved, which closed, and which are unchanged. A rising "unchanged mainline gaps" count across rounds is a strong stagnation indicator that should feed into a softer circuit breaker (warning at N=2, hard stop at N=3) rather than waiting for the hard stagnation breaker to fire.

                  6. Reviews caught real issues with low false-positive rate, but the issues they caught shrank in importance over time.

                  The severity profile shifted across the five rounds: early reviews concerned the structural contract; later reviews concerned narrowing literal interpretations of one acceptance criterion and re-flagging cleanup hygiene. The methodology spent its last two rounds polishing the edges of a single acceptance criterion while the central acceptance criteria made no progress. The review machinery functioned, but it ran out of high-value work to do. Reviews should be required to bucket findings as mainline-vs-peripheral and to flag when, for two rounds in a row, all closed findings are peripheral. That signal — "we are closing only peripheral issues" — is a clean stagnation criterion that complements the unchanged-mainline-gaps signal.

                  Theme 4: Communication and round contracts

                  7. Round contracts narrowed scope appropriately, but the narrowing was not reconciled with the plan boundaries.

                  Each round contract correctly identified a tractable local objective. Local execution against those contracts was clean. But the contracts implicitly accepted scope reduction that the plan explicitly forbade, and the reviewer correctly refused to recognize that local progress as plan progress. Implementer and reviewer were operating on different success contracts; both were right in their own frame. A round contract should either (a) be a strict subset of the plan-level deliverables and the reviewer is told to grade only that subset, or (b) explicitly request a plan amendment if the round will not advance the plan. Mixing local-scope contracts with plan-scope reviews is what produced the verdict mismatch.

                  8. Per-round summaries were thorough, but they normalized non-progress.

                  The implementer's summaries grew progressively better at justifying why the central block was not advanced each round. By the final round, the deferral justification was a polished, internally-consistent paragraph that referenced earlier rounds and a pending user decision. This is good craftsmanship in writing, but it had the effect of making non-progress feel like a stable, defensible state and partially masked the underlying problem. Round summaries should include a quantitative "tasks closed this round" and "tasks closed cumulatively / total tasks" header at the top. Once that ratio plateaus across two rounds, the loop should trigger a planning check-in rather than a sixth round.

                  What worked well

                  • Clean local rounds. Each round was contained, the contract was followed, the validations described were what was actually performed, and the diff between rounds was small and focused.
                  • Honest blocker reporting. The implementer surfaced the external blocker as soon as it was confirmed and proposed concrete resolution paths rather than hiding it.
                  • Specific, citation-rich reviews. The reviewer consistently pointed to exact citations and the precise mismatch between expected and observed behavior, which made each round's "next action" unambiguous.
                  • Circuit breaker eventually fired. The stagnation breaker did catch the loop and force an exit, preventing an unbounded sequence of low-value rounds. The improvement opportunity is making that safety net trigger sooner and more gracefully.

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    No labels
                    No labels

                    Type

                    No type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' RLCR: detect external blockers, allow min/max plan scope, surface stagnation via delta reviews · Issue #180 · PolyArch/humanize · GitHub
                      Skip to content

                      RLCR: detect external blockers, allow min/max plan scope, surface stagnation via delta reviews #180

                      Description

                      @donglinz

                      Context

                      During an RLCR session, the loop ran five rounds (Round 0 → Round 4) before the stagnation circuit breaker forced exit. The plan declared an all-or-nothing scope (lower bound = upper bound = full target deliverable). An external execution-environment blocker appeared from Round 2 onward and was explicitly surfaced with resolution options in the Round 3 summary's update request. The loop continued anyway for two more rounds; the reviewer's verdict drifted from ADVANCED (Rounds 0–2) to STALLED (Rounds 3–4) as the mainline gap list barely changed. Local round contracts were well-scoped and cleanly executed, but they accepted scope reduction the plan forbade, so the reviewer never credited the work as plan-level progress. The methodology produced clean local rounds and the circuit breaker did eventually fire — but it fired after three rounds of diminishing returns.

                      This issue summarizes the methodology-improvement suggestions distilled from a post-loop analysis. All content is sanitized; no project-specific details below.

                      Suggestions

                      Theme 1: Detecting and reacting to external blockers early

                      1. Treat reported "environment-bound" deferrals as a plan-level event, not a per-round footnote.

                      Once the implementer reports an unresolvable external constraint that prevents touching the mainline deliverable, every subsequent round will inherit the same constraint and produce only peripheral work. In the observed session, the implementer surfaced the blocker as early as Round 2 and explicitly listed four resolution options in a Round 3 update request, but the loop continued for two more rounds anyway. When a round summary contains an explicit "Goal Tracker Update Request" presenting decision options that only the user can resolve, the methodology should pause the implement-review loop and route to a human-decision gate. Continuing the loop after a clearly-articulated blocker is a form of busy-waiting that the circuit breaker eventually catches, but a dedicated escalation gate would catch it sooner and more cleanly.

                      2. Distinguish "implementer cannot proceed" from "implementer chose to defer."

                      Round summaries listed the same set of unimplemented tasks as "deferred" for environment reasons across multiple rounds. The reviewer flagged this each time as "unjustified deferrals" because, per the plan boundaries, no deferrals were valid. Both sides were correct under their own framing, but the loop had no way to converge. The methodology should require the round summary to label each deferred item with one of three causes (scope choice / dependency on earlier in-loop work / external blocker) and require the reviewer to acknowledge that taxonomy. External-blocker items should automatically trigger the escalation gate from suggestion #1 rather than recurring as "unjustified deferral" findings.

                      Theme 2: Plan rigidity vs. execution reality

                      3. A plan whose lower bound equals its upper bound is brittle under any unexpected friction.

                      The plan collapsed the acceptable outcome range to a single point: full feature parity, no partial credit. Any friction (environment, missing tools, scope discovery) renders the entire loop incapable of producing a "complete" result, and the reviewer is structurally forced to return "incomplete" every time. Every review across all five rounds reported zero acceptance criteria fully addressed (or a small partial), even though substantive structural work was landed. This makes the reviewer's per-round verdict almost non-informative — it cannot distinguish "great round" from "bad round" because both report the same overall progress percentage. Plans should declare both a target scope and a minimum acceptable scope, with an explicit decision rule for what to do when execution can hit the minimum but not the target. The reviewer can then meaningfully grade rounds against the minimum boundary while still pointing toward the target.

                      4. Validate that the execution environment supports the plan before the loop starts.

                      The plan presumed a working hardware-and-software stack appropriate to the target task. The actual execution environment was missing key prerequisites, and this was discovered only after the first round of work was underway. Rounds 2 onward were entirely shaped by working around the environment instead of executing the plan. The plan-generation phase should produce an explicit "environment prerequisites" checklist, and the loop should run an environment-probe pre-flight (executing minimal smoke checks for each prerequisite) before Round 0. If a prerequisite fails the probe, the plan is revised or the user is asked before any implement-review cycles run.

                      Theme 3: Review effectiveness and stagnation detection

                      5. Reviews were specific and actionable, but the stagnation signal was buried.

                      The reviewer's findings were consistently concrete: precise citations, exact failure modes, and clear "next steps" lists. However, across rounds, large portions of the "required next implementation plan" text were near-verbatim copies of prior rounds. A reader skimming any single review would not realize that the mainline gap list had been almost identical for three consecutive rounds. The signal that the loop was circling lived in the diff between rounds, not in any single round. The reviewer prompt should include a directive to compute a "delta from previous round" summary — explicitly listing which mainline gaps moved, which closed, and which are unchanged. A rising "unchanged mainline gaps" count across rounds is a strong stagnation indicator that should feed into a softer circuit breaker (warning at N=2, hard stop at N=3) rather than waiting for the hard stagnation breaker to fire.

                      6. Reviews caught real issues with low false-positive rate, but the issues they caught shrank in importance over time.

                      The severity profile shifted across the five rounds: early reviews concerned the structural contract; later reviews concerned narrowing literal interpretations of one acceptance criterion and re-flagging cleanup hygiene. The methodology spent its last two rounds polishing the edges of a single acceptance criterion while the central acceptance criteria made no progress. The review machinery functioned, but it ran out of high-value work to do. Reviews should be required to bucket findings as mainline-vs-peripheral and to flag when, for two rounds in a row, all closed findings are peripheral. That signal — "we are closing only peripheral issues" — is a clean stagnation criterion that complements the unchanged-mainline-gaps signal.

                      Theme 4: Communication and round contracts

                      7. Round contracts narrowed scope appropriately, but the narrowing was not reconciled with the plan boundaries.

                      Each round contract correctly identified a tractable local objective. Local execution against those contracts was clean. But the contracts implicitly accepted scope reduction that the plan explicitly forbade, and the reviewer correctly refused to recognize that local progress as plan progress. Implementer and reviewer were operating on different success contracts; both were right in their own frame. A round contract should either (a) be a strict subset of the plan-level deliverables and the reviewer is told to grade only that subset, or (b) explicitly request a plan amendment if the round will not advance the plan. Mixing local-scope contracts with plan-scope reviews is what produced the verdict mismatch.

                      8. Per-round summaries were thorough, but they normalized non-progress.

                      The implementer's summaries grew progressively better at justifying why the central block was not advanced each round. By the final round, the deferral justification was a polished, internally-consistent paragraph that referenced earlier rounds and a pending user decision. This is good craftsmanship in writing, but it had the effect of making non-progress feel like a stable, defensible state and partially masked the underlying problem. Round summaries should include a quantitative "tasks closed this round" and "tasks closed cumulatively / total tasks" header at the top. Once that ratio plateaus across two rounds, the loop should trigger a planning check-in rather than a sixth round.

                      What worked well

                      • Clean local rounds. Each round was contained, the contract was followed, the validations described were what was actually performed, and the diff between rounds was small and focused.
                      • Honest blocker reporting. The implementer surfaced the external blocker as soon as it was confirmed and proposed concrete resolution paths rather than hiding it.
                      • Specific, citation-rich reviews. The reviewer consistently pointed to exact citations and the precise mismatch between expected and observed behavior, which made each round's "next action" unambiguous.
                      • Circuit breaker eventually fired. The stagnation breaker did catch the loop and force an exit, preventing an unbounded sequence of low-value rounds. The improvement opportunity is making that safety net trigger sooner and more gracefully.

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        No labels
                        No labels

                        Type

                        No type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' RLCR: detect external blockers, allow min/max plan scope, surface stagnation via delta reviews · Issue #180 · PolyArch/humanize · GitHub
                          Skip to content

                          RLCR: detect external blockers, allow min/max plan scope, surface stagnation via delta reviews #180

                          Description

                          @donglinz

                          Context

                          During an RLCR session, the loop ran five rounds (Round 0 → Round 4) before the stagnation circuit breaker forced exit. The plan declared an all-or-nothing scope (lower bound = upper bound = full target deliverable). An external execution-environment blocker appeared from Round 2 onward and was explicitly surfaced with resolution options in the Round 3 summary's update request. The loop continued anyway for two more rounds; the reviewer's verdict drifted from ADVANCED (Rounds 0–2) to STALLED (Rounds 3–4) as the mainline gap list barely changed. Local round contracts were well-scoped and cleanly executed, but they accepted scope reduction the plan forbade, so the reviewer never credited the work as plan-level progress. The methodology produced clean local rounds and the circuit breaker did eventually fire — but it fired after three rounds of diminishing returns.

                          This issue summarizes the methodology-improvement suggestions distilled from a post-loop analysis. All content is sanitized; no project-specific details below.

                          Suggestions

                          Theme 1: Detecting and reacting to external blockers early

                          1. Treat reported "environment-bound" deferrals as a plan-level event, not a per-round footnote.

                          Once the implementer reports an unresolvable external constraint that prevents touching the mainline deliverable, every subsequent round will inherit the same constraint and produce only peripheral work. In the observed session, the implementer surfaced the blocker as early as Round 2 and explicitly listed four resolution options in a Round 3 update request, but the loop continued for two more rounds anyway. When a round summary contains an explicit "Goal Tracker Update Request" presenting decision options that only the user can resolve, the methodology should pause the implement-review loop and route to a human-decision gate. Continuing the loop after a clearly-articulated blocker is a form of busy-waiting that the circuit breaker eventually catches, but a dedicated escalation gate would catch it sooner and more cleanly.

                          2. Distinguish "implementer cannot proceed" from "implementer chose to defer."

                          Round summaries listed the same set of unimplemented tasks as "deferred" for environment reasons across multiple rounds. The reviewer flagged this each time as "unjustified deferrals" because, per the plan boundaries, no deferrals were valid. Both sides were correct under their own framing, but the loop had no way to converge. The methodology should require the round summary to label each deferred item with one of three causes (scope choice / dependency on earlier in-loop work / external blocker) and require the reviewer to acknowledge that taxonomy. External-blocker items should automatically trigger the escalation gate from suggestion #1 rather than recurring as "unjustified deferral" findings.

                          Theme 2: Plan rigidity vs. execution reality

                          3. A plan whose lower bound equals its upper bound is brittle under any unexpected friction.

                          The plan collapsed the acceptable outcome range to a single point: full feature parity, no partial credit. Any friction (environment, missing tools, scope discovery) renders the entire loop incapable of producing a "complete" result, and the reviewer is structurally forced to return "incomplete" every time. Every review across all five rounds reported zero acceptance criteria fully addressed (or a small partial), even though substantive structural work was landed. This makes the reviewer's per-round verdict almost non-informative — it cannot distinguish "great round" from "bad round" because both report the same overall progress percentage. Plans should declare both a target scope and a minimum acceptable scope, with an explicit decision rule for what to do when execution can hit the minimum but not the target. The reviewer can then meaningfully grade rounds against the minimum boundary while still pointing toward the target.

                          4. Validate that the execution environment supports the plan before the loop starts.

                          The plan presumed a working hardware-and-software stack appropriate to the target task. The actual execution environment was missing key prerequisites, and this was discovered only after the first round of work was underway. Rounds 2 onward were entirely shaped by working around the environment instead of executing the plan. The plan-generation phase should produce an explicit "environment prerequisites" checklist, and the loop should run an environment-probe pre-flight (executing minimal smoke checks for each prerequisite) before Round 0. If a prerequisite fails the probe, the plan is revised or the user is asked before any implement-review cycles run.

                          Theme 3: Review effectiveness and stagnation detection

                          5. Reviews were specific and actionable, but the stagnation signal was buried.

                          The reviewer's findings were consistently concrete: precise citations, exact failure modes, and clear "next steps" lists. However, across rounds, large portions of the "required next implementation plan" text were near-verbatim copies of prior rounds. A reader skimming any single review would not realize that the mainline gap list had been almost identical for three consecutive rounds. The signal that the loop was circling lived in the diff between rounds, not in any single round. The reviewer prompt should include a directive to compute a "delta from previous round" summary — explicitly listing which mainline gaps moved, which closed, and which are unchanged. A rising "unchanged mainline gaps" count across rounds is a strong stagnation indicator that should feed into a softer circuit breaker (warning at N=2, hard stop at N=3) rather than waiting for the hard stagnation breaker to fire.

                          6. Reviews caught real issues with low false-positive rate, but the issues they caught shrank in importance over time.

                          The severity profile shifted across the five rounds: early reviews concerned the structural contract; later reviews concerned narrowing literal interpretations of one acceptance criterion and re-flagging cleanup hygiene. The methodology spent its last two rounds polishing the edges of a single acceptance criterion while the central acceptance criteria made no progress. The review machinery functioned, but it ran out of high-value work to do. Reviews should be required to bucket findings as mainline-vs-peripheral and to flag when, for two rounds in a row, all closed findings are peripheral. That signal — "we are closing only peripheral issues" — is a clean stagnation criterion that complements the unchanged-mainline-gaps signal.

                          Theme 4: Communication and round contracts

                          7. Round contracts narrowed scope appropriately, but the narrowing was not reconciled with the plan boundaries.

                          Each round contract correctly identified a tractable local objective. Local execution against those contracts was clean. But the contracts implicitly accepted scope reduction that the plan explicitly forbade, and the reviewer correctly refused to recognize that local progress as plan progress. Implementer and reviewer were operating on different success contracts; both were right in their own frame. A round contract should either (a) be a strict subset of the plan-level deliverables and the reviewer is told to grade only that subset, or (b) explicitly request a plan amendment if the round will not advance the plan. Mixing local-scope contracts with plan-scope reviews is what produced the verdict mismatch.

                          8. Per-round summaries were thorough, but they normalized non-progress.

                          The implementer's summaries grew progressively better at justifying why the central block was not advanced each round. By the final round, the deferral justification was a polished, internally-consistent paragraph that referenced earlier rounds and a pending user decision. This is good craftsmanship in writing, but it had the effect of making non-progress feel like a stable, defensible state and partially masked the underlying problem. Round summaries should include a quantitative "tasks closed this round" and "tasks closed cumulatively / total tasks" header at the top. Once that ratio plateaus across two rounds, the loop should trigger a planning check-in rather than a sixth round.

                          What worked well

                          • Clean local rounds. Each round was contained, the contract was followed, the validations described were what was actually performed, and the diff between rounds was small and focused.
                          • Honest blocker reporting. The implementer surfaced the external blocker as soon as it was confirmed and proposed concrete resolution paths rather than hiding it.
                          • Specific, citation-rich reviews. The reviewer consistently pointed to exact citations and the precise mismatch between expected and observed behavior, which made each round's "next action" unambiguous.
                          • Circuit breaker eventually fired. The stagnation breaker did catch the loop and force an exit, preventing an unbounded sequence of low-value rounds. The improvement opportunity is making that safety net trigger sooner and more gracefully.

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            No labels
                            No labels

                            Type

                            No type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); RLCR: detect external blockers, allow min/max plan scope, surface stagnation via delta reviews · Issue #180 · PolyArch/humanize · GitHub
                              Skip to content

                              RLCR: detect external blockers, allow min/max plan scope, surface stagnation via delta reviews #180

                              Description

                              @donglinz

                              Context

                              During an RLCR session, the loop ran five rounds (Round 0 → Round 4) before the stagnation circuit breaker forced exit. The plan declared an all-or-nothing scope (lower bound = upper bound = full target deliverable). An external execution-environment blocker appeared from Round 2 onward and was explicitly surfaced with resolution options in the Round 3 summary's update request. The loop continued anyway for two more rounds; the reviewer's verdict drifted from ADVANCED (Rounds 0–2) to STALLED (Rounds 3–4) as the mainline gap list barely changed. Local round contracts were well-scoped and cleanly executed, but they accepted scope reduction the plan forbade, so the reviewer never credited the work as plan-level progress. The methodology produced clean local rounds and the circuit breaker did eventually fire — but it fired after three rounds of diminishing returns.

                              This issue summarizes the methodology-improvement suggestions distilled from a post-loop analysis. All content is sanitized; no project-specific details below.

                              Suggestions

                              Theme 1: Detecting and reacting to external blockers early

                              1. Treat reported "environment-bound" deferrals as a plan-level event, not a per-round footnote.

                              Once the implementer reports an unresolvable external constraint that prevents touching the mainline deliverable, every subsequent round will inherit the same constraint and produce only peripheral work. In the observed session, the implementer surfaced the blocker as early as Round 2 and explicitly listed four resolution options in a Round 3 update request, but the loop continued for two more rounds anyway. When a round summary contains an explicit "Goal Tracker Update Request" presenting decision options that only the user can resolve, the methodology should pause the implement-review loop and route to a human-decision gate. Continuing the loop after a clearly-articulated blocker is a form of busy-waiting that the circuit breaker eventually catches, but a dedicated escalation gate would catch it sooner and more cleanly.

                              2. Distinguish "implementer cannot proceed" from "implementer chose to defer."

                              Round summaries listed the same set of unimplemented tasks as "deferred" for environment reasons across multiple rounds. The reviewer flagged this each time as "unjustified deferrals" because, per the plan boundaries, no deferrals were valid. Both sides were correct under their own framing, but the loop had no way to converge. The methodology should require the round summary to label each deferred item with one of three causes (scope choice / dependency on earlier in-loop work / external blocker) and require the reviewer to acknowledge that taxonomy. External-blocker items should automatically trigger the escalation gate from suggestion #1 rather than recurring as "unjustified deferral" findings.

                              Theme 2: Plan rigidity vs. execution reality

                              3. A plan whose lower bound equals its upper bound is brittle under any unexpected friction.

                              The plan collapsed the acceptable outcome range to a single point: full feature parity, no partial credit. Any friction (environment, missing tools, scope discovery) renders the entire loop incapable of producing a "complete" result, and the reviewer is structurally forced to return "incomplete" every time. Every review across all five rounds reported zero acceptance criteria fully addressed (or a small partial), even though substantive structural work was landed. This makes the reviewer's per-round verdict almost non-informative — it cannot distinguish "great round" from "bad round" because both report the same overall progress percentage. Plans should declare both a target scope and a minimum acceptable scope, with an explicit decision rule for what to do when execution can hit the minimum but not the target. The reviewer can then meaningfully grade rounds against the minimum boundary while still pointing toward the target.

                              4. Validate that the execution environment supports the plan before the loop starts.

                              The plan presumed a working hardware-and-software stack appropriate to the target task. The actual execution environment was missing key prerequisites, and this was discovered only after the first round of work was underway. Rounds 2 onward were entirely shaped by working around the environment instead of executing the plan. The plan-generation phase should produce an explicit "environment prerequisites" checklist, and the loop should run an environment-probe pre-flight (executing minimal smoke checks for each prerequisite) before Round 0. If a prerequisite fails the probe, the plan is revised or the user is asked before any implement-review cycles run.

                              Theme 3: Review effectiveness and stagnation detection

                              5. Reviews were specific and actionable, but the stagnation signal was buried.

                              The reviewer's findings were consistently concrete: precise citations, exact failure modes, and clear "next steps" lists. However, across rounds, large portions of the "required next implementation plan" text were near-verbatim copies of prior rounds. A reader skimming any single review would not realize that the mainline gap list had been almost identical for three consecutive rounds. The signal that the loop was circling lived in the diff between rounds, not in any single round. The reviewer prompt should include a directive to compute a "delta from previous round" summary — explicitly listing which mainline gaps moved, which closed, and which are unchanged. A rising "unchanged mainline gaps" count across rounds is a strong stagnation indicator that should feed into a softer circuit breaker (warning at N=2, hard stop at N=3) rather than waiting for the hard stagnation breaker to fire.

                              6. Reviews caught real issues with low false-positive rate, but the issues they caught shrank in importance over time.

                              The severity profile shifted across the five rounds: early reviews concerned the structural contract; later reviews concerned narrowing literal interpretations of one acceptance criterion and re-flagging cleanup hygiene. The methodology spent its last two rounds polishing the edges of a single acceptance criterion while the central acceptance criteria made no progress. The review machinery functioned, but it ran out of high-value work to do. Reviews should be required to bucket findings as mainline-vs-peripheral and to flag when, for two rounds in a row, all closed findings are peripheral. That signal — "we are closing only peripheral issues" — is a clean stagnation criterion that complements the unchanged-mainline-gaps signal.

                              Theme 4: Communication and round contracts

                              7. Round contracts narrowed scope appropriately, but the narrowing was not reconciled with the plan boundaries.

                              Each round contract correctly identified a tractable local objective. Local execution against those contracts was clean. But the contracts implicitly accepted scope reduction that the plan explicitly forbade, and the reviewer correctly refused to recognize that local progress as plan progress. Implementer and reviewer were operating on different success contracts; both were right in their own frame. A round contract should either (a) be a strict subset of the plan-level deliverables and the reviewer is told to grade only that subset, or (b) explicitly request a plan amendment if the round will not advance the plan. Mixing local-scope contracts with plan-scope reviews is what produced the verdict mismatch.

                              8. Per-round summaries were thorough, but they normalized non-progress.

                              The implementer's summaries grew progressively better at justifying why the central block was not advanced each round. By the final round, the deferral justification was a polished, internally-consistent paragraph that referenced earlier rounds and a pending user decision. This is good craftsmanship in writing, but it had the effect of making non-progress feel like a stable, defensible state and partially masked the underlying problem. Round summaries should include a quantitative "tasks closed this round" and "tasks closed cumulatively / total tasks" header at the top. Once that ratio plateaus across two rounds, the loop should trigger a planning check-in rather than a sixth round.

                              What worked well

                              • Clean local rounds. Each round was contained, the contract was followed, the validations described were what was actually performed, and the diff between rounds was small and focused.
                              • Honest blocker reporting. The implementer surfaced the external blocker as soon as it was confirmed and proposed concrete resolution paths rather than hiding it.
                              • Specific, citation-rich reviews. The reviewer consistently pointed to exact citations and the precise mismatch between expected and observed behavior, which made each round's "next action" unambiguous.
                              • Circuit breaker eventually fired. The stagnation breaker did catch the loop and force an exit, preventing an unbounded sequence of low-value rounds. The improvement opportunity is making that safety net trigger sooner and more gracefully.

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                No labels
                                No labels

                                Type

                                No type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions