Skip to content

Qualitative consequences, positive framing, and consistent length targets - #477

Closed
neoneye wants to merge 1 commit into
mainfrom
fix/identify-qualitative-consequences
Closed

Qualitative consequences, positive framing, and consistent length targets#477
neoneye wants to merge 1 commit into
mainfrom
fix/identify-qualitative-consequences

Conversation

@neoneye

Copy link
Copy Markdown
Member

Summary

  • Qualitative consequences: Replace all number-evidence constraints with "Exact numbers will be determined further downstream — here the goal is to articulate the cause-effect relationship clearly." Section 5: "NO specific numbers, percentages, or monetary amounts — describe effects qualitatively."
  • Positive framing: Replace "Do NOT include 'Controls ... vs.'" (same as prior PRs)
  • Consistent length targets: "2-3 sentences" in all three sources (field desc ×2, system prompt). review_lever "One sentence (20-40 words)" matches system prompt section 6.

PR #475's "Never invent" constraint made fabrication 4x worse — models interpreted it as a permission floor and derived numbers by arithmetic from plan figures. This PR removes the number discussion entirely.

Test plan

🤖 Generated with Claude Code

…gets
Replace number-focused constraints with qualitative framing: "Exact
numbers will be determined further downstream — here the goal is to
articulate the cause-effect relationship clearly." This avoids the
permission-floor effect where models derive numbers by arithmetic
from real plan figures (4x HKD regression in PR #475).
Section 5 prohibition changed from "NO fabricated statistics without
evidence" to "NO specific numbers, percentages, or monetary amounts —
describe effects qualitatively."
All length targets consistent: consequences "2-3 sentences" in both
field descriptions and system prompt. review_lever "one sentence
(20-40 words)" in both field description and system prompt section 6.
Replace negative prohibition ("Do NOT include 'Controls ... vs.'")
with positive framing.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@neoneye

Copy link
Copy Markdown
MemberAuthor

Self-improve iteration — analysis 72

Verdict: CONDITIONAL

Wins:

  • Fabricated numbers: -48% overall, haiku -70% (20→6)
  • Review template lock ("All three options...") appears resolved (85%→0%)
  • gpt-oss-20b 5/5 success (gpt-oss-20b parasomnia timeout is infrastructure noise)
  • No content quality regressions on successful plans

Residual:

  • 3 plan timeouts (gpt-oss-20b parasomnia, gpt-5-nano sovereign_identity/silo) — likely infrastructure noise
  • qwen3-30b unchanged at 3 pct claims — "Exact numbers will be determined further downstream" may anchor it
  • 6 remaining haiku pct claims in engineering-domain plans

Conditions: Confirm timeouts are noise (re-run), and consider removing the "downstream" anchor phrase.

@neoneye

Copy link
Copy Markdown
MemberAuthor

Closing — not convincingly better than baseline. Will refine: keep stricter review_lever and shorter consequences, but allow plan-provided numbers instead of banning all numbers outright.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@neoneye