Skip to content

Adaptive behavior v1: learn from feedback and prove one useful improvement #117

Description

@JRichlen

Build an agent that improves how it helps a person over time. Start with review: adapt depth, domain lenses and output to the user's goal. Extend to learning and coding after demonstrating benefit.

Lifecycle: understand → select behaviors → act → collect feedback → compare with baseline → approve or reject → monitor and roll back. Preferences remain scoped, inspectable, correctable and deletable.

Now

Control mechanisms have been exercised; user benefit and judge calibration remain unproven. Runtime expansion is paused. Detailed lab evidence stays private.

Next: one useful experiment

  • Freeze a small corpus around “issues contain too much context.” Set thresholds for less editing and easier comprehension without losing required facts or actions.
  • Compare baseline and one concise-review candidate on held-out tasks. Retain outputs, deterministic checks, blinded human labels and total usage. Calibrate an independent frontier judge against human labels; external grading requires separately authorized inputs and spending limits.
  • Report improvement, regression or inconclusive evidence. Only a supported result advances to an approved, scoped install with demonstrated rollback.

Fit and boundaries

Reuse agent-compiler for composition, recurrence-detector for candidate signals, and existing verification, packaging and curation patterns. Evaluate the framework in #115 and calibration fix in #114 before adding another harness. Both remain unqualified for behavioral benefit.

Actor qualification follows the separate upstream plan: pinned local Qwen 27B, direct 32K/one-sequence baseline, then measured context lanes; dispatcher continuity and Aperture follow their own acceptance gates. Capacity is unqualified, AgentWorld is parked, and the actor route has no external fallback. The request contract stays in #101/#107; this plugin consumes that capability.

Later

Learning/coding transfer and opt-in curation automation. First run should preview a natural-language schedule, scope and limits, then deploy only after approval; repeated setup must not duplicate it, and pause/removal must work. Trigger semantics (#84), broader testing architecture (#89), and the optional cost benchmark (#116) are deferred references, not additional active workstreams.

Tracking: update this checklist at meaningful checkpoints. Preserve historical discussions (#85, #102); keep protocols in artifacts and avoid child-ticket sprawl.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions