Build an agent that improves how it helps a person over time. Start with review: adapt depth, domain lenses and output to the user's goal. Extend to learning and coding after demonstrating benefit.
Lifecycle: understand → select behaviors → act → collect feedback → compare with baseline → approve or reject → monitor and roll back. Preferences remain scoped, inspectable, correctable and deletable.
Now
Control mechanisms have been exercised; user benefit and judge calibration remain unproven. Runtime expansion is paused. Detailed lab evidence stays private.
Next: one useful experiment
Fit and boundaries
Reuse agent-compiler for composition, recurrence-detector for candidate signals, and existing verification, packaging and curation patterns. Evaluate the framework in #115 and calibration fix in #114 before adding another harness. Both remain unqualified for behavioral benefit.
Actor qualification follows the separate upstream plan: pinned local Qwen 27B, direct 32K/one-sequence baseline, then measured context lanes; dispatcher continuity and Aperture follow their own acceptance gates. Capacity is unqualified, AgentWorld is parked, and the actor route has no external fallback. The request contract stays in #101/#107; this plugin consumes that capability.
Later
Learning/coding transfer and opt-in curation automation. First run should preview a natural-language schedule, scope and limits, then deploy only after approval; repeated setup must not duplicate it, and pause/removal must work. Trigger semantics (#84), broader testing architecture (#89), and the optional cost benchmark (#116) are deferred references, not additional active workstreams.
Tracking: update this checklist at meaningful checkpoints. Preserve historical discussions (#85, #102); keep protocols in artifacts and avoid child-ticket sprawl.
Build an agent that improves how it helps a person over time. Start with review: adapt depth, domain lenses and output to the user's goal. Extend to learning and coding after demonstrating benefit.
Lifecycle: understand → select behaviors → act → collect feedback → compare with baseline → approve or reject → monitor and roll back. Preferences remain scoped, inspectable, correctable and deletable.
Now
Control mechanisms have been exercised; user benefit and judge calibration remain unproven. Runtime expansion is paused. Detailed lab evidence stays private.
Next: one useful experiment
Fit and boundaries
Reuse
agent-compilerfor composition,recurrence-detectorfor candidate signals, and existing verification, packaging and curation patterns. Evaluate the framework in #115 and calibration fix in #114 before adding another harness. Both remain unqualified for behavioral benefit.Actor qualification follows the separate upstream plan: pinned local Qwen 27B, direct 32K/one-sequence baseline, then measured context lanes; dispatcher continuity and Aperture follow their own acceptance gates. Capacity is unqualified, AgentWorld is parked, and the actor route has no external fallback. The request contract stays in #101/#107; this plugin consumes that capability.
Later
Learning/coding transfer and opt-in curation automation. First run should preview a natural-language schedule, scope and limits, then deploy only after approval; repeated setup must not duplicate it, and pause/removal must work. Trigger semantics (#84), broader testing architecture (#89), and the optional cost benchmark (#116) are deferred references, not additional active workstreams.
Tracking: update this checklist at meaningful checkpoints. Preserve historical discussions (#85, #102); keep protocols in artifacts and avoid child-ticket sprawl.