Autonomous Agents - #5

Merged
timothyerwin merged 13 commits into
mainfrom
feat/personal-agents
Aug 30, 2026
Merged

Autonomous Agents#5
timothyerwin merged 13 commits into
mainfrom
feat/personal-agents

Conversation

@timothyerwin

Copy link
Copy Markdown
Contributor

Summary

Adds Personal Agents: long-lived agent identities with a durable objective/subgoal/facts store, a delegation policy that governs what their executive loop may do (allowed tools, consequence classes, filesystem/browser/MCP resource grants, budgets), ask_user suspend/resume for human-in-the-loop interactions, durable cron/interval/one-shot triggers polled by the gateway, and a browser broker that leases exclusive browser control to one run at a time.

New surfaces: memcode personal CLI (create, approve-policy, resources, triggers, answer, status, cockpit), gw_agent/gw_* admin tools, and the gateway's internal personal wake route.

Review fixes included

This branch went through a review pass (/code-review high) before commit; the following were found and fixed:

  • Duplicate-answer replay: personal answer now refuses to re-resume an already-answered interaction, and a resumed run that re-suspends on a new question now marks its prior continuation resolved — closing a window where a retried personal answer call could re-execute real tool side effects.
  • Triggers silently dead after hot-add: the gateway now always registers the personal wake route and always starts the trigger poll loop (it already re-reads config every tick), so a Personal Agent created after boot has its triggers picked up without a gateway restart.
  • Expired browser lease still granted mutation rights: CanMutate/OwnPage now check lease expiry like Authenticate already did.
  • Resource grant ID collisions: IDs now use a monotonic timestamp instead of len(locator), matching the cockpit's existing scheme.
  • Delegation policy narrowing gaps: IsRestriction/NarrowPolicy now also checks filesystem roots, browser origins, MCP tools, budget fields, and quiet hours, not just tools/consequence-classes/concurrency/depth.
  • Silent no-op on bad trigger id: personal triggers pause/resume now errors instead of reporting success when nothing matched.
  • Stale tool description: gw_agent no longer advertises the unsupported "inspect" action.
  • Non-atomic suspension writes: suspension continuation files are now written via atomicfile.WriteFile.
  • Per-tick DB churn: the trigger poll loop now reuses one DB connection per agent across ticks and filters due triggers in SQL instead of opening/closing a connection and Go-filtering every row every 15s.

Testing

  • go build ./... — clean
  • go vet ./... — clean
  • go test ./... — all pass except two pre-existing, unrelated failures in internal/guard caused by a symlink under desktop/node_modules/@electron/notarize breaking go list (not touched by this branch)

… triggers, browser broker)
Implements the Personal Agents feature: long-lived agent identities with
their own durable objective, subgoals, and facts store; a delegation
policy model (allowed tools, consequence classes, filesystem/browser/MCP
grants, budgets) that governs what a Personal Agent's executive loop may
do; ask_user suspend/resume for human-in-the-loop interactions; durable
triggers (cron/interval/one-shot) polled by the gateway to wake an agent;
resource grants (filesystem, mcp, command, channel); and a browser broker
that leases exclusive browser control to one run at a time.
New surfaces: `memcode personal` CLI (create, approve-policy, resources,
triggers, answer, status, cockpit), gw_agent/gw_* admin tools, and the
gateway's internal personal wake route.
Includes fixes from review of this branch:
- personal answer refuses to re-resume an already-answered interaction,
and a resumed run that re-suspends now marks its prior continuation
resolved — closing a window where a retried answer could replay tool
side effects.
- The gateway always registers the personal wake route and always starts
the trigger poll loop, so a Personal Agent added after boot has its
triggers picked up without a restart.
- Browser broker leases now check expiry in CanMutate/OwnPage, matching
Authenticate, so an expired lease can no longer authorize mutation.
- Resource grant IDs use a monotonic timestamp instead of locator length,
avoiding same-length-locator collisions.
- Delegation policy narrowing (IsRestriction/NarrowPolicy) now also
checks filesystem roots, browser origins, MCP tools, budget fields, and
quiet hours, not just tools/consequence-classes/concurrency/depth.
- `personal triggers pause/resume` errors on an unknown trigger id instead
of reporting success silently.
- gw_agent's description no longer advertises the unsupported "inspect"
action.
- Suspension continuation files are written atomically.
- The trigger poll loop reuses one DB connection per agent across ticks
and filters due triggers in SQL, instead of opening/closing a
connection and scanning every trigger row every 15s.
… CI lint
Closes the biggest gap from review: the Personal Agent executive had no way
to reach browser/MCP/shell/filesystem capability — it was a second, weaker
agent runtime sitting beside the real memcode engine instead of commanding
it. Two new executive tools fix that:
- delegate: spawns a policy-scoped worker as a real detached `memcode run`
job (internal/jobs.SpawnWithSpec), using the delegation.go scaffolding
(ExecutionEnvelope, ValidateDelegation, PrepareRunDirectory) that already
existed but was never called from anywhere. The worker gets whatever
toolsets it's granted — browser, MCP, shell, filesystem, skills — same as
any ordinary memcode agent, not the executive's fixed 7 tools. Requested
toolsets/consequences must be a subset of the agent's own approved policy;
MaxDelegationDepth gates whether delegation is allowed at all.
- check_delegate: reads a delegated job's status/result back on a later
wake (RunOnce is one bounded wake, so the result can't arrive in the same
call) and closes out its action-journal entry.
Also makes jobs.SpawnSpec.ToolPolicy a REAL restriction instead of recorded-
only metadata: new hidden --allow-tools/--deny-tools flags on `memcode run`
bind to the same SetToolPolicy enforcement an ordinary gateway-bound agent
gets from its config, and SpawnWithSpec now passes them through argv.
Scope notes, called out explicitly rather than implied as done:
- "browser" toolset reuses the existing ephemeral Chrome session (the same
one --chrome already drives) — NOT the user's own already-running Chrome.
Attaching to a persistent/remote Chrome needs the broker/RemoteConfig
wiring in internal/browser, which is declared but has zero callers and
zero implementation; that's separate follow-up work, not done here.
- Filesystem/MCP resource grants are not OS-sandboxed to a spawned worker
(memcode has no such sandbox anywhere) — governance here is the same
approval/audit model as the rest of the codebase (tool-level allow/deny +
the action journal), not a hard jail.
Also fixes two CI lint failures on the open PR:
- internal/gateway/state/state.go: removed formatTime, an unused helper
(staticcheck U1000).
- internal/personal/store.go: decapitalized a Go error string (ST1005).
Corrects course on the browser gap flagged after the last review: existing-
Chrome access is core to Personal Agents, not deferrable follow-up. A
Personal Agent's whole premise — acting as the user across their real
accounts (Gmail, LinkedIn, an ATS, internal dashboards) — requires the
user's actual signed-in session, not a fresh ephemeral profile with no
cookies. Per docs/design/personal-agents.md's already-specified "browser
broker trust boundary": Personal Agents connect to the user's existing
Chrome through a gateway-owned broker and a permission-protected local
socket, fail closed on any connection/auth problem, and never silently fall
back to another profile.
- internal/browser/broker/{server,client}.go: exposes the (already-reviewed)
in-process Broker over a Unix socket, so a delegated worker — a SEPARATE
OS process spawned via jobs.SpawnWithSpec — can Acquire/Release/OwnPage/
CanMutate against the SAME broker the gateway owns, not a fresh one.
Socket is 0600 inside gwconfig.Dir() (0700). Round-trip tested, including
that a concurrent Acquire is correctly rejected and a release unblocks the
next one.
- internal/gateway/server: the gateway now owns one *broker.Broker for its
whole lifetime and serves it on that socket unconditionally at startup
(cheap — a local listener), so it's there the moment a delegate call needs
it, no restart required after `memcode personal browser setup`.
- internal/personal: browserModeFor's "browser" toolset now defaults to
BrowserExistingChrome, not ephemeral — "browser:ephemeral" is the explicit
opt-down for tasks that genuinely want a fresh, logged-out profile. Before
spawning anything, the delegate tool verifies the broker is reachable and
fails the tool call closed (a normal tool_result error, not a spawned job)
if it isn't — verified by test, alongside the success path asserting the
spawned job's SpawnSpec actually carries BrowserMode=existing_chrome.
- internal/jobs: SpawnWithSpec passes existing_chrome through as
--browser-session (+ --browser-agent/--browser-run for the lease identity)
instead of --chrome.
- cmd/run.go: --browser-session existing_chrome acquires the broker lease
itself (via the client) and, only on success, wires an ad-hoc
chrome-devtools-mcp --autoConnect MCP server into the session
(Session.SetExtraMCPServers, new) — the SAME chrome-devtools-mcp package
this browser experience is built on (see internal/browser/remote.go). On
any failure it returns an error and exits; it never calls
SetBrowserEnabled(true) (ephemeral) as a fallback.
- cmd/personal_browser.go: `memcode personal browser setup` checks npx and
the broker socket, then attempts a REAL bounded (10s) chrome-devtools-mcp
connection and reports the tool count or the actual connect errors — it
does not pretend to click Chrome's own "Allow" dialog, which only the user
can do.
- Also fixes a real bug found while testing this: ValidateDelegation treated
an empty parent.AllowedTools as "allows no tools by name", the opposite of
the len(...)>0-means-restricted convention Executive.allowedTools already
uses — every delegate call with a name-unrestricted policy was rejected.
Scope note, stated rather than implied: the literal Chrome-side consent flow
(the "Allow" dialog, actually clicking it) cannot be exercised or verified
in this environment — no live Chrome, no human to click it. Everything
up to that point (broker lifecycle, lease exclusivity, fail-closed wiring,
the actual chrome-devtools-mcp CLI invocation, verified against the real
published package) is built and tested; the dialog itself needs the user's
own machine.
…setup via the cockpit
Three separate complaints from review, addressed at the right layer each:
1. "It's just a file ref" — `resources add` required a redundant --mode read
(already the default) and a mandatory <type> argument even for the
overwhelming common case. A bare path is now enough:
memcode personal resources add jobhunt ~/resume.md
Non-filesystem grants (mcp/command/channel) still need the type spelled
out, since there's nothing to infer it from. Also fixes the real bug this
surfaced: CanonicalFilesystemGrant rejected any non-directory path, even
though PathWithinGrant already handles a single-file grant correctly
(rel == "." on an exact match) — you could not actually grant a single
file like a resume before this, only whole directories. `personal create`
also gained a repeatable --grant flag to fold "create it, then give it
something to read" into one command.
2. "Not this CLI-command-argument garbage" — Personal Agent state (objective,
policy, resources) lived ONLY in personal.db, reachable exclusively
through bespoke commands, unlike every other piece of memcode config
(gateway.yaml, .mcp.json, CLAUDE.md) which is a plain file. WriteConfigMirror
now regenerates objective.md/policy.yaml/resources.yaml in the agent's home
on every mutation — real, readable, diffable files. The run journal
(actions/triggers/interactions) deliberately stays in SQLite: it needs
atomic claim/complete semantics under concurrent access (gateway wake loop,
CLI, and cockpit can all touch the same agent) that flat files don't get
for free — a correctness reason, not a habit.
3. "It should gather requirements, do HITL, pre-approve permissions, set up
the runtime — like our other agents already do via telegram/model setup"
— `memcode personal` (no args) is ALREADY the interactive cockpit (like
`memcode admin`) with the pa_* tools wired for exactly this. The actual gap
was its system prompt: personalAdminDoctrine listed tool mechanics with no
instruction to gather-then-propose-then-approve. Rewrote it to require a
real walkthrough for first-time creation — reason about what resources/
toolsets/consequences the objective needs (asking the user rather than
guessing), present the concrete plan in plain language BEFORE touching
anything, then apply everything at once on approval, including wake
cadence (trigger) — an agent with no trigger and no plan to ever be woken
is dead on arrival, so cadence is not an optional afterthought either.
A prior pass at this (adding a NEW deterministic `memcode personal setup`
CLI wizard, mirroring `gateway setup`'s bufio-prompt mechanism) was wrong
and reverted before commit — that reproduces the exact CLI-surface-
sprawl complaint instead of fixing it. The cockpit conversation IS the
wizard; it needed a better prompt, not a new command.
Root cause of the whole CLI-surface complaint: the interactive cockpit
(`memcode personal`, no args) is supposed to be the entire interface, but
its tool set had NO way to create a new Personal Agent. pa_objective only
supported show/set, and set required the objective to already exist —
creation was only reachable through the CLI's `personal create`. That's why
every explanation kept surfacing a CLI command: until now, it was the only
path that actually worked.
- internal/agent/tools/personal.go: new pa_create tool (agent name +
objective). Documented as the required first call for a brand-new agent.
- cmd/personal_cockpit.go: paCreate mirrors personalCreate's steps (register
in gateway.yaml, open the store, create the objective) and is dispatched
BEFORE the paStore existence gate, since the agent doesn't exist yet.
pa_resource's grant action now defaults type=filesystem/mode=read when
omitted too (same inference the CLI's `resources add` already got), and
pa_resource/pa_policy's grant/revoke/stage/approve paths now call
WriteConfigMirror — that was wired into the CLI paths last commit but not
into the cockpit paths, which are the ones actually used.
- cmd/personal.go: every `memcode personal` subcommand is now Hidden — they
remain fully callable for scripts, but `memcode personal --help` no longer
reads like a CLI to memorize. The interface is `memcode personal` with no
args, then plain conversation.
- internal/doctrine/prompts.go: the cockpit's own system prompt now says
outright never to tell the user to run a CLI command — it does the work
itself, right there — and that pa_create, not pa_objective, is the first
call for a new agent.
Verified end to end: pa_create then a bare-path pa_resource grant (no type,
no mode) both succeed through personalExecute exactly as the cockpit would
call them, no CLI involved.
Hiding the subcommands last commit wasn't enough — they still existed as a
parallel, scriptable way to do everything, which is real duplicate logic
(and already caused drift: a fix landed in one path and not the other,
twice, in this same review). Deleted them outright:
cmd/personal_resources.go, personal_policy.go, personal_triggers.go,
personal_status.go, personal_browser.go, personal_test.go — all removed.
cmd/personal.go cut down to just the cockpit launcher (~90 lines from
~370): `memcode personal` opens the interactive session, full stop.
Two operations the CLI covered had no pa_* equivalent yet, so they'd have
been silently lost:
- pa_doctor: the health check (home layout, objective, approved policy,
generated workspace, sandbox availability, trigger/pending-interaction
counts) that `personal doctor` used to run.
- pa_browser_setup: the existing-Chrome prerequisite check + real, bounded
chrome-devtools-mcp connection attempt that `personal browser setup` used
to run. Same logic, now a tool call instead of a command — it still
doesn't (can't) click Chrome's own consent dialog for the user.
Both dispatch through personalExecute exactly like every other pa_* tool.
internal/agent/tools/personal.go: added PaDoctor and PaBrowserSetup to the
registry.
cmd/personal_cmd_test.go: rewritten from scratch to call personalExecute
directly — no CLI args, no cobra Execute(), no rootCmd. This IS the
interface being tested now, so this is what the tests exercise. Also
asserts the config-mirror files (policy.yaml, resources.yaml) actually
reflect approvals/revokes done through the tool path, since that's the one
that matters.
Verified: `memcode personal create foo bar` no longer resolves to anything
— it just opens the cockpit with those words ignored as stray args, because
there is no "create" subcommand left to match. The interface is
conversation with pa_* tools behind it, not CLI argument syntax.
objective.md/policy.yaml/resources.yaml split one agent's config across
three files by table for no real reason — a person has to know to look in
three places for one agent's setup. Consolidated into a single config.yaml
(objective, policies, resources as sections) with proper snake_case field
names. Same authority split as before (this is a mirror; policy hash-
approval and the run journal stay in personal.db, and why, is unchanged) —
just one file to read instead of three.
Consolidation step 1 of folding Personal Agents into the ordinary agent
system. There were three partial suspend/resume designs in the tree:
1. internal/agent/runtime/continuation.go — typed, atomic, tested, and with
ZERO production callers.
2. jobs.Job's InteractionID/WaitingReason/ContinuationVersion/WaitingAt/
ResumedAt — declared, never written by anything.
3. A hand-rolled map[string]any in the personal executive — the only one
actually running, and (until a fix earlier in this branch) the only one
that wasn't crash-safe.
New internal/agent/continuation is the single implementation. It lives in its
own small package rather than in internal/agent/runtime so the executive can
suspend without importing the whole session runtime (and so neither side ends
up depending on the other).
It has to serve two genuinely different callers, which is why the previous
attempt at a shared API didn't fit:
- An interactive session already holds the conversation, so it only needs the
missing assistant/answer pair back.
- An unattended executive keeps NO transcript — it rebuilds context from
durable state each wake — so its continuation must carry the whole
conversation. Hence the optional Messages field.
They also differ on when marking resolved is safe, so building the resume
messages is split from marking:
- ResumeMessages() builds without marking. The executive uses this and marks
only once the resumed run reaches a terminal state, so a transient model
error leaves the answer re-giveable instead of stranding it.
- Resolve() does both, for a caller with a human right there who can simply
ask again.
Also deleted the five never-written Job fields (2 above). StatusWaiting stays:
it is a meaningful status in the job state machine and check_delegate
correctly treats it as non-terminal.
Tests: ported both original round-trip/validation tests, plus new coverage for
the transcript-carrying path and MarkResolved. The executive's suspend/resume
test now asserts through the continuation API instead of a hardcoded filename.
… admin
Consolidation step 2. "Personal Agent" is no longer a kind of agent. There is
one Agent abstraction; autonomy is orthogonal settings on it, managed through
the admin cockpit that already manages agents.
CONFIG. gwconfig.Agent drops Kind and gains four settings:
objective — the durable outcome it works toward
autonomous — whether it may act on that unprompted
browser — ephemeral (default) | existing_chrome
paused — stop unattended wakes without deleting anything
objective and autonomous are deliberately SEPARATE grants, which is the
correction that motivated this design. They answer different questions ("what
is it for" vs "may it act unasked"), and all four combinations are meaningful:
autonomous + objective → unattended objective pursuit (the old Personal Agent)
autonomous, no objective→ scheduled work UNDER GOVERNANCE — this is new, and
closes the long-standing gap where a cron-fired
agent ran unattended with no policy gate, no action
journal, and no way to pause and ask
objective, not autonomous → a goal you work on together; wakes only on demand
neither → an ordinary conversational agent, exactly as before
Autonomy therefore gates GOVERNANCE, not capability, and applies to any run of
the agent whether or not an objective exists.
SCHEDULING. The second cron implementation is gone as a user-facing concept.
Recurring cadence is an ordinary `schedules:` entry: set agent=<name>, leave
deliver_to empty, and gw_schedule routes the wake to the agent itself via the
renamed internal sink (personal → agent). The DB-backed loop now fires only
the agent's OWN self-scheduled next-wakes (schedule_wake, "come back in 45
minutes") — a genuinely different thing from human-authored cadence, and the
only part that needs to be writable from inside a run.
TOOLS. internal/agent/tools/personal.go (pa_*) is deleted. The admin registry
gains gw_policy, gw_grant, gw_wake, gw_inbox, gw_answer, gw_journal, gw_doctor,
gw_browser; gw_agent gains objective/autonomous/browser/pause/resume actions.
Everything dispatches through the existing adminExecute + approval gate, and
gw_schedule's shared BuildSchedule validation now covers agent cadence too, so
the surfaces cannot drift the way the parallel path did.
COCKPIT. `memcode personal` is deleted outright — cmd/personal.go,
cmd/personal_cockpit.go, Session.SetPersonal, personalMode, the personal and
personal_admin doctrines. The Personal setup walkthrough is merged into
adminDoctrine, generalized, and now teaches the two gates explicitly: granting
an objective is not granting autonomy, and the second must be confirmed on its
own.
EXECUTIVE. Executive.Objective is read from configuration rather than the
store, so the objective has ONE source that a human edits and the gateway
hot-reloads. The per-agent config.yaml mirror drops the objective for the same
reason — it lives in gateway.yaml, which is already a readable file.
Tests are rewritten against adminExecute (there is no other surface). New
coverage asserts the orthogonality directly, since a single overloaded switch
is exactly what this change exists to prevent: an objective alone must not
confer autonomy, a non-affirmative value must not grant it by typo, an
autonomous agent with no objective is still governed but refuses to invent
work, and an autonomous agent's schedule defaults to the agent route.
…ard both
Consolidation step 3, finishing the merge.
RENAME. internal/personal -> internal/agent/autonomy. The package was never
about a "personal" species; it is the machinery an agent uses when running
unattended toward an objective. Terminology in comments follows.
SECOND CRON PARSER DELETED. NextDue now understands only the kinds an AGENT
writes for itself from inside a run (one_shot / next_wake — always a single
future instant, e.g. schedule_wake "come back in 45 minutes"). Recurring
cadence a human configures is not a trigger at all: it is an ordinary gateway
schedule delivering to agent:<name>, validated once by gwconfig. With
recurring kinds gone, ClaimDueTrigger's reschedule branch was unreachable, so
firing a wake now simply completes it — and the atomic `last_fired_at IS ?`
claim still keeps two gateway processes from double-firing.
GUARDS (internal/guard/singletons_test.go). The duplication this branch kept
producing was not a one-off: fixes landed on one path and not the other,
repeatedly. Four invariants now fail the build instead:
- TestSingleCronParser — only internal/gateway parses cron. (This one caught a
real leftover while being written: the autonomy package was still importing
robfig/cron, which is what prompted the deletion above.)
- TestSingleSuspensionImplementation — no hand-rolled continuation formats
outside internal/agent/continuation.
- TestNoSecondCockpit — no SetPersonal/personalMode/pa_* anywhere.
- TestNoAgentKind — autonomy is orthogonal settings, never a kind
discriminator; a "kind" field is what made Personal a separate species.
The guards deliberately skip vendored forks, .memcode session snapshots, and
desktop/node_modules (the same symlink that breaks the two pre-existing
go-list guard failures).
Store tests updated: the trigger test now exercises a real self-scheduled wake
and asserts it completes rather than reschedules, and NextDue's test asserts
that interval/cron are REJECTED — accepting them again would rebuild the
second scheduler.
Consolidation step 4. An agent had two unrelated places to put what it knew:
a structured `facts` table only reachable through bespoke tools, and the
memory.md every memcode agent already has and the runtime already injects. One
of those a human can read. Collapse to it.
BLOCKING PREREQUISITE, done first. delegate/check_delegate were using the
facts table as a keyed index — writing "delegation.<job-id>" so a later wake
could find the action to close out. Its own comment admitted why ("facts are
the only durable log delegate can write to without a schema migration"). This
is that migration: actions gain a job_id column plus LinkActionJob /
ActionForJob, and the link now lives on the action it actually describes.
Removing facts before this would have broken check_delegate silently.
note_fact becomes `remember`: one plain sentence appended to memory.md, read
back into every wake's state summary. Append-only and deduplicated, because
every line replays into the model on each future wake — a re-learned fact must
not grow the file forever.
Deleted with it: environment.go (EnvironmentModel/StructuredFact, referenced
nowhere but its own test) and UsableForExternalRepresentation.
The tradeoff is deliberate and documented in memory.go rather than left
implicit. Prose cannot distinguish "you told me this" from "I inferred it from
your resume" from "a website said so", nor mark a claim safe to assert on the
user's behalf, nor mark it stale. That distinction becomes load-bearing the
moment an agent fills in a form or sends a message stating something about the
user. It is cheap to drop TODAY only because the gate was never wired —
nothing read Confirmed to gate anything; the sole caller of
UsableForExternalRepresentation was its own test. When external
representation needs provenance, it should come back as its own
machine-checkable thing in the store alongside policies and the action
journal, NOT by reviving this table.
… as a mode
Consolidation step 5, finishing the merge.
MIGRATION. YAML silently ignores unknown fields, so an existing
`kind: personal` agent would have loaded as an ordinary one — still listed,
apparently fine, and never waking again. For a change that removes an agent's
authority to run on its own, silence is the wrong failure. gwconfig.Agent keeps
a LegacyKind field for the sole purpose of rejecting it, with the one-line fix
in the error:
agent "demo" still uses the removed `kind: personal` setting. Autonomy is now
explicit: replace it with `autonomous: true` (and an `objective:` ...), or
just delete the `kind:` line if it should not. Its home ... is untouched
Verified against the real local config, which has exactly that agent. The
guard's TestNoAgentKind is refined to allow this one mention while still
banning any behavioral branch on it.
DOCS. docs/personal-agents.md -> docs/autonomous-agents.md, rewritten around
the two orthogonal grants and the four combinations they produce, with the
scheduling split (your cadence = an ordinary schedule; the agent's own next
wake = self-scheduled), the fail-closed existing-Chrome rule, and the
memory.md provenance limitation stated rather than glossed.
docs/design/personal-agents.md -> docs/design/autonomous-agents.md keeps the
original design contract but opens with a note on what the implementation
revised and why, including the orthogonality correction. It is a design record;
rewriting history there would lose the reasoning.
README drops "three ways to run it" (there are two) and describes autonomy as a
setting rather than a product tier, leading with the case that actually
motivated it: a scheduled agent that is finally policy-gated, journaled, and
able to stop and ask.
The store file was still named personal.db. It holds state for any agent
running unattended, not a species of agent — rename it, and note in place that
it is opened lazily so an ordinary conversational agent never grows one.
That claim is now a test rather than a comment: TestOrdinaryAgentGetsNoAutonomyStore
creates a plain agent, exercises the admin surface against it, and asserts no
agent.db appears — the governance machinery costs nothing until asked for.
Also covers the remaining orthogonality row: an agent WITH an objective but
WITHOUT autonomy still wakes on demand, reaching the policy gate rather than
being refused for lacking autonomy. Autonomy governs unprompted action, not
whether a human may ask.
All four combinations of (objective, autonomous) now have a test.
@timothyerwintimothyerwin changed the title Personal Agents: durable objectives, delegation policy, triggers, browser brokerAutonomous AgentsAug 30, 2026
@timothyerwin
timothyerwin merged commit 8a2e6b3 into mainAug 30, 2026
2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@timothyerwin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Autonomous Agents - #5

Merged
timothyerwin merged 13 commits into
mainfrom
feat/personal-agents
Aug 30, 2026
Merged

Autonomous Agents#5
timothyerwin merged 13 commits into
mainfrom
feat/personal-agents

Conversation

@timothyerwin

Copy link
Copy Markdown
Contributor

Summary

Adds Personal Agents: long-lived agent identities with a durable objective/subgoal/facts store, a delegation policy that governs what their executive loop may do (allowed tools, consequence classes, filesystem/browser/MCP resource grants, budgets), ask_user suspend/resume for human-in-the-loop interactions, durable cron/interval/one-shot triggers polled by the gateway, and a browser broker that leases exclusive browser control to one run at a time.

New surfaces: memcode personal CLI (create, approve-policy, resources, triggers, answer, status, cockpit), gw_agent/gw_* admin tools, and the gateway's internal personal wake route.

Review fixes included

This branch went through a review pass (/code-review high) before commit; the following were found and fixed:

  • Duplicate-answer replay: personal answer now refuses to re-resume an already-answered interaction, and a resumed run that re-suspends on a new question now marks its prior continuation resolved — closing a window where a retried personal answer call could re-execute real tool side effects.
  • Triggers silently dead after hot-add: the gateway now always registers the personal wake route and always starts the trigger poll loop (it already re-reads config every tick), so a Personal Agent created after boot has its triggers picked up without a gateway restart.
  • Expired browser lease still granted mutation rights: CanMutate/OwnPage now check lease expiry like Authenticate already did.
  • Resource grant ID collisions: IDs now use a monotonic timestamp instead of len(locator), matching the cockpit's existing scheme.
  • Delegation policy narrowing gaps: IsRestriction/NarrowPolicy now also checks filesystem roots, browser origins, MCP tools, budget fields, and quiet hours, not just tools/consequence-classes/concurrency/depth.
  • Silent no-op on bad trigger id: personal triggers pause/resume now errors instead of reporting success when nothing matched.
  • Stale tool description: gw_agent no longer advertises the unsupported "inspect" action.
  • Non-atomic suspension writes: suspension continuation files are now written via atomicfile.WriteFile.
  • Per-tick DB churn: the trigger poll loop now reuses one DB connection per agent across ticks and filters due triggers in SQL instead of opening/closing a connection and Go-filtering every row every 15s.

Testing

  • go build ./... — clean
  • go vet ./... — clean
  • go test ./... — all pass except two pre-existing, unrelated failures in internal/guard caused by a symlink under desktop/node_modules/@electron/notarize breaking go list (not touched by this branch)

… triggers, browser broker)
Implements the Personal Agents feature: long-lived agent identities with
their own durable objective, subgoals, and facts store; a delegation
policy model (allowed tools, consequence classes, filesystem/browser/MCP
grants, budgets) that governs what a Personal Agent's executive loop may
do; ask_user suspend/resume for human-in-the-loop interactions; durable
triggers (cron/interval/one-shot) polled by the gateway to wake an agent;
resource grants (filesystem, mcp, command, channel); and a browser broker
that leases exclusive browser control to one run at a time.
New surfaces: `memcode personal` CLI (create, approve-policy, resources,
triggers, answer, status, cockpit), gw_agent/gw_* admin tools, and the
gateway's internal personal wake route.
Includes fixes from review of this branch:
- personal answer refuses to re-resume an already-answered interaction,
and a resumed run that re-suspends now marks its prior continuation
resolved — closing a window where a retried answer could replay tool
side effects.
- The gateway always registers the personal wake route and always starts
the trigger poll loop, so a Personal Agent added after boot has its
triggers picked up without a restart.
- Browser broker leases now check expiry in CanMutate/OwnPage, matching
Authenticate, so an expired lease can no longer authorize mutation.
- Resource grant IDs use a monotonic timestamp instead of locator length,
avoiding same-length-locator collisions.
- Delegation policy narrowing (IsRestriction/NarrowPolicy) now also
checks filesystem roots, browser origins, MCP tools, budget fields, and
quiet hours, not just tools/consequence-classes/concurrency/depth.
- `personal triggers pause/resume` errors on an unknown trigger id instead
of reporting success silently.
- gw_agent's description no longer advertises the unsupported "inspect"
action.
- Suspension continuation files are written atomically.
- The trigger poll loop reuses one DB connection per agent across ticks
and filters due triggers in SQL, instead of opening/closing a
connection and scanning every trigger row every 15s.
… CI lint
Closes the biggest gap from review: the Personal Agent executive had no way
to reach browser/MCP/shell/filesystem capability — it was a second, weaker
agent runtime sitting beside the real memcode engine instead of commanding
it. Two new executive tools fix that:
- delegate: spawns a policy-scoped worker as a real detached `memcode run`
job (internal/jobs.SpawnWithSpec), using the delegation.go scaffolding
(ExecutionEnvelope, ValidateDelegation, PrepareRunDirectory) that already
existed but was never called from anywhere. The worker gets whatever
toolsets it's granted — browser, MCP, shell, filesystem, skills — same as
any ordinary memcode agent, not the executive's fixed 7 tools. Requested
toolsets/consequences must be a subset of the agent's own approved policy;
MaxDelegationDepth gates whether delegation is allowed at all.
- check_delegate: reads a delegated job's status/result back on a later
wake (RunOnce is one bounded wake, so the result can't arrive in the same
call) and closes out its action-journal entry.
Also makes jobs.SpawnSpec.ToolPolicy a REAL restriction instead of recorded-
only metadata: new hidden --allow-tools/--deny-tools flags on `memcode run`
bind to the same SetToolPolicy enforcement an ordinary gateway-bound agent
gets from its config, and SpawnWithSpec now passes them through argv.
Scope notes, called out explicitly rather than implied as done:
- "browser" toolset reuses the existing ephemeral Chrome session (the same
one --chrome already drives) — NOT the user's own already-running Chrome.
Attaching to a persistent/remote Chrome needs the broker/RemoteConfig
wiring in internal/browser, which is declared but has zero callers and
zero implementation; that's separate follow-up work, not done here.
- Filesystem/MCP resource grants are not OS-sandboxed to a spawned worker
(memcode has no such sandbox anywhere) — governance here is the same
approval/audit model as the rest of the codebase (tool-level allow/deny +
the action journal), not a hard jail.
Also fixes two CI lint failures on the open PR:
- internal/gateway/state/state.go: removed formatTime, an unused helper
(staticcheck U1000).
- internal/personal/store.go: decapitalized a Go error string (ST1005).
Corrects course on the browser gap flagged after the last review: existing-
Chrome access is core to Personal Agents, not deferrable follow-up. A
Personal Agent's whole premise — acting as the user across their real
accounts (Gmail, LinkedIn, an ATS, internal dashboards) — requires the
user's actual signed-in session, not a fresh ephemeral profile with no
cookies. Per docs/design/personal-agents.md's already-specified "browser
broker trust boundary": Personal Agents connect to the user's existing
Chrome through a gateway-owned broker and a permission-protected local
socket, fail closed on any connection/auth problem, and never silently fall
back to another profile.
- internal/browser/broker/{server,client}.go: exposes the (already-reviewed)
in-process Broker over a Unix socket, so a delegated worker — a SEPARATE
OS process spawned via jobs.SpawnWithSpec — can Acquire/Release/OwnPage/
CanMutate against the SAME broker the gateway owns, not a fresh one.
Socket is 0600 inside gwconfig.Dir() (0700). Round-trip tested, including
that a concurrent Acquire is correctly rejected and a release unblocks the
next one.
- internal/gateway/server: the gateway now owns one *broker.Broker for its
whole lifetime and serves it on that socket unconditionally at startup
(cheap — a local listener), so it's there the moment a delegate call needs
it, no restart required after `memcode personal browser setup`.
- internal/personal: browserModeFor's "browser" toolset now defaults to
BrowserExistingChrome, not ephemeral — "browser:ephemeral" is the explicit
opt-down for tasks that genuinely want a fresh, logged-out profile. Before
spawning anything, the delegate tool verifies the broker is reachable and
fails the tool call closed (a normal tool_result error, not a spawned job)
if it isn't — verified by test, alongside the success path asserting the
spawned job's SpawnSpec actually carries BrowserMode=existing_chrome.
- internal/jobs: SpawnWithSpec passes existing_chrome through as
--browser-session (+ --browser-agent/--browser-run for the lease identity)
instead of --chrome.
- cmd/run.go: --browser-session existing_chrome acquires the broker lease
itself (via the client) and, only on success, wires an ad-hoc
chrome-devtools-mcp --autoConnect MCP server into the session
(Session.SetExtraMCPServers, new) — the SAME chrome-devtools-mcp package
this browser experience is built on (see internal/browser/remote.go). On
any failure it returns an error and exits; it never calls
SetBrowserEnabled(true) (ephemeral) as a fallback.
- cmd/personal_browser.go: `memcode personal browser setup` checks npx and
the broker socket, then attempts a REAL bounded (10s) chrome-devtools-mcp
connection and reports the tool count or the actual connect errors — it
does not pretend to click Chrome's own "Allow" dialog, which only the user
can do.
- Also fixes a real bug found while testing this: ValidateDelegation treated
an empty parent.AllowedTools as "allows no tools by name", the opposite of
the len(...)>0-means-restricted convention Executive.allowedTools already
uses — every delegate call with a name-unrestricted policy was rejected.
Scope note, stated rather than implied: the literal Chrome-side consent flow
(the "Allow" dialog, actually clicking it) cannot be exercised or verified
in this environment — no live Chrome, no human to click it. Everything
up to that point (broker lifecycle, lease exclusivity, fail-closed wiring,
the actual chrome-devtools-mcp CLI invocation, verified against the real
published package) is built and tested; the dialog itself needs the user's
own machine.
…setup via the cockpit
Three separate complaints from review, addressed at the right layer each:
1. "It's just a file ref" — `resources add` required a redundant --mode read
(already the default) and a mandatory <type> argument even for the
overwhelming common case. A bare path is now enough:
memcode personal resources add jobhunt ~/resume.md
Non-filesystem grants (mcp/command/channel) still need the type spelled
out, since there's nothing to infer it from. Also fixes the real bug this
surfaced: CanonicalFilesystemGrant rejected any non-directory path, even
though PathWithinGrant already handles a single-file grant correctly
(rel == "." on an exact match) — you could not actually grant a single
file like a resume before this, only whole directories. `personal create`
also gained a repeatable --grant flag to fold "create it, then give it
something to read" into one command.
2. "Not this CLI-command-argument garbage" — Personal Agent state (objective,
policy, resources) lived ONLY in personal.db, reachable exclusively
through bespoke commands, unlike every other piece of memcode config
(gateway.yaml, .mcp.json, CLAUDE.md) which is a plain file. WriteConfigMirror
now regenerates objective.md/policy.yaml/resources.yaml in the agent's home
on every mutation — real, readable, diffable files. The run journal
(actions/triggers/interactions) deliberately stays in SQLite: it needs
atomic claim/complete semantics under concurrent access (gateway wake loop,
CLI, and cockpit can all touch the same agent) that flat files don't get
for free — a correctness reason, not a habit.
3. "It should gather requirements, do HITL, pre-approve permissions, set up
the runtime — like our other agents already do via telegram/model setup"
— `memcode personal` (no args) is ALREADY the interactive cockpit (like
`memcode admin`) with the pa_* tools wired for exactly this. The actual gap
was its system prompt: personalAdminDoctrine listed tool mechanics with no
instruction to gather-then-propose-then-approve. Rewrote it to require a
real walkthrough for first-time creation — reason about what resources/
toolsets/consequences the objective needs (asking the user rather than
guessing), present the concrete plan in plain language BEFORE touching
anything, then apply everything at once on approval, including wake
cadence (trigger) — an agent with no trigger and no plan to ever be woken
is dead on arrival, so cadence is not an optional afterthought either.
A prior pass at this (adding a NEW deterministic `memcode personal setup`
CLI wizard, mirroring `gateway setup`'s bufio-prompt mechanism) was wrong
and reverted before commit — that reproduces the exact CLI-surface-
sprawl complaint instead of fixing it. The cockpit conversation IS the
wizard; it needed a better prompt, not a new command.
Root cause of the whole CLI-surface complaint: the interactive cockpit
(`memcode personal`, no args) is supposed to be the entire interface, but
its tool set had NO way to create a new Personal Agent. pa_objective only
supported show/set, and set required the objective to already exist —
creation was only reachable through the CLI's `personal create`. That's why
every explanation kept surfacing a CLI command: until now, it was the only
path that actually worked.
- internal/agent/tools/personal.go: new pa_create tool (agent name +
objective). Documented as the required first call for a brand-new agent.
- cmd/personal_cockpit.go: paCreate mirrors personalCreate's steps (register
in gateway.yaml, open the store, create the objective) and is dispatched
BEFORE the paStore existence gate, since the agent doesn't exist yet.
pa_resource's grant action now defaults type=filesystem/mode=read when
omitted too (same inference the CLI's `resources add` already got), and
pa_resource/pa_policy's grant/revoke/stage/approve paths now call
WriteConfigMirror — that was wired into the CLI paths last commit but not
into the cockpit paths, which are the ones actually used.
- cmd/personal.go: every `memcode personal` subcommand is now Hidden — they
remain fully callable for scripts, but `memcode personal --help` no longer
reads like a CLI to memorize. The interface is `memcode personal` with no
args, then plain conversation.
- internal/doctrine/prompts.go: the cockpit's own system prompt now says
outright never to tell the user to run a CLI command — it does the work
itself, right there — and that pa_create, not pa_objective, is the first
call for a new agent.
Verified end to end: pa_create then a bare-path pa_resource grant (no type,
no mode) both succeed through personalExecute exactly as the cockpit would
call them, no CLI involved.
Hiding the subcommands last commit wasn't enough — they still existed as a
parallel, scriptable way to do everything, which is real duplicate logic
(and already caused drift: a fix landed in one path and not the other,
twice, in this same review). Deleted them outright:
cmd/personal_resources.go, personal_policy.go, personal_triggers.go,
personal_status.go, personal_browser.go, personal_test.go — all removed.
cmd/personal.go cut down to just the cockpit launcher (~90 lines from
~370): `memcode personal` opens the interactive session, full stop.
Two operations the CLI covered had no pa_* equivalent yet, so they'd have
been silently lost:
- pa_doctor: the health check (home layout, objective, approved policy,
generated workspace, sandbox availability, trigger/pending-interaction
counts) that `personal doctor` used to run.
- pa_browser_setup: the existing-Chrome prerequisite check + real, bounded
chrome-devtools-mcp connection attempt that `personal browser setup` used
to run. Same logic, now a tool call instead of a command — it still
doesn't (can't) click Chrome's own consent dialog for the user.
Both dispatch through personalExecute exactly like every other pa_* tool.
internal/agent/tools/personal.go: added PaDoctor and PaBrowserSetup to the
registry.
cmd/personal_cmd_test.go: rewritten from scratch to call personalExecute
directly — no CLI args, no cobra Execute(), no rootCmd. This IS the
interface being tested now, so this is what the tests exercise. Also
asserts the config-mirror files (policy.yaml, resources.yaml) actually
reflect approvals/revokes done through the tool path, since that's the one
that matters.
Verified: `memcode personal create foo bar` no longer resolves to anything
— it just opens the cockpit with those words ignored as stray args, because
there is no "create" subcommand left to match. The interface is
conversation with pa_* tools behind it, not CLI argument syntax.
objective.md/policy.yaml/resources.yaml split one agent's config across
three files by table for no real reason — a person has to know to look in
three places for one agent's setup. Consolidated into a single config.yaml
(objective, policies, resources as sections) with proper snake_case field
names. Same authority split as before (this is a mirror; policy hash-
approval and the run journal stay in personal.db, and why, is unchanged) —
just one file to read instead of three.
Consolidation step 1 of folding Personal Agents into the ordinary agent
system. There were three partial suspend/resume designs in the tree:
1. internal/agent/runtime/continuation.go — typed, atomic, tested, and with
ZERO production callers.
2. jobs.Job's InteractionID/WaitingReason/ContinuationVersion/WaitingAt/
ResumedAt — declared, never written by anything.
3. A hand-rolled map[string]any in the personal executive — the only one
actually running, and (until a fix earlier in this branch) the only one
that wasn't crash-safe.
New internal/agent/continuation is the single implementation. It lives in its
own small package rather than in internal/agent/runtime so the executive can
suspend without importing the whole session runtime (and so neither side ends
up depending on the other).
It has to serve two genuinely different callers, which is why the previous
attempt at a shared API didn't fit:
- An interactive session already holds the conversation, so it only needs the
missing assistant/answer pair back.
- An unattended executive keeps NO transcript — it rebuilds context from
durable state each wake — so its continuation must carry the whole
conversation. Hence the optional Messages field.
They also differ on when marking resolved is safe, so building the resume
messages is split from marking:
- ResumeMessages() builds without marking. The executive uses this and marks
only once the resumed run reaches a terminal state, so a transient model
error leaves the answer re-giveable instead of stranding it.
- Resolve() does both, for a caller with a human right there who can simply
ask again.
Also deleted the five never-written Job fields (2 above). StatusWaiting stays:
it is a meaningful status in the job state machine and check_delegate
correctly treats it as non-terminal.
Tests: ported both original round-trip/validation tests, plus new coverage for
the transcript-carrying path and MarkResolved. The executive's suspend/resume
test now asserts through the continuation API instead of a hardcoded filename.
… admin
Consolidation step 2. "Personal Agent" is no longer a kind of agent. There is
one Agent abstraction; autonomy is orthogonal settings on it, managed through
the admin cockpit that already manages agents.
CONFIG. gwconfig.Agent drops Kind and gains four settings:
objective — the durable outcome it works toward
autonomous — whether it may act on that unprompted
browser — ephemeral (default) | existing_chrome
paused — stop unattended wakes without deleting anything
objective and autonomous are deliberately SEPARATE grants, which is the
correction that motivated this design. They answer different questions ("what
is it for" vs "may it act unasked"), and all four combinations are meaningful:
autonomous + objective → unattended objective pursuit (the old Personal Agent)
autonomous, no objective→ scheduled work UNDER GOVERNANCE — this is new, and
closes the long-standing gap where a cron-fired
agent ran unattended with no policy gate, no action
journal, and no way to pause and ask
objective, not autonomous → a goal you work on together; wakes only on demand
neither → an ordinary conversational agent, exactly as before
Autonomy therefore gates GOVERNANCE, not capability, and applies to any run of
the agent whether or not an objective exists.
SCHEDULING. The second cron implementation is gone as a user-facing concept.
Recurring cadence is an ordinary `schedules:` entry: set agent=<name>, leave
deliver_to empty, and gw_schedule routes the wake to the agent itself via the
renamed internal sink (personal → agent). The DB-backed loop now fires only
the agent's OWN self-scheduled next-wakes (schedule_wake, "come back in 45
minutes") — a genuinely different thing from human-authored cadence, and the
only part that needs to be writable from inside a run.
TOOLS. internal/agent/tools/personal.go (pa_*) is deleted. The admin registry
gains gw_policy, gw_grant, gw_wake, gw_inbox, gw_answer, gw_journal, gw_doctor,
gw_browser; gw_agent gains objective/autonomous/browser/pause/resume actions.
Everything dispatches through the existing adminExecute + approval gate, and
gw_schedule's shared BuildSchedule validation now covers agent cadence too, so
the surfaces cannot drift the way the parallel path did.
COCKPIT. `memcode personal` is deleted outright — cmd/personal.go,
cmd/personal_cockpit.go, Session.SetPersonal, personalMode, the personal and
personal_admin doctrines. The Personal setup walkthrough is merged into
adminDoctrine, generalized, and now teaches the two gates explicitly: granting
an objective is not granting autonomy, and the second must be confirmed on its
own.
EXECUTIVE. Executive.Objective is read from configuration rather than the
store, so the objective has ONE source that a human edits and the gateway
hot-reloads. The per-agent config.yaml mirror drops the objective for the same
reason — it lives in gateway.yaml, which is already a readable file.
Tests are rewritten against adminExecute (there is no other surface). New
coverage asserts the orthogonality directly, since a single overloaded switch
is exactly what this change exists to prevent: an objective alone must not
confer autonomy, a non-affirmative value must not grant it by typo, an
autonomous agent with no objective is still governed but refuses to invent
work, and an autonomous agent's schedule defaults to the agent route.
…ard both
Consolidation step 3, finishing the merge.
RENAME. internal/personal -> internal/agent/autonomy. The package was never
about a "personal" species; it is the machinery an agent uses when running
unattended toward an objective. Terminology in comments follows.
SECOND CRON PARSER DELETED. NextDue now understands only the kinds an AGENT
writes for itself from inside a run (one_shot / next_wake — always a single
future instant, e.g. schedule_wake "come back in 45 minutes"). Recurring
cadence a human configures is not a trigger at all: it is an ordinary gateway
schedule delivering to agent:<name>, validated once by gwconfig. With
recurring kinds gone, ClaimDueTrigger's reschedule branch was unreachable, so
firing a wake now simply completes it — and the atomic `last_fired_at IS ?`
claim still keeps two gateway processes from double-firing.
GUARDS (internal/guard/singletons_test.go). The duplication this branch kept
producing was not a one-off: fixes landed on one path and not the other,
repeatedly. Four invariants now fail the build instead:
- TestSingleCronParser — only internal/gateway parses cron. (This one caught a
real leftover while being written: the autonomy package was still importing
robfig/cron, which is what prompted the deletion above.)
- TestSingleSuspensionImplementation — no hand-rolled continuation formats
outside internal/agent/continuation.
- TestNoSecondCockpit — no SetPersonal/personalMode/pa_* anywhere.
- TestNoAgentKind — autonomy is orthogonal settings, never a kind
discriminator; a "kind" field is what made Personal a separate species.
The guards deliberately skip vendored forks, .memcode session snapshots, and
desktop/node_modules (the same symlink that breaks the two pre-existing
go-list guard failures).
Store tests updated: the trigger test now exercises a real self-scheduled wake
and asserts it completes rather than reschedules, and NextDue's test asserts
that interval/cron are REJECTED — accepting them again would rebuild the
second scheduler.
Consolidation step 4. An agent had two unrelated places to put what it knew:
a structured `facts` table only reachable through bespoke tools, and the
memory.md every memcode agent already has and the runtime already injects. One
of those a human can read. Collapse to it.
BLOCKING PREREQUISITE, done first. delegate/check_delegate were using the
facts table as a keyed index — writing "delegation.<job-id>" so a later wake
could find the action to close out. Its own comment admitted why ("facts are
the only durable log delegate can write to without a schema migration"). This
is that migration: actions gain a job_id column plus LinkActionJob /
ActionForJob, and the link now lives on the action it actually describes.
Removing facts before this would have broken check_delegate silently.
note_fact becomes `remember`: one plain sentence appended to memory.md, read
back into every wake's state summary. Append-only and deduplicated, because
every line replays into the model on each future wake — a re-learned fact must
not grow the file forever.
Deleted with it: environment.go (EnvironmentModel/StructuredFact, referenced
nowhere but its own test) and UsableForExternalRepresentation.
The tradeoff is deliberate and documented in memory.go rather than left
implicit. Prose cannot distinguish "you told me this" from "I inferred it from
your resume" from "a website said so", nor mark a claim safe to assert on the
user's behalf, nor mark it stale. That distinction becomes load-bearing the
moment an agent fills in a form or sends a message stating something about the
user. It is cheap to drop TODAY only because the gate was never wired —
nothing read Confirmed to gate anything; the sole caller of
UsableForExternalRepresentation was its own test. When external
representation needs provenance, it should come back as its own
machine-checkable thing in the store alongside policies and the action
journal, NOT by reviving this table.
… as a mode
Consolidation step 5, finishing the merge.
MIGRATION. YAML silently ignores unknown fields, so an existing
`kind: personal` agent would have loaded as an ordinary one — still listed,
apparently fine, and never waking again. For a change that removes an agent's
authority to run on its own, silence is the wrong failure. gwconfig.Agent keeps
a LegacyKind field for the sole purpose of rejecting it, with the one-line fix
in the error:
agent "demo" still uses the removed `kind: personal` setting. Autonomy is now
explicit: replace it with `autonomous: true` (and an `objective:` ...), or
just delete the `kind:` line if it should not. Its home ... is untouched
Verified against the real local config, which has exactly that agent. The
guard's TestNoAgentKind is refined to allow this one mention while still
banning any behavioral branch on it.
DOCS. docs/personal-agents.md -> docs/autonomous-agents.md, rewritten around
the two orthogonal grants and the four combinations they produce, with the
scheduling split (your cadence = an ordinary schedule; the agent's own next
wake = self-scheduled), the fail-closed existing-Chrome rule, and the
memory.md provenance limitation stated rather than glossed.
docs/design/personal-agents.md -> docs/design/autonomous-agents.md keeps the
original design contract but opens with a note on what the implementation
revised and why, including the orthogonality correction. It is a design record;
rewriting history there would lose the reasoning.
README drops "three ways to run it" (there are two) and describes autonomy as a
setting rather than a product tier, leading with the case that actually
motivated it: a scheduled agent that is finally policy-gated, journaled, and
able to stop and ask.
The store file was still named personal.db. It holds state for any agent
running unattended, not a species of agent — rename it, and note in place that
it is opened lazily so an ordinary conversational agent never grows one.
That claim is now a test rather than a comment: TestOrdinaryAgentGetsNoAutonomyStore
creates a plain agent, exercises the admin surface against it, and asserts no
agent.db appears — the governance machinery costs nothing until asked for.
Also covers the remaining orthogonality row: an agent WITH an objective but
WITHOUT autonomy still wakes on demand, reaching the policy gate rather than
being refused for lacking autonomy. Autonomy governs unprompted action, not
whether a human may ask.
All four combinations of (objective, autonomous) now have a test.
@timothyerwintimothyerwin changed the title Personal Agents: durable objectives, delegation policy, triggers, browser brokerAutonomous AgentsAug 30, 2026
@timothyerwin
timothyerwin merged commit 8a2e6b3 into mainAug 30, 2026
2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@timothyerwin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Autonomous Agents - #5

Merged
timothyerwin merged 13 commits into
mainfrom
feat/personal-agents
Aug 30, 2026
Merged

Autonomous Agents#5
timothyerwin merged 13 commits into
mainfrom
feat/personal-agents

Conversation

@timothyerwin

Copy link
Copy Markdown
Contributor

Summary

Adds Personal Agents: long-lived agent identities with a durable objective/subgoal/facts store, a delegation policy that governs what their executive loop may do (allowed tools, consequence classes, filesystem/browser/MCP resource grants, budgets), ask_user suspend/resume for human-in-the-loop interactions, durable cron/interval/one-shot triggers polled by the gateway, and a browser broker that leases exclusive browser control to one run at a time.

New surfaces: memcode personal CLI (create, approve-policy, resources, triggers, answer, status, cockpit), gw_agent/gw_* admin tools, and the gateway's internal personal wake route.

Review fixes included

This branch went through a review pass (/code-review high) before commit; the following were found and fixed:

  • Duplicate-answer replay: personal answer now refuses to re-resume an already-answered interaction, and a resumed run that re-suspends on a new question now marks its prior continuation resolved — closing a window where a retried personal answer call could re-execute real tool side effects.
  • Triggers silently dead after hot-add: the gateway now always registers the personal wake route and always starts the trigger poll loop (it already re-reads config every tick), so a Personal Agent created after boot has its triggers picked up without a gateway restart.
  • Expired browser lease still granted mutation rights: CanMutate/OwnPage now check lease expiry like Authenticate already did.
  • Resource grant ID collisions: IDs now use a monotonic timestamp instead of len(locator), matching the cockpit's existing scheme.
  • Delegation policy narrowing gaps: IsRestriction/NarrowPolicy now also checks filesystem roots, browser origins, MCP tools, budget fields, and quiet hours, not just tools/consequence-classes/concurrency/depth.
  • Silent no-op on bad trigger id: personal triggers pause/resume now errors instead of reporting success when nothing matched.
  • Stale tool description: gw_agent no longer advertises the unsupported "inspect" action.
  • Non-atomic suspension writes: suspension continuation files are now written via atomicfile.WriteFile.
  • Per-tick DB churn: the trigger poll loop now reuses one DB connection per agent across ticks and filters due triggers in SQL instead of opening/closing a connection and Go-filtering every row every 15s.

Testing

  • go build ./... — clean
  • go vet ./... — clean
  • go test ./... — all pass except two pre-existing, unrelated failures in internal/guard caused by a symlink under desktop/node_modules/@electron/notarize breaking go list (not touched by this branch)

… triggers, browser broker)
Implements the Personal Agents feature: long-lived agent identities with
their own durable objective, subgoals, and facts store; a delegation
policy model (allowed tools, consequence classes, filesystem/browser/MCP
grants, budgets) that governs what a Personal Agent's executive loop may
do; ask_user suspend/resume for human-in-the-loop interactions; durable
triggers (cron/interval/one-shot) polled by the gateway to wake an agent;
resource grants (filesystem, mcp, command, channel); and a browser broker
that leases exclusive browser control to one run at a time.
New surfaces: `memcode personal` CLI (create, approve-policy, resources,
triggers, answer, status, cockpit), gw_agent/gw_* admin tools, and the
gateway's internal personal wake route.
Includes fixes from review of this branch:
- personal answer refuses to re-resume an already-answered interaction,
and a resumed run that re-suspends now marks its prior continuation
resolved — closing a window where a retried answer could replay tool
side effects.
- The gateway always registers the personal wake route and always starts
the trigger poll loop, so a Personal Agent added after boot has its
triggers picked up without a restart.
- Browser broker leases now check expiry in CanMutate/OwnPage, matching
Authenticate, so an expired lease can no longer authorize mutation.
- Resource grant IDs use a monotonic timestamp instead of locator length,
avoiding same-length-locator collisions.
- Delegation policy narrowing (IsRestriction/NarrowPolicy) now also
checks filesystem roots, browser origins, MCP tools, budget fields, and
quiet hours, not just tools/consequence-classes/concurrency/depth.
- `personal triggers pause/resume` errors on an unknown trigger id instead
of reporting success silently.
- gw_agent's description no longer advertises the unsupported "inspect"
action.
- Suspension continuation files are written atomically.
- The trigger poll loop reuses one DB connection per agent across ticks
and filters due triggers in SQL, instead of opening/closing a
connection and scanning every trigger row every 15s.
… CI lint
Closes the biggest gap from review: the Personal Agent executive had no way
to reach browser/MCP/shell/filesystem capability — it was a second, weaker
agent runtime sitting beside the real memcode engine instead of commanding
it. Two new executive tools fix that:
- delegate: spawns a policy-scoped worker as a real detached `memcode run`
job (internal/jobs.SpawnWithSpec), using the delegation.go scaffolding
(ExecutionEnvelope, ValidateDelegation, PrepareRunDirectory) that already
existed but was never called from anywhere. The worker gets whatever
toolsets it's granted — browser, MCP, shell, filesystem, skills — same as
any ordinary memcode agent, not the executive's fixed 7 tools. Requested
toolsets/consequences must be a subset of the agent's own approved policy;
MaxDelegationDepth gates whether delegation is allowed at all.
- check_delegate: reads a delegated job's status/result back on a later
wake (RunOnce is one bounded wake, so the result can't arrive in the same
call) and closes out its action-journal entry.
Also makes jobs.SpawnSpec.ToolPolicy a REAL restriction instead of recorded-
only metadata: new hidden --allow-tools/--deny-tools flags on `memcode run`
bind to the same SetToolPolicy enforcement an ordinary gateway-bound agent
gets from its config, and SpawnWithSpec now passes them through argv.
Scope notes, called out explicitly rather than implied as done:
- "browser" toolset reuses the existing ephemeral Chrome session (the same
one --chrome already drives) — NOT the user's own already-running Chrome.
Attaching to a persistent/remote Chrome needs the broker/RemoteConfig
wiring in internal/browser, which is declared but has zero callers and
zero implementation; that's separate follow-up work, not done here.
- Filesystem/MCP resource grants are not OS-sandboxed to a spawned worker
(memcode has no such sandbox anywhere) — governance here is the same
approval/audit model as the rest of the codebase (tool-level allow/deny +
the action journal), not a hard jail.
Also fixes two CI lint failures on the open PR:
- internal/gateway/state/state.go: removed formatTime, an unused helper
(staticcheck U1000).
- internal/personal/store.go: decapitalized a Go error string (ST1005).
Corrects course on the browser gap flagged after the last review: existing-
Chrome access is core to Personal Agents, not deferrable follow-up. A
Personal Agent's whole premise — acting as the user across their real
accounts (Gmail, LinkedIn, an ATS, internal dashboards) — requires the
user's actual signed-in session, not a fresh ephemeral profile with no
cookies. Per docs/design/personal-agents.md's already-specified "browser
broker trust boundary": Personal Agents connect to the user's existing
Chrome through a gateway-owned broker and a permission-protected local
socket, fail closed on any connection/auth problem, and never silently fall
back to another profile.
- internal/browser/broker/{server,client}.go: exposes the (already-reviewed)
in-process Broker over a Unix socket, so a delegated worker — a SEPARATE
OS process spawned via jobs.SpawnWithSpec — can Acquire/Release/OwnPage/
CanMutate against the SAME broker the gateway owns, not a fresh one.
Socket is 0600 inside gwconfig.Dir() (0700). Round-trip tested, including
that a concurrent Acquire is correctly rejected and a release unblocks the
next one.
- internal/gateway/server: the gateway now owns one *broker.Broker for its
whole lifetime and serves it on that socket unconditionally at startup
(cheap — a local listener), so it's there the moment a delegate call needs
it, no restart required after `memcode personal browser setup`.
- internal/personal: browserModeFor's "browser" toolset now defaults to
BrowserExistingChrome, not ephemeral — "browser:ephemeral" is the explicit
opt-down for tasks that genuinely want a fresh, logged-out profile. Before
spawning anything, the delegate tool verifies the broker is reachable and
fails the tool call closed (a normal tool_result error, not a spawned job)
if it isn't — verified by test, alongside the success path asserting the
spawned job's SpawnSpec actually carries BrowserMode=existing_chrome.
- internal/jobs: SpawnWithSpec passes existing_chrome through as
--browser-session (+ --browser-agent/--browser-run for the lease identity)
instead of --chrome.
- cmd/run.go: --browser-session existing_chrome acquires the broker lease
itself (via the client) and, only on success, wires an ad-hoc
chrome-devtools-mcp --autoConnect MCP server into the session
(Session.SetExtraMCPServers, new) — the SAME chrome-devtools-mcp package
this browser experience is built on (see internal/browser/remote.go). On
any failure it returns an error and exits; it never calls
SetBrowserEnabled(true) (ephemeral) as a fallback.
- cmd/personal_browser.go: `memcode personal browser setup` checks npx and
the broker socket, then attempts a REAL bounded (10s) chrome-devtools-mcp
connection and reports the tool count or the actual connect errors — it
does not pretend to click Chrome's own "Allow" dialog, which only the user
can do.
- Also fixes a real bug found while testing this: ValidateDelegation treated
an empty parent.AllowedTools as "allows no tools by name", the opposite of
the len(...)>0-means-restricted convention Executive.allowedTools already
uses — every delegate call with a name-unrestricted policy was rejected.
Scope note, stated rather than implied: the literal Chrome-side consent flow
(the "Allow" dialog, actually clicking it) cannot be exercised or verified
in this environment — no live Chrome, no human to click it. Everything
up to that point (broker lifecycle, lease exclusivity, fail-closed wiring,
the actual chrome-devtools-mcp CLI invocation, verified against the real
published package) is built and tested; the dialog itself needs the user's
own machine.
…setup via the cockpit
Three separate complaints from review, addressed at the right layer each:
1. "It's just a file ref" — `resources add` required a redundant --mode read
(already the default) and a mandatory <type> argument even for the
overwhelming common case. A bare path is now enough:
memcode personal resources add jobhunt ~/resume.md
Non-filesystem grants (mcp/command/channel) still need the type spelled
out, since there's nothing to infer it from. Also fixes the real bug this
surfaced: CanonicalFilesystemGrant rejected any non-directory path, even
though PathWithinGrant already handles a single-file grant correctly
(rel == "." on an exact match) — you could not actually grant a single
file like a resume before this, only whole directories. `personal create`
also gained a repeatable --grant flag to fold "create it, then give it
something to read" into one command.
2. "Not this CLI-command-argument garbage" — Personal Agent state (objective,
policy, resources) lived ONLY in personal.db, reachable exclusively
through bespoke commands, unlike every other piece of memcode config
(gateway.yaml, .mcp.json, CLAUDE.md) which is a plain file. WriteConfigMirror
now regenerates objective.md/policy.yaml/resources.yaml in the agent's home
on every mutation — real, readable, diffable files. The run journal
(actions/triggers/interactions) deliberately stays in SQLite: it needs
atomic claim/complete semantics under concurrent access (gateway wake loop,
CLI, and cockpit can all touch the same agent) that flat files don't get
for free — a correctness reason, not a habit.
3. "It should gather requirements, do HITL, pre-approve permissions, set up
the runtime — like our other agents already do via telegram/model setup"
— `memcode personal` (no args) is ALREADY the interactive cockpit (like
`memcode admin`) with the pa_* tools wired for exactly this. The actual gap
was its system prompt: personalAdminDoctrine listed tool mechanics with no
instruction to gather-then-propose-then-approve. Rewrote it to require a
real walkthrough for first-time creation — reason about what resources/
toolsets/consequences the objective needs (asking the user rather than
guessing), present the concrete plan in plain language BEFORE touching
anything, then apply everything at once on approval, including wake
cadence (trigger) — an agent with no trigger and no plan to ever be woken
is dead on arrival, so cadence is not an optional afterthought either.
A prior pass at this (adding a NEW deterministic `memcode personal setup`
CLI wizard, mirroring `gateway setup`'s bufio-prompt mechanism) was wrong
and reverted before commit — that reproduces the exact CLI-surface-
sprawl complaint instead of fixing it. The cockpit conversation IS the
wizard; it needed a better prompt, not a new command.
Root cause of the whole CLI-surface complaint: the interactive cockpit
(`memcode personal`, no args) is supposed to be the entire interface, but
its tool set had NO way to create a new Personal Agent. pa_objective only
supported show/set, and set required the objective to already exist —
creation was only reachable through the CLI's `personal create`. That's why
every explanation kept surfacing a CLI command: until now, it was the only
path that actually worked.
- internal/agent/tools/personal.go: new pa_create tool (agent name +
objective). Documented as the required first call for a brand-new agent.
- cmd/personal_cockpit.go: paCreate mirrors personalCreate's steps (register
in gateway.yaml, open the store, create the objective) and is dispatched
BEFORE the paStore existence gate, since the agent doesn't exist yet.
pa_resource's grant action now defaults type=filesystem/mode=read when
omitted too (same inference the CLI's `resources add` already got), and
pa_resource/pa_policy's grant/revoke/stage/approve paths now call
WriteConfigMirror — that was wired into the CLI paths last commit but not
into the cockpit paths, which are the ones actually used.
- cmd/personal.go: every `memcode personal` subcommand is now Hidden — they
remain fully callable for scripts, but `memcode personal --help` no longer
reads like a CLI to memorize. The interface is `memcode personal` with no
args, then plain conversation.
- internal/doctrine/prompts.go: the cockpit's own system prompt now says
outright never to tell the user to run a CLI command — it does the work
itself, right there — and that pa_create, not pa_objective, is the first
call for a new agent.
Verified end to end: pa_create then a bare-path pa_resource grant (no type,
no mode) both succeed through personalExecute exactly as the cockpit would
call them, no CLI involved.
Hiding the subcommands last commit wasn't enough — they still existed as a
parallel, scriptable way to do everything, which is real duplicate logic
(and already caused drift: a fix landed in one path and not the other,
twice, in this same review). Deleted them outright:
cmd/personal_resources.go, personal_policy.go, personal_triggers.go,
personal_status.go, personal_browser.go, personal_test.go — all removed.
cmd/personal.go cut down to just the cockpit launcher (~90 lines from
~370): `memcode personal` opens the interactive session, full stop.
Two operations the CLI covered had no pa_* equivalent yet, so they'd have
been silently lost:
- pa_doctor: the health check (home layout, objective, approved policy,
generated workspace, sandbox availability, trigger/pending-interaction
counts) that `personal doctor` used to run.
- pa_browser_setup: the existing-Chrome prerequisite check + real, bounded
chrome-devtools-mcp connection attempt that `personal browser setup` used
to run. Same logic, now a tool call instead of a command — it still
doesn't (can't) click Chrome's own consent dialog for the user.
Both dispatch through personalExecute exactly like every other pa_* tool.
internal/agent/tools/personal.go: added PaDoctor and PaBrowserSetup to the
registry.
cmd/personal_cmd_test.go: rewritten from scratch to call personalExecute
directly — no CLI args, no cobra Execute(), no rootCmd. This IS the
interface being tested now, so this is what the tests exercise. Also
asserts the config-mirror files (policy.yaml, resources.yaml) actually
reflect approvals/revokes done through the tool path, since that's the one
that matters.
Verified: `memcode personal create foo bar` no longer resolves to anything
— it just opens the cockpit with those words ignored as stray args, because
there is no "create" subcommand left to match. The interface is
conversation with pa_* tools behind it, not CLI argument syntax.
objective.md/policy.yaml/resources.yaml split one agent's config across
three files by table for no real reason — a person has to know to look in
three places for one agent's setup. Consolidated into a single config.yaml
(objective, policies, resources as sections) with proper snake_case field
names. Same authority split as before (this is a mirror; policy hash-
approval and the run journal stay in personal.db, and why, is unchanged) —
just one file to read instead of three.
Consolidation step 1 of folding Personal Agents into the ordinary agent
system. There were three partial suspend/resume designs in the tree:
1. internal/agent/runtime/continuation.go — typed, atomic, tested, and with
ZERO production callers.
2. jobs.Job's InteractionID/WaitingReason/ContinuationVersion/WaitingAt/
ResumedAt — declared, never written by anything.
3. A hand-rolled map[string]any in the personal executive — the only one
actually running, and (until a fix earlier in this branch) the only one
that wasn't crash-safe.
New internal/agent/continuation is the single implementation. It lives in its
own small package rather than in internal/agent/runtime so the executive can
suspend without importing the whole session runtime (and so neither side ends
up depending on the other).
It has to serve two genuinely different callers, which is why the previous
attempt at a shared API didn't fit:
- An interactive session already holds the conversation, so it only needs the
missing assistant/answer pair back.
- An unattended executive keeps NO transcript — it rebuilds context from
durable state each wake — so its continuation must carry the whole
conversation. Hence the optional Messages field.
They also differ on when marking resolved is safe, so building the resume
messages is split from marking:
- ResumeMessages() builds without marking. The executive uses this and marks
only once the resumed run reaches a terminal state, so a transient model
error leaves the answer re-giveable instead of stranding it.
- Resolve() does both, for a caller with a human right there who can simply
ask again.
Also deleted the five never-written Job fields (2 above). StatusWaiting stays:
it is a meaningful status in the job state machine and check_delegate
correctly treats it as non-terminal.
Tests: ported both original round-trip/validation tests, plus new coverage for
the transcript-carrying path and MarkResolved. The executive's suspend/resume
test now asserts through the continuation API instead of a hardcoded filename.
… admin
Consolidation step 2. "Personal Agent" is no longer a kind of agent. There is
one Agent abstraction; autonomy is orthogonal settings on it, managed through
the admin cockpit that already manages agents.
CONFIG. gwconfig.Agent drops Kind and gains four settings:
objective — the durable outcome it works toward
autonomous — whether it may act on that unprompted
browser — ephemeral (default) | existing_chrome
paused — stop unattended wakes without deleting anything
objective and autonomous are deliberately SEPARATE grants, which is the
correction that motivated this design. They answer different questions ("what
is it for" vs "may it act unasked"), and all four combinations are meaningful:
autonomous + objective → unattended objective pursuit (the old Personal Agent)
autonomous, no objective→ scheduled work UNDER GOVERNANCE — this is new, and
closes the long-standing gap where a cron-fired
agent ran unattended with no policy gate, no action
journal, and no way to pause and ask
objective, not autonomous → a goal you work on together; wakes only on demand
neither → an ordinary conversational agent, exactly as before
Autonomy therefore gates GOVERNANCE, not capability, and applies to any run of
the agent whether or not an objective exists.
SCHEDULING. The second cron implementation is gone as a user-facing concept.
Recurring cadence is an ordinary `schedules:` entry: set agent=<name>, leave
deliver_to empty, and gw_schedule routes the wake to the agent itself via the
renamed internal sink (personal → agent). The DB-backed loop now fires only
the agent's OWN self-scheduled next-wakes (schedule_wake, "come back in 45
minutes") — a genuinely different thing from human-authored cadence, and the
only part that needs to be writable from inside a run.
TOOLS. internal/agent/tools/personal.go (pa_*) is deleted. The admin registry
gains gw_policy, gw_grant, gw_wake, gw_inbox, gw_answer, gw_journal, gw_doctor,
gw_browser; gw_agent gains objective/autonomous/browser/pause/resume actions.
Everything dispatches through the existing adminExecute + approval gate, and
gw_schedule's shared BuildSchedule validation now covers agent cadence too, so
the surfaces cannot drift the way the parallel path did.
COCKPIT. `memcode personal` is deleted outright — cmd/personal.go,
cmd/personal_cockpit.go, Session.SetPersonal, personalMode, the personal and
personal_admin doctrines. The Personal setup walkthrough is merged into
adminDoctrine, generalized, and now teaches the two gates explicitly: granting
an objective is not granting autonomy, and the second must be confirmed on its
own.
EXECUTIVE. Executive.Objective is read from configuration rather than the
store, so the objective has ONE source that a human edits and the gateway
hot-reloads. The per-agent config.yaml mirror drops the objective for the same
reason — it lives in gateway.yaml, which is already a readable file.
Tests are rewritten against adminExecute (there is no other surface). New
coverage asserts the orthogonality directly, since a single overloaded switch
is exactly what this change exists to prevent: an objective alone must not
confer autonomy, a non-affirmative value must not grant it by typo, an
autonomous agent with no objective is still governed but refuses to invent
work, and an autonomous agent's schedule defaults to the agent route.
…ard both
Consolidation step 3, finishing the merge.
RENAME. internal/personal -> internal/agent/autonomy. The package was never
about a "personal" species; it is the machinery an agent uses when running
unattended toward an objective. Terminology in comments follows.
SECOND CRON PARSER DELETED. NextDue now understands only the kinds an AGENT
writes for itself from inside a run (one_shot / next_wake — always a single
future instant, e.g. schedule_wake "come back in 45 minutes"). Recurring
cadence a human configures is not a trigger at all: it is an ordinary gateway
schedule delivering to agent:<name>, validated once by gwconfig. With
recurring kinds gone, ClaimDueTrigger's reschedule branch was unreachable, so
firing a wake now simply completes it — and the atomic `last_fired_at IS ?`
claim still keeps two gateway processes from double-firing.
GUARDS (internal/guard/singletons_test.go). The duplication this branch kept
producing was not a one-off: fixes landed on one path and not the other,
repeatedly. Four invariants now fail the build instead:
- TestSingleCronParser — only internal/gateway parses cron. (This one caught a
real leftover while being written: the autonomy package was still importing
robfig/cron, which is what prompted the deletion above.)
- TestSingleSuspensionImplementation — no hand-rolled continuation formats
outside internal/agent/continuation.
- TestNoSecondCockpit — no SetPersonal/personalMode/pa_* anywhere.
- TestNoAgentKind — autonomy is orthogonal settings, never a kind
discriminator; a "kind" field is what made Personal a separate species.
The guards deliberately skip vendored forks, .memcode session snapshots, and
desktop/node_modules (the same symlink that breaks the two pre-existing
go-list guard failures).
Store tests updated: the trigger test now exercises a real self-scheduled wake
and asserts it completes rather than reschedules, and NextDue's test asserts
that interval/cron are REJECTED — accepting them again would rebuild the
second scheduler.
Consolidation step 4. An agent had two unrelated places to put what it knew:
a structured `facts` table only reachable through bespoke tools, and the
memory.md every memcode agent already has and the runtime already injects. One
of those a human can read. Collapse to it.
BLOCKING PREREQUISITE, done first. delegate/check_delegate were using the
facts table as a keyed index — writing "delegation.<job-id>" so a later wake
could find the action to close out. Its own comment admitted why ("facts are
the only durable log delegate can write to without a schema migration"). This
is that migration: actions gain a job_id column plus LinkActionJob /
ActionForJob, and the link now lives on the action it actually describes.
Removing facts before this would have broken check_delegate silently.
note_fact becomes `remember`: one plain sentence appended to memory.md, read
back into every wake's state summary. Append-only and deduplicated, because
every line replays into the model on each future wake — a re-learned fact must
not grow the file forever.
Deleted with it: environment.go (EnvironmentModel/StructuredFact, referenced
nowhere but its own test) and UsableForExternalRepresentation.
The tradeoff is deliberate and documented in memory.go rather than left
implicit. Prose cannot distinguish "you told me this" from "I inferred it from
your resume" from "a website said so", nor mark a claim safe to assert on the
user's behalf, nor mark it stale. That distinction becomes load-bearing the
moment an agent fills in a form or sends a message stating something about the
user. It is cheap to drop TODAY only because the gate was never wired —
nothing read Confirmed to gate anything; the sole caller of
UsableForExternalRepresentation was its own test. When external
representation needs provenance, it should come back as its own
machine-checkable thing in the store alongside policies and the action
journal, NOT by reviving this table.
… as a mode
Consolidation step 5, finishing the merge.
MIGRATION. YAML silently ignores unknown fields, so an existing
`kind: personal` agent would have loaded as an ordinary one — still listed,
apparently fine, and never waking again. For a change that removes an agent's
authority to run on its own, silence is the wrong failure. gwconfig.Agent keeps
a LegacyKind field for the sole purpose of rejecting it, with the one-line fix
in the error:
agent "demo" still uses the removed `kind: personal` setting. Autonomy is now
explicit: replace it with `autonomous: true` (and an `objective:` ...), or
just delete the `kind:` line if it should not. Its home ... is untouched
Verified against the real local config, which has exactly that agent. The
guard's TestNoAgentKind is refined to allow this one mention while still
banning any behavioral branch on it.
DOCS. docs/personal-agents.md -> docs/autonomous-agents.md, rewritten around
the two orthogonal grants and the four combinations they produce, with the
scheduling split (your cadence = an ordinary schedule; the agent's own next
wake = self-scheduled), the fail-closed existing-Chrome rule, and the
memory.md provenance limitation stated rather than glossed.
docs/design/personal-agents.md -> docs/design/autonomous-agents.md keeps the
original design contract but opens with a note on what the implementation
revised and why, including the orthogonality correction. It is a design record;
rewriting history there would lose the reasoning.
README drops "three ways to run it" (there are two) and describes autonomy as a
setting rather than a product tier, leading with the case that actually
motivated it: a scheduled agent that is finally policy-gated, journaled, and
able to stop and ask.
The store file was still named personal.db. It holds state for any agent
running unattended, not a species of agent — rename it, and note in place that
it is opened lazily so an ordinary conversational agent never grows one.
That claim is now a test rather than a comment: TestOrdinaryAgentGetsNoAutonomyStore
creates a plain agent, exercises the admin surface against it, and asserts no
agent.db appears — the governance machinery costs nothing until asked for.
Also covers the remaining orthogonality row: an agent WITH an objective but
WITHOUT autonomy still wakes on demand, reaching the policy gate rather than
being refused for lacking autonomy. Autonomy governs unprompted action, not
whether a human may ask.
All four combinations of (objective, autonomous) now have a test.
@timothyerwintimothyerwin changed the title Personal Agents: durable objectives, delegation policy, triggers, browser brokerAutonomous AgentsAug 30, 2026
@timothyerwin
timothyerwin merged commit 8a2e6b3 into mainAug 30, 2026
2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@timothyerwin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Autonomous Agents - #5

Merged
timothyerwin merged 13 commits into
mainfrom
feat/personal-agents
Aug 30, 2026
Merged

Autonomous Agents#5
timothyerwin merged 13 commits into
mainfrom
feat/personal-agents

Conversation

@timothyerwin

Copy link
Copy Markdown
Contributor

Summary

Adds Personal Agents: long-lived agent identities with a durable objective/subgoal/facts store, a delegation policy that governs what their executive loop may do (allowed tools, consequence classes, filesystem/browser/MCP resource grants, budgets), ask_user suspend/resume for human-in-the-loop interactions, durable cron/interval/one-shot triggers polled by the gateway, and a browser broker that leases exclusive browser control to one run at a time.

New surfaces: memcode personal CLI (create, approve-policy, resources, triggers, answer, status, cockpit), gw_agent/gw_* admin tools, and the gateway's internal personal wake route.

Review fixes included

This branch went through a review pass (/code-review high) before commit; the following were found and fixed:

  • Duplicate-answer replay: personal answer now refuses to re-resume an already-answered interaction, and a resumed run that re-suspends on a new question now marks its prior continuation resolved — closing a window where a retried personal answer call could re-execute real tool side effects.
  • Triggers silently dead after hot-add: the gateway now always registers the personal wake route and always starts the trigger poll loop (it already re-reads config every tick), so a Personal Agent created after boot has its triggers picked up without a gateway restart.
  • Expired browser lease still granted mutation rights: CanMutate/OwnPage now check lease expiry like Authenticate already did.
  • Resource grant ID collisions: IDs now use a monotonic timestamp instead of len(locator), matching the cockpit's existing scheme.
  • Delegation policy narrowing gaps: IsRestriction/NarrowPolicy now also checks filesystem roots, browser origins, MCP tools, budget fields, and quiet hours, not just tools/consequence-classes/concurrency/depth.
  • Silent no-op on bad trigger id: personal triggers pause/resume now errors instead of reporting success when nothing matched.
  • Stale tool description: gw_agent no longer advertises the unsupported "inspect" action.
  • Non-atomic suspension writes: suspension continuation files are now written via atomicfile.WriteFile.
  • Per-tick DB churn: the trigger poll loop now reuses one DB connection per agent across ticks and filters due triggers in SQL instead of opening/closing a connection and Go-filtering every row every 15s.

Testing

  • go build ./... — clean
  • go vet ./... — clean
  • go test ./... — all pass except two pre-existing, unrelated failures in internal/guard caused by a symlink under desktop/node_modules/@electron/notarize breaking go list (not touched by this branch)

… triggers, browser broker)
Implements the Personal Agents feature: long-lived agent identities with
their own durable objective, subgoals, and facts store; a delegation
policy model (allowed tools, consequence classes, filesystem/browser/MCP
grants, budgets) that governs what a Personal Agent's executive loop may
do; ask_user suspend/resume for human-in-the-loop interactions; durable
triggers (cron/interval/one-shot) polled by the gateway to wake an agent;
resource grants (filesystem, mcp, command, channel); and a browser broker
that leases exclusive browser control to one run at a time.
New surfaces: `memcode personal` CLI (create, approve-policy, resources,
triggers, answer, status, cockpit), gw_agent/gw_* admin tools, and the
gateway's internal personal wake route.
Includes fixes from review of this branch:
- personal answer refuses to re-resume an already-answered interaction,
and a resumed run that re-suspends now marks its prior continuation
resolved — closing a window where a retried answer could replay tool
side effects.
- The gateway always registers the personal wake route and always starts
the trigger poll loop, so a Personal Agent added after boot has its
triggers picked up without a restart.
- Browser broker leases now check expiry in CanMutate/OwnPage, matching
Authenticate, so an expired lease can no longer authorize mutation.
- Resource grant IDs use a monotonic timestamp instead of locator length,
avoiding same-length-locator collisions.
- Delegation policy narrowing (IsRestriction/NarrowPolicy) now also
checks filesystem roots, browser origins, MCP tools, budget fields, and
quiet hours, not just tools/consequence-classes/concurrency/depth.
- `personal triggers pause/resume` errors on an unknown trigger id instead
of reporting success silently.
- gw_agent's description no longer advertises the unsupported "inspect"
action.
- Suspension continuation files are written atomically.
- The trigger poll loop reuses one DB connection per agent across ticks
and filters due triggers in SQL, instead of opening/closing a
connection and scanning every trigger row every 15s.
… CI lint
Closes the biggest gap from review: the Personal Agent executive had no way
to reach browser/MCP/shell/filesystem capability — it was a second, weaker
agent runtime sitting beside the real memcode engine instead of commanding
it. Two new executive tools fix that:
- delegate: spawns a policy-scoped worker as a real detached `memcode run`
job (internal/jobs.SpawnWithSpec), using the delegation.go scaffolding
(ExecutionEnvelope, ValidateDelegation, PrepareRunDirectory) that already
existed but was never called from anywhere. The worker gets whatever
toolsets it's granted — browser, MCP, shell, filesystem, skills — same as
any ordinary memcode agent, not the executive's fixed 7 tools. Requested
toolsets/consequences must be a subset of the agent's own approved policy;
MaxDelegationDepth gates whether delegation is allowed at all.
- check_delegate: reads a delegated job's status/result back on a later
wake (RunOnce is one bounded wake, so the result can't arrive in the same
call) and closes out its action-journal entry.
Also makes jobs.SpawnSpec.ToolPolicy a REAL restriction instead of recorded-
only metadata: new hidden --allow-tools/--deny-tools flags on `memcode run`
bind to the same SetToolPolicy enforcement an ordinary gateway-bound agent
gets from its config, and SpawnWithSpec now passes them through argv.
Scope notes, called out explicitly rather than implied as done:
- "browser" toolset reuses the existing ephemeral Chrome session (the same
one --chrome already drives) — NOT the user's own already-running Chrome.
Attaching to a persistent/remote Chrome needs the broker/RemoteConfig
wiring in internal/browser, which is declared but has zero callers and
zero implementation; that's separate follow-up work, not done here.
- Filesystem/MCP resource grants are not OS-sandboxed to a spawned worker
(memcode has no such sandbox anywhere) — governance here is the same
approval/audit model as the rest of the codebase (tool-level allow/deny +
the action journal), not a hard jail.
Also fixes two CI lint failures on the open PR:
- internal/gateway/state/state.go: removed formatTime, an unused helper
(staticcheck U1000).
- internal/personal/store.go: decapitalized a Go error string (ST1005).
Corrects course on the browser gap flagged after the last review: existing-
Chrome access is core to Personal Agents, not deferrable follow-up. A
Personal Agent's whole premise — acting as the user across their real
accounts (Gmail, LinkedIn, an ATS, internal dashboards) — requires the
user's actual signed-in session, not a fresh ephemeral profile with no
cookies. Per docs/design/personal-agents.md's already-specified "browser
broker trust boundary": Personal Agents connect to the user's existing
Chrome through a gateway-owned broker and a permission-protected local
socket, fail closed on any connection/auth problem, and never silently fall
back to another profile.
- internal/browser/broker/{server,client}.go: exposes the (already-reviewed)
in-process Broker over a Unix socket, so a delegated worker — a SEPARATE
OS process spawned via jobs.SpawnWithSpec — can Acquire/Release/OwnPage/
CanMutate against the SAME broker the gateway owns, not a fresh one.
Socket is 0600 inside gwconfig.Dir() (0700). Round-trip tested, including
that a concurrent Acquire is correctly rejected and a release unblocks the
next one.
- internal/gateway/server: the gateway now owns one *broker.Broker for its
whole lifetime and serves it on that socket unconditionally at startup
(cheap — a local listener), so it's there the moment a delegate call needs
it, no restart required after `memcode personal browser setup`.
- internal/personal: browserModeFor's "browser" toolset now defaults to
BrowserExistingChrome, not ephemeral — "browser:ephemeral" is the explicit
opt-down for tasks that genuinely want a fresh, logged-out profile. Before
spawning anything, the delegate tool verifies the broker is reachable and
fails the tool call closed (a normal tool_result error, not a spawned job)
if it isn't — verified by test, alongside the success path asserting the
spawned job's SpawnSpec actually carries BrowserMode=existing_chrome.
- internal/jobs: SpawnWithSpec passes existing_chrome through as
--browser-session (+ --browser-agent/--browser-run for the lease identity)
instead of --chrome.
- cmd/run.go: --browser-session existing_chrome acquires the broker lease
itself (via the client) and, only on success, wires an ad-hoc
chrome-devtools-mcp --autoConnect MCP server into the session
(Session.SetExtraMCPServers, new) — the SAME chrome-devtools-mcp package
this browser experience is built on (see internal/browser/remote.go). On
any failure it returns an error and exits; it never calls
SetBrowserEnabled(true) (ephemeral) as a fallback.
- cmd/personal_browser.go: `memcode personal browser setup` checks npx and
the broker socket, then attempts a REAL bounded (10s) chrome-devtools-mcp
connection and reports the tool count or the actual connect errors — it
does not pretend to click Chrome's own "Allow" dialog, which only the user
can do.
- Also fixes a real bug found while testing this: ValidateDelegation treated
an empty parent.AllowedTools as "allows no tools by name", the opposite of
the len(...)>0-means-restricted convention Executive.allowedTools already
uses — every delegate call with a name-unrestricted policy was rejected.
Scope note, stated rather than implied: the literal Chrome-side consent flow
(the "Allow" dialog, actually clicking it) cannot be exercised or verified
in this environment — no live Chrome, no human to click it. Everything
up to that point (broker lifecycle, lease exclusivity, fail-closed wiring,
the actual chrome-devtools-mcp CLI invocation, verified against the real
published package) is built and tested; the dialog itself needs the user's
own machine.
…setup via the cockpit
Three separate complaints from review, addressed at the right layer each:
1. "It's just a file ref" — `resources add` required a redundant --mode read
(already the default) and a mandatory <type> argument even for the
overwhelming common case. A bare path is now enough:
memcode personal resources add jobhunt ~/resume.md
Non-filesystem grants (mcp/command/channel) still need the type spelled
out, since there's nothing to infer it from. Also fixes the real bug this
surfaced: CanonicalFilesystemGrant rejected any non-directory path, even
though PathWithinGrant already handles a single-file grant correctly
(rel == "." on an exact match) — you could not actually grant a single
file like a resume before this, only whole directories. `personal create`
also gained a repeatable --grant flag to fold "create it, then give it
something to read" into one command.
2. "Not this CLI-command-argument garbage" — Personal Agent state (objective,
policy, resources) lived ONLY in personal.db, reachable exclusively
through bespoke commands, unlike every other piece of memcode config
(gateway.yaml, .mcp.json, CLAUDE.md) which is a plain file. WriteConfigMirror
now regenerates objective.md/policy.yaml/resources.yaml in the agent's home
on every mutation — real, readable, diffable files. The run journal
(actions/triggers/interactions) deliberately stays in SQLite: it needs
atomic claim/complete semantics under concurrent access (gateway wake loop,
CLI, and cockpit can all touch the same agent) that flat files don't get
for free — a correctness reason, not a habit.
3. "It should gather requirements, do HITL, pre-approve permissions, set up
the runtime — like our other agents already do via telegram/model setup"
— `memcode personal` (no args) is ALREADY the interactive cockpit (like
`memcode admin`) with the pa_* tools wired for exactly this. The actual gap
was its system prompt: personalAdminDoctrine listed tool mechanics with no
instruction to gather-then-propose-then-approve. Rewrote it to require a
real walkthrough for first-time creation — reason about what resources/
toolsets/consequences the objective needs (asking the user rather than
guessing), present the concrete plan in plain language BEFORE touching
anything, then apply everything at once on approval, including wake
cadence (trigger) — an agent with no trigger and no plan to ever be woken
is dead on arrival, so cadence is not an optional afterthought either.
A prior pass at this (adding a NEW deterministic `memcode personal setup`
CLI wizard, mirroring `gateway setup`'s bufio-prompt mechanism) was wrong
and reverted before commit — that reproduces the exact CLI-surface-
sprawl complaint instead of fixing it. The cockpit conversation IS the
wizard; it needed a better prompt, not a new command.
Root cause of the whole CLI-surface complaint: the interactive cockpit
(`memcode personal`, no args) is supposed to be the entire interface, but
its tool set had NO way to create a new Personal Agent. pa_objective only
supported show/set, and set required the objective to already exist —
creation was only reachable through the CLI's `personal create`. That's why
every explanation kept surfacing a CLI command: until now, it was the only
path that actually worked.
- internal/agent/tools/personal.go: new pa_create tool (agent name +
objective). Documented as the required first call for a brand-new agent.
- cmd/personal_cockpit.go: paCreate mirrors personalCreate's steps (register
in gateway.yaml, open the store, create the objective) and is dispatched
BEFORE the paStore existence gate, since the agent doesn't exist yet.
pa_resource's grant action now defaults type=filesystem/mode=read when
omitted too (same inference the CLI's `resources add` already got), and
pa_resource/pa_policy's grant/revoke/stage/approve paths now call
WriteConfigMirror — that was wired into the CLI paths last commit but not
into the cockpit paths, which are the ones actually used.
- cmd/personal.go: every `memcode personal` subcommand is now Hidden — they
remain fully callable for scripts, but `memcode personal --help` no longer
reads like a CLI to memorize. The interface is `memcode personal` with no
args, then plain conversation.
- internal/doctrine/prompts.go: the cockpit's own system prompt now says
outright never to tell the user to run a CLI command — it does the work
itself, right there — and that pa_create, not pa_objective, is the first
call for a new agent.
Verified end to end: pa_create then a bare-path pa_resource grant (no type,
no mode) both succeed through personalExecute exactly as the cockpit would
call them, no CLI involved.
Hiding the subcommands last commit wasn't enough — they still existed as a
parallel, scriptable way to do everything, which is real duplicate logic
(and already caused drift: a fix landed in one path and not the other,
twice, in this same review). Deleted them outright:
cmd/personal_resources.go, personal_policy.go, personal_triggers.go,
personal_status.go, personal_browser.go, personal_test.go — all removed.
cmd/personal.go cut down to just the cockpit launcher (~90 lines from
~370): `memcode personal` opens the interactive session, full stop.
Two operations the CLI covered had no pa_* equivalent yet, so they'd have
been silently lost:
- pa_doctor: the health check (home layout, objective, approved policy,
generated workspace, sandbox availability, trigger/pending-interaction
counts) that `personal doctor` used to run.
- pa_browser_setup: the existing-Chrome prerequisite check + real, bounded
chrome-devtools-mcp connection attempt that `personal browser setup` used
to run. Same logic, now a tool call instead of a command — it still
doesn't (can't) click Chrome's own consent dialog for the user.
Both dispatch through personalExecute exactly like every other pa_* tool.
internal/agent/tools/personal.go: added PaDoctor and PaBrowserSetup to the
registry.
cmd/personal_cmd_test.go: rewritten from scratch to call personalExecute
directly — no CLI args, no cobra Execute(), no rootCmd. This IS the
interface being tested now, so this is what the tests exercise. Also
asserts the config-mirror files (policy.yaml, resources.yaml) actually
reflect approvals/revokes done through the tool path, since that's the one
that matters.
Verified: `memcode personal create foo bar` no longer resolves to anything
— it just opens the cockpit with those words ignored as stray args, because
there is no "create" subcommand left to match. The interface is
conversation with pa_* tools behind it, not CLI argument syntax.
objective.md/policy.yaml/resources.yaml split one agent's config across
three files by table for no real reason — a person has to know to look in
three places for one agent's setup. Consolidated into a single config.yaml
(objective, policies, resources as sections) with proper snake_case field
names. Same authority split as before (this is a mirror; policy hash-
approval and the run journal stay in personal.db, and why, is unchanged) —
just one file to read instead of three.
Consolidation step 1 of folding Personal Agents into the ordinary agent
system. There were three partial suspend/resume designs in the tree:
1. internal/agent/runtime/continuation.go — typed, atomic, tested, and with
ZERO production callers.
2. jobs.Job's InteractionID/WaitingReason/ContinuationVersion/WaitingAt/
ResumedAt — declared, never written by anything.
3. A hand-rolled map[string]any in the personal executive — the only one
actually running, and (until a fix earlier in this branch) the only one
that wasn't crash-safe.
New internal/agent/continuation is the single implementation. It lives in its
own small package rather than in internal/agent/runtime so the executive can
suspend without importing the whole session runtime (and so neither side ends
up depending on the other).
It has to serve two genuinely different callers, which is why the previous
attempt at a shared API didn't fit:
- An interactive session already holds the conversation, so it only needs the
missing assistant/answer pair back.
- An unattended executive keeps NO transcript — it rebuilds context from
durable state each wake — so its continuation must carry the whole
conversation. Hence the optional Messages field.
They also differ on when marking resolved is safe, so building the resume
messages is split from marking:
- ResumeMessages() builds without marking. The executive uses this and marks
only once the resumed run reaches a terminal state, so a transient model
error leaves the answer re-giveable instead of stranding it.
- Resolve() does both, for a caller with a human right there who can simply
ask again.
Also deleted the five never-written Job fields (2 above). StatusWaiting stays:
it is a meaningful status in the job state machine and check_delegate
correctly treats it as non-terminal.
Tests: ported both original round-trip/validation tests, plus new coverage for
the transcript-carrying path and MarkResolved. The executive's suspend/resume
test now asserts through the continuation API instead of a hardcoded filename.
… admin
Consolidation step 2. "Personal Agent" is no longer a kind of agent. There is
one Agent abstraction; autonomy is orthogonal settings on it, managed through
the admin cockpit that already manages agents.
CONFIG. gwconfig.Agent drops Kind and gains four settings:
objective — the durable outcome it works toward
autonomous — whether it may act on that unprompted
browser — ephemeral (default) | existing_chrome
paused — stop unattended wakes without deleting anything
objective and autonomous are deliberately SEPARATE grants, which is the
correction that motivated this design. They answer different questions ("what
is it for" vs "may it act unasked"), and all four combinations are meaningful:
autonomous + objective → unattended objective pursuit (the old Personal Agent)
autonomous, no objective→ scheduled work UNDER GOVERNANCE — this is new, and
closes the long-standing gap where a cron-fired
agent ran unattended with no policy gate, no action
journal, and no way to pause and ask
objective, not autonomous → a goal you work on together; wakes only on demand
neither → an ordinary conversational agent, exactly as before
Autonomy therefore gates GOVERNANCE, not capability, and applies to any run of
the agent whether or not an objective exists.
SCHEDULING. The second cron implementation is gone as a user-facing concept.
Recurring cadence is an ordinary `schedules:` entry: set agent=<name>, leave
deliver_to empty, and gw_schedule routes the wake to the agent itself via the
renamed internal sink (personal → agent). The DB-backed loop now fires only
the agent's OWN self-scheduled next-wakes (schedule_wake, "come back in 45
minutes") — a genuinely different thing from human-authored cadence, and the
only part that needs to be writable from inside a run.
TOOLS. internal/agent/tools/personal.go (pa_*) is deleted. The admin registry
gains gw_policy, gw_grant, gw_wake, gw_inbox, gw_answer, gw_journal, gw_doctor,
gw_browser; gw_agent gains objective/autonomous/browser/pause/resume actions.
Everything dispatches through the existing adminExecute + approval gate, and
gw_schedule's shared BuildSchedule validation now covers agent cadence too, so
the surfaces cannot drift the way the parallel path did.
COCKPIT. `memcode personal` is deleted outright — cmd/personal.go,
cmd/personal_cockpit.go, Session.SetPersonal, personalMode, the personal and
personal_admin doctrines. The Personal setup walkthrough is merged into
adminDoctrine, generalized, and now teaches the two gates explicitly: granting
an objective is not granting autonomy, and the second must be confirmed on its
own.
EXECUTIVE. Executive.Objective is read from configuration rather than the
store, so the objective has ONE source that a human edits and the gateway
hot-reloads. The per-agent config.yaml mirror drops the objective for the same
reason — it lives in gateway.yaml, which is already a readable file.
Tests are rewritten against adminExecute (there is no other surface). New
coverage asserts the orthogonality directly, since a single overloaded switch
is exactly what this change exists to prevent: an objective alone must not
confer autonomy, a non-affirmative value must not grant it by typo, an
autonomous agent with no objective is still governed but refuses to invent
work, and an autonomous agent's schedule defaults to the agent route.
…ard both
Consolidation step 3, finishing the merge.
RENAME. internal/personal -> internal/agent/autonomy. The package was never
about a "personal" species; it is the machinery an agent uses when running
unattended toward an objective. Terminology in comments follows.
SECOND CRON PARSER DELETED. NextDue now understands only the kinds an AGENT
writes for itself from inside a run (one_shot / next_wake — always a single
future instant, e.g. schedule_wake "come back in 45 minutes"). Recurring
cadence a human configures is not a trigger at all: it is an ordinary gateway
schedule delivering to agent:<name>, validated once by gwconfig. With
recurring kinds gone, ClaimDueTrigger's reschedule branch was unreachable, so
firing a wake now simply completes it — and the atomic `last_fired_at IS ?`
claim still keeps two gateway processes from double-firing.
GUARDS (internal/guard/singletons_test.go). The duplication this branch kept
producing was not a one-off: fixes landed on one path and not the other,
repeatedly. Four invariants now fail the build instead:
- TestSingleCronParser — only internal/gateway parses cron. (This one caught a
real leftover while being written: the autonomy package was still importing
robfig/cron, which is what prompted the deletion above.)
- TestSingleSuspensionImplementation — no hand-rolled continuation formats
outside internal/agent/continuation.
- TestNoSecondCockpit — no SetPersonal/personalMode/pa_* anywhere.
- TestNoAgentKind — autonomy is orthogonal settings, never a kind
discriminator; a "kind" field is what made Personal a separate species.
The guards deliberately skip vendored forks, .memcode session snapshots, and
desktop/node_modules (the same symlink that breaks the two pre-existing
go-list guard failures).
Store tests updated: the trigger test now exercises a real self-scheduled wake
and asserts it completes rather than reschedules, and NextDue's test asserts
that interval/cron are REJECTED — accepting them again would rebuild the
second scheduler.
Consolidation step 4. An agent had two unrelated places to put what it knew:
a structured `facts` table only reachable through bespoke tools, and the
memory.md every memcode agent already has and the runtime already injects. One
of those a human can read. Collapse to it.
BLOCKING PREREQUISITE, done first. delegate/check_delegate were using the
facts table as a keyed index — writing "delegation.<job-id>" so a later wake
could find the action to close out. Its own comment admitted why ("facts are
the only durable log delegate can write to without a schema migration"). This
is that migration: actions gain a job_id column plus LinkActionJob /
ActionForJob, and the link now lives on the action it actually describes.
Removing facts before this would have broken check_delegate silently.
note_fact becomes `remember`: one plain sentence appended to memory.md, read
back into every wake's state summary. Append-only and deduplicated, because
every line replays into the model on each future wake — a re-learned fact must
not grow the file forever.
Deleted with it: environment.go (EnvironmentModel/StructuredFact, referenced
nowhere but its own test) and UsableForExternalRepresentation.
The tradeoff is deliberate and documented in memory.go rather than left
implicit. Prose cannot distinguish "you told me this" from "I inferred it from
your resume" from "a website said so", nor mark a claim safe to assert on the
user's behalf, nor mark it stale. That distinction becomes load-bearing the
moment an agent fills in a form or sends a message stating something about the
user. It is cheap to drop TODAY only because the gate was never wired —
nothing read Confirmed to gate anything; the sole caller of
UsableForExternalRepresentation was its own test. When external
representation needs provenance, it should come back as its own
machine-checkable thing in the store alongside policies and the action
journal, NOT by reviving this table.
… as a mode
Consolidation step 5, finishing the merge.
MIGRATION. YAML silently ignores unknown fields, so an existing
`kind: personal` agent would have loaded as an ordinary one — still listed,
apparently fine, and never waking again. For a change that removes an agent's
authority to run on its own, silence is the wrong failure. gwconfig.Agent keeps
a LegacyKind field for the sole purpose of rejecting it, with the one-line fix
in the error:
agent "demo" still uses the removed `kind: personal` setting. Autonomy is now
explicit: replace it with `autonomous: true` (and an `objective:` ...), or
just delete the `kind:` line if it should not. Its home ... is untouched
Verified against the real local config, which has exactly that agent. The
guard's TestNoAgentKind is refined to allow this one mention while still
banning any behavioral branch on it.
DOCS. docs/personal-agents.md -> docs/autonomous-agents.md, rewritten around
the two orthogonal grants and the four combinations they produce, with the
scheduling split (your cadence = an ordinary schedule; the agent's own next
wake = self-scheduled), the fail-closed existing-Chrome rule, and the
memory.md provenance limitation stated rather than glossed.
docs/design/personal-agents.md -> docs/design/autonomous-agents.md keeps the
original design contract but opens with a note on what the implementation
revised and why, including the orthogonality correction. It is a design record;
rewriting history there would lose the reasoning.
README drops "three ways to run it" (there are two) and describes autonomy as a
setting rather than a product tier, leading with the case that actually
motivated it: a scheduled agent that is finally policy-gated, journaled, and
able to stop and ask.
The store file was still named personal.db. It holds state for any agent
running unattended, not a species of agent — rename it, and note in place that
it is opened lazily so an ordinary conversational agent never grows one.
That claim is now a test rather than a comment: TestOrdinaryAgentGetsNoAutonomyStore
creates a plain agent, exercises the admin surface against it, and asserts no
agent.db appears — the governance machinery costs nothing until asked for.
Also covers the remaining orthogonality row: an agent WITH an objective but
WITHOUT autonomy still wakes on demand, reaching the policy gate rather than
being refused for lacking autonomy. Autonomy governs unprompted action, not
whether a human may ask.
All four combinations of (objective, autonomous) now have a test.
@timothyerwintimothyerwin changed the title Personal Agents: durable objectives, delegation policy, triggers, browser brokerAutonomous AgentsAug 30, 2026
@timothyerwin
timothyerwin merged commit 8a2e6b3 into mainAug 30, 2026
2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@timothyerwin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Autonomous Agents - #5

Merged
timothyerwin merged 13 commits into
mainfrom
feat/personal-agents
Aug 30, 2026
Merged

Autonomous Agents#5
timothyerwin merged 13 commits into
mainfrom
feat/personal-agents

Conversation

@timothyerwin

Copy link
Copy Markdown
Contributor

Summary

Adds Personal Agents: long-lived agent identities with a durable objective/subgoal/facts store, a delegation policy that governs what their executive loop may do (allowed tools, consequence classes, filesystem/browser/MCP resource grants, budgets), ask_user suspend/resume for human-in-the-loop interactions, durable cron/interval/one-shot triggers polled by the gateway, and a browser broker that leases exclusive browser control to one run at a time.

New surfaces: memcode personal CLI (create, approve-policy, resources, triggers, answer, status, cockpit), gw_agent/gw_* admin tools, and the gateway's internal personal wake route.

Review fixes included

This branch went through a review pass (/code-review high) before commit; the following were found and fixed:

  • Duplicate-answer replay: personal answer now refuses to re-resume an already-answered interaction, and a resumed run that re-suspends on a new question now marks its prior continuation resolved — closing a window where a retried personal answer call could re-execute real tool side effects.
  • Triggers silently dead after hot-add: the gateway now always registers the personal wake route and always starts the trigger poll loop (it already re-reads config every tick), so a Personal Agent created after boot has its triggers picked up without a gateway restart.
  • Expired browser lease still granted mutation rights: CanMutate/OwnPage now check lease expiry like Authenticate already did.
  • Resource grant ID collisions: IDs now use a monotonic timestamp instead of len(locator), matching the cockpit's existing scheme.
  • Delegation policy narrowing gaps: IsRestriction/NarrowPolicy now also checks filesystem roots, browser origins, MCP tools, budget fields, and quiet hours, not just tools/consequence-classes/concurrency/depth.
  • Silent no-op on bad trigger id: personal triggers pause/resume now errors instead of reporting success when nothing matched.
  • Stale tool description: gw_agent no longer advertises the unsupported "inspect" action.
  • Non-atomic suspension writes: suspension continuation files are now written via atomicfile.WriteFile.
  • Per-tick DB churn: the trigger poll loop now reuses one DB connection per agent across ticks and filters due triggers in SQL instead of opening/closing a connection and Go-filtering every row every 15s.

Testing

  • go build ./... — clean
  • go vet ./... — clean
  • go test ./... — all pass except two pre-existing, unrelated failures in internal/guard caused by a symlink under desktop/node_modules/@electron/notarize breaking go list (not touched by this branch)

… triggers, browser broker)
Implements the Personal Agents feature: long-lived agent identities with
their own durable objective, subgoals, and facts store; a delegation
policy model (allowed tools, consequence classes, filesystem/browser/MCP
grants, budgets) that governs what a Personal Agent's executive loop may
do; ask_user suspend/resume for human-in-the-loop interactions; durable
triggers (cron/interval/one-shot) polled by the gateway to wake an agent;
resource grants (filesystem, mcp, command, channel); and a browser broker
that leases exclusive browser control to one run at a time.
New surfaces: `memcode personal` CLI (create, approve-policy, resources,
triggers, answer, status, cockpit), gw_agent/gw_* admin tools, and the
gateway's internal personal wake route.
Includes fixes from review of this branch:
- personal answer refuses to re-resume an already-answered interaction,
and a resumed run that re-suspends now marks its prior continuation
resolved — closing a window where a retried answer could replay tool
side effects.
- The gateway always registers the personal wake route and always starts
the trigger poll loop, so a Personal Agent added after boot has its
triggers picked up without a restart.
- Browser broker leases now check expiry in CanMutate/OwnPage, matching
Authenticate, so an expired lease can no longer authorize mutation.
- Resource grant IDs use a monotonic timestamp instead of locator length,
avoiding same-length-locator collisions.
- Delegation policy narrowing (IsRestriction/NarrowPolicy) now also
checks filesystem roots, browser origins, MCP tools, budget fields, and
quiet hours, not just tools/consequence-classes/concurrency/depth.
- `personal triggers pause/resume` errors on an unknown trigger id instead
of reporting success silently.
- gw_agent's description no longer advertises the unsupported "inspect"
action.
- Suspension continuation files are written atomically.
- The trigger poll loop reuses one DB connection per agent across ticks
and filters due triggers in SQL, instead of opening/closing a
connection and scanning every trigger row every 15s.
… CI lint
Closes the biggest gap from review: the Personal Agent executive had no way
to reach browser/MCP/shell/filesystem capability — it was a second, weaker
agent runtime sitting beside the real memcode engine instead of commanding
it. Two new executive tools fix that:
- delegate: spawns a policy-scoped worker as a real detached `memcode run`
job (internal/jobs.SpawnWithSpec), using the delegation.go scaffolding
(ExecutionEnvelope, ValidateDelegation, PrepareRunDirectory) that already
existed but was never called from anywhere. The worker gets whatever
toolsets it's granted — browser, MCP, shell, filesystem, skills — same as
any ordinary memcode agent, not the executive's fixed 7 tools. Requested
toolsets/consequences must be a subset of the agent's own approved policy;
MaxDelegationDepth gates whether delegation is allowed at all.
- check_delegate: reads a delegated job's status/result back on a later
wake (RunOnce is one bounded wake, so the result can't arrive in the same
call) and closes out its action-journal entry.
Also makes jobs.SpawnSpec.ToolPolicy a REAL restriction instead of recorded-
only metadata: new hidden --allow-tools/--deny-tools flags on `memcode run`
bind to the same SetToolPolicy enforcement an ordinary gateway-bound agent
gets from its config, and SpawnWithSpec now passes them through argv.
Scope notes, called out explicitly rather than implied as done:
- "browser" toolset reuses the existing ephemeral Chrome session (the same
one --chrome already drives) — NOT the user's own already-running Chrome.
Attaching to a persistent/remote Chrome needs the broker/RemoteConfig
wiring in internal/browser, which is declared but has zero callers and
zero implementation; that's separate follow-up work, not done here.
- Filesystem/MCP resource grants are not OS-sandboxed to a spawned worker
(memcode has no such sandbox anywhere) — governance here is the same
approval/audit model as the rest of the codebase (tool-level allow/deny +
the action journal), not a hard jail.
Also fixes two CI lint failures on the open PR:
- internal/gateway/state/state.go: removed formatTime, an unused helper
(staticcheck U1000).
- internal/personal/store.go: decapitalized a Go error string (ST1005).
Corrects course on the browser gap flagged after the last review: existing-
Chrome access is core to Personal Agents, not deferrable follow-up. A
Personal Agent's whole premise — acting as the user across their real
accounts (Gmail, LinkedIn, an ATS, internal dashboards) — requires the
user's actual signed-in session, not a fresh ephemeral profile with no
cookies. Per docs/design/personal-agents.md's already-specified "browser
broker trust boundary": Personal Agents connect to the user's existing
Chrome through a gateway-owned broker and a permission-protected local
socket, fail closed on any connection/auth problem, and never silently fall
back to another profile.
- internal/browser/broker/{server,client}.go: exposes the (already-reviewed)
in-process Broker over a Unix socket, so a delegated worker — a SEPARATE
OS process spawned via jobs.SpawnWithSpec — can Acquire/Release/OwnPage/
CanMutate against the SAME broker the gateway owns, not a fresh one.
Socket is 0600 inside gwconfig.Dir() (0700). Round-trip tested, including
that a concurrent Acquire is correctly rejected and a release unblocks the
next one.
- internal/gateway/server: the gateway now owns one *broker.Broker for its
whole lifetime and serves it on that socket unconditionally at startup
(cheap — a local listener), so it's there the moment a delegate call needs
it, no restart required after `memcode personal browser setup`.
- internal/personal: browserModeFor's "browser" toolset now defaults to
BrowserExistingChrome, not ephemeral — "browser:ephemeral" is the explicit
opt-down for tasks that genuinely want a fresh, logged-out profile. Before
spawning anything, the delegate tool verifies the broker is reachable and
fails the tool call closed (a normal tool_result error, not a spawned job)
if it isn't — verified by test, alongside the success path asserting the
spawned job's SpawnSpec actually carries BrowserMode=existing_chrome.
- internal/jobs: SpawnWithSpec passes existing_chrome through as
--browser-session (+ --browser-agent/--browser-run for the lease identity)
instead of --chrome.
- cmd/run.go: --browser-session existing_chrome acquires the broker lease
itself (via the client) and, only on success, wires an ad-hoc
chrome-devtools-mcp --autoConnect MCP server into the session
(Session.SetExtraMCPServers, new) — the SAME chrome-devtools-mcp package
this browser experience is built on (see internal/browser/remote.go). On
any failure it returns an error and exits; it never calls
SetBrowserEnabled(true) (ephemeral) as a fallback.
- cmd/personal_browser.go: `memcode personal browser setup` checks npx and
the broker socket, then attempts a REAL bounded (10s) chrome-devtools-mcp
connection and reports the tool count or the actual connect errors — it
does not pretend to click Chrome's own "Allow" dialog, which only the user
can do.
- Also fixes a real bug found while testing this: ValidateDelegation treated
an empty parent.AllowedTools as "allows no tools by name", the opposite of
the len(...)>0-means-restricted convention Executive.allowedTools already
uses — every delegate call with a name-unrestricted policy was rejected.
Scope note, stated rather than implied: the literal Chrome-side consent flow
(the "Allow" dialog, actually clicking it) cannot be exercised or verified
in this environment — no live Chrome, no human to click it. Everything
up to that point (broker lifecycle, lease exclusivity, fail-closed wiring,
the actual chrome-devtools-mcp CLI invocation, verified against the real
published package) is built and tested; the dialog itself needs the user's
own machine.
…setup via the cockpit
Three separate complaints from review, addressed at the right layer each:
1. "It's just a file ref" — `resources add` required a redundant --mode read
(already the default) and a mandatory <type> argument even for the
overwhelming common case. A bare path is now enough:
memcode personal resources add jobhunt ~/resume.md
Non-filesystem grants (mcp/command/channel) still need the type spelled
out, since there's nothing to infer it from. Also fixes the real bug this
surfaced: CanonicalFilesystemGrant rejected any non-directory path, even
though PathWithinGrant already handles a single-file grant correctly
(rel == "." on an exact match) — you could not actually grant a single
file like a resume before this, only whole directories. `personal create`
also gained a repeatable --grant flag to fold "create it, then give it
something to read" into one command.
2. "Not this CLI-command-argument garbage" — Personal Agent state (objective,
policy, resources) lived ONLY in personal.db, reachable exclusively
through bespoke commands, unlike every other piece of memcode config
(gateway.yaml, .mcp.json, CLAUDE.md) which is a plain file. WriteConfigMirror
now regenerates objective.md/policy.yaml/resources.yaml in the agent's home
on every mutation — real, readable, diffable files. The run journal
(actions/triggers/interactions) deliberately stays in SQLite: it needs
atomic claim/complete semantics under concurrent access (gateway wake loop,
CLI, and cockpit can all touch the same agent) that flat files don't get
for free — a correctness reason, not a habit.
3. "It should gather requirements, do HITL, pre-approve permissions, set up
the runtime — like our other agents already do via telegram/model setup"
— `memcode personal` (no args) is ALREADY the interactive cockpit (like
`memcode admin`) with the pa_* tools wired for exactly this. The actual gap
was its system prompt: personalAdminDoctrine listed tool mechanics with no
instruction to gather-then-propose-then-approve. Rewrote it to require a
real walkthrough for first-time creation — reason about what resources/
toolsets/consequences the objective needs (asking the user rather than
guessing), present the concrete plan in plain language BEFORE touching
anything, then apply everything at once on approval, including wake
cadence (trigger) — an agent with no trigger and no plan to ever be woken
is dead on arrival, so cadence is not an optional afterthought either.
A prior pass at this (adding a NEW deterministic `memcode personal setup`
CLI wizard, mirroring `gateway setup`'s bufio-prompt mechanism) was wrong
and reverted before commit — that reproduces the exact CLI-surface-
sprawl complaint instead of fixing it. The cockpit conversation IS the
wizard; it needed a better prompt, not a new command.
Root cause of the whole CLI-surface complaint: the interactive cockpit
(`memcode personal`, no args) is supposed to be the entire interface, but
its tool set had NO way to create a new Personal Agent. pa_objective only
supported show/set, and set required the objective to already exist —
creation was only reachable through the CLI's `personal create`. That's why
every explanation kept surfacing a CLI command: until now, it was the only
path that actually worked.
- internal/agent/tools/personal.go: new pa_create tool (agent name +
objective). Documented as the required first call for a brand-new agent.
- cmd/personal_cockpit.go: paCreate mirrors personalCreate's steps (register
in gateway.yaml, open the store, create the objective) and is dispatched
BEFORE the paStore existence gate, since the agent doesn't exist yet.
pa_resource's grant action now defaults type=filesystem/mode=read when
omitted too (same inference the CLI's `resources add` already got), and
pa_resource/pa_policy's grant/revoke/stage/approve paths now call
WriteConfigMirror — that was wired into the CLI paths last commit but not
into the cockpit paths, which are the ones actually used.
- cmd/personal.go: every `memcode personal` subcommand is now Hidden — they
remain fully callable for scripts, but `memcode personal --help` no longer
reads like a CLI to memorize. The interface is `memcode personal` with no
args, then plain conversation.
- internal/doctrine/prompts.go: the cockpit's own system prompt now says
outright never to tell the user to run a CLI command — it does the work
itself, right there — and that pa_create, not pa_objective, is the first
call for a new agent.
Verified end to end: pa_create then a bare-path pa_resource grant (no type,
no mode) both succeed through personalExecute exactly as the cockpit would
call them, no CLI involved.
Hiding the subcommands last commit wasn't enough — they still existed as a
parallel, scriptable way to do everything, which is real duplicate logic
(and already caused drift: a fix landed in one path and not the other,
twice, in this same review). Deleted them outright:
cmd/personal_resources.go, personal_policy.go, personal_triggers.go,
personal_status.go, personal_browser.go, personal_test.go — all removed.
cmd/personal.go cut down to just the cockpit launcher (~90 lines from
~370): `memcode personal` opens the interactive session, full stop.
Two operations the CLI covered had no pa_* equivalent yet, so they'd have
been silently lost:
- pa_doctor: the health check (home layout, objective, approved policy,
generated workspace, sandbox availability, trigger/pending-interaction
counts) that `personal doctor` used to run.
- pa_browser_setup: the existing-Chrome prerequisite check + real, bounded
chrome-devtools-mcp connection attempt that `personal browser setup` used
to run. Same logic, now a tool call instead of a command — it still
doesn't (can't) click Chrome's own consent dialog for the user.
Both dispatch through personalExecute exactly like every other pa_* tool.
internal/agent/tools/personal.go: added PaDoctor and PaBrowserSetup to the
registry.
cmd/personal_cmd_test.go: rewritten from scratch to call personalExecute
directly — no CLI args, no cobra Execute(), no rootCmd. This IS the
interface being tested now, so this is what the tests exercise. Also
asserts the config-mirror files (policy.yaml, resources.yaml) actually
reflect approvals/revokes done through the tool path, since that's the one
that matters.
Verified: `memcode personal create foo bar` no longer resolves to anything
— it just opens the cockpit with those words ignored as stray args, because
there is no "create" subcommand left to match. The interface is
conversation with pa_* tools behind it, not CLI argument syntax.
objective.md/policy.yaml/resources.yaml split one agent's config across
three files by table for no real reason — a person has to know to look in
three places for one agent's setup. Consolidated into a single config.yaml
(objective, policies, resources as sections) with proper snake_case field
names. Same authority split as before (this is a mirror; policy hash-
approval and the run journal stay in personal.db, and why, is unchanged) —
just one file to read instead of three.
Consolidation step 1 of folding Personal Agents into the ordinary agent
system. There were three partial suspend/resume designs in the tree:
1. internal/agent/runtime/continuation.go — typed, atomic, tested, and with
ZERO production callers.
2. jobs.Job's InteractionID/WaitingReason/ContinuationVersion/WaitingAt/
ResumedAt — declared, never written by anything.
3. A hand-rolled map[string]any in the personal executive — the only one
actually running, and (until a fix earlier in this branch) the only one
that wasn't crash-safe.
New internal/agent/continuation is the single implementation. It lives in its
own small package rather than in internal/agent/runtime so the executive can
suspend without importing the whole session runtime (and so neither side ends
up depending on the other).
It has to serve two genuinely different callers, which is why the previous
attempt at a shared API didn't fit:
- An interactive session already holds the conversation, so it only needs the
missing assistant/answer pair back.
- An unattended executive keeps NO transcript — it rebuilds context from
durable state each wake — so its continuation must carry the whole
conversation. Hence the optional Messages field.
They also differ on when marking resolved is safe, so building the resume
messages is split from marking:
- ResumeMessages() builds without marking. The executive uses this and marks
only once the resumed run reaches a terminal state, so a transient model
error leaves the answer re-giveable instead of stranding it.
- Resolve() does both, for a caller with a human right there who can simply
ask again.
Also deleted the five never-written Job fields (2 above). StatusWaiting stays:
it is a meaningful status in the job state machine and check_delegate
correctly treats it as non-terminal.
Tests: ported both original round-trip/validation tests, plus new coverage for
the transcript-carrying path and MarkResolved. The executive's suspend/resume
test now asserts through the continuation API instead of a hardcoded filename.
… admin
Consolidation step 2. "Personal Agent" is no longer a kind of agent. There is
one Agent abstraction; autonomy is orthogonal settings on it, managed through
the admin cockpit that already manages agents.
CONFIG. gwconfig.Agent drops Kind and gains four settings:
objective — the durable outcome it works toward
autonomous — whether it may act on that unprompted
browser — ephemeral (default) | existing_chrome
paused — stop unattended wakes without deleting anything
objective and autonomous are deliberately SEPARATE grants, which is the
correction that motivated this design. They answer different questions ("what
is it for" vs "may it act unasked"), and all four combinations are meaningful:
autonomous + objective → unattended objective pursuit (the old Personal Agent)
autonomous, no objective→ scheduled work UNDER GOVERNANCE — this is new, and
closes the long-standing gap where a cron-fired
agent ran unattended with no policy gate, no action
journal, and no way to pause and ask
objective, not autonomous → a goal you work on together; wakes only on demand
neither → an ordinary conversational agent, exactly as before
Autonomy therefore gates GOVERNANCE, not capability, and applies to any run of
the agent whether or not an objective exists.
SCHEDULING. The second cron implementation is gone as a user-facing concept.
Recurring cadence is an ordinary `schedules:` entry: set agent=<name>, leave
deliver_to empty, and gw_schedule routes the wake to the agent itself via the
renamed internal sink (personal → agent). The DB-backed loop now fires only
the agent's OWN self-scheduled next-wakes (schedule_wake, "come back in 45
minutes") — a genuinely different thing from human-authored cadence, and the
only part that needs to be writable from inside a run.
TOOLS. internal/agent/tools/personal.go (pa_*) is deleted. The admin registry
gains gw_policy, gw_grant, gw_wake, gw_inbox, gw_answer, gw_journal, gw_doctor,
gw_browser; gw_agent gains objective/autonomous/browser/pause/resume actions.
Everything dispatches through the existing adminExecute + approval gate, and
gw_schedule's shared BuildSchedule validation now covers agent cadence too, so
the surfaces cannot drift the way the parallel path did.
COCKPIT. `memcode personal` is deleted outright — cmd/personal.go,
cmd/personal_cockpit.go, Session.SetPersonal, personalMode, the personal and
personal_admin doctrines. The Personal setup walkthrough is merged into
adminDoctrine, generalized, and now teaches the two gates explicitly: granting
an objective is not granting autonomy, and the second must be confirmed on its
own.
EXECUTIVE. Executive.Objective is read from configuration rather than the
store, so the objective has ONE source that a human edits and the gateway
hot-reloads. The per-agent config.yaml mirror drops the objective for the same
reason — it lives in gateway.yaml, which is already a readable file.
Tests are rewritten against adminExecute (there is no other surface). New
coverage asserts the orthogonality directly, since a single overloaded switch
is exactly what this change exists to prevent: an objective alone must not
confer autonomy, a non-affirmative value must not grant it by typo, an
autonomous agent with no objective is still governed but refuses to invent
work, and an autonomous agent's schedule defaults to the agent route.
…ard both
Consolidation step 3, finishing the merge.
RENAME. internal/personal -> internal/agent/autonomy. The package was never
about a "personal" species; it is the machinery an agent uses when running
unattended toward an objective. Terminology in comments follows.
SECOND CRON PARSER DELETED. NextDue now understands only the kinds an AGENT
writes for itself from inside a run (one_shot / next_wake — always a single
future instant, e.g. schedule_wake "come back in 45 minutes"). Recurring
cadence a human configures is not a trigger at all: it is an ordinary gateway
schedule delivering to agent:<name>, validated once by gwconfig. With
recurring kinds gone, ClaimDueTrigger's reschedule branch was unreachable, so
firing a wake now simply completes it — and the atomic `last_fired_at IS ?`
claim still keeps two gateway processes from double-firing.
GUARDS (internal/guard/singletons_test.go). The duplication this branch kept
producing was not a one-off: fixes landed on one path and not the other,
repeatedly. Four invariants now fail the build instead:
- TestSingleCronParser — only internal/gateway parses cron. (This one caught a
real leftover while being written: the autonomy package was still importing
robfig/cron, which is what prompted the deletion above.)
- TestSingleSuspensionImplementation — no hand-rolled continuation formats
outside internal/agent/continuation.
- TestNoSecondCockpit — no SetPersonal/personalMode/pa_* anywhere.
- TestNoAgentKind — autonomy is orthogonal settings, never a kind
discriminator; a "kind" field is what made Personal a separate species.
The guards deliberately skip vendored forks, .memcode session snapshots, and
desktop/node_modules (the same symlink that breaks the two pre-existing
go-list guard failures).
Store tests updated: the trigger test now exercises a real self-scheduled wake
and asserts it completes rather than reschedules, and NextDue's test asserts
that interval/cron are REJECTED — accepting them again would rebuild the
second scheduler.
Consolidation step 4. An agent had two unrelated places to put what it knew:
a structured `facts` table only reachable through bespoke tools, and the
memory.md every memcode agent already has and the runtime already injects. One
of those a human can read. Collapse to it.
BLOCKING PREREQUISITE, done first. delegate/check_delegate were using the
facts table as a keyed index — writing "delegation.<job-id>" so a later wake
could find the action to close out. Its own comment admitted why ("facts are
the only durable log delegate can write to without a schema migration"). This
is that migration: actions gain a job_id column plus LinkActionJob /
ActionForJob, and the link now lives on the action it actually describes.
Removing facts before this would have broken check_delegate silently.
note_fact becomes `remember`: one plain sentence appended to memory.md, read
back into every wake's state summary. Append-only and deduplicated, because
every line replays into the model on each future wake — a re-learned fact must
not grow the file forever.
Deleted with it: environment.go (EnvironmentModel/StructuredFact, referenced
nowhere but its own test) and UsableForExternalRepresentation.
The tradeoff is deliberate and documented in memory.go rather than left
implicit. Prose cannot distinguish "you told me this" from "I inferred it from
your resume" from "a website said so", nor mark a claim safe to assert on the
user's behalf, nor mark it stale. That distinction becomes load-bearing the
moment an agent fills in a form or sends a message stating something about the
user. It is cheap to drop TODAY only because the gate was never wired —
nothing read Confirmed to gate anything; the sole caller of
UsableForExternalRepresentation was its own test. When external
representation needs provenance, it should come back as its own
machine-checkable thing in the store alongside policies and the action
journal, NOT by reviving this table.
… as a mode
Consolidation step 5, finishing the merge.
MIGRATION. YAML silently ignores unknown fields, so an existing
`kind: personal` agent would have loaded as an ordinary one — still listed,
apparently fine, and never waking again. For a change that removes an agent's
authority to run on its own, silence is the wrong failure. gwconfig.Agent keeps
a LegacyKind field for the sole purpose of rejecting it, with the one-line fix
in the error:
agent "demo" still uses the removed `kind: personal` setting. Autonomy is now
explicit: replace it with `autonomous: true` (and an `objective:` ...), or
just delete the `kind:` line if it should not. Its home ... is untouched
Verified against the real local config, which has exactly that agent. The
guard's TestNoAgentKind is refined to allow this one mention while still
banning any behavioral branch on it.
DOCS. docs/personal-agents.md -> docs/autonomous-agents.md, rewritten around
the two orthogonal grants and the four combinations they produce, with the
scheduling split (your cadence = an ordinary schedule; the agent's own next
wake = self-scheduled), the fail-closed existing-Chrome rule, and the
memory.md provenance limitation stated rather than glossed.
docs/design/personal-agents.md -> docs/design/autonomous-agents.md keeps the
original design contract but opens with a note on what the implementation
revised and why, including the orthogonality correction. It is a design record;
rewriting history there would lose the reasoning.
README drops "three ways to run it" (there are two) and describes autonomy as a
setting rather than a product tier, leading with the case that actually
motivated it: a scheduled agent that is finally policy-gated, journaled, and
able to stop and ask.
The store file was still named personal.db. It holds state for any agent
running unattended, not a species of agent — rename it, and note in place that
it is opened lazily so an ordinary conversational agent never grows one.
That claim is now a test rather than a comment: TestOrdinaryAgentGetsNoAutonomyStore
creates a plain agent, exercises the admin surface against it, and asserts no
agent.db appears — the governance machinery costs nothing until asked for.
Also covers the remaining orthogonality row: an agent WITH an objective but
WITHOUT autonomy still wakes on demand, reaching the policy gate rather than
being refused for lacking autonomy. Autonomy governs unprompted action, not
whether a human may ask.
All four combinations of (objective, autonomous) now have a test.
@timothyerwintimothyerwin changed the title Personal Agents: durable objectives, delegation policy, triggers, browser brokerAutonomous AgentsAug 30, 2026
@timothyerwin
timothyerwin merged commit 8a2e6b3 into mainAug 30, 2026
2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@timothyerwin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Autonomous Agents - #5

Merged
timothyerwin merged 13 commits into
mainfrom
feat/personal-agents
Aug 30, 2026
Merged

Autonomous Agents#5
timothyerwin merged 13 commits into
mainfrom
feat/personal-agents

Conversation

@timothyerwin

Copy link
Copy Markdown
Contributor

Summary

Adds Personal Agents: long-lived agent identities with a durable objective/subgoal/facts store, a delegation policy that governs what their executive loop may do (allowed tools, consequence classes, filesystem/browser/MCP resource grants, budgets), ask_user suspend/resume for human-in-the-loop interactions, durable cron/interval/one-shot triggers polled by the gateway, and a browser broker that leases exclusive browser control to one run at a time.

New surfaces: memcode personal CLI (create, approve-policy, resources, triggers, answer, status, cockpit), gw_agent/gw_* admin tools, and the gateway's internal personal wake route.

Review fixes included

This branch went through a review pass (/code-review high) before commit; the following were found and fixed:

  • Duplicate-answer replay: personal answer now refuses to re-resume an already-answered interaction, and a resumed run that re-suspends on a new question now marks its prior continuation resolved — closing a window where a retried personal answer call could re-execute real tool side effects.
  • Triggers silently dead after hot-add: the gateway now always registers the personal wake route and always starts the trigger poll loop (it already re-reads config every tick), so a Personal Agent created after boot has its triggers picked up without a gateway restart.
  • Expired browser lease still granted mutation rights: CanMutate/OwnPage now check lease expiry like Authenticate already did.
  • Resource grant ID collisions: IDs now use a monotonic timestamp instead of len(locator), matching the cockpit's existing scheme.
  • Delegation policy narrowing gaps: IsRestriction/NarrowPolicy now also checks filesystem roots, browser origins, MCP tools, budget fields, and quiet hours, not just tools/consequence-classes/concurrency/depth.
  • Silent no-op on bad trigger id: personal triggers pause/resume now errors instead of reporting success when nothing matched.
  • Stale tool description: gw_agent no longer advertises the unsupported "inspect" action.
  • Non-atomic suspension writes: suspension continuation files are now written via atomicfile.WriteFile.
  • Per-tick DB churn: the trigger poll loop now reuses one DB connection per agent across ticks and filters due triggers in SQL instead of opening/closing a connection and Go-filtering every row every 15s.

Testing

  • go build ./... — clean
  • go vet ./... — clean
  • go test ./... — all pass except two pre-existing, unrelated failures in internal/guard caused by a symlink under desktop/node_modules/@electron/notarize breaking go list (not touched by this branch)

… triggers, browser broker)
Implements the Personal Agents feature: long-lived agent identities with
their own durable objective, subgoals, and facts store; a delegation
policy model (allowed tools, consequence classes, filesystem/browser/MCP
grants, budgets) that governs what a Personal Agent's executive loop may
do; ask_user suspend/resume for human-in-the-loop interactions; durable
triggers (cron/interval/one-shot) polled by the gateway to wake an agent;
resource grants (filesystem, mcp, command, channel); and a browser broker
that leases exclusive browser control to one run at a time.
New surfaces: `memcode personal` CLI (create, approve-policy, resources,
triggers, answer, status, cockpit), gw_agent/gw_* admin tools, and the
gateway's internal personal wake route.
Includes fixes from review of this branch:
- personal answer refuses to re-resume an already-answered interaction,
and a resumed run that re-suspends now marks its prior continuation
resolved — closing a window where a retried answer could replay tool
side effects.
- The gateway always registers the personal wake route and always starts
the trigger poll loop, so a Personal Agent added after boot has its
triggers picked up without a restart.
- Browser broker leases now check expiry in CanMutate/OwnPage, matching
Authenticate, so an expired lease can no longer authorize mutation.
- Resource grant IDs use a monotonic timestamp instead of locator length,
avoiding same-length-locator collisions.
- Delegation policy narrowing (IsRestriction/NarrowPolicy) now also
checks filesystem roots, browser origins, MCP tools, budget fields, and
quiet hours, not just tools/consequence-classes/concurrency/depth.
- `personal triggers pause/resume` errors on an unknown trigger id instead
of reporting success silently.
- gw_agent's description no longer advertises the unsupported "inspect"
action.
- Suspension continuation files are written atomically.
- The trigger poll loop reuses one DB connection per agent across ticks
and filters due triggers in SQL, instead of opening/closing a
connection and scanning every trigger row every 15s.
… CI lint
Closes the biggest gap from review: the Personal Agent executive had no way
to reach browser/MCP/shell/filesystem capability — it was a second, weaker
agent runtime sitting beside the real memcode engine instead of commanding
it. Two new executive tools fix that:
- delegate: spawns a policy-scoped worker as a real detached `memcode run`
job (internal/jobs.SpawnWithSpec), using the delegation.go scaffolding
(ExecutionEnvelope, ValidateDelegation, PrepareRunDirectory) that already
existed but was never called from anywhere. The worker gets whatever
toolsets it's granted — browser, MCP, shell, filesystem, skills — same as
any ordinary memcode agent, not the executive's fixed 7 tools. Requested
toolsets/consequences must be a subset of the agent's own approved policy;
MaxDelegationDepth gates whether delegation is allowed at all.
- check_delegate: reads a delegated job's status/result back on a later
wake (RunOnce is one bounded wake, so the result can't arrive in the same
call) and closes out its action-journal entry.
Also makes jobs.SpawnSpec.ToolPolicy a REAL restriction instead of recorded-
only metadata: new hidden --allow-tools/--deny-tools flags on `memcode run`
bind to the same SetToolPolicy enforcement an ordinary gateway-bound agent
gets from its config, and SpawnWithSpec now passes them through argv.
Scope notes, called out explicitly rather than implied as done:
- "browser" toolset reuses the existing ephemeral Chrome session (the same
one --chrome already drives) — NOT the user's own already-running Chrome.
Attaching to a persistent/remote Chrome needs the broker/RemoteConfig
wiring in internal/browser, which is declared but has zero callers and
zero implementation; that's separate follow-up work, not done here.
- Filesystem/MCP resource grants are not OS-sandboxed to a spawned worker
(memcode has no such sandbox anywhere) — governance here is the same
approval/audit model as the rest of the codebase (tool-level allow/deny +
the action journal), not a hard jail.
Also fixes two CI lint failures on the open PR:
- internal/gateway/state/state.go: removed formatTime, an unused helper
(staticcheck U1000).
- internal/personal/store.go: decapitalized a Go error string (ST1005).
Corrects course on the browser gap flagged after the last review: existing-
Chrome access is core to Personal Agents, not deferrable follow-up. A
Personal Agent's whole premise — acting as the user across their real
accounts (Gmail, LinkedIn, an ATS, internal dashboards) — requires the
user's actual signed-in session, not a fresh ephemeral profile with no
cookies. Per docs/design/personal-agents.md's already-specified "browser
broker trust boundary": Personal Agents connect to the user's existing
Chrome through a gateway-owned broker and a permission-protected local
socket, fail closed on any connection/auth problem, and never silently fall
back to another profile.
- internal/browser/broker/{server,client}.go: exposes the (already-reviewed)
in-process Broker over a Unix socket, so a delegated worker — a SEPARATE
OS process spawned via jobs.SpawnWithSpec — can Acquire/Release/OwnPage/
CanMutate against the SAME broker the gateway owns, not a fresh one.
Socket is 0600 inside gwconfig.Dir() (0700). Round-trip tested, including
that a concurrent Acquire is correctly rejected and a release unblocks the
next one.
- internal/gateway/server: the gateway now owns one *broker.Broker for its
whole lifetime and serves it on that socket unconditionally at startup
(cheap — a local listener), so it's there the moment a delegate call needs
it, no restart required after `memcode personal browser setup`.
- internal/personal: browserModeFor's "browser" toolset now defaults to
BrowserExistingChrome, not ephemeral — "browser:ephemeral" is the explicit
opt-down for tasks that genuinely want a fresh, logged-out profile. Before
spawning anything, the delegate tool verifies the broker is reachable and
fails the tool call closed (a normal tool_result error, not a spawned job)
if it isn't — verified by test, alongside the success path asserting the
spawned job's SpawnSpec actually carries BrowserMode=existing_chrome.
- internal/jobs: SpawnWithSpec passes existing_chrome through as
--browser-session (+ --browser-agent/--browser-run for the lease identity)
instead of --chrome.
- cmd/run.go: --browser-session existing_chrome acquires the broker lease
itself (via the client) and, only on success, wires an ad-hoc
chrome-devtools-mcp --autoConnect MCP server into the session
(Session.SetExtraMCPServers, new) — the SAME chrome-devtools-mcp package
this browser experience is built on (see internal/browser/remote.go). On
any failure it returns an error and exits; it never calls
SetBrowserEnabled(true) (ephemeral) as a fallback.
- cmd/personal_browser.go: `memcode personal browser setup` checks npx and
the broker socket, then attempts a REAL bounded (10s) chrome-devtools-mcp
connection and reports the tool count or the actual connect errors — it
does not pretend to click Chrome's own "Allow" dialog, which only the user
can do.
- Also fixes a real bug found while testing this: ValidateDelegation treated
an empty parent.AllowedTools as "allows no tools by name", the opposite of
the len(...)>0-means-restricted convention Executive.allowedTools already
uses — every delegate call with a name-unrestricted policy was rejected.
Scope note, stated rather than implied: the literal Chrome-side consent flow
(the "Allow" dialog, actually clicking it) cannot be exercised or verified
in this environment — no live Chrome, no human to click it. Everything
up to that point (broker lifecycle, lease exclusivity, fail-closed wiring,
the actual chrome-devtools-mcp CLI invocation, verified against the real
published package) is built and tested; the dialog itself needs the user's
own machine.
…setup via the cockpit
Three separate complaints from review, addressed at the right layer each:
1. "It's just a file ref" — `resources add` required a redundant --mode read
(already the default) and a mandatory <type> argument even for the
overwhelming common case. A bare path is now enough:
memcode personal resources add jobhunt ~/resume.md
Non-filesystem grants (mcp/command/channel) still need the type spelled
out, since there's nothing to infer it from. Also fixes the real bug this
surfaced: CanonicalFilesystemGrant rejected any non-directory path, even
though PathWithinGrant already handles a single-file grant correctly
(rel == "." on an exact match) — you could not actually grant a single
file like a resume before this, only whole directories. `personal create`
also gained a repeatable --grant flag to fold "create it, then give it
something to read" into one command.
2. "Not this CLI-command-argument garbage" — Personal Agent state (objective,
policy, resources) lived ONLY in personal.db, reachable exclusively
through bespoke commands, unlike every other piece of memcode config
(gateway.yaml, .mcp.json, CLAUDE.md) which is a plain file. WriteConfigMirror
now regenerates objective.md/policy.yaml/resources.yaml in the agent's home
on every mutation — real, readable, diffable files. The run journal
(actions/triggers/interactions) deliberately stays in SQLite: it needs
atomic claim/complete semantics under concurrent access (gateway wake loop,
CLI, and cockpit can all touch the same agent) that flat files don't get
for free — a correctness reason, not a habit.
3. "It should gather requirements, do HITL, pre-approve permissions, set up
the runtime — like our other agents already do via telegram/model setup"
— `memcode personal` (no args) is ALREADY the interactive cockpit (like
`memcode admin`) with the pa_* tools wired for exactly this. The actual gap
was its system prompt: personalAdminDoctrine listed tool mechanics with no
instruction to gather-then-propose-then-approve. Rewrote it to require a
real walkthrough for first-time creation — reason about what resources/
toolsets/consequences the objective needs (asking the user rather than
guessing), present the concrete plan in plain language BEFORE touching
anything, then apply everything at once on approval, including wake
cadence (trigger) — an agent with no trigger and no plan to ever be woken
is dead on arrival, so cadence is not an optional afterthought either.
A prior pass at this (adding a NEW deterministic `memcode personal setup`
CLI wizard, mirroring `gateway setup`'s bufio-prompt mechanism) was wrong
and reverted before commit — that reproduces the exact CLI-surface-
sprawl complaint instead of fixing it. The cockpit conversation IS the
wizard; it needed a better prompt, not a new command.
Root cause of the whole CLI-surface complaint: the interactive cockpit
(`memcode personal`, no args) is supposed to be the entire interface, but
its tool set had NO way to create a new Personal Agent. pa_objective only
supported show/set, and set required the objective to already exist —
creation was only reachable through the CLI's `personal create`. That's why
every explanation kept surfacing a CLI command: until now, it was the only
path that actually worked.
- internal/agent/tools/personal.go: new pa_create tool (agent name +
objective). Documented as the required first call for a brand-new agent.
- cmd/personal_cockpit.go: paCreate mirrors personalCreate's steps (register
in gateway.yaml, open the store, create the objective) and is dispatched
BEFORE the paStore existence gate, since the agent doesn't exist yet.
pa_resource's grant action now defaults type=filesystem/mode=read when
omitted too (same inference the CLI's `resources add` already got), and
pa_resource/pa_policy's grant/revoke/stage/approve paths now call
WriteConfigMirror — that was wired into the CLI paths last commit but not
into the cockpit paths, which are the ones actually used.
- cmd/personal.go: every `memcode personal` subcommand is now Hidden — they
remain fully callable for scripts, but `memcode personal --help` no longer
reads like a CLI to memorize. The interface is `memcode personal` with no
args, then plain conversation.
- internal/doctrine/prompts.go: the cockpit's own system prompt now says
outright never to tell the user to run a CLI command — it does the work
itself, right there — and that pa_create, not pa_objective, is the first
call for a new agent.
Verified end to end: pa_create then a bare-path pa_resource grant (no type,
no mode) both succeed through personalExecute exactly as the cockpit would
call them, no CLI involved.
Hiding the subcommands last commit wasn't enough — they still existed as a
parallel, scriptable way to do everything, which is real duplicate logic
(and already caused drift: a fix landed in one path and not the other,
twice, in this same review). Deleted them outright:
cmd/personal_resources.go, personal_policy.go, personal_triggers.go,
personal_status.go, personal_browser.go, personal_test.go — all removed.
cmd/personal.go cut down to just the cockpit launcher (~90 lines from
~370): `memcode personal` opens the interactive session, full stop.
Two operations the CLI covered had no pa_* equivalent yet, so they'd have
been silently lost:
- pa_doctor: the health check (home layout, objective, approved policy,
generated workspace, sandbox availability, trigger/pending-interaction
counts) that `personal doctor` used to run.
- pa_browser_setup: the existing-Chrome prerequisite check + real, bounded
chrome-devtools-mcp connection attempt that `personal browser setup` used
to run. Same logic, now a tool call instead of a command — it still
doesn't (can't) click Chrome's own consent dialog for the user.
Both dispatch through personalExecute exactly like every other pa_* tool.
internal/agent/tools/personal.go: added PaDoctor and PaBrowserSetup to the
registry.
cmd/personal_cmd_test.go: rewritten from scratch to call personalExecute
directly — no CLI args, no cobra Execute(), no rootCmd. This IS the
interface being tested now, so this is what the tests exercise. Also
asserts the config-mirror files (policy.yaml, resources.yaml) actually
reflect approvals/revokes done through the tool path, since that's the one
that matters.
Verified: `memcode personal create foo bar` no longer resolves to anything
— it just opens the cockpit with those words ignored as stray args, because
there is no "create" subcommand left to match. The interface is
conversation with pa_* tools behind it, not CLI argument syntax.
objective.md/policy.yaml/resources.yaml split one agent's config across
three files by table for no real reason — a person has to know to look in
three places for one agent's setup. Consolidated into a single config.yaml
(objective, policies, resources as sections) with proper snake_case field
names. Same authority split as before (this is a mirror; policy hash-
approval and the run journal stay in personal.db, and why, is unchanged) —
just one file to read instead of three.
Consolidation step 1 of folding Personal Agents into the ordinary agent
system. There were three partial suspend/resume designs in the tree:
1. internal/agent/runtime/continuation.go — typed, atomic, tested, and with
ZERO production callers.
2. jobs.Job's InteractionID/WaitingReason/ContinuationVersion/WaitingAt/
ResumedAt — declared, never written by anything.
3. A hand-rolled map[string]any in the personal executive — the only one
actually running, and (until a fix earlier in this branch) the only one
that wasn't crash-safe.
New internal/agent/continuation is the single implementation. It lives in its
own small package rather than in internal/agent/runtime so the executive can
suspend without importing the whole session runtime (and so neither side ends
up depending on the other).
It has to serve two genuinely different callers, which is why the previous
attempt at a shared API didn't fit:
- An interactive session already holds the conversation, so it only needs the
missing assistant/answer pair back.
- An unattended executive keeps NO transcript — it rebuilds context from
durable state each wake — so its continuation must carry the whole
conversation. Hence the optional Messages field.
They also differ on when marking resolved is safe, so building the resume
messages is split from marking:
- ResumeMessages() builds without marking. The executive uses this and marks
only once the resumed run reaches a terminal state, so a transient model
error leaves the answer re-giveable instead of stranding it.
- Resolve() does both, for a caller with a human right there who can simply
ask again.
Also deleted the five never-written Job fields (2 above). StatusWaiting stays:
it is a meaningful status in the job state machine and check_delegate
correctly treats it as non-terminal.
Tests: ported both original round-trip/validation tests, plus new coverage for
the transcript-carrying path and MarkResolved. The executive's suspend/resume
test now asserts through the continuation API instead of a hardcoded filename.
… admin
Consolidation step 2. "Personal Agent" is no longer a kind of agent. There is
one Agent abstraction; autonomy is orthogonal settings on it, managed through
the admin cockpit that already manages agents.
CONFIG. gwconfig.Agent drops Kind and gains four settings:
objective — the durable outcome it works toward
autonomous — whether it may act on that unprompted
browser — ephemeral (default) | existing_chrome
paused — stop unattended wakes without deleting anything
objective and autonomous are deliberately SEPARATE grants, which is the
correction that motivated this design. They answer different questions ("what
is it for" vs "may it act unasked"), and all four combinations are meaningful:
autonomous + objective → unattended objective pursuit (the old Personal Agent)
autonomous, no objective→ scheduled work UNDER GOVERNANCE — this is new, and
closes the long-standing gap where a cron-fired
agent ran unattended with no policy gate, no action
journal, and no way to pause and ask
objective, not autonomous → a goal you work on together; wakes only on demand
neither → an ordinary conversational agent, exactly as before
Autonomy therefore gates GOVERNANCE, not capability, and applies to any run of
the agent whether or not an objective exists.
SCHEDULING. The second cron implementation is gone as a user-facing concept.
Recurring cadence is an ordinary `schedules:` entry: set agent=<name>, leave
deliver_to empty, and gw_schedule routes the wake to the agent itself via the
renamed internal sink (personal → agent). The DB-backed loop now fires only
the agent's OWN self-scheduled next-wakes (schedule_wake, "come back in 45
minutes") — a genuinely different thing from human-authored cadence, and the
only part that needs to be writable from inside a run.
TOOLS. internal/agent/tools/personal.go (pa_*) is deleted. The admin registry
gains gw_policy, gw_grant, gw_wake, gw_inbox, gw_answer, gw_journal, gw_doctor,
gw_browser; gw_agent gains objective/autonomous/browser/pause/resume actions.
Everything dispatches through the existing adminExecute + approval gate, and
gw_schedule's shared BuildSchedule validation now covers agent cadence too, so
the surfaces cannot drift the way the parallel path did.
COCKPIT. `memcode personal` is deleted outright — cmd/personal.go,
cmd/personal_cockpit.go, Session.SetPersonal, personalMode, the personal and
personal_admin doctrines. The Personal setup walkthrough is merged into
adminDoctrine, generalized, and now teaches the two gates explicitly: granting
an objective is not granting autonomy, and the second must be confirmed on its
own.
EXECUTIVE. Executive.Objective is read from configuration rather than the
store, so the objective has ONE source that a human edits and the gateway
hot-reloads. The per-agent config.yaml mirror drops the objective for the same
reason — it lives in gateway.yaml, which is already a readable file.
Tests are rewritten against adminExecute (there is no other surface). New
coverage asserts the orthogonality directly, since a single overloaded switch
is exactly what this change exists to prevent: an objective alone must not
confer autonomy, a non-affirmative value must not grant it by typo, an
autonomous agent with no objective is still governed but refuses to invent
work, and an autonomous agent's schedule defaults to the agent route.
…ard both
Consolidation step 3, finishing the merge.
RENAME. internal/personal -> internal/agent/autonomy. The package was never
about a "personal" species; it is the machinery an agent uses when running
unattended toward an objective. Terminology in comments follows.
SECOND CRON PARSER DELETED. NextDue now understands only the kinds an AGENT
writes for itself from inside a run (one_shot / next_wake — always a single
future instant, e.g. schedule_wake "come back in 45 minutes"). Recurring
cadence a human configures is not a trigger at all: it is an ordinary gateway
schedule delivering to agent:<name>, validated once by gwconfig. With
recurring kinds gone, ClaimDueTrigger's reschedule branch was unreachable, so
firing a wake now simply completes it — and the atomic `last_fired_at IS ?`
claim still keeps two gateway processes from double-firing.
GUARDS (internal/guard/singletons_test.go). The duplication this branch kept
producing was not a one-off: fixes landed on one path and not the other,
repeatedly. Four invariants now fail the build instead:
- TestSingleCronParser — only internal/gateway parses cron. (This one caught a
real leftover while being written: the autonomy package was still importing
robfig/cron, which is what prompted the deletion above.)
- TestSingleSuspensionImplementation — no hand-rolled continuation formats
outside internal/agent/continuation.
- TestNoSecondCockpit — no SetPersonal/personalMode/pa_* anywhere.
- TestNoAgentKind — autonomy is orthogonal settings, never a kind
discriminator; a "kind" field is what made Personal a separate species.
The guards deliberately skip vendored forks, .memcode session snapshots, and
desktop/node_modules (the same symlink that breaks the two pre-existing
go-list guard failures).
Store tests updated: the trigger test now exercises a real self-scheduled wake
and asserts it completes rather than reschedules, and NextDue's test asserts
that interval/cron are REJECTED — accepting them again would rebuild the
second scheduler.
Consolidation step 4. An agent had two unrelated places to put what it knew:
a structured `facts` table only reachable through bespoke tools, and the
memory.md every memcode agent already has and the runtime already injects. One
of those a human can read. Collapse to it.
BLOCKING PREREQUISITE, done first. delegate/check_delegate were using the
facts table as a keyed index — writing "delegation.<job-id>" so a later wake
could find the action to close out. Its own comment admitted why ("facts are
the only durable log delegate can write to without a schema migration"). This
is that migration: actions gain a job_id column plus LinkActionJob /
ActionForJob, and the link now lives on the action it actually describes.
Removing facts before this would have broken check_delegate silently.
note_fact becomes `remember`: one plain sentence appended to memory.md, read
back into every wake's state summary. Append-only and deduplicated, because
every line replays into the model on each future wake — a re-learned fact must
not grow the file forever.
Deleted with it: environment.go (EnvironmentModel/StructuredFact, referenced
nowhere but its own test) and UsableForExternalRepresentation.
The tradeoff is deliberate and documented in memory.go rather than left
implicit. Prose cannot distinguish "you told me this" from "I inferred it from
your resume" from "a website said so", nor mark a claim safe to assert on the
user's behalf, nor mark it stale. That distinction becomes load-bearing the
moment an agent fills in a form or sends a message stating something about the
user. It is cheap to drop TODAY only because the gate was never wired —
nothing read Confirmed to gate anything; the sole caller of
UsableForExternalRepresentation was its own test. When external
representation needs provenance, it should come back as its own
machine-checkable thing in the store alongside policies and the action
journal, NOT by reviving this table.
… as a mode
Consolidation step 5, finishing the merge.
MIGRATION. YAML silently ignores unknown fields, so an existing
`kind: personal` agent would have loaded as an ordinary one — still listed,
apparently fine, and never waking again. For a change that removes an agent's
authority to run on its own, silence is the wrong failure. gwconfig.Agent keeps
a LegacyKind field for the sole purpose of rejecting it, with the one-line fix
in the error:
agent "demo" still uses the removed `kind: personal` setting. Autonomy is now
explicit: replace it with `autonomous: true` (and an `objective:` ...), or
just delete the `kind:` line if it should not. Its home ... is untouched
Verified against the real local config, which has exactly that agent. The
guard's TestNoAgentKind is refined to allow this one mention while still
banning any behavioral branch on it.
DOCS. docs/personal-agents.md -> docs/autonomous-agents.md, rewritten around
the two orthogonal grants and the four combinations they produce, with the
scheduling split (your cadence = an ordinary schedule; the agent's own next
wake = self-scheduled), the fail-closed existing-Chrome rule, and the
memory.md provenance limitation stated rather than glossed.
docs/design/personal-agents.md -> docs/design/autonomous-agents.md keeps the
original design contract but opens with a note on what the implementation
revised and why, including the orthogonality correction. It is a design record;
rewriting history there would lose the reasoning.
README drops "three ways to run it" (there are two) and describes autonomy as a
setting rather than a product tier, leading with the case that actually
motivated it: a scheduled agent that is finally policy-gated, journaled, and
able to stop and ask.
The store file was still named personal.db. It holds state for any agent
running unattended, not a species of agent — rename it, and note in place that
it is opened lazily so an ordinary conversational agent never grows one.
That claim is now a test rather than a comment: TestOrdinaryAgentGetsNoAutonomyStore
creates a plain agent, exercises the admin surface against it, and asserts no
agent.db appears — the governance machinery costs nothing until asked for.
Also covers the remaining orthogonality row: an agent WITH an objective but
WITHOUT autonomy still wakes on demand, reaching the policy gate rather than
being refused for lacking autonomy. Autonomy governs unprompted action, not
whether a human may ask.
All four combinations of (objective, autonomous) now have a test.
@timothyerwintimothyerwin changed the title Personal Agents: durable objectives, delegation policy, triggers, browser brokerAutonomous AgentsAug 30, 2026
@timothyerwin
timothyerwin merged commit 8a2e6b3 into mainAug 30, 2026
2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@timothyerwin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Autonomous Agents - #5

Merged
timothyerwin merged 13 commits into
mainfrom
feat/personal-agents
Aug 30, 2026
Merged

Autonomous Agents#5
timothyerwin merged 13 commits into
mainfrom
feat/personal-agents

Conversation

@timothyerwin

Copy link
Copy Markdown
Contributor

Summary

Adds Personal Agents: long-lived agent identities with a durable objective/subgoal/facts store, a delegation policy that governs what their executive loop may do (allowed tools, consequence classes, filesystem/browser/MCP resource grants, budgets), ask_user suspend/resume for human-in-the-loop interactions, durable cron/interval/one-shot triggers polled by the gateway, and a browser broker that leases exclusive browser control to one run at a time.

New surfaces: memcode personal CLI (create, approve-policy, resources, triggers, answer, status, cockpit), gw_agent/gw_* admin tools, and the gateway's internal personal wake route.

Review fixes included

This branch went through a review pass (/code-review high) before commit; the following were found and fixed:

  • Duplicate-answer replay: personal answer now refuses to re-resume an already-answered interaction, and a resumed run that re-suspends on a new question now marks its prior continuation resolved — closing a window where a retried personal answer call could re-execute real tool side effects.
  • Triggers silently dead after hot-add: the gateway now always registers the personal wake route and always starts the trigger poll loop (it already re-reads config every tick), so a Personal Agent created after boot has its triggers picked up without a gateway restart.
  • Expired browser lease still granted mutation rights: CanMutate/OwnPage now check lease expiry like Authenticate already did.
  • Resource grant ID collisions: IDs now use a monotonic timestamp instead of len(locator), matching the cockpit's existing scheme.
  • Delegation policy narrowing gaps: IsRestriction/NarrowPolicy now also checks filesystem roots, browser origins, MCP tools, budget fields, and quiet hours, not just tools/consequence-classes/concurrency/depth.
  • Silent no-op on bad trigger id: personal triggers pause/resume now errors instead of reporting success when nothing matched.
  • Stale tool description: gw_agent no longer advertises the unsupported "inspect" action.
  • Non-atomic suspension writes: suspension continuation files are now written via atomicfile.WriteFile.
  • Per-tick DB churn: the trigger poll loop now reuses one DB connection per agent across ticks and filters due triggers in SQL instead of opening/closing a connection and Go-filtering every row every 15s.

Testing

  • go build ./... — clean
  • go vet ./... — clean
  • go test ./... — all pass except two pre-existing, unrelated failures in internal/guard caused by a symlink under desktop/node_modules/@electron/notarize breaking go list (not touched by this branch)

… triggers, browser broker)
Implements the Personal Agents feature: long-lived agent identities with
their own durable objective, subgoals, and facts store; a delegation
policy model (allowed tools, consequence classes, filesystem/browser/MCP
grants, budgets) that governs what a Personal Agent's executive loop may
do; ask_user suspend/resume for human-in-the-loop interactions; durable
triggers (cron/interval/one-shot) polled by the gateway to wake an agent;
resource grants (filesystem, mcp, command, channel); and a browser broker
that leases exclusive browser control to one run at a time.
New surfaces: `memcode personal` CLI (create, approve-policy, resources,
triggers, answer, status, cockpit), gw_agent/gw_* admin tools, and the
gateway's internal personal wake route.
Includes fixes from review of this branch:
- personal answer refuses to re-resume an already-answered interaction,
and a resumed run that re-suspends now marks its prior continuation
resolved — closing a window where a retried answer could replay tool
side effects.
- The gateway always registers the personal wake route and always starts
the trigger poll loop, so a Personal Agent added after boot has its
triggers picked up without a restart.
- Browser broker leases now check expiry in CanMutate/OwnPage, matching
Authenticate, so an expired lease can no longer authorize mutation.
- Resource grant IDs use a monotonic timestamp instead of locator length,
avoiding same-length-locator collisions.
- Delegation policy narrowing (IsRestriction/NarrowPolicy) now also
checks filesystem roots, browser origins, MCP tools, budget fields, and
quiet hours, not just tools/consequence-classes/concurrency/depth.
- `personal triggers pause/resume` errors on an unknown trigger id instead
of reporting success silently.
- gw_agent's description no longer advertises the unsupported "inspect"
action.
- Suspension continuation files are written atomically.
- The trigger poll loop reuses one DB connection per agent across ticks
and filters due triggers in SQL, instead of opening/closing a
connection and scanning every trigger row every 15s.
… CI lint
Closes the biggest gap from review: the Personal Agent executive had no way
to reach browser/MCP/shell/filesystem capability — it was a second, weaker
agent runtime sitting beside the real memcode engine instead of commanding
it. Two new executive tools fix that:
- delegate: spawns a policy-scoped worker as a real detached `memcode run`
job (internal/jobs.SpawnWithSpec), using the delegation.go scaffolding
(ExecutionEnvelope, ValidateDelegation, PrepareRunDirectory) that already
existed but was never called from anywhere. The worker gets whatever
toolsets it's granted — browser, MCP, shell, filesystem, skills — same as
any ordinary memcode agent, not the executive's fixed 7 tools. Requested
toolsets/consequences must be a subset of the agent's own approved policy;
MaxDelegationDepth gates whether delegation is allowed at all.
- check_delegate: reads a delegated job's status/result back on a later
wake (RunOnce is one bounded wake, so the result can't arrive in the same
call) and closes out its action-journal entry.
Also makes jobs.SpawnSpec.ToolPolicy a REAL restriction instead of recorded-
only metadata: new hidden --allow-tools/--deny-tools flags on `memcode run`
bind to the same SetToolPolicy enforcement an ordinary gateway-bound agent
gets from its config, and SpawnWithSpec now passes them through argv.
Scope notes, called out explicitly rather than implied as done:
- "browser" toolset reuses the existing ephemeral Chrome session (the same
one --chrome already drives) — NOT the user's own already-running Chrome.
Attaching to a persistent/remote Chrome needs the broker/RemoteConfig
wiring in internal/browser, which is declared but has zero callers and
zero implementation; that's separate follow-up work, not done here.
- Filesystem/MCP resource grants are not OS-sandboxed to a spawned worker
(memcode has no such sandbox anywhere) — governance here is the same
approval/audit model as the rest of the codebase (tool-level allow/deny +
the action journal), not a hard jail.
Also fixes two CI lint failures on the open PR:
- internal/gateway/state/state.go: removed formatTime, an unused helper
(staticcheck U1000).
- internal/personal/store.go: decapitalized a Go error string (ST1005).
Corrects course on the browser gap flagged after the last review: existing-
Chrome access is core to Personal Agents, not deferrable follow-up. A
Personal Agent's whole premise — acting as the user across their real
accounts (Gmail, LinkedIn, an ATS, internal dashboards) — requires the
user's actual signed-in session, not a fresh ephemeral profile with no
cookies. Per docs/design/personal-agents.md's already-specified "browser
broker trust boundary": Personal Agents connect to the user's existing
Chrome through a gateway-owned broker and a permission-protected local
socket, fail closed on any connection/auth problem, and never silently fall
back to another profile.
- internal/browser/broker/{server,client}.go: exposes the (already-reviewed)
in-process Broker over a Unix socket, so a delegated worker — a SEPARATE
OS process spawned via jobs.SpawnWithSpec — can Acquire/Release/OwnPage/
CanMutate against the SAME broker the gateway owns, not a fresh one.
Socket is 0600 inside gwconfig.Dir() (0700). Round-trip tested, including
that a concurrent Acquire is correctly rejected and a release unblocks the
next one.
- internal/gateway/server: the gateway now owns one *broker.Broker for its
whole lifetime and serves it on that socket unconditionally at startup
(cheap — a local listener), so it's there the moment a delegate call needs
it, no restart required after `memcode personal browser setup`.
- internal/personal: browserModeFor's "browser" toolset now defaults to
BrowserExistingChrome, not ephemeral — "browser:ephemeral" is the explicit
opt-down for tasks that genuinely want a fresh, logged-out profile. Before
spawning anything, the delegate tool verifies the broker is reachable and
fails the tool call closed (a normal tool_result error, not a spawned job)
if it isn't — verified by test, alongside the success path asserting the
spawned job's SpawnSpec actually carries BrowserMode=existing_chrome.
- internal/jobs: SpawnWithSpec passes existing_chrome through as
--browser-session (+ --browser-agent/--browser-run for the lease identity)
instead of --chrome.
- cmd/run.go: --browser-session existing_chrome acquires the broker lease
itself (via the client) and, only on success, wires an ad-hoc
chrome-devtools-mcp --autoConnect MCP server into the session
(Session.SetExtraMCPServers, new) — the SAME chrome-devtools-mcp package
this browser experience is built on (see internal/browser/remote.go). On
any failure it returns an error and exits; it never calls
SetBrowserEnabled(true) (ephemeral) as a fallback.
- cmd/personal_browser.go: `memcode personal browser setup` checks npx and
the broker socket, then attempts a REAL bounded (10s) chrome-devtools-mcp
connection and reports the tool count or the actual connect errors — it
does not pretend to click Chrome's own "Allow" dialog, which only the user
can do.
- Also fixes a real bug found while testing this: ValidateDelegation treated
an empty parent.AllowedTools as "allows no tools by name", the opposite of
the len(...)>0-means-restricted convention Executive.allowedTools already
uses — every delegate call with a name-unrestricted policy was rejected.
Scope note, stated rather than implied: the literal Chrome-side consent flow
(the "Allow" dialog, actually clicking it) cannot be exercised or verified
in this environment — no live Chrome, no human to click it. Everything
up to that point (broker lifecycle, lease exclusivity, fail-closed wiring,
the actual chrome-devtools-mcp CLI invocation, verified against the real
published package) is built and tested; the dialog itself needs the user's
own machine.
…setup via the cockpit
Three separate complaints from review, addressed at the right layer each:
1. "It's just a file ref" — `resources add` required a redundant --mode read
(already the default) and a mandatory <type> argument even for the
overwhelming common case. A bare path is now enough:
memcode personal resources add jobhunt ~/resume.md
Non-filesystem grants (mcp/command/channel) still need the type spelled
out, since there's nothing to infer it from. Also fixes the real bug this
surfaced: CanonicalFilesystemGrant rejected any non-directory path, even
though PathWithinGrant already handles a single-file grant correctly
(rel == "." on an exact match) — you could not actually grant a single
file like a resume before this, only whole directories. `personal create`
also gained a repeatable --grant flag to fold "create it, then give it
something to read" into one command.
2. "Not this CLI-command-argument garbage" — Personal Agent state (objective,
policy, resources) lived ONLY in personal.db, reachable exclusively
through bespoke commands, unlike every other piece of memcode config
(gateway.yaml, .mcp.json, CLAUDE.md) which is a plain file. WriteConfigMirror
now regenerates objective.md/policy.yaml/resources.yaml in the agent's home
on every mutation — real, readable, diffable files. The run journal
(actions/triggers/interactions) deliberately stays in SQLite: it needs
atomic claim/complete semantics under concurrent access (gateway wake loop,
CLI, and cockpit can all touch the same agent) that flat files don't get
for free — a correctness reason, not a habit.
3. "It should gather requirements, do HITL, pre-approve permissions, set up
the runtime — like our other agents already do via telegram/model setup"
— `memcode personal` (no args) is ALREADY the interactive cockpit (like
`memcode admin`) with the pa_* tools wired for exactly this. The actual gap
was its system prompt: personalAdminDoctrine listed tool mechanics with no
instruction to gather-then-propose-then-approve. Rewrote it to require a
real walkthrough for first-time creation — reason about what resources/
toolsets/consequences the objective needs (asking the user rather than
guessing), present the concrete plan in plain language BEFORE touching
anything, then apply everything at once on approval, including wake
cadence (trigger) — an agent with no trigger and no plan to ever be woken
is dead on arrival, so cadence is not an optional afterthought either.
A prior pass at this (adding a NEW deterministic `memcode personal setup`
CLI wizard, mirroring `gateway setup`'s bufio-prompt mechanism) was wrong
and reverted before commit — that reproduces the exact CLI-surface-
sprawl complaint instead of fixing it. The cockpit conversation IS the
wizard; it needed a better prompt, not a new command.
Root cause of the whole CLI-surface complaint: the interactive cockpit
(`memcode personal`, no args) is supposed to be the entire interface, but
its tool set had NO way to create a new Personal Agent. pa_objective only
supported show/set, and set required the objective to already exist —
creation was only reachable through the CLI's `personal create`. That's why
every explanation kept surfacing a CLI command: until now, it was the only
path that actually worked.
- internal/agent/tools/personal.go: new pa_create tool (agent name +
objective). Documented as the required first call for a brand-new agent.
- cmd/personal_cockpit.go: paCreate mirrors personalCreate's steps (register
in gateway.yaml, open the store, create the objective) and is dispatched
BEFORE the paStore existence gate, since the agent doesn't exist yet.
pa_resource's grant action now defaults type=filesystem/mode=read when
omitted too (same inference the CLI's `resources add` already got), and
pa_resource/pa_policy's grant/revoke/stage/approve paths now call
WriteConfigMirror — that was wired into the CLI paths last commit but not
into the cockpit paths, which are the ones actually used.
- cmd/personal.go: every `memcode personal` subcommand is now Hidden — they
remain fully callable for scripts, but `memcode personal --help` no longer
reads like a CLI to memorize. The interface is `memcode personal` with no
args, then plain conversation.
- internal/doctrine/prompts.go: the cockpit's own system prompt now says
outright never to tell the user to run a CLI command — it does the work
itself, right there — and that pa_create, not pa_objective, is the first
call for a new agent.
Verified end to end: pa_create then a bare-path pa_resource grant (no type,
no mode) both succeed through personalExecute exactly as the cockpit would
call them, no CLI involved.
Hiding the subcommands last commit wasn't enough — they still existed as a
parallel, scriptable way to do everything, which is real duplicate logic
(and already caused drift: a fix landed in one path and not the other,
twice, in this same review). Deleted them outright:
cmd/personal_resources.go, personal_policy.go, personal_triggers.go,
personal_status.go, personal_browser.go, personal_test.go — all removed.
cmd/personal.go cut down to just the cockpit launcher (~90 lines from
~370): `memcode personal` opens the interactive session, full stop.
Two operations the CLI covered had no pa_* equivalent yet, so they'd have
been silently lost:
- pa_doctor: the health check (home layout, objective, approved policy,
generated workspace, sandbox availability, trigger/pending-interaction
counts) that `personal doctor` used to run.
- pa_browser_setup: the existing-Chrome prerequisite check + real, bounded
chrome-devtools-mcp connection attempt that `personal browser setup` used
to run. Same logic, now a tool call instead of a command — it still
doesn't (can't) click Chrome's own consent dialog for the user.
Both dispatch through personalExecute exactly like every other pa_* tool.
internal/agent/tools/personal.go: added PaDoctor and PaBrowserSetup to the
registry.
cmd/personal_cmd_test.go: rewritten from scratch to call personalExecute
directly — no CLI args, no cobra Execute(), no rootCmd. This IS the
interface being tested now, so this is what the tests exercise. Also
asserts the config-mirror files (policy.yaml, resources.yaml) actually
reflect approvals/revokes done through the tool path, since that's the one
that matters.
Verified: `memcode personal create foo bar` no longer resolves to anything
— it just opens the cockpit with those words ignored as stray args, because
there is no "create" subcommand left to match. The interface is
conversation with pa_* tools behind it, not CLI argument syntax.
objective.md/policy.yaml/resources.yaml split one agent's config across
three files by table for no real reason — a person has to know to look in
three places for one agent's setup. Consolidated into a single config.yaml
(objective, policies, resources as sections) with proper snake_case field
names. Same authority split as before (this is a mirror; policy hash-
approval and the run journal stay in personal.db, and why, is unchanged) —
just one file to read instead of three.
Consolidation step 1 of folding Personal Agents into the ordinary agent
system. There were three partial suspend/resume designs in the tree:
1. internal/agent/runtime/continuation.go — typed, atomic, tested, and with
ZERO production callers.
2. jobs.Job's InteractionID/WaitingReason/ContinuationVersion/WaitingAt/
ResumedAt — declared, never written by anything.
3. A hand-rolled map[string]any in the personal executive — the only one
actually running, and (until a fix earlier in this branch) the only one
that wasn't crash-safe.
New internal/agent/continuation is the single implementation. It lives in its
own small package rather than in internal/agent/runtime so the executive can
suspend without importing the whole session runtime (and so neither side ends
up depending on the other).
It has to serve two genuinely different callers, which is why the previous
attempt at a shared API didn't fit:
- An interactive session already holds the conversation, so it only needs the
missing assistant/answer pair back.
- An unattended executive keeps NO transcript — it rebuilds context from
durable state each wake — so its continuation must carry the whole
conversation. Hence the optional Messages field.
They also differ on when marking resolved is safe, so building the resume
messages is split from marking:
- ResumeMessages() builds without marking. The executive uses this and marks
only once the resumed run reaches a terminal state, so a transient model
error leaves the answer re-giveable instead of stranding it.
- Resolve() does both, for a caller with a human right there who can simply
ask again.
Also deleted the five never-written Job fields (2 above). StatusWaiting stays:
it is a meaningful status in the job state machine and check_delegate
correctly treats it as non-terminal.
Tests: ported both original round-trip/validation tests, plus new coverage for
the transcript-carrying path and MarkResolved. The executive's suspend/resume
test now asserts through the continuation API instead of a hardcoded filename.
… admin
Consolidation step 2. "Personal Agent" is no longer a kind of agent. There is
one Agent abstraction; autonomy is orthogonal settings on it, managed through
the admin cockpit that already manages agents.
CONFIG. gwconfig.Agent drops Kind and gains four settings:
objective — the durable outcome it works toward
autonomous — whether it may act on that unprompted
browser — ephemeral (default) | existing_chrome
paused — stop unattended wakes without deleting anything
objective and autonomous are deliberately SEPARATE grants, which is the
correction that motivated this design. They answer different questions ("what
is it for" vs "may it act unasked"), and all four combinations are meaningful:
autonomous + objective → unattended objective pursuit (the old Personal Agent)
autonomous, no objective→ scheduled work UNDER GOVERNANCE — this is new, and
closes the long-standing gap where a cron-fired
agent ran unattended with no policy gate, no action
journal, and no way to pause and ask
objective, not autonomous → a goal you work on together; wakes only on demand
neither → an ordinary conversational agent, exactly as before
Autonomy therefore gates GOVERNANCE, not capability, and applies to any run of
the agent whether or not an objective exists.
SCHEDULING. The second cron implementation is gone as a user-facing concept.
Recurring cadence is an ordinary `schedules:` entry: set agent=<name>, leave
deliver_to empty, and gw_schedule routes the wake to the agent itself via the
renamed internal sink (personal → agent). The DB-backed loop now fires only
the agent's OWN self-scheduled next-wakes (schedule_wake, "come back in 45
minutes") — a genuinely different thing from human-authored cadence, and the
only part that needs to be writable from inside a run.
TOOLS. internal/agent/tools/personal.go (pa_*) is deleted. The admin registry
gains gw_policy, gw_grant, gw_wake, gw_inbox, gw_answer, gw_journal, gw_doctor,
gw_browser; gw_agent gains objective/autonomous/browser/pause/resume actions.
Everything dispatches through the existing adminExecute + approval gate, and
gw_schedule's shared BuildSchedule validation now covers agent cadence too, so
the surfaces cannot drift the way the parallel path did.
COCKPIT. `memcode personal` is deleted outright — cmd/personal.go,
cmd/personal_cockpit.go, Session.SetPersonal, personalMode, the personal and
personal_admin doctrines. The Personal setup walkthrough is merged into
adminDoctrine, generalized, and now teaches the two gates explicitly: granting
an objective is not granting autonomy, and the second must be confirmed on its
own.
EXECUTIVE. Executive.Objective is read from configuration rather than the
store, so the objective has ONE source that a human edits and the gateway
hot-reloads. The per-agent config.yaml mirror drops the objective for the same
reason — it lives in gateway.yaml, which is already a readable file.
Tests are rewritten against adminExecute (there is no other surface). New
coverage asserts the orthogonality directly, since a single overloaded switch
is exactly what this change exists to prevent: an objective alone must not
confer autonomy, a non-affirmative value must not grant it by typo, an
autonomous agent with no objective is still governed but refuses to invent
work, and an autonomous agent's schedule defaults to the agent route.
…ard both
Consolidation step 3, finishing the merge.
RENAME. internal/personal -> internal/agent/autonomy. The package was never
about a "personal" species; it is the machinery an agent uses when running
unattended toward an objective. Terminology in comments follows.
SECOND CRON PARSER DELETED. NextDue now understands only the kinds an AGENT
writes for itself from inside a run (one_shot / next_wake — always a single
future instant, e.g. schedule_wake "come back in 45 minutes"). Recurring
cadence a human configures is not a trigger at all: it is an ordinary gateway
schedule delivering to agent:<name>, validated once by gwconfig. With
recurring kinds gone, ClaimDueTrigger's reschedule branch was unreachable, so
firing a wake now simply completes it — and the atomic `last_fired_at IS ?`
claim still keeps two gateway processes from double-firing.
GUARDS (internal/guard/singletons_test.go). The duplication this branch kept
producing was not a one-off: fixes landed on one path and not the other,
repeatedly. Four invariants now fail the build instead:
- TestSingleCronParser — only internal/gateway parses cron. (This one caught a
real leftover while being written: the autonomy package was still importing
robfig/cron, which is what prompted the deletion above.)
- TestSingleSuspensionImplementation — no hand-rolled continuation formats
outside internal/agent/continuation.
- TestNoSecondCockpit — no SetPersonal/personalMode/pa_* anywhere.
- TestNoAgentKind — autonomy is orthogonal settings, never a kind
discriminator; a "kind" field is what made Personal a separate species.
The guards deliberately skip vendored forks, .memcode session snapshots, and
desktop/node_modules (the same symlink that breaks the two pre-existing
go-list guard failures).
Store tests updated: the trigger test now exercises a real self-scheduled wake
and asserts it completes rather than reschedules, and NextDue's test asserts
that interval/cron are REJECTED — accepting them again would rebuild the
second scheduler.
Consolidation step 4. An agent had two unrelated places to put what it knew:
a structured `facts` table only reachable through bespoke tools, and the
memory.md every memcode agent already has and the runtime already injects. One
of those a human can read. Collapse to it.
BLOCKING PREREQUISITE, done first. delegate/check_delegate were using the
facts table as a keyed index — writing "delegation.<job-id>" so a later wake
could find the action to close out. Its own comment admitted why ("facts are
the only durable log delegate can write to without a schema migration"). This
is that migration: actions gain a job_id column plus LinkActionJob /
ActionForJob, and the link now lives on the action it actually describes.
Removing facts before this would have broken check_delegate silently.
note_fact becomes `remember`: one plain sentence appended to memory.md, read
back into every wake's state summary. Append-only and deduplicated, because
every line replays into the model on each future wake — a re-learned fact must
not grow the file forever.
Deleted with it: environment.go (EnvironmentModel/StructuredFact, referenced
nowhere but its own test) and UsableForExternalRepresentation.
The tradeoff is deliberate and documented in memory.go rather than left
implicit. Prose cannot distinguish "you told me this" from "I inferred it from
your resume" from "a website said so", nor mark a claim safe to assert on the
user's behalf, nor mark it stale. That distinction becomes load-bearing the
moment an agent fills in a form or sends a message stating something about the
user. It is cheap to drop TODAY only because the gate was never wired —
nothing read Confirmed to gate anything; the sole caller of
UsableForExternalRepresentation was its own test. When external
representation needs provenance, it should come back as its own
machine-checkable thing in the store alongside policies and the action
journal, NOT by reviving this table.
… as a mode
Consolidation step 5, finishing the merge.
MIGRATION. YAML silently ignores unknown fields, so an existing
`kind: personal` agent would have loaded as an ordinary one — still listed,
apparently fine, and never waking again. For a change that removes an agent's
authority to run on its own, silence is the wrong failure. gwconfig.Agent keeps
a LegacyKind field for the sole purpose of rejecting it, with the one-line fix
in the error:
agent "demo" still uses the removed `kind: personal` setting. Autonomy is now
explicit: replace it with `autonomous: true` (and an `objective:` ...), or
just delete the `kind:` line if it should not. Its home ... is untouched
Verified against the real local config, which has exactly that agent. The
guard's TestNoAgentKind is refined to allow this one mention while still
banning any behavioral branch on it.
DOCS. docs/personal-agents.md -> docs/autonomous-agents.md, rewritten around
the two orthogonal grants and the four combinations they produce, with the
scheduling split (your cadence = an ordinary schedule; the agent's own next
wake = self-scheduled), the fail-closed existing-Chrome rule, and the
memory.md provenance limitation stated rather than glossed.
docs/design/personal-agents.md -> docs/design/autonomous-agents.md keeps the
original design contract but opens with a note on what the implementation
revised and why, including the orthogonality correction. It is a design record;
rewriting history there would lose the reasoning.
README drops "three ways to run it" (there are two) and describes autonomy as a
setting rather than a product tier, leading with the case that actually
motivated it: a scheduled agent that is finally policy-gated, journaled, and
able to stop and ask.
The store file was still named personal.db. It holds state for any agent
running unattended, not a species of agent — rename it, and note in place that
it is opened lazily so an ordinary conversational agent never grows one.
That claim is now a test rather than a comment: TestOrdinaryAgentGetsNoAutonomyStore
creates a plain agent, exercises the admin surface against it, and asserts no
agent.db appears — the governance machinery costs nothing until asked for.
Also covers the remaining orthogonality row: an agent WITH an objective but
WITHOUT autonomy still wakes on demand, reaching the policy gate rather than
being refused for lacking autonomy. Autonomy governs unprompted action, not
whether a human may ask.
All four combinations of (objective, autonomous) now have a test.
@timothyerwintimothyerwin changed the title Personal Agents: durable objectives, delegation policy, triggers, browser brokerAutonomous AgentsAug 30, 2026
@timothyerwin
timothyerwin merged commit 8a2e6b3 into mainAug 30, 2026
2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@timothyerwin
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Autonomous Agents - #5

Merged
timothyerwin merged 13 commits into
mainfrom
feat/personal-agents
Aug 30, 2026
Merged

Autonomous Agents#5
timothyerwin merged 13 commits into
mainfrom
feat/personal-agents

Conversation

@timothyerwin

Copy link
Copy Markdown
Contributor

Summary

Adds Personal Agents: long-lived agent identities with a durable objective/subgoal/facts store, a delegation policy that governs what their executive loop may do (allowed tools, consequence classes, filesystem/browser/MCP resource grants, budgets), ask_user suspend/resume for human-in-the-loop interactions, durable cron/interval/one-shot triggers polled by the gateway, and a browser broker that leases exclusive browser control to one run at a time.

New surfaces: memcode personal CLI (create, approve-policy, resources, triggers, answer, status, cockpit), gw_agent/gw_* admin tools, and the gateway's internal personal wake route.

Review fixes included

This branch went through a review pass (/code-review high) before commit; the following were found and fixed:

  • Duplicate-answer replay: personal answer now refuses to re-resume an already-answered interaction, and a resumed run that re-suspends on a new question now marks its prior continuation resolved — closing a window where a retried personal answer call could re-execute real tool side effects.
  • Triggers silently dead after hot-add: the gateway now always registers the personal wake route and always starts the trigger poll loop (it already re-reads config every tick), so a Personal Agent created after boot has its triggers picked up without a gateway restart.
  • Expired browser lease still granted mutation rights: CanMutate/OwnPage now check lease expiry like Authenticate already did.
  • Resource grant ID collisions: IDs now use a monotonic timestamp instead of len(locator), matching the cockpit's existing scheme.
  • Delegation policy narrowing gaps: IsRestriction/NarrowPolicy now also checks filesystem roots, browser origins, MCP tools, budget fields, and quiet hours, not just tools/consequence-classes/concurrency/depth.
  • Silent no-op on bad trigger id: personal triggers pause/resume now errors instead of reporting success when nothing matched.
  • Stale tool description: gw_agent no longer advertises the unsupported "inspect" action.
  • Non-atomic suspension writes: suspension continuation files are now written via atomicfile.WriteFile.
  • Per-tick DB churn: the trigger poll loop now reuses one DB connection per agent across ticks and filters due triggers in SQL instead of opening/closing a connection and Go-filtering every row every 15s.

Testing

  • go build ./... — clean
  • go vet ./... — clean
  • go test ./... — all pass except two pre-existing, unrelated failures in internal/guard caused by a symlink under desktop/node_modules/@electron/notarize breaking go list (not touched by this branch)

… triggers, browser broker)
Implements the Personal Agents feature: long-lived agent identities with
their own durable objective, subgoals, and facts store; a delegation
policy model (allowed tools, consequence classes, filesystem/browser/MCP
grants, budgets) that governs what a Personal Agent's executive loop may
do; ask_user suspend/resume for human-in-the-loop interactions; durable
triggers (cron/interval/one-shot) polled by the gateway to wake an agent;
resource grants (filesystem, mcp, command, channel); and a browser broker
that leases exclusive browser control to one run at a time.
New surfaces: `memcode personal` CLI (create, approve-policy, resources,
triggers, answer, status, cockpit), gw_agent/gw_* admin tools, and the
gateway's internal personal wake route.
Includes fixes from review of this branch:
- personal answer refuses to re-resume an already-answered interaction,
and a resumed run that re-suspends now marks its prior continuation
resolved — closing a window where a retried answer could replay tool
side effects.
- The gateway always registers the personal wake route and always starts
the trigger poll loop, so a Personal Agent added after boot has its
triggers picked up without a restart.
- Browser broker leases now check expiry in CanMutate/OwnPage, matching
Authenticate, so an expired lease can no longer authorize mutation.
- Resource grant IDs use a monotonic timestamp instead of locator length,
avoiding same-length-locator collisions.
- Delegation policy narrowing (IsRestriction/NarrowPolicy) now also
checks filesystem roots, browser origins, MCP tools, budget fields, and
quiet hours, not just tools/consequence-classes/concurrency/depth.
- `personal triggers pause/resume` errors on an unknown trigger id instead
of reporting success silently.
- gw_agent's description no longer advertises the unsupported "inspect"
action.
- Suspension continuation files are written atomically.
- The trigger poll loop reuses one DB connection per agent across ticks
and filters due triggers in SQL, instead of opening/closing a
connection and scanning every trigger row every 15s.
… CI lint
Closes the biggest gap from review: the Personal Agent executive had no way
to reach browser/MCP/shell/filesystem capability — it was a second, weaker
agent runtime sitting beside the real memcode engine instead of commanding
it. Two new executive tools fix that:
- delegate: spawns a policy-scoped worker as a real detached `memcode run`
job (internal/jobs.SpawnWithSpec), using the delegation.go scaffolding
(ExecutionEnvelope, ValidateDelegation, PrepareRunDirectory) that already
existed but was never called from anywhere. The worker gets whatever
toolsets it's granted — browser, MCP, shell, filesystem, skills — same as
any ordinary memcode agent, not the executive's fixed 7 tools. Requested
toolsets/consequences must be a subset of the agent's own approved policy;
MaxDelegationDepth gates whether delegation is allowed at all.
- check_delegate: reads a delegated job's status/result back on a later
wake (RunOnce is one bounded wake, so the result can't arrive in the same
call) and closes out its action-journal entry.
Also makes jobs.SpawnSpec.ToolPolicy a REAL restriction instead of recorded-
only metadata: new hidden --allow-tools/--deny-tools flags on `memcode run`
bind to the same SetToolPolicy enforcement an ordinary gateway-bound agent
gets from its config, and SpawnWithSpec now passes them through argv.
Scope notes, called out explicitly rather than implied as done:
- "browser" toolset reuses the existing ephemeral Chrome session (the same
one --chrome already drives) — NOT the user's own already-running Chrome.
Attaching to a persistent/remote Chrome needs the broker/RemoteConfig
wiring in internal/browser, which is declared but has zero callers and
zero implementation; that's separate follow-up work, not done here.
- Filesystem/MCP resource grants are not OS-sandboxed to a spawned worker
(memcode has no such sandbox anywhere) — governance here is the same
approval/audit model as the rest of the codebase (tool-level allow/deny +
the action journal), not a hard jail.
Also fixes two CI lint failures on the open PR:
- internal/gateway/state/state.go: removed formatTime, an unused helper
(staticcheck U1000).
- internal/personal/store.go: decapitalized a Go error string (ST1005).
Corrects course on the browser gap flagged after the last review: existing-
Chrome access is core to Personal Agents, not deferrable follow-up. A
Personal Agent's whole premise — acting as the user across their real
accounts (Gmail, LinkedIn, an ATS, internal dashboards) — requires the
user's actual signed-in session, not a fresh ephemeral profile with no
cookies. Per docs/design/personal-agents.md's already-specified "browser
broker trust boundary": Personal Agents connect to the user's existing
Chrome through a gateway-owned broker and a permission-protected local
socket, fail closed on any connection/auth problem, and never silently fall
back to another profile.
- internal/browser/broker/{server,client}.go: exposes the (already-reviewed)
in-process Broker over a Unix socket, so a delegated worker — a SEPARATE
OS process spawned via jobs.SpawnWithSpec — can Acquire/Release/OwnPage/
CanMutate against the SAME broker the gateway owns, not a fresh one.
Socket is 0600 inside gwconfig.Dir() (0700). Round-trip tested, including
that a concurrent Acquire is correctly rejected and a release unblocks the
next one.
- internal/gateway/server: the gateway now owns one *broker.Broker for its
whole lifetime and serves it on that socket unconditionally at startup
(cheap — a local listener), so it's there the moment a delegate call needs
it, no restart required after `memcode personal browser setup`.
- internal/personal: browserModeFor's "browser" toolset now defaults to
BrowserExistingChrome, not ephemeral — "browser:ephemeral" is the explicit
opt-down for tasks that genuinely want a fresh, logged-out profile. Before
spawning anything, the delegate tool verifies the broker is reachable and
fails the tool call closed (a normal tool_result error, not a spawned job)
if it isn't — verified by test, alongside the success path asserting the
spawned job's SpawnSpec actually carries BrowserMode=existing_chrome.
- internal/jobs: SpawnWithSpec passes existing_chrome through as
--browser-session (+ --browser-agent/--browser-run for the lease identity)
instead of --chrome.
- cmd/run.go: --browser-session existing_chrome acquires the broker lease
itself (via the client) and, only on success, wires an ad-hoc
chrome-devtools-mcp --autoConnect MCP server into the session
(Session.SetExtraMCPServers, new) — the SAME chrome-devtools-mcp package
this browser experience is built on (see internal/browser/remote.go). On
any failure it returns an error and exits; it never calls
SetBrowserEnabled(true) (ephemeral) as a fallback.
- cmd/personal_browser.go: `memcode personal browser setup` checks npx and
the broker socket, then attempts a REAL bounded (10s) chrome-devtools-mcp
connection and reports the tool count or the actual connect errors — it
does not pretend to click Chrome's own "Allow" dialog, which only the user
can do.
- Also fixes a real bug found while testing this: ValidateDelegation treated
an empty parent.AllowedTools as "allows no tools by name", the opposite of
the len(...)>0-means-restricted convention Executive.allowedTools already
uses — every delegate call with a name-unrestricted policy was rejected.
Scope note, stated rather than implied: the literal Chrome-side consent flow
(the "Allow" dialog, actually clicking it) cannot be exercised or verified
in this environment — no live Chrome, no human to click it. Everything
up to that point (broker lifecycle, lease exclusivity, fail-closed wiring,
the actual chrome-devtools-mcp CLI invocation, verified against the real
published package) is built and tested; the dialog itself needs the user's
own machine.
…setup via the cockpit
Three separate complaints from review, addressed at the right layer each:
1. "It's just a file ref" — `resources add` required a redundant --mode read
(already the default) and a mandatory <type> argument even for the
overwhelming common case. A bare path is now enough:
memcode personal resources add jobhunt ~/resume.md
Non-filesystem grants (mcp/command/channel) still need the type spelled
out, since there's nothing to infer it from. Also fixes the real bug this
surfaced: CanonicalFilesystemGrant rejected any non-directory path, even
though PathWithinGrant already handles a single-file grant correctly
(rel == "." on an exact match) — you could not actually grant a single
file like a resume before this, only whole directories. `personal create`
also gained a repeatable --grant flag to fold "create it, then give it
something to read" into one command.
2. "Not this CLI-command-argument garbage" — Personal Agent state (objective,
policy, resources) lived ONLY in personal.db, reachable exclusively
through bespoke commands, unlike every other piece of memcode config
(gateway.yaml, .mcp.json, CLAUDE.md) which is a plain file. WriteConfigMirror
now regenerates objective.md/policy.yaml/resources.yaml in the agent's home
on every mutation — real, readable, diffable files. The run journal
(actions/triggers/interactions) deliberately stays in SQLite: it needs
atomic claim/complete semantics under concurrent access (gateway wake loop,
CLI, and cockpit can all touch the same agent) that flat files don't get
for free — a correctness reason, not a habit.
3. "It should gather requirements, do HITL, pre-approve permissions, set up
the runtime — like our other agents already do via telegram/model setup"
— `memcode personal` (no args) is ALREADY the interactive cockpit (like
`memcode admin`) with the pa_* tools wired for exactly this. The actual gap
was its system prompt: personalAdminDoctrine listed tool mechanics with no
instruction to gather-then-propose-then-approve. Rewrote it to require a
real walkthrough for first-time creation — reason about what resources/
toolsets/consequences the objective needs (asking the user rather than
guessing), present the concrete plan in plain language BEFORE touching
anything, then apply everything at once on approval, including wake
cadence (trigger) — an agent with no trigger and no plan to ever be woken
is dead on arrival, so cadence is not an optional afterthought either.
A prior pass at this (adding a NEW deterministic `memcode personal setup`
CLI wizard, mirroring `gateway setup`'s bufio-prompt mechanism) was wrong
and reverted before commit — that reproduces the exact CLI-surface-
sprawl complaint instead of fixing it. The cockpit conversation IS the
wizard; it needed a better prompt, not a new command.
Root cause of the whole CLI-surface complaint: the interactive cockpit
(`memcode personal`, no args) is supposed to be the entire interface, but
its tool set had NO way to create a new Personal Agent. pa_objective only
supported show/set, and set required the objective to already exist —
creation was only reachable through the CLI's `personal create`. That's why
every explanation kept surfacing a CLI command: until now, it was the only
path that actually worked.
- internal/agent/tools/personal.go: new pa_create tool (agent name +
objective). Documented as the required first call for a brand-new agent.
- cmd/personal_cockpit.go: paCreate mirrors personalCreate's steps (register
in gateway.yaml, open the store, create the objective) and is dispatched
BEFORE the paStore existence gate, since the agent doesn't exist yet.
pa_resource's grant action now defaults type=filesystem/mode=read when
omitted too (same inference the CLI's `resources add` already got), and
pa_resource/pa_policy's grant/revoke/stage/approve paths now call
WriteConfigMirror — that was wired into the CLI paths last commit but not
into the cockpit paths, which are the ones actually used.
- cmd/personal.go: every `memcode personal` subcommand is now Hidden — they
remain fully callable for scripts, but `memcode personal --help` no longer
reads like a CLI to memorize. The interface is `memcode personal` with no
args, then plain conversation.
- internal/doctrine/prompts.go: the cockpit's own system prompt now says
outright never to tell the user to run a CLI command — it does the work
itself, right there — and that pa_create, not pa_objective, is the first
call for a new agent.
Verified end to end: pa_create then a bare-path pa_resource grant (no type,
no mode) both succeed through personalExecute exactly as the cockpit would
call them, no CLI involved.
Hiding the subcommands last commit wasn't enough — they still existed as a
parallel, scriptable way to do everything, which is real duplicate logic
(and already caused drift: a fix landed in one path and not the other,
twice, in this same review). Deleted them outright:
cmd/personal_resources.go, personal_policy.go, personal_triggers.go,
personal_status.go, personal_browser.go, personal_test.go — all removed.
cmd/personal.go cut down to just the cockpit launcher (~90 lines from
~370): `memcode personal` opens the interactive session, full stop.
Two operations the CLI covered had no pa_* equivalent yet, so they'd have
been silently lost:
- pa_doctor: the health check (home layout, objective, approved policy,
generated workspace, sandbox availability, trigger/pending-interaction
counts) that `personal doctor` used to run.
- pa_browser_setup: the existing-Chrome prerequisite check + real, bounded
chrome-devtools-mcp connection attempt that `personal browser setup` used
to run. Same logic, now a tool call instead of a command — it still
doesn't (can't) click Chrome's own consent dialog for the user.
Both dispatch through personalExecute exactly like every other pa_* tool.
internal/agent/tools/personal.go: added PaDoctor and PaBrowserSetup to the
registry.
cmd/personal_cmd_test.go: rewritten from scratch to call personalExecute
directly — no CLI args, no cobra Execute(), no rootCmd. This IS the
interface being tested now, so this is what the tests exercise. Also
asserts the config-mirror files (policy.yaml, resources.yaml) actually
reflect approvals/revokes done through the tool path, since that's the one
that matters.
Verified: `memcode personal create foo bar` no longer resolves to anything
— it just opens the cockpit with those words ignored as stray args, because
there is no "create" subcommand left to match. The interface is
conversation with pa_* tools behind it, not CLI argument syntax.
objective.md/policy.yaml/resources.yaml split one agent's config across
three files by table for no real reason — a person has to know to look in
three places for one agent's setup. Consolidated into a single config.yaml
(objective, policies, resources as sections) with proper snake_case field
names. Same authority split as before (this is a mirror; policy hash-
approval and the run journal stay in personal.db, and why, is unchanged) —
just one file to read instead of three.
Consolidation step 1 of folding Personal Agents into the ordinary agent
system. There were three partial suspend/resume designs in the tree:
1. internal/agent/runtime/continuation.go — typed, atomic, tested, and with
ZERO production callers.
2. jobs.Job's InteractionID/WaitingReason/ContinuationVersion/WaitingAt/
ResumedAt — declared, never written by anything.
3. A hand-rolled map[string]any in the personal executive — the only one
actually running, and (until a fix earlier in this branch) the only one
that wasn't crash-safe.
New internal/agent/continuation is the single implementation. It lives in its
own small package rather than in internal/agent/runtime so the executive can
suspend without importing the whole session runtime (and so neither side ends
up depending on the other).
It has to serve two genuinely different callers, which is why the previous
attempt at a shared API didn't fit:
- An interactive session already holds the conversation, so it only needs the
missing assistant/answer pair back.
- An unattended executive keeps NO transcript — it rebuilds context from
durable state each wake — so its continuation must carry the whole
conversation. Hence the optional Messages field.
They also differ on when marking resolved is safe, so building the resume
messages is split from marking:
- ResumeMessages() builds without marking. The executive uses this and marks
only once the resumed run reaches a terminal state, so a transient model
error leaves the answer re-giveable instead of stranding it.
- Resolve() does both, for a caller with a human right there who can simply
ask again.
Also deleted the five never-written Job fields (2 above). StatusWaiting stays:
it is a meaningful status in the job state machine and check_delegate
correctly treats it as non-terminal.
Tests: ported both original round-trip/validation tests, plus new coverage for
the transcript-carrying path and MarkResolved. The executive's suspend/resume
test now asserts through the continuation API instead of a hardcoded filename.
… admin
Consolidation step 2. "Personal Agent" is no longer a kind of agent. There is
one Agent abstraction; autonomy is orthogonal settings on it, managed through
the admin cockpit that already manages agents.
CONFIG. gwconfig.Agent drops Kind and gains four settings:
objective — the durable outcome it works toward
autonomous — whether it may act on that unprompted
browser — ephemeral (default) | existing_chrome
paused — stop unattended wakes without deleting anything
objective and autonomous are deliberately SEPARATE grants, which is the
correction that motivated this design. They answer different questions ("what
is it for" vs "may it act unasked"), and all four combinations are meaningful:
autonomous + objective → unattended objective pursuit (the old Personal Agent)
autonomous, no objective→ scheduled work UNDER GOVERNANCE — this is new, and
closes the long-standing gap where a cron-fired
agent ran unattended with no policy gate, no action
journal, and no way to pause and ask
objective, not autonomous → a goal you work on together; wakes only on demand
neither → an ordinary conversational agent, exactly as before
Autonomy therefore gates GOVERNANCE, not capability, and applies to any run of
the agent whether or not an objective exists.
SCHEDULING. The second cron implementation is gone as a user-facing concept.
Recurring cadence is an ordinary `schedules:` entry: set agent=<name>, leave
deliver_to empty, and gw_schedule routes the wake to the agent itself via the
renamed internal sink (personal → agent). The DB-backed loop now fires only
the agent's OWN self-scheduled next-wakes (schedule_wake, "come back in 45
minutes") — a genuinely different thing from human-authored cadence, and the
only part that needs to be writable from inside a run.
TOOLS. internal/agent/tools/personal.go (pa_*) is deleted. The admin registry
gains gw_policy, gw_grant, gw_wake, gw_inbox, gw_answer, gw_journal, gw_doctor,
gw_browser; gw_agent gains objective/autonomous/browser/pause/resume actions.
Everything dispatches through the existing adminExecute + approval gate, and
gw_schedule's shared BuildSchedule validation now covers agent cadence too, so
the surfaces cannot drift the way the parallel path did.
COCKPIT. `memcode personal` is deleted outright — cmd/personal.go,
cmd/personal_cockpit.go, Session.SetPersonal, personalMode, the personal and
personal_admin doctrines. The Personal setup walkthrough is merged into
adminDoctrine, generalized, and now teaches the two gates explicitly: granting
an objective is not granting autonomy, and the second must be confirmed on its
own.
EXECUTIVE. Executive.Objective is read from configuration rather than the
store, so the objective has ONE source that a human edits and the gateway
hot-reloads. The per-agent config.yaml mirror drops the objective for the same
reason — it lives in gateway.yaml, which is already a readable file.
Tests are rewritten against adminExecute (there is no other surface). New
coverage asserts the orthogonality directly, since a single overloaded switch
is exactly what this change exists to prevent: an objective alone must not
confer autonomy, a non-affirmative value must not grant it by typo, an
autonomous agent with no objective is still governed but refuses to invent
work, and an autonomous agent's schedule defaults to the agent route.
…ard both
Consolidation step 3, finishing the merge.
RENAME. internal/personal -> internal/agent/autonomy. The package was never
about a "personal" species; it is the machinery an agent uses when running
unattended toward an objective. Terminology in comments follows.
SECOND CRON PARSER DELETED. NextDue now understands only the kinds an AGENT
writes for itself from inside a run (one_shot / next_wake — always a single
future instant, e.g. schedule_wake "come back in 45 minutes"). Recurring
cadence a human configures is not a trigger at all: it is an ordinary gateway
schedule delivering to agent:<name>, validated once by gwconfig. With
recurring kinds gone, ClaimDueTrigger's reschedule branch was unreachable, so
firing a wake now simply completes it — and the atomic `last_fired_at IS ?`
claim still keeps two gateway processes from double-firing.
GUARDS (internal/guard/singletons_test.go). The duplication this branch kept
producing was not a one-off: fixes landed on one path and not the other,
repeatedly. Four invariants now fail the build instead:
- TestSingleCronParser — only internal/gateway parses cron. (This one caught a
real leftover while being written: the autonomy package was still importing
robfig/cron, which is what prompted the deletion above.)
- TestSingleSuspensionImplementation — no hand-rolled continuation formats
outside internal/agent/continuation.
- TestNoSecondCockpit — no SetPersonal/personalMode/pa_* anywhere.
- TestNoAgentKind — autonomy is orthogonal settings, never a kind
discriminator; a "kind" field is what made Personal a separate species.
The guards deliberately skip vendored forks, .memcode session snapshots, and
desktop/node_modules (the same symlink that breaks the two pre-existing
go-list guard failures).
Store tests updated: the trigger test now exercises a real self-scheduled wake
and asserts it completes rather than reschedules, and NextDue's test asserts
that interval/cron are REJECTED — accepting them again would rebuild the
second scheduler.
Consolidation step 4. An agent had two unrelated places to put what it knew:
a structured `facts` table only reachable through bespoke tools, and the
memory.md every memcode agent already has and the runtime already injects. One
of those a human can read. Collapse to it.
BLOCKING PREREQUISITE, done first. delegate/check_delegate were using the
facts table as a keyed index — writing "delegation.<job-id>" so a later wake
could find the action to close out. Its own comment admitted why ("facts are
the only durable log delegate can write to without a schema migration"). This
is that migration: actions gain a job_id column plus LinkActionJob /
ActionForJob, and the link now lives on the action it actually describes.
Removing facts before this would have broken check_delegate silently.
note_fact becomes `remember`: one plain sentence appended to memory.md, read
back into every wake's state summary. Append-only and deduplicated, because
every line replays into the model on each future wake — a re-learned fact must
not grow the file forever.
Deleted with it: environment.go (EnvironmentModel/StructuredFact, referenced
nowhere but its own test) and UsableForExternalRepresentation.
The tradeoff is deliberate and documented in memory.go rather than left
implicit. Prose cannot distinguish "you told me this" from "I inferred it from
your resume" from "a website said so", nor mark a claim safe to assert on the
user's behalf, nor mark it stale. That distinction becomes load-bearing the
moment an agent fills in a form or sends a message stating something about the
user. It is cheap to drop TODAY only because the gate was never wired —
nothing read Confirmed to gate anything; the sole caller of
UsableForExternalRepresentation was its own test. When external
representation needs provenance, it should come back as its own
machine-checkable thing in the store alongside policies and the action
journal, NOT by reviving this table.
… as a mode
Consolidation step 5, finishing the merge.
MIGRATION. YAML silently ignores unknown fields, so an existing
`kind: personal` agent would have loaded as an ordinary one — still listed,
apparently fine, and never waking again. For a change that removes an agent's
authority to run on its own, silence is the wrong failure. gwconfig.Agent keeps
a LegacyKind field for the sole purpose of rejecting it, with the one-line fix
in the error:
agent "demo" still uses the removed `kind: personal` setting. Autonomy is now
explicit: replace it with `autonomous: true` (and an `objective:` ...), or
just delete the `kind:` line if it should not. Its home ... is untouched
Verified against the real local config, which has exactly that agent. The
guard's TestNoAgentKind is refined to allow this one mention while still
banning any behavioral branch on it.
DOCS. docs/personal-agents.md -> docs/autonomous-agents.md, rewritten around
the two orthogonal grants and the four combinations they produce, with the
scheduling split (your cadence = an ordinary schedule; the agent's own next
wake = self-scheduled), the fail-closed existing-Chrome rule, and the
memory.md provenance limitation stated rather than glossed.
docs/design/personal-agents.md -> docs/design/autonomous-agents.md keeps the
original design contract but opens with a note on what the implementation
revised and why, including the orthogonality correction. It is a design record;
rewriting history there would lose the reasoning.
README drops "three ways to run it" (there are two) and describes autonomy as a
setting rather than a product tier, leading with the case that actually
motivated it: a scheduled agent that is finally policy-gated, journaled, and
able to stop and ask.
The store file was still named personal.db. It holds state for any agent
running unattended, not a species of agent — rename it, and note in place that
it is opened lazily so an ordinary conversational agent never grows one.
That claim is now a test rather than a comment: TestOrdinaryAgentGetsNoAutonomyStore
creates a plain agent, exercises the admin surface against it, and asserts no
agent.db appears — the governance machinery costs nothing until asked for.
Also covers the remaining orthogonality row: an agent WITH an objective but
WITHOUT autonomy still wakes on demand, reaching the policy gate rather than
being refused for lacking autonomy. Autonomy governs unprompted action, not
whether a human may ask.
All four combinations of (objective, autonomous) now have a test.
@timothyerwintimothyerwin changed the title Personal Agents: durable objectives, delegation policy, triggers, browser brokerAutonomous AgentsAug 30, 2026
@timothyerwin
timothyerwin merged commit 8a2e6b3 into mainAug 30, 2026
2 checks passed
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant

@timothyerwin