Skip to content

Repository files navigation

OpenMem — Agent Memory System for AI Coding Agents

Persistent semantic memory for AI coding agents. OpenMem indexes your agents' real session history into a local LanceDB vector store, retrieves relevant memories on demand, reflects on sessions to extract facts and lessons, and — since v2.x — exposes the whole memory layer as a native MCP server so any MCP-capable client can remember / recall without file-based integration.

Works with Claude Code and Codex CLI session formats today (evidence-based parsers), tolerates absent history gracefully elsewhere, and includes a generic file-based fallback.

Status: actively developed. The storage layer, parsers, learning cycle, MCP server, and evaluation harness are implemented and covered by a 225-test suite. LLM-backed reflection activates automatically when an API key is present; without one everything runs in documented heuristic mode. Retrieval quality is measured, not claimed — see Evaluation.

What It Does (verified)

  • Real session parsing — typed-record JSONL parsers built from on-disk evidence of Claude Code (~/.claude/projects/**/*.jsonl) and Codex CLI (~/.codex/sessions/**/rollout-*.jsonl), with noise filtering (auth-error spam, CLI echoes, sidechain transcripts) and malformed-line tolerance. Formats are documented in doc/session_formats.md.
  • Semantic memory store — LanceDB with fixed-size vector columns, deterministic IDs, float64 importance scores, automatic schema migration, BGE cross-encoder reranking when the ml extra is installed, and an honest keyword-fallback search when it is not.
  • Tiered memory — daily → weekly → long-term consolidation with stable, process-independent content hashing; re-runs are idempotent.
  • Reflection loop — per-session analysis producing facts, improvements, and memories. Mode-tagged llm or heuristic; malformed LLM output falls back visibly instead of silently.
  • Outcome-grounded improvements — improvement queue items can only be completed with linked evidence (memory id / session id / explicit user confirmation). No self-completion theater.
  • MCP serverremember, recall, context, profile, stats, forget tools over stdio; see doc/mcp_integration.md.
  • Evaluation harness — golden retrieval benchmark with recall@k / MRR / nDCG@k / fallout metrics and a regression gate wired into the test suite.
  • 11 agent integrations — skill/context-file installation for Claude Code, Codex CLI, Cursor, VS Code, Windsurf, Qwen Code, OpenCode, Antigravity IDE, Kilo CLI, OpenClaw, plus a generic adapter — all generated from a single template source (bin/generate_skills.py) so they cannot drift apart.

Quick Start

git clone https://github.com/ncdevshiv/openmem.git
cd openmem
# create an environment and install (core deps only)
python -m venv .venv && .venv\Scripts\activate # Windows
pip install -e .# initialize the store and check health
python main.py status
# index your real agent history and run one learning cycle
python main.py run-cycle
# search your own memory
python main.py search "deepseek harness"# measure retrieval quality (writes data/eval/latest.json)
python main.py eval

Optional extras:

pip install -e ".[ml]"# torch + sentence-transformers + transformers (embeddings & reranker)
pip install -e ".[mcp]"# MCP server support (mcp>=2.0.0)
pip install -e ".[llm]"# litellm for LLM-backed reflection

Copy config.example.json to config.json if you want to override defaults; the repo never ships your machine-specific config.

Enabling AI features

LLM reflection activates automatically when a provider is reachable:

ProviderEnvironment variableDefault model
OpenAIOPENAI_API_KEYgpt-4o-mini
AnthropicANTHROPIC_API_KEYclaude-sonnet-4-20250514
GeminiGEMINI_API_KEYgemini-pro
OllamaOLLAMA_BASE_URLllama3

Without keys, zero network calls are made and reflections are tagged "mode": "heuristic". With keys, reflections are tagged "mode": "llm" and cycle reports include a reflection_modes summary. Malformed LLM output is rejected and falls back with a visible warning.

Commands

python main.py status # system health
python main.py run-cycle # full learning cycle (idempotent)
python main.py run-cycle --full # re-index from scratch window
python main.py search <query> # semantic/keyword memory search
python main.py eval [--report P] # retrieval benchmark (markdown + JSON)
python main.py profile # learned user profile
python main.py stats # store statistics
python main.py --agents # list supported adapters
python main.py --skill <agent> # install skill files for an agent

Evaluation

Retrieval quality is measured against a versioned golden set (16 queries over a deterministic 36-memory corpus: exact-term, paraphrase, and negative classes) with a regression gate wired into the test suite. Baseline (keyword-fallback mode, no reranker):

classqueriesrecall@5MRRnDCG@5fallout@5
exact_term60.9721.0001.0000.333
paraphrase61.0000.8890.9170.300
negative40.0000.0000.0000.000
aggregate160.7400.7080.7190.237

Thresholds and rationale live in eval/BASELINE.md; the gate fails CI if retrieval regresses below them. The embedder-free path is powered by a vendored ENR lexical layer (memory_store/enr_lexical.py): Porter-stemmed BM25 with a coord factor, positional phrase matching for quoted spans, cached inverted index, and per-result score_details explanations. Remaining known limits (polysemy across domains, vector-mode negative silence) are logged there as the exact targets for future semantic work.

Architecture

┌────────────────────────────────────────────────────────────┐
│ Your agent Any MCP-capable client │
│ Claude Code etc. (remember/recall/context/profile/...) │
└───────┬───────────────────────────┬────────────────────────┘
│ skill files + │ stdio (mcp>=2.0)
│ context injection │
┌───────▼───────────┐ ┌────────▼─────────┐
│ agents/* │ │ mcp_server.py │
│ real session │ └────────┬─────────┘
│ parsers + skills │ │
└───────┬───────────┘ ▼
▼ memory_store (same core)
┌────────────────────────────────────────────────────────────┐
│ learning_loop/ scheduler · indexer · patterns · reflect │
├────────────────────────────────────────────────────────────┤
│ memory_store/ vector_db (LanceDB) · tiers · user model │
│ skill generator · retrieval metrics │
├────────────────────────────────────────────────────────────┤
│ core/llm.py litellm wrapper — lazy, network-free init │
├────────────────────────────────────────────────────────────┤
│ eval/ golden corpus · queries · runner · gate │
└────────────────────────────────────────────────────────────┘

Directory map:

main.py entry point (delegates to bin/launcher.py)
mcp_server.py MCP stdio server
openmem_cli.py console-script wrapper (pip install -e .)
agents/ adapter contract (base.py) + per-agent parsers/skills
memory_store/ vector DB, tier manager, user model, skill gen, metrics
learning_loop/ scheduler, conversation indexer, pattern recognizer,
reflection engine
autonomous/ EXPERIMENTAL optimizer/evolution scaffolds (not yet
wired into the cycle — see roadmap)
core/ provider-agnostic LLM abstraction
eval/ golden corpus, queries, runner, baseline
bin/ launcher, installer, config/skill generators
doc/ session format inventory, MCP integration guide
tests/ 225-test suite (unit + integration + gates)

Privacy

Everything is local-first. Indexed content lives in data/lancedb/ (gitignored), the MCP server talks over stdio, and no network call happens unless you explicitly configure an LLM provider. Tests run hermetically in temp directories and provably never touch the live store.

Testing

python -m unittest discover -s tests # 225 tests

Includes unit tests for every module, parser fixtures mirroring real on-disk formats, MCP end-to-end subprocess tests, LLM boundary mocks, a retrieval regression gate, and leakage checks proving the suite leaves the live store byte-identical.

Roadmap

  • Reranker/embedder integration for semantic retrieval (levers already identified by the eval: morphology-aware matching, IDF weighting, semantic tie-breaking)
  • Wire autonomous/ evolution scaffolds to real fitness signals from the eval harness
  • Sleep-time consolidation ("dream cycles") measured by eval lift
  • Cross-agent shared memory namespaces via MCP

License

MIT — © Shivam Tiwari

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
GitHub - ncdevshiv/openmem · GitHub
Skip to content

Repository files navigation

OpenMem — Agent Memory System for AI Coding Agents

Persistent semantic memory for AI coding agents. OpenMem indexes your agents' real session history into a local LanceDB vector store, retrieves relevant memories on demand, reflects on sessions to extract facts and lessons, and — since v2.x — exposes the whole memory layer as a native MCP server so any MCP-capable client can remember / recall without file-based integration.

Works with Claude Code and Codex CLI session formats today (evidence-based parsers), tolerates absent history gracefully elsewhere, and includes a generic file-based fallback.

Status: actively developed. The storage layer, parsers, learning cycle, MCP server, and evaluation harness are implemented and covered by a 225-test suite. LLM-backed reflection activates automatically when an API key is present; without one everything runs in documented heuristic mode. Retrieval quality is measured, not claimed — see Evaluation.

What It Does (verified)

  • Real session parsing — typed-record JSONL parsers built from on-disk evidence of Claude Code (~/.claude/projects/**/*.jsonl) and Codex CLI (~/.codex/sessions/**/rollout-*.jsonl), with noise filtering (auth-error spam, CLI echoes, sidechain transcripts) and malformed-line tolerance. Formats are documented in doc/session_formats.md.
  • Semantic memory store — LanceDB with fixed-size vector columns, deterministic IDs, float64 importance scores, automatic schema migration, BGE cross-encoder reranking when the ml extra is installed, and an honest keyword-fallback search when it is not.
  • Tiered memory — daily → weekly → long-term consolidation with stable, process-independent content hashing; re-runs are idempotent.
  • Reflection loop — per-session analysis producing facts, improvements, and memories. Mode-tagged llm or heuristic; malformed LLM output falls back visibly instead of silently.
  • Outcome-grounded improvements — improvement queue items can only be completed with linked evidence (memory id / session id / explicit user confirmation). No self-completion theater.
  • MCP serverremember, recall, context, profile, stats, forget tools over stdio; see doc/mcp_integration.md.
  • Evaluation harness — golden retrieval benchmark with recall@k / MRR / nDCG@k / fallout metrics and a regression gate wired into the test suite.
  • 11 agent integrations — skill/context-file installation for Claude Code, Codex CLI, Cursor, VS Code, Windsurf, Qwen Code, OpenCode, Antigravity IDE, Kilo CLI, OpenClaw, plus a generic adapter — all generated from a single template source (bin/generate_skills.py) so they cannot drift apart.

Quick Start

git clone https://github.com/ncdevshiv/openmem.git
cd openmem
# create an environment and install (core deps only)
python -m venv .venv && .venv\Scripts\activate # Windows
pip install -e .# initialize the store and check health
python main.py status
# index your real agent history and run one learning cycle
python main.py run-cycle
# search your own memory
python main.py search "deepseek harness"# measure retrieval quality (writes data/eval/latest.json)
python main.py eval

Optional extras:

pip install -e ".[ml]"# torch + sentence-transformers + transformers (embeddings & reranker)
pip install -e ".[mcp]"# MCP server support (mcp>=2.0.0)
pip install -e ".[llm]"# litellm for LLM-backed reflection

Copy config.example.json to config.json if you want to override defaults; the repo never ships your machine-specific config.

Enabling AI features

LLM reflection activates automatically when a provider is reachable:

ProviderEnvironment variableDefault model
OpenAIOPENAI_API_KEYgpt-4o-mini
AnthropicANTHROPIC_API_KEYclaude-sonnet-4-20250514
GeminiGEMINI_API_KEYgemini-pro
OllamaOLLAMA_BASE_URLllama3

Without keys, zero network calls are made and reflections are tagged "mode": "heuristic". With keys, reflections are tagged "mode": "llm" and cycle reports include a reflection_modes summary. Malformed LLM output is rejected and falls back with a visible warning.

Commands

python main.py status # system health
python main.py run-cycle # full learning cycle (idempotent)
python main.py run-cycle --full # re-index from scratch window
python main.py search <query> # semantic/keyword memory search
python main.py eval [--report P] # retrieval benchmark (markdown + JSON)
python main.py profile # learned user profile
python main.py stats # store statistics
python main.py --agents # list supported adapters
python main.py --skill <agent> # install skill files for an agent

Evaluation

Retrieval quality is measured against a versioned golden set (16 queries over a deterministic 36-memory corpus: exact-term, paraphrase, and negative classes) with a regression gate wired into the test suite. Baseline (keyword-fallback mode, no reranker):

classqueriesrecall@5MRRnDCG@5fallout@5
exact_term60.9721.0001.0000.333
paraphrase61.0000.8890.9170.300
negative40.0000.0000.0000.000
aggregate160.7400.7080.7190.237

Thresholds and rationale live in eval/BASELINE.md; the gate fails CI if retrieval regresses below them. The embedder-free path is powered by a vendored ENR lexical layer (memory_store/enr_lexical.py): Porter-stemmed BM25 with a coord factor, positional phrase matching for quoted spans, cached inverted index, and per-result score_details explanations. Remaining known limits (polysemy across domains, vector-mode negative silence) are logged there as the exact targets for future semantic work.

Architecture

┌────────────────────────────────────────────────────────────┐
│ Your agent Any MCP-capable client │
│ Claude Code etc. (remember/recall/context/profile/...) │
└───────┬───────────────────────────┬────────────────────────┘
│ skill files + │ stdio (mcp>=2.0)
│ context injection │
┌───────▼───────────┐ ┌────────▼─────────┐
│ agents/* │ │ mcp_server.py │
│ real session │ └────────┬─────────┘
│ parsers + skills │ │
└───────┬───────────┘ ▼
▼ memory_store (same core)
┌────────────────────────────────────────────────────────────┐
│ learning_loop/ scheduler · indexer · patterns · reflect │
├────────────────────────────────────────────────────────────┤
│ memory_store/ vector_db (LanceDB) · tiers · user model │
│ skill generator · retrieval metrics │
├────────────────────────────────────────────────────────────┤
│ core/llm.py litellm wrapper — lazy, network-free init │
├────────────────────────────────────────────────────────────┤
│ eval/ golden corpus · queries · runner · gate │
└────────────────────────────────────────────────────────────┘

Directory map:

main.py entry point (delegates to bin/launcher.py)
mcp_server.py MCP stdio server
openmem_cli.py console-script wrapper (pip install -e .)
agents/ adapter contract (base.py) + per-agent parsers/skills
memory_store/ vector DB, tier manager, user model, skill gen, metrics
learning_loop/ scheduler, conversation indexer, pattern recognizer,
reflection engine
autonomous/ EXPERIMENTAL optimizer/evolution scaffolds (not yet
wired into the cycle — see roadmap)
core/ provider-agnostic LLM abstraction
eval/ golden corpus, queries, runner, baseline
bin/ launcher, installer, config/skill generators
doc/ session format inventory, MCP integration guide
tests/ 225-test suite (unit + integration + gates)

Privacy

Everything is local-first. Indexed content lives in data/lancedb/ (gitignored), the MCP server talks over stdio, and no network call happens unless you explicitly configure an LLM provider. Tests run hermetically in temp directories and provably never touch the live store.

Testing

python -m unittest discover -s tests # 225 tests

Includes unit tests for every module, parser fixtures mirroring real on-disk formats, MCP end-to-end subprocess tests, LLM boundary mocks, a retrieval regression gate, and leakage checks proving the suite leaves the live store byte-identical.

Roadmap

  • Reranker/embedder integration for semantic retrieval (levers already identified by the eval: morphology-aware matching, IDF weighting, semantic tie-breaking)
  • Wire autonomous/ evolution scaffolds to real fitness signals from the eval harness
  • Sleep-time consolidation ("dream cycles") measured by eval lift
  • Cross-agent shared memory namespaces via MCP

License

MIT — © Shivam Tiwari

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - ncdevshiv/openmem · GitHub
Skip to content

Repository files navigation

OpenMem — Agent Memory System for AI Coding Agents

Persistent semantic memory for AI coding agents. OpenMem indexes your agents' real session history into a local LanceDB vector store, retrieves relevant memories on demand, reflects on sessions to extract facts and lessons, and — since v2.x — exposes the whole memory layer as a native MCP server so any MCP-capable client can remember / recall without file-based integration.

Works with Claude Code and Codex CLI session formats today (evidence-based parsers), tolerates absent history gracefully elsewhere, and includes a generic file-based fallback.

Status: actively developed. The storage layer, parsers, learning cycle, MCP server, and evaluation harness are implemented and covered by a 225-test suite. LLM-backed reflection activates automatically when an API key is present; without one everything runs in documented heuristic mode. Retrieval quality is measured, not claimed — see Evaluation.

What It Does (verified)

  • Real session parsing — typed-record JSONL parsers built from on-disk evidence of Claude Code (~/.claude/projects/**/*.jsonl) and Codex CLI (~/.codex/sessions/**/rollout-*.jsonl), with noise filtering (auth-error spam, CLI echoes, sidechain transcripts) and malformed-line tolerance. Formats are documented in doc/session_formats.md.
  • Semantic memory store — LanceDB with fixed-size vector columns, deterministic IDs, float64 importance scores, automatic schema migration, BGE cross-encoder reranking when the ml extra is installed, and an honest keyword-fallback search when it is not.
  • Tiered memory — daily → weekly → long-term consolidation with stable, process-independent content hashing; re-runs are idempotent.
  • Reflection loop — per-session analysis producing facts, improvements, and memories. Mode-tagged llm or heuristic; malformed LLM output falls back visibly instead of silently.
  • Outcome-grounded improvements — improvement queue items can only be completed with linked evidence (memory id / session id / explicit user confirmation). No self-completion theater.
  • MCP serverremember, recall, context, profile, stats, forget tools over stdio; see doc/mcp_integration.md.
  • Evaluation harness — golden retrieval benchmark with recall@k / MRR / nDCG@k / fallout metrics and a regression gate wired into the test suite.
  • 11 agent integrations — skill/context-file installation for Claude Code, Codex CLI, Cursor, VS Code, Windsurf, Qwen Code, OpenCode, Antigravity IDE, Kilo CLI, OpenClaw, plus a generic adapter — all generated from a single template source (bin/generate_skills.py) so they cannot drift apart.

Quick Start

git clone https://github.com/ncdevshiv/openmem.git
cd openmem
# create an environment and install (core deps only)
python -m venv .venv && .venv\Scripts\activate # Windows
pip install -e .# initialize the store and check health
python main.py status
# index your real agent history and run one learning cycle
python main.py run-cycle
# search your own memory
python main.py search "deepseek harness"# measure retrieval quality (writes data/eval/latest.json)
python main.py eval

Optional extras:

pip install -e ".[ml]"# torch + sentence-transformers + transformers (embeddings & reranker)
pip install -e ".[mcp]"# MCP server support (mcp>=2.0.0)
pip install -e ".[llm]"# litellm for LLM-backed reflection

Copy config.example.json to config.json if you want to override defaults; the repo never ships your machine-specific config.

Enabling AI features

LLM reflection activates automatically when a provider is reachable:

ProviderEnvironment variableDefault model
OpenAIOPENAI_API_KEYgpt-4o-mini
AnthropicANTHROPIC_API_KEYclaude-sonnet-4-20250514
GeminiGEMINI_API_KEYgemini-pro
OllamaOLLAMA_BASE_URLllama3

Without keys, zero network calls are made and reflections are tagged "mode": "heuristic". With keys, reflections are tagged "mode": "llm" and cycle reports include a reflection_modes summary. Malformed LLM output is rejected and falls back with a visible warning.

Commands

python main.py status # system health
python main.py run-cycle # full learning cycle (idempotent)
python main.py run-cycle --full # re-index from scratch window
python main.py search <query> # semantic/keyword memory search
python main.py eval [--report P] # retrieval benchmark (markdown + JSON)
python main.py profile # learned user profile
python main.py stats # store statistics
python main.py --agents # list supported adapters
python main.py --skill <agent> # install skill files for an agent

Evaluation

Retrieval quality is measured against a versioned golden set (16 queries over a deterministic 36-memory corpus: exact-term, paraphrase, and negative classes) with a regression gate wired into the test suite. Baseline (keyword-fallback mode, no reranker):

classqueriesrecall@5MRRnDCG@5fallout@5
exact_term60.9721.0001.0000.333
paraphrase61.0000.8890.9170.300
negative40.0000.0000.0000.000
aggregate160.7400.7080.7190.237

Thresholds and rationale live in eval/BASELINE.md; the gate fails CI if retrieval regresses below them. The embedder-free path is powered by a vendored ENR lexical layer (memory_store/enr_lexical.py): Porter-stemmed BM25 with a coord factor, positional phrase matching for quoted spans, cached inverted index, and per-result score_details explanations. Remaining known limits (polysemy across domains, vector-mode negative silence) are logged there as the exact targets for future semantic work.

Architecture

┌────────────────────────────────────────────────────────────┐
│ Your agent Any MCP-capable client │
│ Claude Code etc. (remember/recall/context/profile/...) │
└───────┬───────────────────────────┬────────────────────────┘
│ skill files + │ stdio (mcp>=2.0)
│ context injection │
┌───────▼───────────┐ ┌────────▼─────────┐
│ agents/* │ │ mcp_server.py │
│ real session │ └────────┬─────────┘
│ parsers + skills │ │
└───────┬───────────┘ ▼
▼ memory_store (same core)
┌────────────────────────────────────────────────────────────┐
│ learning_loop/ scheduler · indexer · patterns · reflect │
├────────────────────────────────────────────────────────────┤
│ memory_store/ vector_db (LanceDB) · tiers · user model │
│ skill generator · retrieval metrics │
├────────────────────────────────────────────────────────────┤
│ core/llm.py litellm wrapper — lazy, network-free init │
├────────────────────────────────────────────────────────────┤
│ eval/ golden corpus · queries · runner · gate │
└────────────────────────────────────────────────────────────┘

Directory map:

main.py entry point (delegates to bin/launcher.py)
mcp_server.py MCP stdio server
openmem_cli.py console-script wrapper (pip install -e .)
agents/ adapter contract (base.py) + per-agent parsers/skills
memory_store/ vector DB, tier manager, user model, skill gen, metrics
learning_loop/ scheduler, conversation indexer, pattern recognizer,
reflection engine
autonomous/ EXPERIMENTAL optimizer/evolution scaffolds (not yet
wired into the cycle — see roadmap)
core/ provider-agnostic LLM abstraction
eval/ golden corpus, queries, runner, baseline
bin/ launcher, installer, config/skill generators
doc/ session format inventory, MCP integration guide
tests/ 225-test suite (unit + integration + gates)

Privacy

Everything is local-first. Indexed content lives in data/lancedb/ (gitignored), the MCP server talks over stdio, and no network call happens unless you explicitly configure an LLM provider. Tests run hermetically in temp directories and provably never touch the live store.

Testing

python -m unittest discover -s tests # 225 tests

Includes unit tests for every module, parser fixtures mirroring real on-disk formats, MCP end-to-end subprocess tests, LLM boundary mocks, a retrieval regression gate, and leakage checks proving the suite leaves the live store byte-identical.

Roadmap

  • Reranker/embedder integration for semantic retrieval (levers already identified by the eval: morphology-aware matching, IDF weighting, semantic tie-breaking)
  • Wire autonomous/ evolution scaffolds to real fitness signals from the eval harness
  • Sleep-time consolidation ("dream cycles") measured by eval lift
  • Cross-agent shared memory namespaces via MCP

License

MIT — © Shivam Tiwari

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - ncdevshiv/openmem · GitHub
Skip to content

Repository files navigation

OpenMem — Agent Memory System for AI Coding Agents

Persistent semantic memory for AI coding agents. OpenMem indexes your agents' real session history into a local LanceDB vector store, retrieves relevant memories on demand, reflects on sessions to extract facts and lessons, and — since v2.x — exposes the whole memory layer as a native MCP server so any MCP-capable client can remember / recall without file-based integration.

Works with Claude Code and Codex CLI session formats today (evidence-based parsers), tolerates absent history gracefully elsewhere, and includes a generic file-based fallback.

Status: actively developed. The storage layer, parsers, learning cycle, MCP server, and evaluation harness are implemented and covered by a 225-test suite. LLM-backed reflection activates automatically when an API key is present; without one everything runs in documented heuristic mode. Retrieval quality is measured, not claimed — see Evaluation.

What It Does (verified)

  • Real session parsing — typed-record JSONL parsers built from on-disk evidence of Claude Code (~/.claude/projects/**/*.jsonl) and Codex CLI (~/.codex/sessions/**/rollout-*.jsonl), with noise filtering (auth-error spam, CLI echoes, sidechain transcripts) and malformed-line tolerance. Formats are documented in doc/session_formats.md.
  • Semantic memory store — LanceDB with fixed-size vector columns, deterministic IDs, float64 importance scores, automatic schema migration, BGE cross-encoder reranking when the ml extra is installed, and an honest keyword-fallback search when it is not.
  • Tiered memory — daily → weekly → long-term consolidation with stable, process-independent content hashing; re-runs are idempotent.
  • Reflection loop — per-session analysis producing facts, improvements, and memories. Mode-tagged llm or heuristic; malformed LLM output falls back visibly instead of silently.
  • Outcome-grounded improvements — improvement queue items can only be completed with linked evidence (memory id / session id / explicit user confirmation). No self-completion theater.
  • MCP serverremember, recall, context, profile, stats, forget tools over stdio; see doc/mcp_integration.md.
  • Evaluation harness — golden retrieval benchmark with recall@k / MRR / nDCG@k / fallout metrics and a regression gate wired into the test suite.
  • 11 agent integrations — skill/context-file installation for Claude Code, Codex CLI, Cursor, VS Code, Windsurf, Qwen Code, OpenCode, Antigravity IDE, Kilo CLI, OpenClaw, plus a generic adapter — all generated from a single template source (bin/generate_skills.py) so they cannot drift apart.

Quick Start

git clone https://github.com/ncdevshiv/openmem.git
cd openmem
# create an environment and install (core deps only)
python -m venv .venv && .venv\Scripts\activate # Windows
pip install -e .# initialize the store and check health
python main.py status
# index your real agent history and run one learning cycle
python main.py run-cycle
# search your own memory
python main.py search "deepseek harness"# measure retrieval quality (writes data/eval/latest.json)
python main.py eval

Optional extras:

pip install -e ".[ml]"# torch + sentence-transformers + transformers (embeddings & reranker)
pip install -e ".[mcp]"# MCP server support (mcp>=2.0.0)
pip install -e ".[llm]"# litellm for LLM-backed reflection

Copy config.example.json to config.json if you want to override defaults; the repo never ships your machine-specific config.

Enabling AI features

LLM reflection activates automatically when a provider is reachable:

ProviderEnvironment variableDefault model
OpenAIOPENAI_API_KEYgpt-4o-mini
AnthropicANTHROPIC_API_KEYclaude-sonnet-4-20250514
GeminiGEMINI_API_KEYgemini-pro
OllamaOLLAMA_BASE_URLllama3

Without keys, zero network calls are made and reflections are tagged "mode": "heuristic". With keys, reflections are tagged "mode": "llm" and cycle reports include a reflection_modes summary. Malformed LLM output is rejected and falls back with a visible warning.

Commands

python main.py status # system health
python main.py run-cycle # full learning cycle (idempotent)
python main.py run-cycle --full # re-index from scratch window
python main.py search <query> # semantic/keyword memory search
python main.py eval [--report P] # retrieval benchmark (markdown + JSON)
python main.py profile # learned user profile
python main.py stats # store statistics
python main.py --agents # list supported adapters
python main.py --skill <agent> # install skill files for an agent

Evaluation

Retrieval quality is measured against a versioned golden set (16 queries over a deterministic 36-memory corpus: exact-term, paraphrase, and negative classes) with a regression gate wired into the test suite. Baseline (keyword-fallback mode, no reranker):

classqueriesrecall@5MRRnDCG@5fallout@5
exact_term60.9721.0001.0000.333
paraphrase61.0000.8890.9170.300
negative40.0000.0000.0000.000
aggregate160.7400.7080.7190.237

Thresholds and rationale live in eval/BASELINE.md; the gate fails CI if retrieval regresses below them. The embedder-free path is powered by a vendored ENR lexical layer (memory_store/enr_lexical.py): Porter-stemmed BM25 with a coord factor, positional phrase matching for quoted spans, cached inverted index, and per-result score_details explanations. Remaining known limits (polysemy across domains, vector-mode negative silence) are logged there as the exact targets for future semantic work.

Architecture

┌────────────────────────────────────────────────────────────┐
│ Your agent Any MCP-capable client │
│ Claude Code etc. (remember/recall/context/profile/...) │
└───────┬───────────────────────────┬────────────────────────┘
│ skill files + │ stdio (mcp>=2.0)
│ context injection │
┌───────▼───────────┐ ┌────────▼─────────┐
│ agents/* │ │ mcp_server.py │
│ real session │ └────────┬─────────┘
│ parsers + skills │ │
└───────┬───────────┘ ▼
▼ memory_store (same core)
┌────────────────────────────────────────────────────────────┐
│ learning_loop/ scheduler · indexer · patterns · reflect │
├────────────────────────────────────────────────────────────┤
│ memory_store/ vector_db (LanceDB) · tiers · user model │
│ skill generator · retrieval metrics │
├────────────────────────────────────────────────────────────┤
│ core/llm.py litellm wrapper — lazy, network-free init │
├────────────────────────────────────────────────────────────┤
│ eval/ golden corpus · queries · runner · gate │
└────────────────────────────────────────────────────────────┘

Directory map:

main.py entry point (delegates to bin/launcher.py)
mcp_server.py MCP stdio server
openmem_cli.py console-script wrapper (pip install -e .)
agents/ adapter contract (base.py) + per-agent parsers/skills
memory_store/ vector DB, tier manager, user model, skill gen, metrics
learning_loop/ scheduler, conversation indexer, pattern recognizer,
reflection engine
autonomous/ EXPERIMENTAL optimizer/evolution scaffolds (not yet
wired into the cycle — see roadmap)
core/ provider-agnostic LLM abstraction
eval/ golden corpus, queries, runner, baseline
bin/ launcher, installer, config/skill generators
doc/ session format inventory, MCP integration guide
tests/ 225-test suite (unit + integration + gates)

Privacy

Everything is local-first. Indexed content lives in data/lancedb/ (gitignored), the MCP server talks over stdio, and no network call happens unless you explicitly configure an LLM provider. Tests run hermetically in temp directories and provably never touch the live store.

Testing

python -m unittest discover -s tests # 225 tests

Includes unit tests for every module, parser fixtures mirroring real on-disk formats, MCP end-to-end subprocess tests, LLM boundary mocks, a retrieval regression gate, and leakage checks proving the suite leaves the live store byte-identical.

Roadmap

  • Reranker/embedder integration for semantic retrieval (levers already identified by the eval: morphology-aware matching, IDF weighting, semantic tie-breaking)
  • Wire autonomous/ evolution scaffolds to real fitness signals from the eval harness
  • Sleep-time consolidation ("dream cycles") measured by eval lift
  • Cross-agent shared memory namespaces via MCP

License

MIT — © Shivam Tiwari

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' GitHub - ncdevshiv/openmem · GitHub
Skip to content

Repository files navigation

OpenMem — Agent Memory System for AI Coding Agents

Persistent semantic memory for AI coding agents. OpenMem indexes your agents' real session history into a local LanceDB vector store, retrieves relevant memories on demand, reflects on sessions to extract facts and lessons, and — since v2.x — exposes the whole memory layer as a native MCP server so any MCP-capable client can remember / recall without file-based integration.

Works with Claude Code and Codex CLI session formats today (evidence-based parsers), tolerates absent history gracefully elsewhere, and includes a generic file-based fallback.

Status: actively developed. The storage layer, parsers, learning cycle, MCP server, and evaluation harness are implemented and covered by a 225-test suite. LLM-backed reflection activates automatically when an API key is present; without one everything runs in documented heuristic mode. Retrieval quality is measured, not claimed — see Evaluation.

What It Does (verified)

  • Real session parsing — typed-record JSONL parsers built from on-disk evidence of Claude Code (~/.claude/projects/**/*.jsonl) and Codex CLI (~/.codex/sessions/**/rollout-*.jsonl), with noise filtering (auth-error spam, CLI echoes, sidechain transcripts) and malformed-line tolerance. Formats are documented in doc/session_formats.md.
  • Semantic memory store — LanceDB with fixed-size vector columns, deterministic IDs, float64 importance scores, automatic schema migration, BGE cross-encoder reranking when the ml extra is installed, and an honest keyword-fallback search when it is not.
  • Tiered memory — daily → weekly → long-term consolidation with stable, process-independent content hashing; re-runs are idempotent.
  • Reflection loop — per-session analysis producing facts, improvements, and memories. Mode-tagged llm or heuristic; malformed LLM output falls back visibly instead of silently.
  • Outcome-grounded improvements — improvement queue items can only be completed with linked evidence (memory id / session id / explicit user confirmation). No self-completion theater.
  • MCP serverremember, recall, context, profile, stats, forget tools over stdio; see doc/mcp_integration.md.
  • Evaluation harness — golden retrieval benchmark with recall@k / MRR / nDCG@k / fallout metrics and a regression gate wired into the test suite.
  • 11 agent integrations — skill/context-file installation for Claude Code, Codex CLI, Cursor, VS Code, Windsurf, Qwen Code, OpenCode, Antigravity IDE, Kilo CLI, OpenClaw, plus a generic adapter — all generated from a single template source (bin/generate_skills.py) so they cannot drift apart.

Quick Start

git clone https://github.com/ncdevshiv/openmem.git
cd openmem
# create an environment and install (core deps only)
python -m venv .venv && .venv\Scripts\activate # Windows
pip install -e .# initialize the store and check health
python main.py status
# index your real agent history and run one learning cycle
python main.py run-cycle
# search your own memory
python main.py search "deepseek harness"# measure retrieval quality (writes data/eval/latest.json)
python main.py eval

Optional extras:

pip install -e ".[ml]"# torch + sentence-transformers + transformers (embeddings & reranker)
pip install -e ".[mcp]"# MCP server support (mcp>=2.0.0)
pip install -e ".[llm]"# litellm for LLM-backed reflection

Copy config.example.json to config.json if you want to override defaults; the repo never ships your machine-specific config.

Enabling AI features

LLM reflection activates automatically when a provider is reachable:

ProviderEnvironment variableDefault model
OpenAIOPENAI_API_KEYgpt-4o-mini
AnthropicANTHROPIC_API_KEYclaude-sonnet-4-20250514
GeminiGEMINI_API_KEYgemini-pro
OllamaOLLAMA_BASE_URLllama3

Without keys, zero network calls are made and reflections are tagged "mode": "heuristic". With keys, reflections are tagged "mode": "llm" and cycle reports include a reflection_modes summary. Malformed LLM output is rejected and falls back with a visible warning.

Commands

python main.py status # system health
python main.py run-cycle # full learning cycle (idempotent)
python main.py run-cycle --full # re-index from scratch window
python main.py search <query> # semantic/keyword memory search
python main.py eval [--report P] # retrieval benchmark (markdown + JSON)
python main.py profile # learned user profile
python main.py stats # store statistics
python main.py --agents # list supported adapters
python main.py --skill <agent> # install skill files for an agent

Evaluation

Retrieval quality is measured against a versioned golden set (16 queries over a deterministic 36-memory corpus: exact-term, paraphrase, and negative classes) with a regression gate wired into the test suite. Baseline (keyword-fallback mode, no reranker):

classqueriesrecall@5MRRnDCG@5fallout@5
exact_term60.9721.0001.0000.333
paraphrase61.0000.8890.9170.300
negative40.0000.0000.0000.000
aggregate160.7400.7080.7190.237

Thresholds and rationale live in eval/BASELINE.md; the gate fails CI if retrieval regresses below them. The embedder-free path is powered by a vendored ENR lexical layer (memory_store/enr_lexical.py): Porter-stemmed BM25 with a coord factor, positional phrase matching for quoted spans, cached inverted index, and per-result score_details explanations. Remaining known limits (polysemy across domains, vector-mode negative silence) are logged there as the exact targets for future semantic work.

Architecture

┌────────────────────────────────────────────────────────────┐
│ Your agent Any MCP-capable client │
│ Claude Code etc. (remember/recall/context/profile/...) │
└───────┬───────────────────────────┬────────────────────────┘
│ skill files + │ stdio (mcp>=2.0)
│ context injection │
┌───────▼───────────┐ ┌────────▼─────────┐
│ agents/* │ │ mcp_server.py │
│ real session │ └────────┬─────────┘
│ parsers + skills │ │
└───────┬───────────┘ ▼
▼ memory_store (same core)
┌────────────────────────────────────────────────────────────┐
│ learning_loop/ scheduler · indexer · patterns · reflect │
├────────────────────────────────────────────────────────────┤
│ memory_store/ vector_db (LanceDB) · tiers · user model │
│ skill generator · retrieval metrics │
├────────────────────────────────────────────────────────────┤
│ core/llm.py litellm wrapper — lazy, network-free init │
├────────────────────────────────────────────────────────────┤
│ eval/ golden corpus · queries · runner · gate │
└────────────────────────────────────────────────────────────┘

Directory map:

main.py entry point (delegates to bin/launcher.py)
mcp_server.py MCP stdio server
openmem_cli.py console-script wrapper (pip install -e .)
agents/ adapter contract (base.py) + per-agent parsers/skills
memory_store/ vector DB, tier manager, user model, skill gen, metrics
learning_loop/ scheduler, conversation indexer, pattern recognizer,
reflection engine
autonomous/ EXPERIMENTAL optimizer/evolution scaffolds (not yet
wired into the cycle — see roadmap)
core/ provider-agnostic LLM abstraction
eval/ golden corpus, queries, runner, baseline
bin/ launcher, installer, config/skill generators
doc/ session format inventory, MCP integration guide
tests/ 225-test suite (unit + integration + gates)

Privacy

Everything is local-first. Indexed content lives in data/lancedb/ (gitignored), the MCP server talks over stdio, and no network call happens unless you explicitly configure an LLM provider. Tests run hermetically in temp directories and provably never touch the live store.

Testing

python -m unittest discover -s tests # 225 tests

Includes unit tests for every module, parser fixtures mirroring real on-disk formats, MCP end-to-end subprocess tests, LLM boundary mocks, a retrieval regression gate, and leakage checks proving the suite leaves the live store byte-identical.

Roadmap

  • Reranker/embedder integration for semantic retrieval (levers already identified by the eval: morphology-aware matching, IDF weighting, semantic tie-breaking)
  • Wire autonomous/ evolution scaffolds to real fitness signals from the eval harness
  • Sleep-time consolidation ("dream cycles") measured by eval lift
  • Cross-agent shared memory namespaces via MCP

License

MIT — © Shivam Tiwari

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - ncdevshiv/openmem · GitHub
Skip to content

Repository files navigation

OpenMem — Agent Memory System for AI Coding Agents

Persistent semantic memory for AI coding agents. OpenMem indexes your agents' real session history into a local LanceDB vector store, retrieves relevant memories on demand, reflects on sessions to extract facts and lessons, and — since v2.x — exposes the whole memory layer as a native MCP server so any MCP-capable client can remember / recall without file-based integration.

Works with Claude Code and Codex CLI session formats today (evidence-based parsers), tolerates absent history gracefully elsewhere, and includes a generic file-based fallback.

Status: actively developed. The storage layer, parsers, learning cycle, MCP server, and evaluation harness are implemented and covered by a 225-test suite. LLM-backed reflection activates automatically when an API key is present; without one everything runs in documented heuristic mode. Retrieval quality is measured, not claimed — see Evaluation.

What It Does (verified)

  • Real session parsing — typed-record JSONL parsers built from on-disk evidence of Claude Code (~/.claude/projects/**/*.jsonl) and Codex CLI (~/.codex/sessions/**/rollout-*.jsonl), with noise filtering (auth-error spam, CLI echoes, sidechain transcripts) and malformed-line tolerance. Formats are documented in doc/session_formats.md.
  • Semantic memory store — LanceDB with fixed-size vector columns, deterministic IDs, float64 importance scores, automatic schema migration, BGE cross-encoder reranking when the ml extra is installed, and an honest keyword-fallback search when it is not.
  • Tiered memory — daily → weekly → long-term consolidation with stable, process-independent content hashing; re-runs are idempotent.
  • Reflection loop — per-session analysis producing facts, improvements, and memories. Mode-tagged llm or heuristic; malformed LLM output falls back visibly instead of silently.
  • Outcome-grounded improvements — improvement queue items can only be completed with linked evidence (memory id / session id / explicit user confirmation). No self-completion theater.
  • MCP serverremember, recall, context, profile, stats, forget tools over stdio; see doc/mcp_integration.md.
  • Evaluation harness — golden retrieval benchmark with recall@k / MRR / nDCG@k / fallout metrics and a regression gate wired into the test suite.
  • 11 agent integrations — skill/context-file installation for Claude Code, Codex CLI, Cursor, VS Code, Windsurf, Qwen Code, OpenCode, Antigravity IDE, Kilo CLI, OpenClaw, plus a generic adapter — all generated from a single template source (bin/generate_skills.py) so they cannot drift apart.

Quick Start

git clone https://github.com/ncdevshiv/openmem.git
cd openmem
# create an environment and install (core deps only)
python -m venv .venv && .venv\Scripts\activate # Windows
pip install -e .# initialize the store and check health
python main.py status
# index your real agent history and run one learning cycle
python main.py run-cycle
# search your own memory
python main.py search "deepseek harness"# measure retrieval quality (writes data/eval/latest.json)
python main.py eval

Optional extras:

pip install -e ".[ml]"# torch + sentence-transformers + transformers (embeddings & reranker)
pip install -e ".[mcp]"# MCP server support (mcp>=2.0.0)
pip install -e ".[llm]"# litellm for LLM-backed reflection

Copy config.example.json to config.json if you want to override defaults; the repo never ships your machine-specific config.

Enabling AI features

LLM reflection activates automatically when a provider is reachable:

ProviderEnvironment variableDefault model
OpenAIOPENAI_API_KEYgpt-4o-mini
AnthropicANTHROPIC_API_KEYclaude-sonnet-4-20250514
GeminiGEMINI_API_KEYgemini-pro
OllamaOLLAMA_BASE_URLllama3

Without keys, zero network calls are made and reflections are tagged "mode": "heuristic". With keys, reflections are tagged "mode": "llm" and cycle reports include a reflection_modes summary. Malformed LLM output is rejected and falls back with a visible warning.

Commands

python main.py status # system health
python main.py run-cycle # full learning cycle (idempotent)
python main.py run-cycle --full # re-index from scratch window
python main.py search <query> # semantic/keyword memory search
python main.py eval [--report P] # retrieval benchmark (markdown + JSON)
python main.py profile # learned user profile
python main.py stats # store statistics
python main.py --agents # list supported adapters
python main.py --skill <agent> # install skill files for an agent

Evaluation

Retrieval quality is measured against a versioned golden set (16 queries over a deterministic 36-memory corpus: exact-term, paraphrase, and negative classes) with a regression gate wired into the test suite. Baseline (keyword-fallback mode, no reranker):

classqueriesrecall@5MRRnDCG@5fallout@5
exact_term60.9721.0001.0000.333
paraphrase61.0000.8890.9170.300
negative40.0000.0000.0000.000
aggregate160.7400.7080.7190.237

Thresholds and rationale live in eval/BASELINE.md; the gate fails CI if retrieval regresses below them. The embedder-free path is powered by a vendored ENR lexical layer (memory_store/enr_lexical.py): Porter-stemmed BM25 with a coord factor, positional phrase matching for quoted spans, cached inverted index, and per-result score_details explanations. Remaining known limits (polysemy across domains, vector-mode negative silence) are logged there as the exact targets for future semantic work.

Architecture

┌────────────────────────────────────────────────────────────┐
│ Your agent Any MCP-capable client │
│ Claude Code etc. (remember/recall/context/profile/...) │
└───────┬───────────────────────────┬────────────────────────┘
│ skill files + │ stdio (mcp>=2.0)
│ context injection │
┌───────▼───────────┐ ┌────────▼─────────┐
│ agents/* │ │ mcp_server.py │
│ real session │ └────────┬─────────┘
│ parsers + skills │ │
└───────┬───────────┘ ▼
▼ memory_store (same core)
┌────────────────────────────────────────────────────────────┐
│ learning_loop/ scheduler · indexer · patterns · reflect │
├────────────────────────────────────────────────────────────┤
│ memory_store/ vector_db (LanceDB) · tiers · user model │
│ skill generator · retrieval metrics │
├────────────────────────────────────────────────────────────┤
│ core/llm.py litellm wrapper — lazy, network-free init │
├────────────────────────────────────────────────────────────┤
│ eval/ golden corpus · queries · runner · gate │
└────────────────────────────────────────────────────────────┘

Directory map:

main.py entry point (delegates to bin/launcher.py)
mcp_server.py MCP stdio server
openmem_cli.py console-script wrapper (pip install -e .)
agents/ adapter contract (base.py) + per-agent parsers/skills
memory_store/ vector DB, tier manager, user model, skill gen, metrics
learning_loop/ scheduler, conversation indexer, pattern recognizer,
reflection engine
autonomous/ EXPERIMENTAL optimizer/evolution scaffolds (not yet
wired into the cycle — see roadmap)
core/ provider-agnostic LLM abstraction
eval/ golden corpus, queries, runner, baseline
bin/ launcher, installer, config/skill generators
doc/ session format inventory, MCP integration guide
tests/ 225-test suite (unit + integration + gates)

Privacy

Everything is local-first. Indexed content lives in data/lancedb/ (gitignored), the MCP server talks over stdio, and no network call happens unless you explicitly configure an LLM provider. Tests run hermetically in temp directories and provably never touch the live store.

Testing

python -m unittest discover -s tests # 225 tests

Includes unit tests for every module, parser fixtures mirroring real on-disk formats, MCP end-to-end subprocess tests, LLM boundary mocks, a retrieval regression gate, and leakage checks proving the suite leaves the live store byte-identical.

Roadmap

  • Reranker/embedder integration for semantic retrieval (levers already identified by the eval: morphology-aware matching, IDF weighting, semantic tie-breaking)
  • Wire autonomous/ evolution scaffolds to real fitness signals from the eval harness
  • Sleep-time consolidation ("dream cycles") measured by eval lift
  • Cross-agent shared memory namespaces via MCP

License

MIT — © Shivam Tiwari

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' GitHub - ncdevshiv/openmem · GitHub
Skip to content

Repository files navigation

OpenMem — Agent Memory System for AI Coding Agents

Persistent semantic memory for AI coding agents. OpenMem indexes your agents' real session history into a local LanceDB vector store, retrieves relevant memories on demand, reflects on sessions to extract facts and lessons, and — since v2.x — exposes the whole memory layer as a native MCP server so any MCP-capable client can remember / recall without file-based integration.

Works with Claude Code and Codex CLI session formats today (evidence-based parsers), tolerates absent history gracefully elsewhere, and includes a generic file-based fallback.

Status: actively developed. The storage layer, parsers, learning cycle, MCP server, and evaluation harness are implemented and covered by a 225-test suite. LLM-backed reflection activates automatically when an API key is present; without one everything runs in documented heuristic mode. Retrieval quality is measured, not claimed — see Evaluation.

What It Does (verified)

  • Real session parsing — typed-record JSONL parsers built from on-disk evidence of Claude Code (~/.claude/projects/**/*.jsonl) and Codex CLI (~/.codex/sessions/**/rollout-*.jsonl), with noise filtering (auth-error spam, CLI echoes, sidechain transcripts) and malformed-line tolerance. Formats are documented in doc/session_formats.md.
  • Semantic memory store — LanceDB with fixed-size vector columns, deterministic IDs, float64 importance scores, automatic schema migration, BGE cross-encoder reranking when the ml extra is installed, and an honest keyword-fallback search when it is not.
  • Tiered memory — daily → weekly → long-term consolidation with stable, process-independent content hashing; re-runs are idempotent.
  • Reflection loop — per-session analysis producing facts, improvements, and memories. Mode-tagged llm or heuristic; malformed LLM output falls back visibly instead of silently.
  • Outcome-grounded improvements — improvement queue items can only be completed with linked evidence (memory id / session id / explicit user confirmation). No self-completion theater.
  • MCP serverremember, recall, context, profile, stats, forget tools over stdio; see doc/mcp_integration.md.
  • Evaluation harness — golden retrieval benchmark with recall@k / MRR / nDCG@k / fallout metrics and a regression gate wired into the test suite.
  • 11 agent integrations — skill/context-file installation for Claude Code, Codex CLI, Cursor, VS Code, Windsurf, Qwen Code, OpenCode, Antigravity IDE, Kilo CLI, OpenClaw, plus a generic adapter — all generated from a single template source (bin/generate_skills.py) so they cannot drift apart.

Quick Start

git clone https://github.com/ncdevshiv/openmem.git
cd openmem
# create an environment and install (core deps only)
python -m venv .venv && .venv\Scripts\activate # Windows
pip install -e .# initialize the store and check health
python main.py status
# index your real agent history and run one learning cycle
python main.py run-cycle
# search your own memory
python main.py search "deepseek harness"# measure retrieval quality (writes data/eval/latest.json)
python main.py eval

Optional extras:

pip install -e ".[ml]"# torch + sentence-transformers + transformers (embeddings & reranker)
pip install -e ".[mcp]"# MCP server support (mcp>=2.0.0)
pip install -e ".[llm]"# litellm for LLM-backed reflection

Copy config.example.json to config.json if you want to override defaults; the repo never ships your machine-specific config.

Enabling AI features

LLM reflection activates automatically when a provider is reachable:

ProviderEnvironment variableDefault model
OpenAIOPENAI_API_KEYgpt-4o-mini
AnthropicANTHROPIC_API_KEYclaude-sonnet-4-20250514
GeminiGEMINI_API_KEYgemini-pro
OllamaOLLAMA_BASE_URLllama3

Without keys, zero network calls are made and reflections are tagged "mode": "heuristic". With keys, reflections are tagged "mode": "llm" and cycle reports include a reflection_modes summary. Malformed LLM output is rejected and falls back with a visible warning.

Commands

python main.py status # system health
python main.py run-cycle # full learning cycle (idempotent)
python main.py run-cycle --full # re-index from scratch window
python main.py search <query> # semantic/keyword memory search
python main.py eval [--report P] # retrieval benchmark (markdown + JSON)
python main.py profile # learned user profile
python main.py stats # store statistics
python main.py --agents # list supported adapters
python main.py --skill <agent> # install skill files for an agent

Evaluation

Retrieval quality is measured against a versioned golden set (16 queries over a deterministic 36-memory corpus: exact-term, paraphrase, and negative classes) with a regression gate wired into the test suite. Baseline (keyword-fallback mode, no reranker):

classqueriesrecall@5MRRnDCG@5fallout@5
exact_term60.9721.0001.0000.333
paraphrase61.0000.8890.9170.300
negative40.0000.0000.0000.000
aggregate160.7400.7080.7190.237

Thresholds and rationale live in eval/BASELINE.md; the gate fails CI if retrieval regresses below them. The embedder-free path is powered by a vendored ENR lexical layer (memory_store/enr_lexical.py): Porter-stemmed BM25 with a coord factor, positional phrase matching for quoted spans, cached inverted index, and per-result score_details explanations. Remaining known limits (polysemy across domains, vector-mode negative silence) are logged there as the exact targets for future semantic work.

Architecture

┌────────────────────────────────────────────────────────────┐
│ Your agent Any MCP-capable client │
│ Claude Code etc. (remember/recall/context/profile/...) │
└───────┬───────────────────────────┬────────────────────────┘
│ skill files + │ stdio (mcp>=2.0)
│ context injection │
┌───────▼───────────┐ ┌────────▼─────────┐
│ agents/* │ │ mcp_server.py │
│ real session │ └────────┬─────────┘
│ parsers + skills │ │
└───────┬───────────┘ ▼
▼ memory_store (same core)
┌────────────────────────────────────────────────────────────┐
│ learning_loop/ scheduler · indexer · patterns · reflect │
├────────────────────────────────────────────────────────────┤
│ memory_store/ vector_db (LanceDB) · tiers · user model │
│ skill generator · retrieval metrics │
├────────────────────────────────────────────────────────────┤
│ core/llm.py litellm wrapper — lazy, network-free init │
├────────────────────────────────────────────────────────────┤
│ eval/ golden corpus · queries · runner · gate │
└────────────────────────────────────────────────────────────┘

Directory map:

main.py entry point (delegates to bin/launcher.py)
mcp_server.py MCP stdio server
openmem_cli.py console-script wrapper (pip install -e .)
agents/ adapter contract (base.py) + per-agent parsers/skills
memory_store/ vector DB, tier manager, user model, skill gen, metrics
learning_loop/ scheduler, conversation indexer, pattern recognizer,
reflection engine
autonomous/ EXPERIMENTAL optimizer/evolution scaffolds (not yet
wired into the cycle — see roadmap)
core/ provider-agnostic LLM abstraction
eval/ golden corpus, queries, runner, baseline
bin/ launcher, installer, config/skill generators
doc/ session format inventory, MCP integration guide
tests/ 225-test suite (unit + integration + gates)

Privacy

Everything is local-first. Indexed content lives in data/lancedb/ (gitignored), the MCP server talks over stdio, and no network call happens unless you explicitly configure an LLM provider. Tests run hermetically in temp directories and provably never touch the live store.

Testing

python -m unittest discover -s tests # 225 tests

Includes unit tests for every module, parser fixtures mirroring real on-disk formats, MCP end-to-end subprocess tests, LLM boundary mocks, a retrieval regression gate, and leakage checks proving the suite leaves the live store byte-identical.

Roadmap

  • Reranker/embedder integration for semantic retrieval (levers already identified by the eval: morphology-aware matching, IDF weighting, semantic tie-breaking)
  • Wire autonomous/ evolution scaffolds to real fitness signals from the eval harness
  • Sleep-time consolidation ("dream cycles") measured by eval lift
  • Cross-agent shared memory namespaces via MCP

License

MIT — © Shivam Tiwari

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); GitHub - ncdevshiv/openmem · GitHub
Skip to content

Repository files navigation

OpenMem — Agent Memory System for AI Coding Agents

Persistent semantic memory for AI coding agents. OpenMem indexes your agents' real session history into a local LanceDB vector store, retrieves relevant memories on demand, reflects on sessions to extract facts and lessons, and — since v2.x — exposes the whole memory layer as a native MCP server so any MCP-capable client can remember / recall without file-based integration.

Works with Claude Code and Codex CLI session formats today (evidence-based parsers), tolerates absent history gracefully elsewhere, and includes a generic file-based fallback.

Status: actively developed. The storage layer, parsers, learning cycle, MCP server, and evaluation harness are implemented and covered by a 225-test suite. LLM-backed reflection activates automatically when an API key is present; without one everything runs in documented heuristic mode. Retrieval quality is measured, not claimed — see Evaluation.

What It Does (verified)

  • Real session parsing — typed-record JSONL parsers built from on-disk evidence of Claude Code (~/.claude/projects/**/*.jsonl) and Codex CLI (~/.codex/sessions/**/rollout-*.jsonl), with noise filtering (auth-error spam, CLI echoes, sidechain transcripts) and malformed-line tolerance. Formats are documented in doc/session_formats.md.
  • Semantic memory store — LanceDB with fixed-size vector columns, deterministic IDs, float64 importance scores, automatic schema migration, BGE cross-encoder reranking when the ml extra is installed, and an honest keyword-fallback search when it is not.
  • Tiered memory — daily → weekly → long-term consolidation with stable, process-independent content hashing; re-runs are idempotent.
  • Reflection loop — per-session analysis producing facts, improvements, and memories. Mode-tagged llm or heuristic; malformed LLM output falls back visibly instead of silently.
  • Outcome-grounded improvements — improvement queue items can only be completed with linked evidence (memory id / session id / explicit user confirmation). No self-completion theater.
  • MCP serverremember, recall, context, profile, stats, forget tools over stdio; see doc/mcp_integration.md.
  • Evaluation harness — golden retrieval benchmark with recall@k / MRR / nDCG@k / fallout metrics and a regression gate wired into the test suite.
  • 11 agent integrations — skill/context-file installation for Claude Code, Codex CLI, Cursor, VS Code, Windsurf, Qwen Code, OpenCode, Antigravity IDE, Kilo CLI, OpenClaw, plus a generic adapter — all generated from a single template source (bin/generate_skills.py) so they cannot drift apart.

Quick Start

git clone https://github.com/ncdevshiv/openmem.git
cd openmem
# create an environment and install (core deps only)
python -m venv .venv && .venv\Scripts\activate # Windows
pip install -e .# initialize the store and check health
python main.py status
# index your real agent history and run one learning cycle
python main.py run-cycle
# search your own memory
python main.py search "deepseek harness"# measure retrieval quality (writes data/eval/latest.json)
python main.py eval

Optional extras:

pip install -e ".[ml]"# torch + sentence-transformers + transformers (embeddings & reranker)
pip install -e ".[mcp]"# MCP server support (mcp>=2.0.0)
pip install -e ".[llm]"# litellm for LLM-backed reflection

Copy config.example.json to config.json if you want to override defaults; the repo never ships your machine-specific config.

Enabling AI features

LLM reflection activates automatically when a provider is reachable:

ProviderEnvironment variableDefault model
OpenAIOPENAI_API_KEYgpt-4o-mini
AnthropicANTHROPIC_API_KEYclaude-sonnet-4-20250514
GeminiGEMINI_API_KEYgemini-pro
OllamaOLLAMA_BASE_URLllama3

Without keys, zero network calls are made and reflections are tagged "mode": "heuristic". With keys, reflections are tagged "mode": "llm" and cycle reports include a reflection_modes summary. Malformed LLM output is rejected and falls back with a visible warning.

Commands

python main.py status # system health
python main.py run-cycle # full learning cycle (idempotent)
python main.py run-cycle --full # re-index from scratch window
python main.py search <query> # semantic/keyword memory search
python main.py eval [--report P] # retrieval benchmark (markdown + JSON)
python main.py profile # learned user profile
python main.py stats # store statistics
python main.py --agents # list supported adapters
python main.py --skill <agent> # install skill files for an agent

Evaluation

Retrieval quality is measured against a versioned golden set (16 queries over a deterministic 36-memory corpus: exact-term, paraphrase, and negative classes) with a regression gate wired into the test suite. Baseline (keyword-fallback mode, no reranker):

classqueriesrecall@5MRRnDCG@5fallout@5
exact_term60.9721.0001.0000.333
paraphrase61.0000.8890.9170.300
negative40.0000.0000.0000.000
aggregate160.7400.7080.7190.237

Thresholds and rationale live in eval/BASELINE.md; the gate fails CI if retrieval regresses below them. The embedder-free path is powered by a vendored ENR lexical layer (memory_store/enr_lexical.py): Porter-stemmed BM25 with a coord factor, positional phrase matching for quoted spans, cached inverted index, and per-result score_details explanations. Remaining known limits (polysemy across domains, vector-mode negative silence) are logged there as the exact targets for future semantic work.

Architecture

┌────────────────────────────────────────────────────────────┐
│ Your agent Any MCP-capable client │
│ Claude Code etc. (remember/recall/context/profile/...) │
└───────┬───────────────────────────┬────────────────────────┘
│ skill files + │ stdio (mcp>=2.0)
│ context injection │
┌───────▼───────────┐ ┌────────▼─────────┐
│ agents/* │ │ mcp_server.py │
│ real session │ └────────┬─────────┘
│ parsers + skills │ │
└───────┬───────────┘ ▼
▼ memory_store (same core)
┌────────────────────────────────────────────────────────────┐
│ learning_loop/ scheduler · indexer · patterns · reflect │
├────────────────────────────────────────────────────────────┤
│ memory_store/ vector_db (LanceDB) · tiers · user model │
│ skill generator · retrieval metrics │
├────────────────────────────────────────────────────────────┤
│ core/llm.py litellm wrapper — lazy, network-free init │
├────────────────────────────────────────────────────────────┤
│ eval/ golden corpus · queries · runner · gate │
└────────────────────────────────────────────────────────────┘

Directory map:

main.py entry point (delegates to bin/launcher.py)
mcp_server.py MCP stdio server
openmem_cli.py console-script wrapper (pip install -e .)
agents/ adapter contract (base.py) + per-agent parsers/skills
memory_store/ vector DB, tier manager, user model, skill gen, metrics
learning_loop/ scheduler, conversation indexer, pattern recognizer,
reflection engine
autonomous/ EXPERIMENTAL optimizer/evolution scaffolds (not yet
wired into the cycle — see roadmap)
core/ provider-agnostic LLM abstraction
eval/ golden corpus, queries, runner, baseline
bin/ launcher, installer, config/skill generators
doc/ session format inventory, MCP integration guide
tests/ 225-test suite (unit + integration + gates)

Privacy

Everything is local-first. Indexed content lives in data/lancedb/ (gitignored), the MCP server talks over stdio, and no network call happens unless you explicitly configure an LLM provider. Tests run hermetically in temp directories and provably never touch the live store.

Testing

python -m unittest discover -s tests # 225 tests

Includes unit tests for every module, parser fixtures mirroring real on-disk formats, MCP end-to-end subprocess tests, LLM boundary mocks, a retrieval regression gate, and leakage checks proving the suite leaves the live store byte-identical.

Roadmap

  • Reranker/embedder integration for semantic retrieval (levers already identified by the eval: morphology-aware matching, IDF weighting, semantic tie-breaking)
  • Wire autonomous/ evolution scaffolds to real fitness signals from the eval harness
  • Sleep-time consolidation ("dream cycles") measured by eval lift
  • Cross-agent shared memory namespaces via MCP

License

MIT — © Shivam Tiwari

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages