Repository files navigation

AI Companion with Long-Term Memory

A conversational AI companion that remembers past conversations using a multi-tier memory architecture. Two retrieval strategies: rolling summary for chat (fast, lossy), hybrid search on permanent facts for precise queries (lossless). 2 LLM calls per ingestion.

Architecture

Conversation messages <-> Ingestion / Retrieval libraries <-> Gemini 2.5 Flash
|
PostgreSQL + Qdrant
|
Memory Pipeline
(Ingestion + Profile + Summary + Facts)

Memory System

Short-term memory: Last 20 messages from the current thread (regardless of ingestion status) — immediate conversational context.

Long-term memory: 5 components:

ComponentDescriptionStorage
User ProfileExplicit facts + implicit traits, updated via add/update/delete opsPostgreSQL
ForesightTime-bounded events with valid_from / valid_until / duration_days, auto-expiredPostgreSQL
Rolling SummaryTwo-tier: Archive (compressed old) + Recent (date-tagged entries)PostgreSQL
Facts TablePermanent, never compressed — every extracted fact stored foreverPostgreSQL (tsvector + GIN)
Facts VectorsEmbeddings for semantic searchQdrant (3072-dim, cosine)

Two Retrieval Strategies

Chat endpoint (/chat) — conversational, general awareness:

 Profile + Foresight + Rolling Summary + Last 20 messages

Query endpoint (/query) — precise recall:

Parallel:
[Thread 1] Gemini embed query
[Thread 2] profile + foresight + messages + keyword search
Then:
Qdrant vector search (~200ms) → RRF fusion with keyword results → top 5 facts

Ingestion Pipeline

2 LLM calls per ingestion (extract + profile update), +1 when compression triggers.

Conversation (20 messages)
|
[1] FETCH CONTEXT (1 DB read)
| Rolling summary for coreference resolution
|
[2] EXTRACT (1 LLM call — Gemini)
| Input: conversation + rolling summary
| Output: consolidated facts (category-tagged, dated) + foresight signals
| Dates sanitized: invalid LLM dates replaced with current_date
|
[3] In parallel (3 threads):
| [Thread 1] PG store — 3 parallel connections:
| | facts INSERT + foresight expire/INSERT + summary read/append
| [Thread 2] Embed facts (Gemini embedding API)
| [Thread 3] Profile update (1 LLM call):
| | LLM receives numbered profile + new facts
| | Returns add/update/delete operations as JSON
| | Operations applied programmatically (not LLM rewrite)
| | Conflicts logged to conflict_log table
|
[4] Qdrant upsert (needs fact_ids from PG + embeddings from Gemini)
|
[5] Compression checks:
If rolling summary >= 90% of 10k token budget:
| → Evict oldest 25% of recent → compress into archive (1 LLM call)
| → Archive capped at 20% of budget
If profile >= 80% of 3k token budget:
→ Merge similar facts (1 LLM call)

Extraction Details

  • 1 LLM call per batch — no segmentation step
  • Produces consolidated facts — dense, self-contained sentences (not atomic fragments)
  • Attribution rules — every fact must name the person, pronouns resolved
  • Frequency tracking — recurring activities captured with explicit frequency
  • Verbatim details prioritized: signs, paintings, book titles, pet behaviors
  • Quality filters — greetings, pleasantries, acknowledgments excluded
  • Prior rolling summary used as context for coreference resolution (with guard: "do NOT extract from this")

Profile Update (add/update/delete operations)

  • LLM receives numbered profile lines: [1] - Rampal is 24 years old
  • Returns structured JSON operations referencing line numbers
  • apply_operations() in Python handles the text editing
  • More predictable than full LLM rewrite — LLM decides what, Python does editing
  • Conflict detection only on profile (facts table is append-only)

Foresight

  • Time-bounded events with valid_from, valid_until, duration_days, evidence
  • Auto-expired during ingestion when valid_until < current_date
  • Examples: travel plans, illness recovery, deadlines, new jobs

Compression

  • Trigger: 90% of 10k token budget
  • Eviction: oldest 25% of recent entries
  • Recursive: new_archive = LLM(existing_archive + evicted_entries)
  • Archive capped at 20% of budget — re-compressed if exceeded
  • Token count: char-based estimate (len/4), no API call

Ingestion Triggers

  • Production: After 20 unprocessed messages accumulate
  • Periodic: Every 10 minutes, checks for threads with 4+ old unprocessed messages
  • Manual: POST /threads/{thread_id}/ingest
  • Benchmark: After each session

Data Storage

PostgreSQL

TablePurpose
user_profileExplicit facts + implicit traits (bullet-point format)
foresightTime-bounded events with validity windows + duration_days
conversation_summariesTwo-tier: archive_text + recent_text + token_count
factsPermanent fact store with tsvector index for keyword search
conflict_logDetected contradictions (category, old/new value, resolution)
chat_threadsChat session metadata
chat_messagesRaw messages with ingestion status
query_logsQuery + response audit trail

Qdrant

CollectionPurpose
facts3072-dim embeddings with payload: fact_text, conversation_date, date_int, category, fact_id

Project Structure

main.py # Ingestion pipeline orchestrator
config.py # Environment config + memory budgets
db.py # PostgreSQL operations (compound queries, connection pooling)
models.py # Data classes (UserProfile, Foresight, ConflictLog, etc.)
vector_store.py # Qdrant facts collection (search, upsert, rebuild)
gemini.py # Gemini client wrapper
run_locomo.py # LoCoMo benchmark runner
locomo/ # LoCoMo benchmark dataset (https://github.com/snap-research/locomo.git)
ingestion/
extractor.py # Single-call fact + foresight extraction (LLM)
profile_extractor.py # Profile update via add/update/delete ops (LLM)
profile_ops.py # Apply operations to profile text (pure Python)
profile_manager.py # Profile compression when over budget (LLM)
summary_manager.py # Summary compression when over budget (LLM)
prompts/
extraction.txt # Fact + foresight extraction prompt
profile_update.txt # Profile add/update/delete operations prompt
summary_compression.txt # Recursive archive compression prompt
profile_compression.txt # Profile compaction (merge similar) prompt
retrieval/
fetch_mem_service.py # retrieve_for_query() + hybrid search + context composers
vectorize_service.py # Gemini embedding (single + batch)
Dockerfile # Python 3.12 container
docker-compose.yml # Local PostgreSQL + Qdrant setup

Key Features

  • 2 LLM calls per ingestion — extract + profile update (parallel)
  • Two retrieval strategies — rolling summary for chat, hybrid search for queries
  • Permanent facts — never compressed, keyword + vector searchable

Performance (Supabase, Qdrant Cloud)

EndpointContextLLMTotal
Chat~600ms~2.5s~3.1s
Query~1.7s (embed + search)~2.2s~3.9s
EndpointContextLLMTotal
Chat (local)~5ms~2.5s~2.5s
Query (local)~40ms + embed 650ms~2.2s~2.9s

LoCoMo Benchmark Results

Evaluated on the LoCoMo benchmark across 10 sessions, scored by GPT-4 on a 1-5 scale.

MetricScore (%)
Overall (all categories)88.3%
Without adversarial90.2%

Category breakdown:

CategoryDescription
TemporalDate/time-based recall
Multi-hopReasoning across multiple facts
AdversarialDeliberately tricky questions

Adversarial questions are excluded in the filtered score since they test robustness to misleading prompts rather than memory accuracy.

Note on retrieval during the benchmark: although vector embeddings are generated and stored in Qdrant during ingestion, the LoCoMo run does not use the vector service at QA time. run_locomo.py retrieves context via db.get_chat_context only — i.e. rolling summary + foresight (the chat path, no RAG). The hybrid keyword + vector search exposed by retrieval/fetch_mem_service.py is wired up for the /query endpoint for the chat application.


Getting Started

1. Environment Setup

cp .env.example .env
# Edit .env with your credentials:# GEMINI_API_KEY, PG_HOST, PG_PORT, PG_USER, PG_PASSWORD, PG_DB# QDRANT_HOST, QDRANT_PORT (or QDRANT_URL + QDRANT_API_KEY for cloud)

2. Install dependencies

pip install -r requirements.txt

3. Local PostgreSQL + Qdrant (Docker)

docker-compose up -d

4. Run the LoCoMo benchmark

python run_locomo.py --samples 0 --workers 5
# or all 10 samples:
python run_locomo.py

About

No description, website, or topics provided.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

AI Companion with Long-Term Memory

A conversational AI companion that remembers past conversations using a multi-tier memory architecture. Two retrieval strategies: rolling summary for chat (fast, lossy), hybrid search on permanent facts for precise queries (lossless). 2 LLM calls per ingestion.

Architecture

Conversation messages <-> Ingestion / Retrieval libraries <-> Gemini 2.5 Flash
|
PostgreSQL + Qdrant
|
Memory Pipeline
(Ingestion + Profile + Summary + Facts)

Memory System

Short-term memory: Last 20 messages from the current thread (regardless of ingestion status) — immediate conversational context.

Long-term memory: 5 components:

ComponentDescriptionStorage
User ProfileExplicit facts + implicit traits, updated via add/update/delete opsPostgreSQL
ForesightTime-bounded events with valid_from / valid_until / duration_days, auto-expiredPostgreSQL
Rolling SummaryTwo-tier: Archive (compressed old) + Recent (date-tagged entries)PostgreSQL
Facts TablePermanent, never compressed — every extracted fact stored foreverPostgreSQL (tsvector + GIN)
Facts VectorsEmbeddings for semantic searchQdrant (3072-dim, cosine)

Two Retrieval Strategies

Chat endpoint (/chat) — conversational, general awareness:

 Profile + Foresight + Rolling Summary + Last 20 messages

Query endpoint (/query) — precise recall:

Parallel:
[Thread 1] Gemini embed query
[Thread 2] profile + foresight + messages + keyword search
Then:
Qdrant vector search (~200ms) → RRF fusion with keyword results → top 5 facts

Ingestion Pipeline

2 LLM calls per ingestion (extract + profile update), +1 when compression triggers.

Conversation (20 messages)
|
[1] FETCH CONTEXT (1 DB read)
| Rolling summary for coreference resolution
|
[2] EXTRACT (1 LLM call — Gemini)
| Input: conversation + rolling summary
| Output: consolidated facts (category-tagged, dated) + foresight signals
| Dates sanitized: invalid LLM dates replaced with current_date
|
[3] In parallel (3 threads):
| [Thread 1] PG store — 3 parallel connections:
| | facts INSERT + foresight expire/INSERT + summary read/append
| [Thread 2] Embed facts (Gemini embedding API)
| [Thread 3] Profile update (1 LLM call):
| | LLM receives numbered profile + new facts
| | Returns add/update/delete operations as JSON
| | Operations applied programmatically (not LLM rewrite)
| | Conflicts logged to conflict_log table
|
[4] Qdrant upsert (needs fact_ids from PG + embeddings from Gemini)
|
[5] Compression checks:
If rolling summary >= 90% of 10k token budget:
| → Evict oldest 25% of recent → compress into archive (1 LLM call)
| → Archive capped at 20% of budget
If profile >= 80% of 3k token budget:
→ Merge similar facts (1 LLM call)

Extraction Details

  • 1 LLM call per batch — no segmentation step
  • Produces consolidated facts — dense, self-contained sentences (not atomic fragments)
  • Attribution rules — every fact must name the person, pronouns resolved
  • Frequency tracking — recurring activities captured with explicit frequency
  • Verbatim details prioritized: signs, paintings, book titles, pet behaviors
  • Quality filters — greetings, pleasantries, acknowledgments excluded
  • Prior rolling summary used as context for coreference resolution (with guard: "do NOT extract from this")

Profile Update (add/update/delete operations)

  • LLM receives numbered profile lines: [1] - Rampal is 24 years old
  • Returns structured JSON operations referencing line numbers
  • apply_operations() in Python handles the text editing
  • More predictable than full LLM rewrite — LLM decides what, Python does editing
  • Conflict detection only on profile (facts table is append-only)

Foresight

  • Time-bounded events with valid_from, valid_until, duration_days, evidence
  • Auto-expired during ingestion when valid_until < current_date
  • Examples: travel plans, illness recovery, deadlines, new jobs

Compression

  • Trigger: 90% of 10k token budget
  • Eviction: oldest 25% of recent entries
  • Recursive: new_archive = LLM(existing_archive + evicted_entries)
  • Archive capped at 20% of budget — re-compressed if exceeded
  • Token count: char-based estimate (len/4), no API call

Ingestion Triggers

  • Production: After 20 unprocessed messages accumulate
  • Periodic: Every 10 minutes, checks for threads with 4+ old unprocessed messages
  • Manual: POST /threads/{thread_id}/ingest
  • Benchmark: After each session

Data Storage

PostgreSQL

TablePurpose
user_profileExplicit facts + implicit traits (bullet-point format)
foresightTime-bounded events with validity windows + duration_days
conversation_summariesTwo-tier: archive_text + recent_text + token_count
factsPermanent fact store with tsvector index for keyword search
conflict_logDetected contradictions (category, old/new value, resolution)
chat_threadsChat session metadata
chat_messagesRaw messages with ingestion status
query_logsQuery + response audit trail

Qdrant

CollectionPurpose
facts3072-dim embeddings with payload: fact_text, conversation_date, date_int, category, fact_id

Project Structure

main.py # Ingestion pipeline orchestrator
config.py # Environment config + memory budgets
db.py # PostgreSQL operations (compound queries, connection pooling)
models.py # Data classes (UserProfile, Foresight, ConflictLog, etc.)
vector_store.py # Qdrant facts collection (search, upsert, rebuild)
gemini.py # Gemini client wrapper
run_locomo.py # LoCoMo benchmark runner
locomo/ # LoCoMo benchmark dataset (https://github.com/snap-research/locomo.git)
ingestion/
extractor.py # Single-call fact + foresight extraction (LLM)
profile_extractor.py # Profile update via add/update/delete ops (LLM)
profile_ops.py # Apply operations to profile text (pure Python)
profile_manager.py # Profile compression when over budget (LLM)
summary_manager.py # Summary compression when over budget (LLM)
prompts/
extraction.txt # Fact + foresight extraction prompt
profile_update.txt # Profile add/update/delete operations prompt
summary_compression.txt # Recursive archive compression prompt
profile_compression.txt # Profile compaction (merge similar) prompt
retrieval/
fetch_mem_service.py # retrieve_for_query() + hybrid search + context composers
vectorize_service.py # Gemini embedding (single + batch)
Dockerfile # Python 3.12 container
docker-compose.yml # Local PostgreSQL + Qdrant setup

Key Features

  • 2 LLM calls per ingestion — extract + profile update (parallel)
  • Two retrieval strategies — rolling summary for chat, hybrid search for queries
  • Permanent facts — never compressed, keyword + vector searchable

Performance (Supabase, Qdrant Cloud)

EndpointContextLLMTotal
Chat~600ms~2.5s~3.1s
Query~1.7s (embed + search)~2.2s~3.9s
EndpointContextLLMTotal
Chat (local)~5ms~2.5s~2.5s
Query (local)~40ms + embed 650ms~2.2s~2.9s

LoCoMo Benchmark Results

Evaluated on the LoCoMo benchmark across 10 sessions, scored by GPT-4 on a 1-5 scale.

MetricScore (%)
Overall (all categories)88.3%
Without adversarial90.2%

Category breakdown:

CategoryDescription
TemporalDate/time-based recall
Multi-hopReasoning across multiple facts
AdversarialDeliberately tricky questions

Adversarial questions are excluded in the filtered score since they test robustness to misleading prompts rather than memory accuracy.

Note on retrieval during the benchmark: although vector embeddings are generated and stored in Qdrant during ingestion, the LoCoMo run does not use the vector service at QA time. run_locomo.py retrieves context via db.get_chat_context only — i.e. rolling summary + foresight (the chat path, no RAG). The hybrid keyword + vector search exposed by retrieval/fetch_mem_service.py is wired up for the /query endpoint for the chat application.


Getting Started

1. Environment Setup

cp .env.example .env
# Edit .env with your credentials:# GEMINI_API_KEY, PG_HOST, PG_PORT, PG_USER, PG_PASSWORD, PG_DB# QDRANT_HOST, QDRANT_PORT (or QDRANT_URL + QDRANT_API_KEY for cloud)

2. Install dependencies

pip install -r requirements.txt

3. Local PostgreSQL + Qdrant (Docker)

docker-compose up -d

4. Run the LoCoMo benchmark

python run_locomo.py --samples 0 --workers 5
# or all 10 samples:
python run_locomo.py

About

No description, website, or topics provided.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

AI Companion with Long-Term Memory

A conversational AI companion that remembers past conversations using a multi-tier memory architecture. Two retrieval strategies: rolling summary for chat (fast, lossy), hybrid search on permanent facts for precise queries (lossless). 2 LLM calls per ingestion.

Architecture

Conversation messages <-> Ingestion / Retrieval libraries <-> Gemini 2.5 Flash
|
PostgreSQL + Qdrant
|
Memory Pipeline
(Ingestion + Profile + Summary + Facts)

Memory System

Short-term memory: Last 20 messages from the current thread (regardless of ingestion status) — immediate conversational context.

Long-term memory: 5 components:

ComponentDescriptionStorage
User ProfileExplicit facts + implicit traits, updated via add/update/delete opsPostgreSQL
ForesightTime-bounded events with valid_from / valid_until / duration_days, auto-expiredPostgreSQL
Rolling SummaryTwo-tier: Archive (compressed old) + Recent (date-tagged entries)PostgreSQL
Facts TablePermanent, never compressed — every extracted fact stored foreverPostgreSQL (tsvector + GIN)
Facts VectorsEmbeddings for semantic searchQdrant (3072-dim, cosine)

Two Retrieval Strategies

Chat endpoint (/chat) — conversational, general awareness:

 Profile + Foresight + Rolling Summary + Last 20 messages

Query endpoint (/query) — precise recall:

Parallel:
[Thread 1] Gemini embed query
[Thread 2] profile + foresight + messages + keyword search
Then:
Qdrant vector search (~200ms) → RRF fusion with keyword results → top 5 facts

Ingestion Pipeline

2 LLM calls per ingestion (extract + profile update), +1 when compression triggers.

Conversation (20 messages)
|
[1] FETCH CONTEXT (1 DB read)
| Rolling summary for coreference resolution
|
[2] EXTRACT (1 LLM call — Gemini)
| Input: conversation + rolling summary
| Output: consolidated facts (category-tagged, dated) + foresight signals
| Dates sanitized: invalid LLM dates replaced with current_date
|
[3] In parallel (3 threads):
| [Thread 1] PG store — 3 parallel connections:
| | facts INSERT + foresight expire/INSERT + summary read/append
| [Thread 2] Embed facts (Gemini embedding API)
| [Thread 3] Profile update (1 LLM call):
| | LLM receives numbered profile + new facts
| | Returns add/update/delete operations as JSON
| | Operations applied programmatically (not LLM rewrite)
| | Conflicts logged to conflict_log table
|
[4] Qdrant upsert (needs fact_ids from PG + embeddings from Gemini)
|
[5] Compression checks:
If rolling summary >= 90% of 10k token budget:
| → Evict oldest 25% of recent → compress into archive (1 LLM call)
| → Archive capped at 20% of budget
If profile >= 80% of 3k token budget:
→ Merge similar facts (1 LLM call)

Extraction Details

  • 1 LLM call per batch — no segmentation step
  • Produces consolidated facts — dense, self-contained sentences (not atomic fragments)
  • Attribution rules — every fact must name the person, pronouns resolved
  • Frequency tracking — recurring activities captured with explicit frequency
  • Verbatim details prioritized: signs, paintings, book titles, pet behaviors
  • Quality filters — greetings, pleasantries, acknowledgments excluded
  • Prior rolling summary used as context for coreference resolution (with guard: "do NOT extract from this")

Profile Update (add/update/delete operations)

  • LLM receives numbered profile lines: [1] - Rampal is 24 years old
  • Returns structured JSON operations referencing line numbers
  • apply_operations() in Python handles the text editing
  • More predictable than full LLM rewrite — LLM decides what, Python does editing
  • Conflict detection only on profile (facts table is append-only)

Foresight

  • Time-bounded events with valid_from, valid_until, duration_days, evidence
  • Auto-expired during ingestion when valid_until < current_date
  • Examples: travel plans, illness recovery, deadlines, new jobs

Compression

  • Trigger: 90% of 10k token budget
  • Eviction: oldest 25% of recent entries
  • Recursive: new_archive = LLM(existing_archive + evicted_entries)
  • Archive capped at 20% of budget — re-compressed if exceeded
  • Token count: char-based estimate (len/4), no API call

Ingestion Triggers

  • Production: After 20 unprocessed messages accumulate
  • Periodic: Every 10 minutes, checks for threads with 4+ old unprocessed messages
  • Manual: POST /threads/{thread_id}/ingest
  • Benchmark: After each session

Data Storage

PostgreSQL

TablePurpose
user_profileExplicit facts + implicit traits (bullet-point format)
foresightTime-bounded events with validity windows + duration_days
conversation_summariesTwo-tier: archive_text + recent_text + token_count
factsPermanent fact store with tsvector index for keyword search
conflict_logDetected contradictions (category, old/new value, resolution)
chat_threadsChat session metadata
chat_messagesRaw messages with ingestion status
query_logsQuery + response audit trail

Qdrant

CollectionPurpose
facts3072-dim embeddings with payload: fact_text, conversation_date, date_int, category, fact_id

Project Structure

main.py # Ingestion pipeline orchestrator
config.py # Environment config + memory budgets
db.py # PostgreSQL operations (compound queries, connection pooling)
models.py # Data classes (UserProfile, Foresight, ConflictLog, etc.)
vector_store.py # Qdrant facts collection (search, upsert, rebuild)
gemini.py # Gemini client wrapper
run_locomo.py # LoCoMo benchmark runner
locomo/ # LoCoMo benchmark dataset (https://github.com/snap-research/locomo.git)
ingestion/
extractor.py # Single-call fact + foresight extraction (LLM)
profile_extractor.py # Profile update via add/update/delete ops (LLM)
profile_ops.py # Apply operations to profile text (pure Python)
profile_manager.py # Profile compression when over budget (LLM)
summary_manager.py # Summary compression when over budget (LLM)
prompts/
extraction.txt # Fact + foresight extraction prompt
profile_update.txt # Profile add/update/delete operations prompt
summary_compression.txt # Recursive archive compression prompt
profile_compression.txt # Profile compaction (merge similar) prompt
retrieval/
fetch_mem_service.py # retrieve_for_query() + hybrid search + context composers
vectorize_service.py # Gemini embedding (single + batch)
Dockerfile # Python 3.12 container
docker-compose.yml # Local PostgreSQL + Qdrant setup

Key Features

  • 2 LLM calls per ingestion — extract + profile update (parallel)
  • Two retrieval strategies — rolling summary for chat, hybrid search for queries
  • Permanent facts — never compressed, keyword + vector searchable

Performance (Supabase, Qdrant Cloud)

EndpointContextLLMTotal
Chat~600ms~2.5s~3.1s
Query~1.7s (embed + search)~2.2s~3.9s
EndpointContextLLMTotal
Chat (local)~5ms~2.5s~2.5s
Query (local)~40ms + embed 650ms~2.2s~2.9s

LoCoMo Benchmark Results

Evaluated on the LoCoMo benchmark across 10 sessions, scored by GPT-4 on a 1-5 scale.

MetricScore (%)
Overall (all categories)88.3%
Without adversarial90.2%

Category breakdown:

CategoryDescription
TemporalDate/time-based recall
Multi-hopReasoning across multiple facts
AdversarialDeliberately tricky questions

Adversarial questions are excluded in the filtered score since they test robustness to misleading prompts rather than memory accuracy.

Note on retrieval during the benchmark: although vector embeddings are generated and stored in Qdrant during ingestion, the LoCoMo run does not use the vector service at QA time. run_locomo.py retrieves context via db.get_chat_context only — i.e. rolling summary + foresight (the chat path, no RAG). The hybrid keyword + vector search exposed by retrieval/fetch_mem_service.py is wired up for the /query endpoint for the chat application.


Getting Started

1. Environment Setup

cp .env.example .env
# Edit .env with your credentials:# GEMINI_API_KEY, PG_HOST, PG_PORT, PG_USER, PG_PASSWORD, PG_DB# QDRANT_HOST, QDRANT_PORT (or QDRANT_URL + QDRANT_API_KEY for cloud)

2. Install dependencies

pip install -r requirements.txt

3. Local PostgreSQL + Qdrant (Docker)

docker-compose up -d

4. Run the LoCoMo benchmark

python run_locomo.py --samples 0 --workers 5
# or all 10 samples:
python run_locomo.py

About

No description, website, or topics provided.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

AI Companion with Long-Term Memory

A conversational AI companion that remembers past conversations using a multi-tier memory architecture. Two retrieval strategies: rolling summary for chat (fast, lossy), hybrid search on permanent facts for precise queries (lossless). 2 LLM calls per ingestion.

Architecture

Conversation messages <-> Ingestion / Retrieval libraries <-> Gemini 2.5 Flash
|
PostgreSQL + Qdrant
|
Memory Pipeline
(Ingestion + Profile + Summary + Facts)

Memory System

Short-term memory: Last 20 messages from the current thread (regardless of ingestion status) — immediate conversational context.

Long-term memory: 5 components:

ComponentDescriptionStorage
User ProfileExplicit facts + implicit traits, updated via add/update/delete opsPostgreSQL
ForesightTime-bounded events with valid_from / valid_until / duration_days, auto-expiredPostgreSQL
Rolling SummaryTwo-tier: Archive (compressed old) + Recent (date-tagged entries)PostgreSQL
Facts TablePermanent, never compressed — every extracted fact stored foreverPostgreSQL (tsvector + GIN)
Facts VectorsEmbeddings for semantic searchQdrant (3072-dim, cosine)

Two Retrieval Strategies

Chat endpoint (/chat) — conversational, general awareness:

 Profile + Foresight + Rolling Summary + Last 20 messages

Query endpoint (/query) — precise recall:

Parallel:
[Thread 1] Gemini embed query
[Thread 2] profile + foresight + messages + keyword search
Then:
Qdrant vector search (~200ms) → RRF fusion with keyword results → top 5 facts

Ingestion Pipeline

2 LLM calls per ingestion (extract + profile update), +1 when compression triggers.

Conversation (20 messages)
|
[1] FETCH CONTEXT (1 DB read)
| Rolling summary for coreference resolution
|
[2] EXTRACT (1 LLM call — Gemini)
| Input: conversation + rolling summary
| Output: consolidated facts (category-tagged, dated) + foresight signals
| Dates sanitized: invalid LLM dates replaced with current_date
|
[3] In parallel (3 threads):
| [Thread 1] PG store — 3 parallel connections:
| | facts INSERT + foresight expire/INSERT + summary read/append
| [Thread 2] Embed facts (Gemini embedding API)
| [Thread 3] Profile update (1 LLM call):
| | LLM receives numbered profile + new facts
| | Returns add/update/delete operations as JSON
| | Operations applied programmatically (not LLM rewrite)
| | Conflicts logged to conflict_log table
|
[4] Qdrant upsert (needs fact_ids from PG + embeddings from Gemini)
|
[5] Compression checks:
If rolling summary >= 90% of 10k token budget:
| → Evict oldest 25% of recent → compress into archive (1 LLM call)
| → Archive capped at 20% of budget
If profile >= 80% of 3k token budget:
→ Merge similar facts (1 LLM call)

Extraction Details

  • 1 LLM call per batch — no segmentation step
  • Produces consolidated facts — dense, self-contained sentences (not atomic fragments)
  • Attribution rules — every fact must name the person, pronouns resolved
  • Frequency tracking — recurring activities captured with explicit frequency
  • Verbatim details prioritized: signs, paintings, book titles, pet behaviors
  • Quality filters — greetings, pleasantries, acknowledgments excluded
  • Prior rolling summary used as context for coreference resolution (with guard: "do NOT extract from this")

Profile Update (add/update/delete operations)

  • LLM receives numbered profile lines: [1] - Rampal is 24 years old
  • Returns structured JSON operations referencing line numbers
  • apply_operations() in Python handles the text editing
  • More predictable than full LLM rewrite — LLM decides what, Python does editing
  • Conflict detection only on profile (facts table is append-only)

Foresight

  • Time-bounded events with valid_from, valid_until, duration_days, evidence
  • Auto-expired during ingestion when valid_until < current_date
  • Examples: travel plans, illness recovery, deadlines, new jobs

Compression

  • Trigger: 90% of 10k token budget
  • Eviction: oldest 25% of recent entries
  • Recursive: new_archive = LLM(existing_archive + evicted_entries)
  • Archive capped at 20% of budget — re-compressed if exceeded
  • Token count: char-based estimate (len/4), no API call

Ingestion Triggers

  • Production: After 20 unprocessed messages accumulate
  • Periodic: Every 10 minutes, checks for threads with 4+ old unprocessed messages
  • Manual: POST /threads/{thread_id}/ingest
  • Benchmark: After each session

Data Storage

PostgreSQL

TablePurpose
user_profileExplicit facts + implicit traits (bullet-point format)
foresightTime-bounded events with validity windows + duration_days
conversation_summariesTwo-tier: archive_text + recent_text + token_count
factsPermanent fact store with tsvector index for keyword search
conflict_logDetected contradictions (category, old/new value, resolution)
chat_threadsChat session metadata
chat_messagesRaw messages with ingestion status
query_logsQuery + response audit trail

Qdrant

CollectionPurpose
facts3072-dim embeddings with payload: fact_text, conversation_date, date_int, category, fact_id

Project Structure

main.py # Ingestion pipeline orchestrator
config.py # Environment config + memory budgets
db.py # PostgreSQL operations (compound queries, connection pooling)
models.py # Data classes (UserProfile, Foresight, ConflictLog, etc.)
vector_store.py # Qdrant facts collection (search, upsert, rebuild)
gemini.py # Gemini client wrapper
run_locomo.py # LoCoMo benchmark runner
locomo/ # LoCoMo benchmark dataset (https://github.com/snap-research/locomo.git)
ingestion/
extractor.py # Single-call fact + foresight extraction (LLM)
profile_extractor.py # Profile update via add/update/delete ops (LLM)
profile_ops.py # Apply operations to profile text (pure Python)
profile_manager.py # Profile compression when over budget (LLM)
summary_manager.py # Summary compression when over budget (LLM)
prompts/
extraction.txt # Fact + foresight extraction prompt
profile_update.txt # Profile add/update/delete operations prompt
summary_compression.txt # Recursive archive compression prompt
profile_compression.txt # Profile compaction (merge similar) prompt
retrieval/
fetch_mem_service.py # retrieve_for_query() + hybrid search + context composers
vectorize_service.py # Gemini embedding (single + batch)
Dockerfile # Python 3.12 container
docker-compose.yml # Local PostgreSQL + Qdrant setup

Key Features

  • 2 LLM calls per ingestion — extract + profile update (parallel)
  • Two retrieval strategies — rolling summary for chat, hybrid search for queries
  • Permanent facts — never compressed, keyword + vector searchable

Performance (Supabase, Qdrant Cloud)

EndpointContextLLMTotal
Chat~600ms~2.5s~3.1s
Query~1.7s (embed + search)~2.2s~3.9s
EndpointContextLLMTotal
Chat (local)~5ms~2.5s~2.5s
Query (local)~40ms + embed 650ms~2.2s~2.9s

LoCoMo Benchmark Results

Evaluated on the LoCoMo benchmark across 10 sessions, scored by GPT-4 on a 1-5 scale.

MetricScore (%)
Overall (all categories)88.3%
Without adversarial90.2%

Category breakdown:

CategoryDescription
TemporalDate/time-based recall
Multi-hopReasoning across multiple facts
AdversarialDeliberately tricky questions

Adversarial questions are excluded in the filtered score since they test robustness to misleading prompts rather than memory accuracy.

Note on retrieval during the benchmark: although vector embeddings are generated and stored in Qdrant during ingestion, the LoCoMo run does not use the vector service at QA time. run_locomo.py retrieves context via db.get_chat_context only — i.e. rolling summary + foresight (the chat path, no RAG). The hybrid keyword + vector search exposed by retrieval/fetch_mem_service.py is wired up for the /query endpoint for the chat application.


Getting Started

1. Environment Setup

cp .env.example .env
# Edit .env with your credentials:# GEMINI_API_KEY, PG_HOST, PG_PORT, PG_USER, PG_PASSWORD, PG_DB# QDRANT_HOST, QDRANT_PORT (or QDRANT_URL + QDRANT_API_KEY for cloud)

2. Install dependencies

pip install -r requirements.txt

3. Local PostgreSQL + Qdrant (Docker)

docker-compose up -d

4. Run the LoCoMo benchmark

python run_locomo.py --samples 0 --workers 5
# or all 10 samples:
python run_locomo.py

About

No description, website, or topics provided.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

AI Companion with Long-Term Memory

A conversational AI companion that remembers past conversations using a multi-tier memory architecture. Two retrieval strategies: rolling summary for chat (fast, lossy), hybrid search on permanent facts for precise queries (lossless). 2 LLM calls per ingestion.

Architecture

Conversation messages <-> Ingestion / Retrieval libraries <-> Gemini 2.5 Flash
|
PostgreSQL + Qdrant
|
Memory Pipeline
(Ingestion + Profile + Summary + Facts)

Memory System

Short-term memory: Last 20 messages from the current thread (regardless of ingestion status) — immediate conversational context.

Long-term memory: 5 components:

ComponentDescriptionStorage
User ProfileExplicit facts + implicit traits, updated via add/update/delete opsPostgreSQL
ForesightTime-bounded events with valid_from / valid_until / duration_days, auto-expiredPostgreSQL
Rolling SummaryTwo-tier: Archive (compressed old) + Recent (date-tagged entries)PostgreSQL
Facts TablePermanent, never compressed — every extracted fact stored foreverPostgreSQL (tsvector + GIN)
Facts VectorsEmbeddings for semantic searchQdrant (3072-dim, cosine)

Two Retrieval Strategies

Chat endpoint (/chat) — conversational, general awareness:

 Profile + Foresight + Rolling Summary + Last 20 messages

Query endpoint (/query) — precise recall:

Parallel:
[Thread 1] Gemini embed query
[Thread 2] profile + foresight + messages + keyword search
Then:
Qdrant vector search (~200ms) → RRF fusion with keyword results → top 5 facts

Ingestion Pipeline

2 LLM calls per ingestion (extract + profile update), +1 when compression triggers.

Conversation (20 messages)
|
[1] FETCH CONTEXT (1 DB read)
| Rolling summary for coreference resolution
|
[2] EXTRACT (1 LLM call — Gemini)
| Input: conversation + rolling summary
| Output: consolidated facts (category-tagged, dated) + foresight signals
| Dates sanitized: invalid LLM dates replaced with current_date
|
[3] In parallel (3 threads):
| [Thread 1] PG store — 3 parallel connections:
| | facts INSERT + foresight expire/INSERT + summary read/append
| [Thread 2] Embed facts (Gemini embedding API)
| [Thread 3] Profile update (1 LLM call):
| | LLM receives numbered profile + new facts
| | Returns add/update/delete operations as JSON
| | Operations applied programmatically (not LLM rewrite)
| | Conflicts logged to conflict_log table
|
[4] Qdrant upsert (needs fact_ids from PG + embeddings from Gemini)
|
[5] Compression checks:
If rolling summary >= 90% of 10k token budget:
| → Evict oldest 25% of recent → compress into archive (1 LLM call)
| → Archive capped at 20% of budget
If profile >= 80% of 3k token budget:
→ Merge similar facts (1 LLM call)

Extraction Details

  • 1 LLM call per batch — no segmentation step
  • Produces consolidated facts — dense, self-contained sentences (not atomic fragments)
  • Attribution rules — every fact must name the person, pronouns resolved
  • Frequency tracking — recurring activities captured with explicit frequency
  • Verbatim details prioritized: signs, paintings, book titles, pet behaviors
  • Quality filters — greetings, pleasantries, acknowledgments excluded
  • Prior rolling summary used as context for coreference resolution (with guard: "do NOT extract from this")

Profile Update (add/update/delete operations)

  • LLM receives numbered profile lines: [1] - Rampal is 24 years old
  • Returns structured JSON operations referencing line numbers
  • apply_operations() in Python handles the text editing
  • More predictable than full LLM rewrite — LLM decides what, Python does editing
  • Conflict detection only on profile (facts table is append-only)

Foresight

  • Time-bounded events with valid_from, valid_until, duration_days, evidence
  • Auto-expired during ingestion when valid_until < current_date
  • Examples: travel plans, illness recovery, deadlines, new jobs

Compression

  • Trigger: 90% of 10k token budget
  • Eviction: oldest 25% of recent entries
  • Recursive: new_archive = LLM(existing_archive + evicted_entries)
  • Archive capped at 20% of budget — re-compressed if exceeded
  • Token count: char-based estimate (len/4), no API call

Ingestion Triggers

  • Production: After 20 unprocessed messages accumulate
  • Periodic: Every 10 minutes, checks for threads with 4+ old unprocessed messages
  • Manual: POST /threads/{thread_id}/ingest
  • Benchmark: After each session

Data Storage

PostgreSQL

TablePurpose
user_profileExplicit facts + implicit traits (bullet-point format)
foresightTime-bounded events with validity windows + duration_days
conversation_summariesTwo-tier: archive_text + recent_text + token_count
factsPermanent fact store with tsvector index for keyword search
conflict_logDetected contradictions (category, old/new value, resolution)
chat_threadsChat session metadata
chat_messagesRaw messages with ingestion status
query_logsQuery + response audit trail

Qdrant

CollectionPurpose
facts3072-dim embeddings with payload: fact_text, conversation_date, date_int, category, fact_id

Project Structure

main.py # Ingestion pipeline orchestrator
config.py # Environment config + memory budgets
db.py # PostgreSQL operations (compound queries, connection pooling)
models.py # Data classes (UserProfile, Foresight, ConflictLog, etc.)
vector_store.py # Qdrant facts collection (search, upsert, rebuild)
gemini.py # Gemini client wrapper
run_locomo.py # LoCoMo benchmark runner
locomo/ # LoCoMo benchmark dataset (https://github.com/snap-research/locomo.git)
ingestion/
extractor.py # Single-call fact + foresight extraction (LLM)
profile_extractor.py # Profile update via add/update/delete ops (LLM)
profile_ops.py # Apply operations to profile text (pure Python)
profile_manager.py # Profile compression when over budget (LLM)
summary_manager.py # Summary compression when over budget (LLM)
prompts/
extraction.txt # Fact + foresight extraction prompt
profile_update.txt # Profile add/update/delete operations prompt
summary_compression.txt # Recursive archive compression prompt
profile_compression.txt # Profile compaction (merge similar) prompt
retrieval/
fetch_mem_service.py # retrieve_for_query() + hybrid search + context composers
vectorize_service.py # Gemini embedding (single + batch)
Dockerfile # Python 3.12 container
docker-compose.yml # Local PostgreSQL + Qdrant setup

Key Features

  • 2 LLM calls per ingestion — extract + profile update (parallel)
  • Two retrieval strategies — rolling summary for chat, hybrid search for queries
  • Permanent facts — never compressed, keyword + vector searchable

Performance (Supabase, Qdrant Cloud)

EndpointContextLLMTotal
Chat~600ms~2.5s~3.1s
Query~1.7s (embed + search)~2.2s~3.9s
EndpointContextLLMTotal
Chat (local)~5ms~2.5s~2.5s
Query (local)~40ms + embed 650ms~2.2s~2.9s

LoCoMo Benchmark Results

Evaluated on the LoCoMo benchmark across 10 sessions, scored by GPT-4 on a 1-5 scale.

MetricScore (%)
Overall (all categories)88.3%
Without adversarial90.2%

Category breakdown:

CategoryDescription
TemporalDate/time-based recall
Multi-hopReasoning across multiple facts
AdversarialDeliberately tricky questions

Adversarial questions are excluded in the filtered score since they test robustness to misleading prompts rather than memory accuracy.

Note on retrieval during the benchmark: although vector embeddings are generated and stored in Qdrant during ingestion, the LoCoMo run does not use the vector service at QA time. run_locomo.py retrieves context via db.get_chat_context only — i.e. rolling summary + foresight (the chat path, no RAG). The hybrid keyword + vector search exposed by retrieval/fetch_mem_service.py is wired up for the /query endpoint for the chat application.


Getting Started

1. Environment Setup

cp .env.example .env
# Edit .env with your credentials:# GEMINI_API_KEY, PG_HOST, PG_PORT, PG_USER, PG_PASSWORD, PG_DB# QDRANT_HOST, QDRANT_PORT (or QDRANT_URL + QDRANT_API_KEY for cloud)

2. Install dependencies

pip install -r requirements.txt

3. Local PostgreSQL + Qdrant (Docker)

docker-compose up -d

4. Run the LoCoMo benchmark

python run_locomo.py --samples 0 --workers 5
# or all 10 samples:
python run_locomo.py

About

No description, website, or topics provided.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

AI Companion with Long-Term Memory

A conversational AI companion that remembers past conversations using a multi-tier memory architecture. Two retrieval strategies: rolling summary for chat (fast, lossy), hybrid search on permanent facts for precise queries (lossless). 2 LLM calls per ingestion.

Architecture

Conversation messages <-> Ingestion / Retrieval libraries <-> Gemini 2.5 Flash
|
PostgreSQL + Qdrant
|
Memory Pipeline
(Ingestion + Profile + Summary + Facts)

Memory System

Short-term memory: Last 20 messages from the current thread (regardless of ingestion status) — immediate conversational context.

Long-term memory: 5 components:

ComponentDescriptionStorage
User ProfileExplicit facts + implicit traits, updated via add/update/delete opsPostgreSQL
ForesightTime-bounded events with valid_from / valid_until / duration_days, auto-expiredPostgreSQL
Rolling SummaryTwo-tier: Archive (compressed old) + Recent (date-tagged entries)PostgreSQL
Facts TablePermanent, never compressed — every extracted fact stored foreverPostgreSQL (tsvector + GIN)
Facts VectorsEmbeddings for semantic searchQdrant (3072-dim, cosine)

Two Retrieval Strategies

Chat endpoint (/chat) — conversational, general awareness:

 Profile + Foresight + Rolling Summary + Last 20 messages

Query endpoint (/query) — precise recall:

Parallel:
[Thread 1] Gemini embed query
[Thread 2] profile + foresight + messages + keyword search
Then:
Qdrant vector search (~200ms) → RRF fusion with keyword results → top 5 facts

Ingestion Pipeline

2 LLM calls per ingestion (extract + profile update), +1 when compression triggers.

Conversation (20 messages)
|
[1] FETCH CONTEXT (1 DB read)
| Rolling summary for coreference resolution
|
[2] EXTRACT (1 LLM call — Gemini)
| Input: conversation + rolling summary
| Output: consolidated facts (category-tagged, dated) + foresight signals
| Dates sanitized: invalid LLM dates replaced with current_date
|
[3] In parallel (3 threads):
| [Thread 1] PG store — 3 parallel connections:
| | facts INSERT + foresight expire/INSERT + summary read/append
| [Thread 2] Embed facts (Gemini embedding API)
| [Thread 3] Profile update (1 LLM call):
| | LLM receives numbered profile + new facts
| | Returns add/update/delete operations as JSON
| | Operations applied programmatically (not LLM rewrite)
| | Conflicts logged to conflict_log table
|
[4] Qdrant upsert (needs fact_ids from PG + embeddings from Gemini)
|
[5] Compression checks:
If rolling summary >= 90% of 10k token budget:
| → Evict oldest 25% of recent → compress into archive (1 LLM call)
| → Archive capped at 20% of budget
If profile >= 80% of 3k token budget:
→ Merge similar facts (1 LLM call)

Extraction Details

  • 1 LLM call per batch — no segmentation step
  • Produces consolidated facts — dense, self-contained sentences (not atomic fragments)
  • Attribution rules — every fact must name the person, pronouns resolved
  • Frequency tracking — recurring activities captured with explicit frequency
  • Verbatim details prioritized: signs, paintings, book titles, pet behaviors
  • Quality filters — greetings, pleasantries, acknowledgments excluded
  • Prior rolling summary used as context for coreference resolution (with guard: "do NOT extract from this")

Profile Update (add/update/delete operations)

  • LLM receives numbered profile lines: [1] - Rampal is 24 years old
  • Returns structured JSON operations referencing line numbers
  • apply_operations() in Python handles the text editing
  • More predictable than full LLM rewrite — LLM decides what, Python does editing
  • Conflict detection only on profile (facts table is append-only)

Foresight

  • Time-bounded events with valid_from, valid_until, duration_days, evidence
  • Auto-expired during ingestion when valid_until < current_date
  • Examples: travel plans, illness recovery, deadlines, new jobs

Compression

  • Trigger: 90% of 10k token budget
  • Eviction: oldest 25% of recent entries
  • Recursive: new_archive = LLM(existing_archive + evicted_entries)
  • Archive capped at 20% of budget — re-compressed if exceeded
  • Token count: char-based estimate (len/4), no API call

Ingestion Triggers

  • Production: After 20 unprocessed messages accumulate
  • Periodic: Every 10 minutes, checks for threads with 4+ old unprocessed messages
  • Manual: POST /threads/{thread_id}/ingest
  • Benchmark: After each session

Data Storage

PostgreSQL

TablePurpose
user_profileExplicit facts + implicit traits (bullet-point format)
foresightTime-bounded events with validity windows + duration_days
conversation_summariesTwo-tier: archive_text + recent_text + token_count
factsPermanent fact store with tsvector index for keyword search
conflict_logDetected contradictions (category, old/new value, resolution)
chat_threadsChat session metadata
chat_messagesRaw messages with ingestion status
query_logsQuery + response audit trail

Qdrant

CollectionPurpose
facts3072-dim embeddings with payload: fact_text, conversation_date, date_int, category, fact_id

Project Structure

main.py # Ingestion pipeline orchestrator
config.py # Environment config + memory budgets
db.py # PostgreSQL operations (compound queries, connection pooling)
models.py # Data classes (UserProfile, Foresight, ConflictLog, etc.)
vector_store.py # Qdrant facts collection (search, upsert, rebuild)
gemini.py # Gemini client wrapper
run_locomo.py # LoCoMo benchmark runner
locomo/ # LoCoMo benchmark dataset (https://github.com/snap-research/locomo.git)
ingestion/
extractor.py # Single-call fact + foresight extraction (LLM)
profile_extractor.py # Profile update via add/update/delete ops (LLM)
profile_ops.py # Apply operations to profile text (pure Python)
profile_manager.py # Profile compression when over budget (LLM)
summary_manager.py # Summary compression when over budget (LLM)
prompts/
extraction.txt # Fact + foresight extraction prompt
profile_update.txt # Profile add/update/delete operations prompt
summary_compression.txt # Recursive archive compression prompt
profile_compression.txt # Profile compaction (merge similar) prompt
retrieval/
fetch_mem_service.py # retrieve_for_query() + hybrid search + context composers
vectorize_service.py # Gemini embedding (single + batch)
Dockerfile # Python 3.12 container
docker-compose.yml # Local PostgreSQL + Qdrant setup

Key Features

  • 2 LLM calls per ingestion — extract + profile update (parallel)
  • Two retrieval strategies — rolling summary for chat, hybrid search for queries
  • Permanent facts — never compressed, keyword + vector searchable

Performance (Supabase, Qdrant Cloud)

EndpointContextLLMTotal
Chat~600ms~2.5s~3.1s
Query~1.7s (embed + search)~2.2s~3.9s
EndpointContextLLMTotal
Chat (local)~5ms~2.5s~2.5s
Query (local)~40ms + embed 650ms~2.2s~2.9s

LoCoMo Benchmark Results

Evaluated on the LoCoMo benchmark across 10 sessions, scored by GPT-4 on a 1-5 scale.

MetricScore (%)
Overall (all categories)88.3%
Without adversarial90.2%

Category breakdown:

CategoryDescription
TemporalDate/time-based recall
Multi-hopReasoning across multiple facts
AdversarialDeliberately tricky questions

Adversarial questions are excluded in the filtered score since they test robustness to misleading prompts rather than memory accuracy.

Note on retrieval during the benchmark: although vector embeddings are generated and stored in Qdrant during ingestion, the LoCoMo run does not use the vector service at QA time. run_locomo.py retrieves context via db.get_chat_context only — i.e. rolling summary + foresight (the chat path, no RAG). The hybrid keyword + vector search exposed by retrieval/fetch_mem_service.py is wired up for the /query endpoint for the chat application.


Getting Started

1. Environment Setup

cp .env.example .env
# Edit .env with your credentials:# GEMINI_API_KEY, PG_HOST, PG_PORT, PG_USER, PG_PASSWORD, PG_DB# QDRANT_HOST, QDRANT_PORT (or QDRANT_URL + QDRANT_API_KEY for cloud)

2. Install dependencies

pip install -r requirements.txt

3. Local PostgreSQL + Qdrant (Docker)

docker-compose up -d

4. Run the LoCoMo benchmark

python run_locomo.py --samples 0 --workers 5
# or all 10 samples:
python run_locomo.py

About

No description, website, or topics provided.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

AI Companion with Long-Term Memory

A conversational AI companion that remembers past conversations using a multi-tier memory architecture. Two retrieval strategies: rolling summary for chat (fast, lossy), hybrid search on permanent facts for precise queries (lossless). 2 LLM calls per ingestion.

Architecture

Conversation messages <-> Ingestion / Retrieval libraries <-> Gemini 2.5 Flash
|
PostgreSQL + Qdrant
|
Memory Pipeline
(Ingestion + Profile + Summary + Facts)

Memory System

Short-term memory: Last 20 messages from the current thread (regardless of ingestion status) — immediate conversational context.

Long-term memory: 5 components:

ComponentDescriptionStorage
User ProfileExplicit facts + implicit traits, updated via add/update/delete opsPostgreSQL
ForesightTime-bounded events with valid_from / valid_until / duration_days, auto-expiredPostgreSQL
Rolling SummaryTwo-tier: Archive (compressed old) + Recent (date-tagged entries)PostgreSQL
Facts TablePermanent, never compressed — every extracted fact stored foreverPostgreSQL (tsvector + GIN)
Facts VectorsEmbeddings for semantic searchQdrant (3072-dim, cosine)

Two Retrieval Strategies

Chat endpoint (/chat) — conversational, general awareness:

 Profile + Foresight + Rolling Summary + Last 20 messages

Query endpoint (/query) — precise recall:

Parallel:
[Thread 1] Gemini embed query
[Thread 2] profile + foresight + messages + keyword search
Then:
Qdrant vector search (~200ms) → RRF fusion with keyword results → top 5 facts

Ingestion Pipeline

2 LLM calls per ingestion (extract + profile update), +1 when compression triggers.

Conversation (20 messages)
|
[1] FETCH CONTEXT (1 DB read)
| Rolling summary for coreference resolution
|
[2] EXTRACT (1 LLM call — Gemini)
| Input: conversation + rolling summary
| Output: consolidated facts (category-tagged, dated) + foresight signals
| Dates sanitized: invalid LLM dates replaced with current_date
|
[3] In parallel (3 threads):
| [Thread 1] PG store — 3 parallel connections:
| | facts INSERT + foresight expire/INSERT + summary read/append
| [Thread 2] Embed facts (Gemini embedding API)
| [Thread 3] Profile update (1 LLM call):
| | LLM receives numbered profile + new facts
| | Returns add/update/delete operations as JSON
| | Operations applied programmatically (not LLM rewrite)
| | Conflicts logged to conflict_log table
|
[4] Qdrant upsert (needs fact_ids from PG + embeddings from Gemini)
|
[5] Compression checks:
If rolling summary >= 90% of 10k token budget:
| → Evict oldest 25% of recent → compress into archive (1 LLM call)
| → Archive capped at 20% of budget
If profile >= 80% of 3k token budget:
→ Merge similar facts (1 LLM call)

Extraction Details

  • 1 LLM call per batch — no segmentation step
  • Produces consolidated facts — dense, self-contained sentences (not atomic fragments)
  • Attribution rules — every fact must name the person, pronouns resolved
  • Frequency tracking — recurring activities captured with explicit frequency
  • Verbatim details prioritized: signs, paintings, book titles, pet behaviors
  • Quality filters — greetings, pleasantries, acknowledgments excluded
  • Prior rolling summary used as context for coreference resolution (with guard: "do NOT extract from this")

Profile Update (add/update/delete operations)

  • LLM receives numbered profile lines: [1] - Rampal is 24 years old
  • Returns structured JSON operations referencing line numbers
  • apply_operations() in Python handles the text editing
  • More predictable than full LLM rewrite — LLM decides what, Python does editing
  • Conflict detection only on profile (facts table is append-only)

Foresight

  • Time-bounded events with valid_from, valid_until, duration_days, evidence
  • Auto-expired during ingestion when valid_until < current_date
  • Examples: travel plans, illness recovery, deadlines, new jobs

Compression

  • Trigger: 90% of 10k token budget
  • Eviction: oldest 25% of recent entries
  • Recursive: new_archive = LLM(existing_archive + evicted_entries)
  • Archive capped at 20% of budget — re-compressed if exceeded
  • Token count: char-based estimate (len/4), no API call

Ingestion Triggers

  • Production: After 20 unprocessed messages accumulate
  • Periodic: Every 10 minutes, checks for threads with 4+ old unprocessed messages
  • Manual: POST /threads/{thread_id}/ingest
  • Benchmark: After each session

Data Storage

PostgreSQL

TablePurpose
user_profileExplicit facts + implicit traits (bullet-point format)
foresightTime-bounded events with validity windows + duration_days
conversation_summariesTwo-tier: archive_text + recent_text + token_count
factsPermanent fact store with tsvector index for keyword search
conflict_logDetected contradictions (category, old/new value, resolution)
chat_threadsChat session metadata
chat_messagesRaw messages with ingestion status
query_logsQuery + response audit trail

Qdrant

CollectionPurpose
facts3072-dim embeddings with payload: fact_text, conversation_date, date_int, category, fact_id

Project Structure

main.py # Ingestion pipeline orchestrator
config.py # Environment config + memory budgets
db.py # PostgreSQL operations (compound queries, connection pooling)
models.py # Data classes (UserProfile, Foresight, ConflictLog, etc.)
vector_store.py # Qdrant facts collection (search, upsert, rebuild)
gemini.py # Gemini client wrapper
run_locomo.py # LoCoMo benchmark runner
locomo/ # LoCoMo benchmark dataset (https://github.com/snap-research/locomo.git)
ingestion/
extractor.py # Single-call fact + foresight extraction (LLM)
profile_extractor.py # Profile update via add/update/delete ops (LLM)
profile_ops.py # Apply operations to profile text (pure Python)
profile_manager.py # Profile compression when over budget (LLM)
summary_manager.py # Summary compression when over budget (LLM)
prompts/
extraction.txt # Fact + foresight extraction prompt
profile_update.txt # Profile add/update/delete operations prompt
summary_compression.txt # Recursive archive compression prompt
profile_compression.txt # Profile compaction (merge similar) prompt
retrieval/
fetch_mem_service.py # retrieve_for_query() + hybrid search + context composers
vectorize_service.py # Gemini embedding (single + batch)
Dockerfile # Python 3.12 container
docker-compose.yml # Local PostgreSQL + Qdrant setup

Key Features

  • 2 LLM calls per ingestion — extract + profile update (parallel)
  • Two retrieval strategies — rolling summary for chat, hybrid search for queries
  • Permanent facts — never compressed, keyword + vector searchable

Performance (Supabase, Qdrant Cloud)

EndpointContextLLMTotal
Chat~600ms~2.5s~3.1s
Query~1.7s (embed + search)~2.2s~3.9s
EndpointContextLLMTotal
Chat (local)~5ms~2.5s~2.5s
Query (local)~40ms + embed 650ms~2.2s~2.9s

LoCoMo Benchmark Results

Evaluated on the LoCoMo benchmark across 10 sessions, scored by GPT-4 on a 1-5 scale.

MetricScore (%)
Overall (all categories)88.3%
Without adversarial90.2%

Category breakdown:

CategoryDescription
TemporalDate/time-based recall
Multi-hopReasoning across multiple facts
AdversarialDeliberately tricky questions

Adversarial questions are excluded in the filtered score since they test robustness to misleading prompts rather than memory accuracy.

Note on retrieval during the benchmark: although vector embeddings are generated and stored in Qdrant during ingestion, the LoCoMo run does not use the vector service at QA time. run_locomo.py retrieves context via db.get_chat_context only — i.e. rolling summary + foresight (the chat path, no RAG). The hybrid keyword + vector search exposed by retrieval/fetch_mem_service.py is wired up for the /query endpoint for the chat application.


Getting Started

1. Environment Setup

cp .env.example .env
# Edit .env with your credentials:# GEMINI_API_KEY, PG_HOST, PG_PORT, PG_USER, PG_PASSWORD, PG_DB# QDRANT_HOST, QDRANT_PORT (or QDRANT_URL + QDRANT_API_KEY for cloud)

2. Install dependencies

pip install -r requirements.txt

3. Local PostgreSQL + Qdrant (Docker)

docker-compose up -d

4. Run the LoCoMo benchmark

python run_locomo.py --samples 0 --workers 5
# or all 10 samples:
python run_locomo.py

About

No description, website, or topics provided.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

AI Companion with Long-Term Memory

A conversational AI companion that remembers past conversations using a multi-tier memory architecture. Two retrieval strategies: rolling summary for chat (fast, lossy), hybrid search on permanent facts for precise queries (lossless). 2 LLM calls per ingestion.

Architecture

Conversation messages <-> Ingestion / Retrieval libraries <-> Gemini 2.5 Flash
|
PostgreSQL + Qdrant
|
Memory Pipeline
(Ingestion + Profile + Summary + Facts)

Memory System

Short-term memory: Last 20 messages from the current thread (regardless of ingestion status) — immediate conversational context.

Long-term memory: 5 components:

ComponentDescriptionStorage
User ProfileExplicit facts + implicit traits, updated via add/update/delete opsPostgreSQL
ForesightTime-bounded events with valid_from / valid_until / duration_days, auto-expiredPostgreSQL
Rolling SummaryTwo-tier: Archive (compressed old) + Recent (date-tagged entries)PostgreSQL
Facts TablePermanent, never compressed — every extracted fact stored foreverPostgreSQL (tsvector + GIN)
Facts VectorsEmbeddings for semantic searchQdrant (3072-dim, cosine)

Two Retrieval Strategies

Chat endpoint (/chat) — conversational, general awareness:

 Profile + Foresight + Rolling Summary + Last 20 messages

Query endpoint (/query) — precise recall:

Parallel:
[Thread 1] Gemini embed query
[Thread 2] profile + foresight + messages + keyword search
Then:
Qdrant vector search (~200ms) → RRF fusion with keyword results → top 5 facts

Ingestion Pipeline

2 LLM calls per ingestion (extract + profile update), +1 when compression triggers.

Conversation (20 messages)
|
[1] FETCH CONTEXT (1 DB read)
| Rolling summary for coreference resolution
|
[2] EXTRACT (1 LLM call — Gemini)
| Input: conversation + rolling summary
| Output: consolidated facts (category-tagged, dated) + foresight signals
| Dates sanitized: invalid LLM dates replaced with current_date
|
[3] In parallel (3 threads):
| [Thread 1] PG store — 3 parallel connections:
| | facts INSERT + foresight expire/INSERT + summary read/append
| [Thread 2] Embed facts (Gemini embedding API)
| [Thread 3] Profile update (1 LLM call):
| | LLM receives numbered profile + new facts
| | Returns add/update/delete operations as JSON
| | Operations applied programmatically (not LLM rewrite)
| | Conflicts logged to conflict_log table
|
[4] Qdrant upsert (needs fact_ids from PG + embeddings from Gemini)
|
[5] Compression checks:
If rolling summary >= 90% of 10k token budget:
| → Evict oldest 25% of recent → compress into archive (1 LLM call)
| → Archive capped at 20% of budget
If profile >= 80% of 3k token budget:
→ Merge similar facts (1 LLM call)

Extraction Details

  • 1 LLM call per batch — no segmentation step
  • Produces consolidated facts — dense, self-contained sentences (not atomic fragments)
  • Attribution rules — every fact must name the person, pronouns resolved
  • Frequency tracking — recurring activities captured with explicit frequency
  • Verbatim details prioritized: signs, paintings, book titles, pet behaviors
  • Quality filters — greetings, pleasantries, acknowledgments excluded
  • Prior rolling summary used as context for coreference resolution (with guard: "do NOT extract from this")

Profile Update (add/update/delete operations)

  • LLM receives numbered profile lines: [1] - Rampal is 24 years old
  • Returns structured JSON operations referencing line numbers
  • apply_operations() in Python handles the text editing
  • More predictable than full LLM rewrite — LLM decides what, Python does editing
  • Conflict detection only on profile (facts table is append-only)

Foresight

  • Time-bounded events with valid_from, valid_until, duration_days, evidence
  • Auto-expired during ingestion when valid_until < current_date
  • Examples: travel plans, illness recovery, deadlines, new jobs

Compression

  • Trigger: 90% of 10k token budget
  • Eviction: oldest 25% of recent entries
  • Recursive: new_archive = LLM(existing_archive + evicted_entries)
  • Archive capped at 20% of budget — re-compressed if exceeded
  • Token count: char-based estimate (len/4), no API call

Ingestion Triggers

  • Production: After 20 unprocessed messages accumulate
  • Periodic: Every 10 minutes, checks for threads with 4+ old unprocessed messages
  • Manual: POST /threads/{thread_id}/ingest
  • Benchmark: After each session

Data Storage

PostgreSQL

TablePurpose
user_profileExplicit facts + implicit traits (bullet-point format)
foresightTime-bounded events with validity windows + duration_days
conversation_summariesTwo-tier: archive_text + recent_text + token_count
factsPermanent fact store with tsvector index for keyword search
conflict_logDetected contradictions (category, old/new value, resolution)
chat_threadsChat session metadata
chat_messagesRaw messages with ingestion status
query_logsQuery + response audit trail

Qdrant

CollectionPurpose
facts3072-dim embeddings with payload: fact_text, conversation_date, date_int, category, fact_id

Project Structure

main.py # Ingestion pipeline orchestrator
config.py # Environment config + memory budgets
db.py # PostgreSQL operations (compound queries, connection pooling)
models.py # Data classes (UserProfile, Foresight, ConflictLog, etc.)
vector_store.py # Qdrant facts collection (search, upsert, rebuild)
gemini.py # Gemini client wrapper
run_locomo.py # LoCoMo benchmark runner
locomo/ # LoCoMo benchmark dataset (https://github.com/snap-research/locomo.git)
ingestion/
extractor.py # Single-call fact + foresight extraction (LLM)
profile_extractor.py # Profile update via add/update/delete ops (LLM)
profile_ops.py # Apply operations to profile text (pure Python)
profile_manager.py # Profile compression when over budget (LLM)
summary_manager.py # Summary compression when over budget (LLM)
prompts/
extraction.txt # Fact + foresight extraction prompt
profile_update.txt # Profile add/update/delete operations prompt
summary_compression.txt # Recursive archive compression prompt
profile_compression.txt # Profile compaction (merge similar) prompt
retrieval/
fetch_mem_service.py # retrieve_for_query() + hybrid search + context composers
vectorize_service.py # Gemini embedding (single + batch)
Dockerfile # Python 3.12 container
docker-compose.yml # Local PostgreSQL + Qdrant setup

Key Features

  • 2 LLM calls per ingestion — extract + profile update (parallel)
  • Two retrieval strategies — rolling summary for chat, hybrid search for queries
  • Permanent facts — never compressed, keyword + vector searchable

Performance (Supabase, Qdrant Cloud)

EndpointContextLLMTotal
Chat~600ms~2.5s~3.1s
Query~1.7s (embed + search)~2.2s~3.9s
EndpointContextLLMTotal
Chat (local)~5ms~2.5s~2.5s
Query (local)~40ms + embed 650ms~2.2s~2.9s

LoCoMo Benchmark Results

Evaluated on the LoCoMo benchmark across 10 sessions, scored by GPT-4 on a 1-5 scale.

MetricScore (%)
Overall (all categories)88.3%
Without adversarial90.2%

Category breakdown:

CategoryDescription
TemporalDate/time-based recall
Multi-hopReasoning across multiple facts
AdversarialDeliberately tricky questions

Adversarial questions are excluded in the filtered score since they test robustness to misleading prompts rather than memory accuracy.

Note on retrieval during the benchmark: although vector embeddings are generated and stored in Qdrant during ingestion, the LoCoMo run does not use the vector service at QA time. run_locomo.py retrieves context via db.get_chat_context only — i.e. rolling summary + foresight (the chat path, no RAG). The hybrid keyword + vector search exposed by retrieval/fetch_mem_service.py is wired up for the /query endpoint for the chat application.


Getting Started

1. Environment Setup

cp .env.example .env
# Edit .env with your credentials:# GEMINI_API_KEY, PG_HOST, PG_PORT, PG_USER, PG_PASSWORD, PG_DB# QDRANT_HOST, QDRANT_PORT (or QDRANT_URL + QDRANT_API_KEY for cloud)

2. Install dependencies

pip install -r requirements.txt

3. Local PostgreSQL + Qdrant (Docker)

docker-compose up -d

4. Run the LoCoMo benchmark

python run_locomo.py --samples 0 --workers 5
# or all 10 samples:
python run_locomo.py

About

No description, website, or topics provided.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages