Repository files navigation

If you have feedback about puffgres, or are interested in building tools for the future of film and TV production, A24 Labs is hiring software engineers — email lgelfond@a24films.com for more.

puffgres (beta) is a logical replication service that keeps Postgres entities mirrored in turbopuffer. Rather than duplicating application code every time you modify a vector (and risking partial successes that keep data out of sync), your Postgres changes automatically update.

A bit of puffgres' design philosophy:

  • You should not need extra database calls to keep vectors up to date. Upserting rows in your primary database and a secondary vector database is bound to produce drift (forgetting to add parallel / compensating calls) and hard-to-detect failures (i.e. just one of the two calls succeeds). puffgres lets us "derive" state, making Postgres the source of truth and keeping Turbopuffer in sync.
  • The service handles at-least once delivery. Developers should not need to consider batching, retry logic, backfills, or change data capture in any of the code that they write. The service maintains its own state in a dedicated puffgres schema inside your source Postgres, and can stop/start/resume at any time without losing changes (even if they are slightly out of date). Co-locating state with the source means PITR restores naturally roll the two together.
  • Sync is maintained through "configs" which link Postgres tables to turbopuffer namespaces. Each defines a mapping, and a TypeScript-based "transform," which lets us easily do operations like tokenization, embedding, and other manipulation.
  • Configs and transforms are immutable. We avoid an abundance of thorny cases that come from letting us change a mapping (i.e. rows produced with two different set of transforms.). If we want to make a change, we should "tombstone" the old one and create a new one.

Read our docs to get started.

Install

Install the CLI as a native binary — under the hood verifies you have the Rust toolchain and builds from source + adds puffgres to your PATH.

curl -fsSL https://raw.githubusercontent.com/a24films/puffgres/main/install.sh | sh

Use with coding agents

The docs are published as one file at https://a24films.github.io/puffgres/AGENTS.md. Install it as a skill — paste one of these:

# Claude Code
mkdir -p ~/.claude/skills/puffgres && curl -fsSL https://a24films.github.io/puffgres/AGENTS.md -o ~/.claude/skills/puffgres/SKILL.md
# Codex
mkdir -p ~/.codex/skills/puffgres && curl -fsSL https://a24films.github.io/puffgres/AGENTS.md -o ~/.codex/skills/puffgres/SKILL.md

Why puffgres?

We built this because we needed vector embeddings internally and had read compelling evidence pgvector was a bad solution because of performance hits to maintain indexes, poor filtered queries, etc. Our naive / base solution was very hacky; we kept a separate table everytime we kept something in turbopuffer that kept an id, turbopuffer_updated_at, and updated_at and would simply embed / upsert whenever updated_at was more recent. This meant a full table scan whenever our pipeline ran (very inefficient) and effectively polling for changes in turbopuffer. It didn't handle deletes, required tons of duplicative code, and meant all updates only happened when the pipeline ran.

This was inspired by two tech talks: Martin Kleppmann's Turning the database inside out with Apache Samza, that argues strongly for making changes in one place and having derived data act like a materialized view, and Bryan Cantrill's Sharpening the Axe: The Primacy of Toolmaking, which suggests companies are well-suited investing in (and releasing) generic tools when they find themselves doing repeated work.

I built a very hacky version of this over a weekend, starting with a detailed spec and working through it with Claude. When we decided to use it in-house, I broke it up into PRs, and, especially at the beginning, ran each through code review + traditional testing, CI, etc. We've been running puffgres internally for a bit now without issue, and after talking with the turbopuffer team figured others might find use in it as well.

Performance

These benchmarks are a bit artificial, in isolation on a GitHub action (ubuntu-latest, the 4-core x86 / 16GB RAM runner). In practice, we've used puffgres on tables of a few hundred thousand rows, although it should scale much beyond this. If you implement this at large scale or hit bumps, feel free to shoot me an email. Initial benchmarking on GitHub Actions runners shows:

  • Throughput: >600K events/sec sustained over 100M events
  • Batch latency: p50 <10µs, p99 <100µs across 100K transactions
  • Recovery: <60ms to resume from checkpoint after crash
  • Memory: <160 bytes/event at scale, and total memory usage grows sub-linearly with event volume
  • Router fanout: >1M source events/sec across 1000 configs

Development

Package Organization

puffgres is a Rust workspace divided into several crates under crates/:

  • pg — Postgres setup, creates / manages publication and slot, generates schema files from table definitions, and runs backfill queries to bring in data
  • replication — the change data capture stream. Decodes Postgres logical replication protocol, manages caching relations, schema changes + in-transaction batching.
  • core — routes change events to their respective configs, runs transforms via a TypeScript subprocess, manages retry logic / dead letter queue, and calls to the puff client
  • puff — turbopuffer API client, light wrapper around rs-puff
  • state — stores persistent information (i.e. streaming replication checkpoints, backfill progress, failed entry queue) in a dedicated Postgres schema (by default: puffgres).
  • cli — the puffgres binary, handles subcommands, environment setup, orchestration, and default / template files.
  • config — definitions, parsing, validation, and hashing of config files.
  • debug — light server / web UI to inspect turbopuffer / WAL contents, just easier than the turbopuffer dashboard / making a bunch of cURLs

Documentation lives in docs/ and is built with mdbook.

Install

Build the binary from source:

  1. Install Just.
  2. Install the Rust toolchain.
  3. Build and install puffgres into ~/.cargo/bin:
just install
# To overwrite an existing install
just reinstall

Testing

# Unit tests
cargo test --workspace --lib
# Integration tests (requires Postgres)
cargo test --workspace

Benchmarks

cargo bench --package replication --bench decoder_bench

Fuzzing

Fuzz targets live in fuzz/ and use cargo-fuzz.

cargo install cargo-fuzz
cargo +nightly fuzz run fuzz_decoder
# Regenerate seed corpuscd fuzz && cargo run --bin generate_seeds

About

Keep Postgres entities synced with turbopuffer using logical replication (beta)

Resources

Stars

50 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

If you have feedback about puffgres, or are interested in building tools for the future of film and TV production, A24 Labs is hiring software engineers — email lgelfond@a24films.com for more.

puffgres (beta) is a logical replication service that keeps Postgres entities mirrored in turbopuffer. Rather than duplicating application code every time you modify a vector (and risking partial successes that keep data out of sync), your Postgres changes automatically update.

A bit of puffgres' design philosophy:

  • You should not need extra database calls to keep vectors up to date. Upserting rows in your primary database and a secondary vector database is bound to produce drift (forgetting to add parallel / compensating calls) and hard-to-detect failures (i.e. just one of the two calls succeeds). puffgres lets us "derive" state, making Postgres the source of truth and keeping Turbopuffer in sync.
  • The service handles at-least once delivery. Developers should not need to consider batching, retry logic, backfills, or change data capture in any of the code that they write. The service maintains its own state in a dedicated puffgres schema inside your source Postgres, and can stop/start/resume at any time without losing changes (even if they are slightly out of date). Co-locating state with the source means PITR restores naturally roll the two together.
  • Sync is maintained through "configs" which link Postgres tables to turbopuffer namespaces. Each defines a mapping, and a TypeScript-based "transform," which lets us easily do operations like tokenization, embedding, and other manipulation.
  • Configs and transforms are immutable. We avoid an abundance of thorny cases that come from letting us change a mapping (i.e. rows produced with two different set of transforms.). If we want to make a change, we should "tombstone" the old one and create a new one.

Read our docs to get started.

Install

Install the CLI as a native binary — under the hood verifies you have the Rust toolchain and builds from source + adds puffgres to your PATH.

curl -fsSL https://raw.githubusercontent.com/a24films/puffgres/main/install.sh | sh

Use with coding agents

The docs are published as one file at https://a24films.github.io/puffgres/AGENTS.md. Install it as a skill — paste one of these:

# Claude Code
mkdir -p ~/.claude/skills/puffgres && curl -fsSL https://a24films.github.io/puffgres/AGENTS.md -o ~/.claude/skills/puffgres/SKILL.md
# Codex
mkdir -p ~/.codex/skills/puffgres && curl -fsSL https://a24films.github.io/puffgres/AGENTS.md -o ~/.codex/skills/puffgres/SKILL.md

Why puffgres?

We built this because we needed vector embeddings internally and had read compelling evidence pgvector was a bad solution because of performance hits to maintain indexes, poor filtered queries, etc. Our naive / base solution was very hacky; we kept a separate table everytime we kept something in turbopuffer that kept an id, turbopuffer_updated_at, and updated_at and would simply embed / upsert whenever updated_at was more recent. This meant a full table scan whenever our pipeline ran (very inefficient) and effectively polling for changes in turbopuffer. It didn't handle deletes, required tons of duplicative code, and meant all updates only happened when the pipeline ran.

This was inspired by two tech talks: Martin Kleppmann's Turning the database inside out with Apache Samza, that argues strongly for making changes in one place and having derived data act like a materialized view, and Bryan Cantrill's Sharpening the Axe: The Primacy of Toolmaking, which suggests companies are well-suited investing in (and releasing) generic tools when they find themselves doing repeated work.

I built a very hacky version of this over a weekend, starting with a detailed spec and working through it with Claude. When we decided to use it in-house, I broke it up into PRs, and, especially at the beginning, ran each through code review + traditional testing, CI, etc. We've been running puffgres internally for a bit now without issue, and after talking with the turbopuffer team figured others might find use in it as well.

Performance

These benchmarks are a bit artificial, in isolation on a GitHub action (ubuntu-latest, the 4-core x86 / 16GB RAM runner). In practice, we've used puffgres on tables of a few hundred thousand rows, although it should scale much beyond this. If you implement this at large scale or hit bumps, feel free to shoot me an email. Initial benchmarking on GitHub Actions runners shows:

  • Throughput: >600K events/sec sustained over 100M events
  • Batch latency: p50 <10µs, p99 <100µs across 100K transactions
  • Recovery: <60ms to resume from checkpoint after crash
  • Memory: <160 bytes/event at scale, and total memory usage grows sub-linearly with event volume
  • Router fanout: >1M source events/sec across 1000 configs

Development

Package Organization

puffgres is a Rust workspace divided into several crates under crates/:

  • pg — Postgres setup, creates / manages publication and slot, generates schema files from table definitions, and runs backfill queries to bring in data
  • replication — the change data capture stream. Decodes Postgres logical replication protocol, manages caching relations, schema changes + in-transaction batching.
  • core — routes change events to their respective configs, runs transforms via a TypeScript subprocess, manages retry logic / dead letter queue, and calls to the puff client
  • puff — turbopuffer API client, light wrapper around rs-puff
  • state — stores persistent information (i.e. streaming replication checkpoints, backfill progress, failed entry queue) in a dedicated Postgres schema (by default: puffgres).
  • cli — the puffgres binary, handles subcommands, environment setup, orchestration, and default / template files.
  • config — definitions, parsing, validation, and hashing of config files.
  • debug — light server / web UI to inspect turbopuffer / WAL contents, just easier than the turbopuffer dashboard / making a bunch of cURLs

Documentation lives in docs/ and is built with mdbook.

Install

Build the binary from source:

  1. Install Just.
  2. Install the Rust toolchain.
  3. Build and install puffgres into ~/.cargo/bin:
just install
# To overwrite an existing install
just reinstall

Testing

# Unit tests
cargo test --workspace --lib
# Integration tests (requires Postgres)
cargo test --workspace

Benchmarks

cargo bench --package replication --bench decoder_bench

Fuzzing

Fuzz targets live in fuzz/ and use cargo-fuzz.

cargo install cargo-fuzz
cargo +nightly fuzz run fuzz_decoder
# Regenerate seed corpuscd fuzz && cargo run --bin generate_seeds

About

Keep Postgres entities synced with turbopuffer using logical replication (beta)

Resources

Stars

50 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

If you have feedback about puffgres, or are interested in building tools for the future of film and TV production, A24 Labs is hiring software engineers — email lgelfond@a24films.com for more.

puffgres (beta) is a logical replication service that keeps Postgres entities mirrored in turbopuffer. Rather than duplicating application code every time you modify a vector (and risking partial successes that keep data out of sync), your Postgres changes automatically update.

A bit of puffgres' design philosophy:

  • You should not need extra database calls to keep vectors up to date. Upserting rows in your primary database and a secondary vector database is bound to produce drift (forgetting to add parallel / compensating calls) and hard-to-detect failures (i.e. just one of the two calls succeeds). puffgres lets us "derive" state, making Postgres the source of truth and keeping Turbopuffer in sync.
  • The service handles at-least once delivery. Developers should not need to consider batching, retry logic, backfills, or change data capture in any of the code that they write. The service maintains its own state in a dedicated puffgres schema inside your source Postgres, and can stop/start/resume at any time without losing changes (even if they are slightly out of date). Co-locating state with the source means PITR restores naturally roll the two together.
  • Sync is maintained through "configs" which link Postgres tables to turbopuffer namespaces. Each defines a mapping, and a TypeScript-based "transform," which lets us easily do operations like tokenization, embedding, and other manipulation.
  • Configs and transforms are immutable. We avoid an abundance of thorny cases that come from letting us change a mapping (i.e. rows produced with two different set of transforms.). If we want to make a change, we should "tombstone" the old one and create a new one.

Read our docs to get started.

Install

Install the CLI as a native binary — under the hood verifies you have the Rust toolchain and builds from source + adds puffgres to your PATH.

curl -fsSL https://raw.githubusercontent.com/a24films/puffgres/main/install.sh | sh

Use with coding agents

The docs are published as one file at https://a24films.github.io/puffgres/AGENTS.md. Install it as a skill — paste one of these:

# Claude Code
mkdir -p ~/.claude/skills/puffgres && curl -fsSL https://a24films.github.io/puffgres/AGENTS.md -o ~/.claude/skills/puffgres/SKILL.md
# Codex
mkdir -p ~/.codex/skills/puffgres && curl -fsSL https://a24films.github.io/puffgres/AGENTS.md -o ~/.codex/skills/puffgres/SKILL.md

Why puffgres?

We built this because we needed vector embeddings internally and had read compelling evidence pgvector was a bad solution because of performance hits to maintain indexes, poor filtered queries, etc. Our naive / base solution was very hacky; we kept a separate table everytime we kept something in turbopuffer that kept an id, turbopuffer_updated_at, and updated_at and would simply embed / upsert whenever updated_at was more recent. This meant a full table scan whenever our pipeline ran (very inefficient) and effectively polling for changes in turbopuffer. It didn't handle deletes, required tons of duplicative code, and meant all updates only happened when the pipeline ran.

This was inspired by two tech talks: Martin Kleppmann's Turning the database inside out with Apache Samza, that argues strongly for making changes in one place and having derived data act like a materialized view, and Bryan Cantrill's Sharpening the Axe: The Primacy of Toolmaking, which suggests companies are well-suited investing in (and releasing) generic tools when they find themselves doing repeated work.

I built a very hacky version of this over a weekend, starting with a detailed spec and working through it with Claude. When we decided to use it in-house, I broke it up into PRs, and, especially at the beginning, ran each through code review + traditional testing, CI, etc. We've been running puffgres internally for a bit now without issue, and after talking with the turbopuffer team figured others might find use in it as well.

Performance

These benchmarks are a bit artificial, in isolation on a GitHub action (ubuntu-latest, the 4-core x86 / 16GB RAM runner). In practice, we've used puffgres on tables of a few hundred thousand rows, although it should scale much beyond this. If you implement this at large scale or hit bumps, feel free to shoot me an email. Initial benchmarking on GitHub Actions runners shows:

  • Throughput: >600K events/sec sustained over 100M events
  • Batch latency: p50 <10µs, p99 <100µs across 100K transactions
  • Recovery: <60ms to resume from checkpoint after crash
  • Memory: <160 bytes/event at scale, and total memory usage grows sub-linearly with event volume
  • Router fanout: >1M source events/sec across 1000 configs

Development

Package Organization

puffgres is a Rust workspace divided into several crates under crates/:

  • pg — Postgres setup, creates / manages publication and slot, generates schema files from table definitions, and runs backfill queries to bring in data
  • replication — the change data capture stream. Decodes Postgres logical replication protocol, manages caching relations, schema changes + in-transaction batching.
  • core — routes change events to their respective configs, runs transforms via a TypeScript subprocess, manages retry logic / dead letter queue, and calls to the puff client
  • puff — turbopuffer API client, light wrapper around rs-puff
  • state — stores persistent information (i.e. streaming replication checkpoints, backfill progress, failed entry queue) in a dedicated Postgres schema (by default: puffgres).
  • cli — the puffgres binary, handles subcommands, environment setup, orchestration, and default / template files.
  • config — definitions, parsing, validation, and hashing of config files.
  • debug — light server / web UI to inspect turbopuffer / WAL contents, just easier than the turbopuffer dashboard / making a bunch of cURLs

Documentation lives in docs/ and is built with mdbook.

Install

Build the binary from source:

  1. Install Just.
  2. Install the Rust toolchain.
  3. Build and install puffgres into ~/.cargo/bin:
just install
# To overwrite an existing install
just reinstall

Testing

# Unit tests
cargo test --workspace --lib
# Integration tests (requires Postgres)
cargo test --workspace

Benchmarks

cargo bench --package replication --bench decoder_bench

Fuzzing

Fuzz targets live in fuzz/ and use cargo-fuzz.

cargo install cargo-fuzz
cargo +nightly fuzz run fuzz_decoder
# Regenerate seed corpuscd fuzz && cargo run --bin generate_seeds

About

Keep Postgres entities synced with turbopuffer using logical replication (beta)

Resources

Stars

50 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

If you have feedback about puffgres, or are interested in building tools for the future of film and TV production, A24 Labs is hiring software engineers — email lgelfond@a24films.com for more.

puffgres (beta) is a logical replication service that keeps Postgres entities mirrored in turbopuffer. Rather than duplicating application code every time you modify a vector (and risking partial successes that keep data out of sync), your Postgres changes automatically update.

A bit of puffgres' design philosophy:

  • You should not need extra database calls to keep vectors up to date. Upserting rows in your primary database and a secondary vector database is bound to produce drift (forgetting to add parallel / compensating calls) and hard-to-detect failures (i.e. just one of the two calls succeeds). puffgres lets us "derive" state, making Postgres the source of truth and keeping Turbopuffer in sync.
  • The service handles at-least once delivery. Developers should not need to consider batching, retry logic, backfills, or change data capture in any of the code that they write. The service maintains its own state in a dedicated puffgres schema inside your source Postgres, and can stop/start/resume at any time without losing changes (even if they are slightly out of date). Co-locating state with the source means PITR restores naturally roll the two together.
  • Sync is maintained through "configs" which link Postgres tables to turbopuffer namespaces. Each defines a mapping, and a TypeScript-based "transform," which lets us easily do operations like tokenization, embedding, and other manipulation.
  • Configs and transforms are immutable. We avoid an abundance of thorny cases that come from letting us change a mapping (i.e. rows produced with two different set of transforms.). If we want to make a change, we should "tombstone" the old one and create a new one.

Read our docs to get started.

Install

Install the CLI as a native binary — under the hood verifies you have the Rust toolchain and builds from source + adds puffgres to your PATH.

curl -fsSL https://raw.githubusercontent.com/a24films/puffgres/main/install.sh | sh

Use with coding agents

The docs are published as one file at https://a24films.github.io/puffgres/AGENTS.md. Install it as a skill — paste one of these:

# Claude Code
mkdir -p ~/.claude/skills/puffgres && curl -fsSL https://a24films.github.io/puffgres/AGENTS.md -o ~/.claude/skills/puffgres/SKILL.md
# Codex
mkdir -p ~/.codex/skills/puffgres && curl -fsSL https://a24films.github.io/puffgres/AGENTS.md -o ~/.codex/skills/puffgres/SKILL.md

Why puffgres?

We built this because we needed vector embeddings internally and had read compelling evidence pgvector was a bad solution because of performance hits to maintain indexes, poor filtered queries, etc. Our naive / base solution was very hacky; we kept a separate table everytime we kept something in turbopuffer that kept an id, turbopuffer_updated_at, and updated_at and would simply embed / upsert whenever updated_at was more recent. This meant a full table scan whenever our pipeline ran (very inefficient) and effectively polling for changes in turbopuffer. It didn't handle deletes, required tons of duplicative code, and meant all updates only happened when the pipeline ran.

This was inspired by two tech talks: Martin Kleppmann's Turning the database inside out with Apache Samza, that argues strongly for making changes in one place and having derived data act like a materialized view, and Bryan Cantrill's Sharpening the Axe: The Primacy of Toolmaking, which suggests companies are well-suited investing in (and releasing) generic tools when they find themselves doing repeated work.

I built a very hacky version of this over a weekend, starting with a detailed spec and working through it with Claude. When we decided to use it in-house, I broke it up into PRs, and, especially at the beginning, ran each through code review + traditional testing, CI, etc. We've been running puffgres internally for a bit now without issue, and after talking with the turbopuffer team figured others might find use in it as well.

Performance

These benchmarks are a bit artificial, in isolation on a GitHub action (ubuntu-latest, the 4-core x86 / 16GB RAM runner). In practice, we've used puffgres on tables of a few hundred thousand rows, although it should scale much beyond this. If you implement this at large scale or hit bumps, feel free to shoot me an email. Initial benchmarking on GitHub Actions runners shows:

  • Throughput: >600K events/sec sustained over 100M events
  • Batch latency: p50 <10µs, p99 <100µs across 100K transactions
  • Recovery: <60ms to resume from checkpoint after crash
  • Memory: <160 bytes/event at scale, and total memory usage grows sub-linearly with event volume
  • Router fanout: >1M source events/sec across 1000 configs

Development

Package Organization

puffgres is a Rust workspace divided into several crates under crates/:

  • pg — Postgres setup, creates / manages publication and slot, generates schema files from table definitions, and runs backfill queries to bring in data
  • replication — the change data capture stream. Decodes Postgres logical replication protocol, manages caching relations, schema changes + in-transaction batching.
  • core — routes change events to their respective configs, runs transforms via a TypeScript subprocess, manages retry logic / dead letter queue, and calls to the puff client
  • puff — turbopuffer API client, light wrapper around rs-puff
  • state — stores persistent information (i.e. streaming replication checkpoints, backfill progress, failed entry queue) in a dedicated Postgres schema (by default: puffgres).
  • cli — the puffgres binary, handles subcommands, environment setup, orchestration, and default / template files.
  • config — definitions, parsing, validation, and hashing of config files.
  • debug — light server / web UI to inspect turbopuffer / WAL contents, just easier than the turbopuffer dashboard / making a bunch of cURLs

Documentation lives in docs/ and is built with mdbook.

Install

Build the binary from source:

  1. Install Just.
  2. Install the Rust toolchain.
  3. Build and install puffgres into ~/.cargo/bin:
just install
# To overwrite an existing install
just reinstall

Testing

# Unit tests
cargo test --workspace --lib
# Integration tests (requires Postgres)
cargo test --workspace

Benchmarks

cargo bench --package replication --bench decoder_bench

Fuzzing

Fuzz targets live in fuzz/ and use cargo-fuzz.

cargo install cargo-fuzz
cargo +nightly fuzz run fuzz_decoder
# Regenerate seed corpuscd fuzz && cargo run --bin generate_seeds

About

Keep Postgres entities synced with turbopuffer using logical replication (beta)

Resources

Stars

50 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

If you have feedback about puffgres, or are interested in building tools for the future of film and TV production, A24 Labs is hiring software engineers — email lgelfond@a24films.com for more.

puffgres (beta) is a logical replication service that keeps Postgres entities mirrored in turbopuffer. Rather than duplicating application code every time you modify a vector (and risking partial successes that keep data out of sync), your Postgres changes automatically update.

A bit of puffgres' design philosophy:

  • You should not need extra database calls to keep vectors up to date. Upserting rows in your primary database and a secondary vector database is bound to produce drift (forgetting to add parallel / compensating calls) and hard-to-detect failures (i.e. just one of the two calls succeeds). puffgres lets us "derive" state, making Postgres the source of truth and keeping Turbopuffer in sync.
  • The service handles at-least once delivery. Developers should not need to consider batching, retry logic, backfills, or change data capture in any of the code that they write. The service maintains its own state in a dedicated puffgres schema inside your source Postgres, and can stop/start/resume at any time without losing changes (even if they are slightly out of date). Co-locating state with the source means PITR restores naturally roll the two together.
  • Sync is maintained through "configs" which link Postgres tables to turbopuffer namespaces. Each defines a mapping, and a TypeScript-based "transform," which lets us easily do operations like tokenization, embedding, and other manipulation.
  • Configs and transforms are immutable. We avoid an abundance of thorny cases that come from letting us change a mapping (i.e. rows produced with two different set of transforms.). If we want to make a change, we should "tombstone" the old one and create a new one.

Read our docs to get started.

Install

Install the CLI as a native binary — under the hood verifies you have the Rust toolchain and builds from source + adds puffgres to your PATH.

curl -fsSL https://raw.githubusercontent.com/a24films/puffgres/main/install.sh | sh

Use with coding agents

The docs are published as one file at https://a24films.github.io/puffgres/AGENTS.md. Install it as a skill — paste one of these:

# Claude Code
mkdir -p ~/.claude/skills/puffgres && curl -fsSL https://a24films.github.io/puffgres/AGENTS.md -o ~/.claude/skills/puffgres/SKILL.md
# Codex
mkdir -p ~/.codex/skills/puffgres && curl -fsSL https://a24films.github.io/puffgres/AGENTS.md -o ~/.codex/skills/puffgres/SKILL.md

Why puffgres?

We built this because we needed vector embeddings internally and had read compelling evidence pgvector was a bad solution because of performance hits to maintain indexes, poor filtered queries, etc. Our naive / base solution was very hacky; we kept a separate table everytime we kept something in turbopuffer that kept an id, turbopuffer_updated_at, and updated_at and would simply embed / upsert whenever updated_at was more recent. This meant a full table scan whenever our pipeline ran (very inefficient) and effectively polling for changes in turbopuffer. It didn't handle deletes, required tons of duplicative code, and meant all updates only happened when the pipeline ran.

This was inspired by two tech talks: Martin Kleppmann's Turning the database inside out with Apache Samza, that argues strongly for making changes in one place and having derived data act like a materialized view, and Bryan Cantrill's Sharpening the Axe: The Primacy of Toolmaking, which suggests companies are well-suited investing in (and releasing) generic tools when they find themselves doing repeated work.

I built a very hacky version of this over a weekend, starting with a detailed spec and working through it with Claude. When we decided to use it in-house, I broke it up into PRs, and, especially at the beginning, ran each through code review + traditional testing, CI, etc. We've been running puffgres internally for a bit now without issue, and after talking with the turbopuffer team figured others might find use in it as well.

Performance

These benchmarks are a bit artificial, in isolation on a GitHub action (ubuntu-latest, the 4-core x86 / 16GB RAM runner). In practice, we've used puffgres on tables of a few hundred thousand rows, although it should scale much beyond this. If you implement this at large scale or hit bumps, feel free to shoot me an email. Initial benchmarking on GitHub Actions runners shows:

  • Throughput: >600K events/sec sustained over 100M events
  • Batch latency: p50 <10µs, p99 <100µs across 100K transactions
  • Recovery: <60ms to resume from checkpoint after crash
  • Memory: <160 bytes/event at scale, and total memory usage grows sub-linearly with event volume
  • Router fanout: >1M source events/sec across 1000 configs

Development

Package Organization

puffgres is a Rust workspace divided into several crates under crates/:

  • pg — Postgres setup, creates / manages publication and slot, generates schema files from table definitions, and runs backfill queries to bring in data
  • replication — the change data capture stream. Decodes Postgres logical replication protocol, manages caching relations, schema changes + in-transaction batching.
  • core — routes change events to their respective configs, runs transforms via a TypeScript subprocess, manages retry logic / dead letter queue, and calls to the puff client
  • puff — turbopuffer API client, light wrapper around rs-puff
  • state — stores persistent information (i.e. streaming replication checkpoints, backfill progress, failed entry queue) in a dedicated Postgres schema (by default: puffgres).
  • cli — the puffgres binary, handles subcommands, environment setup, orchestration, and default / template files.
  • config — definitions, parsing, validation, and hashing of config files.
  • debug — light server / web UI to inspect turbopuffer / WAL contents, just easier than the turbopuffer dashboard / making a bunch of cURLs

Documentation lives in docs/ and is built with mdbook.

Install

Build the binary from source:

  1. Install Just.
  2. Install the Rust toolchain.
  3. Build and install puffgres into ~/.cargo/bin:
just install
# To overwrite an existing install
just reinstall

Testing

# Unit tests
cargo test --workspace --lib
# Integration tests (requires Postgres)
cargo test --workspace

Benchmarks

cargo bench --package replication --bench decoder_bench

Fuzzing

Fuzz targets live in fuzz/ and use cargo-fuzz.

cargo install cargo-fuzz
cargo +nightly fuzz run fuzz_decoder
# Regenerate seed corpuscd fuzz && cargo run --bin generate_seeds

About

Keep Postgres entities synced with turbopuffer using logical replication (beta)

Resources

Stars

50 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

If you have feedback about puffgres, or are interested in building tools for the future of film and TV production, A24 Labs is hiring software engineers — email lgelfond@a24films.com for more.

puffgres (beta) is a logical replication service that keeps Postgres entities mirrored in turbopuffer. Rather than duplicating application code every time you modify a vector (and risking partial successes that keep data out of sync), your Postgres changes automatically update.

A bit of puffgres' design philosophy:

  • You should not need extra database calls to keep vectors up to date. Upserting rows in your primary database and a secondary vector database is bound to produce drift (forgetting to add parallel / compensating calls) and hard-to-detect failures (i.e. just one of the two calls succeeds). puffgres lets us "derive" state, making Postgres the source of truth and keeping Turbopuffer in sync.
  • The service handles at-least once delivery. Developers should not need to consider batching, retry logic, backfills, or change data capture in any of the code that they write. The service maintains its own state in a dedicated puffgres schema inside your source Postgres, and can stop/start/resume at any time without losing changes (even if they are slightly out of date). Co-locating state with the source means PITR restores naturally roll the two together.
  • Sync is maintained through "configs" which link Postgres tables to turbopuffer namespaces. Each defines a mapping, and a TypeScript-based "transform," which lets us easily do operations like tokenization, embedding, and other manipulation.
  • Configs and transforms are immutable. We avoid an abundance of thorny cases that come from letting us change a mapping (i.e. rows produced with two different set of transforms.). If we want to make a change, we should "tombstone" the old one and create a new one.

Read our docs to get started.

Install

Install the CLI as a native binary — under the hood verifies you have the Rust toolchain and builds from source + adds puffgres to your PATH.

curl -fsSL https://raw.githubusercontent.com/a24films/puffgres/main/install.sh | sh

Use with coding agents

The docs are published as one file at https://a24films.github.io/puffgres/AGENTS.md. Install it as a skill — paste one of these:

# Claude Code
mkdir -p ~/.claude/skills/puffgres && curl -fsSL https://a24films.github.io/puffgres/AGENTS.md -o ~/.claude/skills/puffgres/SKILL.md
# Codex
mkdir -p ~/.codex/skills/puffgres && curl -fsSL https://a24films.github.io/puffgres/AGENTS.md -o ~/.codex/skills/puffgres/SKILL.md

Why puffgres?

We built this because we needed vector embeddings internally and had read compelling evidence pgvector was a bad solution because of performance hits to maintain indexes, poor filtered queries, etc. Our naive / base solution was very hacky; we kept a separate table everytime we kept something in turbopuffer that kept an id, turbopuffer_updated_at, and updated_at and would simply embed / upsert whenever updated_at was more recent. This meant a full table scan whenever our pipeline ran (very inefficient) and effectively polling for changes in turbopuffer. It didn't handle deletes, required tons of duplicative code, and meant all updates only happened when the pipeline ran.

This was inspired by two tech talks: Martin Kleppmann's Turning the database inside out with Apache Samza, that argues strongly for making changes in one place and having derived data act like a materialized view, and Bryan Cantrill's Sharpening the Axe: The Primacy of Toolmaking, which suggests companies are well-suited investing in (and releasing) generic tools when they find themselves doing repeated work.

I built a very hacky version of this over a weekend, starting with a detailed spec and working through it with Claude. When we decided to use it in-house, I broke it up into PRs, and, especially at the beginning, ran each through code review + traditional testing, CI, etc. We've been running puffgres internally for a bit now without issue, and after talking with the turbopuffer team figured others might find use in it as well.

Performance

These benchmarks are a bit artificial, in isolation on a GitHub action (ubuntu-latest, the 4-core x86 / 16GB RAM runner). In practice, we've used puffgres on tables of a few hundred thousand rows, although it should scale much beyond this. If you implement this at large scale or hit bumps, feel free to shoot me an email. Initial benchmarking on GitHub Actions runners shows:

  • Throughput: >600K events/sec sustained over 100M events
  • Batch latency: p50 <10µs, p99 <100µs across 100K transactions
  • Recovery: <60ms to resume from checkpoint after crash
  • Memory: <160 bytes/event at scale, and total memory usage grows sub-linearly with event volume
  • Router fanout: >1M source events/sec across 1000 configs

Development

Package Organization

puffgres is a Rust workspace divided into several crates under crates/:

  • pg — Postgres setup, creates / manages publication and slot, generates schema files from table definitions, and runs backfill queries to bring in data
  • replication — the change data capture stream. Decodes Postgres logical replication protocol, manages caching relations, schema changes + in-transaction batching.
  • core — routes change events to their respective configs, runs transforms via a TypeScript subprocess, manages retry logic / dead letter queue, and calls to the puff client
  • puff — turbopuffer API client, light wrapper around rs-puff
  • state — stores persistent information (i.e. streaming replication checkpoints, backfill progress, failed entry queue) in a dedicated Postgres schema (by default: puffgres).
  • cli — the puffgres binary, handles subcommands, environment setup, orchestration, and default / template files.
  • config — definitions, parsing, validation, and hashing of config files.
  • debug — light server / web UI to inspect turbopuffer / WAL contents, just easier than the turbopuffer dashboard / making a bunch of cURLs

Documentation lives in docs/ and is built with mdbook.

Install

Build the binary from source:

  1. Install Just.
  2. Install the Rust toolchain.
  3. Build and install puffgres into ~/.cargo/bin:
just install
# To overwrite an existing install
just reinstall

Testing

# Unit tests
cargo test --workspace --lib
# Integration tests (requires Postgres)
cargo test --workspace

Benchmarks

cargo bench --package replication --bench decoder_bench

Fuzzing

Fuzz targets live in fuzz/ and use cargo-fuzz.

cargo install cargo-fuzz
cargo +nightly fuzz run fuzz_decoder
# Regenerate seed corpuscd fuzz && cargo run --bin generate_seeds

About

Keep Postgres entities synced with turbopuffer using logical replication (beta)

Resources

Stars

50 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

If you have feedback about puffgres, or are interested in building tools for the future of film and TV production, A24 Labs is hiring software engineers — email lgelfond@a24films.com for more.

puffgres (beta) is a logical replication service that keeps Postgres entities mirrored in turbopuffer. Rather than duplicating application code every time you modify a vector (and risking partial successes that keep data out of sync), your Postgres changes automatically update.

A bit of puffgres' design philosophy:

  • You should not need extra database calls to keep vectors up to date. Upserting rows in your primary database and a secondary vector database is bound to produce drift (forgetting to add parallel / compensating calls) and hard-to-detect failures (i.e. just one of the two calls succeeds). puffgres lets us "derive" state, making Postgres the source of truth and keeping Turbopuffer in sync.
  • The service handles at-least once delivery. Developers should not need to consider batching, retry logic, backfills, or change data capture in any of the code that they write. The service maintains its own state in a dedicated puffgres schema inside your source Postgres, and can stop/start/resume at any time without losing changes (even if they are slightly out of date). Co-locating state with the source means PITR restores naturally roll the two together.
  • Sync is maintained through "configs" which link Postgres tables to turbopuffer namespaces. Each defines a mapping, and a TypeScript-based "transform," which lets us easily do operations like tokenization, embedding, and other manipulation.
  • Configs and transforms are immutable. We avoid an abundance of thorny cases that come from letting us change a mapping (i.e. rows produced with two different set of transforms.). If we want to make a change, we should "tombstone" the old one and create a new one.

Read our docs to get started.

Install

Install the CLI as a native binary — under the hood verifies you have the Rust toolchain and builds from source + adds puffgres to your PATH.

curl -fsSL https://raw.githubusercontent.com/a24films/puffgres/main/install.sh | sh

Use with coding agents

The docs are published as one file at https://a24films.github.io/puffgres/AGENTS.md. Install it as a skill — paste one of these:

# Claude Code
mkdir -p ~/.claude/skills/puffgres && curl -fsSL https://a24films.github.io/puffgres/AGENTS.md -o ~/.claude/skills/puffgres/SKILL.md
# Codex
mkdir -p ~/.codex/skills/puffgres && curl -fsSL https://a24films.github.io/puffgres/AGENTS.md -o ~/.codex/skills/puffgres/SKILL.md

Why puffgres?

We built this because we needed vector embeddings internally and had read compelling evidence pgvector was a bad solution because of performance hits to maintain indexes, poor filtered queries, etc. Our naive / base solution was very hacky; we kept a separate table everytime we kept something in turbopuffer that kept an id, turbopuffer_updated_at, and updated_at and would simply embed / upsert whenever updated_at was more recent. This meant a full table scan whenever our pipeline ran (very inefficient) and effectively polling for changes in turbopuffer. It didn't handle deletes, required tons of duplicative code, and meant all updates only happened when the pipeline ran.

This was inspired by two tech talks: Martin Kleppmann's Turning the database inside out with Apache Samza, that argues strongly for making changes in one place and having derived data act like a materialized view, and Bryan Cantrill's Sharpening the Axe: The Primacy of Toolmaking, which suggests companies are well-suited investing in (and releasing) generic tools when they find themselves doing repeated work.

I built a very hacky version of this over a weekend, starting with a detailed spec and working through it with Claude. When we decided to use it in-house, I broke it up into PRs, and, especially at the beginning, ran each through code review + traditional testing, CI, etc. We've been running puffgres internally for a bit now without issue, and after talking with the turbopuffer team figured others might find use in it as well.

Performance

These benchmarks are a bit artificial, in isolation on a GitHub action (ubuntu-latest, the 4-core x86 / 16GB RAM runner). In practice, we've used puffgres on tables of a few hundred thousand rows, although it should scale much beyond this. If you implement this at large scale or hit bumps, feel free to shoot me an email. Initial benchmarking on GitHub Actions runners shows:

  • Throughput: >600K events/sec sustained over 100M events
  • Batch latency: p50 <10µs, p99 <100µs across 100K transactions
  • Recovery: <60ms to resume from checkpoint after crash
  • Memory: <160 bytes/event at scale, and total memory usage grows sub-linearly with event volume
  • Router fanout: >1M source events/sec across 1000 configs

Development

Package Organization

puffgres is a Rust workspace divided into several crates under crates/:

  • pg — Postgres setup, creates / manages publication and slot, generates schema files from table definitions, and runs backfill queries to bring in data
  • replication — the change data capture stream. Decodes Postgres logical replication protocol, manages caching relations, schema changes + in-transaction batching.
  • core — routes change events to their respective configs, runs transforms via a TypeScript subprocess, manages retry logic / dead letter queue, and calls to the puff client
  • puff — turbopuffer API client, light wrapper around rs-puff
  • state — stores persistent information (i.e. streaming replication checkpoints, backfill progress, failed entry queue) in a dedicated Postgres schema (by default: puffgres).
  • cli — the puffgres binary, handles subcommands, environment setup, orchestration, and default / template files.
  • config — definitions, parsing, validation, and hashing of config files.
  • debug — light server / web UI to inspect turbopuffer / WAL contents, just easier than the turbopuffer dashboard / making a bunch of cURLs

Documentation lives in docs/ and is built with mdbook.

Install

Build the binary from source:

  1. Install Just.
  2. Install the Rust toolchain.
  3. Build and install puffgres into ~/.cargo/bin:
just install
# To overwrite an existing install
just reinstall

Testing

# Unit tests
cargo test --workspace --lib
# Integration tests (requires Postgres)
cargo test --workspace

Benchmarks

cargo bench --package replication --bench decoder_bench

Fuzzing

Fuzz targets live in fuzz/ and use cargo-fuzz.

cargo install cargo-fuzz
cargo +nightly fuzz run fuzz_decoder
# Regenerate seed corpuscd fuzz && cargo run --bin generate_seeds

About

Keep Postgres entities synced with turbopuffer using logical replication (beta)

Resources

Stars

50 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

If you have feedback about puffgres, or are interested in building tools for the future of film and TV production, A24 Labs is hiring software engineers — email lgelfond@a24films.com for more.

puffgres (beta) is a logical replication service that keeps Postgres entities mirrored in turbopuffer. Rather than duplicating application code every time you modify a vector (and risking partial successes that keep data out of sync), your Postgres changes automatically update.

A bit of puffgres' design philosophy:

  • You should not need extra database calls to keep vectors up to date. Upserting rows in your primary database and a secondary vector database is bound to produce drift (forgetting to add parallel / compensating calls) and hard-to-detect failures (i.e. just one of the two calls succeeds). puffgres lets us "derive" state, making Postgres the source of truth and keeping Turbopuffer in sync.
  • The service handles at-least once delivery. Developers should not need to consider batching, retry logic, backfills, or change data capture in any of the code that they write. The service maintains its own state in a dedicated puffgres schema inside your source Postgres, and can stop/start/resume at any time without losing changes (even if they are slightly out of date). Co-locating state with the source means PITR restores naturally roll the two together.
  • Sync is maintained through "configs" which link Postgres tables to turbopuffer namespaces. Each defines a mapping, and a TypeScript-based "transform," which lets us easily do operations like tokenization, embedding, and other manipulation.
  • Configs and transforms are immutable. We avoid an abundance of thorny cases that come from letting us change a mapping (i.e. rows produced with two different set of transforms.). If we want to make a change, we should "tombstone" the old one and create a new one.

Read our docs to get started.

Install

Install the CLI as a native binary — under the hood verifies you have the Rust toolchain and builds from source + adds puffgres to your PATH.

curl -fsSL https://raw.githubusercontent.com/a24films/puffgres/main/install.sh | sh

Use with coding agents

The docs are published as one file at https://a24films.github.io/puffgres/AGENTS.md. Install it as a skill — paste one of these:

# Claude Code
mkdir -p ~/.claude/skills/puffgres && curl -fsSL https://a24films.github.io/puffgres/AGENTS.md -o ~/.claude/skills/puffgres/SKILL.md
# Codex
mkdir -p ~/.codex/skills/puffgres && curl -fsSL https://a24films.github.io/puffgres/AGENTS.md -o ~/.codex/skills/puffgres/SKILL.md

Why puffgres?

We built this because we needed vector embeddings internally and had read compelling evidence pgvector was a bad solution because of performance hits to maintain indexes, poor filtered queries, etc. Our naive / base solution was very hacky; we kept a separate table everytime we kept something in turbopuffer that kept an id, turbopuffer_updated_at, and updated_at and would simply embed / upsert whenever updated_at was more recent. This meant a full table scan whenever our pipeline ran (very inefficient) and effectively polling for changes in turbopuffer. It didn't handle deletes, required tons of duplicative code, and meant all updates only happened when the pipeline ran.

This was inspired by two tech talks: Martin Kleppmann's Turning the database inside out with Apache Samza, that argues strongly for making changes in one place and having derived data act like a materialized view, and Bryan Cantrill's Sharpening the Axe: The Primacy of Toolmaking, which suggests companies are well-suited investing in (and releasing) generic tools when they find themselves doing repeated work.

I built a very hacky version of this over a weekend, starting with a detailed spec and working through it with Claude. When we decided to use it in-house, I broke it up into PRs, and, especially at the beginning, ran each through code review + traditional testing, CI, etc. We've been running puffgres internally for a bit now without issue, and after talking with the turbopuffer team figured others might find use in it as well.

Performance

These benchmarks are a bit artificial, in isolation on a GitHub action (ubuntu-latest, the 4-core x86 / 16GB RAM runner). In practice, we've used puffgres on tables of a few hundred thousand rows, although it should scale much beyond this. If you implement this at large scale or hit bumps, feel free to shoot me an email. Initial benchmarking on GitHub Actions runners shows:

  • Throughput: >600K events/sec sustained over 100M events
  • Batch latency: p50 <10µs, p99 <100µs across 100K transactions
  • Recovery: <60ms to resume from checkpoint after crash
  • Memory: <160 bytes/event at scale, and total memory usage grows sub-linearly with event volume
  • Router fanout: >1M source events/sec across 1000 configs

Development

Package Organization

puffgres is a Rust workspace divided into several crates under crates/:

  • pg — Postgres setup, creates / manages publication and slot, generates schema files from table definitions, and runs backfill queries to bring in data
  • replication — the change data capture stream. Decodes Postgres logical replication protocol, manages caching relations, schema changes + in-transaction batching.
  • core — routes change events to their respective configs, runs transforms via a TypeScript subprocess, manages retry logic / dead letter queue, and calls to the puff client
  • puff — turbopuffer API client, light wrapper around rs-puff
  • state — stores persistent information (i.e. streaming replication checkpoints, backfill progress, failed entry queue) in a dedicated Postgres schema (by default: puffgres).
  • cli — the puffgres binary, handles subcommands, environment setup, orchestration, and default / template files.
  • config — definitions, parsing, validation, and hashing of config files.
  • debug — light server / web UI to inspect turbopuffer / WAL contents, just easier than the turbopuffer dashboard / making a bunch of cURLs

Documentation lives in docs/ and is built with mdbook.

Install

Build the binary from source:

  1. Install Just.
  2. Install the Rust toolchain.
  3. Build and install puffgres into ~/.cargo/bin:
just install
# To overwrite an existing install
just reinstall

Testing

# Unit tests
cargo test --workspace --lib
# Integration tests (requires Postgres)
cargo test --workspace

Benchmarks

cargo bench --package replication --bench decoder_bench

Fuzzing

Fuzz targets live in fuzz/ and use cargo-fuzz.

cargo install cargo-fuzz
cargo +nightly fuzz run fuzz_decoder
# Regenerate seed corpuscd fuzz && cargo run --bin generate_seeds

About

Keep Postgres entities synced with turbopuffer using logical replication (beta)

Resources

Stars

50 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages