Reduce sync work: faster snapshot projection + benchmarks - #86

Draft
hahn-kev-bot wants to merge 16 commits into
mainfrom
reduce-sync-work
Draft

Reduce sync work: faster snapshot projection + benchmarks#86
hahn-kev-bot wants to merge 16 commits into
mainfrom
reduce-sync-work

Conversation

@hahn-kev-bot

Copy link
Copy Markdown
Collaborator

Overview

Reduces the work done during sync (AddRangeFromSyncSnapshotWorker.UpdateSnapshotsCrdtRepository.AddSnapshots) and adds a benchmark suite to measure it. On the CreateWords workload at 10k changes this branch takes sync from ~1.96 s / 976 MB allocated down to ~0.99 s / 538 MB — roughly 2× faster and ~45% less memory.

What changed

Snapshot pre-load (SnapshotWorker / DataModel)

  • UpdateSnapshots now bulk-loads the relevant current snapshots (with their Commit) into a cache keyed by entity id, and SnapshotWorker reads full snapshots straight from that cache instead of issuing a FindSnapshot DB round-trip per cache hit.

Fast raw-SQL projection (FastProjection, new)

  • Snapshot rows are still inserted through EF unchanged, but the projected tables are now populated with hand-written raw SQL INSERT ... ON CONFLICT(pk) DO UPDATE (one upsert per entity row) instead of going through EF's change tracker (FindAsync / SetValues / graph tracking).
  • Everything the SQL needs — table/column names (correctly delimited), primary key, the SnapshotId shadow FK, value converters — is derived from the EF model, so there is no per-entity code.
  • It dedups to the latest snapshot per entity, runs deletes before upserts (children-first) then upserts (parents-first) for FK/unique-constraint safety, and reuses the caller's transaction.
  • FastProjection is an injectable singleton; its per-type SQL metadata cache lives on an internal ConcurrentDictionary on CrdtConfig, shared across repositories/contexts. This replaces the previous EF change-tracker projection path in CrdtRepository.AddSnapshots, which is removed.

Benchmarks (new SIL.Harmony.Benchmarks project)

  • BenchmarkDotNet suite with a DataModelSyncBenchmarks (7 sync workloads) and an AddSnapshotsBenchmarks that isolates the persist step, [MemoryDiagnoser] enabled. Run with dotnet run -c Release --project src/SIL.Harmony.Benchmarks.

Testing

  • Full test suite passes: 241 passing. The only failures are the 6 DataModelPerformanceBenchmarks timing-threshold tests, which also fail on main (environmental, not caused by this change).

Notes for reviewers

  • The projection is SQLite-specific in a couple of spots that matter for FK correctness (GUIDs stored as uppercase TEXT; ON CONFLICT after a SELECT needs the SQLite WHERE true disambiguator in the code history — the current per-query path uses VALUES). If other providers are ever targeted, FastProjection would need revisiting.
  • Known limitation: an intra-batch self-reference (e.g. Word.AntonymId pointing at another new Word in the same batch) is not ordered; it's a nullable SET NULL FK and not exercised by current workloads.

hahn-kevand others added 11 commits July 21, 2026 10:42
Previously AddSnapshots projected each snapshot by calling FindAsync per
entity, which issued one database query per snapshot (and, on an initial
sync of new data, every query returned null after a round-trip).
Pre-load the projected rows that already exist for the batch with a single
tracked query per object type. ProjectSnapshot then resolves existing
entities from the change tracker and skips the lookup entirely for entities
that have no projected row yet, collapsing N queries down to roughly one
per distinct object type.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019zEmz7jPRPF6Lv8h6YAWBW
Adds a BenchmarkDotNet suite that measures CrdtRepository.AddSnapshots on
its own, across 7 workloads mirroring DataModelSyncBenchmarks. Expensive DB
seeding runs once in a template DB; each iteration forks the DB and recomputes
the snapshot batch so no EF-tracked state leaks across iterations.
- SnapshotWorker.ComputeSnapshotsToPersist: returns the exact snapshot list
UpdateSnapshots would persist, without writing it.
- DataModelTestBase: internal CreateRepository() and CrdtConfig accessors.
- BenchmarkWorkloadBuilders: shared commit builders extracted from
DataModelSyncBenchmarks (+ BuildUpdateExisting).
- Program.cs: run both suites via BenchmarkSwitcher (handles --filter/args).
- Remove leftover Console.WriteLine debug lines from AddSnapshots.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two experimental fast AddSnapshots implementations that keep the EF snapshot
insert unchanged but populate projected tables with raw INSERT ... ON CONFLICT
upserts instead of going through EF's change tracker:
- FAST: one upsert command per entity row
- FAST_JSON: one command per entity type, rows passed as a single JSON array
expanded with SQLite json_each/json_extract
FastProjection derives table/column names, primary key, the SnapshotId shadow
FK, and value converters from the EF model (no per-entity code). It dedups to
the latest snapshot per entity, runs deletes before upserts (children-first)
then upserts (parents-first) for FK/unique-constraint safety, and reuses the
caller's transaction.
CrdtRepository.AddSnapshots now selects via #if FAST_JSON / #elif FAST / #else.
Program.cs adds a third FAST_JSON benchmark job and DataModelSyncBenchmarks
enables [MemoryDiagnoser].
Benchmarks (CreateWords, 1000): both fast paths ~35% faster and ~32% fewer
allocations than baseline; per-query vs JSON-batch shows no measurable
difference against in-memory SQLite.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The JSON-batch projection benchmarked identically to the per-query path against
in-memory SQLite (same time and allocations), so remove it and keep only the
per-query raw-SQL upsert path. FastProjection loses the useJsonBatch parameter
and all json_each/json_extract code; CrdtRepository.AddSnapshots collapses to
#if FAST / #else; the benchmark drops the FAST_JSON job.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Remove the EF change-tracker slow path and the #if FAST conditional so
AddSnapshots always uses FastProjection. Deletes the now-dead slow-path helpers
(ProjectSnapshot, GetEntityEntry, LoadExistingEntityIds, LoadExistingEntities).
The benchmark collapses to a single job since FAST vs DEFAULT are now identical.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
FastProjection is now an injected singleton (registered in AddCrdtDataCore and
resolved into CrdtRepository via ActivatorUtilities) instead of a static class.
Its per-type projected-table SQL metadata cache moves from a static field onto
an internal ConcurrentDictionary on CrdtConfig, so it's shared across
repositories/contexts and tied to config lifetime.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
auto_review:
drafts: true
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch reduce-sync-work

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

# Conflicts:
#	src/SIL.Harmony/Config/HarmonyConfig.cs
#	src/SIL.Harmony/SnapshotWorker.cs
Notify DI interceptors and HarmonyConfig.OnProjectedEntitiesChanged after projected SQL with the latest upsert or delete per entity.
Keep the slnx migration from main and include SIL.Harmony.Benchmarks in the solution.
Main now uses Microsoft.Testing.Platform, so solution-wide dotnet test was launching the Benchmarks exe and failing on unknown MTP flags.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@hahn-kev-bot@hahn-kev@claude
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Reduce sync work: faster snapshot projection + benchmarks - #86

Draft
hahn-kev-bot wants to merge 16 commits into
mainfrom
reduce-sync-work
Draft

Reduce sync work: faster snapshot projection + benchmarks#86
hahn-kev-bot wants to merge 16 commits into
mainfrom
reduce-sync-work

Conversation

@hahn-kev-bot

Copy link
Copy Markdown
Collaborator

Overview

Reduces the work done during sync (AddRangeFromSyncSnapshotWorker.UpdateSnapshotsCrdtRepository.AddSnapshots) and adds a benchmark suite to measure it. On the CreateWords workload at 10k changes this branch takes sync from ~1.96 s / 976 MB allocated down to ~0.99 s / 538 MB — roughly 2× faster and ~45% less memory.

What changed

Snapshot pre-load (SnapshotWorker / DataModel)

  • UpdateSnapshots now bulk-loads the relevant current snapshots (with their Commit) into a cache keyed by entity id, and SnapshotWorker reads full snapshots straight from that cache instead of issuing a FindSnapshot DB round-trip per cache hit.

Fast raw-SQL projection (FastProjection, new)

  • Snapshot rows are still inserted through EF unchanged, but the projected tables are now populated with hand-written raw SQL INSERT ... ON CONFLICT(pk) DO UPDATE (one upsert per entity row) instead of going through EF's change tracker (FindAsync / SetValues / graph tracking).
  • Everything the SQL needs — table/column names (correctly delimited), primary key, the SnapshotId shadow FK, value converters — is derived from the EF model, so there is no per-entity code.
  • It dedups to the latest snapshot per entity, runs deletes before upserts (children-first) then upserts (parents-first) for FK/unique-constraint safety, and reuses the caller's transaction.
  • FastProjection is an injectable singleton; its per-type SQL metadata cache lives on an internal ConcurrentDictionary on CrdtConfig, shared across repositories/contexts. This replaces the previous EF change-tracker projection path in CrdtRepository.AddSnapshots, which is removed.

Benchmarks (new SIL.Harmony.Benchmarks project)

  • BenchmarkDotNet suite with a DataModelSyncBenchmarks (7 sync workloads) and an AddSnapshotsBenchmarks that isolates the persist step, [MemoryDiagnoser] enabled. Run with dotnet run -c Release --project src/SIL.Harmony.Benchmarks.

Testing

  • Full test suite passes: 241 passing. The only failures are the 6 DataModelPerformanceBenchmarks timing-threshold tests, which also fail on main (environmental, not caused by this change).

Notes for reviewers

  • The projection is SQLite-specific in a couple of spots that matter for FK correctness (GUIDs stored as uppercase TEXT; ON CONFLICT after a SELECT needs the SQLite WHERE true disambiguator in the code history — the current per-query path uses VALUES). If other providers are ever targeted, FastProjection would need revisiting.
  • Known limitation: an intra-batch self-reference (e.g. Word.AntonymId pointing at another new Word in the same batch) is not ordered; it's a nullable SET NULL FK and not exercised by current workloads.

hahn-kevand others added 11 commits July 21, 2026 10:42
Previously AddSnapshots projected each snapshot by calling FindAsync per
entity, which issued one database query per snapshot (and, on an initial
sync of new data, every query returned null after a round-trip).
Pre-load the projected rows that already exist for the batch with a single
tracked query per object type. ProjectSnapshot then resolves existing
entities from the change tracker and skips the lookup entirely for entities
that have no projected row yet, collapsing N queries down to roughly one
per distinct object type.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019zEmz7jPRPF6Lv8h6YAWBW
Adds a BenchmarkDotNet suite that measures CrdtRepository.AddSnapshots on
its own, across 7 workloads mirroring DataModelSyncBenchmarks. Expensive DB
seeding runs once in a template DB; each iteration forks the DB and recomputes
the snapshot batch so no EF-tracked state leaks across iterations.
- SnapshotWorker.ComputeSnapshotsToPersist: returns the exact snapshot list
UpdateSnapshots would persist, without writing it.
- DataModelTestBase: internal CreateRepository() and CrdtConfig accessors.
- BenchmarkWorkloadBuilders: shared commit builders extracted from
DataModelSyncBenchmarks (+ BuildUpdateExisting).
- Program.cs: run both suites via BenchmarkSwitcher (handles --filter/args).
- Remove leftover Console.WriteLine debug lines from AddSnapshots.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two experimental fast AddSnapshots implementations that keep the EF snapshot
insert unchanged but populate projected tables with raw INSERT ... ON CONFLICT
upserts instead of going through EF's change tracker:
- FAST: one upsert command per entity row
- FAST_JSON: one command per entity type, rows passed as a single JSON array
expanded with SQLite json_each/json_extract
FastProjection derives table/column names, primary key, the SnapshotId shadow
FK, and value converters from the EF model (no per-entity code). It dedups to
the latest snapshot per entity, runs deletes before upserts (children-first)
then upserts (parents-first) for FK/unique-constraint safety, and reuses the
caller's transaction.
CrdtRepository.AddSnapshots now selects via #if FAST_JSON / #elif FAST / #else.
Program.cs adds a third FAST_JSON benchmark job and DataModelSyncBenchmarks
enables [MemoryDiagnoser].
Benchmarks (CreateWords, 1000): both fast paths ~35% faster and ~32% fewer
allocations than baseline; per-query vs JSON-batch shows no measurable
difference against in-memory SQLite.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The JSON-batch projection benchmarked identically to the per-query path against
in-memory SQLite (same time and allocations), so remove it and keep only the
per-query raw-SQL upsert path. FastProjection loses the useJsonBatch parameter
and all json_each/json_extract code; CrdtRepository.AddSnapshots collapses to
#if FAST / #else; the benchmark drops the FAST_JSON job.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Remove the EF change-tracker slow path and the #if FAST conditional so
AddSnapshots always uses FastProjection. Deletes the now-dead slow-path helpers
(ProjectSnapshot, GetEntityEntry, LoadExistingEntityIds, LoadExistingEntities).
The benchmark collapses to a single job since FAST vs DEFAULT are now identical.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
FastProjection is now an injected singleton (registered in AddCrdtDataCore and
resolved into CrdtRepository via ActivatorUtilities) instead of a static class.
Its per-type projected-table SQL metadata cache moves from a static field onto
an internal ConcurrentDictionary on CrdtConfig, so it's shared across
repositories/contexts and tied to config lifetime.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
auto_review:
drafts: true
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch reduce-sync-work

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

# Conflicts:
#	src/SIL.Harmony/Config/HarmonyConfig.cs
#	src/SIL.Harmony/SnapshotWorker.cs
Notify DI interceptors and HarmonyConfig.OnProjectedEntitiesChanged after projected SQL with the latest upsert or delete per entity.
Keep the slnx migration from main and include SIL.Harmony.Benchmarks in the solution.
Main now uses Microsoft.Testing.Platform, so solution-wide dotnet test was launching the Benchmarks exe and failing on unknown MTP flags.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@hahn-kev-bot@hahn-kev@claude
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Reduce sync work: faster snapshot projection + benchmarks - #86

Draft
hahn-kev-bot wants to merge 16 commits into
mainfrom
reduce-sync-work
Draft

Reduce sync work: faster snapshot projection + benchmarks#86
hahn-kev-bot wants to merge 16 commits into
mainfrom
reduce-sync-work

Conversation

@hahn-kev-bot

Copy link
Copy Markdown
Collaborator

Overview

Reduces the work done during sync (AddRangeFromSyncSnapshotWorker.UpdateSnapshotsCrdtRepository.AddSnapshots) and adds a benchmark suite to measure it. On the CreateWords workload at 10k changes this branch takes sync from ~1.96 s / 976 MB allocated down to ~0.99 s / 538 MB — roughly 2× faster and ~45% less memory.

What changed

Snapshot pre-load (SnapshotWorker / DataModel)

  • UpdateSnapshots now bulk-loads the relevant current snapshots (with their Commit) into a cache keyed by entity id, and SnapshotWorker reads full snapshots straight from that cache instead of issuing a FindSnapshot DB round-trip per cache hit.

Fast raw-SQL projection (FastProjection, new)

  • Snapshot rows are still inserted through EF unchanged, but the projected tables are now populated with hand-written raw SQL INSERT ... ON CONFLICT(pk) DO UPDATE (one upsert per entity row) instead of going through EF's change tracker (FindAsync / SetValues / graph tracking).
  • Everything the SQL needs — table/column names (correctly delimited), primary key, the SnapshotId shadow FK, value converters — is derived from the EF model, so there is no per-entity code.
  • It dedups to the latest snapshot per entity, runs deletes before upserts (children-first) then upserts (parents-first) for FK/unique-constraint safety, and reuses the caller's transaction.
  • FastProjection is an injectable singleton; its per-type SQL metadata cache lives on an internal ConcurrentDictionary on CrdtConfig, shared across repositories/contexts. This replaces the previous EF change-tracker projection path in CrdtRepository.AddSnapshots, which is removed.

Benchmarks (new SIL.Harmony.Benchmarks project)

  • BenchmarkDotNet suite with a DataModelSyncBenchmarks (7 sync workloads) and an AddSnapshotsBenchmarks that isolates the persist step, [MemoryDiagnoser] enabled. Run with dotnet run -c Release --project src/SIL.Harmony.Benchmarks.

Testing

  • Full test suite passes: 241 passing. The only failures are the 6 DataModelPerformanceBenchmarks timing-threshold tests, which also fail on main (environmental, not caused by this change).

Notes for reviewers

  • The projection is SQLite-specific in a couple of spots that matter for FK correctness (GUIDs stored as uppercase TEXT; ON CONFLICT after a SELECT needs the SQLite WHERE true disambiguator in the code history — the current per-query path uses VALUES). If other providers are ever targeted, FastProjection would need revisiting.
  • Known limitation: an intra-batch self-reference (e.g. Word.AntonymId pointing at another new Word in the same batch) is not ordered; it's a nullable SET NULL FK and not exercised by current workloads.

hahn-kevand others added 11 commits July 21, 2026 10:42
Previously AddSnapshots projected each snapshot by calling FindAsync per
entity, which issued one database query per snapshot (and, on an initial
sync of new data, every query returned null after a round-trip).
Pre-load the projected rows that already exist for the batch with a single
tracked query per object type. ProjectSnapshot then resolves existing
entities from the change tracker and skips the lookup entirely for entities
that have no projected row yet, collapsing N queries down to roughly one
per distinct object type.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019zEmz7jPRPF6Lv8h6YAWBW
Adds a BenchmarkDotNet suite that measures CrdtRepository.AddSnapshots on
its own, across 7 workloads mirroring DataModelSyncBenchmarks. Expensive DB
seeding runs once in a template DB; each iteration forks the DB and recomputes
the snapshot batch so no EF-tracked state leaks across iterations.
- SnapshotWorker.ComputeSnapshotsToPersist: returns the exact snapshot list
UpdateSnapshots would persist, without writing it.
- DataModelTestBase: internal CreateRepository() and CrdtConfig accessors.
- BenchmarkWorkloadBuilders: shared commit builders extracted from
DataModelSyncBenchmarks (+ BuildUpdateExisting).
- Program.cs: run both suites via BenchmarkSwitcher (handles --filter/args).
- Remove leftover Console.WriteLine debug lines from AddSnapshots.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two experimental fast AddSnapshots implementations that keep the EF snapshot
insert unchanged but populate projected tables with raw INSERT ... ON CONFLICT
upserts instead of going through EF's change tracker:
- FAST: one upsert command per entity row
- FAST_JSON: one command per entity type, rows passed as a single JSON array
expanded with SQLite json_each/json_extract
FastProjection derives table/column names, primary key, the SnapshotId shadow
FK, and value converters from the EF model (no per-entity code). It dedups to
the latest snapshot per entity, runs deletes before upserts (children-first)
then upserts (parents-first) for FK/unique-constraint safety, and reuses the
caller's transaction.
CrdtRepository.AddSnapshots now selects via #if FAST_JSON / #elif FAST / #else.
Program.cs adds a third FAST_JSON benchmark job and DataModelSyncBenchmarks
enables [MemoryDiagnoser].
Benchmarks (CreateWords, 1000): both fast paths ~35% faster and ~32% fewer
allocations than baseline; per-query vs JSON-batch shows no measurable
difference against in-memory SQLite.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The JSON-batch projection benchmarked identically to the per-query path against
in-memory SQLite (same time and allocations), so remove it and keep only the
per-query raw-SQL upsert path. FastProjection loses the useJsonBatch parameter
and all json_each/json_extract code; CrdtRepository.AddSnapshots collapses to
#if FAST / #else; the benchmark drops the FAST_JSON job.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Remove the EF change-tracker slow path and the #if FAST conditional so
AddSnapshots always uses FastProjection. Deletes the now-dead slow-path helpers
(ProjectSnapshot, GetEntityEntry, LoadExistingEntityIds, LoadExistingEntities).
The benchmark collapses to a single job since FAST vs DEFAULT are now identical.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
FastProjection is now an injected singleton (registered in AddCrdtDataCore and
resolved into CrdtRepository via ActivatorUtilities) instead of a static class.
Its per-type projected-table SQL metadata cache moves from a static field onto
an internal ConcurrentDictionary on CrdtConfig, so it's shared across
repositories/contexts and tied to config lifetime.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
auto_review:
drafts: true
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch reduce-sync-work

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

# Conflicts:
#	src/SIL.Harmony/Config/HarmonyConfig.cs
#	src/SIL.Harmony/SnapshotWorker.cs
Notify DI interceptors and HarmonyConfig.OnProjectedEntitiesChanged after projected SQL with the latest upsert or delete per entity.
Keep the slnx migration from main and include SIL.Harmony.Benchmarks in the solution.
Main now uses Microsoft.Testing.Platform, so solution-wide dotnet test was launching the Benchmarks exe and failing on unknown MTP flags.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@hahn-kev-bot@hahn-kev@claude
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Reduce sync work: faster snapshot projection + benchmarks - #86

Draft
hahn-kev-bot wants to merge 16 commits into
mainfrom
reduce-sync-work
Draft

Reduce sync work: faster snapshot projection + benchmarks#86
hahn-kev-bot wants to merge 16 commits into
mainfrom
reduce-sync-work

Conversation

@hahn-kev-bot

Copy link
Copy Markdown
Collaborator

Overview

Reduces the work done during sync (AddRangeFromSyncSnapshotWorker.UpdateSnapshotsCrdtRepository.AddSnapshots) and adds a benchmark suite to measure it. On the CreateWords workload at 10k changes this branch takes sync from ~1.96 s / 976 MB allocated down to ~0.99 s / 538 MB — roughly 2× faster and ~45% less memory.

What changed

Snapshot pre-load (SnapshotWorker / DataModel)

  • UpdateSnapshots now bulk-loads the relevant current snapshots (with their Commit) into a cache keyed by entity id, and SnapshotWorker reads full snapshots straight from that cache instead of issuing a FindSnapshot DB round-trip per cache hit.

Fast raw-SQL projection (FastProjection, new)

  • Snapshot rows are still inserted through EF unchanged, but the projected tables are now populated with hand-written raw SQL INSERT ... ON CONFLICT(pk) DO UPDATE (one upsert per entity row) instead of going through EF's change tracker (FindAsync / SetValues / graph tracking).
  • Everything the SQL needs — table/column names (correctly delimited), primary key, the SnapshotId shadow FK, value converters — is derived from the EF model, so there is no per-entity code.
  • It dedups to the latest snapshot per entity, runs deletes before upserts (children-first) then upserts (parents-first) for FK/unique-constraint safety, and reuses the caller's transaction.
  • FastProjection is an injectable singleton; its per-type SQL metadata cache lives on an internal ConcurrentDictionary on CrdtConfig, shared across repositories/contexts. This replaces the previous EF change-tracker projection path in CrdtRepository.AddSnapshots, which is removed.

Benchmarks (new SIL.Harmony.Benchmarks project)

  • BenchmarkDotNet suite with a DataModelSyncBenchmarks (7 sync workloads) and an AddSnapshotsBenchmarks that isolates the persist step, [MemoryDiagnoser] enabled. Run with dotnet run -c Release --project src/SIL.Harmony.Benchmarks.

Testing

  • Full test suite passes: 241 passing. The only failures are the 6 DataModelPerformanceBenchmarks timing-threshold tests, which also fail on main (environmental, not caused by this change).

Notes for reviewers

  • The projection is SQLite-specific in a couple of spots that matter for FK correctness (GUIDs stored as uppercase TEXT; ON CONFLICT after a SELECT needs the SQLite WHERE true disambiguator in the code history — the current per-query path uses VALUES). If other providers are ever targeted, FastProjection would need revisiting.
  • Known limitation: an intra-batch self-reference (e.g. Word.AntonymId pointing at another new Word in the same batch) is not ordered; it's a nullable SET NULL FK and not exercised by current workloads.

hahn-kevand others added 11 commits July 21, 2026 10:42
Previously AddSnapshots projected each snapshot by calling FindAsync per
entity, which issued one database query per snapshot (and, on an initial
sync of new data, every query returned null after a round-trip).
Pre-load the projected rows that already exist for the batch with a single
tracked query per object type. ProjectSnapshot then resolves existing
entities from the change tracker and skips the lookup entirely for entities
that have no projected row yet, collapsing N queries down to roughly one
per distinct object type.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019zEmz7jPRPF6Lv8h6YAWBW
Adds a BenchmarkDotNet suite that measures CrdtRepository.AddSnapshots on
its own, across 7 workloads mirroring DataModelSyncBenchmarks. Expensive DB
seeding runs once in a template DB; each iteration forks the DB and recomputes
the snapshot batch so no EF-tracked state leaks across iterations.
- SnapshotWorker.ComputeSnapshotsToPersist: returns the exact snapshot list
UpdateSnapshots would persist, without writing it.
- DataModelTestBase: internal CreateRepository() and CrdtConfig accessors.
- BenchmarkWorkloadBuilders: shared commit builders extracted from
DataModelSyncBenchmarks (+ BuildUpdateExisting).
- Program.cs: run both suites via BenchmarkSwitcher (handles --filter/args).
- Remove leftover Console.WriteLine debug lines from AddSnapshots.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two experimental fast AddSnapshots implementations that keep the EF snapshot
insert unchanged but populate projected tables with raw INSERT ... ON CONFLICT
upserts instead of going through EF's change tracker:
- FAST: one upsert command per entity row
- FAST_JSON: one command per entity type, rows passed as a single JSON array
expanded with SQLite json_each/json_extract
FastProjection derives table/column names, primary key, the SnapshotId shadow
FK, and value converters from the EF model (no per-entity code). It dedups to
the latest snapshot per entity, runs deletes before upserts (children-first)
then upserts (parents-first) for FK/unique-constraint safety, and reuses the
caller's transaction.
CrdtRepository.AddSnapshots now selects via #if FAST_JSON / #elif FAST / #else.
Program.cs adds a third FAST_JSON benchmark job and DataModelSyncBenchmarks
enables [MemoryDiagnoser].
Benchmarks (CreateWords, 1000): both fast paths ~35% faster and ~32% fewer
allocations than baseline; per-query vs JSON-batch shows no measurable
difference against in-memory SQLite.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The JSON-batch projection benchmarked identically to the per-query path against
in-memory SQLite (same time and allocations), so remove it and keep only the
per-query raw-SQL upsert path. FastProjection loses the useJsonBatch parameter
and all json_each/json_extract code; CrdtRepository.AddSnapshots collapses to
#if FAST / #else; the benchmark drops the FAST_JSON job.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Remove the EF change-tracker slow path and the #if FAST conditional so
AddSnapshots always uses FastProjection. Deletes the now-dead slow-path helpers
(ProjectSnapshot, GetEntityEntry, LoadExistingEntityIds, LoadExistingEntities).
The benchmark collapses to a single job since FAST vs DEFAULT are now identical.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
FastProjection is now an injected singleton (registered in AddCrdtDataCore and
resolved into CrdtRepository via ActivatorUtilities) instead of a static class.
Its per-type projected-table SQL metadata cache moves from a static field onto
an internal ConcurrentDictionary on CrdtConfig, so it's shared across
repositories/contexts and tied to config lifetime.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
auto_review:
drafts: true
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch reduce-sync-work

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

# Conflicts:
#	src/SIL.Harmony/Config/HarmonyConfig.cs
#	src/SIL.Harmony/SnapshotWorker.cs
Notify DI interceptors and HarmonyConfig.OnProjectedEntitiesChanged after projected SQL with the latest upsert or delete per entity.
Keep the slnx migration from main and include SIL.Harmony.Benchmarks in the solution.
Main now uses Microsoft.Testing.Platform, so solution-wide dotnet test was launching the Benchmarks exe and failing on unknown MTP flags.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@hahn-kev-bot@hahn-kev@claude
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Reduce sync work: faster snapshot projection + benchmarks - #86

Draft
hahn-kev-bot wants to merge 16 commits into
mainfrom
reduce-sync-work
Draft

Reduce sync work: faster snapshot projection + benchmarks#86
hahn-kev-bot wants to merge 16 commits into
mainfrom
reduce-sync-work

Conversation

@hahn-kev-bot

Copy link
Copy Markdown
Collaborator

Overview

Reduces the work done during sync (AddRangeFromSyncSnapshotWorker.UpdateSnapshotsCrdtRepository.AddSnapshots) and adds a benchmark suite to measure it. On the CreateWords workload at 10k changes this branch takes sync from ~1.96 s / 976 MB allocated down to ~0.99 s / 538 MB — roughly 2× faster and ~45% less memory.

What changed

Snapshot pre-load (SnapshotWorker / DataModel)

  • UpdateSnapshots now bulk-loads the relevant current snapshots (with their Commit) into a cache keyed by entity id, and SnapshotWorker reads full snapshots straight from that cache instead of issuing a FindSnapshot DB round-trip per cache hit.

Fast raw-SQL projection (FastProjection, new)

  • Snapshot rows are still inserted through EF unchanged, but the projected tables are now populated with hand-written raw SQL INSERT ... ON CONFLICT(pk) DO UPDATE (one upsert per entity row) instead of going through EF's change tracker (FindAsync / SetValues / graph tracking).
  • Everything the SQL needs — table/column names (correctly delimited), primary key, the SnapshotId shadow FK, value converters — is derived from the EF model, so there is no per-entity code.
  • It dedups to the latest snapshot per entity, runs deletes before upserts (children-first) then upserts (parents-first) for FK/unique-constraint safety, and reuses the caller's transaction.
  • FastProjection is an injectable singleton; its per-type SQL metadata cache lives on an internal ConcurrentDictionary on CrdtConfig, shared across repositories/contexts. This replaces the previous EF change-tracker projection path in CrdtRepository.AddSnapshots, which is removed.

Benchmarks (new SIL.Harmony.Benchmarks project)

  • BenchmarkDotNet suite with a DataModelSyncBenchmarks (7 sync workloads) and an AddSnapshotsBenchmarks that isolates the persist step, [MemoryDiagnoser] enabled. Run with dotnet run -c Release --project src/SIL.Harmony.Benchmarks.

Testing

  • Full test suite passes: 241 passing. The only failures are the 6 DataModelPerformanceBenchmarks timing-threshold tests, which also fail on main (environmental, not caused by this change).

Notes for reviewers

  • The projection is SQLite-specific in a couple of spots that matter for FK correctness (GUIDs stored as uppercase TEXT; ON CONFLICT after a SELECT needs the SQLite WHERE true disambiguator in the code history — the current per-query path uses VALUES). If other providers are ever targeted, FastProjection would need revisiting.
  • Known limitation: an intra-batch self-reference (e.g. Word.AntonymId pointing at another new Word in the same batch) is not ordered; it's a nullable SET NULL FK and not exercised by current workloads.

hahn-kevand others added 11 commits July 21, 2026 10:42
Previously AddSnapshots projected each snapshot by calling FindAsync per
entity, which issued one database query per snapshot (and, on an initial
sync of new data, every query returned null after a round-trip).
Pre-load the projected rows that already exist for the batch with a single
tracked query per object type. ProjectSnapshot then resolves existing
entities from the change tracker and skips the lookup entirely for entities
that have no projected row yet, collapsing N queries down to roughly one
per distinct object type.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019zEmz7jPRPF6Lv8h6YAWBW
Adds a BenchmarkDotNet suite that measures CrdtRepository.AddSnapshots on
its own, across 7 workloads mirroring DataModelSyncBenchmarks. Expensive DB
seeding runs once in a template DB; each iteration forks the DB and recomputes
the snapshot batch so no EF-tracked state leaks across iterations.
- SnapshotWorker.ComputeSnapshotsToPersist: returns the exact snapshot list
UpdateSnapshots would persist, without writing it.
- DataModelTestBase: internal CreateRepository() and CrdtConfig accessors.
- BenchmarkWorkloadBuilders: shared commit builders extracted from
DataModelSyncBenchmarks (+ BuildUpdateExisting).
- Program.cs: run both suites via BenchmarkSwitcher (handles --filter/args).
- Remove leftover Console.WriteLine debug lines from AddSnapshots.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two experimental fast AddSnapshots implementations that keep the EF snapshot
insert unchanged but populate projected tables with raw INSERT ... ON CONFLICT
upserts instead of going through EF's change tracker:
- FAST: one upsert command per entity row
- FAST_JSON: one command per entity type, rows passed as a single JSON array
expanded with SQLite json_each/json_extract
FastProjection derives table/column names, primary key, the SnapshotId shadow
FK, and value converters from the EF model (no per-entity code). It dedups to
the latest snapshot per entity, runs deletes before upserts (children-first)
then upserts (parents-first) for FK/unique-constraint safety, and reuses the
caller's transaction.
CrdtRepository.AddSnapshots now selects via #if FAST_JSON / #elif FAST / #else.
Program.cs adds a third FAST_JSON benchmark job and DataModelSyncBenchmarks
enables [MemoryDiagnoser].
Benchmarks (CreateWords, 1000): both fast paths ~35% faster and ~32% fewer
allocations than baseline; per-query vs JSON-batch shows no measurable
difference against in-memory SQLite.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The JSON-batch projection benchmarked identically to the per-query path against
in-memory SQLite (same time and allocations), so remove it and keep only the
per-query raw-SQL upsert path. FastProjection loses the useJsonBatch parameter
and all json_each/json_extract code; CrdtRepository.AddSnapshots collapses to
#if FAST / #else; the benchmark drops the FAST_JSON job.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Remove the EF change-tracker slow path and the #if FAST conditional so
AddSnapshots always uses FastProjection. Deletes the now-dead slow-path helpers
(ProjectSnapshot, GetEntityEntry, LoadExistingEntityIds, LoadExistingEntities).
The benchmark collapses to a single job since FAST vs DEFAULT are now identical.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
FastProjection is now an injected singleton (registered in AddCrdtDataCore and
resolved into CrdtRepository via ActivatorUtilities) instead of a static class.
Its per-type projected-table SQL metadata cache moves from a static field onto
an internal ConcurrentDictionary on CrdtConfig, so it's shared across
repositories/contexts and tied to config lifetime.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
auto_review:
drafts: true
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch reduce-sync-work

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

# Conflicts:
#	src/SIL.Harmony/Config/HarmonyConfig.cs
#	src/SIL.Harmony/SnapshotWorker.cs
Notify DI interceptors and HarmonyConfig.OnProjectedEntitiesChanged after projected SQL with the latest upsert or delete per entity.
Keep the slnx migration from main and include SIL.Harmony.Benchmarks in the solution.
Main now uses Microsoft.Testing.Platform, so solution-wide dotnet test was launching the Benchmarks exe and failing on unknown MTP flags.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@hahn-kev-bot@hahn-kev@claude
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Reduce sync work: faster snapshot projection + benchmarks - #86

Draft
hahn-kev-bot wants to merge 16 commits into
mainfrom
reduce-sync-work
Draft

Reduce sync work: faster snapshot projection + benchmarks#86
hahn-kev-bot wants to merge 16 commits into
mainfrom
reduce-sync-work

Conversation

@hahn-kev-bot

Copy link
Copy Markdown
Collaborator

Overview

Reduces the work done during sync (AddRangeFromSyncSnapshotWorker.UpdateSnapshotsCrdtRepository.AddSnapshots) and adds a benchmark suite to measure it. On the CreateWords workload at 10k changes this branch takes sync from ~1.96 s / 976 MB allocated down to ~0.99 s / 538 MB — roughly 2× faster and ~45% less memory.

What changed

Snapshot pre-load (SnapshotWorker / DataModel)

  • UpdateSnapshots now bulk-loads the relevant current snapshots (with their Commit) into a cache keyed by entity id, and SnapshotWorker reads full snapshots straight from that cache instead of issuing a FindSnapshot DB round-trip per cache hit.

Fast raw-SQL projection (FastProjection, new)

  • Snapshot rows are still inserted through EF unchanged, but the projected tables are now populated with hand-written raw SQL INSERT ... ON CONFLICT(pk) DO UPDATE (one upsert per entity row) instead of going through EF's change tracker (FindAsync / SetValues / graph tracking).
  • Everything the SQL needs — table/column names (correctly delimited), primary key, the SnapshotId shadow FK, value converters — is derived from the EF model, so there is no per-entity code.
  • It dedups to the latest snapshot per entity, runs deletes before upserts (children-first) then upserts (parents-first) for FK/unique-constraint safety, and reuses the caller's transaction.
  • FastProjection is an injectable singleton; its per-type SQL metadata cache lives on an internal ConcurrentDictionary on CrdtConfig, shared across repositories/contexts. This replaces the previous EF change-tracker projection path in CrdtRepository.AddSnapshots, which is removed.

Benchmarks (new SIL.Harmony.Benchmarks project)

  • BenchmarkDotNet suite with a DataModelSyncBenchmarks (7 sync workloads) and an AddSnapshotsBenchmarks that isolates the persist step, [MemoryDiagnoser] enabled. Run with dotnet run -c Release --project src/SIL.Harmony.Benchmarks.

Testing

  • Full test suite passes: 241 passing. The only failures are the 6 DataModelPerformanceBenchmarks timing-threshold tests, which also fail on main (environmental, not caused by this change).

Notes for reviewers

  • The projection is SQLite-specific in a couple of spots that matter for FK correctness (GUIDs stored as uppercase TEXT; ON CONFLICT after a SELECT needs the SQLite WHERE true disambiguator in the code history — the current per-query path uses VALUES). If other providers are ever targeted, FastProjection would need revisiting.
  • Known limitation: an intra-batch self-reference (e.g. Word.AntonymId pointing at another new Word in the same batch) is not ordered; it's a nullable SET NULL FK and not exercised by current workloads.

hahn-kevand others added 11 commits July 21, 2026 10:42
Previously AddSnapshots projected each snapshot by calling FindAsync per
entity, which issued one database query per snapshot (and, on an initial
sync of new data, every query returned null after a round-trip).
Pre-load the projected rows that already exist for the batch with a single
tracked query per object type. ProjectSnapshot then resolves existing
entities from the change tracker and skips the lookup entirely for entities
that have no projected row yet, collapsing N queries down to roughly one
per distinct object type.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019zEmz7jPRPF6Lv8h6YAWBW
Adds a BenchmarkDotNet suite that measures CrdtRepository.AddSnapshots on
its own, across 7 workloads mirroring DataModelSyncBenchmarks. Expensive DB
seeding runs once in a template DB; each iteration forks the DB and recomputes
the snapshot batch so no EF-tracked state leaks across iterations.
- SnapshotWorker.ComputeSnapshotsToPersist: returns the exact snapshot list
UpdateSnapshots would persist, without writing it.
- DataModelTestBase: internal CreateRepository() and CrdtConfig accessors.
- BenchmarkWorkloadBuilders: shared commit builders extracted from
DataModelSyncBenchmarks (+ BuildUpdateExisting).
- Program.cs: run both suites via BenchmarkSwitcher (handles --filter/args).
- Remove leftover Console.WriteLine debug lines from AddSnapshots.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two experimental fast AddSnapshots implementations that keep the EF snapshot
insert unchanged but populate projected tables with raw INSERT ... ON CONFLICT
upserts instead of going through EF's change tracker:
- FAST: one upsert command per entity row
- FAST_JSON: one command per entity type, rows passed as a single JSON array
expanded with SQLite json_each/json_extract
FastProjection derives table/column names, primary key, the SnapshotId shadow
FK, and value converters from the EF model (no per-entity code). It dedups to
the latest snapshot per entity, runs deletes before upserts (children-first)
then upserts (parents-first) for FK/unique-constraint safety, and reuses the
caller's transaction.
CrdtRepository.AddSnapshots now selects via #if FAST_JSON / #elif FAST / #else.
Program.cs adds a third FAST_JSON benchmark job and DataModelSyncBenchmarks
enables [MemoryDiagnoser].
Benchmarks (CreateWords, 1000): both fast paths ~35% faster and ~32% fewer
allocations than baseline; per-query vs JSON-batch shows no measurable
difference against in-memory SQLite.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The JSON-batch projection benchmarked identically to the per-query path against
in-memory SQLite (same time and allocations), so remove it and keep only the
per-query raw-SQL upsert path. FastProjection loses the useJsonBatch parameter
and all json_each/json_extract code; CrdtRepository.AddSnapshots collapses to
#if FAST / #else; the benchmark drops the FAST_JSON job.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Remove the EF change-tracker slow path and the #if FAST conditional so
AddSnapshots always uses FastProjection. Deletes the now-dead slow-path helpers
(ProjectSnapshot, GetEntityEntry, LoadExistingEntityIds, LoadExistingEntities).
The benchmark collapses to a single job since FAST vs DEFAULT are now identical.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
FastProjection is now an injected singleton (registered in AddCrdtDataCore and
resolved into CrdtRepository via ActivatorUtilities) instead of a static class.
Its per-type projected-table SQL metadata cache moves from a static field onto
an internal ConcurrentDictionary on CrdtConfig, so it's shared across
repositories/contexts and tied to config lifetime.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
auto_review:
drafts: true
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch reduce-sync-work

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

# Conflicts:
#	src/SIL.Harmony/Config/HarmonyConfig.cs
#	src/SIL.Harmony/SnapshotWorker.cs
Notify DI interceptors and HarmonyConfig.OnProjectedEntitiesChanged after projected SQL with the latest upsert or delete per entity.
Keep the slnx migration from main and include SIL.Harmony.Benchmarks in the solution.
Main now uses Microsoft.Testing.Platform, so solution-wide dotnet test was launching the Benchmarks exe and failing on unknown MTP flags.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@hahn-kev-bot@hahn-kev@claude
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Reduce sync work: faster snapshot projection + benchmarks - #86

Draft
hahn-kev-bot wants to merge 16 commits into
mainfrom
reduce-sync-work
Draft

Reduce sync work: faster snapshot projection + benchmarks#86
hahn-kev-bot wants to merge 16 commits into
mainfrom
reduce-sync-work

Conversation

@hahn-kev-bot

Copy link
Copy Markdown
Collaborator

Overview

Reduces the work done during sync (AddRangeFromSyncSnapshotWorker.UpdateSnapshotsCrdtRepository.AddSnapshots) and adds a benchmark suite to measure it. On the CreateWords workload at 10k changes this branch takes sync from ~1.96 s / 976 MB allocated down to ~0.99 s / 538 MB — roughly 2× faster and ~45% less memory.

What changed

Snapshot pre-load (SnapshotWorker / DataModel)

  • UpdateSnapshots now bulk-loads the relevant current snapshots (with their Commit) into a cache keyed by entity id, and SnapshotWorker reads full snapshots straight from that cache instead of issuing a FindSnapshot DB round-trip per cache hit.

Fast raw-SQL projection (FastProjection, new)

  • Snapshot rows are still inserted through EF unchanged, but the projected tables are now populated with hand-written raw SQL INSERT ... ON CONFLICT(pk) DO UPDATE (one upsert per entity row) instead of going through EF's change tracker (FindAsync / SetValues / graph tracking).
  • Everything the SQL needs — table/column names (correctly delimited), primary key, the SnapshotId shadow FK, value converters — is derived from the EF model, so there is no per-entity code.
  • It dedups to the latest snapshot per entity, runs deletes before upserts (children-first) then upserts (parents-first) for FK/unique-constraint safety, and reuses the caller's transaction.
  • FastProjection is an injectable singleton; its per-type SQL metadata cache lives on an internal ConcurrentDictionary on CrdtConfig, shared across repositories/contexts. This replaces the previous EF change-tracker projection path in CrdtRepository.AddSnapshots, which is removed.

Benchmarks (new SIL.Harmony.Benchmarks project)

  • BenchmarkDotNet suite with a DataModelSyncBenchmarks (7 sync workloads) and an AddSnapshotsBenchmarks that isolates the persist step, [MemoryDiagnoser] enabled. Run with dotnet run -c Release --project src/SIL.Harmony.Benchmarks.

Testing

  • Full test suite passes: 241 passing. The only failures are the 6 DataModelPerformanceBenchmarks timing-threshold tests, which also fail on main (environmental, not caused by this change).

Notes for reviewers

  • The projection is SQLite-specific in a couple of spots that matter for FK correctness (GUIDs stored as uppercase TEXT; ON CONFLICT after a SELECT needs the SQLite WHERE true disambiguator in the code history — the current per-query path uses VALUES). If other providers are ever targeted, FastProjection would need revisiting.
  • Known limitation: an intra-batch self-reference (e.g. Word.AntonymId pointing at another new Word in the same batch) is not ordered; it's a nullable SET NULL FK and not exercised by current workloads.

hahn-kevand others added 11 commits July 21, 2026 10:42
Previously AddSnapshots projected each snapshot by calling FindAsync per
entity, which issued one database query per snapshot (and, on an initial
sync of new data, every query returned null after a round-trip).
Pre-load the projected rows that already exist for the batch with a single
tracked query per object type. ProjectSnapshot then resolves existing
entities from the change tracker and skips the lookup entirely for entities
that have no projected row yet, collapsing N queries down to roughly one
per distinct object type.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019zEmz7jPRPF6Lv8h6YAWBW
Adds a BenchmarkDotNet suite that measures CrdtRepository.AddSnapshots on
its own, across 7 workloads mirroring DataModelSyncBenchmarks. Expensive DB
seeding runs once in a template DB; each iteration forks the DB and recomputes
the snapshot batch so no EF-tracked state leaks across iterations.
- SnapshotWorker.ComputeSnapshotsToPersist: returns the exact snapshot list
UpdateSnapshots would persist, without writing it.
- DataModelTestBase: internal CreateRepository() and CrdtConfig accessors.
- BenchmarkWorkloadBuilders: shared commit builders extracted from
DataModelSyncBenchmarks (+ BuildUpdateExisting).
- Program.cs: run both suites via BenchmarkSwitcher (handles --filter/args).
- Remove leftover Console.WriteLine debug lines from AddSnapshots.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two experimental fast AddSnapshots implementations that keep the EF snapshot
insert unchanged but populate projected tables with raw INSERT ... ON CONFLICT
upserts instead of going through EF's change tracker:
- FAST: one upsert command per entity row
- FAST_JSON: one command per entity type, rows passed as a single JSON array
expanded with SQLite json_each/json_extract
FastProjection derives table/column names, primary key, the SnapshotId shadow
FK, and value converters from the EF model (no per-entity code). It dedups to
the latest snapshot per entity, runs deletes before upserts (children-first)
then upserts (parents-first) for FK/unique-constraint safety, and reuses the
caller's transaction.
CrdtRepository.AddSnapshots now selects via #if FAST_JSON / #elif FAST / #else.
Program.cs adds a third FAST_JSON benchmark job and DataModelSyncBenchmarks
enables [MemoryDiagnoser].
Benchmarks (CreateWords, 1000): both fast paths ~35% faster and ~32% fewer
allocations than baseline; per-query vs JSON-batch shows no measurable
difference against in-memory SQLite.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The JSON-batch projection benchmarked identically to the per-query path against
in-memory SQLite (same time and allocations), so remove it and keep only the
per-query raw-SQL upsert path. FastProjection loses the useJsonBatch parameter
and all json_each/json_extract code; CrdtRepository.AddSnapshots collapses to
#if FAST / #else; the benchmark drops the FAST_JSON job.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Remove the EF change-tracker slow path and the #if FAST conditional so
AddSnapshots always uses FastProjection. Deletes the now-dead slow-path helpers
(ProjectSnapshot, GetEntityEntry, LoadExistingEntityIds, LoadExistingEntities).
The benchmark collapses to a single job since FAST vs DEFAULT are now identical.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
FastProjection is now an injected singleton (registered in AddCrdtDataCore and
resolved into CrdtRepository via ActivatorUtilities) instead of a static class.
Its per-type projected-table SQL metadata cache moves from a static field onto
an internal ConcurrentDictionary on CrdtConfig, so it's shared across
repositories/contexts and tied to config lifetime.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
auto_review:
drafts: true
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch reduce-sync-work

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

# Conflicts:
#	src/SIL.Harmony/Config/HarmonyConfig.cs
#	src/SIL.Harmony/SnapshotWorker.cs
Notify DI interceptors and HarmonyConfig.OnProjectedEntitiesChanged after projected SQL with the latest upsert or delete per entity.
Keep the slnx migration from main and include SIL.Harmony.Benchmarks in the solution.
Main now uses Microsoft.Testing.Platform, so solution-wide dotnet test was launching the Benchmarks exe and failing on unknown MTP flags.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@hahn-kev-bot@hahn-kev@claude
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Reduce sync work: faster snapshot projection + benchmarks - #86

Draft
hahn-kev-bot wants to merge 16 commits into
mainfrom
reduce-sync-work
Draft

Reduce sync work: faster snapshot projection + benchmarks#86
hahn-kev-bot wants to merge 16 commits into
mainfrom
reduce-sync-work

Conversation

@hahn-kev-bot

Copy link
Copy Markdown
Collaborator

Overview

Reduces the work done during sync (AddRangeFromSyncSnapshotWorker.UpdateSnapshotsCrdtRepository.AddSnapshots) and adds a benchmark suite to measure it. On the CreateWords workload at 10k changes this branch takes sync from ~1.96 s / 976 MB allocated down to ~0.99 s / 538 MB — roughly 2× faster and ~45% less memory.

What changed

Snapshot pre-load (SnapshotWorker / DataModel)

  • UpdateSnapshots now bulk-loads the relevant current snapshots (with their Commit) into a cache keyed by entity id, and SnapshotWorker reads full snapshots straight from that cache instead of issuing a FindSnapshot DB round-trip per cache hit.

Fast raw-SQL projection (FastProjection, new)

  • Snapshot rows are still inserted through EF unchanged, but the projected tables are now populated with hand-written raw SQL INSERT ... ON CONFLICT(pk) DO UPDATE (one upsert per entity row) instead of going through EF's change tracker (FindAsync / SetValues / graph tracking).
  • Everything the SQL needs — table/column names (correctly delimited), primary key, the SnapshotId shadow FK, value converters — is derived from the EF model, so there is no per-entity code.
  • It dedups to the latest snapshot per entity, runs deletes before upserts (children-first) then upserts (parents-first) for FK/unique-constraint safety, and reuses the caller's transaction.
  • FastProjection is an injectable singleton; its per-type SQL metadata cache lives on an internal ConcurrentDictionary on CrdtConfig, shared across repositories/contexts. This replaces the previous EF change-tracker projection path in CrdtRepository.AddSnapshots, which is removed.

Benchmarks (new SIL.Harmony.Benchmarks project)

  • BenchmarkDotNet suite with a DataModelSyncBenchmarks (7 sync workloads) and an AddSnapshotsBenchmarks that isolates the persist step, [MemoryDiagnoser] enabled. Run with dotnet run -c Release --project src/SIL.Harmony.Benchmarks.

Testing

  • Full test suite passes: 241 passing. The only failures are the 6 DataModelPerformanceBenchmarks timing-threshold tests, which also fail on main (environmental, not caused by this change).

Notes for reviewers

  • The projection is SQLite-specific in a couple of spots that matter for FK correctness (GUIDs stored as uppercase TEXT; ON CONFLICT after a SELECT needs the SQLite WHERE true disambiguator in the code history — the current per-query path uses VALUES). If other providers are ever targeted, FastProjection would need revisiting.
  • Known limitation: an intra-batch self-reference (e.g. Word.AntonymId pointing at another new Word in the same batch) is not ordered; it's a nullable SET NULL FK and not exercised by current workloads.

hahn-kevand others added 11 commits July 21, 2026 10:42
Previously AddSnapshots projected each snapshot by calling FindAsync per
entity, which issued one database query per snapshot (and, on an initial
sync of new data, every query returned null after a round-trip).
Pre-load the projected rows that already exist for the batch with a single
tracked query per object type. ProjectSnapshot then resolves existing
entities from the change tracker and skips the lookup entirely for entities
that have no projected row yet, collapsing N queries down to roughly one
per distinct object type.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019zEmz7jPRPF6Lv8h6YAWBW
Adds a BenchmarkDotNet suite that measures CrdtRepository.AddSnapshots on
its own, across 7 workloads mirroring DataModelSyncBenchmarks. Expensive DB
seeding runs once in a template DB; each iteration forks the DB and recomputes
the snapshot batch so no EF-tracked state leaks across iterations.
- SnapshotWorker.ComputeSnapshotsToPersist: returns the exact snapshot list
UpdateSnapshots would persist, without writing it.
- DataModelTestBase: internal CreateRepository() and CrdtConfig accessors.
- BenchmarkWorkloadBuilders: shared commit builders extracted from
DataModelSyncBenchmarks (+ BuildUpdateExisting).
- Program.cs: run both suites via BenchmarkSwitcher (handles --filter/args).
- Remove leftover Console.WriteLine debug lines from AddSnapshots.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two experimental fast AddSnapshots implementations that keep the EF snapshot
insert unchanged but populate projected tables with raw INSERT ... ON CONFLICT
upserts instead of going through EF's change tracker:
- FAST: one upsert command per entity row
- FAST_JSON: one command per entity type, rows passed as a single JSON array
expanded with SQLite json_each/json_extract
FastProjection derives table/column names, primary key, the SnapshotId shadow
FK, and value converters from the EF model (no per-entity code). It dedups to
the latest snapshot per entity, runs deletes before upserts (children-first)
then upserts (parents-first) for FK/unique-constraint safety, and reuses the
caller's transaction.
CrdtRepository.AddSnapshots now selects via #if FAST_JSON / #elif FAST / #else.
Program.cs adds a third FAST_JSON benchmark job and DataModelSyncBenchmarks
enables [MemoryDiagnoser].
Benchmarks (CreateWords, 1000): both fast paths ~35% faster and ~32% fewer
allocations than baseline; per-query vs JSON-batch shows no measurable
difference against in-memory SQLite.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The JSON-batch projection benchmarked identically to the per-query path against
in-memory SQLite (same time and allocations), so remove it and keep only the
per-query raw-SQL upsert path. FastProjection loses the useJsonBatch parameter
and all json_each/json_extract code; CrdtRepository.AddSnapshots collapses to
#if FAST / #else; the benchmark drops the FAST_JSON job.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Remove the EF change-tracker slow path and the #if FAST conditional so
AddSnapshots always uses FastProjection. Deletes the now-dead slow-path helpers
(ProjectSnapshot, GetEntityEntry, LoadExistingEntityIds, LoadExistingEntities).
The benchmark collapses to a single job since FAST vs DEFAULT are now identical.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
FastProjection is now an injected singleton (registered in AddCrdtDataCore and
resolved into CrdtRepository via ActivatorUtilities) instead of a static class.
Its per-type projected-table SQL metadata cache moves from a static field onto
an internal ConcurrentDictionary on CrdtConfig, so it's shared across
repositories/contexts and tied to config lifetime.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@coderabbitai

coderabbitaiBot commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
auto_review:
drafts: true
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch reduce-sync-work

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

# Conflicts:
#	src/SIL.Harmony/Config/HarmonyConfig.cs
#	src/SIL.Harmony/SnapshotWorker.cs
Notify DI interceptors and HarmonyConfig.OnProjectedEntitiesChanged after projected SQL with the latest upsert or delete per entity.
Keep the slnx migration from main and include SIL.Harmony.Benchmarks in the solution.
Main now uses Microsoft.Testing.Platform, so solution-wide dotnet test was launching the Benchmarks exe and failing on unknown MTP flags.
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@hahn-kev-bot@hahn-kev@claude