Restore lazy flag to KVStore::remove - #4189

Merged
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
TheBlueMatt:2025-10-lazy-again
Oct 30, 2025
Merged

Restore lazy flag to KVStore::remove#4189
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
TheBlueMatt:2025-10-lazy-again

Conversation

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

A user pointed out, when looking to upgrade to LDK 0.2, that the
lazy flag is actually quite important for performance when using
a MonitorUpdatingPersister, especially in synchronous persistence
mode.

Thus, we add it back here.

@TheBlueMattTheBlueMatt added this to the 0.2 milestone Oct 30, 2025
@ldk-reviews-bot

ldk-reviews-bot commented Oct 30, 2025

Copy link
Copy Markdown

👋 Thanks for assigning @joostjager as a reviewer!
I'll wait for their review and will help manage the review process.
Once they submit their review, I'll check if a second reviewer would be helpful.

@TheBlueMattTheBlueMatt linked an issue Oct 30, 2025 that may be closed by this pull request
@codecov

codecovBot commented Oct 30, 2025

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 56.25000% with 21 lines in your changes missing coverage. Please review.
✅ Project coverage is 88.84%. Comparing base (02a9af9) to head (0f9548b).
⚠️ Report is 8 commits behind head on main.

Files with missing linesPatch %Lines
lightning-persister/src/fs_store.rs52.38%5 Missing and 5 partials ⚠️
lightning/src/util/persist.rs68.75%3 Missing and 2 partials ⚠️
lightning-background-processor/src/lib.rs0.00%2 Missing ⚠️
lightning/src/util/test_utils.rs60.00%2 Missing ⚠️
lightning-liquidity/src/lsps2/service.rs0.00%1 Missing ⚠️
lightning-liquidity/src/lsps5/service.rs0.00%1 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #4189 +/- ##
==========================================
- Coverage 88.87% 88.84% -0.04% 
==========================================
Files 180 180 Lines 137863 137870 +7 Branches 137863 137870 +7 ==========================================
- Hits 122522 122485 -37 - Misses 12532 12573 +41 - Partials 2809 2812 +3 
FlagCoverage Δ
fuzzing21.44% <0.00%> (+0.58%)⬆️
tests88.68% <56.25%> (-0.04%)⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

/// potentially get lost on crash after the method returns. Therefore, this flag should only be
/// set for `remove` operations that can be safely replayed at a later time.
///
/// All removal operations must complete in a consistent total order with [`Self::write`]s

@tnulltnullOct 30, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not sure if this would even work. For example in FilesystemStore, if we simply call remove and leave the decision on when to sync the changes to disk to the OS, how could we be certain that the ordering is preserved? IIRC we basically concluded this can't be guaranteed, especially since different guarantees on different platforms might vary?

@tnulltnullOct 30, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that man 2 unlink states:

If the name was the last link to a file but any processes still have the file open, the file will remain in existence until the last file descriptor referring to it is closed.

That means that if we have a concurrent read, we may defer the actual deletion, allowing it to interact with following writes, e.g.:

| t1 | t2 | t3 |
| READ | unlink | |
| READ | ... | WRITE |
| READ | ... | |
| READ | ... | SYNC |
| READ | SYNC | |

Although, given we use rename for write, I do wonder if the unlink would simply get lost here as it would apply to the original file that is dropped already anyways?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To avoid this, maybe it is possible to constrain lazy removes to keys that won't ever be written again in the future?

I've read the context of this PR now, and it seems the perf issue was around monitor updates. I think those are never re-written?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not sure if this would even work. For example in FilesystemStore, if we simply call remove and leave the decision on when to sync the changes to disk to the OS, how could we be certain that the ordering is preserved?

Filesystems provide an order, the only thing they dont provide without an fsync is any kinds of guarantee its on disk. I don't think this is a problem.

Although, given we use rename for write, I do wonder if the unlink would simply get lost here as it would apply to the original file that is dropped already anyways?

Yes, that is how it should work on any reasonable filesystem. In theory its possible for some filesystems to fail the write rename part because the file still exists, but that's unrelated to the remove, that's just the read+write being at the same time.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Filesystems provide an order, the only thing they dont provide without an fsync is any kinds of guarantee its on disk. I don't think this is a problem.

A file that should have been removed, but is still there. Is that not a pb?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's the explicit point of the lazy flag - it allows a store to not guarantee that the entry will be removed if there's a crash/ill-timed restart.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Discussed this offline, man 3p unlink had me convinced this is safe to do on POSIX.

@ldk-reviews-bot

Copy link
Copy Markdown

👋 The first review has been submitted!

Do you think this PR is ready for a second reviewer? If so, click here to assign a second reviewer.

@tnull
tnull requested a review from joostjagerOctober 30, 2025 09:42

@joostjagerjoostjager left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are there more details on why/when the lazy flag is important?

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

In this case its important for perf when removing a large number of monitor updates on an fsstore (requiring fsync for each can add up rather substantially eg if we're removing 1k mon updates), but thinking about it more I think its also important for the same case in the async design - if you have a KVStore that handles ordering (eg like the locks in the fsstore/vss store) then the lazy flag allows you to spawn-and-forget a removal, rather than the callsite having to "block" waiting on your removal to finish.

@joostjager

Copy link
Copy Markdown
Contributor

1k updates is a lot. Do you think it adds much over let's say 50 updates? Maybe that also makes the perf problem go away without lazy flag...

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

1k updates seems entirely reasonable for a node doing lots of forwarding. ChannelMonitors can easily be a few thousand times larger than ChannelMonitorUpdates, so wanting to amortize over more ChannelMonitorUpdates seems very reasonable (the startup cost of more ChannelMonitorUpdates is pretty low, or at least is if your KVStore read latency is low or once we load them in parallel).

Not being able to pick a reasonable update count just because of an API limitation in how we do removals seems like a pretty weird limitation, no?

@joostjager

Copy link
Copy Markdown
Contributor

Not being able to pick a reasonable update count just because of an API limitation in how we do removals seems like a pretty weird limitation, no?

The question was just whether 1000 is reasonable, and you made it clear that it is 👍

joostjager
joostjager previously approved these changes Oct 30, 2025
This reverts commit 561da4c.
A user pointed out, when looking to upgrade to LDK 0.2, that the
`lazy` flag is actually quite important for performance when using
a `MonitorUpdatingPersister`, especially in synchronous persistence
mode.
Thus, we add it back here.
Fixeslightningdevkit#4188
In the previous commit we reverted
561da4c. One of the motivations
for it (in addition to `lazy` removals being somewhat less, though
still arguably useful in an async context) was that the ordering
requirements of `lazy` removals is somewhat unclear.
Here we simply default to the simplest safe option, requiring a
total order across all `write` and `remove` operations to the same
key, `lazy` or not.
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Fixed rustfmt

$ git diff-tree -U1 9973d780a 0f9548bf8
diff --git a/lightning/src/util/persist.rs b/lightning/src/util/persist.rs
index 78fdba2113..5d34603c96 100644
--- a/lightning/src/util/persist.rs+++ b/lightning/src/util/persist.rs@@ -1084,4 +1084,3 @@ where
let latest_update_id = current_monitor.get_latest_update_id();
- self- .cleanup_stale_updates_for_monitor_to(&monitor_key, latest_update_id, lazy)+ self.cleanup_stale_updates_for_monitor_to(&monitor_key, latest_update_id, lazy)
.await?;

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Only a rustfmt change since @joostjager ack'd, so landing.

@TheBlueMatt
TheBlueMatt merged commit d53d6b4 into lightningdevkit:mainOct 30, 2025
23 of 25 checks passed
@TheBlueMattTheBlueMatt mentioned this pull request Oct 30, 2025
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Backported to 0.2 in #4193

@domZippilli

Copy link
Copy Markdown
Contributor

🥳

@wvanlint

Copy link
Copy Markdown
Contributor

Thanks for landing this!

I think the comments above covered everything. We use the MonitorUpdatingPersister with maximum_pending_updates = 1000 for efficiency due to the high forwarding volume, and an update_persisted_channel call can trigger channel monitor update consolidation when maximum_pending_updates is reached. This consolidation results in maximum_pending_updates sequential KVStore::remove calls, which caused issues when it's performed in a non-lazy fashion. In our case in 0.1, it blocked the Tokio runtime (for ~7 ms * 1000 = 7s), but I assume it will affect the caller in the async design as well as Matt mentioned.

I was curious if there are possible simplifications, such as all remove calls being considered lazy or if remove can be constrained to keys that won't ever be written again in the future as Joost mentioned. But I see there are requirements coming from #4059 (comment) as well.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add lazy persistence back to KVStore::delete

6 participants

@TheBlueMatt@ldk-reviews-bot@joostjager@domZippilli@wvanlint@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Restore lazy flag to KVStore::remove - #4189

Merged
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
TheBlueMatt:2025-10-lazy-again
Oct 30, 2025
Merged

Restore lazy flag to KVStore::remove#4189
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
TheBlueMatt:2025-10-lazy-again

Conversation

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

A user pointed out, when looking to upgrade to LDK 0.2, that the
lazy flag is actually quite important for performance when using
a MonitorUpdatingPersister, especially in synchronous persistence
mode.

Thus, we add it back here.

@TheBlueMattTheBlueMatt added this to the 0.2 milestone Oct 30, 2025
@ldk-reviews-bot

ldk-reviews-bot commented Oct 30, 2025

Copy link
Copy Markdown

👋 Thanks for assigning @joostjager as a reviewer!
I'll wait for their review and will help manage the review process.
Once they submit their review, I'll check if a second reviewer would be helpful.

@TheBlueMattTheBlueMatt linked an issue Oct 30, 2025 that may be closed by this pull request
@codecov

codecovBot commented Oct 30, 2025

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 56.25000% with 21 lines in your changes missing coverage. Please review.
✅ Project coverage is 88.84%. Comparing base (02a9af9) to head (0f9548b).
⚠️ Report is 8 commits behind head on main.

Files with missing linesPatch %Lines
lightning-persister/src/fs_store.rs52.38%5 Missing and 5 partials ⚠️
lightning/src/util/persist.rs68.75%3 Missing and 2 partials ⚠️
lightning-background-processor/src/lib.rs0.00%2 Missing ⚠️
lightning/src/util/test_utils.rs60.00%2 Missing ⚠️
lightning-liquidity/src/lsps2/service.rs0.00%1 Missing ⚠️
lightning-liquidity/src/lsps5/service.rs0.00%1 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #4189 +/- ##
==========================================
- Coverage 88.87% 88.84% -0.04% 
==========================================
Files 180 180 Lines 137863 137870 +7 Branches 137863 137870 +7 ==========================================
- Hits 122522 122485 -37 - Misses 12532 12573 +41 - Partials 2809 2812 +3 
FlagCoverage Δ
fuzzing21.44% <0.00%> (+0.58%)⬆️
tests88.68% <56.25%> (-0.04%)⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

/// potentially get lost on crash after the method returns. Therefore, this flag should only be
/// set for `remove` operations that can be safely replayed at a later time.
///
/// All removal operations must complete in a consistent total order with [`Self::write`]s

@tnulltnullOct 30, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not sure if this would even work. For example in FilesystemStore, if we simply call remove and leave the decision on when to sync the changes to disk to the OS, how could we be certain that the ordering is preserved? IIRC we basically concluded this can't be guaranteed, especially since different guarantees on different platforms might vary?

@tnulltnullOct 30, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that man 2 unlink states:

If the name was the last link to a file but any processes still have the file open, the file will remain in existence until the last file descriptor referring to it is closed.

That means that if we have a concurrent read, we may defer the actual deletion, allowing it to interact with following writes, e.g.:

| t1 | t2 | t3 |
| READ | unlink | |
| READ | ... | WRITE |
| READ | ... | |
| READ | ... | SYNC |
| READ | SYNC | |

Although, given we use rename for write, I do wonder if the unlink would simply get lost here as it would apply to the original file that is dropped already anyways?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To avoid this, maybe it is possible to constrain lazy removes to keys that won't ever be written again in the future?

I've read the context of this PR now, and it seems the perf issue was around monitor updates. I think those are never re-written?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not sure if this would even work. For example in FilesystemStore, if we simply call remove and leave the decision on when to sync the changes to disk to the OS, how could we be certain that the ordering is preserved?

Filesystems provide an order, the only thing they dont provide without an fsync is any kinds of guarantee its on disk. I don't think this is a problem.

Although, given we use rename for write, I do wonder if the unlink would simply get lost here as it would apply to the original file that is dropped already anyways?

Yes, that is how it should work on any reasonable filesystem. In theory its possible for some filesystems to fail the write rename part because the file still exists, but that's unrelated to the remove, that's just the read+write being at the same time.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Filesystems provide an order, the only thing they dont provide without an fsync is any kinds of guarantee its on disk. I don't think this is a problem.

A file that should have been removed, but is still there. Is that not a pb?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's the explicit point of the lazy flag - it allows a store to not guarantee that the entry will be removed if there's a crash/ill-timed restart.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Discussed this offline, man 3p unlink had me convinced this is safe to do on POSIX.

@ldk-reviews-bot

Copy link
Copy Markdown

👋 The first review has been submitted!

Do you think this PR is ready for a second reviewer? If so, click here to assign a second reviewer.

@tnull
tnull requested a review from joostjagerOctober 30, 2025 09:42

@joostjagerjoostjager left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are there more details on why/when the lazy flag is important?

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

In this case its important for perf when removing a large number of monitor updates on an fsstore (requiring fsync for each can add up rather substantially eg if we're removing 1k mon updates), but thinking about it more I think its also important for the same case in the async design - if you have a KVStore that handles ordering (eg like the locks in the fsstore/vss store) then the lazy flag allows you to spawn-and-forget a removal, rather than the callsite having to "block" waiting on your removal to finish.

@joostjager

Copy link
Copy Markdown
Contributor

1k updates is a lot. Do you think it adds much over let's say 50 updates? Maybe that also makes the perf problem go away without lazy flag...

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

1k updates seems entirely reasonable for a node doing lots of forwarding. ChannelMonitors can easily be a few thousand times larger than ChannelMonitorUpdates, so wanting to amortize over more ChannelMonitorUpdates seems very reasonable (the startup cost of more ChannelMonitorUpdates is pretty low, or at least is if your KVStore read latency is low or once we load them in parallel).

Not being able to pick a reasonable update count just because of an API limitation in how we do removals seems like a pretty weird limitation, no?

@joostjager

Copy link
Copy Markdown
Contributor

Not being able to pick a reasonable update count just because of an API limitation in how we do removals seems like a pretty weird limitation, no?

The question was just whether 1000 is reasonable, and you made it clear that it is 👍

joostjager
joostjager previously approved these changes Oct 30, 2025
This reverts commit 561da4c.
A user pointed out, when looking to upgrade to LDK 0.2, that the
`lazy` flag is actually quite important for performance when using
a `MonitorUpdatingPersister`, especially in synchronous persistence
mode.
Thus, we add it back here.
Fixeslightningdevkit#4188
In the previous commit we reverted
561da4c. One of the motivations
for it (in addition to `lazy` removals being somewhat less, though
still arguably useful in an async context) was that the ordering
requirements of `lazy` removals is somewhat unclear.
Here we simply default to the simplest safe option, requiring a
total order across all `write` and `remove` operations to the same
key, `lazy` or not.
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Fixed rustfmt

$ git diff-tree -U1 9973d780a 0f9548bf8
diff --git a/lightning/src/util/persist.rs b/lightning/src/util/persist.rs
index 78fdba2113..5d34603c96 100644
--- a/lightning/src/util/persist.rs+++ b/lightning/src/util/persist.rs@@ -1084,4 +1084,3 @@ where
let latest_update_id = current_monitor.get_latest_update_id();
- self- .cleanup_stale_updates_for_monitor_to(&monitor_key, latest_update_id, lazy)+ self.cleanup_stale_updates_for_monitor_to(&monitor_key, latest_update_id, lazy)
.await?;

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Only a rustfmt change since @joostjager ack'd, so landing.

@TheBlueMatt
TheBlueMatt merged commit d53d6b4 into lightningdevkit:mainOct 30, 2025
23 of 25 checks passed
@TheBlueMattTheBlueMatt mentioned this pull request Oct 30, 2025
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Backported to 0.2 in #4193

@domZippilli

Copy link
Copy Markdown
Contributor

🥳

@wvanlint

Copy link
Copy Markdown
Contributor

Thanks for landing this!

I think the comments above covered everything. We use the MonitorUpdatingPersister with maximum_pending_updates = 1000 for efficiency due to the high forwarding volume, and an update_persisted_channel call can trigger channel monitor update consolidation when maximum_pending_updates is reached. This consolidation results in maximum_pending_updates sequential KVStore::remove calls, which caused issues when it's performed in a non-lazy fashion. In our case in 0.1, it blocked the Tokio runtime (for ~7 ms * 1000 = 7s), but I assume it will affect the caller in the async design as well as Matt mentioned.

I was curious if there are possible simplifications, such as all remove calls being considered lazy or if remove can be constrained to keys that won't ever be written again in the future as Joost mentioned. But I see there are requirements coming from #4059 (comment) as well.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add lazy persistence back to KVStore::delete

6 participants

@TheBlueMatt@ldk-reviews-bot@joostjager@domZippilli@wvanlint@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Restore lazy flag to KVStore::remove - #4189

Merged
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
TheBlueMatt:2025-10-lazy-again
Oct 30, 2025
Merged

Restore lazy flag to KVStore::remove#4189
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
TheBlueMatt:2025-10-lazy-again

Conversation

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

A user pointed out, when looking to upgrade to LDK 0.2, that the
lazy flag is actually quite important for performance when using
a MonitorUpdatingPersister, especially in synchronous persistence
mode.

Thus, we add it back here.

@TheBlueMattTheBlueMatt added this to the 0.2 milestone Oct 30, 2025
@ldk-reviews-bot

ldk-reviews-bot commented Oct 30, 2025

Copy link
Copy Markdown

👋 Thanks for assigning @joostjager as a reviewer!
I'll wait for their review and will help manage the review process.
Once they submit their review, I'll check if a second reviewer would be helpful.

@TheBlueMattTheBlueMatt linked an issue Oct 30, 2025 that may be closed by this pull request
@codecov

codecovBot commented Oct 30, 2025

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 56.25000% with 21 lines in your changes missing coverage. Please review.
✅ Project coverage is 88.84%. Comparing base (02a9af9) to head (0f9548b).
⚠️ Report is 8 commits behind head on main.

Files with missing linesPatch %Lines
lightning-persister/src/fs_store.rs52.38%5 Missing and 5 partials ⚠️
lightning/src/util/persist.rs68.75%3 Missing and 2 partials ⚠️
lightning-background-processor/src/lib.rs0.00%2 Missing ⚠️
lightning/src/util/test_utils.rs60.00%2 Missing ⚠️
lightning-liquidity/src/lsps2/service.rs0.00%1 Missing ⚠️
lightning-liquidity/src/lsps5/service.rs0.00%1 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #4189 +/- ##
==========================================
- Coverage 88.87% 88.84% -0.04% 
==========================================
Files 180 180 Lines 137863 137870 +7 Branches 137863 137870 +7 ==========================================
- Hits 122522 122485 -37 - Misses 12532 12573 +41 - Partials 2809 2812 +3 
FlagCoverage Δ
fuzzing21.44% <0.00%> (+0.58%)⬆️
tests88.68% <56.25%> (-0.04%)⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

/// potentially get lost on crash after the method returns. Therefore, this flag should only be
/// set for `remove` operations that can be safely replayed at a later time.
///
/// All removal operations must complete in a consistent total order with [`Self::write`]s

@tnulltnullOct 30, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not sure if this would even work. For example in FilesystemStore, if we simply call remove and leave the decision on when to sync the changes to disk to the OS, how could we be certain that the ordering is preserved? IIRC we basically concluded this can't be guaranteed, especially since different guarantees on different platforms might vary?

@tnulltnullOct 30, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that man 2 unlink states:

If the name was the last link to a file but any processes still have the file open, the file will remain in existence until the last file descriptor referring to it is closed.

That means that if we have a concurrent read, we may defer the actual deletion, allowing it to interact with following writes, e.g.:

| t1 | t2 | t3 |
| READ | unlink | |
| READ | ... | WRITE |
| READ | ... | |
| READ | ... | SYNC |
| READ | SYNC | |

Although, given we use rename for write, I do wonder if the unlink would simply get lost here as it would apply to the original file that is dropped already anyways?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To avoid this, maybe it is possible to constrain lazy removes to keys that won't ever be written again in the future?

I've read the context of this PR now, and it seems the perf issue was around monitor updates. I think those are never re-written?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not sure if this would even work. For example in FilesystemStore, if we simply call remove and leave the decision on when to sync the changes to disk to the OS, how could we be certain that the ordering is preserved?

Filesystems provide an order, the only thing they dont provide without an fsync is any kinds of guarantee its on disk. I don't think this is a problem.

Although, given we use rename for write, I do wonder if the unlink would simply get lost here as it would apply to the original file that is dropped already anyways?

Yes, that is how it should work on any reasonable filesystem. In theory its possible for some filesystems to fail the write rename part because the file still exists, but that's unrelated to the remove, that's just the read+write being at the same time.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Filesystems provide an order, the only thing they dont provide without an fsync is any kinds of guarantee its on disk. I don't think this is a problem.

A file that should have been removed, but is still there. Is that not a pb?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's the explicit point of the lazy flag - it allows a store to not guarantee that the entry will be removed if there's a crash/ill-timed restart.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Discussed this offline, man 3p unlink had me convinced this is safe to do on POSIX.

@ldk-reviews-bot

Copy link
Copy Markdown

👋 The first review has been submitted!

Do you think this PR is ready for a second reviewer? If so, click here to assign a second reviewer.

@tnull
tnull requested a review from joostjagerOctober 30, 2025 09:42

@joostjagerjoostjager left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are there more details on why/when the lazy flag is important?

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

In this case its important for perf when removing a large number of monitor updates on an fsstore (requiring fsync for each can add up rather substantially eg if we're removing 1k mon updates), but thinking about it more I think its also important for the same case in the async design - if you have a KVStore that handles ordering (eg like the locks in the fsstore/vss store) then the lazy flag allows you to spawn-and-forget a removal, rather than the callsite having to "block" waiting on your removal to finish.

@joostjager

Copy link
Copy Markdown
Contributor

1k updates is a lot. Do you think it adds much over let's say 50 updates? Maybe that also makes the perf problem go away without lazy flag...

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

1k updates seems entirely reasonable for a node doing lots of forwarding. ChannelMonitors can easily be a few thousand times larger than ChannelMonitorUpdates, so wanting to amortize over more ChannelMonitorUpdates seems very reasonable (the startup cost of more ChannelMonitorUpdates is pretty low, or at least is if your KVStore read latency is low or once we load them in parallel).

Not being able to pick a reasonable update count just because of an API limitation in how we do removals seems like a pretty weird limitation, no?

@joostjager

Copy link
Copy Markdown
Contributor

Not being able to pick a reasonable update count just because of an API limitation in how we do removals seems like a pretty weird limitation, no?

The question was just whether 1000 is reasonable, and you made it clear that it is 👍

joostjager
joostjager previously approved these changes Oct 30, 2025
This reverts commit 561da4c.
A user pointed out, when looking to upgrade to LDK 0.2, that the
`lazy` flag is actually quite important for performance when using
a `MonitorUpdatingPersister`, especially in synchronous persistence
mode.
Thus, we add it back here.
Fixeslightningdevkit#4188
In the previous commit we reverted
561da4c. One of the motivations
for it (in addition to `lazy` removals being somewhat less, though
still arguably useful in an async context) was that the ordering
requirements of `lazy` removals is somewhat unclear.
Here we simply default to the simplest safe option, requiring a
total order across all `write` and `remove` operations to the same
key, `lazy` or not.
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Fixed rustfmt

$ git diff-tree -U1 9973d780a 0f9548bf8
diff --git a/lightning/src/util/persist.rs b/lightning/src/util/persist.rs
index 78fdba2113..5d34603c96 100644
--- a/lightning/src/util/persist.rs+++ b/lightning/src/util/persist.rs@@ -1084,4 +1084,3 @@ where
let latest_update_id = current_monitor.get_latest_update_id();
- self- .cleanup_stale_updates_for_monitor_to(&monitor_key, latest_update_id, lazy)+ self.cleanup_stale_updates_for_monitor_to(&monitor_key, latest_update_id, lazy)
.await?;

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Only a rustfmt change since @joostjager ack'd, so landing.

@TheBlueMatt
TheBlueMatt merged commit d53d6b4 into lightningdevkit:mainOct 30, 2025
23 of 25 checks passed
@TheBlueMattTheBlueMatt mentioned this pull request Oct 30, 2025
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Backported to 0.2 in #4193

@domZippilli

Copy link
Copy Markdown
Contributor

🥳

@wvanlint

Copy link
Copy Markdown
Contributor

Thanks for landing this!

I think the comments above covered everything. We use the MonitorUpdatingPersister with maximum_pending_updates = 1000 for efficiency due to the high forwarding volume, and an update_persisted_channel call can trigger channel monitor update consolidation when maximum_pending_updates is reached. This consolidation results in maximum_pending_updates sequential KVStore::remove calls, which caused issues when it's performed in a non-lazy fashion. In our case in 0.1, it blocked the Tokio runtime (for ~7 ms * 1000 = 7s), but I assume it will affect the caller in the async design as well as Matt mentioned.

I was curious if there are possible simplifications, such as all remove calls being considered lazy or if remove can be constrained to keys that won't ever be written again in the future as Joost mentioned. But I see there are requirements coming from #4059 (comment) as well.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add lazy persistence back to KVStore::delete

6 participants

@TheBlueMatt@ldk-reviews-bot@joostjager@domZippilli@wvanlint@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Restore lazy flag to KVStore::remove - #4189

Merged
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
TheBlueMatt:2025-10-lazy-again
Oct 30, 2025
Merged

Restore lazy flag to KVStore::remove#4189
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
TheBlueMatt:2025-10-lazy-again

Conversation

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

A user pointed out, when looking to upgrade to LDK 0.2, that the
lazy flag is actually quite important for performance when using
a MonitorUpdatingPersister, especially in synchronous persistence
mode.

Thus, we add it back here.

@TheBlueMattTheBlueMatt added this to the 0.2 milestone Oct 30, 2025
@ldk-reviews-bot

ldk-reviews-bot commented Oct 30, 2025

Copy link
Copy Markdown

👋 Thanks for assigning @joostjager as a reviewer!
I'll wait for their review and will help manage the review process.
Once they submit their review, I'll check if a second reviewer would be helpful.

@TheBlueMattTheBlueMatt linked an issue Oct 30, 2025 that may be closed by this pull request
@codecov

codecovBot commented Oct 30, 2025

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 56.25000% with 21 lines in your changes missing coverage. Please review.
✅ Project coverage is 88.84%. Comparing base (02a9af9) to head (0f9548b).
⚠️ Report is 8 commits behind head on main.

Files with missing linesPatch %Lines
lightning-persister/src/fs_store.rs52.38%5 Missing and 5 partials ⚠️
lightning/src/util/persist.rs68.75%3 Missing and 2 partials ⚠️
lightning-background-processor/src/lib.rs0.00%2 Missing ⚠️
lightning/src/util/test_utils.rs60.00%2 Missing ⚠️
lightning-liquidity/src/lsps2/service.rs0.00%1 Missing ⚠️
lightning-liquidity/src/lsps5/service.rs0.00%1 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #4189 +/- ##
==========================================
- Coverage 88.87% 88.84% -0.04% 
==========================================
Files 180 180 Lines 137863 137870 +7 Branches 137863 137870 +7 ==========================================
- Hits 122522 122485 -37 - Misses 12532 12573 +41 - Partials 2809 2812 +3 
FlagCoverage Δ
fuzzing21.44% <0.00%> (+0.58%)⬆️
tests88.68% <56.25%> (-0.04%)⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

/// potentially get lost on crash after the method returns. Therefore, this flag should only be
/// set for `remove` operations that can be safely replayed at a later time.
///
/// All removal operations must complete in a consistent total order with [`Self::write`]s

@tnulltnullOct 30, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not sure if this would even work. For example in FilesystemStore, if we simply call remove and leave the decision on when to sync the changes to disk to the OS, how could we be certain that the ordering is preserved? IIRC we basically concluded this can't be guaranteed, especially since different guarantees on different platforms might vary?

@tnulltnullOct 30, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that man 2 unlink states:

If the name was the last link to a file but any processes still have the file open, the file will remain in existence until the last file descriptor referring to it is closed.

That means that if we have a concurrent read, we may defer the actual deletion, allowing it to interact with following writes, e.g.:

| t1 | t2 | t3 |
| READ | unlink | |
| READ | ... | WRITE |
| READ | ... | |
| READ | ... | SYNC |
| READ | SYNC | |

Although, given we use rename for write, I do wonder if the unlink would simply get lost here as it would apply to the original file that is dropped already anyways?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To avoid this, maybe it is possible to constrain lazy removes to keys that won't ever be written again in the future?

I've read the context of this PR now, and it seems the perf issue was around monitor updates. I think those are never re-written?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not sure if this would even work. For example in FilesystemStore, if we simply call remove and leave the decision on when to sync the changes to disk to the OS, how could we be certain that the ordering is preserved?

Filesystems provide an order, the only thing they dont provide without an fsync is any kinds of guarantee its on disk. I don't think this is a problem.

Although, given we use rename for write, I do wonder if the unlink would simply get lost here as it would apply to the original file that is dropped already anyways?

Yes, that is how it should work on any reasonable filesystem. In theory its possible for some filesystems to fail the write rename part because the file still exists, but that's unrelated to the remove, that's just the read+write being at the same time.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Filesystems provide an order, the only thing they dont provide without an fsync is any kinds of guarantee its on disk. I don't think this is a problem.

A file that should have been removed, but is still there. Is that not a pb?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's the explicit point of the lazy flag - it allows a store to not guarantee that the entry will be removed if there's a crash/ill-timed restart.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Discussed this offline, man 3p unlink had me convinced this is safe to do on POSIX.

@ldk-reviews-bot

Copy link
Copy Markdown

👋 The first review has been submitted!

Do you think this PR is ready for a second reviewer? If so, click here to assign a second reviewer.

@tnull
tnull requested a review from joostjagerOctober 30, 2025 09:42

@joostjagerjoostjager left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are there more details on why/when the lazy flag is important?

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

In this case its important for perf when removing a large number of monitor updates on an fsstore (requiring fsync for each can add up rather substantially eg if we're removing 1k mon updates), but thinking about it more I think its also important for the same case in the async design - if you have a KVStore that handles ordering (eg like the locks in the fsstore/vss store) then the lazy flag allows you to spawn-and-forget a removal, rather than the callsite having to "block" waiting on your removal to finish.

@joostjager

Copy link
Copy Markdown
Contributor

1k updates is a lot. Do you think it adds much over let's say 50 updates? Maybe that also makes the perf problem go away without lazy flag...

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

1k updates seems entirely reasonable for a node doing lots of forwarding. ChannelMonitors can easily be a few thousand times larger than ChannelMonitorUpdates, so wanting to amortize over more ChannelMonitorUpdates seems very reasonable (the startup cost of more ChannelMonitorUpdates is pretty low, or at least is if your KVStore read latency is low or once we load them in parallel).

Not being able to pick a reasonable update count just because of an API limitation in how we do removals seems like a pretty weird limitation, no?

@joostjager

Copy link
Copy Markdown
Contributor

Not being able to pick a reasonable update count just because of an API limitation in how we do removals seems like a pretty weird limitation, no?

The question was just whether 1000 is reasonable, and you made it clear that it is 👍

joostjager
joostjager previously approved these changes Oct 30, 2025
This reverts commit 561da4c.
A user pointed out, when looking to upgrade to LDK 0.2, that the
`lazy` flag is actually quite important for performance when using
a `MonitorUpdatingPersister`, especially in synchronous persistence
mode.
Thus, we add it back here.
Fixeslightningdevkit#4188
In the previous commit we reverted
561da4c. One of the motivations
for it (in addition to `lazy` removals being somewhat less, though
still arguably useful in an async context) was that the ordering
requirements of `lazy` removals is somewhat unclear.
Here we simply default to the simplest safe option, requiring a
total order across all `write` and `remove` operations to the same
key, `lazy` or not.
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Fixed rustfmt

$ git diff-tree -U1 9973d780a 0f9548bf8
diff --git a/lightning/src/util/persist.rs b/lightning/src/util/persist.rs
index 78fdba2113..5d34603c96 100644
--- a/lightning/src/util/persist.rs+++ b/lightning/src/util/persist.rs@@ -1084,4 +1084,3 @@ where
let latest_update_id = current_monitor.get_latest_update_id();
- self- .cleanup_stale_updates_for_monitor_to(&monitor_key, latest_update_id, lazy)+ self.cleanup_stale_updates_for_monitor_to(&monitor_key, latest_update_id, lazy)
.await?;

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Only a rustfmt change since @joostjager ack'd, so landing.

@TheBlueMatt
TheBlueMatt merged commit d53d6b4 into lightningdevkit:mainOct 30, 2025
23 of 25 checks passed
@TheBlueMattTheBlueMatt mentioned this pull request Oct 30, 2025
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Backported to 0.2 in #4193

@domZippilli

Copy link
Copy Markdown
Contributor

🥳

@wvanlint

Copy link
Copy Markdown
Contributor

Thanks for landing this!

I think the comments above covered everything. We use the MonitorUpdatingPersister with maximum_pending_updates = 1000 for efficiency due to the high forwarding volume, and an update_persisted_channel call can trigger channel monitor update consolidation when maximum_pending_updates is reached. This consolidation results in maximum_pending_updates sequential KVStore::remove calls, which caused issues when it's performed in a non-lazy fashion. In our case in 0.1, it blocked the Tokio runtime (for ~7 ms * 1000 = 7s), but I assume it will affect the caller in the async design as well as Matt mentioned.

I was curious if there are possible simplifications, such as all remove calls being considered lazy or if remove can be constrained to keys that won't ever be written again in the future as Joost mentioned. But I see there are requirements coming from #4059 (comment) as well.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add lazy persistence back to KVStore::delete

6 participants

@TheBlueMatt@ldk-reviews-bot@joostjager@domZippilli@wvanlint@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Restore lazy flag to KVStore::remove - #4189

Merged
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
TheBlueMatt:2025-10-lazy-again
Oct 30, 2025
Merged

Restore lazy flag to KVStore::remove#4189
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
TheBlueMatt:2025-10-lazy-again

Conversation

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

A user pointed out, when looking to upgrade to LDK 0.2, that the
lazy flag is actually quite important for performance when using
a MonitorUpdatingPersister, especially in synchronous persistence
mode.

Thus, we add it back here.

@TheBlueMattTheBlueMatt added this to the 0.2 milestone Oct 30, 2025
@ldk-reviews-bot

ldk-reviews-bot commented Oct 30, 2025

Copy link
Copy Markdown

👋 Thanks for assigning @joostjager as a reviewer!
I'll wait for their review and will help manage the review process.
Once they submit their review, I'll check if a second reviewer would be helpful.

@TheBlueMattTheBlueMatt linked an issue Oct 30, 2025 that may be closed by this pull request
@codecov

codecovBot commented Oct 30, 2025

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 56.25000% with 21 lines in your changes missing coverage. Please review.
✅ Project coverage is 88.84%. Comparing base (02a9af9) to head (0f9548b).
⚠️ Report is 8 commits behind head on main.

Files with missing linesPatch %Lines
lightning-persister/src/fs_store.rs52.38%5 Missing and 5 partials ⚠️
lightning/src/util/persist.rs68.75%3 Missing and 2 partials ⚠️
lightning-background-processor/src/lib.rs0.00%2 Missing ⚠️
lightning/src/util/test_utils.rs60.00%2 Missing ⚠️
lightning-liquidity/src/lsps2/service.rs0.00%1 Missing ⚠️
lightning-liquidity/src/lsps5/service.rs0.00%1 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #4189 +/- ##
==========================================
- Coverage 88.87% 88.84% -0.04% 
==========================================
Files 180 180 Lines 137863 137870 +7 Branches 137863 137870 +7 ==========================================
- Hits 122522 122485 -37 - Misses 12532 12573 +41 - Partials 2809 2812 +3 
FlagCoverage Δ
fuzzing21.44% <0.00%> (+0.58%)⬆️
tests88.68% <56.25%> (-0.04%)⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

/// potentially get lost on crash after the method returns. Therefore, this flag should only be
/// set for `remove` operations that can be safely replayed at a later time.
///
/// All removal operations must complete in a consistent total order with [`Self::write`]s

@tnulltnullOct 30, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not sure if this would even work. For example in FilesystemStore, if we simply call remove and leave the decision on when to sync the changes to disk to the OS, how could we be certain that the ordering is preserved? IIRC we basically concluded this can't be guaranteed, especially since different guarantees on different platforms might vary?

@tnulltnullOct 30, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that man 2 unlink states:

If the name was the last link to a file but any processes still have the file open, the file will remain in existence until the last file descriptor referring to it is closed.

That means that if we have a concurrent read, we may defer the actual deletion, allowing it to interact with following writes, e.g.:

| t1 | t2 | t3 |
| READ | unlink | |
| READ | ... | WRITE |
| READ | ... | |
| READ | ... | SYNC |
| READ | SYNC | |

Although, given we use rename for write, I do wonder if the unlink would simply get lost here as it would apply to the original file that is dropped already anyways?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To avoid this, maybe it is possible to constrain lazy removes to keys that won't ever be written again in the future?

I've read the context of this PR now, and it seems the perf issue was around monitor updates. I think those are never re-written?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not sure if this would even work. For example in FilesystemStore, if we simply call remove and leave the decision on when to sync the changes to disk to the OS, how could we be certain that the ordering is preserved?

Filesystems provide an order, the only thing they dont provide without an fsync is any kinds of guarantee its on disk. I don't think this is a problem.

Although, given we use rename for write, I do wonder if the unlink would simply get lost here as it would apply to the original file that is dropped already anyways?

Yes, that is how it should work on any reasonable filesystem. In theory its possible for some filesystems to fail the write rename part because the file still exists, but that's unrelated to the remove, that's just the read+write being at the same time.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Filesystems provide an order, the only thing they dont provide without an fsync is any kinds of guarantee its on disk. I don't think this is a problem.

A file that should have been removed, but is still there. Is that not a pb?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's the explicit point of the lazy flag - it allows a store to not guarantee that the entry will be removed if there's a crash/ill-timed restart.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Discussed this offline, man 3p unlink had me convinced this is safe to do on POSIX.

@ldk-reviews-bot

Copy link
Copy Markdown

👋 The first review has been submitted!

Do you think this PR is ready for a second reviewer? If so, click here to assign a second reviewer.

@tnull
tnull requested a review from joostjagerOctober 30, 2025 09:42

@joostjagerjoostjager left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are there more details on why/when the lazy flag is important?

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

In this case its important for perf when removing a large number of monitor updates on an fsstore (requiring fsync for each can add up rather substantially eg if we're removing 1k mon updates), but thinking about it more I think its also important for the same case in the async design - if you have a KVStore that handles ordering (eg like the locks in the fsstore/vss store) then the lazy flag allows you to spawn-and-forget a removal, rather than the callsite having to "block" waiting on your removal to finish.

@joostjager

Copy link
Copy Markdown
Contributor

1k updates is a lot. Do you think it adds much over let's say 50 updates? Maybe that also makes the perf problem go away without lazy flag...

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

1k updates seems entirely reasonable for a node doing lots of forwarding. ChannelMonitors can easily be a few thousand times larger than ChannelMonitorUpdates, so wanting to amortize over more ChannelMonitorUpdates seems very reasonable (the startup cost of more ChannelMonitorUpdates is pretty low, or at least is if your KVStore read latency is low or once we load them in parallel).

Not being able to pick a reasonable update count just because of an API limitation in how we do removals seems like a pretty weird limitation, no?

@joostjager

Copy link
Copy Markdown
Contributor

Not being able to pick a reasonable update count just because of an API limitation in how we do removals seems like a pretty weird limitation, no?

The question was just whether 1000 is reasonable, and you made it clear that it is 👍

joostjager
joostjager previously approved these changes Oct 30, 2025
This reverts commit 561da4c.
A user pointed out, when looking to upgrade to LDK 0.2, that the
`lazy` flag is actually quite important for performance when using
a `MonitorUpdatingPersister`, especially in synchronous persistence
mode.
Thus, we add it back here.
Fixeslightningdevkit#4188
In the previous commit we reverted
561da4c. One of the motivations
for it (in addition to `lazy` removals being somewhat less, though
still arguably useful in an async context) was that the ordering
requirements of `lazy` removals is somewhat unclear.
Here we simply default to the simplest safe option, requiring a
total order across all `write` and `remove` operations to the same
key, `lazy` or not.
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Fixed rustfmt

$ git diff-tree -U1 9973d780a 0f9548bf8
diff --git a/lightning/src/util/persist.rs b/lightning/src/util/persist.rs
index 78fdba2113..5d34603c96 100644
--- a/lightning/src/util/persist.rs+++ b/lightning/src/util/persist.rs@@ -1084,4 +1084,3 @@ where
let latest_update_id = current_monitor.get_latest_update_id();
- self- .cleanup_stale_updates_for_monitor_to(&monitor_key, latest_update_id, lazy)+ self.cleanup_stale_updates_for_monitor_to(&monitor_key, latest_update_id, lazy)
.await?;

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Only a rustfmt change since @joostjager ack'd, so landing.

@TheBlueMatt
TheBlueMatt merged commit d53d6b4 into lightningdevkit:mainOct 30, 2025
23 of 25 checks passed
@TheBlueMattTheBlueMatt mentioned this pull request Oct 30, 2025
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Backported to 0.2 in #4193

@domZippilli

Copy link
Copy Markdown
Contributor

🥳

@wvanlint

Copy link
Copy Markdown
Contributor

Thanks for landing this!

I think the comments above covered everything. We use the MonitorUpdatingPersister with maximum_pending_updates = 1000 for efficiency due to the high forwarding volume, and an update_persisted_channel call can trigger channel monitor update consolidation when maximum_pending_updates is reached. This consolidation results in maximum_pending_updates sequential KVStore::remove calls, which caused issues when it's performed in a non-lazy fashion. In our case in 0.1, it blocked the Tokio runtime (for ~7 ms * 1000 = 7s), but I assume it will affect the caller in the async design as well as Matt mentioned.

I was curious if there are possible simplifications, such as all remove calls being considered lazy or if remove can be constrained to keys that won't ever be written again in the future as Joost mentioned. But I see there are requirements coming from #4059 (comment) as well.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add lazy persistence back to KVStore::delete

6 participants

@TheBlueMatt@ldk-reviews-bot@joostjager@domZippilli@wvanlint@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Restore lazy flag to KVStore::remove - #4189

Merged
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
TheBlueMatt:2025-10-lazy-again
Oct 30, 2025
Merged

Restore lazy flag to KVStore::remove#4189
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
TheBlueMatt:2025-10-lazy-again

Conversation

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

A user pointed out, when looking to upgrade to LDK 0.2, that the
lazy flag is actually quite important for performance when using
a MonitorUpdatingPersister, especially in synchronous persistence
mode.

Thus, we add it back here.

@TheBlueMattTheBlueMatt added this to the 0.2 milestone Oct 30, 2025
@ldk-reviews-bot

ldk-reviews-bot commented Oct 30, 2025

Copy link
Copy Markdown

👋 Thanks for assigning @joostjager as a reviewer!
I'll wait for their review and will help manage the review process.
Once they submit their review, I'll check if a second reviewer would be helpful.

@TheBlueMattTheBlueMatt linked an issue Oct 30, 2025 that may be closed by this pull request
@codecov

codecovBot commented Oct 30, 2025

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 56.25000% with 21 lines in your changes missing coverage. Please review.
✅ Project coverage is 88.84%. Comparing base (02a9af9) to head (0f9548b).
⚠️ Report is 8 commits behind head on main.

Files with missing linesPatch %Lines
lightning-persister/src/fs_store.rs52.38%5 Missing and 5 partials ⚠️
lightning/src/util/persist.rs68.75%3 Missing and 2 partials ⚠️
lightning-background-processor/src/lib.rs0.00%2 Missing ⚠️
lightning/src/util/test_utils.rs60.00%2 Missing ⚠️
lightning-liquidity/src/lsps2/service.rs0.00%1 Missing ⚠️
lightning-liquidity/src/lsps5/service.rs0.00%1 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #4189 +/- ##
==========================================
- Coverage 88.87% 88.84% -0.04% 
==========================================
Files 180 180 Lines 137863 137870 +7 Branches 137863 137870 +7 ==========================================
- Hits 122522 122485 -37 - Misses 12532 12573 +41 - Partials 2809 2812 +3 
FlagCoverage Δ
fuzzing21.44% <0.00%> (+0.58%)⬆️
tests88.68% <56.25%> (-0.04%)⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

/// potentially get lost on crash after the method returns. Therefore, this flag should only be
/// set for `remove` operations that can be safely replayed at a later time.
///
/// All removal operations must complete in a consistent total order with [`Self::write`]s

@tnulltnullOct 30, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not sure if this would even work. For example in FilesystemStore, if we simply call remove and leave the decision on when to sync the changes to disk to the OS, how could we be certain that the ordering is preserved? IIRC we basically concluded this can't be guaranteed, especially since different guarantees on different platforms might vary?

@tnulltnullOct 30, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that man 2 unlink states:

If the name was the last link to a file but any processes still have the file open, the file will remain in existence until the last file descriptor referring to it is closed.

That means that if we have a concurrent read, we may defer the actual deletion, allowing it to interact with following writes, e.g.:

| t1 | t2 | t3 |
| READ | unlink | |
| READ | ... | WRITE |
| READ | ... | |
| READ | ... | SYNC |
| READ | SYNC | |

Although, given we use rename for write, I do wonder if the unlink would simply get lost here as it would apply to the original file that is dropped already anyways?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To avoid this, maybe it is possible to constrain lazy removes to keys that won't ever be written again in the future?

I've read the context of this PR now, and it seems the perf issue was around monitor updates. I think those are never re-written?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not sure if this would even work. For example in FilesystemStore, if we simply call remove and leave the decision on when to sync the changes to disk to the OS, how could we be certain that the ordering is preserved?

Filesystems provide an order, the only thing they dont provide without an fsync is any kinds of guarantee its on disk. I don't think this is a problem.

Although, given we use rename for write, I do wonder if the unlink would simply get lost here as it would apply to the original file that is dropped already anyways?

Yes, that is how it should work on any reasonable filesystem. In theory its possible for some filesystems to fail the write rename part because the file still exists, but that's unrelated to the remove, that's just the read+write being at the same time.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Filesystems provide an order, the only thing they dont provide without an fsync is any kinds of guarantee its on disk. I don't think this is a problem.

A file that should have been removed, but is still there. Is that not a pb?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's the explicit point of the lazy flag - it allows a store to not guarantee that the entry will be removed if there's a crash/ill-timed restart.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Discussed this offline, man 3p unlink had me convinced this is safe to do on POSIX.

@ldk-reviews-bot

Copy link
Copy Markdown

👋 The first review has been submitted!

Do you think this PR is ready for a second reviewer? If so, click here to assign a second reviewer.

@tnull
tnull requested a review from joostjagerOctober 30, 2025 09:42

@joostjagerjoostjager left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are there more details on why/when the lazy flag is important?

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

In this case its important for perf when removing a large number of monitor updates on an fsstore (requiring fsync for each can add up rather substantially eg if we're removing 1k mon updates), but thinking about it more I think its also important for the same case in the async design - if you have a KVStore that handles ordering (eg like the locks in the fsstore/vss store) then the lazy flag allows you to spawn-and-forget a removal, rather than the callsite having to "block" waiting on your removal to finish.

@joostjager

Copy link
Copy Markdown
Contributor

1k updates is a lot. Do you think it adds much over let's say 50 updates? Maybe that also makes the perf problem go away without lazy flag...

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

1k updates seems entirely reasonable for a node doing lots of forwarding. ChannelMonitors can easily be a few thousand times larger than ChannelMonitorUpdates, so wanting to amortize over more ChannelMonitorUpdates seems very reasonable (the startup cost of more ChannelMonitorUpdates is pretty low, or at least is if your KVStore read latency is low or once we load them in parallel).

Not being able to pick a reasonable update count just because of an API limitation in how we do removals seems like a pretty weird limitation, no?

@joostjager

Copy link
Copy Markdown
Contributor

Not being able to pick a reasonable update count just because of an API limitation in how we do removals seems like a pretty weird limitation, no?

The question was just whether 1000 is reasonable, and you made it clear that it is 👍

joostjager
joostjager previously approved these changes Oct 30, 2025
This reverts commit 561da4c.
A user pointed out, when looking to upgrade to LDK 0.2, that the
`lazy` flag is actually quite important for performance when using
a `MonitorUpdatingPersister`, especially in synchronous persistence
mode.
Thus, we add it back here.
Fixeslightningdevkit#4188
In the previous commit we reverted
561da4c. One of the motivations
for it (in addition to `lazy` removals being somewhat less, though
still arguably useful in an async context) was that the ordering
requirements of `lazy` removals is somewhat unclear.
Here we simply default to the simplest safe option, requiring a
total order across all `write` and `remove` operations to the same
key, `lazy` or not.
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Fixed rustfmt

$ git diff-tree -U1 9973d780a 0f9548bf8
diff --git a/lightning/src/util/persist.rs b/lightning/src/util/persist.rs
index 78fdba2113..5d34603c96 100644
--- a/lightning/src/util/persist.rs+++ b/lightning/src/util/persist.rs@@ -1084,4 +1084,3 @@ where
let latest_update_id = current_monitor.get_latest_update_id();
- self- .cleanup_stale_updates_for_monitor_to(&monitor_key, latest_update_id, lazy)+ self.cleanup_stale_updates_for_monitor_to(&monitor_key, latest_update_id, lazy)
.await?;

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Only a rustfmt change since @joostjager ack'd, so landing.

@TheBlueMatt
TheBlueMatt merged commit d53d6b4 into lightningdevkit:mainOct 30, 2025
23 of 25 checks passed
@TheBlueMattTheBlueMatt mentioned this pull request Oct 30, 2025
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Backported to 0.2 in #4193

@domZippilli

Copy link
Copy Markdown
Contributor

🥳

@wvanlint

Copy link
Copy Markdown
Contributor

Thanks for landing this!

I think the comments above covered everything. We use the MonitorUpdatingPersister with maximum_pending_updates = 1000 for efficiency due to the high forwarding volume, and an update_persisted_channel call can trigger channel monitor update consolidation when maximum_pending_updates is reached. This consolidation results in maximum_pending_updates sequential KVStore::remove calls, which caused issues when it's performed in a non-lazy fashion. In our case in 0.1, it blocked the Tokio runtime (for ~7 ms * 1000 = 7s), but I assume it will affect the caller in the async design as well as Matt mentioned.

I was curious if there are possible simplifications, such as all remove calls being considered lazy or if remove can be constrained to keys that won't ever be written again in the future as Joost mentioned. But I see there are requirements coming from #4059 (comment) as well.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add lazy persistence back to KVStore::delete

6 participants

@TheBlueMatt@ldk-reviews-bot@joostjager@domZippilli@wvanlint@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Restore lazy flag to KVStore::remove - #4189

Merged
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
TheBlueMatt:2025-10-lazy-again
Oct 30, 2025
Merged

Restore lazy flag to KVStore::remove#4189
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
TheBlueMatt:2025-10-lazy-again

Conversation

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

A user pointed out, when looking to upgrade to LDK 0.2, that the
lazy flag is actually quite important for performance when using
a MonitorUpdatingPersister, especially in synchronous persistence
mode.

Thus, we add it back here.

@TheBlueMattTheBlueMatt added this to the 0.2 milestone Oct 30, 2025
@ldk-reviews-bot

ldk-reviews-bot commented Oct 30, 2025

Copy link
Copy Markdown

👋 Thanks for assigning @joostjager as a reviewer!
I'll wait for their review and will help manage the review process.
Once they submit their review, I'll check if a second reviewer would be helpful.

@TheBlueMattTheBlueMatt linked an issue Oct 30, 2025 that may be closed by this pull request
@codecov

codecovBot commented Oct 30, 2025

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 56.25000% with 21 lines in your changes missing coverage. Please review.
✅ Project coverage is 88.84%. Comparing base (02a9af9) to head (0f9548b).
⚠️ Report is 8 commits behind head on main.

Files with missing linesPatch %Lines
lightning-persister/src/fs_store.rs52.38%5 Missing and 5 partials ⚠️
lightning/src/util/persist.rs68.75%3 Missing and 2 partials ⚠️
lightning-background-processor/src/lib.rs0.00%2 Missing ⚠️
lightning/src/util/test_utils.rs60.00%2 Missing ⚠️
lightning-liquidity/src/lsps2/service.rs0.00%1 Missing ⚠️
lightning-liquidity/src/lsps5/service.rs0.00%1 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #4189 +/- ##
==========================================
- Coverage 88.87% 88.84% -0.04% 
==========================================
Files 180 180 Lines 137863 137870 +7 Branches 137863 137870 +7 ==========================================
- Hits 122522 122485 -37 - Misses 12532 12573 +41 - Partials 2809 2812 +3 
FlagCoverage Δ
fuzzing21.44% <0.00%> (+0.58%)⬆️
tests88.68% <56.25%> (-0.04%)⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

/// potentially get lost on crash after the method returns. Therefore, this flag should only be
/// set for `remove` operations that can be safely replayed at a later time.
///
/// All removal operations must complete in a consistent total order with [`Self::write`]s

@tnulltnullOct 30, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not sure if this would even work. For example in FilesystemStore, if we simply call remove and leave the decision on when to sync the changes to disk to the OS, how could we be certain that the ordering is preserved? IIRC we basically concluded this can't be guaranteed, especially since different guarantees on different platforms might vary?

@tnulltnullOct 30, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that man 2 unlink states:

If the name was the last link to a file but any processes still have the file open, the file will remain in existence until the last file descriptor referring to it is closed.

That means that if we have a concurrent read, we may defer the actual deletion, allowing it to interact with following writes, e.g.:

| t1 | t2 | t3 |
| READ | unlink | |
| READ | ... | WRITE |
| READ | ... | |
| READ | ... | SYNC |
| READ | SYNC | |

Although, given we use rename for write, I do wonder if the unlink would simply get lost here as it would apply to the original file that is dropped already anyways?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To avoid this, maybe it is possible to constrain lazy removes to keys that won't ever be written again in the future?

I've read the context of this PR now, and it seems the perf issue was around monitor updates. I think those are never re-written?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not sure if this would even work. For example in FilesystemStore, if we simply call remove and leave the decision on when to sync the changes to disk to the OS, how could we be certain that the ordering is preserved?

Filesystems provide an order, the only thing they dont provide without an fsync is any kinds of guarantee its on disk. I don't think this is a problem.

Although, given we use rename for write, I do wonder if the unlink would simply get lost here as it would apply to the original file that is dropped already anyways?

Yes, that is how it should work on any reasonable filesystem. In theory its possible for some filesystems to fail the write rename part because the file still exists, but that's unrelated to the remove, that's just the read+write being at the same time.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Filesystems provide an order, the only thing they dont provide without an fsync is any kinds of guarantee its on disk. I don't think this is a problem.

A file that should have been removed, but is still there. Is that not a pb?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's the explicit point of the lazy flag - it allows a store to not guarantee that the entry will be removed if there's a crash/ill-timed restart.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Discussed this offline, man 3p unlink had me convinced this is safe to do on POSIX.

@ldk-reviews-bot

Copy link
Copy Markdown

👋 The first review has been submitted!

Do you think this PR is ready for a second reviewer? If so, click here to assign a second reviewer.

@tnull
tnull requested a review from joostjagerOctober 30, 2025 09:42

@joostjagerjoostjager left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are there more details on why/when the lazy flag is important?

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

In this case its important for perf when removing a large number of monitor updates on an fsstore (requiring fsync for each can add up rather substantially eg if we're removing 1k mon updates), but thinking about it more I think its also important for the same case in the async design - if you have a KVStore that handles ordering (eg like the locks in the fsstore/vss store) then the lazy flag allows you to spawn-and-forget a removal, rather than the callsite having to "block" waiting on your removal to finish.

@joostjager

Copy link
Copy Markdown
Contributor

1k updates is a lot. Do you think it adds much over let's say 50 updates? Maybe that also makes the perf problem go away without lazy flag...

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

1k updates seems entirely reasonable for a node doing lots of forwarding. ChannelMonitors can easily be a few thousand times larger than ChannelMonitorUpdates, so wanting to amortize over more ChannelMonitorUpdates seems very reasonable (the startup cost of more ChannelMonitorUpdates is pretty low, or at least is if your KVStore read latency is low or once we load them in parallel).

Not being able to pick a reasonable update count just because of an API limitation in how we do removals seems like a pretty weird limitation, no?

@joostjager

Copy link
Copy Markdown
Contributor

Not being able to pick a reasonable update count just because of an API limitation in how we do removals seems like a pretty weird limitation, no?

The question was just whether 1000 is reasonable, and you made it clear that it is 👍

joostjager
joostjager previously approved these changes Oct 30, 2025
This reverts commit 561da4c.
A user pointed out, when looking to upgrade to LDK 0.2, that the
`lazy` flag is actually quite important for performance when using
a `MonitorUpdatingPersister`, especially in synchronous persistence
mode.
Thus, we add it back here.
Fixeslightningdevkit#4188
In the previous commit we reverted
561da4c. One of the motivations
for it (in addition to `lazy` removals being somewhat less, though
still arguably useful in an async context) was that the ordering
requirements of `lazy` removals is somewhat unclear.
Here we simply default to the simplest safe option, requiring a
total order across all `write` and `remove` operations to the same
key, `lazy` or not.
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Fixed rustfmt

$ git diff-tree -U1 9973d780a 0f9548bf8
diff --git a/lightning/src/util/persist.rs b/lightning/src/util/persist.rs
index 78fdba2113..5d34603c96 100644
--- a/lightning/src/util/persist.rs+++ b/lightning/src/util/persist.rs@@ -1084,4 +1084,3 @@ where
let latest_update_id = current_monitor.get_latest_update_id();
- self- .cleanup_stale_updates_for_monitor_to(&monitor_key, latest_update_id, lazy)+ self.cleanup_stale_updates_for_monitor_to(&monitor_key, latest_update_id, lazy)
.await?;

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Only a rustfmt change since @joostjager ack'd, so landing.

@TheBlueMatt
TheBlueMatt merged commit d53d6b4 into lightningdevkit:mainOct 30, 2025
23 of 25 checks passed
@TheBlueMattTheBlueMatt mentioned this pull request Oct 30, 2025
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Backported to 0.2 in #4193

@domZippilli

Copy link
Copy Markdown
Contributor

🥳

@wvanlint

Copy link
Copy Markdown
Contributor

Thanks for landing this!

I think the comments above covered everything. We use the MonitorUpdatingPersister with maximum_pending_updates = 1000 for efficiency due to the high forwarding volume, and an update_persisted_channel call can trigger channel monitor update consolidation when maximum_pending_updates is reached. This consolidation results in maximum_pending_updates sequential KVStore::remove calls, which caused issues when it's performed in a non-lazy fashion. In our case in 0.1, it blocked the Tokio runtime (for ~7 ms * 1000 = 7s), but I assume it will affect the caller in the async design as well as Matt mentioned.

I was curious if there are possible simplifications, such as all remove calls being considered lazy or if remove can be constrained to keys that won't ever be written again in the future as Joost mentioned. But I see there are requirements coming from #4059 (comment) as well.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add lazy persistence back to KVStore::delete

6 participants

@TheBlueMatt@ldk-reviews-bot@joostjager@domZippilli@wvanlint@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Restore lazy flag to KVStore::remove - #4189

Merged
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
TheBlueMatt:2025-10-lazy-again
Oct 30, 2025
Merged

Restore lazy flag to KVStore::remove#4189
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
TheBlueMatt:2025-10-lazy-again

Conversation

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

A user pointed out, when looking to upgrade to LDK 0.2, that the
lazy flag is actually quite important for performance when using
a MonitorUpdatingPersister, especially in synchronous persistence
mode.

Thus, we add it back here.

@TheBlueMattTheBlueMatt added this to the 0.2 milestone Oct 30, 2025
@ldk-reviews-bot

ldk-reviews-bot commented Oct 30, 2025

Copy link
Copy Markdown

👋 Thanks for assigning @joostjager as a reviewer!
I'll wait for their review and will help manage the review process.
Once they submit their review, I'll check if a second reviewer would be helpful.

@TheBlueMattTheBlueMatt linked an issue Oct 30, 2025 that may be closed by this pull request
@codecov

codecovBot commented Oct 30, 2025

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 56.25000% with 21 lines in your changes missing coverage. Please review.
✅ Project coverage is 88.84%. Comparing base (02a9af9) to head (0f9548b).
⚠️ Report is 8 commits behind head on main.

Files with missing linesPatch %Lines
lightning-persister/src/fs_store.rs52.38%5 Missing and 5 partials ⚠️
lightning/src/util/persist.rs68.75%3 Missing and 2 partials ⚠️
lightning-background-processor/src/lib.rs0.00%2 Missing ⚠️
lightning/src/util/test_utils.rs60.00%2 Missing ⚠️
lightning-liquidity/src/lsps2/service.rs0.00%1 Missing ⚠️
lightning-liquidity/src/lsps5/service.rs0.00%1 Missing ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #4189 +/- ##
==========================================
- Coverage 88.87% 88.84% -0.04% 
==========================================
Files 180 180 Lines 137863 137870 +7 Branches 137863 137870 +7 ==========================================
- Hits 122522 122485 -37 - Misses 12532 12573 +41 - Partials 2809 2812 +3 
FlagCoverage Δ
fuzzing21.44% <0.00%> (+0.58%)⬆️
tests88.68% <56.25%> (-0.04%)⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

/// potentially get lost on crash after the method returns. Therefore, this flag should only be
/// set for `remove` operations that can be safely replayed at a later time.
///
/// All removal operations must complete in a consistent total order with [`Self::write`]s

@tnulltnullOct 30, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not sure if this would even work. For example in FilesystemStore, if we simply call remove and leave the decision on when to sync the changes to disk to the OS, how could we be certain that the ordering is preserved? IIRC we basically concluded this can't be guaranteed, especially since different guarantees on different platforms might vary?

@tnulltnullOct 30, 2025

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note that man 2 unlink states:

If the name was the last link to a file but any processes still have the file open, the file will remain in existence until the last file descriptor referring to it is closed.

That means that if we have a concurrent read, we may defer the actual deletion, allowing it to interact with following writes, e.g.:

| t1 | t2 | t3 |
| READ | unlink | |
| READ | ... | WRITE |
| READ | ... | |
| READ | ... | SYNC |
| READ | SYNC | |

Although, given we use rename for write, I do wonder if the unlink would simply get lost here as it would apply to the original file that is dropped already anyways?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To avoid this, maybe it is possible to constrain lazy removes to keys that won't ever be written again in the future?

I've read the context of this PR now, and it seems the perf issue was around monitor updates. I think those are never re-written?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not sure if this would even work. For example in FilesystemStore, if we simply call remove and leave the decision on when to sync the changes to disk to the OS, how could we be certain that the ordering is preserved?

Filesystems provide an order, the only thing they dont provide without an fsync is any kinds of guarantee its on disk. I don't think this is a problem.

Although, given we use rename for write, I do wonder if the unlink would simply get lost here as it would apply to the original file that is dropped already anyways?

Yes, that is how it should work on any reasonable filesystem. In theory its possible for some filesystems to fail the write rename part because the file still exists, but that's unrelated to the remove, that's just the read+write being at the same time.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Filesystems provide an order, the only thing they dont provide without an fsync is any kinds of guarantee its on disk. I don't think this is a problem.

A file that should have been removed, but is still there. Is that not a pb?

Copy link
Copy Markdown
CollaboratorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's the explicit point of the lazy flag - it allows a store to not guarantee that the entry will be removed if there's a crash/ill-timed restart.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Discussed this offline, man 3p unlink had me convinced this is safe to do on POSIX.

@ldk-reviews-bot

Copy link
Copy Markdown

👋 The first review has been submitted!

Do you think this PR is ready for a second reviewer? If so, click here to assign a second reviewer.

@tnull
tnull requested a review from joostjagerOctober 30, 2025 09:42

@joostjagerjoostjager left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are there more details on why/when the lazy flag is important?

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

In this case its important for perf when removing a large number of monitor updates on an fsstore (requiring fsync for each can add up rather substantially eg if we're removing 1k mon updates), but thinking about it more I think its also important for the same case in the async design - if you have a KVStore that handles ordering (eg like the locks in the fsstore/vss store) then the lazy flag allows you to spawn-and-forget a removal, rather than the callsite having to "block" waiting on your removal to finish.

@joostjager

Copy link
Copy Markdown
Contributor

1k updates is a lot. Do you think it adds much over let's say 50 updates? Maybe that also makes the perf problem go away without lazy flag...

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

1k updates seems entirely reasonable for a node doing lots of forwarding. ChannelMonitors can easily be a few thousand times larger than ChannelMonitorUpdates, so wanting to amortize over more ChannelMonitorUpdates seems very reasonable (the startup cost of more ChannelMonitorUpdates is pretty low, or at least is if your KVStore read latency is low or once we load them in parallel).

Not being able to pick a reasonable update count just because of an API limitation in how we do removals seems like a pretty weird limitation, no?

@joostjager

Copy link
Copy Markdown
Contributor

Not being able to pick a reasonable update count just because of an API limitation in how we do removals seems like a pretty weird limitation, no?

The question was just whether 1000 is reasonable, and you made it clear that it is 👍

joostjager
joostjager previously approved these changes Oct 30, 2025
This reverts commit 561da4c.
A user pointed out, when looking to upgrade to LDK 0.2, that the
`lazy` flag is actually quite important for performance when using
a `MonitorUpdatingPersister`, especially in synchronous persistence
mode.
Thus, we add it back here.
Fixeslightningdevkit#4188
In the previous commit we reverted
561da4c. One of the motivations
for it (in addition to `lazy` removals being somewhat less, though
still arguably useful in an async context) was that the ordering
requirements of `lazy` removals is somewhat unclear.
Here we simply default to the simplest safe option, requiring a
total order across all `write` and `remove` operations to the same
key, `lazy` or not.
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Fixed rustfmt

$ git diff-tree -U1 9973d780a 0f9548bf8
diff --git a/lightning/src/util/persist.rs b/lightning/src/util/persist.rs
index 78fdba2113..5d34603c96 100644
--- a/lightning/src/util/persist.rs+++ b/lightning/src/util/persist.rs@@ -1084,4 +1084,3 @@ where
let latest_update_id = current_monitor.get_latest_update_id();
- self- .cleanup_stale_updates_for_monitor_to(&monitor_key, latest_update_id, lazy)+ self.cleanup_stale_updates_for_monitor_to(&monitor_key, latest_update_id, lazy)
.await?;

@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Only a rustfmt change since @joostjager ack'd, so landing.

@TheBlueMatt
TheBlueMatt merged commit d53d6b4 into lightningdevkit:mainOct 30, 2025
23 of 25 checks passed
@TheBlueMattTheBlueMatt mentioned this pull request Oct 30, 2025
@TheBlueMatt

Copy link
Copy Markdown
CollaboratorAuthor

Backported to 0.2 in #4193

@domZippilli

Copy link
Copy Markdown
Contributor

🥳

@wvanlint

Copy link
Copy Markdown
Contributor

Thanks for landing this!

I think the comments above covered everything. We use the MonitorUpdatingPersister with maximum_pending_updates = 1000 for efficiency due to the high forwarding volume, and an update_persisted_channel call can trigger channel monitor update consolidation when maximum_pending_updates is reached. This consolidation results in maximum_pending_updates sequential KVStore::remove calls, which caused issues when it's performed in a non-lazy fashion. In our case in 0.1, it blocked the Tokio runtime (for ~7 ms * 1000 = 7s), but I assume it will affect the caller in the async design as well as Matt mentioned.

I was curious if there are possible simplifications, such as all remove calls being considered lazy or if remove can be constrained to keys that won't ever be written again in the future as Joost mentioned. But I see there are requirements coming from #4059 (comment) as well.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add lazy persistence back to KVStore::delete

6 participants

@TheBlueMatt@ldk-reviews-bot@joostjager@domZippilli@wvanlint@tnull