Optimize ChannelMonitor persistence on block connections. - #2966

Merged
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
G8XSU:2647-distribute
Jun 20, 2024
Merged

Optimize ChannelMonitor persistence on block connections.#2966
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
G8XSU:2647-distribute

Conversation

@G8XSU

@G8XSUG8XSU commented Mar 25, 2024

Copy link
Copy Markdown
Contributor

Currently, every block connection triggers the persistence of all
ChannelMonitors with an updated best_block. This approach poses
challenges for large node operators managing thousands of channels.
Furthermore, it leads to a thundering herd problem
(https://en.wikipedia.org/wiki/Thundering_herd_problem), overwhelming
the storage with simultaneous requests.

To address this issue, we now persist ChannelMonitors at a
regular cadence, spreading their persistence across blocks to
mitigate spikes in write operations.

Outcome: After doing this, Ldk's IO footprint should be reduced
by ~50 times. The processing time required to sync each block
will be significantly reduced, particularly for nodes with 1000s
of channels, as write latency plays a significant role in this process.
As a result, the Node/ChainMonitor will be blocked for a shorter
duration, leading to further efficiency gains.

Note that this will also increase time taken to sync during startup, a node
will now have to sync 25 blocks per channel on average, since monitors
can be at most 50 blocks out-of-date.

Based on #2957

Tasks:

  • Don't pause events for chainsync persistence Don't pause events for chainsync persistence #2957 and base it on that.
  • Concept/Approach Ack
  • Decide a good default for partition_factor [50 seems like a good number, every 50 blocks is ~8hours, this will reduce IO by a factor of 50 and at the same time shouldn't be a lot for mobile nodes to sync up, given they are routinely expected to sync this much after every night.]
  • Write more tests for persistence with partition_factor.
  • (Not a priority) Don't trigger chain-sync writes for closed channels/monitors. (This is next level of optimization and only offers very little improvement compared to rest of the changes, this PR without this will cut IO by 50 times, and even if we do this item, this further optimization will only reduce IO by 1-5%, so this can be done as followup and not urgent.)
  • (Not a priority) Maybe we can make partition_factor user-configurable. (We can do this separately and if needed, as our default should be sane enough for now.)

Closes#2647

@wpaulino

Copy link
Copy Markdown
Contributor

Is this something we might want to consider not doing on mobile? Thinking that we won't be able to RBF onchain claims properly if the fee estimator is broken and we're not persisting the most recent feerate we tried within the OnchainTxHandler.

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

I guess we should/could consider always persisting if there's pending claims (eg channel has been closed but has balances to claim)? Alternatively, we could always persist if we only have < 5 channels.

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

What's the status here @G8XSU?

@G8XSU

G8XSU commented Jun 4, 2024

Copy link
Copy Markdown
ContributorAuthor

Yes makes sense, we can always persist if there are pending claims.

I am looking for a concept/approach ack here before I proceed with rest of the changes. Are we in the right direction about how to distribute?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

I think wpaulino raised a good point and we should do something to ensure we regularly persist monitors on mobile (like what I suggested above), but otherwise concept ACK.

@G8XSU
G8XSUforce-pushed the 2647-distribute branch 2 times, most recently from 13f3e59 to d8b1203CompareJune 17, 2024 19:46
Comment threadlightning/src/chain/chainmonitor.rs Outdated
let funding_txid_u32 = u32::from_be_bytes([funding_txid_hash_bytes[0], funding_txid_hash_bytes[1], funding_txid_hash_bytes[2], funding_txid_hash_bytes[3]]);
funding_txid_u32.wrapping_add(best_height.unwrap_or_default())
};
const CHAINSYNC_MONITOR_PARTITION_FACTOR: u32 = 50; // ~ 8hours

@G8XSUG8XSUJun 17, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

50 seems like a good number, every 50 blocks is ~8hours, this will reduce IO by ~50 times and at the same time shouldn't be a lot for mobile nodes to sync up(if they use listen), given they are routinely expected to sync this much after every night.
for confirm users they are expected to sync all watched transactions on restart/reload in any case

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

50 seems fine to me, but can we do min(50, channel_count)? That way nodes with relatively few channels won't be paying the cost of startup needing a lot of replay. Specifically, I'm thinking mobile nodes or other nodes that might have few channels probably can pay the sync cost and might restart more often and don't want to pay the irregular-sync startup cost.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good 👍
Now accounting for small nodes/mobile for partition_factor separately.
Changed to piecewise function for bit more predictability for users compared to min.

It is helpful to assert that chain-sync did trigger a monitor
persist.
@G8XSU
G8XSUforce-pushed the 2647-distribute branch 2 times, most recently from abae637 to 2f29569CompareJune 17, 2024 19:55
@G8XSU
G8XSU marked this pull request as ready for review June 17, 2024 20:10
@G8XSU
G8XSU requested a review from TheBlueMattJune 17, 2024 20:10
@G8XSU

Copy link
Copy Markdown
ContributorAuthor

Marking this PR ready for review.

Comment threadlightning/src/chain/chainmonitor.rs
@G8XSU
G8XSU requested a review from wpaulinoJune 18, 2024 04:09

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, feel free to squash the fixup commit and lets find another reviewer.

Comment threadlightning/src/chain/chainmonitor.rs
Currently, every block connection triggers the persistence of all
ChannelMonitors with an updated best_block. This approach poses
challenges for large node operators managing thousands of channels.
Furthermore, it leads to a thundering herd problem
(https://en.wikipedia.org/wiki/Thundering_herd_problem), overwhelming
the storage with simultaneous requests.
To address this issue, we now persist ChannelMonitors at a
regular cadence, spreading their persistence across blocks to
mitigate spikes in write operations.
Outcome: After doing this, Ldk's IO footprint should be reduced
by ~50 times. The processing time required to sync each block
will be significantly reduced, particularly for nodes with 1000s
of channels, as write latency plays a significant role in this process.
As a result, the Node/ChainMonitor will be blocked for a shorter
duration, leading to further efficiency gains.
@G8XSU

Copy link
Copy Markdown
ContributorAuthor

Squashed Fixup commit.

@tnull
tnull self-requested a review June 20, 2024 07:17

@arik-soarik-so left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks good to me!

fn update_monitor_with_chain_data<FN>(
&self, header: &Header, txdata: &TransactionData, process: FN, funding_outpoint: &OutPoint,
monitor_state: &MonitorHolder<ChannelSigner>
&self, header: &Header, best_height: Option<u32>, txdata: &TransactionData, process: FN, funding_outpoint: &OutPoint,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what's the reason best_height is used to offset the modulus?

@G8XSUG8XSUJun 20, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We use it to distribute monitor persistence across time.

@tnull
tnull removed their request for review June 20, 2024 08:42
@TheBlueMatt
TheBlueMatt merged commit 07d991c into lightningdevkit:mainJun 20, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Persistence] Don't persist ALL channel_monitors on every bitcoin block connection.

4 participants

@G8XSU@wpaulino@TheBlueMatt@arik-so
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Optimize ChannelMonitor persistence on block connections. - #2966

Merged
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
G8XSU:2647-distribute
Jun 20, 2024
Merged

Optimize ChannelMonitor persistence on block connections.#2966
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
G8XSU:2647-distribute

Conversation

@G8XSU

@G8XSUG8XSU commented Mar 25, 2024

Copy link
Copy Markdown
Contributor

Currently, every block connection triggers the persistence of all
ChannelMonitors with an updated best_block. This approach poses
challenges for large node operators managing thousands of channels.
Furthermore, it leads to a thundering herd problem
(https://en.wikipedia.org/wiki/Thundering_herd_problem), overwhelming
the storage with simultaneous requests.

To address this issue, we now persist ChannelMonitors at a
regular cadence, spreading their persistence across blocks to
mitigate spikes in write operations.

Outcome: After doing this, Ldk's IO footprint should be reduced
by ~50 times. The processing time required to sync each block
will be significantly reduced, particularly for nodes with 1000s
of channels, as write latency plays a significant role in this process.
As a result, the Node/ChainMonitor will be blocked for a shorter
duration, leading to further efficiency gains.

Note that this will also increase time taken to sync during startup, a node
will now have to sync 25 blocks per channel on average, since monitors
can be at most 50 blocks out-of-date.

Based on #2957

Tasks:

  • Don't pause events for chainsync persistence Don't pause events for chainsync persistence #2957 and base it on that.
  • Concept/Approach Ack
  • Decide a good default for partition_factor [50 seems like a good number, every 50 blocks is ~8hours, this will reduce IO by a factor of 50 and at the same time shouldn't be a lot for mobile nodes to sync up, given they are routinely expected to sync this much after every night.]
  • Write more tests for persistence with partition_factor.
  • (Not a priority) Don't trigger chain-sync writes for closed channels/monitors. (This is next level of optimization and only offers very little improvement compared to rest of the changes, this PR without this will cut IO by 50 times, and even if we do this item, this further optimization will only reduce IO by 1-5%, so this can be done as followup and not urgent.)
  • (Not a priority) Maybe we can make partition_factor user-configurable. (We can do this separately and if needed, as our default should be sane enough for now.)

Closes#2647

@wpaulino

Copy link
Copy Markdown
Contributor

Is this something we might want to consider not doing on mobile? Thinking that we won't be able to RBF onchain claims properly if the fee estimator is broken and we're not persisting the most recent feerate we tried within the OnchainTxHandler.

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

I guess we should/could consider always persisting if there's pending claims (eg channel has been closed but has balances to claim)? Alternatively, we could always persist if we only have < 5 channels.

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

What's the status here @G8XSU?

@G8XSU

G8XSU commented Jun 4, 2024

Copy link
Copy Markdown
ContributorAuthor

Yes makes sense, we can always persist if there are pending claims.

I am looking for a concept/approach ack here before I proceed with rest of the changes. Are we in the right direction about how to distribute?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

I think wpaulino raised a good point and we should do something to ensure we regularly persist monitors on mobile (like what I suggested above), but otherwise concept ACK.

@G8XSU
G8XSUforce-pushed the 2647-distribute branch 2 times, most recently from 13f3e59 to d8b1203CompareJune 17, 2024 19:46
Comment threadlightning/src/chain/chainmonitor.rs Outdated
let funding_txid_u32 = u32::from_be_bytes([funding_txid_hash_bytes[0], funding_txid_hash_bytes[1], funding_txid_hash_bytes[2], funding_txid_hash_bytes[3]]);
funding_txid_u32.wrapping_add(best_height.unwrap_or_default())
};
const CHAINSYNC_MONITOR_PARTITION_FACTOR: u32 = 50; // ~ 8hours

@G8XSUG8XSUJun 17, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

50 seems like a good number, every 50 blocks is ~8hours, this will reduce IO by ~50 times and at the same time shouldn't be a lot for mobile nodes to sync up(if they use listen), given they are routinely expected to sync this much after every night.
for confirm users they are expected to sync all watched transactions on restart/reload in any case

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

50 seems fine to me, but can we do min(50, channel_count)? That way nodes with relatively few channels won't be paying the cost of startup needing a lot of replay. Specifically, I'm thinking mobile nodes or other nodes that might have few channels probably can pay the sync cost and might restart more often and don't want to pay the irregular-sync startup cost.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good 👍
Now accounting for small nodes/mobile for partition_factor separately.
Changed to piecewise function for bit more predictability for users compared to min.

It is helpful to assert that chain-sync did trigger a monitor
persist.
@G8XSU
G8XSUforce-pushed the 2647-distribute branch 2 times, most recently from abae637 to 2f29569CompareJune 17, 2024 19:55
@G8XSU
G8XSU marked this pull request as ready for review June 17, 2024 20:10
@G8XSU
G8XSU requested a review from TheBlueMattJune 17, 2024 20:10
@G8XSU

Copy link
Copy Markdown
ContributorAuthor

Marking this PR ready for review.

Comment threadlightning/src/chain/chainmonitor.rs
@G8XSU
G8XSU requested a review from wpaulinoJune 18, 2024 04:09

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, feel free to squash the fixup commit and lets find another reviewer.

Comment threadlightning/src/chain/chainmonitor.rs
Currently, every block connection triggers the persistence of all
ChannelMonitors with an updated best_block. This approach poses
challenges for large node operators managing thousands of channels.
Furthermore, it leads to a thundering herd problem
(https://en.wikipedia.org/wiki/Thundering_herd_problem), overwhelming
the storage with simultaneous requests.
To address this issue, we now persist ChannelMonitors at a
regular cadence, spreading their persistence across blocks to
mitigate spikes in write operations.
Outcome: After doing this, Ldk's IO footprint should be reduced
by ~50 times. The processing time required to sync each block
will be significantly reduced, particularly for nodes with 1000s
of channels, as write latency plays a significant role in this process.
As a result, the Node/ChainMonitor will be blocked for a shorter
duration, leading to further efficiency gains.
@G8XSU

Copy link
Copy Markdown
ContributorAuthor

Squashed Fixup commit.

@tnull
tnull self-requested a review June 20, 2024 07:17

@arik-soarik-so left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks good to me!

fn update_monitor_with_chain_data<FN>(
&self, header: &Header, txdata: &TransactionData, process: FN, funding_outpoint: &OutPoint,
monitor_state: &MonitorHolder<ChannelSigner>
&self, header: &Header, best_height: Option<u32>, txdata: &TransactionData, process: FN, funding_outpoint: &OutPoint,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what's the reason best_height is used to offset the modulus?

@G8XSUG8XSUJun 20, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We use it to distribute monitor persistence across time.

@tnull
tnull removed their request for review June 20, 2024 08:42
@TheBlueMatt
TheBlueMatt merged commit 07d991c into lightningdevkit:mainJun 20, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Persistence] Don't persist ALL channel_monitors on every bitcoin block connection.

4 participants

@G8XSU@wpaulino@TheBlueMatt@arik-so
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Optimize ChannelMonitor persistence on block connections. - #2966

Merged
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
G8XSU:2647-distribute
Jun 20, 2024
Merged

Optimize ChannelMonitor persistence on block connections.#2966
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
G8XSU:2647-distribute

Conversation

@G8XSU

@G8XSUG8XSU commented Mar 25, 2024

Copy link
Copy Markdown
Contributor

Currently, every block connection triggers the persistence of all
ChannelMonitors with an updated best_block. This approach poses
challenges for large node operators managing thousands of channels.
Furthermore, it leads to a thundering herd problem
(https://en.wikipedia.org/wiki/Thundering_herd_problem), overwhelming
the storage with simultaneous requests.

To address this issue, we now persist ChannelMonitors at a
regular cadence, spreading their persistence across blocks to
mitigate spikes in write operations.

Outcome: After doing this, Ldk's IO footprint should be reduced
by ~50 times. The processing time required to sync each block
will be significantly reduced, particularly for nodes with 1000s
of channels, as write latency plays a significant role in this process.
As a result, the Node/ChainMonitor will be blocked for a shorter
duration, leading to further efficiency gains.

Note that this will also increase time taken to sync during startup, a node
will now have to sync 25 blocks per channel on average, since monitors
can be at most 50 blocks out-of-date.

Based on #2957

Tasks:

  • Don't pause events for chainsync persistence Don't pause events for chainsync persistence #2957 and base it on that.
  • Concept/Approach Ack
  • Decide a good default for partition_factor [50 seems like a good number, every 50 blocks is ~8hours, this will reduce IO by a factor of 50 and at the same time shouldn't be a lot for mobile nodes to sync up, given they are routinely expected to sync this much after every night.]
  • Write more tests for persistence with partition_factor.
  • (Not a priority) Don't trigger chain-sync writes for closed channels/monitors. (This is next level of optimization and only offers very little improvement compared to rest of the changes, this PR without this will cut IO by 50 times, and even if we do this item, this further optimization will only reduce IO by 1-5%, so this can be done as followup and not urgent.)
  • (Not a priority) Maybe we can make partition_factor user-configurable. (We can do this separately and if needed, as our default should be sane enough for now.)

Closes#2647

@wpaulino

Copy link
Copy Markdown
Contributor

Is this something we might want to consider not doing on mobile? Thinking that we won't be able to RBF onchain claims properly if the fee estimator is broken and we're not persisting the most recent feerate we tried within the OnchainTxHandler.

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

I guess we should/could consider always persisting if there's pending claims (eg channel has been closed but has balances to claim)? Alternatively, we could always persist if we only have < 5 channels.

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

What's the status here @G8XSU?

@G8XSU

G8XSU commented Jun 4, 2024

Copy link
Copy Markdown
ContributorAuthor

Yes makes sense, we can always persist if there are pending claims.

I am looking for a concept/approach ack here before I proceed with rest of the changes. Are we in the right direction about how to distribute?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

I think wpaulino raised a good point and we should do something to ensure we regularly persist monitors on mobile (like what I suggested above), but otherwise concept ACK.

@G8XSU
G8XSUforce-pushed the 2647-distribute branch 2 times, most recently from 13f3e59 to d8b1203CompareJune 17, 2024 19:46
Comment threadlightning/src/chain/chainmonitor.rs Outdated
let funding_txid_u32 = u32::from_be_bytes([funding_txid_hash_bytes[0], funding_txid_hash_bytes[1], funding_txid_hash_bytes[2], funding_txid_hash_bytes[3]]);
funding_txid_u32.wrapping_add(best_height.unwrap_or_default())
};
const CHAINSYNC_MONITOR_PARTITION_FACTOR: u32 = 50; // ~ 8hours

@G8XSUG8XSUJun 17, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

50 seems like a good number, every 50 blocks is ~8hours, this will reduce IO by ~50 times and at the same time shouldn't be a lot for mobile nodes to sync up(if they use listen), given they are routinely expected to sync this much after every night.
for confirm users they are expected to sync all watched transactions on restart/reload in any case

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

50 seems fine to me, but can we do min(50, channel_count)? That way nodes with relatively few channels won't be paying the cost of startup needing a lot of replay. Specifically, I'm thinking mobile nodes or other nodes that might have few channels probably can pay the sync cost and might restart more often and don't want to pay the irregular-sync startup cost.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good 👍
Now accounting for small nodes/mobile for partition_factor separately.
Changed to piecewise function for bit more predictability for users compared to min.

It is helpful to assert that chain-sync did trigger a monitor
persist.
@G8XSU
G8XSUforce-pushed the 2647-distribute branch 2 times, most recently from abae637 to 2f29569CompareJune 17, 2024 19:55
@G8XSU
G8XSU marked this pull request as ready for review June 17, 2024 20:10
@G8XSU
G8XSU requested a review from TheBlueMattJune 17, 2024 20:10
@G8XSU

Copy link
Copy Markdown
ContributorAuthor

Marking this PR ready for review.

Comment threadlightning/src/chain/chainmonitor.rs
@G8XSU
G8XSU requested a review from wpaulinoJune 18, 2024 04:09

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, feel free to squash the fixup commit and lets find another reviewer.

Comment threadlightning/src/chain/chainmonitor.rs
Currently, every block connection triggers the persistence of all
ChannelMonitors with an updated best_block. This approach poses
challenges for large node operators managing thousands of channels.
Furthermore, it leads to a thundering herd problem
(https://en.wikipedia.org/wiki/Thundering_herd_problem), overwhelming
the storage with simultaneous requests.
To address this issue, we now persist ChannelMonitors at a
regular cadence, spreading their persistence across blocks to
mitigate spikes in write operations.
Outcome: After doing this, Ldk's IO footprint should be reduced
by ~50 times. The processing time required to sync each block
will be significantly reduced, particularly for nodes with 1000s
of channels, as write latency plays a significant role in this process.
As a result, the Node/ChainMonitor will be blocked for a shorter
duration, leading to further efficiency gains.
@G8XSU

Copy link
Copy Markdown
ContributorAuthor

Squashed Fixup commit.

@tnull
tnull self-requested a review June 20, 2024 07:17

@arik-soarik-so left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks good to me!

fn update_monitor_with_chain_data<FN>(
&self, header: &Header, txdata: &TransactionData, process: FN, funding_outpoint: &OutPoint,
monitor_state: &MonitorHolder<ChannelSigner>
&self, header: &Header, best_height: Option<u32>, txdata: &TransactionData, process: FN, funding_outpoint: &OutPoint,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what's the reason best_height is used to offset the modulus?

@G8XSUG8XSUJun 20, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We use it to distribute monitor persistence across time.

@tnull
tnull removed their request for review June 20, 2024 08:42
@TheBlueMatt
TheBlueMatt merged commit 07d991c into lightningdevkit:mainJun 20, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Persistence] Don't persist ALL channel_monitors on every bitcoin block connection.

4 participants

@G8XSU@wpaulino@TheBlueMatt@arik-so
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Optimize ChannelMonitor persistence on block connections. - #2966

Merged
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
G8XSU:2647-distribute
Jun 20, 2024
Merged

Optimize ChannelMonitor persistence on block connections.#2966
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
G8XSU:2647-distribute

Conversation

@G8XSU

@G8XSUG8XSU commented Mar 25, 2024

Copy link
Copy Markdown
Contributor

Currently, every block connection triggers the persistence of all
ChannelMonitors with an updated best_block. This approach poses
challenges for large node operators managing thousands of channels.
Furthermore, it leads to a thundering herd problem
(https://en.wikipedia.org/wiki/Thundering_herd_problem), overwhelming
the storage with simultaneous requests.

To address this issue, we now persist ChannelMonitors at a
regular cadence, spreading their persistence across blocks to
mitigate spikes in write operations.

Outcome: After doing this, Ldk's IO footprint should be reduced
by ~50 times. The processing time required to sync each block
will be significantly reduced, particularly for nodes with 1000s
of channels, as write latency plays a significant role in this process.
As a result, the Node/ChainMonitor will be blocked for a shorter
duration, leading to further efficiency gains.

Note that this will also increase time taken to sync during startup, a node
will now have to sync 25 blocks per channel on average, since monitors
can be at most 50 blocks out-of-date.

Based on #2957

Tasks:

  • Don't pause events for chainsync persistence Don't pause events for chainsync persistence #2957 and base it on that.
  • Concept/Approach Ack
  • Decide a good default for partition_factor [50 seems like a good number, every 50 blocks is ~8hours, this will reduce IO by a factor of 50 and at the same time shouldn't be a lot for mobile nodes to sync up, given they are routinely expected to sync this much after every night.]
  • Write more tests for persistence with partition_factor.
  • (Not a priority) Don't trigger chain-sync writes for closed channels/monitors. (This is next level of optimization and only offers very little improvement compared to rest of the changes, this PR without this will cut IO by 50 times, and even if we do this item, this further optimization will only reduce IO by 1-5%, so this can be done as followup and not urgent.)
  • (Not a priority) Maybe we can make partition_factor user-configurable. (We can do this separately and if needed, as our default should be sane enough for now.)

Closes#2647

@wpaulino

Copy link
Copy Markdown
Contributor

Is this something we might want to consider not doing on mobile? Thinking that we won't be able to RBF onchain claims properly if the fee estimator is broken and we're not persisting the most recent feerate we tried within the OnchainTxHandler.

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

I guess we should/could consider always persisting if there's pending claims (eg channel has been closed but has balances to claim)? Alternatively, we could always persist if we only have < 5 channels.

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

What's the status here @G8XSU?

@G8XSU

G8XSU commented Jun 4, 2024

Copy link
Copy Markdown
ContributorAuthor

Yes makes sense, we can always persist if there are pending claims.

I am looking for a concept/approach ack here before I proceed with rest of the changes. Are we in the right direction about how to distribute?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

I think wpaulino raised a good point and we should do something to ensure we regularly persist monitors on mobile (like what I suggested above), but otherwise concept ACK.

@G8XSU
G8XSUforce-pushed the 2647-distribute branch 2 times, most recently from 13f3e59 to d8b1203CompareJune 17, 2024 19:46
Comment threadlightning/src/chain/chainmonitor.rs Outdated
let funding_txid_u32 = u32::from_be_bytes([funding_txid_hash_bytes[0], funding_txid_hash_bytes[1], funding_txid_hash_bytes[2], funding_txid_hash_bytes[3]]);
funding_txid_u32.wrapping_add(best_height.unwrap_or_default())
};
const CHAINSYNC_MONITOR_PARTITION_FACTOR: u32 = 50; // ~ 8hours

@G8XSUG8XSUJun 17, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

50 seems like a good number, every 50 blocks is ~8hours, this will reduce IO by ~50 times and at the same time shouldn't be a lot for mobile nodes to sync up(if they use listen), given they are routinely expected to sync this much after every night.
for confirm users they are expected to sync all watched transactions on restart/reload in any case

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

50 seems fine to me, but can we do min(50, channel_count)? That way nodes with relatively few channels won't be paying the cost of startup needing a lot of replay. Specifically, I'm thinking mobile nodes or other nodes that might have few channels probably can pay the sync cost and might restart more often and don't want to pay the irregular-sync startup cost.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good 👍
Now accounting for small nodes/mobile for partition_factor separately.
Changed to piecewise function for bit more predictability for users compared to min.

It is helpful to assert that chain-sync did trigger a monitor
persist.
@G8XSU
G8XSUforce-pushed the 2647-distribute branch 2 times, most recently from abae637 to 2f29569CompareJune 17, 2024 19:55
@G8XSU
G8XSU marked this pull request as ready for review June 17, 2024 20:10
@G8XSU
G8XSU requested a review from TheBlueMattJune 17, 2024 20:10
@G8XSU

Copy link
Copy Markdown
ContributorAuthor

Marking this PR ready for review.

Comment threadlightning/src/chain/chainmonitor.rs
@G8XSU
G8XSU requested a review from wpaulinoJune 18, 2024 04:09

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, feel free to squash the fixup commit and lets find another reviewer.

Comment threadlightning/src/chain/chainmonitor.rs
Currently, every block connection triggers the persistence of all
ChannelMonitors with an updated best_block. This approach poses
challenges for large node operators managing thousands of channels.
Furthermore, it leads to a thundering herd problem
(https://en.wikipedia.org/wiki/Thundering_herd_problem), overwhelming
the storage with simultaneous requests.
To address this issue, we now persist ChannelMonitors at a
regular cadence, spreading their persistence across blocks to
mitigate spikes in write operations.
Outcome: After doing this, Ldk's IO footprint should be reduced
by ~50 times. The processing time required to sync each block
will be significantly reduced, particularly for nodes with 1000s
of channels, as write latency plays a significant role in this process.
As a result, the Node/ChainMonitor will be blocked for a shorter
duration, leading to further efficiency gains.
@G8XSU

Copy link
Copy Markdown
ContributorAuthor

Squashed Fixup commit.

@tnull
tnull self-requested a review June 20, 2024 07:17

@arik-soarik-so left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks good to me!

fn update_monitor_with_chain_data<FN>(
&self, header: &Header, txdata: &TransactionData, process: FN, funding_outpoint: &OutPoint,
monitor_state: &MonitorHolder<ChannelSigner>
&self, header: &Header, best_height: Option<u32>, txdata: &TransactionData, process: FN, funding_outpoint: &OutPoint,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what's the reason best_height is used to offset the modulus?

@G8XSUG8XSUJun 20, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We use it to distribute monitor persistence across time.

@tnull
tnull removed their request for review June 20, 2024 08:42
@TheBlueMatt
TheBlueMatt merged commit 07d991c into lightningdevkit:mainJun 20, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Persistence] Don't persist ALL channel_monitors on every bitcoin block connection.

4 participants

@G8XSU@wpaulino@TheBlueMatt@arik-so
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Optimize ChannelMonitor persistence on block connections. - #2966

Merged
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
G8XSU:2647-distribute
Jun 20, 2024
Merged

Optimize ChannelMonitor persistence on block connections.#2966
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
G8XSU:2647-distribute

Conversation

@G8XSU

@G8XSUG8XSU commented Mar 25, 2024

Copy link
Copy Markdown
Contributor

Currently, every block connection triggers the persistence of all
ChannelMonitors with an updated best_block. This approach poses
challenges for large node operators managing thousands of channels.
Furthermore, it leads to a thundering herd problem
(https://en.wikipedia.org/wiki/Thundering_herd_problem), overwhelming
the storage with simultaneous requests.

To address this issue, we now persist ChannelMonitors at a
regular cadence, spreading their persistence across blocks to
mitigate spikes in write operations.

Outcome: After doing this, Ldk's IO footprint should be reduced
by ~50 times. The processing time required to sync each block
will be significantly reduced, particularly for nodes with 1000s
of channels, as write latency plays a significant role in this process.
As a result, the Node/ChainMonitor will be blocked for a shorter
duration, leading to further efficiency gains.

Note that this will also increase time taken to sync during startup, a node
will now have to sync 25 blocks per channel on average, since monitors
can be at most 50 blocks out-of-date.

Based on #2957

Tasks:

  • Don't pause events for chainsync persistence Don't pause events for chainsync persistence #2957 and base it on that.
  • Concept/Approach Ack
  • Decide a good default for partition_factor [50 seems like a good number, every 50 blocks is ~8hours, this will reduce IO by a factor of 50 and at the same time shouldn't be a lot for mobile nodes to sync up, given they are routinely expected to sync this much after every night.]
  • Write more tests for persistence with partition_factor.
  • (Not a priority) Don't trigger chain-sync writes for closed channels/monitors. (This is next level of optimization and only offers very little improvement compared to rest of the changes, this PR without this will cut IO by 50 times, and even if we do this item, this further optimization will only reduce IO by 1-5%, so this can be done as followup and not urgent.)
  • (Not a priority) Maybe we can make partition_factor user-configurable. (We can do this separately and if needed, as our default should be sane enough for now.)

Closes#2647

@wpaulino

Copy link
Copy Markdown
Contributor

Is this something we might want to consider not doing on mobile? Thinking that we won't be able to RBF onchain claims properly if the fee estimator is broken and we're not persisting the most recent feerate we tried within the OnchainTxHandler.

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

I guess we should/could consider always persisting if there's pending claims (eg channel has been closed but has balances to claim)? Alternatively, we could always persist if we only have < 5 channels.

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

What's the status here @G8XSU?

@G8XSU

G8XSU commented Jun 4, 2024

Copy link
Copy Markdown
ContributorAuthor

Yes makes sense, we can always persist if there are pending claims.

I am looking for a concept/approach ack here before I proceed with rest of the changes. Are we in the right direction about how to distribute?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

I think wpaulino raised a good point and we should do something to ensure we regularly persist monitors on mobile (like what I suggested above), but otherwise concept ACK.

@G8XSU
G8XSUforce-pushed the 2647-distribute branch 2 times, most recently from 13f3e59 to d8b1203CompareJune 17, 2024 19:46
Comment threadlightning/src/chain/chainmonitor.rs Outdated
let funding_txid_u32 = u32::from_be_bytes([funding_txid_hash_bytes[0], funding_txid_hash_bytes[1], funding_txid_hash_bytes[2], funding_txid_hash_bytes[3]]);
funding_txid_u32.wrapping_add(best_height.unwrap_or_default())
};
const CHAINSYNC_MONITOR_PARTITION_FACTOR: u32 = 50; // ~ 8hours

@G8XSUG8XSUJun 17, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

50 seems like a good number, every 50 blocks is ~8hours, this will reduce IO by ~50 times and at the same time shouldn't be a lot for mobile nodes to sync up(if they use listen), given they are routinely expected to sync this much after every night.
for confirm users they are expected to sync all watched transactions on restart/reload in any case

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

50 seems fine to me, but can we do min(50, channel_count)? That way nodes with relatively few channels won't be paying the cost of startup needing a lot of replay. Specifically, I'm thinking mobile nodes or other nodes that might have few channels probably can pay the sync cost and might restart more often and don't want to pay the irregular-sync startup cost.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good 👍
Now accounting for small nodes/mobile for partition_factor separately.
Changed to piecewise function for bit more predictability for users compared to min.

It is helpful to assert that chain-sync did trigger a monitor
persist.
@G8XSU
G8XSUforce-pushed the 2647-distribute branch 2 times, most recently from abae637 to 2f29569CompareJune 17, 2024 19:55
@G8XSU
G8XSU marked this pull request as ready for review June 17, 2024 20:10
@G8XSU
G8XSU requested a review from TheBlueMattJune 17, 2024 20:10
@G8XSU

Copy link
Copy Markdown
ContributorAuthor

Marking this PR ready for review.

Comment threadlightning/src/chain/chainmonitor.rs
@G8XSU
G8XSU requested a review from wpaulinoJune 18, 2024 04:09

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, feel free to squash the fixup commit and lets find another reviewer.

Comment threadlightning/src/chain/chainmonitor.rs
Currently, every block connection triggers the persistence of all
ChannelMonitors with an updated best_block. This approach poses
challenges for large node operators managing thousands of channels.
Furthermore, it leads to a thundering herd problem
(https://en.wikipedia.org/wiki/Thundering_herd_problem), overwhelming
the storage with simultaneous requests.
To address this issue, we now persist ChannelMonitors at a
regular cadence, spreading their persistence across blocks to
mitigate spikes in write operations.
Outcome: After doing this, Ldk's IO footprint should be reduced
by ~50 times. The processing time required to sync each block
will be significantly reduced, particularly for nodes with 1000s
of channels, as write latency plays a significant role in this process.
As a result, the Node/ChainMonitor will be blocked for a shorter
duration, leading to further efficiency gains.
@G8XSU

Copy link
Copy Markdown
ContributorAuthor

Squashed Fixup commit.

@tnull
tnull self-requested a review June 20, 2024 07:17

@arik-soarik-so left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks good to me!

fn update_monitor_with_chain_data<FN>(
&self, header: &Header, txdata: &TransactionData, process: FN, funding_outpoint: &OutPoint,
monitor_state: &MonitorHolder<ChannelSigner>
&self, header: &Header, best_height: Option<u32>, txdata: &TransactionData, process: FN, funding_outpoint: &OutPoint,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what's the reason best_height is used to offset the modulus?

@G8XSUG8XSUJun 20, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We use it to distribute monitor persistence across time.

@tnull
tnull removed their request for review June 20, 2024 08:42
@TheBlueMatt
TheBlueMatt merged commit 07d991c into lightningdevkit:mainJun 20, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Persistence] Don't persist ALL channel_monitors on every bitcoin block connection.

4 participants

@G8XSU@wpaulino@TheBlueMatt@arik-so
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Optimize ChannelMonitor persistence on block connections. - #2966

Merged
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
G8XSU:2647-distribute
Jun 20, 2024
Merged

Optimize ChannelMonitor persistence on block connections.#2966
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
G8XSU:2647-distribute

Conversation

@G8XSU

@G8XSUG8XSU commented Mar 25, 2024

Copy link
Copy Markdown
Contributor

Currently, every block connection triggers the persistence of all
ChannelMonitors with an updated best_block. This approach poses
challenges for large node operators managing thousands of channels.
Furthermore, it leads to a thundering herd problem
(https://en.wikipedia.org/wiki/Thundering_herd_problem), overwhelming
the storage with simultaneous requests.

To address this issue, we now persist ChannelMonitors at a
regular cadence, spreading their persistence across blocks to
mitigate spikes in write operations.

Outcome: After doing this, Ldk's IO footprint should be reduced
by ~50 times. The processing time required to sync each block
will be significantly reduced, particularly for nodes with 1000s
of channels, as write latency plays a significant role in this process.
As a result, the Node/ChainMonitor will be blocked for a shorter
duration, leading to further efficiency gains.

Note that this will also increase time taken to sync during startup, a node
will now have to sync 25 blocks per channel on average, since monitors
can be at most 50 blocks out-of-date.

Based on #2957

Tasks:

  • Don't pause events for chainsync persistence Don't pause events for chainsync persistence #2957 and base it on that.
  • Concept/Approach Ack
  • Decide a good default for partition_factor [50 seems like a good number, every 50 blocks is ~8hours, this will reduce IO by a factor of 50 and at the same time shouldn't be a lot for mobile nodes to sync up, given they are routinely expected to sync this much after every night.]
  • Write more tests for persistence with partition_factor.
  • (Not a priority) Don't trigger chain-sync writes for closed channels/monitors. (This is next level of optimization and only offers very little improvement compared to rest of the changes, this PR without this will cut IO by 50 times, and even if we do this item, this further optimization will only reduce IO by 1-5%, so this can be done as followup and not urgent.)
  • (Not a priority) Maybe we can make partition_factor user-configurable. (We can do this separately and if needed, as our default should be sane enough for now.)

Closes#2647

@wpaulino

Copy link
Copy Markdown
Contributor

Is this something we might want to consider not doing on mobile? Thinking that we won't be able to RBF onchain claims properly if the fee estimator is broken and we're not persisting the most recent feerate we tried within the OnchainTxHandler.

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

I guess we should/could consider always persisting if there's pending claims (eg channel has been closed but has balances to claim)? Alternatively, we could always persist if we only have < 5 channels.

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

What's the status here @G8XSU?

@G8XSU

G8XSU commented Jun 4, 2024

Copy link
Copy Markdown
ContributorAuthor

Yes makes sense, we can always persist if there are pending claims.

I am looking for a concept/approach ack here before I proceed with rest of the changes. Are we in the right direction about how to distribute?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

I think wpaulino raised a good point and we should do something to ensure we regularly persist monitors on mobile (like what I suggested above), but otherwise concept ACK.

@G8XSU
G8XSUforce-pushed the 2647-distribute branch 2 times, most recently from 13f3e59 to d8b1203CompareJune 17, 2024 19:46
Comment threadlightning/src/chain/chainmonitor.rs Outdated
let funding_txid_u32 = u32::from_be_bytes([funding_txid_hash_bytes[0], funding_txid_hash_bytes[1], funding_txid_hash_bytes[2], funding_txid_hash_bytes[3]]);
funding_txid_u32.wrapping_add(best_height.unwrap_or_default())
};
const CHAINSYNC_MONITOR_PARTITION_FACTOR: u32 = 50; // ~ 8hours

@G8XSUG8XSUJun 17, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

50 seems like a good number, every 50 blocks is ~8hours, this will reduce IO by ~50 times and at the same time shouldn't be a lot for mobile nodes to sync up(if they use listen), given they are routinely expected to sync this much after every night.
for confirm users they are expected to sync all watched transactions on restart/reload in any case

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

50 seems fine to me, but can we do min(50, channel_count)? That way nodes with relatively few channels won't be paying the cost of startup needing a lot of replay. Specifically, I'm thinking mobile nodes or other nodes that might have few channels probably can pay the sync cost and might restart more often and don't want to pay the irregular-sync startup cost.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good 👍
Now accounting for small nodes/mobile for partition_factor separately.
Changed to piecewise function for bit more predictability for users compared to min.

It is helpful to assert that chain-sync did trigger a monitor
persist.
@G8XSU
G8XSUforce-pushed the 2647-distribute branch 2 times, most recently from abae637 to 2f29569CompareJune 17, 2024 19:55
@G8XSU
G8XSU marked this pull request as ready for review June 17, 2024 20:10
@G8XSU
G8XSU requested a review from TheBlueMattJune 17, 2024 20:10
@G8XSU

Copy link
Copy Markdown
ContributorAuthor

Marking this PR ready for review.

Comment threadlightning/src/chain/chainmonitor.rs
@G8XSU
G8XSU requested a review from wpaulinoJune 18, 2024 04:09

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, feel free to squash the fixup commit and lets find another reviewer.

Comment threadlightning/src/chain/chainmonitor.rs
Currently, every block connection triggers the persistence of all
ChannelMonitors with an updated best_block. This approach poses
challenges for large node operators managing thousands of channels.
Furthermore, it leads to a thundering herd problem
(https://en.wikipedia.org/wiki/Thundering_herd_problem), overwhelming
the storage with simultaneous requests.
To address this issue, we now persist ChannelMonitors at a
regular cadence, spreading their persistence across blocks to
mitigate spikes in write operations.
Outcome: After doing this, Ldk's IO footprint should be reduced
by ~50 times. The processing time required to sync each block
will be significantly reduced, particularly for nodes with 1000s
of channels, as write latency plays a significant role in this process.
As a result, the Node/ChainMonitor will be blocked for a shorter
duration, leading to further efficiency gains.
@G8XSU

Copy link
Copy Markdown
ContributorAuthor

Squashed Fixup commit.

@tnull
tnull self-requested a review June 20, 2024 07:17

@arik-soarik-so left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks good to me!

fn update_monitor_with_chain_data<FN>(
&self, header: &Header, txdata: &TransactionData, process: FN, funding_outpoint: &OutPoint,
monitor_state: &MonitorHolder<ChannelSigner>
&self, header: &Header, best_height: Option<u32>, txdata: &TransactionData, process: FN, funding_outpoint: &OutPoint,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what's the reason best_height is used to offset the modulus?

@G8XSUG8XSUJun 20, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We use it to distribute monitor persistence across time.

@tnull
tnull removed their request for review June 20, 2024 08:42
@TheBlueMatt
TheBlueMatt merged commit 07d991c into lightningdevkit:mainJun 20, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Persistence] Don't persist ALL channel_monitors on every bitcoin block connection.

4 participants

@G8XSU@wpaulino@TheBlueMatt@arik-so
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Optimize ChannelMonitor persistence on block connections. - #2966

Merged
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
G8XSU:2647-distribute
Jun 20, 2024
Merged

Optimize ChannelMonitor persistence on block connections.#2966
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
G8XSU:2647-distribute

Conversation

@G8XSU

@G8XSUG8XSU commented Mar 25, 2024

Copy link
Copy Markdown
Contributor

Currently, every block connection triggers the persistence of all
ChannelMonitors with an updated best_block. This approach poses
challenges for large node operators managing thousands of channels.
Furthermore, it leads to a thundering herd problem
(https://en.wikipedia.org/wiki/Thundering_herd_problem), overwhelming
the storage with simultaneous requests.

To address this issue, we now persist ChannelMonitors at a
regular cadence, spreading their persistence across blocks to
mitigate spikes in write operations.

Outcome: After doing this, Ldk's IO footprint should be reduced
by ~50 times. The processing time required to sync each block
will be significantly reduced, particularly for nodes with 1000s
of channels, as write latency plays a significant role in this process.
As a result, the Node/ChainMonitor will be blocked for a shorter
duration, leading to further efficiency gains.

Note that this will also increase time taken to sync during startup, a node
will now have to sync 25 blocks per channel on average, since monitors
can be at most 50 blocks out-of-date.

Based on #2957

Tasks:

  • Don't pause events for chainsync persistence Don't pause events for chainsync persistence #2957 and base it on that.
  • Concept/Approach Ack
  • Decide a good default for partition_factor [50 seems like a good number, every 50 blocks is ~8hours, this will reduce IO by a factor of 50 and at the same time shouldn't be a lot for mobile nodes to sync up, given they are routinely expected to sync this much after every night.]
  • Write more tests for persistence with partition_factor.
  • (Not a priority) Don't trigger chain-sync writes for closed channels/monitors. (This is next level of optimization and only offers very little improvement compared to rest of the changes, this PR without this will cut IO by 50 times, and even if we do this item, this further optimization will only reduce IO by 1-5%, so this can be done as followup and not urgent.)
  • (Not a priority) Maybe we can make partition_factor user-configurable. (We can do this separately and if needed, as our default should be sane enough for now.)

Closes#2647

@wpaulino

Copy link
Copy Markdown
Contributor

Is this something we might want to consider not doing on mobile? Thinking that we won't be able to RBF onchain claims properly if the fee estimator is broken and we're not persisting the most recent feerate we tried within the OnchainTxHandler.

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

I guess we should/could consider always persisting if there's pending claims (eg channel has been closed but has balances to claim)? Alternatively, we could always persist if we only have < 5 channels.

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

What's the status here @G8XSU?

@G8XSU

G8XSU commented Jun 4, 2024

Copy link
Copy Markdown
ContributorAuthor

Yes makes sense, we can always persist if there are pending claims.

I am looking for a concept/approach ack here before I proceed with rest of the changes. Are we in the right direction about how to distribute?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

I think wpaulino raised a good point and we should do something to ensure we regularly persist monitors on mobile (like what I suggested above), but otherwise concept ACK.

@G8XSU
G8XSUforce-pushed the 2647-distribute branch 2 times, most recently from 13f3e59 to d8b1203CompareJune 17, 2024 19:46
Comment threadlightning/src/chain/chainmonitor.rs Outdated
let funding_txid_u32 = u32::from_be_bytes([funding_txid_hash_bytes[0], funding_txid_hash_bytes[1], funding_txid_hash_bytes[2], funding_txid_hash_bytes[3]]);
funding_txid_u32.wrapping_add(best_height.unwrap_or_default())
};
const CHAINSYNC_MONITOR_PARTITION_FACTOR: u32 = 50; // ~ 8hours

@G8XSUG8XSUJun 17, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

50 seems like a good number, every 50 blocks is ~8hours, this will reduce IO by ~50 times and at the same time shouldn't be a lot for mobile nodes to sync up(if they use listen), given they are routinely expected to sync this much after every night.
for confirm users they are expected to sync all watched transactions on restart/reload in any case

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

50 seems fine to me, but can we do min(50, channel_count)? That way nodes with relatively few channels won't be paying the cost of startup needing a lot of replay. Specifically, I'm thinking mobile nodes or other nodes that might have few channels probably can pay the sync cost and might restart more often and don't want to pay the irregular-sync startup cost.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good 👍
Now accounting for small nodes/mobile for partition_factor separately.
Changed to piecewise function for bit more predictability for users compared to min.

It is helpful to assert that chain-sync did trigger a monitor
persist.
@G8XSU
G8XSUforce-pushed the 2647-distribute branch 2 times, most recently from abae637 to 2f29569CompareJune 17, 2024 19:55
@G8XSU
G8XSU marked this pull request as ready for review June 17, 2024 20:10
@G8XSU
G8XSU requested a review from TheBlueMattJune 17, 2024 20:10
@G8XSU

Copy link
Copy Markdown
ContributorAuthor

Marking this PR ready for review.

Comment threadlightning/src/chain/chainmonitor.rs
@G8XSU
G8XSU requested a review from wpaulinoJune 18, 2024 04:09

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, feel free to squash the fixup commit and lets find another reviewer.

Comment threadlightning/src/chain/chainmonitor.rs
Currently, every block connection triggers the persistence of all
ChannelMonitors with an updated best_block. This approach poses
challenges for large node operators managing thousands of channels.
Furthermore, it leads to a thundering herd problem
(https://en.wikipedia.org/wiki/Thundering_herd_problem), overwhelming
the storage with simultaneous requests.
To address this issue, we now persist ChannelMonitors at a
regular cadence, spreading their persistence across blocks to
mitigate spikes in write operations.
Outcome: After doing this, Ldk's IO footprint should be reduced
by ~50 times. The processing time required to sync each block
will be significantly reduced, particularly for nodes with 1000s
of channels, as write latency plays a significant role in this process.
As a result, the Node/ChainMonitor will be blocked for a shorter
duration, leading to further efficiency gains.
@G8XSU

Copy link
Copy Markdown
ContributorAuthor

Squashed Fixup commit.

@tnull
tnull self-requested a review June 20, 2024 07:17

@arik-soarik-so left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks good to me!

fn update_monitor_with_chain_data<FN>(
&self, header: &Header, txdata: &TransactionData, process: FN, funding_outpoint: &OutPoint,
monitor_state: &MonitorHolder<ChannelSigner>
&self, header: &Header, best_height: Option<u32>, txdata: &TransactionData, process: FN, funding_outpoint: &OutPoint,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what's the reason best_height is used to offset the modulus?

@G8XSUG8XSUJun 20, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We use it to distribute monitor persistence across time.

@tnull
tnull removed their request for review June 20, 2024 08:42
@TheBlueMatt
TheBlueMatt merged commit 07d991c into lightningdevkit:mainJun 20, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Persistence] Don't persist ALL channel_monitors on every bitcoin block connection.

4 participants

@G8XSU@wpaulino@TheBlueMatt@arik-so
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Optimize ChannelMonitor persistence on block connections. - #2966

Merged
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
G8XSU:2647-distribute
Jun 20, 2024
Merged

Optimize ChannelMonitor persistence on block connections.#2966
TheBlueMatt merged 2 commits into
lightningdevkit:mainfrom
G8XSU:2647-distribute

Conversation

@G8XSU

@G8XSUG8XSU commented Mar 25, 2024

Copy link
Copy Markdown
Contributor

Currently, every block connection triggers the persistence of all
ChannelMonitors with an updated best_block. This approach poses
challenges for large node operators managing thousands of channels.
Furthermore, it leads to a thundering herd problem
(https://en.wikipedia.org/wiki/Thundering_herd_problem), overwhelming
the storage with simultaneous requests.

To address this issue, we now persist ChannelMonitors at a
regular cadence, spreading their persistence across blocks to
mitigate spikes in write operations.

Outcome: After doing this, Ldk's IO footprint should be reduced
by ~50 times. The processing time required to sync each block
will be significantly reduced, particularly for nodes with 1000s
of channels, as write latency plays a significant role in this process.
As a result, the Node/ChainMonitor will be blocked for a shorter
duration, leading to further efficiency gains.

Note that this will also increase time taken to sync during startup, a node
will now have to sync 25 blocks per channel on average, since monitors
can be at most 50 blocks out-of-date.

Based on #2957

Tasks:

  • Don't pause events for chainsync persistence Don't pause events for chainsync persistence #2957 and base it on that.
  • Concept/Approach Ack
  • Decide a good default for partition_factor [50 seems like a good number, every 50 blocks is ~8hours, this will reduce IO by a factor of 50 and at the same time shouldn't be a lot for mobile nodes to sync up, given they are routinely expected to sync this much after every night.]
  • Write more tests for persistence with partition_factor.
  • (Not a priority) Don't trigger chain-sync writes for closed channels/monitors. (This is next level of optimization and only offers very little improvement compared to rest of the changes, this PR without this will cut IO by 50 times, and even if we do this item, this further optimization will only reduce IO by 1-5%, so this can be done as followup and not urgent.)
  • (Not a priority) Maybe we can make partition_factor user-configurable. (We can do this separately and if needed, as our default should be sane enough for now.)

Closes#2647

@wpaulino

Copy link
Copy Markdown
Contributor

Is this something we might want to consider not doing on mobile? Thinking that we won't be able to RBF onchain claims properly if the fee estimator is broken and we're not persisting the most recent feerate we tried within the OnchainTxHandler.

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

I guess we should/could consider always persisting if there's pending claims (eg channel has been closed but has balances to claim)? Alternatively, we could always persist if we only have < 5 channels.

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

What's the status here @G8XSU?

@G8XSU

G8XSU commented Jun 4, 2024

Copy link
Copy Markdown
ContributorAuthor

Yes makes sense, we can always persist if there are pending claims.

I am looking for a concept/approach ack here before I proceed with rest of the changes. Are we in the right direction about how to distribute?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

I think wpaulino raised a good point and we should do something to ensure we regularly persist monitors on mobile (like what I suggested above), but otherwise concept ACK.

@G8XSU
G8XSUforce-pushed the 2647-distribute branch 2 times, most recently from 13f3e59 to d8b1203CompareJune 17, 2024 19:46
Comment threadlightning/src/chain/chainmonitor.rs Outdated
let funding_txid_u32 = u32::from_be_bytes([funding_txid_hash_bytes[0], funding_txid_hash_bytes[1], funding_txid_hash_bytes[2], funding_txid_hash_bytes[3]]);
funding_txid_u32.wrapping_add(best_height.unwrap_or_default())
};
const CHAINSYNC_MONITOR_PARTITION_FACTOR: u32 = 50; // ~ 8hours

@G8XSUG8XSUJun 17, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

50 seems like a good number, every 50 blocks is ~8hours, this will reduce IO by ~50 times and at the same time shouldn't be a lot for mobile nodes to sync up(if they use listen), given they are routinely expected to sync this much after every night.
for confirm users they are expected to sync all watched transactions on restart/reload in any case

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

50 seems fine to me, but can we do min(50, channel_count)? That way nodes with relatively few channels won't be paying the cost of startup needing a lot of replay. Specifically, I'm thinking mobile nodes or other nodes that might have few channels probably can pay the sync cost and might restart more often and don't want to pay the irregular-sync startup cost.

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good 👍
Now accounting for small nodes/mobile for partition_factor separately.
Changed to piecewise function for bit more predictability for users compared to min.

It is helpful to assert that chain-sync did trigger a monitor
persist.
@G8XSU
G8XSUforce-pushed the 2647-distribute branch 2 times, most recently from abae637 to 2f29569CompareJune 17, 2024 19:55
@G8XSU
G8XSU marked this pull request as ready for review June 17, 2024 20:10
@G8XSU
G8XSU requested a review from TheBlueMattJune 17, 2024 20:10
@G8XSU

Copy link
Copy Markdown
ContributorAuthor

Marking this PR ready for review.

Comment threadlightning/src/chain/chainmonitor.rs
@G8XSU
G8XSU requested a review from wpaulinoJune 18, 2024 04:09

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, feel free to squash the fixup commit and lets find another reviewer.

Comment threadlightning/src/chain/chainmonitor.rs
Currently, every block connection triggers the persistence of all
ChannelMonitors with an updated best_block. This approach poses
challenges for large node operators managing thousands of channels.
Furthermore, it leads to a thundering herd problem
(https://en.wikipedia.org/wiki/Thundering_herd_problem), overwhelming
the storage with simultaneous requests.
To address this issue, we now persist ChannelMonitors at a
regular cadence, spreading their persistence across blocks to
mitigate spikes in write operations.
Outcome: After doing this, Ldk's IO footprint should be reduced
by ~50 times. The processing time required to sync each block
will be significantly reduced, particularly for nodes with 1000s
of channels, as write latency plays a significant role in this process.
As a result, the Node/ChainMonitor will be blocked for a shorter
duration, leading to further efficiency gains.
@G8XSU

Copy link
Copy Markdown
ContributorAuthor

Squashed Fixup commit.

@tnull
tnull self-requested a review June 20, 2024 07:17

@arik-soarik-so left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looks good to me!

fn update_monitor_with_chain_data<FN>(
&self, header: &Header, txdata: &TransactionData, process: FN, funding_outpoint: &OutPoint,
monitor_state: &MonitorHolder<ChannelSigner>
&self, header: &Header, best_height: Option<u32>, txdata: &TransactionData, process: FN, funding_outpoint: &OutPoint,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what's the reason best_height is used to offset the modulus?

@G8XSUG8XSUJun 20, 2024

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We use it to distribute monitor persistence across time.

@tnull
tnull removed their request for review June 20, 2024 08:42
@TheBlueMatt
TheBlueMatt merged commit 07d991c into lightningdevkit:mainJun 20, 2024
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Persistence] Don't persist ALL channel_monitors on every bitcoin block connection.

4 participants

@G8XSU@wpaulino@TheBlueMatt@arik-so