Skip to content

Implement concurrent broadcast tolerance for distributed watchtowers - #679

Merged
TheBlueMatt merged 3 commits into
lightningdevkit:masterfrom
ariard:2020-08-concurrent-watchtowers
Sep 16, 2020
Merged

Implement concurrent broadcast tolerance for distributed watchtowers#679
TheBlueMatt merged 3 commits into
lightningdevkit:masterfrom
ariard:2020-08-concurrent-watchtowers

Conversation

@ariard

Copy link
Copy Markdown

Implement concurrent broadcast tolerance for distributed watchtowers

With a distrbuted watchtowers deployment, where each monitor is plugged
to its own chain view, there is no guarantee that block are going to be
seen in same order. Watchtower may diverge in their acceptance of a
submitted commitment_signed update due to a block timing-out a HTLC
and provoking a subset but yet not seen by the other watchtower subset.
Any update reject by one of the watchtower must block offchain coordinator
to move channel state forward and release revocation secret for previous
state.

In this case, we want any watchtower from the rejection subset to still
be able to claim outputs if the concurrent state, has accepted by the
other subset, is confirming. This improve overall watchtower system
fault-tolerance.

This change stores local commitment transaction unconditionally and fail
the update if there is knowledge of an already signed commitment
transaction (ChannelMonitor.local_tx_signed=true).

Follow-up of #667, which let us easily implement concurrent tolerance and get ride off of confusing OnchainTxHandler/ChannelMonitor comments.

@codecov

codecovBot commented Aug 28, 2020

Copy link
Copy Markdown

Codecov Report

Merging #679 into master will decrease coverage by 0.01%.
The diff coverage is 98.70%.

Impacted file tree graph

@@ Coverage Diff @@## master #679 +/- ##
==========================================
- Coverage 91.92% 91.91% -0.02% 
==========================================
Files 35 35 Lines 20082 20154 +72 ==========================================
+ Hits 18461 18524 +63 - Misses 1621 1630 +9 
Impacted FilesCoverage Δ
lightning/src/ln/functional_tests.rs96.98% <98.63%> (-0.16%)⬇️
lightning/src/ln/channelmonitor.rs94.99% <100.00%> (ø)
lightning/src/ln/onchaintx.rs93.89% <100.00%> (-0.02%)⬇️
lightning/src/ln/channel.rs86.23% <0.00%> (+0.08%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 343aacc...6622ea7. Read the comment docs.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
mem::swap(&mut new_local_commitment_tx, &mut self.current_local_commitment_tx);
self.prev_local_signed_commitment_tx = Some(new_local_commitment_tx);
if self.local_tx_signed {
return Err(MonitorUpdateError("Latest local commitment signed has already been signed, update is rejected"));

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updating before returning an Err probably needs a documentation update to indicate to the user that they should still write out the new version to disk even in case of err (which is a super surprising API).

@ariardariardAug 28, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added a new comment a the top-level API : ManyChannelMonitor::update_monitor

I know this API is surprising, but you rarely have distributed system driven both by a central coordinator and a p2p network.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
if let Err(_) = self.onchain_tx_handler.provide_latest_local_tx(commitment_tx) {
return Err(MonitorUpdateError("Local commitment signed has already been signed, no further update of LOCAL commitment transaction is allowed"));
}
self.onchain_tx_handler.provide_latest_local_tx(commitment_tx);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we assert somewhere in here that prev_local_signed_commitment_tx must have been None before this? I suppose return Err and debug_assert!(false).

@ariardariardAug 28, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't explicitly prune out prev_local_signed_commitment=None.
We could add a round-trip between coordinator and watchtowers with a ChannelMonitorUpdateStep::PrunePrevLocalCommitment but that would be a latency hit for nothing, assuming we never return an Err if we broadcast a local commitment and coordinator fail advancing forward the channel by releasing the secret. Note the broadcast-commitment-then-reject-update is covered in new test (watchtower_alice failing the update_monitor).

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, right.

Comment threadlightning/src/ln/functional_tests.rs Outdated

#[test]
fn test_concurrent_monitor_claim() {
// Watchtower Alice receives block 135, broadcasts state X, rejects state Y.

@TheBlueMattTheBlueMattAug 28, 2020

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you make this comment more chronological? I think this is what threw me off the other day,

Watchtower A receives block, broadcasts state,
then channel receives new state Y, sending it to both watchtowers,
Bob accepts Y, then receives block and broadcasts the latest state Y,
Alice rejects state Y, but Bob has already broadcasted it,
Y confirms.

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't read through the test yet, but I get what you're going for more. A few relatively minor comments but conceptually looks good.

// OnchainTxHandler. After this is set, no future updates to our local commitment transactions
// may occur, and we fail any such monitor updates.
//
// In case of update rejection due to a locally already signed commitment transaction, we

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe:
Note that even when we fail a local commitment transaction update, we still store the update to ensure we can claim from it in case a duplicate copy of this ChannelMonitor broadcasts it.

/// Any spends of outputs which should have been registered which aren't passed to
/// ChannelMonitors via block_connected may result in FUNDS LOSS.
///
/// In case of distributed watchtowers deployment, even if an Err is return, the new version

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This needs more detail about the exact setup of a distributed ChannelMonitor, as several are possible (I think its a misnomer to call it a watchtower as that implies something specific about availability of keys which isn't true today), and probably this comment should be on ChannelMonitor::update_monitor, not the trait (though it could be referenced in the trait), because the trait's return value is only relevant for ChannelManager.

@ariardariardAug 31, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know writing a documentation for a distributed ChannelMonitor is tracked by #604. I intend to do it post-anchor as there are implications for bumping utxo management, namely you should also duplicate those keys across distributed monitors.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
///
/// Should also be used to indicate a failure to update the local persisted copy of the channel
/// monitor.
/// At reception of this error, channel manager MUST force-close the channel and return at

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it makes sense to describe this in these terms - what ChannelManager MUST do isn't interesting to our users, its what we do! Instead, we should describe it as ChannelManagerwill do X.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
/// This failure may also signal a failure to update the local persisted copy of one of
/// the channel monitor instance.
///
/// In case of a distributed watchtower deployment, failures sources can be divergent and

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I mean also because there's no use in any more granular errors :).

@ariard

Copy link
Copy Markdown
Author

Updated, for complete documentation, I will do it post-anchor/chaininterface refactoring.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
/// This failure may also signal a failure to update the local persisted copy of one of
/// the channel monitor instance.
///
/// Note that even when we fail a local commitment transaction update, we still store the

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be phrased as "you must" and probably be more clear (not "we" in the sense of "rust-lightning does"). Also we should distinguish between "ChannelMonitor returned" and "you are returning from a ManyChannelMonitor/chain::Watch", though if you want to just make @jkczyz deal with this in #649 thats ok too.

Antoine Riard added 3 commits September 15, 2020 18:17
With a distrbuted watchtowers deployment, where each monitor is plugged
to its own chain view, there is no guarantee that block are going to be
seen in same order. Watchtower may diverge in their acceptance of a
submitted `commitment_signed` update due to a block timing-out a HTLC
and provoking a subset but yet not seen by the other watchtower subset.
Any update reject by one of the watchtower must block offchain coordinator
to move channel state forward and release revocation secret for previous
state.
In this case, we want any watchtower from the rejection subset to still
be able to claim outputs if the concurrent state, has accepted by the
other subset, is confirming. This improve overall watchtower system
fault-tolerance.
This change stores local commitment transaction unconditionally and fail
the update if there is knowledge of an already signed commitment
transaction (ChannelMonitor.local_tx_signed=true).
Watchower Alice receives block 134, broadcasts state X, rejects state Y.
Watchtower Bob accepts state Y, receives blocks 135, broadcasts state Y.
State Y confirms onchain. Alice must be able to claim outputs.
Sources of the failure may be multiple in case of distributed watchtower
deployment. In either case, the channel manager must return a final
update asking to its channel monitor(s) to broadcast the lastest state
available. Revocation secret must not be released for the faultive
channel.
In the future, we may return wider type of failures to take more
fine-grained processing decision (e.g if local disk failure and
redudant remote channel copy available channel may still be processed
forward).
@TheBlueMatt
TheBlueMatt merged commit 44734be into lightningdevkit:masterSep 16, 2020
TheBlueMatt added a commit that referenced this pull request Sep 16, 2020
Implement concurrent broadcast tolerance for distributed watchtowers
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ariard@TheBlueMatt
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Implement concurrent broadcast tolerance for distributed watchtowers by ariard · Pull Request #679 · lightningdevkit/rust-lightning · GitHub
Skip to content

Implement concurrent broadcast tolerance for distributed watchtowers - #679

Merged
TheBlueMatt merged 3 commits into
lightningdevkit:masterfrom
ariard:2020-08-concurrent-watchtowers
Sep 16, 2020
Merged

Implement concurrent broadcast tolerance for distributed watchtowers#679
TheBlueMatt merged 3 commits into
lightningdevkit:masterfrom
ariard:2020-08-concurrent-watchtowers

Conversation

@ariard

Copy link
Copy Markdown

Implement concurrent broadcast tolerance for distributed watchtowers

With a distrbuted watchtowers deployment, where each monitor is plugged
to its own chain view, there is no guarantee that block are going to be
seen in same order. Watchtower may diverge in their acceptance of a
submitted commitment_signed update due to a block timing-out a HTLC
and provoking a subset but yet not seen by the other watchtower subset.
Any update reject by one of the watchtower must block offchain coordinator
to move channel state forward and release revocation secret for previous
state.

In this case, we want any watchtower from the rejection subset to still
be able to claim outputs if the concurrent state, has accepted by the
other subset, is confirming. This improve overall watchtower system
fault-tolerance.

This change stores local commitment transaction unconditionally and fail
the update if there is knowledge of an already signed commitment
transaction (ChannelMonitor.local_tx_signed=true).

Follow-up of #667, which let us easily implement concurrent tolerance and get ride off of confusing OnchainTxHandler/ChannelMonitor comments.

@codecov

codecovBot commented Aug 28, 2020

Copy link
Copy Markdown

Codecov Report

Merging #679 into master will decrease coverage by 0.01%.
The diff coverage is 98.70%.

Impacted file tree graph

@@ Coverage Diff @@## master #679 +/- ##
==========================================
- Coverage 91.92% 91.91% -0.02% 
==========================================
Files 35 35 Lines 20082 20154 +72 ==========================================
+ Hits 18461 18524 +63 - Misses 1621 1630 +9 
Impacted FilesCoverage Δ
lightning/src/ln/functional_tests.rs96.98% <98.63%> (-0.16%)⬇️
lightning/src/ln/channelmonitor.rs94.99% <100.00%> (ø)
lightning/src/ln/onchaintx.rs93.89% <100.00%> (-0.02%)⬇️
lightning/src/ln/channel.rs86.23% <0.00%> (+0.08%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 343aacc...6622ea7. Read the comment docs.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
mem::swap(&mut new_local_commitment_tx, &mut self.current_local_commitment_tx);
self.prev_local_signed_commitment_tx = Some(new_local_commitment_tx);
if self.local_tx_signed {
return Err(MonitorUpdateError("Latest local commitment signed has already been signed, update is rejected"));

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updating before returning an Err probably needs a documentation update to indicate to the user that they should still write out the new version to disk even in case of err (which is a super surprising API).

@ariardariardAug 28, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added a new comment a the top-level API : ManyChannelMonitor::update_monitor

I know this API is surprising, but you rarely have distributed system driven both by a central coordinator and a p2p network.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
if let Err(_) = self.onchain_tx_handler.provide_latest_local_tx(commitment_tx) {
return Err(MonitorUpdateError("Local commitment signed has already been signed, no further update of LOCAL commitment transaction is allowed"));
}
self.onchain_tx_handler.provide_latest_local_tx(commitment_tx);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we assert somewhere in here that prev_local_signed_commitment_tx must have been None before this? I suppose return Err and debug_assert!(false).

@ariardariardAug 28, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't explicitly prune out prev_local_signed_commitment=None.
We could add a round-trip between coordinator and watchtowers with a ChannelMonitorUpdateStep::PrunePrevLocalCommitment but that would be a latency hit for nothing, assuming we never return an Err if we broadcast a local commitment and coordinator fail advancing forward the channel by releasing the secret. Note the broadcast-commitment-then-reject-update is covered in new test (watchtower_alice failing the update_monitor).

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, right.

Comment threadlightning/src/ln/functional_tests.rs Outdated

#[test]
fn test_concurrent_monitor_claim() {
// Watchtower Alice receives block 135, broadcasts state X, rejects state Y.

@TheBlueMattTheBlueMattAug 28, 2020

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you make this comment more chronological? I think this is what threw me off the other day,

Watchtower A receives block, broadcasts state,
then channel receives new state Y, sending it to both watchtowers,
Bob accepts Y, then receives block and broadcasts the latest state Y,
Alice rejects state Y, but Bob has already broadcasted it,
Y confirms.

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't read through the test yet, but I get what you're going for more. A few relatively minor comments but conceptually looks good.

// OnchainTxHandler. After this is set, no future updates to our local commitment transactions
// may occur, and we fail any such monitor updates.
//
// In case of update rejection due to a locally already signed commitment transaction, we

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe:
Note that even when we fail a local commitment transaction update, we still store the update to ensure we can claim from it in case a duplicate copy of this ChannelMonitor broadcasts it.

/// Any spends of outputs which should have been registered which aren't passed to
/// ChannelMonitors via block_connected may result in FUNDS LOSS.
///
/// In case of distributed watchtowers deployment, even if an Err is return, the new version

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This needs more detail about the exact setup of a distributed ChannelMonitor, as several are possible (I think its a misnomer to call it a watchtower as that implies something specific about availability of keys which isn't true today), and probably this comment should be on ChannelMonitor::update_monitor, not the trait (though it could be referenced in the trait), because the trait's return value is only relevant for ChannelManager.

@ariardariardAug 31, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know writing a documentation for a distributed ChannelMonitor is tracked by #604. I intend to do it post-anchor as there are implications for bumping utxo management, namely you should also duplicate those keys across distributed monitors.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
///
/// Should also be used to indicate a failure to update the local persisted copy of the channel
/// monitor.
/// At reception of this error, channel manager MUST force-close the channel and return at

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it makes sense to describe this in these terms - what ChannelManager MUST do isn't interesting to our users, its what we do! Instead, we should describe it as ChannelManagerwill do X.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
/// This failure may also signal a failure to update the local persisted copy of one of
/// the channel monitor instance.
///
/// In case of a distributed watchtower deployment, failures sources can be divergent and

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I mean also because there's no use in any more granular errors :).

@ariard

Copy link
Copy Markdown
Author

Updated, for complete documentation, I will do it post-anchor/chaininterface refactoring.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
/// This failure may also signal a failure to update the local persisted copy of one of
/// the channel monitor instance.
///
/// Note that even when we fail a local commitment transaction update, we still store the

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be phrased as "you must" and probably be more clear (not "we" in the sense of "rust-lightning does"). Also we should distinguish between "ChannelMonitor returned" and "you are returning from a ManyChannelMonitor/chain::Watch", though if you want to just make @jkczyz deal with this in #649 thats ok too.

Antoine Riard added 3 commits September 15, 2020 18:17
With a distrbuted watchtowers deployment, where each monitor is plugged
to its own chain view, there is no guarantee that block are going to be
seen in same order. Watchtower may diverge in their acceptance of a
submitted `commitment_signed` update due to a block timing-out a HTLC
and provoking a subset but yet not seen by the other watchtower subset.
Any update reject by one of the watchtower must block offchain coordinator
to move channel state forward and release revocation secret for previous
state.
In this case, we want any watchtower from the rejection subset to still
be able to claim outputs if the concurrent state, has accepted by the
other subset, is confirming. This improve overall watchtower system
fault-tolerance.
This change stores local commitment transaction unconditionally and fail
the update if there is knowledge of an already signed commitment
transaction (ChannelMonitor.local_tx_signed=true).
Watchower Alice receives block 134, broadcasts state X, rejects state Y.
Watchtower Bob accepts state Y, receives blocks 135, broadcasts state Y.
State Y confirms onchain. Alice must be able to claim outputs.
Sources of the failure may be multiple in case of distributed watchtower
deployment. In either case, the channel manager must return a final
update asking to its channel monitor(s) to broadcast the lastest state
available. Revocation secret must not be released for the faultive
channel.
In the future, we may return wider type of failures to take more
fine-grained processing decision (e.g if local disk failure and
redudant remote channel copy available channel may still be processed
forward).
@TheBlueMatt
TheBlueMatt merged commit 44734be into lightningdevkit:masterSep 16, 2020
TheBlueMatt added a commit that referenced this pull request Sep 16, 2020
Implement concurrent broadcast tolerance for distributed watchtowers
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ariard@TheBlueMatt
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Implement concurrent broadcast tolerance for distributed watchtowers by ariard · Pull Request #679 · lightningdevkit/rust-lightning · GitHub
Skip to content

Implement concurrent broadcast tolerance for distributed watchtowers - #679

Merged
TheBlueMatt merged 3 commits into
lightningdevkit:masterfrom
ariard:2020-08-concurrent-watchtowers
Sep 16, 2020
Merged

Implement concurrent broadcast tolerance for distributed watchtowers#679
TheBlueMatt merged 3 commits into
lightningdevkit:masterfrom
ariard:2020-08-concurrent-watchtowers

Conversation

@ariard

Copy link
Copy Markdown

Implement concurrent broadcast tolerance for distributed watchtowers

With a distrbuted watchtowers deployment, where each monitor is plugged
to its own chain view, there is no guarantee that block are going to be
seen in same order. Watchtower may diverge in their acceptance of a
submitted commitment_signed update due to a block timing-out a HTLC
and provoking a subset but yet not seen by the other watchtower subset.
Any update reject by one of the watchtower must block offchain coordinator
to move channel state forward and release revocation secret for previous
state.

In this case, we want any watchtower from the rejection subset to still
be able to claim outputs if the concurrent state, has accepted by the
other subset, is confirming. This improve overall watchtower system
fault-tolerance.

This change stores local commitment transaction unconditionally and fail
the update if there is knowledge of an already signed commitment
transaction (ChannelMonitor.local_tx_signed=true).

Follow-up of #667, which let us easily implement concurrent tolerance and get ride off of confusing OnchainTxHandler/ChannelMonitor comments.

@codecov

codecovBot commented Aug 28, 2020

Copy link
Copy Markdown

Codecov Report

Merging #679 into master will decrease coverage by 0.01%.
The diff coverage is 98.70%.

Impacted file tree graph

@@ Coverage Diff @@## master #679 +/- ##
==========================================
- Coverage 91.92% 91.91% -0.02% 
==========================================
Files 35 35 Lines 20082 20154 +72 ==========================================
+ Hits 18461 18524 +63 - Misses 1621 1630 +9 
Impacted FilesCoverage Δ
lightning/src/ln/functional_tests.rs96.98% <98.63%> (-0.16%)⬇️
lightning/src/ln/channelmonitor.rs94.99% <100.00%> (ø)
lightning/src/ln/onchaintx.rs93.89% <100.00%> (-0.02%)⬇️
lightning/src/ln/channel.rs86.23% <0.00%> (+0.08%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 343aacc...6622ea7. Read the comment docs.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
mem::swap(&mut new_local_commitment_tx, &mut self.current_local_commitment_tx);
self.prev_local_signed_commitment_tx = Some(new_local_commitment_tx);
if self.local_tx_signed {
return Err(MonitorUpdateError("Latest local commitment signed has already been signed, update is rejected"));

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updating before returning an Err probably needs a documentation update to indicate to the user that they should still write out the new version to disk even in case of err (which is a super surprising API).

@ariardariardAug 28, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added a new comment a the top-level API : ManyChannelMonitor::update_monitor

I know this API is surprising, but you rarely have distributed system driven both by a central coordinator and a p2p network.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
if let Err(_) = self.onchain_tx_handler.provide_latest_local_tx(commitment_tx) {
return Err(MonitorUpdateError("Local commitment signed has already been signed, no further update of LOCAL commitment transaction is allowed"));
}
self.onchain_tx_handler.provide_latest_local_tx(commitment_tx);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we assert somewhere in here that prev_local_signed_commitment_tx must have been None before this? I suppose return Err and debug_assert!(false).

@ariardariardAug 28, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't explicitly prune out prev_local_signed_commitment=None.
We could add a round-trip between coordinator and watchtowers with a ChannelMonitorUpdateStep::PrunePrevLocalCommitment but that would be a latency hit for nothing, assuming we never return an Err if we broadcast a local commitment and coordinator fail advancing forward the channel by releasing the secret. Note the broadcast-commitment-then-reject-update is covered in new test (watchtower_alice failing the update_monitor).

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, right.

Comment threadlightning/src/ln/functional_tests.rs Outdated

#[test]
fn test_concurrent_monitor_claim() {
// Watchtower Alice receives block 135, broadcasts state X, rejects state Y.

@TheBlueMattTheBlueMattAug 28, 2020

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you make this comment more chronological? I think this is what threw me off the other day,

Watchtower A receives block, broadcasts state,
then channel receives new state Y, sending it to both watchtowers,
Bob accepts Y, then receives block and broadcasts the latest state Y,
Alice rejects state Y, but Bob has already broadcasted it,
Y confirms.

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't read through the test yet, but I get what you're going for more. A few relatively minor comments but conceptually looks good.

// OnchainTxHandler. After this is set, no future updates to our local commitment transactions
// may occur, and we fail any such monitor updates.
//
// In case of update rejection due to a locally already signed commitment transaction, we

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe:
Note that even when we fail a local commitment transaction update, we still store the update to ensure we can claim from it in case a duplicate copy of this ChannelMonitor broadcasts it.

/// Any spends of outputs which should have been registered which aren't passed to
/// ChannelMonitors via block_connected may result in FUNDS LOSS.
///
/// In case of distributed watchtowers deployment, even if an Err is return, the new version

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This needs more detail about the exact setup of a distributed ChannelMonitor, as several are possible (I think its a misnomer to call it a watchtower as that implies something specific about availability of keys which isn't true today), and probably this comment should be on ChannelMonitor::update_monitor, not the trait (though it could be referenced in the trait), because the trait's return value is only relevant for ChannelManager.

@ariardariardAug 31, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know writing a documentation for a distributed ChannelMonitor is tracked by #604. I intend to do it post-anchor as there are implications for bumping utxo management, namely you should also duplicate those keys across distributed monitors.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
///
/// Should also be used to indicate a failure to update the local persisted copy of the channel
/// monitor.
/// At reception of this error, channel manager MUST force-close the channel and return at

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it makes sense to describe this in these terms - what ChannelManager MUST do isn't interesting to our users, its what we do! Instead, we should describe it as ChannelManagerwill do X.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
/// This failure may also signal a failure to update the local persisted copy of one of
/// the channel monitor instance.
///
/// In case of a distributed watchtower deployment, failures sources can be divergent and

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I mean also because there's no use in any more granular errors :).

@ariard

Copy link
Copy Markdown
Author

Updated, for complete documentation, I will do it post-anchor/chaininterface refactoring.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
/// This failure may also signal a failure to update the local persisted copy of one of
/// the channel monitor instance.
///
/// Note that even when we fail a local commitment transaction update, we still store the

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be phrased as "you must" and probably be more clear (not "we" in the sense of "rust-lightning does"). Also we should distinguish between "ChannelMonitor returned" and "you are returning from a ManyChannelMonitor/chain::Watch", though if you want to just make @jkczyz deal with this in #649 thats ok too.

Antoine Riard added 3 commits September 15, 2020 18:17
With a distrbuted watchtowers deployment, where each monitor is plugged
to its own chain view, there is no guarantee that block are going to be
seen in same order. Watchtower may diverge in their acceptance of a
submitted `commitment_signed` update due to a block timing-out a HTLC
and provoking a subset but yet not seen by the other watchtower subset.
Any update reject by one of the watchtower must block offchain coordinator
to move channel state forward and release revocation secret for previous
state.
In this case, we want any watchtower from the rejection subset to still
be able to claim outputs if the concurrent state, has accepted by the
other subset, is confirming. This improve overall watchtower system
fault-tolerance.
This change stores local commitment transaction unconditionally and fail
the update if there is knowledge of an already signed commitment
transaction (ChannelMonitor.local_tx_signed=true).
Watchower Alice receives block 134, broadcasts state X, rejects state Y.
Watchtower Bob accepts state Y, receives blocks 135, broadcasts state Y.
State Y confirms onchain. Alice must be able to claim outputs.
Sources of the failure may be multiple in case of distributed watchtower
deployment. In either case, the channel manager must return a final
update asking to its channel monitor(s) to broadcast the lastest state
available. Revocation secret must not be released for the faultive
channel.
In the future, we may return wider type of failures to take more
fine-grained processing decision (e.g if local disk failure and
redudant remote channel copy available channel may still be processed
forward).
@TheBlueMatt
TheBlueMatt merged commit 44734be into lightningdevkit:masterSep 16, 2020
TheBlueMatt added a commit that referenced this pull request Sep 16, 2020
Implement concurrent broadcast tolerance for distributed watchtowers
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ariard@TheBlueMatt
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Implement concurrent broadcast tolerance for distributed watchtowers by ariard · Pull Request #679 · lightningdevkit/rust-lightning · GitHub
Skip to content

Implement concurrent broadcast tolerance for distributed watchtowers - #679

Merged
TheBlueMatt merged 3 commits into
lightningdevkit:masterfrom
ariard:2020-08-concurrent-watchtowers
Sep 16, 2020
Merged

Implement concurrent broadcast tolerance for distributed watchtowers#679
TheBlueMatt merged 3 commits into
lightningdevkit:masterfrom
ariard:2020-08-concurrent-watchtowers

Conversation

@ariard

Copy link
Copy Markdown

Implement concurrent broadcast tolerance for distributed watchtowers

With a distrbuted watchtowers deployment, where each monitor is plugged
to its own chain view, there is no guarantee that block are going to be
seen in same order. Watchtower may diverge in their acceptance of a
submitted commitment_signed update due to a block timing-out a HTLC
and provoking a subset but yet not seen by the other watchtower subset.
Any update reject by one of the watchtower must block offchain coordinator
to move channel state forward and release revocation secret for previous
state.

In this case, we want any watchtower from the rejection subset to still
be able to claim outputs if the concurrent state, has accepted by the
other subset, is confirming. This improve overall watchtower system
fault-tolerance.

This change stores local commitment transaction unconditionally and fail
the update if there is knowledge of an already signed commitment
transaction (ChannelMonitor.local_tx_signed=true).

Follow-up of #667, which let us easily implement concurrent tolerance and get ride off of confusing OnchainTxHandler/ChannelMonitor comments.

@codecov

codecovBot commented Aug 28, 2020

Copy link
Copy Markdown

Codecov Report

Merging #679 into master will decrease coverage by 0.01%.
The diff coverage is 98.70%.

Impacted file tree graph

@@ Coverage Diff @@## master #679 +/- ##
==========================================
- Coverage 91.92% 91.91% -0.02% 
==========================================
Files 35 35 Lines 20082 20154 +72 ==========================================
+ Hits 18461 18524 +63 - Misses 1621 1630 +9 
Impacted FilesCoverage Δ
lightning/src/ln/functional_tests.rs96.98% <98.63%> (-0.16%)⬇️
lightning/src/ln/channelmonitor.rs94.99% <100.00%> (ø)
lightning/src/ln/onchaintx.rs93.89% <100.00%> (-0.02%)⬇️
lightning/src/ln/channel.rs86.23% <0.00%> (+0.08%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 343aacc...6622ea7. Read the comment docs.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
mem::swap(&mut new_local_commitment_tx, &mut self.current_local_commitment_tx);
self.prev_local_signed_commitment_tx = Some(new_local_commitment_tx);
if self.local_tx_signed {
return Err(MonitorUpdateError("Latest local commitment signed has already been signed, update is rejected"));

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updating before returning an Err probably needs a documentation update to indicate to the user that they should still write out the new version to disk even in case of err (which is a super surprising API).

@ariardariardAug 28, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added a new comment a the top-level API : ManyChannelMonitor::update_monitor

I know this API is surprising, but you rarely have distributed system driven both by a central coordinator and a p2p network.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
if let Err(_) = self.onchain_tx_handler.provide_latest_local_tx(commitment_tx) {
return Err(MonitorUpdateError("Local commitment signed has already been signed, no further update of LOCAL commitment transaction is allowed"));
}
self.onchain_tx_handler.provide_latest_local_tx(commitment_tx);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we assert somewhere in here that prev_local_signed_commitment_tx must have been None before this? I suppose return Err and debug_assert!(false).

@ariardariardAug 28, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't explicitly prune out prev_local_signed_commitment=None.
We could add a round-trip between coordinator and watchtowers with a ChannelMonitorUpdateStep::PrunePrevLocalCommitment but that would be a latency hit for nothing, assuming we never return an Err if we broadcast a local commitment and coordinator fail advancing forward the channel by releasing the secret. Note the broadcast-commitment-then-reject-update is covered in new test (watchtower_alice failing the update_monitor).

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, right.

Comment threadlightning/src/ln/functional_tests.rs Outdated

#[test]
fn test_concurrent_monitor_claim() {
// Watchtower Alice receives block 135, broadcasts state X, rejects state Y.

@TheBlueMattTheBlueMattAug 28, 2020

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you make this comment more chronological? I think this is what threw me off the other day,

Watchtower A receives block, broadcasts state,
then channel receives new state Y, sending it to both watchtowers,
Bob accepts Y, then receives block and broadcasts the latest state Y,
Alice rejects state Y, but Bob has already broadcasted it,
Y confirms.

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't read through the test yet, but I get what you're going for more. A few relatively minor comments but conceptually looks good.

// OnchainTxHandler. After this is set, no future updates to our local commitment transactions
// may occur, and we fail any such monitor updates.
//
// In case of update rejection due to a locally already signed commitment transaction, we

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe:
Note that even when we fail a local commitment transaction update, we still store the update to ensure we can claim from it in case a duplicate copy of this ChannelMonitor broadcasts it.

/// Any spends of outputs which should have been registered which aren't passed to
/// ChannelMonitors via block_connected may result in FUNDS LOSS.
///
/// In case of distributed watchtowers deployment, even if an Err is return, the new version

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This needs more detail about the exact setup of a distributed ChannelMonitor, as several are possible (I think its a misnomer to call it a watchtower as that implies something specific about availability of keys which isn't true today), and probably this comment should be on ChannelMonitor::update_monitor, not the trait (though it could be referenced in the trait), because the trait's return value is only relevant for ChannelManager.

@ariardariardAug 31, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know writing a documentation for a distributed ChannelMonitor is tracked by #604. I intend to do it post-anchor as there are implications for bumping utxo management, namely you should also duplicate those keys across distributed monitors.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
///
/// Should also be used to indicate a failure to update the local persisted copy of the channel
/// monitor.
/// At reception of this error, channel manager MUST force-close the channel and return at

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it makes sense to describe this in these terms - what ChannelManager MUST do isn't interesting to our users, its what we do! Instead, we should describe it as ChannelManagerwill do X.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
/// This failure may also signal a failure to update the local persisted copy of one of
/// the channel monitor instance.
///
/// In case of a distributed watchtower deployment, failures sources can be divergent and

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I mean also because there's no use in any more granular errors :).

@ariard

Copy link
Copy Markdown
Author

Updated, for complete documentation, I will do it post-anchor/chaininterface refactoring.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
/// This failure may also signal a failure to update the local persisted copy of one of
/// the channel monitor instance.
///
/// Note that even when we fail a local commitment transaction update, we still store the

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be phrased as "you must" and probably be more clear (not "we" in the sense of "rust-lightning does"). Also we should distinguish between "ChannelMonitor returned" and "you are returning from a ManyChannelMonitor/chain::Watch", though if you want to just make @jkczyz deal with this in #649 thats ok too.

Antoine Riard added 3 commits September 15, 2020 18:17
With a distrbuted watchtowers deployment, where each monitor is plugged
to its own chain view, there is no guarantee that block are going to be
seen in same order. Watchtower may diverge in their acceptance of a
submitted `commitment_signed` update due to a block timing-out a HTLC
and provoking a subset but yet not seen by the other watchtower subset.
Any update reject by one of the watchtower must block offchain coordinator
to move channel state forward and release revocation secret for previous
state.
In this case, we want any watchtower from the rejection subset to still
be able to claim outputs if the concurrent state, has accepted by the
other subset, is confirming. This improve overall watchtower system
fault-tolerance.
This change stores local commitment transaction unconditionally and fail
the update if there is knowledge of an already signed commitment
transaction (ChannelMonitor.local_tx_signed=true).
Watchower Alice receives block 134, broadcasts state X, rejects state Y.
Watchtower Bob accepts state Y, receives blocks 135, broadcasts state Y.
State Y confirms onchain. Alice must be able to claim outputs.
Sources of the failure may be multiple in case of distributed watchtower
deployment. In either case, the channel manager must return a final
update asking to its channel monitor(s) to broadcast the lastest state
available. Revocation secret must not be released for the faultive
channel.
In the future, we may return wider type of failures to take more
fine-grained processing decision (e.g if local disk failure and
redudant remote channel copy available channel may still be processed
forward).
@TheBlueMatt
TheBlueMatt merged commit 44734be into lightningdevkit:masterSep 16, 2020
TheBlueMatt added a commit that referenced this pull request Sep 16, 2020
Implement concurrent broadcast tolerance for distributed watchtowers
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ariard@TheBlueMatt
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Implement concurrent broadcast tolerance for distributed watchtowers by ariard · Pull Request #679 · lightningdevkit/rust-lightning · GitHub
Skip to content

Implement concurrent broadcast tolerance for distributed watchtowers - #679

Merged
TheBlueMatt merged 3 commits into
lightningdevkit:masterfrom
ariard:2020-08-concurrent-watchtowers
Sep 16, 2020
Merged

Implement concurrent broadcast tolerance for distributed watchtowers#679
TheBlueMatt merged 3 commits into
lightningdevkit:masterfrom
ariard:2020-08-concurrent-watchtowers

Conversation

@ariard

Copy link
Copy Markdown

Implement concurrent broadcast tolerance for distributed watchtowers

With a distrbuted watchtowers deployment, where each monitor is plugged
to its own chain view, there is no guarantee that block are going to be
seen in same order. Watchtower may diverge in their acceptance of a
submitted commitment_signed update due to a block timing-out a HTLC
and provoking a subset but yet not seen by the other watchtower subset.
Any update reject by one of the watchtower must block offchain coordinator
to move channel state forward and release revocation secret for previous
state.

In this case, we want any watchtower from the rejection subset to still
be able to claim outputs if the concurrent state, has accepted by the
other subset, is confirming. This improve overall watchtower system
fault-tolerance.

This change stores local commitment transaction unconditionally and fail
the update if there is knowledge of an already signed commitment
transaction (ChannelMonitor.local_tx_signed=true).

Follow-up of #667, which let us easily implement concurrent tolerance and get ride off of confusing OnchainTxHandler/ChannelMonitor comments.

@codecov

codecovBot commented Aug 28, 2020

Copy link
Copy Markdown

Codecov Report

Merging #679 into master will decrease coverage by 0.01%.
The diff coverage is 98.70%.

Impacted file tree graph

@@ Coverage Diff @@## master #679 +/- ##
==========================================
- Coverage 91.92% 91.91% -0.02% 
==========================================
Files 35 35 Lines 20082 20154 +72 ==========================================
+ Hits 18461 18524 +63 - Misses 1621 1630 +9 
Impacted FilesCoverage Δ
lightning/src/ln/functional_tests.rs96.98% <98.63%> (-0.16%)⬇️
lightning/src/ln/channelmonitor.rs94.99% <100.00%> (ø)
lightning/src/ln/onchaintx.rs93.89% <100.00%> (-0.02%)⬇️
lightning/src/ln/channel.rs86.23% <0.00%> (+0.08%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 343aacc...6622ea7. Read the comment docs.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
mem::swap(&mut new_local_commitment_tx, &mut self.current_local_commitment_tx);
self.prev_local_signed_commitment_tx = Some(new_local_commitment_tx);
if self.local_tx_signed {
return Err(MonitorUpdateError("Latest local commitment signed has already been signed, update is rejected"));

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updating before returning an Err probably needs a documentation update to indicate to the user that they should still write out the new version to disk even in case of err (which is a super surprising API).

@ariardariardAug 28, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added a new comment a the top-level API : ManyChannelMonitor::update_monitor

I know this API is surprising, but you rarely have distributed system driven both by a central coordinator and a p2p network.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
if let Err(_) = self.onchain_tx_handler.provide_latest_local_tx(commitment_tx) {
return Err(MonitorUpdateError("Local commitment signed has already been signed, no further update of LOCAL commitment transaction is allowed"));
}
self.onchain_tx_handler.provide_latest_local_tx(commitment_tx);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we assert somewhere in here that prev_local_signed_commitment_tx must have been None before this? I suppose return Err and debug_assert!(false).

@ariardariardAug 28, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't explicitly prune out prev_local_signed_commitment=None.
We could add a round-trip between coordinator and watchtowers with a ChannelMonitorUpdateStep::PrunePrevLocalCommitment but that would be a latency hit for nothing, assuming we never return an Err if we broadcast a local commitment and coordinator fail advancing forward the channel by releasing the secret. Note the broadcast-commitment-then-reject-update is covered in new test (watchtower_alice failing the update_monitor).

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, right.

Comment threadlightning/src/ln/functional_tests.rs Outdated

#[test]
fn test_concurrent_monitor_claim() {
// Watchtower Alice receives block 135, broadcasts state X, rejects state Y.

@TheBlueMattTheBlueMattAug 28, 2020

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you make this comment more chronological? I think this is what threw me off the other day,

Watchtower A receives block, broadcasts state,
then channel receives new state Y, sending it to both watchtowers,
Bob accepts Y, then receives block and broadcasts the latest state Y,
Alice rejects state Y, but Bob has already broadcasted it,
Y confirms.

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't read through the test yet, but I get what you're going for more. A few relatively minor comments but conceptually looks good.

// OnchainTxHandler. After this is set, no future updates to our local commitment transactions
// may occur, and we fail any such monitor updates.
//
// In case of update rejection due to a locally already signed commitment transaction, we

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe:
Note that even when we fail a local commitment transaction update, we still store the update to ensure we can claim from it in case a duplicate copy of this ChannelMonitor broadcasts it.

/// Any spends of outputs which should have been registered which aren't passed to
/// ChannelMonitors via block_connected may result in FUNDS LOSS.
///
/// In case of distributed watchtowers deployment, even if an Err is return, the new version

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This needs more detail about the exact setup of a distributed ChannelMonitor, as several are possible (I think its a misnomer to call it a watchtower as that implies something specific about availability of keys which isn't true today), and probably this comment should be on ChannelMonitor::update_monitor, not the trait (though it could be referenced in the trait), because the trait's return value is only relevant for ChannelManager.

@ariardariardAug 31, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know writing a documentation for a distributed ChannelMonitor is tracked by #604. I intend to do it post-anchor as there are implications for bumping utxo management, namely you should also duplicate those keys across distributed monitors.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
///
/// Should also be used to indicate a failure to update the local persisted copy of the channel
/// monitor.
/// At reception of this error, channel manager MUST force-close the channel and return at

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it makes sense to describe this in these terms - what ChannelManager MUST do isn't interesting to our users, its what we do! Instead, we should describe it as ChannelManagerwill do X.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
/// This failure may also signal a failure to update the local persisted copy of one of
/// the channel monitor instance.
///
/// In case of a distributed watchtower deployment, failures sources can be divergent and

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I mean also because there's no use in any more granular errors :).

@ariard

Copy link
Copy Markdown
Author

Updated, for complete documentation, I will do it post-anchor/chaininterface refactoring.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
/// This failure may also signal a failure to update the local persisted copy of one of
/// the channel monitor instance.
///
/// Note that even when we fail a local commitment transaction update, we still store the

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be phrased as "you must" and probably be more clear (not "we" in the sense of "rust-lightning does"). Also we should distinguish between "ChannelMonitor returned" and "you are returning from a ManyChannelMonitor/chain::Watch", though if you want to just make @jkczyz deal with this in #649 thats ok too.

Antoine Riard added 3 commits September 15, 2020 18:17
With a distrbuted watchtowers deployment, where each monitor is plugged
to its own chain view, there is no guarantee that block are going to be
seen in same order. Watchtower may diverge in their acceptance of a
submitted `commitment_signed` update due to a block timing-out a HTLC
and provoking a subset but yet not seen by the other watchtower subset.
Any update reject by one of the watchtower must block offchain coordinator
to move channel state forward and release revocation secret for previous
state.
In this case, we want any watchtower from the rejection subset to still
be able to claim outputs if the concurrent state, has accepted by the
other subset, is confirming. This improve overall watchtower system
fault-tolerance.
This change stores local commitment transaction unconditionally and fail
the update if there is knowledge of an already signed commitment
transaction (ChannelMonitor.local_tx_signed=true).
Watchower Alice receives block 134, broadcasts state X, rejects state Y.
Watchtower Bob accepts state Y, receives blocks 135, broadcasts state Y.
State Y confirms onchain. Alice must be able to claim outputs.
Sources of the failure may be multiple in case of distributed watchtower
deployment. In either case, the channel manager must return a final
update asking to its channel monitor(s) to broadcast the lastest state
available. Revocation secret must not be released for the faultive
channel.
In the future, we may return wider type of failures to take more
fine-grained processing decision (e.g if local disk failure and
redudant remote channel copy available channel may still be processed
forward).
@TheBlueMatt
TheBlueMatt merged commit 44734be into lightningdevkit:masterSep 16, 2020
TheBlueMatt added a commit that referenced this pull request Sep 16, 2020
Implement concurrent broadcast tolerance for distributed watchtowers
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ariard@TheBlueMatt
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Implement concurrent broadcast tolerance for distributed watchtowers by ariard · Pull Request #679 · lightningdevkit/rust-lightning · GitHub
Skip to content

Implement concurrent broadcast tolerance for distributed watchtowers - #679

Merged
TheBlueMatt merged 3 commits into
lightningdevkit:masterfrom
ariard:2020-08-concurrent-watchtowers
Sep 16, 2020
Merged

Implement concurrent broadcast tolerance for distributed watchtowers#679
TheBlueMatt merged 3 commits into
lightningdevkit:masterfrom
ariard:2020-08-concurrent-watchtowers

Conversation

@ariard

Copy link
Copy Markdown

Implement concurrent broadcast tolerance for distributed watchtowers

With a distrbuted watchtowers deployment, where each monitor is plugged
to its own chain view, there is no guarantee that block are going to be
seen in same order. Watchtower may diverge in their acceptance of a
submitted commitment_signed update due to a block timing-out a HTLC
and provoking a subset but yet not seen by the other watchtower subset.
Any update reject by one of the watchtower must block offchain coordinator
to move channel state forward and release revocation secret for previous
state.

In this case, we want any watchtower from the rejection subset to still
be able to claim outputs if the concurrent state, has accepted by the
other subset, is confirming. This improve overall watchtower system
fault-tolerance.

This change stores local commitment transaction unconditionally and fail
the update if there is knowledge of an already signed commitment
transaction (ChannelMonitor.local_tx_signed=true).

Follow-up of #667, which let us easily implement concurrent tolerance and get ride off of confusing OnchainTxHandler/ChannelMonitor comments.

@codecov

codecovBot commented Aug 28, 2020

Copy link
Copy Markdown

Codecov Report

Merging #679 into master will decrease coverage by 0.01%.
The diff coverage is 98.70%.

Impacted file tree graph

@@ Coverage Diff @@## master #679 +/- ##
==========================================
- Coverage 91.92% 91.91% -0.02% 
==========================================
Files 35 35 Lines 20082 20154 +72 ==========================================
+ Hits 18461 18524 +63 - Misses 1621 1630 +9 
Impacted FilesCoverage Δ
lightning/src/ln/functional_tests.rs96.98% <98.63%> (-0.16%)⬇️
lightning/src/ln/channelmonitor.rs94.99% <100.00%> (ø)
lightning/src/ln/onchaintx.rs93.89% <100.00%> (-0.02%)⬇️
lightning/src/ln/channel.rs86.23% <0.00%> (+0.08%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 343aacc...6622ea7. Read the comment docs.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
mem::swap(&mut new_local_commitment_tx, &mut self.current_local_commitment_tx);
self.prev_local_signed_commitment_tx = Some(new_local_commitment_tx);
if self.local_tx_signed {
return Err(MonitorUpdateError("Latest local commitment signed has already been signed, update is rejected"));

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updating before returning an Err probably needs a documentation update to indicate to the user that they should still write out the new version to disk even in case of err (which is a super surprising API).

@ariardariardAug 28, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added a new comment a the top-level API : ManyChannelMonitor::update_monitor

I know this API is surprising, but you rarely have distributed system driven both by a central coordinator and a p2p network.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
if let Err(_) = self.onchain_tx_handler.provide_latest_local_tx(commitment_tx) {
return Err(MonitorUpdateError("Local commitment signed has already been signed, no further update of LOCAL commitment transaction is allowed"));
}
self.onchain_tx_handler.provide_latest_local_tx(commitment_tx);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we assert somewhere in here that prev_local_signed_commitment_tx must have been None before this? I suppose return Err and debug_assert!(false).

@ariardariardAug 28, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't explicitly prune out prev_local_signed_commitment=None.
We could add a round-trip between coordinator and watchtowers with a ChannelMonitorUpdateStep::PrunePrevLocalCommitment but that would be a latency hit for nothing, assuming we never return an Err if we broadcast a local commitment and coordinator fail advancing forward the channel by releasing the secret. Note the broadcast-commitment-then-reject-update is covered in new test (watchtower_alice failing the update_monitor).

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, right.

Comment threadlightning/src/ln/functional_tests.rs Outdated

#[test]
fn test_concurrent_monitor_claim() {
// Watchtower Alice receives block 135, broadcasts state X, rejects state Y.

@TheBlueMattTheBlueMattAug 28, 2020

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you make this comment more chronological? I think this is what threw me off the other day,

Watchtower A receives block, broadcasts state,
then channel receives new state Y, sending it to both watchtowers,
Bob accepts Y, then receives block and broadcasts the latest state Y,
Alice rejects state Y, but Bob has already broadcasted it,
Y confirms.

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't read through the test yet, but I get what you're going for more. A few relatively minor comments but conceptually looks good.

// OnchainTxHandler. After this is set, no future updates to our local commitment transactions
// may occur, and we fail any such monitor updates.
//
// In case of update rejection due to a locally already signed commitment transaction, we

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe:
Note that even when we fail a local commitment transaction update, we still store the update to ensure we can claim from it in case a duplicate copy of this ChannelMonitor broadcasts it.

/// Any spends of outputs which should have been registered which aren't passed to
/// ChannelMonitors via block_connected may result in FUNDS LOSS.
///
/// In case of distributed watchtowers deployment, even if an Err is return, the new version

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This needs more detail about the exact setup of a distributed ChannelMonitor, as several are possible (I think its a misnomer to call it a watchtower as that implies something specific about availability of keys which isn't true today), and probably this comment should be on ChannelMonitor::update_monitor, not the trait (though it could be referenced in the trait), because the trait's return value is only relevant for ChannelManager.

@ariardariardAug 31, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know writing a documentation for a distributed ChannelMonitor is tracked by #604. I intend to do it post-anchor as there are implications for bumping utxo management, namely you should also duplicate those keys across distributed monitors.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
///
/// Should also be used to indicate a failure to update the local persisted copy of the channel
/// monitor.
/// At reception of this error, channel manager MUST force-close the channel and return at

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it makes sense to describe this in these terms - what ChannelManager MUST do isn't interesting to our users, its what we do! Instead, we should describe it as ChannelManagerwill do X.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
/// This failure may also signal a failure to update the local persisted copy of one of
/// the channel monitor instance.
///
/// In case of a distributed watchtower deployment, failures sources can be divergent and

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I mean also because there's no use in any more granular errors :).

@ariard

Copy link
Copy Markdown
Author

Updated, for complete documentation, I will do it post-anchor/chaininterface refactoring.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
/// This failure may also signal a failure to update the local persisted copy of one of
/// the channel monitor instance.
///
/// Note that even when we fail a local commitment transaction update, we still store the

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be phrased as "you must" and probably be more clear (not "we" in the sense of "rust-lightning does"). Also we should distinguish between "ChannelMonitor returned" and "you are returning from a ManyChannelMonitor/chain::Watch", though if you want to just make @jkczyz deal with this in #649 thats ok too.

Antoine Riard added 3 commits September 15, 2020 18:17
With a distrbuted watchtowers deployment, where each monitor is plugged
to its own chain view, there is no guarantee that block are going to be
seen in same order. Watchtower may diverge in their acceptance of a
submitted `commitment_signed` update due to a block timing-out a HTLC
and provoking a subset but yet not seen by the other watchtower subset.
Any update reject by one of the watchtower must block offchain coordinator
to move channel state forward and release revocation secret for previous
state.
In this case, we want any watchtower from the rejection subset to still
be able to claim outputs if the concurrent state, has accepted by the
other subset, is confirming. This improve overall watchtower system
fault-tolerance.
This change stores local commitment transaction unconditionally and fail
the update if there is knowledge of an already signed commitment
transaction (ChannelMonitor.local_tx_signed=true).
Watchower Alice receives block 134, broadcasts state X, rejects state Y.
Watchtower Bob accepts state Y, receives blocks 135, broadcasts state Y.
State Y confirms onchain. Alice must be able to claim outputs.
Sources of the failure may be multiple in case of distributed watchtower
deployment. In either case, the channel manager must return a final
update asking to its channel monitor(s) to broadcast the lastest state
available. Revocation secret must not be released for the faultive
channel.
In the future, we may return wider type of failures to take more
fine-grained processing decision (e.g if local disk failure and
redudant remote channel copy available channel may still be processed
forward).
@TheBlueMatt
TheBlueMatt merged commit 44734be into lightningdevkit:masterSep 16, 2020
TheBlueMatt added a commit that referenced this pull request Sep 16, 2020
Implement concurrent broadcast tolerance for distributed watchtowers
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ariard@TheBlueMatt
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Implement concurrent broadcast tolerance for distributed watchtowers by ariard · Pull Request #679 · lightningdevkit/rust-lightning · GitHub
Skip to content

Implement concurrent broadcast tolerance for distributed watchtowers - #679

Merged
TheBlueMatt merged 3 commits into
lightningdevkit:masterfrom
ariard:2020-08-concurrent-watchtowers
Sep 16, 2020
Merged

Implement concurrent broadcast tolerance for distributed watchtowers#679
TheBlueMatt merged 3 commits into
lightningdevkit:masterfrom
ariard:2020-08-concurrent-watchtowers

Conversation

@ariard

Copy link
Copy Markdown

Implement concurrent broadcast tolerance for distributed watchtowers

With a distrbuted watchtowers deployment, where each monitor is plugged
to its own chain view, there is no guarantee that block are going to be
seen in same order. Watchtower may diverge in their acceptance of a
submitted commitment_signed update due to a block timing-out a HTLC
and provoking a subset but yet not seen by the other watchtower subset.
Any update reject by one of the watchtower must block offchain coordinator
to move channel state forward and release revocation secret for previous
state.

In this case, we want any watchtower from the rejection subset to still
be able to claim outputs if the concurrent state, has accepted by the
other subset, is confirming. This improve overall watchtower system
fault-tolerance.

This change stores local commitment transaction unconditionally and fail
the update if there is knowledge of an already signed commitment
transaction (ChannelMonitor.local_tx_signed=true).

Follow-up of #667, which let us easily implement concurrent tolerance and get ride off of confusing OnchainTxHandler/ChannelMonitor comments.

@codecov

codecovBot commented Aug 28, 2020

Copy link
Copy Markdown

Codecov Report

Merging #679 into master will decrease coverage by 0.01%.
The diff coverage is 98.70%.

Impacted file tree graph

@@ Coverage Diff @@## master #679 +/- ##
==========================================
- Coverage 91.92% 91.91% -0.02% 
==========================================
Files 35 35 Lines 20082 20154 +72 ==========================================
+ Hits 18461 18524 +63 - Misses 1621 1630 +9 
Impacted FilesCoverage Δ
lightning/src/ln/functional_tests.rs96.98% <98.63%> (-0.16%)⬇️
lightning/src/ln/channelmonitor.rs94.99% <100.00%> (ø)
lightning/src/ln/onchaintx.rs93.89% <100.00%> (-0.02%)⬇️
lightning/src/ln/channel.rs86.23% <0.00%> (+0.08%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 343aacc...6622ea7. Read the comment docs.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
mem::swap(&mut new_local_commitment_tx, &mut self.current_local_commitment_tx);
self.prev_local_signed_commitment_tx = Some(new_local_commitment_tx);
if self.local_tx_signed {
return Err(MonitorUpdateError("Latest local commitment signed has already been signed, update is rejected"));

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updating before returning an Err probably needs a documentation update to indicate to the user that they should still write out the new version to disk even in case of err (which is a super surprising API).

@ariardariardAug 28, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added a new comment a the top-level API : ManyChannelMonitor::update_monitor

I know this API is surprising, but you rarely have distributed system driven both by a central coordinator and a p2p network.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
if let Err(_) = self.onchain_tx_handler.provide_latest_local_tx(commitment_tx) {
return Err(MonitorUpdateError("Local commitment signed has already been signed, no further update of LOCAL commitment transaction is allowed"));
}
self.onchain_tx_handler.provide_latest_local_tx(commitment_tx);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we assert somewhere in here that prev_local_signed_commitment_tx must have been None before this? I suppose return Err and debug_assert!(false).

@ariardariardAug 28, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't explicitly prune out prev_local_signed_commitment=None.
We could add a round-trip between coordinator and watchtowers with a ChannelMonitorUpdateStep::PrunePrevLocalCommitment but that would be a latency hit for nothing, assuming we never return an Err if we broadcast a local commitment and coordinator fail advancing forward the channel by releasing the secret. Note the broadcast-commitment-then-reject-update is covered in new test (watchtower_alice failing the update_monitor).

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, right.

Comment threadlightning/src/ln/functional_tests.rs Outdated

#[test]
fn test_concurrent_monitor_claim() {
// Watchtower Alice receives block 135, broadcasts state X, rejects state Y.

@TheBlueMattTheBlueMattAug 28, 2020

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you make this comment more chronological? I think this is what threw me off the other day,

Watchtower A receives block, broadcasts state,
then channel receives new state Y, sending it to both watchtowers,
Bob accepts Y, then receives block and broadcasts the latest state Y,
Alice rejects state Y, but Bob has already broadcasted it,
Y confirms.

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't read through the test yet, but I get what you're going for more. A few relatively minor comments but conceptually looks good.

// OnchainTxHandler. After this is set, no future updates to our local commitment transactions
// may occur, and we fail any such monitor updates.
//
// In case of update rejection due to a locally already signed commitment transaction, we

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe:
Note that even when we fail a local commitment transaction update, we still store the update to ensure we can claim from it in case a duplicate copy of this ChannelMonitor broadcasts it.

/// Any spends of outputs which should have been registered which aren't passed to
/// ChannelMonitors via block_connected may result in FUNDS LOSS.
///
/// In case of distributed watchtowers deployment, even if an Err is return, the new version

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This needs more detail about the exact setup of a distributed ChannelMonitor, as several are possible (I think its a misnomer to call it a watchtower as that implies something specific about availability of keys which isn't true today), and probably this comment should be on ChannelMonitor::update_monitor, not the trait (though it could be referenced in the trait), because the trait's return value is only relevant for ChannelManager.

@ariardariardAug 31, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know writing a documentation for a distributed ChannelMonitor is tracked by #604. I intend to do it post-anchor as there are implications for bumping utxo management, namely you should also duplicate those keys across distributed monitors.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
///
/// Should also be used to indicate a failure to update the local persisted copy of the channel
/// monitor.
/// At reception of this error, channel manager MUST force-close the channel and return at

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it makes sense to describe this in these terms - what ChannelManager MUST do isn't interesting to our users, its what we do! Instead, we should describe it as ChannelManagerwill do X.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
/// This failure may also signal a failure to update the local persisted copy of one of
/// the channel monitor instance.
///
/// In case of a distributed watchtower deployment, failures sources can be divergent and

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I mean also because there's no use in any more granular errors :).

@ariard

Copy link
Copy Markdown
Author

Updated, for complete documentation, I will do it post-anchor/chaininterface refactoring.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
/// This failure may also signal a failure to update the local persisted copy of one of
/// the channel monitor instance.
///
/// Note that even when we fail a local commitment transaction update, we still store the

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be phrased as "you must" and probably be more clear (not "we" in the sense of "rust-lightning does"). Also we should distinguish between "ChannelMonitor returned" and "you are returning from a ManyChannelMonitor/chain::Watch", though if you want to just make @jkczyz deal with this in #649 thats ok too.

Antoine Riard added 3 commits September 15, 2020 18:17
With a distrbuted watchtowers deployment, where each monitor is plugged
to its own chain view, there is no guarantee that block are going to be
seen in same order. Watchtower may diverge in their acceptance of a
submitted `commitment_signed` update due to a block timing-out a HTLC
and provoking a subset but yet not seen by the other watchtower subset.
Any update reject by one of the watchtower must block offchain coordinator
to move channel state forward and release revocation secret for previous
state.
In this case, we want any watchtower from the rejection subset to still
be able to claim outputs if the concurrent state, has accepted by the
other subset, is confirming. This improve overall watchtower system
fault-tolerance.
This change stores local commitment transaction unconditionally and fail
the update if there is knowledge of an already signed commitment
transaction (ChannelMonitor.local_tx_signed=true).
Watchower Alice receives block 134, broadcasts state X, rejects state Y.
Watchtower Bob accepts state Y, receives blocks 135, broadcasts state Y.
State Y confirms onchain. Alice must be able to claim outputs.
Sources of the failure may be multiple in case of distributed watchtower
deployment. In either case, the channel manager must return a final
update asking to its channel monitor(s) to broadcast the lastest state
available. Revocation secret must not be released for the faultive
channel.
In the future, we may return wider type of failures to take more
fine-grained processing decision (e.g if local disk failure and
redudant remote channel copy available channel may still be processed
forward).
@TheBlueMatt
TheBlueMatt merged commit 44734be into lightningdevkit:masterSep 16, 2020
TheBlueMatt added a commit that referenced this pull request Sep 16, 2020
Implement concurrent broadcast tolerance for distributed watchtowers
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ariard@TheBlueMatt
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Implement concurrent broadcast tolerance for distributed watchtowers by ariard · Pull Request #679 · lightningdevkit/rust-lightning · GitHub
Skip to content

Implement concurrent broadcast tolerance for distributed watchtowers - #679

Merged
TheBlueMatt merged 3 commits into
lightningdevkit:masterfrom
ariard:2020-08-concurrent-watchtowers
Sep 16, 2020
Merged

Implement concurrent broadcast tolerance for distributed watchtowers#679
TheBlueMatt merged 3 commits into
lightningdevkit:masterfrom
ariard:2020-08-concurrent-watchtowers

Conversation

@ariard

Copy link
Copy Markdown

Implement concurrent broadcast tolerance for distributed watchtowers

With a distrbuted watchtowers deployment, where each monitor is plugged
to its own chain view, there is no guarantee that block are going to be
seen in same order. Watchtower may diverge in their acceptance of a
submitted commitment_signed update due to a block timing-out a HTLC
and provoking a subset but yet not seen by the other watchtower subset.
Any update reject by one of the watchtower must block offchain coordinator
to move channel state forward and release revocation secret for previous
state.

In this case, we want any watchtower from the rejection subset to still
be able to claim outputs if the concurrent state, has accepted by the
other subset, is confirming. This improve overall watchtower system
fault-tolerance.

This change stores local commitment transaction unconditionally and fail
the update if there is knowledge of an already signed commitment
transaction (ChannelMonitor.local_tx_signed=true).

Follow-up of #667, which let us easily implement concurrent tolerance and get ride off of confusing OnchainTxHandler/ChannelMonitor comments.

@codecov

codecovBot commented Aug 28, 2020

Copy link
Copy Markdown

Codecov Report

Merging #679 into master will decrease coverage by 0.01%.
The diff coverage is 98.70%.

Impacted file tree graph

@@ Coverage Diff @@## master #679 +/- ##
==========================================
- Coverage 91.92% 91.91% -0.02% 
==========================================
Files 35 35 Lines 20082 20154 +72 ==========================================
+ Hits 18461 18524 +63 - Misses 1621 1630 +9 
Impacted FilesCoverage Δ
lightning/src/ln/functional_tests.rs96.98% <98.63%> (-0.16%)⬇️
lightning/src/ln/channelmonitor.rs94.99% <100.00%> (ø)
lightning/src/ln/onchaintx.rs93.89% <100.00%> (-0.02%)⬇️
lightning/src/ln/channel.rs86.23% <0.00%> (+0.08%)⬆️

Continue to review full report at Codecov.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 343aacc...6622ea7. Read the comment docs.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
mem::swap(&mut new_local_commitment_tx, &mut self.current_local_commitment_tx);
self.prev_local_signed_commitment_tx = Some(new_local_commitment_tx);
if self.local_tx_signed {
return Err(MonitorUpdateError("Latest local commitment signed has already been signed, update is rejected"));

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updating before returning an Err probably needs a documentation update to indicate to the user that they should still write out the new version to disk even in case of err (which is a super surprising API).

@ariardariardAug 28, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added a new comment a the top-level API : ManyChannelMonitor::update_monitor

I know this API is surprising, but you rarely have distributed system driven both by a central coordinator and a p2p network.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
if let Err(_) = self.onchain_tx_handler.provide_latest_local_tx(commitment_tx) {
return Err(MonitorUpdateError("Local commitment signed has already been signed, no further update of LOCAL commitment transaction is allowed"));
}
self.onchain_tx_handler.provide_latest_local_tx(commitment_tx);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we assert somewhere in here that prev_local_signed_commitment_tx must have been None before this? I suppose return Err and debug_assert!(false).

@ariardariardAug 28, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't explicitly prune out prev_local_signed_commitment=None.
We could add a round-trip between coordinator and watchtowers with a ChannelMonitorUpdateStep::PrunePrevLocalCommitment but that would be a latency hit for nothing, assuming we never return an Err if we broadcast a local commitment and coordinator fail advancing forward the channel by releasing the secret. Note the broadcast-commitment-then-reject-update is covered in new test (watchtower_alice failing the update_monitor).

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, right.

Comment threadlightning/src/ln/functional_tests.rs Outdated

#[test]
fn test_concurrent_monitor_claim() {
// Watchtower Alice receives block 135, broadcasts state X, rejects state Y.

@TheBlueMattTheBlueMattAug 28, 2020

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you make this comment more chronological? I think this is what threw me off the other day,

Watchtower A receives block, broadcasts state,
then channel receives new state Y, sending it to both watchtowers,
Bob accepts Y, then receives block and broadcasts the latest state Y,
Alice rejects state Y, but Bob has already broadcasted it,
Y confirms.

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't read through the test yet, but I get what you're going for more. A few relatively minor comments but conceptually looks good.

// OnchainTxHandler. After this is set, no future updates to our local commitment transactions
// may occur, and we fail any such monitor updates.
//
// In case of update rejection due to a locally already signed commitment transaction, we

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe:
Note that even when we fail a local commitment transaction update, we still store the update to ensure we can claim from it in case a duplicate copy of this ChannelMonitor broadcasts it.

/// Any spends of outputs which should have been registered which aren't passed to
/// ChannelMonitors via block_connected may result in FUNDS LOSS.
///
/// In case of distributed watchtowers deployment, even if an Err is return, the new version

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This needs more detail about the exact setup of a distributed ChannelMonitor, as several are possible (I think its a misnomer to call it a watchtower as that implies something specific about availability of keys which isn't true today), and probably this comment should be on ChannelMonitor::update_monitor, not the trait (though it could be referenced in the trait), because the trait's return value is only relevant for ChannelManager.

@ariardariardAug 31, 2020

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know writing a documentation for a distributed ChannelMonitor is tracked by #604. I intend to do it post-anchor as there are implications for bumping utxo management, namely you should also duplicate those keys across distributed monitors.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
///
/// Should also be used to indicate a failure to update the local persisted copy of the channel
/// monitor.
/// At reception of this error, channel manager MUST force-close the channel and return at

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it makes sense to describe this in these terms - what ChannelManager MUST do isn't interesting to our users, its what we do! Instead, we should describe it as ChannelManagerwill do X.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
/// This failure may also signal a failure to update the local persisted copy of one of
/// the channel monitor instance.
///
/// In case of a distributed watchtower deployment, failures sources can be divergent and

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I mean also because there's no use in any more granular errors :).

@ariard

Copy link
Copy Markdown
Author

Updated, for complete documentation, I will do it post-anchor/chaininterface refactoring.

Comment threadlightning/src/ln/channelmonitor.rs Outdated
/// This failure may also signal a failure to update the local persisted copy of one of
/// the channel monitor instance.
///
/// Note that even when we fail a local commitment transaction update, we still store the

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should be phrased as "you must" and probably be more clear (not "we" in the sense of "rust-lightning does"). Also we should distinguish between "ChannelMonitor returned" and "you are returning from a ManyChannelMonitor/chain::Watch", though if you want to just make @jkczyz deal with this in #649 thats ok too.

Antoine Riard added 3 commits September 15, 2020 18:17
With a distrbuted watchtowers deployment, where each monitor is plugged
to its own chain view, there is no guarantee that block are going to be
seen in same order. Watchtower may diverge in their acceptance of a
submitted `commitment_signed` update due to a block timing-out a HTLC
and provoking a subset but yet not seen by the other watchtower subset.
Any update reject by one of the watchtower must block offchain coordinator
to move channel state forward and release revocation secret for previous
state.
In this case, we want any watchtower from the rejection subset to still
be able to claim outputs if the concurrent state, has accepted by the
other subset, is confirming. This improve overall watchtower system
fault-tolerance.
This change stores local commitment transaction unconditionally and fail
the update if there is knowledge of an already signed commitment
transaction (ChannelMonitor.local_tx_signed=true).
Watchower Alice receives block 134, broadcasts state X, rejects state Y.
Watchtower Bob accepts state Y, receives blocks 135, broadcasts state Y.
State Y confirms onchain. Alice must be able to claim outputs.
Sources of the failure may be multiple in case of distributed watchtower
deployment. In either case, the channel manager must return a final
update asking to its channel monitor(s) to broadcast the lastest state
available. Revocation secret must not be released for the faultive
channel.
In the future, we may return wider type of failures to take more
fine-grained processing decision (e.g if local disk failure and
redudant remote channel copy available channel may still be processed
forward).
@TheBlueMatt
TheBlueMatt merged commit 44734be into lightningdevkit:masterSep 16, 2020
TheBlueMatt added a commit that referenced this pull request Sep 16, 2020
Implement concurrent broadcast tolerance for distributed watchtowers
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
TheBlueMatt added a commit to TheBlueMatt/rust-lightning that referenced this pull request Sep 16, 2020
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@ariard@TheBlueMatt