ensure peer_connected is called before peer_disconnected - #3110

Closed
johncantrell97 wants to merge 2 commits into
lightningdevkit:mainfrom
johncantrell97:2024-06-robust-peer-disconnected-events
Closed

ensure peer_connected is called before peer_disconnected#3110
johncantrell97 wants to merge 2 commits into
lightningdevkit:mainfrom
johncantrell97:2024-06-robust-peer-disconnected-events

Conversation

@johncantrell97

Copy link
Copy Markdown
Contributor

Fixes#3108

Makes sure all message handler's peer_connected methods are called instead of returning early on the first to error.

As for whether or not the user has to call back into socket_disconnected after a PeerManager::read_event, I assume you mean after it returns an Err? I think the user does not have to because read_event will call disconnect_event_internal on any error before returning it to the user.

I took a look at lightning-net-tokio and it appears to be the case over there as well. It does:

if let Disconnect::PeerDisconnected = disconnect_type {
peer_manager.as_ref().socket_disconnected(&our_descriptor);
peer_manager.as_ref().process_events();
}

Only calling socket_disconnected if the disconnection type is one the user detected. If read_event returns an Err it breaks with a disconnection type of Disconnect::CloseConnection and does not call back into socket_disconnected.

Matt seems to think you do have to so I'm probably misunderstanding the original question. Happy to dig into it a bit more with some clarification if I misunderstood.

("Route Handler", self.message_handler.route_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
("Channel Handler", self.message_handler.chan_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
("Onion Handler", self.message_handler.onion_message_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
];

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you don't like this attempt to dry up the handling then I'm find just having separate results where I check them one by one with their own log messages.

@codecov-commenter

codecov-commenter commented Jun 7, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 92.56757% with 11 lines in your changes missing coverage. Please review.

Project coverage is 90.78%. Comparing base (9789152) to head (922c31f).
Report is 1274 commits behind head on main.

Files with missing linesPatch %Lines
lightning/src/ln/peer_handler.rs92.56%10 Missing and 1 partial ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #3110 +/- ##
==========================================
+ Coverage 89.84% 90.78% +0.93% 
==========================================
Files 119 119 Lines 97561 103463 +5902 Branches 97561 103463 +5902 ==========================================
+ Hits 87655 93925 +6270 + Misses 7331 7032 -299 + Partials 2575 2506 -69 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! I think we also need to move the peer_lock.their_features call up - we only call peer_disconnected if that line has been hit (Peer::handshake_complete checks for it) so we want to always hit that immediately before we call peer_connecteds.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from 4a1cade to db3b148CompareJune 7, 2024 16:02
@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

Thanks! I think we also need to move the peer_lock.their_features call up - we only call peer_disconnected if that line has been hit (Peer::handshake_complete checks for it) so we want to always hit that immediately before we call peer_connecteds.

Whoops, fixed it.

Can't be before calls to peer_connected because they pass a reference to msg but as long as we do it before returning it should be okay.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from db3b148 to 3a8f3b2CompareJune 7, 2024 16:18
TheBlueMatt
TheBlueMatt previously approved these changes Jun 7, 2024

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be nice to get a test (which should be pretty easy), but either way LGTM.

@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

Would be nice to get a test (which should be pretty easy), but either way LGTM.

Doesn't look like there's easy way to handle testing it with the existing test message handlers. Should I create new ones that can error on peer_connected and track connected/disconnected have been called or add the functionality to the existing test handlers?

Have used something like mockall for this in the past but without it I'd have to add counters/flags that get updated and a way to check them.

Is this what you had in mind for being able to test it?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

Yea, I was figuring you'd just create a trivial CustomMessageHandler that asserts that connected/disconnecteds all come in order and then errors on connection.

@johncantrell97

johncantrell97 commented Jun 7, 2024

Copy link
Copy Markdown
ContributorAuthor

Yea, I was figuring you'd just create a trivial CustomMessageHandler that asserts that connected/disconnecteds all come in order and then errors on connection.

Hm, using a CustomMessageHandler doesn't really test the fix here since it goes last. One of the issues was the early return causing the later handlers to not get the peer_connected at all. Would have to use multiple new handlers to be able to check the one after an error is returned still gets peer_connected called on it (and disconnected)

I guess at least it would catch the fix for ensuring disconnect is called.

@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

@TheBlueMatt

Added a test that passes but it duplicates a ton of code to handle all of the setup but with the new message handlers :|

not sure if this is okay, looking for feedback on the test and how to do it better if it's not okay.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch 2 times, most recently from 2c4c40a to 7a29c39CompareJune 7, 2024 23:00
@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from 7a29c39 to 922c31fCompareJune 8, 2024 01:22
if let Err(()) = self.message_handler.custom_message_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection) {
log_debug!(logger, "Custom Message Handler decided we couldn't communicate with peer {}", log_pubkey!(their_node_id));

peer_lock.their_features = Some(msg.features);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a bit confused by this, wouldn't that lead to use falsely assuming the handshake succeeded even though one of our handlers rejected it? And there is a window between us dropping the lock and handling the disconnect even where we would deal with it in a 'normal' manner, e.g., accepting further messages, and potentially rebroadcasting etc?

(cc @TheBlueMatt as he requested this change)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hm, if that's true then seems like we'll need to separate "handshake_completed" from "triggered peer_connected" with a new flag on the peer that we can use to decide whether or not to trigger peer_disconnected in do_disconnect and disconnect_event_internal?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yea, kinda, we'll end up forwarding broadcasts, as you point out, which is maybe not ideal, but we shouldn't process any further messages - we're currently in a read processing call, and we require read processing calls for any given peer to be serial, so presumably when we return an error the read-processing pipeline for this peer will stall and we won't get any more reads. We could make that explicit in the docs, however.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mhh, rather than introducing this race-y behavior in the first place, couldn't we just introduce a new handshake_aborted flag and check that alternatively to !peer.handshake_complete in disconnect_event_internal?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's fine too.

use crate::ln::msgs::{Init, LightningError, SocketAddress};
use crate::util::test_utils;


Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: Drop superfluous whitespace.

}
}

struct TestPeerTrackingMessageHandler {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, I believe the alternative would be to add TestCustomMessageHandler and TestOnionMessageHandler to test_utils and use them as part of the default test setup?

@tnull

Copy link
Copy Markdown
Contributor

@johncantrell97 Any interest in finishing this PR?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

Supersceded by #3580.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

peer_disconnected event processing isn't robust if a peer_connected Errs

4 participants

@johncantrell97@codecov-commenter@TheBlueMatt@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

ensure peer_connected is called before peer_disconnected - #3110

Closed
johncantrell97 wants to merge 2 commits into
lightningdevkit:mainfrom
johncantrell97:2024-06-robust-peer-disconnected-events
Closed

ensure peer_connected is called before peer_disconnected#3110
johncantrell97 wants to merge 2 commits into
lightningdevkit:mainfrom
johncantrell97:2024-06-robust-peer-disconnected-events

Conversation

@johncantrell97

Copy link
Copy Markdown
Contributor

Fixes#3108

Makes sure all message handler's peer_connected methods are called instead of returning early on the first to error.

As for whether or not the user has to call back into socket_disconnected after a PeerManager::read_event, I assume you mean after it returns an Err? I think the user does not have to because read_event will call disconnect_event_internal on any error before returning it to the user.

I took a look at lightning-net-tokio and it appears to be the case over there as well. It does:

if let Disconnect::PeerDisconnected = disconnect_type {
peer_manager.as_ref().socket_disconnected(&our_descriptor);
peer_manager.as_ref().process_events();
}

Only calling socket_disconnected if the disconnection type is one the user detected. If read_event returns an Err it breaks with a disconnection type of Disconnect::CloseConnection and does not call back into socket_disconnected.

Matt seems to think you do have to so I'm probably misunderstanding the original question. Happy to dig into it a bit more with some clarification if I misunderstood.

("Route Handler", self.message_handler.route_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
("Channel Handler", self.message_handler.chan_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
("Onion Handler", self.message_handler.onion_message_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
];

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you don't like this attempt to dry up the handling then I'm find just having separate results where I check them one by one with their own log messages.

@codecov-commenter

codecov-commenter commented Jun 7, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 92.56757% with 11 lines in your changes missing coverage. Please review.

Project coverage is 90.78%. Comparing base (9789152) to head (922c31f).
Report is 1274 commits behind head on main.

Files with missing linesPatch %Lines
lightning/src/ln/peer_handler.rs92.56%10 Missing and 1 partial ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #3110 +/- ##
==========================================
+ Coverage 89.84% 90.78% +0.93% 
==========================================
Files 119 119 Lines 97561 103463 +5902 Branches 97561 103463 +5902 ==========================================
+ Hits 87655 93925 +6270 + Misses 7331 7032 -299 + Partials 2575 2506 -69 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! I think we also need to move the peer_lock.their_features call up - we only call peer_disconnected if that line has been hit (Peer::handshake_complete checks for it) so we want to always hit that immediately before we call peer_connecteds.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from 4a1cade to db3b148CompareJune 7, 2024 16:02
@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

Thanks! I think we also need to move the peer_lock.their_features call up - we only call peer_disconnected if that line has been hit (Peer::handshake_complete checks for it) so we want to always hit that immediately before we call peer_connecteds.

Whoops, fixed it.

Can't be before calls to peer_connected because they pass a reference to msg but as long as we do it before returning it should be okay.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from db3b148 to 3a8f3b2CompareJune 7, 2024 16:18
TheBlueMatt
TheBlueMatt previously approved these changes Jun 7, 2024

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be nice to get a test (which should be pretty easy), but either way LGTM.

@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

Would be nice to get a test (which should be pretty easy), but either way LGTM.

Doesn't look like there's easy way to handle testing it with the existing test message handlers. Should I create new ones that can error on peer_connected and track connected/disconnected have been called or add the functionality to the existing test handlers?

Have used something like mockall for this in the past but without it I'd have to add counters/flags that get updated and a way to check them.

Is this what you had in mind for being able to test it?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

Yea, I was figuring you'd just create a trivial CustomMessageHandler that asserts that connected/disconnecteds all come in order and then errors on connection.

@johncantrell97

johncantrell97 commented Jun 7, 2024

Copy link
Copy Markdown
ContributorAuthor

Yea, I was figuring you'd just create a trivial CustomMessageHandler that asserts that connected/disconnecteds all come in order and then errors on connection.

Hm, using a CustomMessageHandler doesn't really test the fix here since it goes last. One of the issues was the early return causing the later handlers to not get the peer_connected at all. Would have to use multiple new handlers to be able to check the one after an error is returned still gets peer_connected called on it (and disconnected)

I guess at least it would catch the fix for ensuring disconnect is called.

@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

@TheBlueMatt

Added a test that passes but it duplicates a ton of code to handle all of the setup but with the new message handlers :|

not sure if this is okay, looking for feedback on the test and how to do it better if it's not okay.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch 2 times, most recently from 2c4c40a to 7a29c39CompareJune 7, 2024 23:00
@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from 7a29c39 to 922c31fCompareJune 8, 2024 01:22
if let Err(()) = self.message_handler.custom_message_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection) {
log_debug!(logger, "Custom Message Handler decided we couldn't communicate with peer {}", log_pubkey!(their_node_id));

peer_lock.their_features = Some(msg.features);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a bit confused by this, wouldn't that lead to use falsely assuming the handshake succeeded even though one of our handlers rejected it? And there is a window between us dropping the lock and handling the disconnect even where we would deal with it in a 'normal' manner, e.g., accepting further messages, and potentially rebroadcasting etc?

(cc @TheBlueMatt as he requested this change)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hm, if that's true then seems like we'll need to separate "handshake_completed" from "triggered peer_connected" with a new flag on the peer that we can use to decide whether or not to trigger peer_disconnected in do_disconnect and disconnect_event_internal?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yea, kinda, we'll end up forwarding broadcasts, as you point out, which is maybe not ideal, but we shouldn't process any further messages - we're currently in a read processing call, and we require read processing calls for any given peer to be serial, so presumably when we return an error the read-processing pipeline for this peer will stall and we won't get any more reads. We could make that explicit in the docs, however.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mhh, rather than introducing this race-y behavior in the first place, couldn't we just introduce a new handshake_aborted flag and check that alternatively to !peer.handshake_complete in disconnect_event_internal?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's fine too.

use crate::ln::msgs::{Init, LightningError, SocketAddress};
use crate::util::test_utils;


Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: Drop superfluous whitespace.

}
}

struct TestPeerTrackingMessageHandler {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, I believe the alternative would be to add TestCustomMessageHandler and TestOnionMessageHandler to test_utils and use them as part of the default test setup?

@tnull

Copy link
Copy Markdown
Contributor

@johncantrell97 Any interest in finishing this PR?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

Supersceded by #3580.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

peer_disconnected event processing isn't robust if a peer_connected Errs

4 participants

@johncantrell97@codecov-commenter@TheBlueMatt@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ensure peer_connected is called before peer_disconnected - #3110

Closed
johncantrell97 wants to merge 2 commits into
lightningdevkit:mainfrom
johncantrell97:2024-06-robust-peer-disconnected-events
Closed

ensure peer_connected is called before peer_disconnected#3110
johncantrell97 wants to merge 2 commits into
lightningdevkit:mainfrom
johncantrell97:2024-06-robust-peer-disconnected-events

Conversation

@johncantrell97

Copy link
Copy Markdown
Contributor

Fixes#3108

Makes sure all message handler's peer_connected methods are called instead of returning early on the first to error.

As for whether or not the user has to call back into socket_disconnected after a PeerManager::read_event, I assume you mean after it returns an Err? I think the user does not have to because read_event will call disconnect_event_internal on any error before returning it to the user.

I took a look at lightning-net-tokio and it appears to be the case over there as well. It does:

if let Disconnect::PeerDisconnected = disconnect_type {
peer_manager.as_ref().socket_disconnected(&our_descriptor);
peer_manager.as_ref().process_events();
}

Only calling socket_disconnected if the disconnection type is one the user detected. If read_event returns an Err it breaks with a disconnection type of Disconnect::CloseConnection and does not call back into socket_disconnected.

Matt seems to think you do have to so I'm probably misunderstanding the original question. Happy to dig into it a bit more with some clarification if I misunderstood.

("Route Handler", self.message_handler.route_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
("Channel Handler", self.message_handler.chan_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
("Onion Handler", self.message_handler.onion_message_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
];

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you don't like this attempt to dry up the handling then I'm find just having separate results where I check them one by one with their own log messages.

@codecov-commenter

codecov-commenter commented Jun 7, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 92.56757% with 11 lines in your changes missing coverage. Please review.

Project coverage is 90.78%. Comparing base (9789152) to head (922c31f).
Report is 1274 commits behind head on main.

Files with missing linesPatch %Lines
lightning/src/ln/peer_handler.rs92.56%10 Missing and 1 partial ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #3110 +/- ##
==========================================
+ Coverage 89.84% 90.78% +0.93% 
==========================================
Files 119 119 Lines 97561 103463 +5902 Branches 97561 103463 +5902 ==========================================
+ Hits 87655 93925 +6270 + Misses 7331 7032 -299 + Partials 2575 2506 -69 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! I think we also need to move the peer_lock.their_features call up - we only call peer_disconnected if that line has been hit (Peer::handshake_complete checks for it) so we want to always hit that immediately before we call peer_connecteds.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from 4a1cade to db3b148CompareJune 7, 2024 16:02
@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

Thanks! I think we also need to move the peer_lock.their_features call up - we only call peer_disconnected if that line has been hit (Peer::handshake_complete checks for it) so we want to always hit that immediately before we call peer_connecteds.

Whoops, fixed it.

Can't be before calls to peer_connected because they pass a reference to msg but as long as we do it before returning it should be okay.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from db3b148 to 3a8f3b2CompareJune 7, 2024 16:18
TheBlueMatt
TheBlueMatt previously approved these changes Jun 7, 2024

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be nice to get a test (which should be pretty easy), but either way LGTM.

@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

Would be nice to get a test (which should be pretty easy), but either way LGTM.

Doesn't look like there's easy way to handle testing it with the existing test message handlers. Should I create new ones that can error on peer_connected and track connected/disconnected have been called or add the functionality to the existing test handlers?

Have used something like mockall for this in the past but without it I'd have to add counters/flags that get updated and a way to check them.

Is this what you had in mind for being able to test it?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

Yea, I was figuring you'd just create a trivial CustomMessageHandler that asserts that connected/disconnecteds all come in order and then errors on connection.

@johncantrell97

johncantrell97 commented Jun 7, 2024

Copy link
Copy Markdown
ContributorAuthor

Yea, I was figuring you'd just create a trivial CustomMessageHandler that asserts that connected/disconnecteds all come in order and then errors on connection.

Hm, using a CustomMessageHandler doesn't really test the fix here since it goes last. One of the issues was the early return causing the later handlers to not get the peer_connected at all. Would have to use multiple new handlers to be able to check the one after an error is returned still gets peer_connected called on it (and disconnected)

I guess at least it would catch the fix for ensuring disconnect is called.

@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

@TheBlueMatt

Added a test that passes but it duplicates a ton of code to handle all of the setup but with the new message handlers :|

not sure if this is okay, looking for feedback on the test and how to do it better if it's not okay.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch 2 times, most recently from 2c4c40a to 7a29c39CompareJune 7, 2024 23:00
@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from 7a29c39 to 922c31fCompareJune 8, 2024 01:22
if let Err(()) = self.message_handler.custom_message_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection) {
log_debug!(logger, "Custom Message Handler decided we couldn't communicate with peer {}", log_pubkey!(their_node_id));

peer_lock.their_features = Some(msg.features);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a bit confused by this, wouldn't that lead to use falsely assuming the handshake succeeded even though one of our handlers rejected it? And there is a window between us dropping the lock and handling the disconnect even where we would deal with it in a 'normal' manner, e.g., accepting further messages, and potentially rebroadcasting etc?

(cc @TheBlueMatt as he requested this change)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hm, if that's true then seems like we'll need to separate "handshake_completed" from "triggered peer_connected" with a new flag on the peer that we can use to decide whether or not to trigger peer_disconnected in do_disconnect and disconnect_event_internal?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yea, kinda, we'll end up forwarding broadcasts, as you point out, which is maybe not ideal, but we shouldn't process any further messages - we're currently in a read processing call, and we require read processing calls for any given peer to be serial, so presumably when we return an error the read-processing pipeline for this peer will stall and we won't get any more reads. We could make that explicit in the docs, however.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mhh, rather than introducing this race-y behavior in the first place, couldn't we just introduce a new handshake_aborted flag and check that alternatively to !peer.handshake_complete in disconnect_event_internal?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's fine too.

use crate::ln::msgs::{Init, LightningError, SocketAddress};
use crate::util::test_utils;


Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: Drop superfluous whitespace.

}
}

struct TestPeerTrackingMessageHandler {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, I believe the alternative would be to add TestCustomMessageHandler and TestOnionMessageHandler to test_utils and use them as part of the default test setup?

@tnull

Copy link
Copy Markdown
Contributor

@johncantrell97 Any interest in finishing this PR?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

Supersceded by #3580.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

peer_disconnected event processing isn't robust if a peer_connected Errs

4 participants

@johncantrell97@codecov-commenter@TheBlueMatt@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ensure peer_connected is called before peer_disconnected - #3110

Closed
johncantrell97 wants to merge 2 commits into
lightningdevkit:mainfrom
johncantrell97:2024-06-robust-peer-disconnected-events
Closed

ensure peer_connected is called before peer_disconnected#3110
johncantrell97 wants to merge 2 commits into
lightningdevkit:mainfrom
johncantrell97:2024-06-robust-peer-disconnected-events

Conversation

@johncantrell97

Copy link
Copy Markdown
Contributor

Fixes#3108

Makes sure all message handler's peer_connected methods are called instead of returning early on the first to error.

As for whether or not the user has to call back into socket_disconnected after a PeerManager::read_event, I assume you mean after it returns an Err? I think the user does not have to because read_event will call disconnect_event_internal on any error before returning it to the user.

I took a look at lightning-net-tokio and it appears to be the case over there as well. It does:

if let Disconnect::PeerDisconnected = disconnect_type {
peer_manager.as_ref().socket_disconnected(&our_descriptor);
peer_manager.as_ref().process_events();
}

Only calling socket_disconnected if the disconnection type is one the user detected. If read_event returns an Err it breaks with a disconnection type of Disconnect::CloseConnection and does not call back into socket_disconnected.

Matt seems to think you do have to so I'm probably misunderstanding the original question. Happy to dig into it a bit more with some clarification if I misunderstood.

("Route Handler", self.message_handler.route_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
("Channel Handler", self.message_handler.chan_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
("Onion Handler", self.message_handler.onion_message_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
];

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you don't like this attempt to dry up the handling then I'm find just having separate results where I check them one by one with their own log messages.

@codecov-commenter

codecov-commenter commented Jun 7, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 92.56757% with 11 lines in your changes missing coverage. Please review.

Project coverage is 90.78%. Comparing base (9789152) to head (922c31f).
Report is 1274 commits behind head on main.

Files with missing linesPatch %Lines
lightning/src/ln/peer_handler.rs92.56%10 Missing and 1 partial ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #3110 +/- ##
==========================================
+ Coverage 89.84% 90.78% +0.93% 
==========================================
Files 119 119 Lines 97561 103463 +5902 Branches 97561 103463 +5902 ==========================================
+ Hits 87655 93925 +6270 + Misses 7331 7032 -299 + Partials 2575 2506 -69 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! I think we also need to move the peer_lock.their_features call up - we only call peer_disconnected if that line has been hit (Peer::handshake_complete checks for it) so we want to always hit that immediately before we call peer_connecteds.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from 4a1cade to db3b148CompareJune 7, 2024 16:02
@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

Thanks! I think we also need to move the peer_lock.their_features call up - we only call peer_disconnected if that line has been hit (Peer::handshake_complete checks for it) so we want to always hit that immediately before we call peer_connecteds.

Whoops, fixed it.

Can't be before calls to peer_connected because they pass a reference to msg but as long as we do it before returning it should be okay.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from db3b148 to 3a8f3b2CompareJune 7, 2024 16:18
TheBlueMatt
TheBlueMatt previously approved these changes Jun 7, 2024

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be nice to get a test (which should be pretty easy), but either way LGTM.

@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

Would be nice to get a test (which should be pretty easy), but either way LGTM.

Doesn't look like there's easy way to handle testing it with the existing test message handlers. Should I create new ones that can error on peer_connected and track connected/disconnected have been called or add the functionality to the existing test handlers?

Have used something like mockall for this in the past but without it I'd have to add counters/flags that get updated and a way to check them.

Is this what you had in mind for being able to test it?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

Yea, I was figuring you'd just create a trivial CustomMessageHandler that asserts that connected/disconnecteds all come in order and then errors on connection.

@johncantrell97

johncantrell97 commented Jun 7, 2024

Copy link
Copy Markdown
ContributorAuthor

Yea, I was figuring you'd just create a trivial CustomMessageHandler that asserts that connected/disconnecteds all come in order and then errors on connection.

Hm, using a CustomMessageHandler doesn't really test the fix here since it goes last. One of the issues was the early return causing the later handlers to not get the peer_connected at all. Would have to use multiple new handlers to be able to check the one after an error is returned still gets peer_connected called on it (and disconnected)

I guess at least it would catch the fix for ensuring disconnect is called.

@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

@TheBlueMatt

Added a test that passes but it duplicates a ton of code to handle all of the setup but with the new message handlers :|

not sure if this is okay, looking for feedback on the test and how to do it better if it's not okay.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch 2 times, most recently from 2c4c40a to 7a29c39CompareJune 7, 2024 23:00
@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from 7a29c39 to 922c31fCompareJune 8, 2024 01:22
if let Err(()) = self.message_handler.custom_message_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection) {
log_debug!(logger, "Custom Message Handler decided we couldn't communicate with peer {}", log_pubkey!(their_node_id));

peer_lock.their_features = Some(msg.features);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a bit confused by this, wouldn't that lead to use falsely assuming the handshake succeeded even though one of our handlers rejected it? And there is a window between us dropping the lock and handling the disconnect even where we would deal with it in a 'normal' manner, e.g., accepting further messages, and potentially rebroadcasting etc?

(cc @TheBlueMatt as he requested this change)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hm, if that's true then seems like we'll need to separate "handshake_completed" from "triggered peer_connected" with a new flag on the peer that we can use to decide whether or not to trigger peer_disconnected in do_disconnect and disconnect_event_internal?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yea, kinda, we'll end up forwarding broadcasts, as you point out, which is maybe not ideal, but we shouldn't process any further messages - we're currently in a read processing call, and we require read processing calls for any given peer to be serial, so presumably when we return an error the read-processing pipeline for this peer will stall and we won't get any more reads. We could make that explicit in the docs, however.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mhh, rather than introducing this race-y behavior in the first place, couldn't we just introduce a new handshake_aborted flag and check that alternatively to !peer.handshake_complete in disconnect_event_internal?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's fine too.

use crate::ln::msgs::{Init, LightningError, SocketAddress};
use crate::util::test_utils;


Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: Drop superfluous whitespace.

}
}

struct TestPeerTrackingMessageHandler {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, I believe the alternative would be to add TestCustomMessageHandler and TestOnionMessageHandler to test_utils and use them as part of the default test setup?

@tnull

Copy link
Copy Markdown
Contributor

@johncantrell97 Any interest in finishing this PR?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

Supersceded by #3580.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

peer_disconnected event processing isn't robust if a peer_connected Errs

4 participants

@johncantrell97@codecov-commenter@TheBlueMatt@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

ensure peer_connected is called before peer_disconnected - #3110

Closed
johncantrell97 wants to merge 2 commits into
lightningdevkit:mainfrom
johncantrell97:2024-06-robust-peer-disconnected-events
Closed

ensure peer_connected is called before peer_disconnected#3110
johncantrell97 wants to merge 2 commits into
lightningdevkit:mainfrom
johncantrell97:2024-06-robust-peer-disconnected-events

Conversation

@johncantrell97

Copy link
Copy Markdown
Contributor

Fixes#3108

Makes sure all message handler's peer_connected methods are called instead of returning early on the first to error.

As for whether or not the user has to call back into socket_disconnected after a PeerManager::read_event, I assume you mean after it returns an Err? I think the user does not have to because read_event will call disconnect_event_internal on any error before returning it to the user.

I took a look at lightning-net-tokio and it appears to be the case over there as well. It does:

if let Disconnect::PeerDisconnected = disconnect_type {
peer_manager.as_ref().socket_disconnected(&our_descriptor);
peer_manager.as_ref().process_events();
}

Only calling socket_disconnected if the disconnection type is one the user detected. If read_event returns an Err it breaks with a disconnection type of Disconnect::CloseConnection and does not call back into socket_disconnected.

Matt seems to think you do have to so I'm probably misunderstanding the original question. Happy to dig into it a bit more with some clarification if I misunderstood.

("Route Handler", self.message_handler.route_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
("Channel Handler", self.message_handler.chan_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
("Onion Handler", self.message_handler.onion_message_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
];

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you don't like this attempt to dry up the handling then I'm find just having separate results where I check them one by one with their own log messages.

@codecov-commenter

codecov-commenter commented Jun 7, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 92.56757% with 11 lines in your changes missing coverage. Please review.

Project coverage is 90.78%. Comparing base (9789152) to head (922c31f).
Report is 1274 commits behind head on main.

Files with missing linesPatch %Lines
lightning/src/ln/peer_handler.rs92.56%10 Missing and 1 partial ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #3110 +/- ##
==========================================
+ Coverage 89.84% 90.78% +0.93% 
==========================================
Files 119 119 Lines 97561 103463 +5902 Branches 97561 103463 +5902 ==========================================
+ Hits 87655 93925 +6270 + Misses 7331 7032 -299 + Partials 2575 2506 -69 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! I think we also need to move the peer_lock.their_features call up - we only call peer_disconnected if that line has been hit (Peer::handshake_complete checks for it) so we want to always hit that immediately before we call peer_connecteds.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from 4a1cade to db3b148CompareJune 7, 2024 16:02
@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

Thanks! I think we also need to move the peer_lock.their_features call up - we only call peer_disconnected if that line has been hit (Peer::handshake_complete checks for it) so we want to always hit that immediately before we call peer_connecteds.

Whoops, fixed it.

Can't be before calls to peer_connected because they pass a reference to msg but as long as we do it before returning it should be okay.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from db3b148 to 3a8f3b2CompareJune 7, 2024 16:18
TheBlueMatt
TheBlueMatt previously approved these changes Jun 7, 2024

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be nice to get a test (which should be pretty easy), but either way LGTM.

@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

Would be nice to get a test (which should be pretty easy), but either way LGTM.

Doesn't look like there's easy way to handle testing it with the existing test message handlers. Should I create new ones that can error on peer_connected and track connected/disconnected have been called or add the functionality to the existing test handlers?

Have used something like mockall for this in the past but without it I'd have to add counters/flags that get updated and a way to check them.

Is this what you had in mind for being able to test it?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

Yea, I was figuring you'd just create a trivial CustomMessageHandler that asserts that connected/disconnecteds all come in order and then errors on connection.

@johncantrell97

johncantrell97 commented Jun 7, 2024

Copy link
Copy Markdown
ContributorAuthor

Yea, I was figuring you'd just create a trivial CustomMessageHandler that asserts that connected/disconnecteds all come in order and then errors on connection.

Hm, using a CustomMessageHandler doesn't really test the fix here since it goes last. One of the issues was the early return causing the later handlers to not get the peer_connected at all. Would have to use multiple new handlers to be able to check the one after an error is returned still gets peer_connected called on it (and disconnected)

I guess at least it would catch the fix for ensuring disconnect is called.

@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

@TheBlueMatt

Added a test that passes but it duplicates a ton of code to handle all of the setup but with the new message handlers :|

not sure if this is okay, looking for feedback on the test and how to do it better if it's not okay.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch 2 times, most recently from 2c4c40a to 7a29c39CompareJune 7, 2024 23:00
@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from 7a29c39 to 922c31fCompareJune 8, 2024 01:22
if let Err(()) = self.message_handler.custom_message_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection) {
log_debug!(logger, "Custom Message Handler decided we couldn't communicate with peer {}", log_pubkey!(their_node_id));

peer_lock.their_features = Some(msg.features);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a bit confused by this, wouldn't that lead to use falsely assuming the handshake succeeded even though one of our handlers rejected it? And there is a window between us dropping the lock and handling the disconnect even where we would deal with it in a 'normal' manner, e.g., accepting further messages, and potentially rebroadcasting etc?

(cc @TheBlueMatt as he requested this change)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hm, if that's true then seems like we'll need to separate "handshake_completed" from "triggered peer_connected" with a new flag on the peer that we can use to decide whether or not to trigger peer_disconnected in do_disconnect and disconnect_event_internal?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yea, kinda, we'll end up forwarding broadcasts, as you point out, which is maybe not ideal, but we shouldn't process any further messages - we're currently in a read processing call, and we require read processing calls for any given peer to be serial, so presumably when we return an error the read-processing pipeline for this peer will stall and we won't get any more reads. We could make that explicit in the docs, however.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mhh, rather than introducing this race-y behavior in the first place, couldn't we just introduce a new handshake_aborted flag and check that alternatively to !peer.handshake_complete in disconnect_event_internal?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's fine too.

use crate::ln::msgs::{Init, LightningError, SocketAddress};
use crate::util::test_utils;


Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: Drop superfluous whitespace.

}
}

struct TestPeerTrackingMessageHandler {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, I believe the alternative would be to add TestCustomMessageHandler and TestOnionMessageHandler to test_utils and use them as part of the default test setup?

@tnull

Copy link
Copy Markdown
Contributor

@johncantrell97 Any interest in finishing this PR?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

Supersceded by #3580.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

peer_disconnected event processing isn't robust if a peer_connected Errs

4 participants

@johncantrell97@codecov-commenter@TheBlueMatt@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ensure peer_connected is called before peer_disconnected - #3110

Closed
johncantrell97 wants to merge 2 commits into
lightningdevkit:mainfrom
johncantrell97:2024-06-robust-peer-disconnected-events
Closed

ensure peer_connected is called before peer_disconnected#3110
johncantrell97 wants to merge 2 commits into
lightningdevkit:mainfrom
johncantrell97:2024-06-robust-peer-disconnected-events

Conversation

@johncantrell97

Copy link
Copy Markdown
Contributor

Fixes#3108

Makes sure all message handler's peer_connected methods are called instead of returning early on the first to error.

As for whether or not the user has to call back into socket_disconnected after a PeerManager::read_event, I assume you mean after it returns an Err? I think the user does not have to because read_event will call disconnect_event_internal on any error before returning it to the user.

I took a look at lightning-net-tokio and it appears to be the case over there as well. It does:

if let Disconnect::PeerDisconnected = disconnect_type {
peer_manager.as_ref().socket_disconnected(&our_descriptor);
peer_manager.as_ref().process_events();
}

Only calling socket_disconnected if the disconnection type is one the user detected. If read_event returns an Err it breaks with a disconnection type of Disconnect::CloseConnection and does not call back into socket_disconnected.

Matt seems to think you do have to so I'm probably misunderstanding the original question. Happy to dig into it a bit more with some clarification if I misunderstood.

("Route Handler", self.message_handler.route_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
("Channel Handler", self.message_handler.chan_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
("Onion Handler", self.message_handler.onion_message_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
];

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you don't like this attempt to dry up the handling then I'm find just having separate results where I check them one by one with their own log messages.

@codecov-commenter

codecov-commenter commented Jun 7, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 92.56757% with 11 lines in your changes missing coverage. Please review.

Project coverage is 90.78%. Comparing base (9789152) to head (922c31f).
Report is 1274 commits behind head on main.

Files with missing linesPatch %Lines
lightning/src/ln/peer_handler.rs92.56%10 Missing and 1 partial ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #3110 +/- ##
==========================================
+ Coverage 89.84% 90.78% +0.93% 
==========================================
Files 119 119 Lines 97561 103463 +5902 Branches 97561 103463 +5902 ==========================================
+ Hits 87655 93925 +6270 + Misses 7331 7032 -299 + Partials 2575 2506 -69 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! I think we also need to move the peer_lock.their_features call up - we only call peer_disconnected if that line has been hit (Peer::handshake_complete checks for it) so we want to always hit that immediately before we call peer_connecteds.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from 4a1cade to db3b148CompareJune 7, 2024 16:02
@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

Thanks! I think we also need to move the peer_lock.their_features call up - we only call peer_disconnected if that line has been hit (Peer::handshake_complete checks for it) so we want to always hit that immediately before we call peer_connecteds.

Whoops, fixed it.

Can't be before calls to peer_connected because they pass a reference to msg but as long as we do it before returning it should be okay.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from db3b148 to 3a8f3b2CompareJune 7, 2024 16:18
TheBlueMatt
TheBlueMatt previously approved these changes Jun 7, 2024

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be nice to get a test (which should be pretty easy), but either way LGTM.

@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

Would be nice to get a test (which should be pretty easy), but either way LGTM.

Doesn't look like there's easy way to handle testing it with the existing test message handlers. Should I create new ones that can error on peer_connected and track connected/disconnected have been called or add the functionality to the existing test handlers?

Have used something like mockall for this in the past but without it I'd have to add counters/flags that get updated and a way to check them.

Is this what you had in mind for being able to test it?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

Yea, I was figuring you'd just create a trivial CustomMessageHandler that asserts that connected/disconnecteds all come in order and then errors on connection.

@johncantrell97

johncantrell97 commented Jun 7, 2024

Copy link
Copy Markdown
ContributorAuthor

Yea, I was figuring you'd just create a trivial CustomMessageHandler that asserts that connected/disconnecteds all come in order and then errors on connection.

Hm, using a CustomMessageHandler doesn't really test the fix here since it goes last. One of the issues was the early return causing the later handlers to not get the peer_connected at all. Would have to use multiple new handlers to be able to check the one after an error is returned still gets peer_connected called on it (and disconnected)

I guess at least it would catch the fix for ensuring disconnect is called.

@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

@TheBlueMatt

Added a test that passes but it duplicates a ton of code to handle all of the setup but with the new message handlers :|

not sure if this is okay, looking for feedback on the test and how to do it better if it's not okay.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch 2 times, most recently from 2c4c40a to 7a29c39CompareJune 7, 2024 23:00
@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from 7a29c39 to 922c31fCompareJune 8, 2024 01:22
if let Err(()) = self.message_handler.custom_message_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection) {
log_debug!(logger, "Custom Message Handler decided we couldn't communicate with peer {}", log_pubkey!(their_node_id));

peer_lock.their_features = Some(msg.features);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a bit confused by this, wouldn't that lead to use falsely assuming the handshake succeeded even though one of our handlers rejected it? And there is a window between us dropping the lock and handling the disconnect even where we would deal with it in a 'normal' manner, e.g., accepting further messages, and potentially rebroadcasting etc?

(cc @TheBlueMatt as he requested this change)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hm, if that's true then seems like we'll need to separate "handshake_completed" from "triggered peer_connected" with a new flag on the peer that we can use to decide whether or not to trigger peer_disconnected in do_disconnect and disconnect_event_internal?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yea, kinda, we'll end up forwarding broadcasts, as you point out, which is maybe not ideal, but we shouldn't process any further messages - we're currently in a read processing call, and we require read processing calls for any given peer to be serial, so presumably when we return an error the read-processing pipeline for this peer will stall and we won't get any more reads. We could make that explicit in the docs, however.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mhh, rather than introducing this race-y behavior in the first place, couldn't we just introduce a new handshake_aborted flag and check that alternatively to !peer.handshake_complete in disconnect_event_internal?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's fine too.

use crate::ln::msgs::{Init, LightningError, SocketAddress};
use crate::util::test_utils;


Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: Drop superfluous whitespace.

}
}

struct TestPeerTrackingMessageHandler {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, I believe the alternative would be to add TestCustomMessageHandler and TestOnionMessageHandler to test_utils and use them as part of the default test setup?

@tnull

Copy link
Copy Markdown
Contributor

@johncantrell97 Any interest in finishing this PR?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

Supersceded by #3580.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

peer_disconnected event processing isn't robust if a peer_connected Errs

4 participants

@johncantrell97@codecov-commenter@TheBlueMatt@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

ensure peer_connected is called before peer_disconnected - #3110

Closed
johncantrell97 wants to merge 2 commits into
lightningdevkit:mainfrom
johncantrell97:2024-06-robust-peer-disconnected-events
Closed

ensure peer_connected is called before peer_disconnected#3110
johncantrell97 wants to merge 2 commits into
lightningdevkit:mainfrom
johncantrell97:2024-06-robust-peer-disconnected-events

Conversation

@johncantrell97

Copy link
Copy Markdown
Contributor

Fixes#3108

Makes sure all message handler's peer_connected methods are called instead of returning early on the first to error.

As for whether or not the user has to call back into socket_disconnected after a PeerManager::read_event, I assume you mean after it returns an Err? I think the user does not have to because read_event will call disconnect_event_internal on any error before returning it to the user.

I took a look at lightning-net-tokio and it appears to be the case over there as well. It does:

if let Disconnect::PeerDisconnected = disconnect_type {
peer_manager.as_ref().socket_disconnected(&our_descriptor);
peer_manager.as_ref().process_events();
}

Only calling socket_disconnected if the disconnection type is one the user detected. If read_event returns an Err it breaks with a disconnection type of Disconnect::CloseConnection and does not call back into socket_disconnected.

Matt seems to think you do have to so I'm probably misunderstanding the original question. Happy to dig into it a bit more with some clarification if I misunderstood.

("Route Handler", self.message_handler.route_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
("Channel Handler", self.message_handler.chan_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
("Onion Handler", self.message_handler.onion_message_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
];

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you don't like this attempt to dry up the handling then I'm find just having separate results where I check them one by one with their own log messages.

@codecov-commenter

codecov-commenter commented Jun 7, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 92.56757% with 11 lines in your changes missing coverage. Please review.

Project coverage is 90.78%. Comparing base (9789152) to head (922c31f).
Report is 1274 commits behind head on main.

Files with missing linesPatch %Lines
lightning/src/ln/peer_handler.rs92.56%10 Missing and 1 partial ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #3110 +/- ##
==========================================
+ Coverage 89.84% 90.78% +0.93% 
==========================================
Files 119 119 Lines 97561 103463 +5902 Branches 97561 103463 +5902 ==========================================
+ Hits 87655 93925 +6270 + Misses 7331 7032 -299 + Partials 2575 2506 -69 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! I think we also need to move the peer_lock.their_features call up - we only call peer_disconnected if that line has been hit (Peer::handshake_complete checks for it) so we want to always hit that immediately before we call peer_connecteds.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from 4a1cade to db3b148CompareJune 7, 2024 16:02
@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

Thanks! I think we also need to move the peer_lock.their_features call up - we only call peer_disconnected if that line has been hit (Peer::handshake_complete checks for it) so we want to always hit that immediately before we call peer_connecteds.

Whoops, fixed it.

Can't be before calls to peer_connected because they pass a reference to msg but as long as we do it before returning it should be okay.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from db3b148 to 3a8f3b2CompareJune 7, 2024 16:18
TheBlueMatt
TheBlueMatt previously approved these changes Jun 7, 2024

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be nice to get a test (which should be pretty easy), but either way LGTM.

@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

Would be nice to get a test (which should be pretty easy), but either way LGTM.

Doesn't look like there's easy way to handle testing it with the existing test message handlers. Should I create new ones that can error on peer_connected and track connected/disconnected have been called or add the functionality to the existing test handlers?

Have used something like mockall for this in the past but without it I'd have to add counters/flags that get updated and a way to check them.

Is this what you had in mind for being able to test it?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

Yea, I was figuring you'd just create a trivial CustomMessageHandler that asserts that connected/disconnecteds all come in order and then errors on connection.

@johncantrell97

johncantrell97 commented Jun 7, 2024

Copy link
Copy Markdown
ContributorAuthor

Yea, I was figuring you'd just create a trivial CustomMessageHandler that asserts that connected/disconnecteds all come in order and then errors on connection.

Hm, using a CustomMessageHandler doesn't really test the fix here since it goes last. One of the issues was the early return causing the later handlers to not get the peer_connected at all. Would have to use multiple new handlers to be able to check the one after an error is returned still gets peer_connected called on it (and disconnected)

I guess at least it would catch the fix for ensuring disconnect is called.

@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

@TheBlueMatt

Added a test that passes but it duplicates a ton of code to handle all of the setup but with the new message handlers :|

not sure if this is okay, looking for feedback on the test and how to do it better if it's not okay.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch 2 times, most recently from 2c4c40a to 7a29c39CompareJune 7, 2024 23:00
@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from 7a29c39 to 922c31fCompareJune 8, 2024 01:22
if let Err(()) = self.message_handler.custom_message_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection) {
log_debug!(logger, "Custom Message Handler decided we couldn't communicate with peer {}", log_pubkey!(their_node_id));

peer_lock.their_features = Some(msg.features);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a bit confused by this, wouldn't that lead to use falsely assuming the handshake succeeded even though one of our handlers rejected it? And there is a window between us dropping the lock and handling the disconnect even where we would deal with it in a 'normal' manner, e.g., accepting further messages, and potentially rebroadcasting etc?

(cc @TheBlueMatt as he requested this change)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hm, if that's true then seems like we'll need to separate "handshake_completed" from "triggered peer_connected" with a new flag on the peer that we can use to decide whether or not to trigger peer_disconnected in do_disconnect and disconnect_event_internal?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yea, kinda, we'll end up forwarding broadcasts, as you point out, which is maybe not ideal, but we shouldn't process any further messages - we're currently in a read processing call, and we require read processing calls for any given peer to be serial, so presumably when we return an error the read-processing pipeline for this peer will stall and we won't get any more reads. We could make that explicit in the docs, however.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mhh, rather than introducing this race-y behavior in the first place, couldn't we just introduce a new handshake_aborted flag and check that alternatively to !peer.handshake_complete in disconnect_event_internal?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's fine too.

use crate::ln::msgs::{Init, LightningError, SocketAddress};
use crate::util::test_utils;


Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: Drop superfluous whitespace.

}
}

struct TestPeerTrackingMessageHandler {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, I believe the alternative would be to add TestCustomMessageHandler and TestOnionMessageHandler to test_utils and use them as part of the default test setup?

@tnull

Copy link
Copy Markdown
Contributor

@johncantrell97 Any interest in finishing this PR?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

Supersceded by #3580.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

peer_disconnected event processing isn't robust if a peer_connected Errs

4 participants

@johncantrell97@codecov-commenter@TheBlueMatt@tnull
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

ensure peer_connected is called before peer_disconnected - #3110

Closed
johncantrell97 wants to merge 2 commits into
lightningdevkit:mainfrom
johncantrell97:2024-06-robust-peer-disconnected-events
Closed

ensure peer_connected is called before peer_disconnected#3110
johncantrell97 wants to merge 2 commits into
lightningdevkit:mainfrom
johncantrell97:2024-06-robust-peer-disconnected-events

Conversation

@johncantrell97

Copy link
Copy Markdown
Contributor

Fixes#3108

Makes sure all message handler's peer_connected methods are called instead of returning early on the first to error.

As for whether or not the user has to call back into socket_disconnected after a PeerManager::read_event, I assume you mean after it returns an Err? I think the user does not have to because read_event will call disconnect_event_internal on any error before returning it to the user.

I took a look at lightning-net-tokio and it appears to be the case over there as well. It does:

if let Disconnect::PeerDisconnected = disconnect_type {
peer_manager.as_ref().socket_disconnected(&our_descriptor);
peer_manager.as_ref().process_events();
}

Only calling socket_disconnected if the disconnection type is one the user detected. If read_event returns an Err it breaks with a disconnection type of Disconnect::CloseConnection and does not call back into socket_disconnected.

Matt seems to think you do have to so I'm probably misunderstanding the original question. Happy to dig into it a bit more with some clarification if I misunderstood.

("Route Handler", self.message_handler.route_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
("Channel Handler", self.message_handler.chan_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
("Onion Handler", self.message_handler.onion_message_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection)),
];

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you don't like this attempt to dry up the handling then I'm find just having separate results where I check them one by one with their own log messages.

@codecov-commenter

codecov-commenter commented Jun 7, 2024

Copy link
Copy Markdown

Codecov Report

Attention: Patch coverage is 92.56757% with 11 lines in your changes missing coverage. Please review.

Project coverage is 90.78%. Comparing base (9789152) to head (922c31f).
Report is 1274 commits behind head on main.

Files with missing linesPatch %Lines
lightning/src/ln/peer_handler.rs92.56%10 Missing and 1 partial ⚠️
Additional details and impacted files
@@ Coverage Diff @@## main #3110 +/- ##
==========================================
+ Coverage 89.84% 90.78% +0.93% 
==========================================
Files 119 119 Lines 97561 103463 +5902 Branches 97561 103463 +5902 ==========================================
+ Hits 87655 93925 +6270 + Misses 7331 7032 -299 + Partials 2575 2506 -69 

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks! I think we also need to move the peer_lock.their_features call up - we only call peer_disconnected if that line has been hit (Peer::handshake_complete checks for it) so we want to always hit that immediately before we call peer_connecteds.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from 4a1cade to db3b148CompareJune 7, 2024 16:02
@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

Thanks! I think we also need to move the peer_lock.their_features call up - we only call peer_disconnected if that line has been hit (Peer::handshake_complete checks for it) so we want to always hit that immediately before we call peer_connecteds.

Whoops, fixed it.

Can't be before calls to peer_connected because they pass a reference to msg but as long as we do it before returning it should be okay.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from db3b148 to 3a8f3b2CompareJune 7, 2024 16:18
TheBlueMatt
TheBlueMatt previously approved these changes Jun 7, 2024

@TheBlueMattTheBlueMatt left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be nice to get a test (which should be pretty easy), but either way LGTM.

@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

Would be nice to get a test (which should be pretty easy), but either way LGTM.

Doesn't look like there's easy way to handle testing it with the existing test message handlers. Should I create new ones that can error on peer_connected and track connected/disconnected have been called or add the functionality to the existing test handlers?

Have used something like mockall for this in the past but without it I'd have to add counters/flags that get updated and a way to check them.

Is this what you had in mind for being able to test it?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

Yea, I was figuring you'd just create a trivial CustomMessageHandler that asserts that connected/disconnecteds all come in order and then errors on connection.

@johncantrell97

johncantrell97 commented Jun 7, 2024

Copy link
Copy Markdown
ContributorAuthor

Yea, I was figuring you'd just create a trivial CustomMessageHandler that asserts that connected/disconnecteds all come in order and then errors on connection.

Hm, using a CustomMessageHandler doesn't really test the fix here since it goes last. One of the issues was the early return causing the later handlers to not get the peer_connected at all. Would have to use multiple new handlers to be able to check the one after an error is returned still gets peer_connected called on it (and disconnected)

I guess at least it would catch the fix for ensuring disconnect is called.

@johncantrell97

Copy link
Copy Markdown
ContributorAuthor

@TheBlueMatt

Added a test that passes but it duplicates a ton of code to handle all of the setup but with the new message handlers :|

not sure if this is okay, looking for feedback on the test and how to do it better if it's not okay.

@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch 2 times, most recently from 2c4c40a to 7a29c39CompareJune 7, 2024 23:00
@johncantrell97
johncantrell97force-pushed the 2024-06-robust-peer-disconnected-events branch from 7a29c39 to 922c31fCompareJune 8, 2024 01:22
if let Err(()) = self.message_handler.custom_message_handler.peer_connected(&their_node_id, &msg, peer_lock.inbound_connection) {
log_debug!(logger, "Custom Message Handler decided we couldn't communicate with peer {}", log_pubkey!(their_node_id));

peer_lock.their_features = Some(msg.features);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm a bit confused by this, wouldn't that lead to use falsely assuming the handshake succeeded even though one of our handlers rejected it? And there is a window between us dropping the lock and handling the disconnect even where we would deal with it in a 'normal' manner, e.g., accepting further messages, and potentially rebroadcasting etc?

(cc @TheBlueMatt as he requested this change)

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hm, if that's true then seems like we'll need to separate "handshake_completed" from "triggered peer_connected" with a new flag on the peer that we can use to decide whether or not to trigger peer_disconnected in do_disconnect and disconnect_event_internal?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yea, kinda, we'll end up forwarding broadcasts, as you point out, which is maybe not ideal, but we shouldn't process any further messages - we're currently in a read processing call, and we require read processing calls for any given peer to be serial, so presumably when we return an error the read-processing pipeline for this peer will stall and we won't get any more reads. We could make that explicit in the docs, however.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mhh, rather than introducing this race-y behavior in the first place, couldn't we just introduce a new handshake_aborted flag and check that alternatively to !peer.handshake_complete in disconnect_event_internal?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's fine too.

use crate::ln::msgs::{Init, LightningError, SocketAddress};
use crate::util::test_utils;


Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: Drop superfluous whitespace.

}
}

struct TestPeerTrackingMessageHandler {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, I believe the alternative would be to add TestCustomMessageHandler and TestOnionMessageHandler to test_utils and use them as part of the default test setup?

@tnull

Copy link
Copy Markdown
Contributor

@johncantrell97 Any interest in finishing this PR?

@TheBlueMatt

Copy link
Copy Markdown
Collaborator

Supersceded by #3580.

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

peer_disconnected event processing isn't robust if a peer_connected Errs

4 participants

@johncantrell97@codecov-commenter@TheBlueMatt@tnull