This repository was archived by the owner on Nov 15, 2023. It is now read-only.

No longer actively open legacy substreams - #7076

Merged
22 commits merged into
paritytech:masterfrom
tomaka:no-longer-open-substream
Oct 16, 2020
Merged

No longer actively open legacy substreams#7076
22 commits merged into
paritytech:masterfrom
tomaka:no-longer-open-substream

Conversation

@tomaka

Copy link
Copy Markdown
Contributor

Tackles bullet number 2 in this comment.

Based on top of #7075
Shouldn't be merged now, as we need to publish a version between #7075 and this.
The intention in this PR is to check whether CI is green to make sure that #7075 is working properly. If we merge a broken version of #7075, then we will have to wait again for a bugfix release.

This PR changes legacy.rs to no longer pro-actively open a legacy substream.
A consequence of this change is that we can no longer establish outgoing connections to nodes that don't have #7075 (hence the need for a release). However we can still receive incoming connections from nodes that don't have #7075.

The diff is quite large because of all the side clean-ups, but the core part of the changes is that legacy.rs no longer emits OutboundSubstreamRequest.

I've opted to keep the timeout system on the listening side as long as #7074 isn't resolved. After 60 seconds of inactivity on the legacy substream, the connection is force-closed, thereby freeing the peerset slot.

@tomakatomaka added A0-please_review Pull request needs code review. B5-clientnoteworthy C3-medium PR touches the given topic and has a medium impact on builders. labels Sep 10, 2020
@tomaka
tomaka marked this pull request as ready for review September 10, 2020 14:26
@tomaka

Copy link
Copy Markdown
ContributorAuthor

Switching to non-draft, as the point is to run the CI.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

I've opted to keep the timeout system on the listening side as long as #7074 isn't resolved. After 60 seconds of inactivity on the legacy substream, the connection is force-closed, thereby freeing the peerset slot.

I've now realized that, since we no longer open legacy substreams, all connections between peers would always close after 60 seconds.

At the moment, the objective of this PR is to prove that #7075 works well. I'll fix that after #7075 is approved or merged.

@mxinden

Copy link
Copy Markdown
Contributor

@tomaka let me know once you would like another review on this pull request.

@tomaka

tomaka commented Sep 15, 2020

Copy link
Copy Markdown
ContributorAuthor

Ready for review.

I've removed the timeout system from the legacy substream entirely.

With notification protocols, the listening side is pro-actively trying to open substreams, which, if they get refused, will result in the keep-alive system closing the connection.

The reason for the existing timeout system for the legacy substream comes from the situation where the dialer doesn't support notification substreams, and we don't know whether it intends to open a legacy substream.
This is no longer relevant, as nodes that don't support notification protocol substreams are too old to be supported.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

Needs a burnin, but after the release of 0.8.24.

ProtocolState::Disabled { .. } | ProtocolState::Poisoned |
ProtocolState::KillAsap => KeepAlive::No,
ProtocolState::Init { .. } | ProtocolState::Normal { .. } => KeepAlive::Yes,
ProtocolState::Opening { .. } | ProtocolState::Disabled { .. } |

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Opening is now No because of the removal of the timeout.

@mxindenmxinden left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Needs a burnin, but after the release of 0.8.24.

👍

num: usize,
err: ProtocolsHandlerUpgrErr<NotificationsHandshakeError>
) {
match (err, num) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it still makes sense to match on num, right?

@tomaka

Copy link
Copy Markdown
ContributorAuthor

Let's wait until Monday to start the burnin, so that more nodes have upgraded.

@tomaka

tomaka commented Sep 21, 2020

Copy link
Copy Markdown
ContributorAuthor

Burnin' report:

  • The percentage of filled slots seems to be unable to go above 70%. While the metrics don't distinguish between in and out slots, my assumption (as expanded in details below) is that all in slots are full while out slots aren't.
  • The number of connections forcibly closed by the local node has dropped to almost 0.
  • The number of connections forcibly closed by the remote has almost doubled.
  • The number of connections closed because of the keep alive timeout has increased from ~5/mn to ~20/mn
  • The number of dialing errors caused by an invalid PeerId is skyrocketting, from ~25/mn to ~200/mn

Force-closing a connection in general is almost always caused by the dialing side not opening a legacy substream in time when the listening side has reserved a slot (see #7074). As expected, the node with this PR consequently almost doesn't force-close connections anymore.
Instead, the situation where the dialer doesn't open substreams when expected is now supposed to trigger the keep-alive timeout rather than force-closing the connection.
Since the node with this PR no longer opens a legacy substream, its outgoing connections are force-closed by nodes on 0.8.23 and below.

I don't really have a strong explanation for the dialing errors caused by an invalid PeerId, other that, since the outgoing slots aren't full, the local node tries to repeatedly try connect to older nodes, which it wouldn't normally have to do.

In other words, so far the observation is consistent with what is expected. It's disappointing that only ~25% of the network (roughly guessing from looking at the telemetry) seems to be using 0.8.24, which is not enough to even fill all the slots of that single burnin node.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

The percentage of filled slots seems to be unable to go above 70%. While the metrics don't distinguish between in and out slots, my assumption (as expanded in details below) is that all in slots are full while out slots aren't.

Filled slots now at 85%, which probably follows the nodes upgrading to 0.8.24.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

The node looks like it is behaving normally.
We only worry concerns the increase in the InvalidPeerId errors. See also #7198.

@tomaka

tomaka commented Sep 25, 2020

Copy link
Copy Markdown
ContributorAuthor

I've been informed that the increase in InvalidPeerId errors corresponds to a period when the node had --reserved-nodes flags. For 2 days, the PR got burned-in without these flags, and it looked normal. I don't think in general that this InvalidPeerId problem is related in any way to this PR, and would go for merging.

@mxinden

Copy link
Copy Markdown
Contributor

and would go for merging.

As far as I understand this pull request requires #7075 to be widely deployed. #7075 is only part of Polkadot v0.8.24. When we merge this pull-request now it will be part of Polkadot v0.8.25.

Are we sure right now that v0.8.24 will be widely deployed once v0.8.25 is released?

@tomakatomaka removed the A0-please_review Pull request needs code review. label Oct 1, 2020
@tomaka

Copy link
Copy Markdown
ContributorAuthor

We had a couple of announcements asking validators to upgrade to at least 0.8.24.
One can see on telemetry that approximately 1/3rd of the network still uses 0.8.23-, which in the absolute is still high but low enough that merging this PR isn't really risky anymore.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

bot merge

@ghost

Copy link
Copy Markdown

Trying merge.

@ghost
ghost merged commit ec18346 into paritytech:masterOct 16, 2020
@tomaka
tomaka deleted the no-longer-open-substream branch October 16, 2020 11:07
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

C3-mediumPR touches the given topic and has a medium impact on builders.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tomaka@mxinden@romanb
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content
This repository was archived by the owner on Nov 15, 2023. It is now read-only.

No longer actively open legacy substreams - #7076

Merged
22 commits merged into
paritytech:masterfrom
tomaka:no-longer-open-substream
Oct 16, 2020
Merged

No longer actively open legacy substreams#7076
22 commits merged into
paritytech:masterfrom
tomaka:no-longer-open-substream

Conversation

@tomaka

Copy link
Copy Markdown
Contributor

Tackles bullet number 2 in this comment.

Based on top of #7075
Shouldn't be merged now, as we need to publish a version between #7075 and this.
The intention in this PR is to check whether CI is green to make sure that #7075 is working properly. If we merge a broken version of #7075, then we will have to wait again for a bugfix release.

This PR changes legacy.rs to no longer pro-actively open a legacy substream.
A consequence of this change is that we can no longer establish outgoing connections to nodes that don't have #7075 (hence the need for a release). However we can still receive incoming connections from nodes that don't have #7075.

The diff is quite large because of all the side clean-ups, but the core part of the changes is that legacy.rs no longer emits OutboundSubstreamRequest.

I've opted to keep the timeout system on the listening side as long as #7074 isn't resolved. After 60 seconds of inactivity on the legacy substream, the connection is force-closed, thereby freeing the peerset slot.

@tomakatomaka added A0-please_review Pull request needs code review. B5-clientnoteworthy C3-medium PR touches the given topic and has a medium impact on builders. labels Sep 10, 2020
@tomaka
tomaka marked this pull request as ready for review September 10, 2020 14:26
@tomaka

Copy link
Copy Markdown
ContributorAuthor

Switching to non-draft, as the point is to run the CI.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

I've opted to keep the timeout system on the listening side as long as #7074 isn't resolved. After 60 seconds of inactivity on the legacy substream, the connection is force-closed, thereby freeing the peerset slot.

I've now realized that, since we no longer open legacy substreams, all connections between peers would always close after 60 seconds.

At the moment, the objective of this PR is to prove that #7075 works well. I'll fix that after #7075 is approved or merged.

@mxinden

Copy link
Copy Markdown
Contributor

@tomaka let me know once you would like another review on this pull request.

@tomaka

tomaka commented Sep 15, 2020

Copy link
Copy Markdown
ContributorAuthor

Ready for review.

I've removed the timeout system from the legacy substream entirely.

With notification protocols, the listening side is pro-actively trying to open substreams, which, if they get refused, will result in the keep-alive system closing the connection.

The reason for the existing timeout system for the legacy substream comes from the situation where the dialer doesn't support notification substreams, and we don't know whether it intends to open a legacy substream.
This is no longer relevant, as nodes that don't support notification protocol substreams are too old to be supported.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

Needs a burnin, but after the release of 0.8.24.

ProtocolState::Disabled { .. } | ProtocolState::Poisoned |
ProtocolState::KillAsap => KeepAlive::No,
ProtocolState::Init { .. } | ProtocolState::Normal { .. } => KeepAlive::Yes,
ProtocolState::Opening { .. } | ProtocolState::Disabled { .. } |

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Opening is now No because of the removal of the timeout.

@mxindenmxinden left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Needs a burnin, but after the release of 0.8.24.

👍

num: usize,
err: ProtocolsHandlerUpgrErr<NotificationsHandshakeError>
) {
match (err, num) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it still makes sense to match on num, right?

@tomaka

Copy link
Copy Markdown
ContributorAuthor

Let's wait until Monday to start the burnin, so that more nodes have upgraded.

@tomaka

tomaka commented Sep 21, 2020

Copy link
Copy Markdown
ContributorAuthor

Burnin' report:

  • The percentage of filled slots seems to be unable to go above 70%. While the metrics don't distinguish between in and out slots, my assumption (as expanded in details below) is that all in slots are full while out slots aren't.
  • The number of connections forcibly closed by the local node has dropped to almost 0.
  • The number of connections forcibly closed by the remote has almost doubled.
  • The number of connections closed because of the keep alive timeout has increased from ~5/mn to ~20/mn
  • The number of dialing errors caused by an invalid PeerId is skyrocketting, from ~25/mn to ~200/mn

Force-closing a connection in general is almost always caused by the dialing side not opening a legacy substream in time when the listening side has reserved a slot (see #7074). As expected, the node with this PR consequently almost doesn't force-close connections anymore.
Instead, the situation where the dialer doesn't open substreams when expected is now supposed to trigger the keep-alive timeout rather than force-closing the connection.
Since the node with this PR no longer opens a legacy substream, its outgoing connections are force-closed by nodes on 0.8.23 and below.

I don't really have a strong explanation for the dialing errors caused by an invalid PeerId, other that, since the outgoing slots aren't full, the local node tries to repeatedly try connect to older nodes, which it wouldn't normally have to do.

In other words, so far the observation is consistent with what is expected. It's disappointing that only ~25% of the network (roughly guessing from looking at the telemetry) seems to be using 0.8.24, which is not enough to even fill all the slots of that single burnin node.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

The percentage of filled slots seems to be unable to go above 70%. While the metrics don't distinguish between in and out slots, my assumption (as expanded in details below) is that all in slots are full while out slots aren't.

Filled slots now at 85%, which probably follows the nodes upgrading to 0.8.24.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

The node looks like it is behaving normally.
We only worry concerns the increase in the InvalidPeerId errors. See also #7198.

@tomaka

tomaka commented Sep 25, 2020

Copy link
Copy Markdown
ContributorAuthor

I've been informed that the increase in InvalidPeerId errors corresponds to a period when the node had --reserved-nodes flags. For 2 days, the PR got burned-in without these flags, and it looked normal. I don't think in general that this InvalidPeerId problem is related in any way to this PR, and would go for merging.

@mxinden

Copy link
Copy Markdown
Contributor

and would go for merging.

As far as I understand this pull request requires #7075 to be widely deployed. #7075 is only part of Polkadot v0.8.24. When we merge this pull-request now it will be part of Polkadot v0.8.25.

Are we sure right now that v0.8.24 will be widely deployed once v0.8.25 is released?

@tomakatomaka removed the A0-please_review Pull request needs code review. label Oct 1, 2020
@tomaka

Copy link
Copy Markdown
ContributorAuthor

We had a couple of announcements asking validators to upgrade to at least 0.8.24.
One can see on telemetry that approximately 1/3rd of the network still uses 0.8.23-, which in the absolute is still high but low enough that merging this PR isn't really risky anymore.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

bot merge

@ghost

Copy link
Copy Markdown

Trying merge.

@ghost
ghost merged commit ec18346 into paritytech:masterOct 16, 2020
@tomaka
tomaka deleted the no-longer-open-substream branch October 16, 2020 11:07
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

C3-mediumPR touches the given topic and has a medium impact on builders.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tomaka@mxinden@romanb
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
This repository was archived by the owner on Nov 15, 2023. It is now read-only.

No longer actively open legacy substreams - #7076

Merged
22 commits merged into
paritytech:masterfrom
tomaka:no-longer-open-substream
Oct 16, 2020
Merged

No longer actively open legacy substreams#7076
22 commits merged into
paritytech:masterfrom
tomaka:no-longer-open-substream

Conversation

@tomaka

Copy link
Copy Markdown
Contributor

Tackles bullet number 2 in this comment.

Based on top of #7075
Shouldn't be merged now, as we need to publish a version between #7075 and this.
The intention in this PR is to check whether CI is green to make sure that #7075 is working properly. If we merge a broken version of #7075, then we will have to wait again for a bugfix release.

This PR changes legacy.rs to no longer pro-actively open a legacy substream.
A consequence of this change is that we can no longer establish outgoing connections to nodes that don't have #7075 (hence the need for a release). However we can still receive incoming connections from nodes that don't have #7075.

The diff is quite large because of all the side clean-ups, but the core part of the changes is that legacy.rs no longer emits OutboundSubstreamRequest.

I've opted to keep the timeout system on the listening side as long as #7074 isn't resolved. After 60 seconds of inactivity on the legacy substream, the connection is force-closed, thereby freeing the peerset slot.

@tomakatomaka added A0-please_review Pull request needs code review. B5-clientnoteworthy C3-medium PR touches the given topic and has a medium impact on builders. labels Sep 10, 2020
@tomaka
tomaka marked this pull request as ready for review September 10, 2020 14:26
@tomaka

Copy link
Copy Markdown
ContributorAuthor

Switching to non-draft, as the point is to run the CI.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

I've opted to keep the timeout system on the listening side as long as #7074 isn't resolved. After 60 seconds of inactivity on the legacy substream, the connection is force-closed, thereby freeing the peerset slot.

I've now realized that, since we no longer open legacy substreams, all connections between peers would always close after 60 seconds.

At the moment, the objective of this PR is to prove that #7075 works well. I'll fix that after #7075 is approved or merged.

@mxinden

Copy link
Copy Markdown
Contributor

@tomaka let me know once you would like another review on this pull request.

@tomaka

tomaka commented Sep 15, 2020

Copy link
Copy Markdown
ContributorAuthor

Ready for review.

I've removed the timeout system from the legacy substream entirely.

With notification protocols, the listening side is pro-actively trying to open substreams, which, if they get refused, will result in the keep-alive system closing the connection.

The reason for the existing timeout system for the legacy substream comes from the situation where the dialer doesn't support notification substreams, and we don't know whether it intends to open a legacy substream.
This is no longer relevant, as nodes that don't support notification protocol substreams are too old to be supported.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

Needs a burnin, but after the release of 0.8.24.

ProtocolState::Disabled { .. } | ProtocolState::Poisoned |
ProtocolState::KillAsap => KeepAlive::No,
ProtocolState::Init { .. } | ProtocolState::Normal { .. } => KeepAlive::Yes,
ProtocolState::Opening { .. } | ProtocolState::Disabled { .. } |

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Opening is now No because of the removal of the timeout.

@mxindenmxinden left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Needs a burnin, but after the release of 0.8.24.

👍

num: usize,
err: ProtocolsHandlerUpgrErr<NotificationsHandshakeError>
) {
match (err, num) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it still makes sense to match on num, right?

@tomaka

Copy link
Copy Markdown
ContributorAuthor

Let's wait until Monday to start the burnin, so that more nodes have upgraded.

@tomaka

tomaka commented Sep 21, 2020

Copy link
Copy Markdown
ContributorAuthor

Burnin' report:

  • The percentage of filled slots seems to be unable to go above 70%. While the metrics don't distinguish between in and out slots, my assumption (as expanded in details below) is that all in slots are full while out slots aren't.
  • The number of connections forcibly closed by the local node has dropped to almost 0.
  • The number of connections forcibly closed by the remote has almost doubled.
  • The number of connections closed because of the keep alive timeout has increased from ~5/mn to ~20/mn
  • The number of dialing errors caused by an invalid PeerId is skyrocketting, from ~25/mn to ~200/mn

Force-closing a connection in general is almost always caused by the dialing side not opening a legacy substream in time when the listening side has reserved a slot (see #7074). As expected, the node with this PR consequently almost doesn't force-close connections anymore.
Instead, the situation where the dialer doesn't open substreams when expected is now supposed to trigger the keep-alive timeout rather than force-closing the connection.
Since the node with this PR no longer opens a legacy substream, its outgoing connections are force-closed by nodes on 0.8.23 and below.

I don't really have a strong explanation for the dialing errors caused by an invalid PeerId, other that, since the outgoing slots aren't full, the local node tries to repeatedly try connect to older nodes, which it wouldn't normally have to do.

In other words, so far the observation is consistent with what is expected. It's disappointing that only ~25% of the network (roughly guessing from looking at the telemetry) seems to be using 0.8.24, which is not enough to even fill all the slots of that single burnin node.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

The percentage of filled slots seems to be unable to go above 70%. While the metrics don't distinguish between in and out slots, my assumption (as expanded in details below) is that all in slots are full while out slots aren't.

Filled slots now at 85%, which probably follows the nodes upgrading to 0.8.24.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

The node looks like it is behaving normally.
We only worry concerns the increase in the InvalidPeerId errors. See also #7198.

@tomaka

tomaka commented Sep 25, 2020

Copy link
Copy Markdown
ContributorAuthor

I've been informed that the increase in InvalidPeerId errors corresponds to a period when the node had --reserved-nodes flags. For 2 days, the PR got burned-in without these flags, and it looked normal. I don't think in general that this InvalidPeerId problem is related in any way to this PR, and would go for merging.

@mxinden

Copy link
Copy Markdown
Contributor

and would go for merging.

As far as I understand this pull request requires #7075 to be widely deployed. #7075 is only part of Polkadot v0.8.24. When we merge this pull-request now it will be part of Polkadot v0.8.25.

Are we sure right now that v0.8.24 will be widely deployed once v0.8.25 is released?

@tomakatomaka removed the A0-please_review Pull request needs code review. label Oct 1, 2020
@tomaka

Copy link
Copy Markdown
ContributorAuthor

We had a couple of announcements asking validators to upgrade to at least 0.8.24.
One can see on telemetry that approximately 1/3rd of the network still uses 0.8.23-, which in the absolute is still high but low enough that merging this PR isn't really risky anymore.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

bot merge

@ghost

Copy link
Copy Markdown

Trying merge.

@ghost
ghost merged commit ec18346 into paritytech:masterOct 16, 2020
@tomaka
tomaka deleted the no-longer-open-substream branch October 16, 2020 11:07
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

C3-mediumPR touches the given topic and has a medium impact on builders.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tomaka@mxinden@romanb
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
This repository was archived by the owner on Nov 15, 2023. It is now read-only.

No longer actively open legacy substreams - #7076

Merged
22 commits merged into
paritytech:masterfrom
tomaka:no-longer-open-substream
Oct 16, 2020
Merged

No longer actively open legacy substreams#7076
22 commits merged into
paritytech:masterfrom
tomaka:no-longer-open-substream

Conversation

@tomaka

Copy link
Copy Markdown
Contributor

Tackles bullet number 2 in this comment.

Based on top of #7075
Shouldn't be merged now, as we need to publish a version between #7075 and this.
The intention in this PR is to check whether CI is green to make sure that #7075 is working properly. If we merge a broken version of #7075, then we will have to wait again for a bugfix release.

This PR changes legacy.rs to no longer pro-actively open a legacy substream.
A consequence of this change is that we can no longer establish outgoing connections to nodes that don't have #7075 (hence the need for a release). However we can still receive incoming connections from nodes that don't have #7075.

The diff is quite large because of all the side clean-ups, but the core part of the changes is that legacy.rs no longer emits OutboundSubstreamRequest.

I've opted to keep the timeout system on the listening side as long as #7074 isn't resolved. After 60 seconds of inactivity on the legacy substream, the connection is force-closed, thereby freeing the peerset slot.

@tomakatomaka added A0-please_review Pull request needs code review. B5-clientnoteworthy C3-medium PR touches the given topic and has a medium impact on builders. labels Sep 10, 2020
@tomaka
tomaka marked this pull request as ready for review September 10, 2020 14:26
@tomaka

Copy link
Copy Markdown
ContributorAuthor

Switching to non-draft, as the point is to run the CI.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

I've opted to keep the timeout system on the listening side as long as #7074 isn't resolved. After 60 seconds of inactivity on the legacy substream, the connection is force-closed, thereby freeing the peerset slot.

I've now realized that, since we no longer open legacy substreams, all connections between peers would always close after 60 seconds.

At the moment, the objective of this PR is to prove that #7075 works well. I'll fix that after #7075 is approved or merged.

@mxinden

Copy link
Copy Markdown
Contributor

@tomaka let me know once you would like another review on this pull request.

@tomaka

tomaka commented Sep 15, 2020

Copy link
Copy Markdown
ContributorAuthor

Ready for review.

I've removed the timeout system from the legacy substream entirely.

With notification protocols, the listening side is pro-actively trying to open substreams, which, if they get refused, will result in the keep-alive system closing the connection.

The reason for the existing timeout system for the legacy substream comes from the situation where the dialer doesn't support notification substreams, and we don't know whether it intends to open a legacy substream.
This is no longer relevant, as nodes that don't support notification protocol substreams are too old to be supported.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

Needs a burnin, but after the release of 0.8.24.

ProtocolState::Disabled { .. } | ProtocolState::Poisoned |
ProtocolState::KillAsap => KeepAlive::No,
ProtocolState::Init { .. } | ProtocolState::Normal { .. } => KeepAlive::Yes,
ProtocolState::Opening { .. } | ProtocolState::Disabled { .. } |

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Opening is now No because of the removal of the timeout.

@mxindenmxinden left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Needs a burnin, but after the release of 0.8.24.

👍

num: usize,
err: ProtocolsHandlerUpgrErr<NotificationsHandshakeError>
) {
match (err, num) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it still makes sense to match on num, right?

@tomaka

Copy link
Copy Markdown
ContributorAuthor

Let's wait until Monday to start the burnin, so that more nodes have upgraded.

@tomaka

tomaka commented Sep 21, 2020

Copy link
Copy Markdown
ContributorAuthor

Burnin' report:

  • The percentage of filled slots seems to be unable to go above 70%. While the metrics don't distinguish between in and out slots, my assumption (as expanded in details below) is that all in slots are full while out slots aren't.
  • The number of connections forcibly closed by the local node has dropped to almost 0.
  • The number of connections forcibly closed by the remote has almost doubled.
  • The number of connections closed because of the keep alive timeout has increased from ~5/mn to ~20/mn
  • The number of dialing errors caused by an invalid PeerId is skyrocketting, from ~25/mn to ~200/mn

Force-closing a connection in general is almost always caused by the dialing side not opening a legacy substream in time when the listening side has reserved a slot (see #7074). As expected, the node with this PR consequently almost doesn't force-close connections anymore.
Instead, the situation where the dialer doesn't open substreams when expected is now supposed to trigger the keep-alive timeout rather than force-closing the connection.
Since the node with this PR no longer opens a legacy substream, its outgoing connections are force-closed by nodes on 0.8.23 and below.

I don't really have a strong explanation for the dialing errors caused by an invalid PeerId, other that, since the outgoing slots aren't full, the local node tries to repeatedly try connect to older nodes, which it wouldn't normally have to do.

In other words, so far the observation is consistent with what is expected. It's disappointing that only ~25% of the network (roughly guessing from looking at the telemetry) seems to be using 0.8.24, which is not enough to even fill all the slots of that single burnin node.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

The percentage of filled slots seems to be unable to go above 70%. While the metrics don't distinguish between in and out slots, my assumption (as expanded in details below) is that all in slots are full while out slots aren't.

Filled slots now at 85%, which probably follows the nodes upgrading to 0.8.24.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

The node looks like it is behaving normally.
We only worry concerns the increase in the InvalidPeerId errors. See also #7198.

@tomaka

tomaka commented Sep 25, 2020

Copy link
Copy Markdown
ContributorAuthor

I've been informed that the increase in InvalidPeerId errors corresponds to a period when the node had --reserved-nodes flags. For 2 days, the PR got burned-in without these flags, and it looked normal. I don't think in general that this InvalidPeerId problem is related in any way to this PR, and would go for merging.

@mxinden

Copy link
Copy Markdown
Contributor

and would go for merging.

As far as I understand this pull request requires #7075 to be widely deployed. #7075 is only part of Polkadot v0.8.24. When we merge this pull-request now it will be part of Polkadot v0.8.25.

Are we sure right now that v0.8.24 will be widely deployed once v0.8.25 is released?

@tomakatomaka removed the A0-please_review Pull request needs code review. label Oct 1, 2020
@tomaka

Copy link
Copy Markdown
ContributorAuthor

We had a couple of announcements asking validators to upgrade to at least 0.8.24.
One can see on telemetry that approximately 1/3rd of the network still uses 0.8.23-, which in the absolute is still high but low enough that merging this PR isn't really risky anymore.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

bot merge

@ghost

Copy link
Copy Markdown

Trying merge.

@ghost
ghost merged commit ec18346 into paritytech:masterOct 16, 2020
@tomaka
tomaka deleted the no-longer-open-substream branch October 16, 2020 11:07
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

C3-mediumPR touches the given topic and has a medium impact on builders.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tomaka@mxinden@romanb
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content
This repository was archived by the owner on Nov 15, 2023. It is now read-only.

No longer actively open legacy substreams - #7076

Merged
22 commits merged into
paritytech:masterfrom
tomaka:no-longer-open-substream
Oct 16, 2020
Merged

No longer actively open legacy substreams#7076
22 commits merged into
paritytech:masterfrom
tomaka:no-longer-open-substream

Conversation

@tomaka

Copy link
Copy Markdown
Contributor

Tackles bullet number 2 in this comment.

Based on top of #7075
Shouldn't be merged now, as we need to publish a version between #7075 and this.
The intention in this PR is to check whether CI is green to make sure that #7075 is working properly. If we merge a broken version of #7075, then we will have to wait again for a bugfix release.

This PR changes legacy.rs to no longer pro-actively open a legacy substream.
A consequence of this change is that we can no longer establish outgoing connections to nodes that don't have #7075 (hence the need for a release). However we can still receive incoming connections from nodes that don't have #7075.

The diff is quite large because of all the side clean-ups, but the core part of the changes is that legacy.rs no longer emits OutboundSubstreamRequest.

I've opted to keep the timeout system on the listening side as long as #7074 isn't resolved. After 60 seconds of inactivity on the legacy substream, the connection is force-closed, thereby freeing the peerset slot.

@tomakatomaka added A0-please_review Pull request needs code review. B5-clientnoteworthy C3-medium PR touches the given topic and has a medium impact on builders. labels Sep 10, 2020
@tomaka
tomaka marked this pull request as ready for review September 10, 2020 14:26
@tomaka

Copy link
Copy Markdown
ContributorAuthor

Switching to non-draft, as the point is to run the CI.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

I've opted to keep the timeout system on the listening side as long as #7074 isn't resolved. After 60 seconds of inactivity on the legacy substream, the connection is force-closed, thereby freeing the peerset slot.

I've now realized that, since we no longer open legacy substreams, all connections between peers would always close after 60 seconds.

At the moment, the objective of this PR is to prove that #7075 works well. I'll fix that after #7075 is approved or merged.

@mxinden

Copy link
Copy Markdown
Contributor

@tomaka let me know once you would like another review on this pull request.

@tomaka

tomaka commented Sep 15, 2020

Copy link
Copy Markdown
ContributorAuthor

Ready for review.

I've removed the timeout system from the legacy substream entirely.

With notification protocols, the listening side is pro-actively trying to open substreams, which, if they get refused, will result in the keep-alive system closing the connection.

The reason for the existing timeout system for the legacy substream comes from the situation where the dialer doesn't support notification substreams, and we don't know whether it intends to open a legacy substream.
This is no longer relevant, as nodes that don't support notification protocol substreams are too old to be supported.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

Needs a burnin, but after the release of 0.8.24.

ProtocolState::Disabled { .. } | ProtocolState::Poisoned |
ProtocolState::KillAsap => KeepAlive::No,
ProtocolState::Init { .. } | ProtocolState::Normal { .. } => KeepAlive::Yes,
ProtocolState::Opening { .. } | ProtocolState::Disabled { .. } |

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Opening is now No because of the removal of the timeout.

@mxindenmxinden left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Needs a burnin, but after the release of 0.8.24.

👍

num: usize,
err: ProtocolsHandlerUpgrErr<NotificationsHandshakeError>
) {
match (err, num) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it still makes sense to match on num, right?

@tomaka

Copy link
Copy Markdown
ContributorAuthor

Let's wait until Monday to start the burnin, so that more nodes have upgraded.

@tomaka

tomaka commented Sep 21, 2020

Copy link
Copy Markdown
ContributorAuthor

Burnin' report:

  • The percentage of filled slots seems to be unable to go above 70%. While the metrics don't distinguish between in and out slots, my assumption (as expanded in details below) is that all in slots are full while out slots aren't.
  • The number of connections forcibly closed by the local node has dropped to almost 0.
  • The number of connections forcibly closed by the remote has almost doubled.
  • The number of connections closed because of the keep alive timeout has increased from ~5/mn to ~20/mn
  • The number of dialing errors caused by an invalid PeerId is skyrocketting, from ~25/mn to ~200/mn

Force-closing a connection in general is almost always caused by the dialing side not opening a legacy substream in time when the listening side has reserved a slot (see #7074). As expected, the node with this PR consequently almost doesn't force-close connections anymore.
Instead, the situation where the dialer doesn't open substreams when expected is now supposed to trigger the keep-alive timeout rather than force-closing the connection.
Since the node with this PR no longer opens a legacy substream, its outgoing connections are force-closed by nodes on 0.8.23 and below.

I don't really have a strong explanation for the dialing errors caused by an invalid PeerId, other that, since the outgoing slots aren't full, the local node tries to repeatedly try connect to older nodes, which it wouldn't normally have to do.

In other words, so far the observation is consistent with what is expected. It's disappointing that only ~25% of the network (roughly guessing from looking at the telemetry) seems to be using 0.8.24, which is not enough to even fill all the slots of that single burnin node.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

The percentage of filled slots seems to be unable to go above 70%. While the metrics don't distinguish between in and out slots, my assumption (as expanded in details below) is that all in slots are full while out slots aren't.

Filled slots now at 85%, which probably follows the nodes upgrading to 0.8.24.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

The node looks like it is behaving normally.
We only worry concerns the increase in the InvalidPeerId errors. See also #7198.

@tomaka

tomaka commented Sep 25, 2020

Copy link
Copy Markdown
ContributorAuthor

I've been informed that the increase in InvalidPeerId errors corresponds to a period when the node had --reserved-nodes flags. For 2 days, the PR got burned-in without these flags, and it looked normal. I don't think in general that this InvalidPeerId problem is related in any way to this PR, and would go for merging.

@mxinden

Copy link
Copy Markdown
Contributor

and would go for merging.

As far as I understand this pull request requires #7075 to be widely deployed. #7075 is only part of Polkadot v0.8.24. When we merge this pull-request now it will be part of Polkadot v0.8.25.

Are we sure right now that v0.8.24 will be widely deployed once v0.8.25 is released?

@tomakatomaka removed the A0-please_review Pull request needs code review. label Oct 1, 2020
@tomaka

Copy link
Copy Markdown
ContributorAuthor

We had a couple of announcements asking validators to upgrade to at least 0.8.24.
One can see on telemetry that approximately 1/3rd of the network still uses 0.8.23-, which in the absolute is still high but low enough that merging this PR isn't really risky anymore.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

bot merge

@ghost

Copy link
Copy Markdown

Trying merge.

@ghost
ghost merged commit ec18346 into paritytech:masterOct 16, 2020
@tomaka
tomaka deleted the no-longer-open-substream branch October 16, 2020 11:07
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

C3-mediumPR touches the given topic and has a medium impact on builders.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tomaka@mxinden@romanb
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
This repository was archived by the owner on Nov 15, 2023. It is now read-only.

No longer actively open legacy substreams - #7076

Merged
22 commits merged into
paritytech:masterfrom
tomaka:no-longer-open-substream
Oct 16, 2020
Merged

No longer actively open legacy substreams#7076
22 commits merged into
paritytech:masterfrom
tomaka:no-longer-open-substream

Conversation

@tomaka

Copy link
Copy Markdown
Contributor

Tackles bullet number 2 in this comment.

Based on top of #7075
Shouldn't be merged now, as we need to publish a version between #7075 and this.
The intention in this PR is to check whether CI is green to make sure that #7075 is working properly. If we merge a broken version of #7075, then we will have to wait again for a bugfix release.

This PR changes legacy.rs to no longer pro-actively open a legacy substream.
A consequence of this change is that we can no longer establish outgoing connections to nodes that don't have #7075 (hence the need for a release). However we can still receive incoming connections from nodes that don't have #7075.

The diff is quite large because of all the side clean-ups, but the core part of the changes is that legacy.rs no longer emits OutboundSubstreamRequest.

I've opted to keep the timeout system on the listening side as long as #7074 isn't resolved. After 60 seconds of inactivity on the legacy substream, the connection is force-closed, thereby freeing the peerset slot.

@tomakatomaka added A0-please_review Pull request needs code review. B5-clientnoteworthy C3-medium PR touches the given topic and has a medium impact on builders. labels Sep 10, 2020
@tomaka
tomaka marked this pull request as ready for review September 10, 2020 14:26
@tomaka

Copy link
Copy Markdown
ContributorAuthor

Switching to non-draft, as the point is to run the CI.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

I've opted to keep the timeout system on the listening side as long as #7074 isn't resolved. After 60 seconds of inactivity on the legacy substream, the connection is force-closed, thereby freeing the peerset slot.

I've now realized that, since we no longer open legacy substreams, all connections between peers would always close after 60 seconds.

At the moment, the objective of this PR is to prove that #7075 works well. I'll fix that after #7075 is approved or merged.

@mxinden

Copy link
Copy Markdown
Contributor

@tomaka let me know once you would like another review on this pull request.

@tomaka

tomaka commented Sep 15, 2020

Copy link
Copy Markdown
ContributorAuthor

Ready for review.

I've removed the timeout system from the legacy substream entirely.

With notification protocols, the listening side is pro-actively trying to open substreams, which, if they get refused, will result in the keep-alive system closing the connection.

The reason for the existing timeout system for the legacy substream comes from the situation where the dialer doesn't support notification substreams, and we don't know whether it intends to open a legacy substream.
This is no longer relevant, as nodes that don't support notification protocol substreams are too old to be supported.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

Needs a burnin, but after the release of 0.8.24.

ProtocolState::Disabled { .. } | ProtocolState::Poisoned |
ProtocolState::KillAsap => KeepAlive::No,
ProtocolState::Init { .. } | ProtocolState::Normal { .. } => KeepAlive::Yes,
ProtocolState::Opening { .. } | ProtocolState::Disabled { .. } |

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Opening is now No because of the removal of the timeout.

@mxindenmxinden left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Needs a burnin, but after the release of 0.8.24.

👍

num: usize,
err: ProtocolsHandlerUpgrErr<NotificationsHandshakeError>
) {
match (err, num) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it still makes sense to match on num, right?

@tomaka

Copy link
Copy Markdown
ContributorAuthor

Let's wait until Monday to start the burnin, so that more nodes have upgraded.

@tomaka

tomaka commented Sep 21, 2020

Copy link
Copy Markdown
ContributorAuthor

Burnin' report:

  • The percentage of filled slots seems to be unable to go above 70%. While the metrics don't distinguish between in and out slots, my assumption (as expanded in details below) is that all in slots are full while out slots aren't.
  • The number of connections forcibly closed by the local node has dropped to almost 0.
  • The number of connections forcibly closed by the remote has almost doubled.
  • The number of connections closed because of the keep alive timeout has increased from ~5/mn to ~20/mn
  • The number of dialing errors caused by an invalid PeerId is skyrocketting, from ~25/mn to ~200/mn

Force-closing a connection in general is almost always caused by the dialing side not opening a legacy substream in time when the listening side has reserved a slot (see #7074). As expected, the node with this PR consequently almost doesn't force-close connections anymore.
Instead, the situation where the dialer doesn't open substreams when expected is now supposed to trigger the keep-alive timeout rather than force-closing the connection.
Since the node with this PR no longer opens a legacy substream, its outgoing connections are force-closed by nodes on 0.8.23 and below.

I don't really have a strong explanation for the dialing errors caused by an invalid PeerId, other that, since the outgoing slots aren't full, the local node tries to repeatedly try connect to older nodes, which it wouldn't normally have to do.

In other words, so far the observation is consistent with what is expected. It's disappointing that only ~25% of the network (roughly guessing from looking at the telemetry) seems to be using 0.8.24, which is not enough to even fill all the slots of that single burnin node.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

The percentage of filled slots seems to be unable to go above 70%. While the metrics don't distinguish between in and out slots, my assumption (as expanded in details below) is that all in slots are full while out slots aren't.

Filled slots now at 85%, which probably follows the nodes upgrading to 0.8.24.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

The node looks like it is behaving normally.
We only worry concerns the increase in the InvalidPeerId errors. See also #7198.

@tomaka

tomaka commented Sep 25, 2020

Copy link
Copy Markdown
ContributorAuthor

I've been informed that the increase in InvalidPeerId errors corresponds to a period when the node had --reserved-nodes flags. For 2 days, the PR got burned-in without these flags, and it looked normal. I don't think in general that this InvalidPeerId problem is related in any way to this PR, and would go for merging.

@mxinden

Copy link
Copy Markdown
Contributor

and would go for merging.

As far as I understand this pull request requires #7075 to be widely deployed. #7075 is only part of Polkadot v0.8.24. When we merge this pull-request now it will be part of Polkadot v0.8.25.

Are we sure right now that v0.8.24 will be widely deployed once v0.8.25 is released?

@tomakatomaka removed the A0-please_review Pull request needs code review. label Oct 1, 2020
@tomaka

Copy link
Copy Markdown
ContributorAuthor

We had a couple of announcements asking validators to upgrade to at least 0.8.24.
One can see on telemetry that approximately 1/3rd of the network still uses 0.8.23-, which in the absolute is still high but low enough that merging this PR isn't really risky anymore.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

bot merge

@ghost

Copy link
Copy Markdown

Trying merge.

@ghost
ghost merged commit ec18346 into paritytech:masterOct 16, 2020
@tomaka
tomaka deleted the no-longer-open-substream branch October 16, 2020 11:07
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

C3-mediumPR touches the given topic and has a medium impact on builders.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tomaka@mxinden@romanb
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content
This repository was archived by the owner on Nov 15, 2023. It is now read-only.

No longer actively open legacy substreams - #7076

Merged
22 commits merged into
paritytech:masterfrom
tomaka:no-longer-open-substream
Oct 16, 2020
Merged

No longer actively open legacy substreams#7076
22 commits merged into
paritytech:masterfrom
tomaka:no-longer-open-substream

Conversation

@tomaka

Copy link
Copy Markdown
Contributor

Tackles bullet number 2 in this comment.

Based on top of #7075
Shouldn't be merged now, as we need to publish a version between #7075 and this.
The intention in this PR is to check whether CI is green to make sure that #7075 is working properly. If we merge a broken version of #7075, then we will have to wait again for a bugfix release.

This PR changes legacy.rs to no longer pro-actively open a legacy substream.
A consequence of this change is that we can no longer establish outgoing connections to nodes that don't have #7075 (hence the need for a release). However we can still receive incoming connections from nodes that don't have #7075.

The diff is quite large because of all the side clean-ups, but the core part of the changes is that legacy.rs no longer emits OutboundSubstreamRequest.

I've opted to keep the timeout system on the listening side as long as #7074 isn't resolved. After 60 seconds of inactivity on the legacy substream, the connection is force-closed, thereby freeing the peerset slot.

@tomakatomaka added A0-please_review Pull request needs code review. B5-clientnoteworthy C3-medium PR touches the given topic and has a medium impact on builders. labels Sep 10, 2020
@tomaka
tomaka marked this pull request as ready for review September 10, 2020 14:26
@tomaka

Copy link
Copy Markdown
ContributorAuthor

Switching to non-draft, as the point is to run the CI.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

I've opted to keep the timeout system on the listening side as long as #7074 isn't resolved. After 60 seconds of inactivity on the legacy substream, the connection is force-closed, thereby freeing the peerset slot.

I've now realized that, since we no longer open legacy substreams, all connections between peers would always close after 60 seconds.

At the moment, the objective of this PR is to prove that #7075 works well. I'll fix that after #7075 is approved or merged.

@mxinden

Copy link
Copy Markdown
Contributor

@tomaka let me know once you would like another review on this pull request.

@tomaka

tomaka commented Sep 15, 2020

Copy link
Copy Markdown
ContributorAuthor

Ready for review.

I've removed the timeout system from the legacy substream entirely.

With notification protocols, the listening side is pro-actively trying to open substreams, which, if they get refused, will result in the keep-alive system closing the connection.

The reason for the existing timeout system for the legacy substream comes from the situation where the dialer doesn't support notification substreams, and we don't know whether it intends to open a legacy substream.
This is no longer relevant, as nodes that don't support notification protocol substreams are too old to be supported.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

Needs a burnin, but after the release of 0.8.24.

ProtocolState::Disabled { .. } | ProtocolState::Poisoned |
ProtocolState::KillAsap => KeepAlive::No,
ProtocolState::Init { .. } | ProtocolState::Normal { .. } => KeepAlive::Yes,
ProtocolState::Opening { .. } | ProtocolState::Disabled { .. } |

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Opening is now No because of the removal of the timeout.

@mxindenmxinden left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Needs a burnin, but after the release of 0.8.24.

👍

num: usize,
err: ProtocolsHandlerUpgrErr<NotificationsHandshakeError>
) {
match (err, num) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it still makes sense to match on num, right?

@tomaka

Copy link
Copy Markdown
ContributorAuthor

Let's wait until Monday to start the burnin, so that more nodes have upgraded.

@tomaka

tomaka commented Sep 21, 2020

Copy link
Copy Markdown
ContributorAuthor

Burnin' report:

  • The percentage of filled slots seems to be unable to go above 70%. While the metrics don't distinguish between in and out slots, my assumption (as expanded in details below) is that all in slots are full while out slots aren't.
  • The number of connections forcibly closed by the local node has dropped to almost 0.
  • The number of connections forcibly closed by the remote has almost doubled.
  • The number of connections closed because of the keep alive timeout has increased from ~5/mn to ~20/mn
  • The number of dialing errors caused by an invalid PeerId is skyrocketting, from ~25/mn to ~200/mn

Force-closing a connection in general is almost always caused by the dialing side not opening a legacy substream in time when the listening side has reserved a slot (see #7074). As expected, the node with this PR consequently almost doesn't force-close connections anymore.
Instead, the situation where the dialer doesn't open substreams when expected is now supposed to trigger the keep-alive timeout rather than force-closing the connection.
Since the node with this PR no longer opens a legacy substream, its outgoing connections are force-closed by nodes on 0.8.23 and below.

I don't really have a strong explanation for the dialing errors caused by an invalid PeerId, other that, since the outgoing slots aren't full, the local node tries to repeatedly try connect to older nodes, which it wouldn't normally have to do.

In other words, so far the observation is consistent with what is expected. It's disappointing that only ~25% of the network (roughly guessing from looking at the telemetry) seems to be using 0.8.24, which is not enough to even fill all the slots of that single burnin node.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

The percentage of filled slots seems to be unable to go above 70%. While the metrics don't distinguish between in and out slots, my assumption (as expanded in details below) is that all in slots are full while out slots aren't.

Filled slots now at 85%, which probably follows the nodes upgrading to 0.8.24.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

The node looks like it is behaving normally.
We only worry concerns the increase in the InvalidPeerId errors. See also #7198.

@tomaka

tomaka commented Sep 25, 2020

Copy link
Copy Markdown
ContributorAuthor

I've been informed that the increase in InvalidPeerId errors corresponds to a period when the node had --reserved-nodes flags. For 2 days, the PR got burned-in without these flags, and it looked normal. I don't think in general that this InvalidPeerId problem is related in any way to this PR, and would go for merging.

@mxinden

Copy link
Copy Markdown
Contributor

and would go for merging.

As far as I understand this pull request requires #7075 to be widely deployed. #7075 is only part of Polkadot v0.8.24. When we merge this pull-request now it will be part of Polkadot v0.8.25.

Are we sure right now that v0.8.24 will be widely deployed once v0.8.25 is released?

@tomakatomaka removed the A0-please_review Pull request needs code review. label Oct 1, 2020
@tomaka

Copy link
Copy Markdown
ContributorAuthor

We had a couple of announcements asking validators to upgrade to at least 0.8.24.
One can see on telemetry that approximately 1/3rd of the network still uses 0.8.23-, which in the absolute is still high but low enough that merging this PR isn't really risky anymore.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

bot merge

@ghost

Copy link
Copy Markdown

Trying merge.

@ghost
ghost merged commit ec18346 into paritytech:masterOct 16, 2020
@tomaka
tomaka deleted the no-longer-open-substream branch October 16, 2020 11:07
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

C3-mediumPR touches the given topic and has a medium impact on builders.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tomaka@mxinden@romanb
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content
This repository was archived by the owner on Nov 15, 2023. It is now read-only.

No longer actively open legacy substreams - #7076

Merged
22 commits merged into
paritytech:masterfrom
tomaka:no-longer-open-substream
Oct 16, 2020
Merged

No longer actively open legacy substreams#7076
22 commits merged into
paritytech:masterfrom
tomaka:no-longer-open-substream

Conversation

@tomaka

Copy link
Copy Markdown
Contributor

Tackles bullet number 2 in this comment.

Based on top of #7075
Shouldn't be merged now, as we need to publish a version between #7075 and this.
The intention in this PR is to check whether CI is green to make sure that #7075 is working properly. If we merge a broken version of #7075, then we will have to wait again for a bugfix release.

This PR changes legacy.rs to no longer pro-actively open a legacy substream.
A consequence of this change is that we can no longer establish outgoing connections to nodes that don't have #7075 (hence the need for a release). However we can still receive incoming connections from nodes that don't have #7075.

The diff is quite large because of all the side clean-ups, but the core part of the changes is that legacy.rs no longer emits OutboundSubstreamRequest.

I've opted to keep the timeout system on the listening side as long as #7074 isn't resolved. After 60 seconds of inactivity on the legacy substream, the connection is force-closed, thereby freeing the peerset slot.

@tomakatomaka added A0-please_review Pull request needs code review. B5-clientnoteworthy C3-medium PR touches the given topic and has a medium impact on builders. labels Sep 10, 2020
@tomaka
tomaka marked this pull request as ready for review September 10, 2020 14:26
@tomaka

Copy link
Copy Markdown
ContributorAuthor

Switching to non-draft, as the point is to run the CI.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

I've opted to keep the timeout system on the listening side as long as #7074 isn't resolved. After 60 seconds of inactivity on the legacy substream, the connection is force-closed, thereby freeing the peerset slot.

I've now realized that, since we no longer open legacy substreams, all connections between peers would always close after 60 seconds.

At the moment, the objective of this PR is to prove that #7075 works well. I'll fix that after #7075 is approved or merged.

@mxinden

Copy link
Copy Markdown
Contributor

@tomaka let me know once you would like another review on this pull request.

@tomaka

tomaka commented Sep 15, 2020

Copy link
Copy Markdown
ContributorAuthor

Ready for review.

I've removed the timeout system from the legacy substream entirely.

With notification protocols, the listening side is pro-actively trying to open substreams, which, if they get refused, will result in the keep-alive system closing the connection.

The reason for the existing timeout system for the legacy substream comes from the situation where the dialer doesn't support notification substreams, and we don't know whether it intends to open a legacy substream.
This is no longer relevant, as nodes that don't support notification protocol substreams are too old to be supported.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

Needs a burnin, but after the release of 0.8.24.

ProtocolState::Disabled { .. } | ProtocolState::Poisoned |
ProtocolState::KillAsap => KeepAlive::No,
ProtocolState::Init { .. } | ProtocolState::Normal { .. } => KeepAlive::Yes,
ProtocolState::Opening { .. } | ProtocolState::Disabled { .. } |

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Opening is now No because of the removal of the timeout.

@mxindenmxinden left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Needs a burnin, but after the release of 0.8.24.

👍

num: usize,
err: ProtocolsHandlerUpgrErr<NotificationsHandshakeError>
) {
match (err, num) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think it still makes sense to match on num, right?

@tomaka

Copy link
Copy Markdown
ContributorAuthor

Let's wait until Monday to start the burnin, so that more nodes have upgraded.

@tomaka

tomaka commented Sep 21, 2020

Copy link
Copy Markdown
ContributorAuthor

Burnin' report:

  • The percentage of filled slots seems to be unable to go above 70%. While the metrics don't distinguish between in and out slots, my assumption (as expanded in details below) is that all in slots are full while out slots aren't.
  • The number of connections forcibly closed by the local node has dropped to almost 0.
  • The number of connections forcibly closed by the remote has almost doubled.
  • The number of connections closed because of the keep alive timeout has increased from ~5/mn to ~20/mn
  • The number of dialing errors caused by an invalid PeerId is skyrocketting, from ~25/mn to ~200/mn

Force-closing a connection in general is almost always caused by the dialing side not opening a legacy substream in time when the listening side has reserved a slot (see #7074). As expected, the node with this PR consequently almost doesn't force-close connections anymore.
Instead, the situation where the dialer doesn't open substreams when expected is now supposed to trigger the keep-alive timeout rather than force-closing the connection.
Since the node with this PR no longer opens a legacy substream, its outgoing connections are force-closed by nodes on 0.8.23 and below.

I don't really have a strong explanation for the dialing errors caused by an invalid PeerId, other that, since the outgoing slots aren't full, the local node tries to repeatedly try connect to older nodes, which it wouldn't normally have to do.

In other words, so far the observation is consistent with what is expected. It's disappointing that only ~25% of the network (roughly guessing from looking at the telemetry) seems to be using 0.8.24, which is not enough to even fill all the slots of that single burnin node.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

The percentage of filled slots seems to be unable to go above 70%. While the metrics don't distinguish between in and out slots, my assumption (as expanded in details below) is that all in slots are full while out slots aren't.

Filled slots now at 85%, which probably follows the nodes upgrading to 0.8.24.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

The node looks like it is behaving normally.
We only worry concerns the increase in the InvalidPeerId errors. See also #7198.

@tomaka

tomaka commented Sep 25, 2020

Copy link
Copy Markdown
ContributorAuthor

I've been informed that the increase in InvalidPeerId errors corresponds to a period when the node had --reserved-nodes flags. For 2 days, the PR got burned-in without these flags, and it looked normal. I don't think in general that this InvalidPeerId problem is related in any way to this PR, and would go for merging.

@mxinden

Copy link
Copy Markdown
Contributor

and would go for merging.

As far as I understand this pull request requires #7075 to be widely deployed. #7075 is only part of Polkadot v0.8.24. When we merge this pull-request now it will be part of Polkadot v0.8.25.

Are we sure right now that v0.8.24 will be widely deployed once v0.8.25 is released?

@tomakatomaka removed the A0-please_review Pull request needs code review. label Oct 1, 2020
@tomaka

Copy link
Copy Markdown
ContributorAuthor

We had a couple of announcements asking validators to upgrade to at least 0.8.24.
One can see on telemetry that approximately 1/3rd of the network still uses 0.8.23-, which in the absolute is still high but low enough that merging this PR isn't really risky anymore.

@tomaka

Copy link
Copy Markdown
ContributorAuthor

bot merge

@ghost

Copy link
Copy Markdown

Trying merge.

@ghost
ghost merged commit ec18346 into paritytech:masterOct 16, 2020
@tomaka
tomaka deleted the no-longer-open-substream branch October 16, 2020 11:07
This pull request was closed.
Sign up for freeto subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

C3-mediumPR touches the given topic and has a medium impact on builders.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@tomaka@mxinden@romanb