Bump MAX_FEC_BLOCKS to 4 - #2787

Closed
CypherGrue wants to merge 1 commit into
LizardByte:masterfrom
CypherGrue:master
Closed

Bump MAX_FEC_BLOCKS to 4#2787
CypherGrue wants to merge 1 commit into
LizardByte:masterfrom
CypherGrue:master

Conversation

@CypherGrue

@CypherGrueCypherGrue commented Jul 1, 2024

Copy link
Copy Markdown

Description

Sunshine implementation arbitrarily limits itself to 3 FEC blocks even though the protocol supports 4. This is fine for typical use cases. However, when dealing with large payloads (low FPS, high bitrate), there is a risk that the resulting packet size exceeds the capability of the error correction code.

This change removes magic constants to make full use of the of the error correcting bandwidth defined by the protocol and supported by Moonlight.

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

Sunshine implementation arbitrarily limits itself to 3 FEC blocks even though the protocol supports 4. This is fine for most use cases. However, when dealing with large payloads (low FPS, high bitrate), there is a risk that the resulting packet size exceeds the capability of the error correction code.
This change removes magic constants to makes full use of the of the protocol.
@CLAassistant

CLAassistant commented Jul 1, 2024

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@cgutman

Copy link
Copy Markdown
Collaborator

We should something like #1466 instead of this. The real problem is that our FEC groups are hardcoded rather than dynamic as they would be in that PR.

The issue is that proper FEC calculation is uncovering our longstanding issues with exhaustion of packet buffers in switches and NICs, so we need some packet pacing changes to support this properly.

@radugrecu97

This comment was marked as off-topic.

@ns6089

Copy link
Copy Markdown
Contributor

@cgutman If feedback-based packet pacing is not ready, maybe in the meantime we can hardcode 1GbE pacing? This should resolve most of the local streaming problems without adding significant amount of latency,

@cgutman

Copy link
Copy Markdown
Collaborator

Yeah, static pacing would probably be enough to unblock #1466

@ns6089

Copy link
Copy Markdown
Contributor

Yeah, static pacing would probably be enough

We have

boolsend_batch(batched_send_info_t &send_info);
boolsend(send_info_t &send_info);

in platform/common.h, if we add

boolsend_batch_paced(batched_send_info_t &send_info,
size_t pacing_block_size,
std::nanoseconds pacing_block_duration)

In your opinion, could it be reused later on for dynamic pacing, or you would rather prefer a different interface?

@ns6089

Copy link
Copy Markdown
Contributor

Or even something as simple as

boolsend_batch_paced(batched_send_info_t &send_info, double bytes_per_second);

Because I don't think we want to do busy-waiting and will have to rely on sleeps with limited precision. So the overall pacing will look like:

  1. Dump 1ms worth of data
  2. Sleep for 1ms
  3. Check how long we actually slept and adjust next batch size
  4. Dump, sleep and repeat

@ns6089

Copy link
Copy Markdown
Contributor

After looking more deeply into this, we may have to do resort to busy spin waiting after all.
RX/TX buffers seem to be more limited than I previously thought, we probably shouldn't batch more than 100 packets at once, maybe even less.
So yeah, this brings us back to the original function

boolsend_batch_paced(batched_send_info_t &send_info,
size_t pacing_block_size_in_packets,
std::nanoseconds pacing_block_duration)

@cgutman If you're ok with this implementation, I can work on it.

@cgutman

cgutman commented Jul 4, 2024

Copy link
Copy Markdown
Collaborator

@ns6089 yep, I agree. I think our batch size will have to be small enough that using OS sleep functionality will overshoot our intended sleep time (definitely on Windows, where the best you can get is rougly ~0.5ms). The way FEC blocks are computed today basically means that the FEC computation itself acts as a little busy wait between blocks, and that's what keeps the current system working today.

At least since we are actually intending to busy wait rather than doing actual work, we can put some code in the body of our busy waiting loop to allow other SMT threads to make progress, like:

#if defined(__i386__) || defined(__x86_64__)
__builtin_ia32_pause();
#elif defined(__aarch64__)
__yield();
#endif

Recent versions of Boost have sp_thread_pause() is exactly what we need, but I think it's only present on Boost 1.83 and later, so we can't rely on it.

@ns6089

Copy link
Copy Markdown
Contributor

At least since we are actually intending to busy wait rather than doing actual work, we can put some code in the body of our busy waiting loop to allow other SMT threads to make progress

Yeah, we can implement some spin_wait(std::chrono::nanoseconds time) function that will do these nop instructions in a loop and check time periodically, possibly dynamically adjusting the amount of instructions between time checks.

The way FEC blocks are computed today basically means that the FEC computation itself acts as a little busy wait between blocks, and that's what keeps the current system working today.

Interesting, we should probably make use of this and do the pacing in stream.cpp. I don't think FEC encoding can be done progressively, but packet headers and especially encryption totally can. @cgutman so unless you have something half-implemented and want to finish it, I will grab your FEC block size optimization PR and try to add pacing on top of it. Later on we should be able to expand it to some proper dynamic flow control, hopefully.

@ns6089ns6089 mentioned this pull request Jul 4, 2024
11 tasks
@ReenigneArcher

Copy link
Copy Markdown
Member

Closing in favor of #2803

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@CypherGrue@CLAassistant@cgutman@radugrecu97@ns6089@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Bump MAX_FEC_BLOCKS to 4 - #2787

Closed
CypherGrue wants to merge 1 commit into
LizardByte:masterfrom
CypherGrue:master
Closed

Bump MAX_FEC_BLOCKS to 4#2787
CypherGrue wants to merge 1 commit into
LizardByte:masterfrom
CypherGrue:master

Conversation

@CypherGrue

@CypherGrueCypherGrue commented Jul 1, 2024

Copy link
Copy Markdown

Description

Sunshine implementation arbitrarily limits itself to 3 FEC blocks even though the protocol supports 4. This is fine for typical use cases. However, when dealing with large payloads (low FPS, high bitrate), there is a risk that the resulting packet size exceeds the capability of the error correction code.

This change removes magic constants to make full use of the of the error correcting bandwidth defined by the protocol and supported by Moonlight.

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

Sunshine implementation arbitrarily limits itself to 3 FEC blocks even though the protocol supports 4. This is fine for most use cases. However, when dealing with large payloads (low FPS, high bitrate), there is a risk that the resulting packet size exceeds the capability of the error correction code.
This change removes magic constants to makes full use of the of the protocol.
@CLAassistant

CLAassistant commented Jul 1, 2024

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@cgutman

Copy link
Copy Markdown
Collaborator

We should something like #1466 instead of this. The real problem is that our FEC groups are hardcoded rather than dynamic as they would be in that PR.

The issue is that proper FEC calculation is uncovering our longstanding issues with exhaustion of packet buffers in switches and NICs, so we need some packet pacing changes to support this properly.

@radugrecu97

This comment was marked as off-topic.

@ns6089

Copy link
Copy Markdown
Contributor

@cgutman If feedback-based packet pacing is not ready, maybe in the meantime we can hardcode 1GbE pacing? This should resolve most of the local streaming problems without adding significant amount of latency,

@cgutman

Copy link
Copy Markdown
Collaborator

Yeah, static pacing would probably be enough to unblock #1466

@ns6089

Copy link
Copy Markdown
Contributor

Yeah, static pacing would probably be enough

We have

boolsend_batch(batched_send_info_t &send_info);
boolsend(send_info_t &send_info);

in platform/common.h, if we add

boolsend_batch_paced(batched_send_info_t &send_info,
size_t pacing_block_size,
std::nanoseconds pacing_block_duration)

In your opinion, could it be reused later on for dynamic pacing, or you would rather prefer a different interface?

@ns6089

Copy link
Copy Markdown
Contributor

Or even something as simple as

boolsend_batch_paced(batched_send_info_t &send_info, double bytes_per_second);

Because I don't think we want to do busy-waiting and will have to rely on sleeps with limited precision. So the overall pacing will look like:

  1. Dump 1ms worth of data
  2. Sleep for 1ms
  3. Check how long we actually slept and adjust next batch size
  4. Dump, sleep and repeat

@ns6089

Copy link
Copy Markdown
Contributor

After looking more deeply into this, we may have to do resort to busy spin waiting after all.
RX/TX buffers seem to be more limited than I previously thought, we probably shouldn't batch more than 100 packets at once, maybe even less.
So yeah, this brings us back to the original function

boolsend_batch_paced(batched_send_info_t &send_info,
size_t pacing_block_size_in_packets,
std::nanoseconds pacing_block_duration)

@cgutman If you're ok with this implementation, I can work on it.

@cgutman

cgutman commented Jul 4, 2024

Copy link
Copy Markdown
Collaborator

@ns6089 yep, I agree. I think our batch size will have to be small enough that using OS sleep functionality will overshoot our intended sleep time (definitely on Windows, where the best you can get is rougly ~0.5ms). The way FEC blocks are computed today basically means that the FEC computation itself acts as a little busy wait between blocks, and that's what keeps the current system working today.

At least since we are actually intending to busy wait rather than doing actual work, we can put some code in the body of our busy waiting loop to allow other SMT threads to make progress, like:

#if defined(__i386__) || defined(__x86_64__)
__builtin_ia32_pause();
#elif defined(__aarch64__)
__yield();
#endif

Recent versions of Boost have sp_thread_pause() is exactly what we need, but I think it's only present on Boost 1.83 and later, so we can't rely on it.

@ns6089

Copy link
Copy Markdown
Contributor

At least since we are actually intending to busy wait rather than doing actual work, we can put some code in the body of our busy waiting loop to allow other SMT threads to make progress

Yeah, we can implement some spin_wait(std::chrono::nanoseconds time) function that will do these nop instructions in a loop and check time periodically, possibly dynamically adjusting the amount of instructions between time checks.

The way FEC blocks are computed today basically means that the FEC computation itself acts as a little busy wait between blocks, and that's what keeps the current system working today.

Interesting, we should probably make use of this and do the pacing in stream.cpp. I don't think FEC encoding can be done progressively, but packet headers and especially encryption totally can. @cgutman so unless you have something half-implemented and want to finish it, I will grab your FEC block size optimization PR and try to add pacing on top of it. Later on we should be able to expand it to some proper dynamic flow control, hopefully.

@ns6089ns6089 mentioned this pull request Jul 4, 2024
11 tasks
@ReenigneArcher

Copy link
Copy Markdown
Member

Closing in favor of #2803

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@CypherGrue@CLAassistant@cgutman@radugrecu97@ns6089@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Bump MAX_FEC_BLOCKS to 4 - #2787

Closed
CypherGrue wants to merge 1 commit into
LizardByte:masterfrom
CypherGrue:master
Closed

Bump MAX_FEC_BLOCKS to 4#2787
CypherGrue wants to merge 1 commit into
LizardByte:masterfrom
CypherGrue:master

Conversation

@CypherGrue

@CypherGrueCypherGrue commented Jul 1, 2024

Copy link
Copy Markdown

Description

Sunshine implementation arbitrarily limits itself to 3 FEC blocks even though the protocol supports 4. This is fine for typical use cases. However, when dealing with large payloads (low FPS, high bitrate), there is a risk that the resulting packet size exceeds the capability of the error correction code.

This change removes magic constants to make full use of the of the error correcting bandwidth defined by the protocol and supported by Moonlight.

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

Sunshine implementation arbitrarily limits itself to 3 FEC blocks even though the protocol supports 4. This is fine for most use cases. However, when dealing with large payloads (low FPS, high bitrate), there is a risk that the resulting packet size exceeds the capability of the error correction code.
This change removes magic constants to makes full use of the of the protocol.
@CLAassistant

CLAassistant commented Jul 1, 2024

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@cgutman

Copy link
Copy Markdown
Collaborator

We should something like #1466 instead of this. The real problem is that our FEC groups are hardcoded rather than dynamic as they would be in that PR.

The issue is that proper FEC calculation is uncovering our longstanding issues with exhaustion of packet buffers in switches and NICs, so we need some packet pacing changes to support this properly.

@radugrecu97

This comment was marked as off-topic.

@ns6089

Copy link
Copy Markdown
Contributor

@cgutman If feedback-based packet pacing is not ready, maybe in the meantime we can hardcode 1GbE pacing? This should resolve most of the local streaming problems without adding significant amount of latency,

@cgutman

Copy link
Copy Markdown
Collaborator

Yeah, static pacing would probably be enough to unblock #1466

@ns6089

Copy link
Copy Markdown
Contributor

Yeah, static pacing would probably be enough

We have

boolsend_batch(batched_send_info_t &send_info);
boolsend(send_info_t &send_info);

in platform/common.h, if we add

boolsend_batch_paced(batched_send_info_t &send_info,
size_t pacing_block_size,
std::nanoseconds pacing_block_duration)

In your opinion, could it be reused later on for dynamic pacing, or you would rather prefer a different interface?

@ns6089

Copy link
Copy Markdown
Contributor

Or even something as simple as

boolsend_batch_paced(batched_send_info_t &send_info, double bytes_per_second);

Because I don't think we want to do busy-waiting and will have to rely on sleeps with limited precision. So the overall pacing will look like:

  1. Dump 1ms worth of data
  2. Sleep for 1ms
  3. Check how long we actually slept and adjust next batch size
  4. Dump, sleep and repeat

@ns6089

Copy link
Copy Markdown
Contributor

After looking more deeply into this, we may have to do resort to busy spin waiting after all.
RX/TX buffers seem to be more limited than I previously thought, we probably shouldn't batch more than 100 packets at once, maybe even less.
So yeah, this brings us back to the original function

boolsend_batch_paced(batched_send_info_t &send_info,
size_t pacing_block_size_in_packets,
std::nanoseconds pacing_block_duration)

@cgutman If you're ok with this implementation, I can work on it.

@cgutman

cgutman commented Jul 4, 2024

Copy link
Copy Markdown
Collaborator

@ns6089 yep, I agree. I think our batch size will have to be small enough that using OS sleep functionality will overshoot our intended sleep time (definitely on Windows, where the best you can get is rougly ~0.5ms). The way FEC blocks are computed today basically means that the FEC computation itself acts as a little busy wait between blocks, and that's what keeps the current system working today.

At least since we are actually intending to busy wait rather than doing actual work, we can put some code in the body of our busy waiting loop to allow other SMT threads to make progress, like:

#if defined(__i386__) || defined(__x86_64__)
__builtin_ia32_pause();
#elif defined(__aarch64__)
__yield();
#endif

Recent versions of Boost have sp_thread_pause() is exactly what we need, but I think it's only present on Boost 1.83 and later, so we can't rely on it.

@ns6089

Copy link
Copy Markdown
Contributor

At least since we are actually intending to busy wait rather than doing actual work, we can put some code in the body of our busy waiting loop to allow other SMT threads to make progress

Yeah, we can implement some spin_wait(std::chrono::nanoseconds time) function that will do these nop instructions in a loop and check time periodically, possibly dynamically adjusting the amount of instructions between time checks.

The way FEC blocks are computed today basically means that the FEC computation itself acts as a little busy wait between blocks, and that's what keeps the current system working today.

Interesting, we should probably make use of this and do the pacing in stream.cpp. I don't think FEC encoding can be done progressively, but packet headers and especially encryption totally can. @cgutman so unless you have something half-implemented and want to finish it, I will grab your FEC block size optimization PR and try to add pacing on top of it. Later on we should be able to expand it to some proper dynamic flow control, hopefully.

@ns6089ns6089 mentioned this pull request Jul 4, 2024
11 tasks
@ReenigneArcher

Copy link
Copy Markdown
Member

Closing in favor of #2803

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@CypherGrue@CLAassistant@cgutman@radugrecu97@ns6089@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Bump MAX_FEC_BLOCKS to 4 - #2787

Closed
CypherGrue wants to merge 1 commit into
LizardByte:masterfrom
CypherGrue:master
Closed

Bump MAX_FEC_BLOCKS to 4#2787
CypherGrue wants to merge 1 commit into
LizardByte:masterfrom
CypherGrue:master

Conversation

@CypherGrue

@CypherGrueCypherGrue commented Jul 1, 2024

Copy link
Copy Markdown

Description

Sunshine implementation arbitrarily limits itself to 3 FEC blocks even though the protocol supports 4. This is fine for typical use cases. However, when dealing with large payloads (low FPS, high bitrate), there is a risk that the resulting packet size exceeds the capability of the error correction code.

This change removes magic constants to make full use of the of the error correcting bandwidth defined by the protocol and supported by Moonlight.

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

Sunshine implementation arbitrarily limits itself to 3 FEC blocks even though the protocol supports 4. This is fine for most use cases. However, when dealing with large payloads (low FPS, high bitrate), there is a risk that the resulting packet size exceeds the capability of the error correction code.
This change removes magic constants to makes full use of the of the protocol.
@CLAassistant

CLAassistant commented Jul 1, 2024

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@cgutman

Copy link
Copy Markdown
Collaborator

We should something like #1466 instead of this. The real problem is that our FEC groups are hardcoded rather than dynamic as they would be in that PR.

The issue is that proper FEC calculation is uncovering our longstanding issues with exhaustion of packet buffers in switches and NICs, so we need some packet pacing changes to support this properly.

@radugrecu97

This comment was marked as off-topic.

@ns6089

Copy link
Copy Markdown
Contributor

@cgutman If feedback-based packet pacing is not ready, maybe in the meantime we can hardcode 1GbE pacing? This should resolve most of the local streaming problems without adding significant amount of latency,

@cgutman

Copy link
Copy Markdown
Collaborator

Yeah, static pacing would probably be enough to unblock #1466

@ns6089

Copy link
Copy Markdown
Contributor

Yeah, static pacing would probably be enough

We have

boolsend_batch(batched_send_info_t &send_info);
boolsend(send_info_t &send_info);

in platform/common.h, if we add

boolsend_batch_paced(batched_send_info_t &send_info,
size_t pacing_block_size,
std::nanoseconds pacing_block_duration)

In your opinion, could it be reused later on for dynamic pacing, or you would rather prefer a different interface?

@ns6089

Copy link
Copy Markdown
Contributor

Or even something as simple as

boolsend_batch_paced(batched_send_info_t &send_info, double bytes_per_second);

Because I don't think we want to do busy-waiting and will have to rely on sleeps with limited precision. So the overall pacing will look like:

  1. Dump 1ms worth of data
  2. Sleep for 1ms
  3. Check how long we actually slept and adjust next batch size
  4. Dump, sleep and repeat

@ns6089

Copy link
Copy Markdown
Contributor

After looking more deeply into this, we may have to do resort to busy spin waiting after all.
RX/TX buffers seem to be more limited than I previously thought, we probably shouldn't batch more than 100 packets at once, maybe even less.
So yeah, this brings us back to the original function

boolsend_batch_paced(batched_send_info_t &send_info,
size_t pacing_block_size_in_packets,
std::nanoseconds pacing_block_duration)

@cgutman If you're ok with this implementation, I can work on it.

@cgutman

cgutman commented Jul 4, 2024

Copy link
Copy Markdown
Collaborator

@ns6089 yep, I agree. I think our batch size will have to be small enough that using OS sleep functionality will overshoot our intended sleep time (definitely on Windows, where the best you can get is rougly ~0.5ms). The way FEC blocks are computed today basically means that the FEC computation itself acts as a little busy wait between blocks, and that's what keeps the current system working today.

At least since we are actually intending to busy wait rather than doing actual work, we can put some code in the body of our busy waiting loop to allow other SMT threads to make progress, like:

#if defined(__i386__) || defined(__x86_64__)
__builtin_ia32_pause();
#elif defined(__aarch64__)
__yield();
#endif

Recent versions of Boost have sp_thread_pause() is exactly what we need, but I think it's only present on Boost 1.83 and later, so we can't rely on it.

@ns6089

Copy link
Copy Markdown
Contributor

At least since we are actually intending to busy wait rather than doing actual work, we can put some code in the body of our busy waiting loop to allow other SMT threads to make progress

Yeah, we can implement some spin_wait(std::chrono::nanoseconds time) function that will do these nop instructions in a loop and check time periodically, possibly dynamically adjusting the amount of instructions between time checks.

The way FEC blocks are computed today basically means that the FEC computation itself acts as a little busy wait between blocks, and that's what keeps the current system working today.

Interesting, we should probably make use of this and do the pacing in stream.cpp. I don't think FEC encoding can be done progressively, but packet headers and especially encryption totally can. @cgutman so unless you have something half-implemented and want to finish it, I will grab your FEC block size optimization PR and try to add pacing on top of it. Later on we should be able to expand it to some proper dynamic flow control, hopefully.

@ns6089ns6089 mentioned this pull request Jul 4, 2024
11 tasks
@ReenigneArcher

Copy link
Copy Markdown
Member

Closing in favor of #2803

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@CypherGrue@CLAassistant@cgutman@radugrecu97@ns6089@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Bump MAX_FEC_BLOCKS to 4 - #2787

Closed
CypherGrue wants to merge 1 commit into
LizardByte:masterfrom
CypherGrue:master
Closed

Bump MAX_FEC_BLOCKS to 4#2787
CypherGrue wants to merge 1 commit into
LizardByte:masterfrom
CypherGrue:master

Conversation

@CypherGrue

@CypherGrueCypherGrue commented Jul 1, 2024

Copy link
Copy Markdown

Description

Sunshine implementation arbitrarily limits itself to 3 FEC blocks even though the protocol supports 4. This is fine for typical use cases. However, when dealing with large payloads (low FPS, high bitrate), there is a risk that the resulting packet size exceeds the capability of the error correction code.

This change removes magic constants to make full use of the of the error correcting bandwidth defined by the protocol and supported by Moonlight.

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

Sunshine implementation arbitrarily limits itself to 3 FEC blocks even though the protocol supports 4. This is fine for most use cases. However, when dealing with large payloads (low FPS, high bitrate), there is a risk that the resulting packet size exceeds the capability of the error correction code.
This change removes magic constants to makes full use of the of the protocol.
@CLAassistant

CLAassistant commented Jul 1, 2024

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@cgutman

Copy link
Copy Markdown
Collaborator

We should something like #1466 instead of this. The real problem is that our FEC groups are hardcoded rather than dynamic as they would be in that PR.

The issue is that proper FEC calculation is uncovering our longstanding issues with exhaustion of packet buffers in switches and NICs, so we need some packet pacing changes to support this properly.

@radugrecu97

This comment was marked as off-topic.

@ns6089

Copy link
Copy Markdown
Contributor

@cgutman If feedback-based packet pacing is not ready, maybe in the meantime we can hardcode 1GbE pacing? This should resolve most of the local streaming problems without adding significant amount of latency,

@cgutman

Copy link
Copy Markdown
Collaborator

Yeah, static pacing would probably be enough to unblock #1466

@ns6089

Copy link
Copy Markdown
Contributor

Yeah, static pacing would probably be enough

We have

boolsend_batch(batched_send_info_t &send_info);
boolsend(send_info_t &send_info);

in platform/common.h, if we add

boolsend_batch_paced(batched_send_info_t &send_info,
size_t pacing_block_size,
std::nanoseconds pacing_block_duration)

In your opinion, could it be reused later on for dynamic pacing, or you would rather prefer a different interface?

@ns6089

Copy link
Copy Markdown
Contributor

Or even something as simple as

boolsend_batch_paced(batched_send_info_t &send_info, double bytes_per_second);

Because I don't think we want to do busy-waiting and will have to rely on sleeps with limited precision. So the overall pacing will look like:

  1. Dump 1ms worth of data
  2. Sleep for 1ms
  3. Check how long we actually slept and adjust next batch size
  4. Dump, sleep and repeat

@ns6089

Copy link
Copy Markdown
Contributor

After looking more deeply into this, we may have to do resort to busy spin waiting after all.
RX/TX buffers seem to be more limited than I previously thought, we probably shouldn't batch more than 100 packets at once, maybe even less.
So yeah, this brings us back to the original function

boolsend_batch_paced(batched_send_info_t &send_info,
size_t pacing_block_size_in_packets,
std::nanoseconds pacing_block_duration)

@cgutman If you're ok with this implementation, I can work on it.

@cgutman

cgutman commented Jul 4, 2024

Copy link
Copy Markdown
Collaborator

@ns6089 yep, I agree. I think our batch size will have to be small enough that using OS sleep functionality will overshoot our intended sleep time (definitely on Windows, where the best you can get is rougly ~0.5ms). The way FEC blocks are computed today basically means that the FEC computation itself acts as a little busy wait between blocks, and that's what keeps the current system working today.

At least since we are actually intending to busy wait rather than doing actual work, we can put some code in the body of our busy waiting loop to allow other SMT threads to make progress, like:

#if defined(__i386__) || defined(__x86_64__)
__builtin_ia32_pause();
#elif defined(__aarch64__)
__yield();
#endif

Recent versions of Boost have sp_thread_pause() is exactly what we need, but I think it's only present on Boost 1.83 and later, so we can't rely on it.

@ns6089

Copy link
Copy Markdown
Contributor

At least since we are actually intending to busy wait rather than doing actual work, we can put some code in the body of our busy waiting loop to allow other SMT threads to make progress

Yeah, we can implement some spin_wait(std::chrono::nanoseconds time) function that will do these nop instructions in a loop and check time periodically, possibly dynamically adjusting the amount of instructions between time checks.

The way FEC blocks are computed today basically means that the FEC computation itself acts as a little busy wait between blocks, and that's what keeps the current system working today.

Interesting, we should probably make use of this and do the pacing in stream.cpp. I don't think FEC encoding can be done progressively, but packet headers and especially encryption totally can. @cgutman so unless you have something half-implemented and want to finish it, I will grab your FEC block size optimization PR and try to add pacing on top of it. Later on we should be able to expand it to some proper dynamic flow control, hopefully.

@ns6089ns6089 mentioned this pull request Jul 4, 2024
11 tasks
@ReenigneArcher

Copy link
Copy Markdown
Member

Closing in favor of #2803

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@CypherGrue@CLAassistant@cgutman@radugrecu97@ns6089@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Bump MAX_FEC_BLOCKS to 4 - #2787

Closed
CypherGrue wants to merge 1 commit into
LizardByte:masterfrom
CypherGrue:master
Closed

Bump MAX_FEC_BLOCKS to 4#2787
CypherGrue wants to merge 1 commit into
LizardByte:masterfrom
CypherGrue:master

Conversation

@CypherGrue

@CypherGrueCypherGrue commented Jul 1, 2024

Copy link
Copy Markdown

Description

Sunshine implementation arbitrarily limits itself to 3 FEC blocks even though the protocol supports 4. This is fine for typical use cases. However, when dealing with large payloads (low FPS, high bitrate), there is a risk that the resulting packet size exceeds the capability of the error correction code.

This change removes magic constants to make full use of the of the error correcting bandwidth defined by the protocol and supported by Moonlight.

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

Sunshine implementation arbitrarily limits itself to 3 FEC blocks even though the protocol supports 4. This is fine for most use cases. However, when dealing with large payloads (low FPS, high bitrate), there is a risk that the resulting packet size exceeds the capability of the error correction code.
This change removes magic constants to makes full use of the of the protocol.
@CLAassistant

CLAassistant commented Jul 1, 2024

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@cgutman

Copy link
Copy Markdown
Collaborator

We should something like #1466 instead of this. The real problem is that our FEC groups are hardcoded rather than dynamic as they would be in that PR.

The issue is that proper FEC calculation is uncovering our longstanding issues with exhaustion of packet buffers in switches and NICs, so we need some packet pacing changes to support this properly.

@radugrecu97

This comment was marked as off-topic.

@ns6089

Copy link
Copy Markdown
Contributor

@cgutman If feedback-based packet pacing is not ready, maybe in the meantime we can hardcode 1GbE pacing? This should resolve most of the local streaming problems without adding significant amount of latency,

@cgutman

Copy link
Copy Markdown
Collaborator

Yeah, static pacing would probably be enough to unblock #1466

@ns6089

Copy link
Copy Markdown
Contributor

Yeah, static pacing would probably be enough

We have

boolsend_batch(batched_send_info_t &send_info);
boolsend(send_info_t &send_info);

in platform/common.h, if we add

boolsend_batch_paced(batched_send_info_t &send_info,
size_t pacing_block_size,
std::nanoseconds pacing_block_duration)

In your opinion, could it be reused later on for dynamic pacing, or you would rather prefer a different interface?

@ns6089

Copy link
Copy Markdown
Contributor

Or even something as simple as

boolsend_batch_paced(batched_send_info_t &send_info, double bytes_per_second);

Because I don't think we want to do busy-waiting and will have to rely on sleeps with limited precision. So the overall pacing will look like:

  1. Dump 1ms worth of data
  2. Sleep for 1ms
  3. Check how long we actually slept and adjust next batch size
  4. Dump, sleep and repeat

@ns6089

Copy link
Copy Markdown
Contributor

After looking more deeply into this, we may have to do resort to busy spin waiting after all.
RX/TX buffers seem to be more limited than I previously thought, we probably shouldn't batch more than 100 packets at once, maybe even less.
So yeah, this brings us back to the original function

boolsend_batch_paced(batched_send_info_t &send_info,
size_t pacing_block_size_in_packets,
std::nanoseconds pacing_block_duration)

@cgutman If you're ok with this implementation, I can work on it.

@cgutman

cgutman commented Jul 4, 2024

Copy link
Copy Markdown
Collaborator

@ns6089 yep, I agree. I think our batch size will have to be small enough that using OS sleep functionality will overshoot our intended sleep time (definitely on Windows, where the best you can get is rougly ~0.5ms). The way FEC blocks are computed today basically means that the FEC computation itself acts as a little busy wait between blocks, and that's what keeps the current system working today.

At least since we are actually intending to busy wait rather than doing actual work, we can put some code in the body of our busy waiting loop to allow other SMT threads to make progress, like:

#if defined(__i386__) || defined(__x86_64__)
__builtin_ia32_pause();
#elif defined(__aarch64__)
__yield();
#endif

Recent versions of Boost have sp_thread_pause() is exactly what we need, but I think it's only present on Boost 1.83 and later, so we can't rely on it.

@ns6089

Copy link
Copy Markdown
Contributor

At least since we are actually intending to busy wait rather than doing actual work, we can put some code in the body of our busy waiting loop to allow other SMT threads to make progress

Yeah, we can implement some spin_wait(std::chrono::nanoseconds time) function that will do these nop instructions in a loop and check time periodically, possibly dynamically adjusting the amount of instructions between time checks.

The way FEC blocks are computed today basically means that the FEC computation itself acts as a little busy wait between blocks, and that's what keeps the current system working today.

Interesting, we should probably make use of this and do the pacing in stream.cpp. I don't think FEC encoding can be done progressively, but packet headers and especially encryption totally can. @cgutman so unless you have something half-implemented and want to finish it, I will grab your FEC block size optimization PR and try to add pacing on top of it. Later on we should be able to expand it to some proper dynamic flow control, hopefully.

@ns6089ns6089 mentioned this pull request Jul 4, 2024
11 tasks
@ReenigneArcher

Copy link
Copy Markdown
Member

Closing in favor of #2803

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@CypherGrue@CLAassistant@cgutman@radugrecu97@ns6089@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Bump MAX_FEC_BLOCKS to 4 - #2787

Closed
CypherGrue wants to merge 1 commit into
LizardByte:masterfrom
CypherGrue:master
Closed

Bump MAX_FEC_BLOCKS to 4#2787
CypherGrue wants to merge 1 commit into
LizardByte:masterfrom
CypherGrue:master

Conversation

@CypherGrue

@CypherGrueCypherGrue commented Jul 1, 2024

Copy link
Copy Markdown

Description

Sunshine implementation arbitrarily limits itself to 3 FEC blocks even though the protocol supports 4. This is fine for typical use cases. However, when dealing with large payloads (low FPS, high bitrate), there is a risk that the resulting packet size exceeds the capability of the error correction code.

This change removes magic constants to make full use of the of the error correcting bandwidth defined by the protocol and supported by Moonlight.

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

Sunshine implementation arbitrarily limits itself to 3 FEC blocks even though the protocol supports 4. This is fine for most use cases. However, when dealing with large payloads (low FPS, high bitrate), there is a risk that the resulting packet size exceeds the capability of the error correction code.
This change removes magic constants to makes full use of the of the protocol.
@CLAassistant

CLAassistant commented Jul 1, 2024

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@cgutman

Copy link
Copy Markdown
Collaborator

We should something like #1466 instead of this. The real problem is that our FEC groups are hardcoded rather than dynamic as they would be in that PR.

The issue is that proper FEC calculation is uncovering our longstanding issues with exhaustion of packet buffers in switches and NICs, so we need some packet pacing changes to support this properly.

@radugrecu97

This comment was marked as off-topic.

@ns6089

Copy link
Copy Markdown
Contributor

@cgutman If feedback-based packet pacing is not ready, maybe in the meantime we can hardcode 1GbE pacing? This should resolve most of the local streaming problems without adding significant amount of latency,

@cgutman

Copy link
Copy Markdown
Collaborator

Yeah, static pacing would probably be enough to unblock #1466

@ns6089

Copy link
Copy Markdown
Contributor

Yeah, static pacing would probably be enough

We have

boolsend_batch(batched_send_info_t &send_info);
boolsend(send_info_t &send_info);

in platform/common.h, if we add

boolsend_batch_paced(batched_send_info_t &send_info,
size_t pacing_block_size,
std::nanoseconds pacing_block_duration)

In your opinion, could it be reused later on for dynamic pacing, or you would rather prefer a different interface?

@ns6089

Copy link
Copy Markdown
Contributor

Or even something as simple as

boolsend_batch_paced(batched_send_info_t &send_info, double bytes_per_second);

Because I don't think we want to do busy-waiting and will have to rely on sleeps with limited precision. So the overall pacing will look like:

  1. Dump 1ms worth of data
  2. Sleep for 1ms
  3. Check how long we actually slept and adjust next batch size
  4. Dump, sleep and repeat

@ns6089

Copy link
Copy Markdown
Contributor

After looking more deeply into this, we may have to do resort to busy spin waiting after all.
RX/TX buffers seem to be more limited than I previously thought, we probably shouldn't batch more than 100 packets at once, maybe even less.
So yeah, this brings us back to the original function

boolsend_batch_paced(batched_send_info_t &send_info,
size_t pacing_block_size_in_packets,
std::nanoseconds pacing_block_duration)

@cgutman If you're ok with this implementation, I can work on it.

@cgutman

cgutman commented Jul 4, 2024

Copy link
Copy Markdown
Collaborator

@ns6089 yep, I agree. I think our batch size will have to be small enough that using OS sleep functionality will overshoot our intended sleep time (definitely on Windows, where the best you can get is rougly ~0.5ms). The way FEC blocks are computed today basically means that the FEC computation itself acts as a little busy wait between blocks, and that's what keeps the current system working today.

At least since we are actually intending to busy wait rather than doing actual work, we can put some code in the body of our busy waiting loop to allow other SMT threads to make progress, like:

#if defined(__i386__) || defined(__x86_64__)
__builtin_ia32_pause();
#elif defined(__aarch64__)
__yield();
#endif

Recent versions of Boost have sp_thread_pause() is exactly what we need, but I think it's only present on Boost 1.83 and later, so we can't rely on it.

@ns6089

Copy link
Copy Markdown
Contributor

At least since we are actually intending to busy wait rather than doing actual work, we can put some code in the body of our busy waiting loop to allow other SMT threads to make progress

Yeah, we can implement some spin_wait(std::chrono::nanoseconds time) function that will do these nop instructions in a loop and check time periodically, possibly dynamically adjusting the amount of instructions between time checks.

The way FEC blocks are computed today basically means that the FEC computation itself acts as a little busy wait between blocks, and that's what keeps the current system working today.

Interesting, we should probably make use of this and do the pacing in stream.cpp. I don't think FEC encoding can be done progressively, but packet headers and especially encryption totally can. @cgutman so unless you have something half-implemented and want to finish it, I will grab your FEC block size optimization PR and try to add pacing on top of it. Later on we should be able to expand it to some proper dynamic flow control, hopefully.

@ns6089ns6089 mentioned this pull request Jul 4, 2024
11 tasks
@ReenigneArcher

Copy link
Copy Markdown
Member

Closing in favor of #2803

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@CypherGrue@CLAassistant@cgutman@radugrecu97@ns6089@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Bump MAX_FEC_BLOCKS to 4 - #2787

Closed
CypherGrue wants to merge 1 commit into
LizardByte:masterfrom
CypherGrue:master
Closed

Bump MAX_FEC_BLOCKS to 4#2787
CypherGrue wants to merge 1 commit into
LizardByte:masterfrom
CypherGrue:master

Conversation

@CypherGrue

@CypherGrueCypherGrue commented Jul 1, 2024

Copy link
Copy Markdown

Description

Sunshine implementation arbitrarily limits itself to 3 FEC blocks even though the protocol supports 4. This is fine for typical use cases. However, when dealing with large payloads (low FPS, high bitrate), there is a risk that the resulting packet size exceeds the capability of the error correction code.

This change removes magic constants to make full use of the of the error correcting bandwidth defined by the protocol and supported by Moonlight.

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

Sunshine implementation arbitrarily limits itself to 3 FEC blocks even though the protocol supports 4. This is fine for most use cases. However, when dealing with large payloads (low FPS, high bitrate), there is a risk that the resulting packet size exceeds the capability of the error correction code.
This change removes magic constants to makes full use of the of the protocol.
@CLAassistant

CLAassistant commented Jul 1, 2024

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@cgutman

Copy link
Copy Markdown
Collaborator

We should something like #1466 instead of this. The real problem is that our FEC groups are hardcoded rather than dynamic as they would be in that PR.

The issue is that proper FEC calculation is uncovering our longstanding issues with exhaustion of packet buffers in switches and NICs, so we need some packet pacing changes to support this properly.

@radugrecu97

This comment was marked as off-topic.

@ns6089

Copy link
Copy Markdown
Contributor

@cgutman If feedback-based packet pacing is not ready, maybe in the meantime we can hardcode 1GbE pacing? This should resolve most of the local streaming problems without adding significant amount of latency,

@cgutman

Copy link
Copy Markdown
Collaborator

Yeah, static pacing would probably be enough to unblock #1466

@ns6089

Copy link
Copy Markdown
Contributor

Yeah, static pacing would probably be enough

We have

boolsend_batch(batched_send_info_t &send_info);
boolsend(send_info_t &send_info);

in platform/common.h, if we add

boolsend_batch_paced(batched_send_info_t &send_info,
size_t pacing_block_size,
std::nanoseconds pacing_block_duration)

In your opinion, could it be reused later on for dynamic pacing, or you would rather prefer a different interface?

@ns6089

Copy link
Copy Markdown
Contributor

Or even something as simple as

boolsend_batch_paced(batched_send_info_t &send_info, double bytes_per_second);

Because I don't think we want to do busy-waiting and will have to rely on sleeps with limited precision. So the overall pacing will look like:

  1. Dump 1ms worth of data
  2. Sleep for 1ms
  3. Check how long we actually slept and adjust next batch size
  4. Dump, sleep and repeat

@ns6089

Copy link
Copy Markdown
Contributor

After looking more deeply into this, we may have to do resort to busy spin waiting after all.
RX/TX buffers seem to be more limited than I previously thought, we probably shouldn't batch more than 100 packets at once, maybe even less.
So yeah, this brings us back to the original function

boolsend_batch_paced(batched_send_info_t &send_info,
size_t pacing_block_size_in_packets,
std::nanoseconds pacing_block_duration)

@cgutman If you're ok with this implementation, I can work on it.

@cgutman

cgutman commented Jul 4, 2024

Copy link
Copy Markdown
Collaborator

@ns6089 yep, I agree. I think our batch size will have to be small enough that using OS sleep functionality will overshoot our intended sleep time (definitely on Windows, where the best you can get is rougly ~0.5ms). The way FEC blocks are computed today basically means that the FEC computation itself acts as a little busy wait between blocks, and that's what keeps the current system working today.

At least since we are actually intending to busy wait rather than doing actual work, we can put some code in the body of our busy waiting loop to allow other SMT threads to make progress, like:

#if defined(__i386__) || defined(__x86_64__)
__builtin_ia32_pause();
#elif defined(__aarch64__)
__yield();
#endif

Recent versions of Boost have sp_thread_pause() is exactly what we need, but I think it's only present on Boost 1.83 and later, so we can't rely on it.

@ns6089

Copy link
Copy Markdown
Contributor

At least since we are actually intending to busy wait rather than doing actual work, we can put some code in the body of our busy waiting loop to allow other SMT threads to make progress

Yeah, we can implement some spin_wait(std::chrono::nanoseconds time) function that will do these nop instructions in a loop and check time periodically, possibly dynamically adjusting the amount of instructions between time checks.

The way FEC blocks are computed today basically means that the FEC computation itself acts as a little busy wait between blocks, and that's what keeps the current system working today.

Interesting, we should probably make use of this and do the pacing in stream.cpp. I don't think FEC encoding can be done progressively, but packet headers and especially encryption totally can. @cgutman so unless you have something half-implemented and want to finish it, I will grab your FEC block size optimization PR and try to add pacing on top of it. Later on we should be able to expand it to some proper dynamic flow control, hopefully.

@ns6089ns6089 mentioned this pull request Jul 4, 2024
11 tasks
@ReenigneArcher

Copy link
Copy Markdown
Member

Closing in favor of #2803

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants

@CypherGrue@CLAassistant@cgutman@radugrecu97@ns6089@ReenigneArcher