feat(win/video): add support for recombined YUV444 encoding - #2760

Closed
ns6089 wants to merge 1 commit into
LizardByte:masterfrom
ns6089:yuv444in420
Closed

feat(win/video): add support for recombined YUV444 encoding#2760
ns6089 wants to merge 1 commit into
LizardByte:masterfrom
ns6089:yuv444in420

Conversation

@ns6089

@ns6089ns6089 commented Jun 26, 2024

Copy link
Copy Markdown
Contributor

Description

The continuation of #2533. It's possible to emulate YUV 4:4:4 on gpus that don't support it natively by doubling the YUV 4:2:0 pixel count and running custom recombination shaders on both encoding and decoding side. Like Microsoft did it in MS-RDPEGFX.

Prototype stage. Requires changes on moonlight's side: I currently have custom libplacebo mpv shader implemented for plvk backend, in the future it should be possible to add Direct3D11 and OpenGL shaders.

https://github.com/ns6089/Sunshine/compare/yuv444..yuv444in420

moonlight-common-c pull request: TBD
moonlight-qt pull request: TBD, testing branch https://github.com/ns6089/moonlight-qt/tree/yuv444in420

What works and what doesn't

  1. First prototype, left half of U_src and V_src planes in Y_out. Good DCT, bad motion compensation.
  2. Second prototype. U_src in Y_out. V_src is spread across U_out and V_out in a pattern that is spatially consistent with Y_out. Good motion compensation, relatively fat DCT on U_out and V_out due to high frequencies.
  3. Third prototype, dropped. Maybe can slightly improve the DCT by running 1/4 of V through averaging low pass filter.

To Do

  • decide what to do with resolutions not divisible by 2
  • decide in which part of the protocol dimension doubling will be taking place, e.g. will the client request the doubled dimension or will it be done implicitly

Screenshot

before
after

Issues Fixed or Closed

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

Awesomely crazy.
Could there be anything worth doing with an emulated 4:2:2 stream then? Like, I don't know, slightly lower recombination overhead, or lower bandwidth requirements?

Or perhaps not hitting encoding limits at higher resolutions. Like, is my understanding correct that this pixel doubling would not allow for 1440p on (say) older VCE versions that max out at 4K?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

I don't think anyone but Intel supports 4:2:2. About 4K limit, 1440p might still work depending on how exactly said limit is implemented, the overall pixel count stays within 4K range.


amd

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

I don't think anyone but Intel supports 4:2:2.

To be honest, I was more thinking of TVs than computers here. It's a mixed bag even there, but still it's not so rare.
But now that you mention pcs, decoding is much lighter on the cpu than encoding. I don't think that would usually be a deal breaker. Or nevertheless, couldn't the client-side recombination just fake to be 4:4:4 then? Or would whatever empty padding you add ruin the image more than the results you could get with just plain 4:2:0?

About 4K limit, 1440p might still work depending on how exactly said limit is implemented, the overall pixel count stays within 4K range.

You mean if the limit is actually implemented like 4096x2160 (usual old amd) vs 4096x4096 (usual old nvidia)?
Or can you really call it a day just as long as the supported total pixel count, whatever the "shape", is 7.372.800 (2560x1440x2) or more?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

You can't encode 4:2:2 on nvidia gpus, implementing a path exclusively for intel will be too expensive.

Or can you really call it a day just as long as the supported total pixel count, whatever the "shape", is 7.372.800 (2560x1440x2) or more?

I'm already calling it a day 😎
Doubling one dimension allows to minimize discontinuities in motion estimation, in contrast to tiling. Current half-naive implementation for example has single motion estimation vertical "seam" in U and V planes.

// Y U V// +-------+ +---+ +---+// | | | | | |// | Y | |UR | |VR |// | | | | | |// +---+---+ +---+ +---+// | | |// |UL |VL |// | | |// +---+---+

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

You can't encode 4:2:2 on nvidia gpus, implementing a path exclusively for intel will be too expensive.

You can't encode 4:4:4 on amd gpus either, and yet this is what this PR is about isn't it?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

Personally, I don't see a point in supporting recombination into 4:2:2
It will still have visible artifacts while having computational overhead close to 4:4:4 and significant amount of additional development time. And this development time will be multiplied by the amount of distinct clients,

@mirh

mirh commented Jun 28, 2024

Copy link
Copy Markdown

I mean, sure, of course this is already miraculous.
I was just trying to think outside the box (4:2:2 is still subpar, but even the worst case scenario starts to be bearable instead).

If any I guess the improvement isn't that clear cut, because unlike with a direct cable connection it's not like there aren't already compression artifacts anyway. So if 4:4:4 couldn't fit in some whatever doubled 4:2:0 4K scenario, just lowering the resolution could also be a possible (and if not any easily immediate) alternative?

@ns6089
ns6089force-pushed the yuv444in420 branch 4 times, most recently from fc48f22 to 3a1115dCompareJune 30, 2024 08:50
@ns6089
ns6089force-pushed the yuv444in420 branch 2 times, most recently from 2cc6a6a to a4ffe24CompareAugust 1, 2024 07:52
@ns6089ns6089 changed the title Support recombined YUV 4:4:4 encoding (Prototype, Windows-only for now) feat(win/video): add support for recombined YUV444 encodingAug 22, 2024
@sonarqubecloud

Copy link
Copy Markdown

@ns6089

Copy link
Copy Markdown
ContributorAuthor

The code in this pull request is Not a Contribution under LizardByte Individual Contributor License Agreement.
The code in this pull request is shared under GNU GENERAL PUBLIC LICENSE Version 3.

The feature itself is completed on sunshine side.

danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 21, 2026
Convert absolute mouse input to relative deltas, accumulating fractional
deltas so slow (sub-pixel) motion still moves the cursor on both axes.
Moonlight TV clients send absolute positions; emitting relative deltas makes
the COSMIC magnifier follow the cursor (PointerMotionAbsolute never updates
the zoom focal point — pop-os/cosmic-comp LizardByte#2760).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (6ad5f97):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Dead reckoning is the only
thing in the smooth motion path (1:1); the estimate snaps to the real
cursor position (published by the KMS cursor plane via platf::kms_cursor_*
atomics) while the client is idle; a raw-saturated axis targets the
matching host edge, with fresh feedback authoritative on that axis —
phantom walls cannot persist. History: v1 walls, v2 partial, v3 jumps,
v4 pseudo-acceleration, v5 zero-slide latch — see CACHYOS.md.
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-9 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 27, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@ns6089@mirh@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat(win/video): add support for recombined YUV444 encoding - #2760

Closed
ns6089 wants to merge 1 commit into
LizardByte:masterfrom
ns6089:yuv444in420
Closed

feat(win/video): add support for recombined YUV444 encoding#2760
ns6089 wants to merge 1 commit into
LizardByte:masterfrom
ns6089:yuv444in420

Conversation

@ns6089

@ns6089ns6089 commented Jun 26, 2024

Copy link
Copy Markdown
Contributor

Description

The continuation of #2533. It's possible to emulate YUV 4:4:4 on gpus that don't support it natively by doubling the YUV 4:2:0 pixel count and running custom recombination shaders on both encoding and decoding side. Like Microsoft did it in MS-RDPEGFX.

Prototype stage. Requires changes on moonlight's side: I currently have custom libplacebo mpv shader implemented for plvk backend, in the future it should be possible to add Direct3D11 and OpenGL shaders.

https://github.com/ns6089/Sunshine/compare/yuv444..yuv444in420

moonlight-common-c pull request: TBD
moonlight-qt pull request: TBD, testing branch https://github.com/ns6089/moonlight-qt/tree/yuv444in420

What works and what doesn't

  1. First prototype, left half of U_src and V_src planes in Y_out. Good DCT, bad motion compensation.
  2. Second prototype. U_src in Y_out. V_src is spread across U_out and V_out in a pattern that is spatially consistent with Y_out. Good motion compensation, relatively fat DCT on U_out and V_out due to high frequencies.
  3. Third prototype, dropped. Maybe can slightly improve the DCT by running 1/4 of V through averaging low pass filter.

To Do

  • decide what to do with resolutions not divisible by 2
  • decide in which part of the protocol dimension doubling will be taking place, e.g. will the client request the doubled dimension or will it be done implicitly

Screenshot

before
after

Issues Fixed or Closed

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

Awesomely crazy.
Could there be anything worth doing with an emulated 4:2:2 stream then? Like, I don't know, slightly lower recombination overhead, or lower bandwidth requirements?

Or perhaps not hitting encoding limits at higher resolutions. Like, is my understanding correct that this pixel doubling would not allow for 1440p on (say) older VCE versions that max out at 4K?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

I don't think anyone but Intel supports 4:2:2. About 4K limit, 1440p might still work depending on how exactly said limit is implemented, the overall pixel count stays within 4K range.


amd

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

I don't think anyone but Intel supports 4:2:2.

To be honest, I was more thinking of TVs than computers here. It's a mixed bag even there, but still it's not so rare.
But now that you mention pcs, decoding is much lighter on the cpu than encoding. I don't think that would usually be a deal breaker. Or nevertheless, couldn't the client-side recombination just fake to be 4:4:4 then? Or would whatever empty padding you add ruin the image more than the results you could get with just plain 4:2:0?

About 4K limit, 1440p might still work depending on how exactly said limit is implemented, the overall pixel count stays within 4K range.

You mean if the limit is actually implemented like 4096x2160 (usual old amd) vs 4096x4096 (usual old nvidia)?
Or can you really call it a day just as long as the supported total pixel count, whatever the "shape", is 7.372.800 (2560x1440x2) or more?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

You can't encode 4:2:2 on nvidia gpus, implementing a path exclusively for intel will be too expensive.

Or can you really call it a day just as long as the supported total pixel count, whatever the "shape", is 7.372.800 (2560x1440x2) or more?

I'm already calling it a day 😎
Doubling one dimension allows to minimize discontinuities in motion estimation, in contrast to tiling. Current half-naive implementation for example has single motion estimation vertical "seam" in U and V planes.

// Y U V// +-------+ +---+ +---+// | | | | | |// | Y | |UR | |VR |// | | | | | |// +---+---+ +---+ +---+// | | |// |UL |VL |// | | |// +---+---+

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

You can't encode 4:2:2 on nvidia gpus, implementing a path exclusively for intel will be too expensive.

You can't encode 4:4:4 on amd gpus either, and yet this is what this PR is about isn't it?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

Personally, I don't see a point in supporting recombination into 4:2:2
It will still have visible artifacts while having computational overhead close to 4:4:4 and significant amount of additional development time. And this development time will be multiplied by the amount of distinct clients,

@mirh

mirh commented Jun 28, 2024

Copy link
Copy Markdown

I mean, sure, of course this is already miraculous.
I was just trying to think outside the box (4:2:2 is still subpar, but even the worst case scenario starts to be bearable instead).

If any I guess the improvement isn't that clear cut, because unlike with a direct cable connection it's not like there aren't already compression artifacts anyway. So if 4:4:4 couldn't fit in some whatever doubled 4:2:0 4K scenario, just lowering the resolution could also be a possible (and if not any easily immediate) alternative?

@ns6089
ns6089force-pushed the yuv444in420 branch 4 times, most recently from fc48f22 to 3a1115dCompareJune 30, 2024 08:50
@ns6089
ns6089force-pushed the yuv444in420 branch 2 times, most recently from 2cc6a6a to a4ffe24CompareAugust 1, 2024 07:52
@ns6089ns6089 changed the title Support recombined YUV 4:4:4 encoding (Prototype, Windows-only for now) feat(win/video): add support for recombined YUV444 encodingAug 22, 2024
@sonarqubecloud

Copy link
Copy Markdown

@ns6089

Copy link
Copy Markdown
ContributorAuthor

The code in this pull request is Not a Contribution under LizardByte Individual Contributor License Agreement.
The code in this pull request is shared under GNU GENERAL PUBLIC LICENSE Version 3.

The feature itself is completed on sunshine side.

danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 21, 2026
Convert absolute mouse input to relative deltas, accumulating fractional
deltas so slow (sub-pixel) motion still moves the cursor on both axes.
Moonlight TV clients send absolute positions; emitting relative deltas makes
the COSMIC magnifier follow the cursor (PointerMotionAbsolute never updates
the zoom focal point — pop-os/cosmic-comp LizardByte#2760).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (6ad5f97):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Dead reckoning is the only
thing in the smooth motion path (1:1); the estimate snaps to the real
cursor position (published by the KMS cursor plane via platf::kms_cursor_*
atomics) while the client is idle; a raw-saturated axis targets the
matching host edge, with fresh feedback authoritative on that axis —
phantom walls cannot persist. History: v1 walls, v2 partial, v3 jumps,
v4 pseudo-acceleration, v5 zero-slide latch — see CACHYOS.md.
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-9 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 27, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@ns6089@mirh@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(win/video): add support for recombined YUV444 encoding - #2760

Closed
ns6089 wants to merge 1 commit into
LizardByte:masterfrom
ns6089:yuv444in420
Closed

feat(win/video): add support for recombined YUV444 encoding#2760
ns6089 wants to merge 1 commit into
LizardByte:masterfrom
ns6089:yuv444in420

Conversation

@ns6089

@ns6089ns6089 commented Jun 26, 2024

Copy link
Copy Markdown
Contributor

Description

The continuation of #2533. It's possible to emulate YUV 4:4:4 on gpus that don't support it natively by doubling the YUV 4:2:0 pixel count and running custom recombination shaders on both encoding and decoding side. Like Microsoft did it in MS-RDPEGFX.

Prototype stage. Requires changes on moonlight's side: I currently have custom libplacebo mpv shader implemented for plvk backend, in the future it should be possible to add Direct3D11 and OpenGL shaders.

https://github.com/ns6089/Sunshine/compare/yuv444..yuv444in420

moonlight-common-c pull request: TBD
moonlight-qt pull request: TBD, testing branch https://github.com/ns6089/moonlight-qt/tree/yuv444in420

What works and what doesn't

  1. First prototype, left half of U_src and V_src planes in Y_out. Good DCT, bad motion compensation.
  2. Second prototype. U_src in Y_out. V_src is spread across U_out and V_out in a pattern that is spatially consistent with Y_out. Good motion compensation, relatively fat DCT on U_out and V_out due to high frequencies.
  3. Third prototype, dropped. Maybe can slightly improve the DCT by running 1/4 of V through averaging low pass filter.

To Do

  • decide what to do with resolutions not divisible by 2
  • decide in which part of the protocol dimension doubling will be taking place, e.g. will the client request the doubled dimension or will it be done implicitly

Screenshot

before
after

Issues Fixed or Closed

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

Awesomely crazy.
Could there be anything worth doing with an emulated 4:2:2 stream then? Like, I don't know, slightly lower recombination overhead, or lower bandwidth requirements?

Or perhaps not hitting encoding limits at higher resolutions. Like, is my understanding correct that this pixel doubling would not allow for 1440p on (say) older VCE versions that max out at 4K?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

I don't think anyone but Intel supports 4:2:2. About 4K limit, 1440p might still work depending on how exactly said limit is implemented, the overall pixel count stays within 4K range.


amd

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

I don't think anyone but Intel supports 4:2:2.

To be honest, I was more thinking of TVs than computers here. It's a mixed bag even there, but still it's not so rare.
But now that you mention pcs, decoding is much lighter on the cpu than encoding. I don't think that would usually be a deal breaker. Or nevertheless, couldn't the client-side recombination just fake to be 4:4:4 then? Or would whatever empty padding you add ruin the image more than the results you could get with just plain 4:2:0?

About 4K limit, 1440p might still work depending on how exactly said limit is implemented, the overall pixel count stays within 4K range.

You mean if the limit is actually implemented like 4096x2160 (usual old amd) vs 4096x4096 (usual old nvidia)?
Or can you really call it a day just as long as the supported total pixel count, whatever the "shape", is 7.372.800 (2560x1440x2) or more?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

You can't encode 4:2:2 on nvidia gpus, implementing a path exclusively for intel will be too expensive.

Or can you really call it a day just as long as the supported total pixel count, whatever the "shape", is 7.372.800 (2560x1440x2) or more?

I'm already calling it a day 😎
Doubling one dimension allows to minimize discontinuities in motion estimation, in contrast to tiling. Current half-naive implementation for example has single motion estimation vertical "seam" in U and V planes.

// Y U V// +-------+ +---+ +---+// | | | | | |// | Y | |UR | |VR |// | | | | | |// +---+---+ +---+ +---+// | | |// |UL |VL |// | | |// +---+---+

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

You can't encode 4:2:2 on nvidia gpus, implementing a path exclusively for intel will be too expensive.

You can't encode 4:4:4 on amd gpus either, and yet this is what this PR is about isn't it?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

Personally, I don't see a point in supporting recombination into 4:2:2
It will still have visible artifacts while having computational overhead close to 4:4:4 and significant amount of additional development time. And this development time will be multiplied by the amount of distinct clients,

@mirh

mirh commented Jun 28, 2024

Copy link
Copy Markdown

I mean, sure, of course this is already miraculous.
I was just trying to think outside the box (4:2:2 is still subpar, but even the worst case scenario starts to be bearable instead).

If any I guess the improvement isn't that clear cut, because unlike with a direct cable connection it's not like there aren't already compression artifacts anyway. So if 4:4:4 couldn't fit in some whatever doubled 4:2:0 4K scenario, just lowering the resolution could also be a possible (and if not any easily immediate) alternative?

@ns6089
ns6089force-pushed the yuv444in420 branch 4 times, most recently from fc48f22 to 3a1115dCompareJune 30, 2024 08:50
@ns6089
ns6089force-pushed the yuv444in420 branch 2 times, most recently from 2cc6a6a to a4ffe24CompareAugust 1, 2024 07:52
@ns6089ns6089 changed the title Support recombined YUV 4:4:4 encoding (Prototype, Windows-only for now) feat(win/video): add support for recombined YUV444 encodingAug 22, 2024
@sonarqubecloud

Copy link
Copy Markdown

@ns6089

Copy link
Copy Markdown
ContributorAuthor

The code in this pull request is Not a Contribution under LizardByte Individual Contributor License Agreement.
The code in this pull request is shared under GNU GENERAL PUBLIC LICENSE Version 3.

The feature itself is completed on sunshine side.

danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 21, 2026
Convert absolute mouse input to relative deltas, accumulating fractional
deltas so slow (sub-pixel) motion still moves the cursor on both axes.
Moonlight TV clients send absolute positions; emitting relative deltas makes
the COSMIC magnifier follow the cursor (PointerMotionAbsolute never updates
the zoom focal point — pop-os/cosmic-comp LizardByte#2760).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (6ad5f97):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Dead reckoning is the only
thing in the smooth motion path (1:1); the estimate snaps to the real
cursor position (published by the KMS cursor plane via platf::kms_cursor_*
atomics) while the client is idle; a raw-saturated axis targets the
matching host edge, with fresh feedback authoritative on that axis —
phantom walls cannot persist. History: v1 walls, v2 partial, v3 jumps,
v4 pseudo-acceleration, v5 zero-slide latch — see CACHYOS.md.
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-9 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 27, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@ns6089@mirh@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(win/video): add support for recombined YUV444 encoding - #2760

Closed
ns6089 wants to merge 1 commit into
LizardByte:masterfrom
ns6089:yuv444in420
Closed

feat(win/video): add support for recombined YUV444 encoding#2760
ns6089 wants to merge 1 commit into
LizardByte:masterfrom
ns6089:yuv444in420

Conversation

@ns6089

@ns6089ns6089 commented Jun 26, 2024

Copy link
Copy Markdown
Contributor

Description

The continuation of #2533. It's possible to emulate YUV 4:4:4 on gpus that don't support it natively by doubling the YUV 4:2:0 pixel count and running custom recombination shaders on both encoding and decoding side. Like Microsoft did it in MS-RDPEGFX.

Prototype stage. Requires changes on moonlight's side: I currently have custom libplacebo mpv shader implemented for plvk backend, in the future it should be possible to add Direct3D11 and OpenGL shaders.

https://github.com/ns6089/Sunshine/compare/yuv444..yuv444in420

moonlight-common-c pull request: TBD
moonlight-qt pull request: TBD, testing branch https://github.com/ns6089/moonlight-qt/tree/yuv444in420

What works and what doesn't

  1. First prototype, left half of U_src and V_src planes in Y_out. Good DCT, bad motion compensation.
  2. Second prototype. U_src in Y_out. V_src is spread across U_out and V_out in a pattern that is spatially consistent with Y_out. Good motion compensation, relatively fat DCT on U_out and V_out due to high frequencies.
  3. Third prototype, dropped. Maybe can slightly improve the DCT by running 1/4 of V through averaging low pass filter.

To Do

  • decide what to do with resolutions not divisible by 2
  • decide in which part of the protocol dimension doubling will be taking place, e.g. will the client request the doubled dimension or will it be done implicitly

Screenshot

before
after

Issues Fixed or Closed

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

Awesomely crazy.
Could there be anything worth doing with an emulated 4:2:2 stream then? Like, I don't know, slightly lower recombination overhead, or lower bandwidth requirements?

Or perhaps not hitting encoding limits at higher resolutions. Like, is my understanding correct that this pixel doubling would not allow for 1440p on (say) older VCE versions that max out at 4K?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

I don't think anyone but Intel supports 4:2:2. About 4K limit, 1440p might still work depending on how exactly said limit is implemented, the overall pixel count stays within 4K range.


amd

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

I don't think anyone but Intel supports 4:2:2.

To be honest, I was more thinking of TVs than computers here. It's a mixed bag even there, but still it's not so rare.
But now that you mention pcs, decoding is much lighter on the cpu than encoding. I don't think that would usually be a deal breaker. Or nevertheless, couldn't the client-side recombination just fake to be 4:4:4 then? Or would whatever empty padding you add ruin the image more than the results you could get with just plain 4:2:0?

About 4K limit, 1440p might still work depending on how exactly said limit is implemented, the overall pixel count stays within 4K range.

You mean if the limit is actually implemented like 4096x2160 (usual old amd) vs 4096x4096 (usual old nvidia)?
Or can you really call it a day just as long as the supported total pixel count, whatever the "shape", is 7.372.800 (2560x1440x2) or more?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

You can't encode 4:2:2 on nvidia gpus, implementing a path exclusively for intel will be too expensive.

Or can you really call it a day just as long as the supported total pixel count, whatever the "shape", is 7.372.800 (2560x1440x2) or more?

I'm already calling it a day 😎
Doubling one dimension allows to minimize discontinuities in motion estimation, in contrast to tiling. Current half-naive implementation for example has single motion estimation vertical "seam" in U and V planes.

// Y U V// +-------+ +---+ +---+// | | | | | |// | Y | |UR | |VR |// | | | | | |// +---+---+ +---+ +---+// | | |// |UL |VL |// | | |// +---+---+

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

You can't encode 4:2:2 on nvidia gpus, implementing a path exclusively for intel will be too expensive.

You can't encode 4:4:4 on amd gpus either, and yet this is what this PR is about isn't it?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

Personally, I don't see a point in supporting recombination into 4:2:2
It will still have visible artifacts while having computational overhead close to 4:4:4 and significant amount of additional development time. And this development time will be multiplied by the amount of distinct clients,

@mirh

mirh commented Jun 28, 2024

Copy link
Copy Markdown

I mean, sure, of course this is already miraculous.
I was just trying to think outside the box (4:2:2 is still subpar, but even the worst case scenario starts to be bearable instead).

If any I guess the improvement isn't that clear cut, because unlike with a direct cable connection it's not like there aren't already compression artifacts anyway. So if 4:4:4 couldn't fit in some whatever doubled 4:2:0 4K scenario, just lowering the resolution could also be a possible (and if not any easily immediate) alternative?

@ns6089
ns6089force-pushed the yuv444in420 branch 4 times, most recently from fc48f22 to 3a1115dCompareJune 30, 2024 08:50
@ns6089
ns6089force-pushed the yuv444in420 branch 2 times, most recently from 2cc6a6a to a4ffe24CompareAugust 1, 2024 07:52
@ns6089ns6089 changed the title Support recombined YUV 4:4:4 encoding (Prototype, Windows-only for now) feat(win/video): add support for recombined YUV444 encodingAug 22, 2024
@sonarqubecloud

Copy link
Copy Markdown

@ns6089

Copy link
Copy Markdown
ContributorAuthor

The code in this pull request is Not a Contribution under LizardByte Individual Contributor License Agreement.
The code in this pull request is shared under GNU GENERAL PUBLIC LICENSE Version 3.

The feature itself is completed on sunshine side.

danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 21, 2026
Convert absolute mouse input to relative deltas, accumulating fractional
deltas so slow (sub-pixel) motion still moves the cursor on both axes.
Moonlight TV clients send absolute positions; emitting relative deltas makes
the COSMIC magnifier follow the cursor (PointerMotionAbsolute never updates
the zoom focal point — pop-os/cosmic-comp LizardByte#2760).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (6ad5f97):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Dead reckoning is the only
thing in the smooth motion path (1:1); the estimate snaps to the real
cursor position (published by the KMS cursor plane via platf::kms_cursor_*
atomics) while the client is idle; a raw-saturated axis targets the
matching host edge, with fresh feedback authoritative on that axis —
phantom walls cannot persist. History: v1 walls, v2 partial, v3 jumps,
v4 pseudo-acceleration, v5 zero-slide latch — see CACHYOS.md.
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-9 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 27, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@ns6089@mirh@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat(win/video): add support for recombined YUV444 encoding - #2760

Closed
ns6089 wants to merge 1 commit into
LizardByte:masterfrom
ns6089:yuv444in420
Closed

feat(win/video): add support for recombined YUV444 encoding#2760
ns6089 wants to merge 1 commit into
LizardByte:masterfrom
ns6089:yuv444in420

Conversation

@ns6089

@ns6089ns6089 commented Jun 26, 2024

Copy link
Copy Markdown
Contributor

Description

The continuation of #2533. It's possible to emulate YUV 4:4:4 on gpus that don't support it natively by doubling the YUV 4:2:0 pixel count and running custom recombination shaders on both encoding and decoding side. Like Microsoft did it in MS-RDPEGFX.

Prototype stage. Requires changes on moonlight's side: I currently have custom libplacebo mpv shader implemented for plvk backend, in the future it should be possible to add Direct3D11 and OpenGL shaders.

https://github.com/ns6089/Sunshine/compare/yuv444..yuv444in420

moonlight-common-c pull request: TBD
moonlight-qt pull request: TBD, testing branch https://github.com/ns6089/moonlight-qt/tree/yuv444in420

What works and what doesn't

  1. First prototype, left half of U_src and V_src planes in Y_out. Good DCT, bad motion compensation.
  2. Second prototype. U_src in Y_out. V_src is spread across U_out and V_out in a pattern that is spatially consistent with Y_out. Good motion compensation, relatively fat DCT on U_out and V_out due to high frequencies.
  3. Third prototype, dropped. Maybe can slightly improve the DCT by running 1/4 of V through averaging low pass filter.

To Do

  • decide what to do with resolutions not divisible by 2
  • decide in which part of the protocol dimension doubling will be taking place, e.g. will the client request the doubled dimension or will it be done implicitly

Screenshot

before
after

Issues Fixed or Closed

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

Awesomely crazy.
Could there be anything worth doing with an emulated 4:2:2 stream then? Like, I don't know, slightly lower recombination overhead, or lower bandwidth requirements?

Or perhaps not hitting encoding limits at higher resolutions. Like, is my understanding correct that this pixel doubling would not allow for 1440p on (say) older VCE versions that max out at 4K?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

I don't think anyone but Intel supports 4:2:2. About 4K limit, 1440p might still work depending on how exactly said limit is implemented, the overall pixel count stays within 4K range.


amd

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

I don't think anyone but Intel supports 4:2:2.

To be honest, I was more thinking of TVs than computers here. It's a mixed bag even there, but still it's not so rare.
But now that you mention pcs, decoding is much lighter on the cpu than encoding. I don't think that would usually be a deal breaker. Or nevertheless, couldn't the client-side recombination just fake to be 4:4:4 then? Or would whatever empty padding you add ruin the image more than the results you could get with just plain 4:2:0?

About 4K limit, 1440p might still work depending on how exactly said limit is implemented, the overall pixel count stays within 4K range.

You mean if the limit is actually implemented like 4096x2160 (usual old amd) vs 4096x4096 (usual old nvidia)?
Or can you really call it a day just as long as the supported total pixel count, whatever the "shape", is 7.372.800 (2560x1440x2) or more?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

You can't encode 4:2:2 on nvidia gpus, implementing a path exclusively for intel will be too expensive.

Or can you really call it a day just as long as the supported total pixel count, whatever the "shape", is 7.372.800 (2560x1440x2) or more?

I'm already calling it a day 😎
Doubling one dimension allows to minimize discontinuities in motion estimation, in contrast to tiling. Current half-naive implementation for example has single motion estimation vertical "seam" in U and V planes.

// Y U V// +-------+ +---+ +---+// | | | | | |// | Y | |UR | |VR |// | | | | | |// +---+---+ +---+ +---+// | | |// |UL |VL |// | | |// +---+---+

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

You can't encode 4:2:2 on nvidia gpus, implementing a path exclusively for intel will be too expensive.

You can't encode 4:4:4 on amd gpus either, and yet this is what this PR is about isn't it?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

Personally, I don't see a point in supporting recombination into 4:2:2
It will still have visible artifacts while having computational overhead close to 4:4:4 and significant amount of additional development time. And this development time will be multiplied by the amount of distinct clients,

@mirh

mirh commented Jun 28, 2024

Copy link
Copy Markdown

I mean, sure, of course this is already miraculous.
I was just trying to think outside the box (4:2:2 is still subpar, but even the worst case scenario starts to be bearable instead).

If any I guess the improvement isn't that clear cut, because unlike with a direct cable connection it's not like there aren't already compression artifacts anyway. So if 4:4:4 couldn't fit in some whatever doubled 4:2:0 4K scenario, just lowering the resolution could also be a possible (and if not any easily immediate) alternative?

@ns6089
ns6089force-pushed the yuv444in420 branch 4 times, most recently from fc48f22 to 3a1115dCompareJune 30, 2024 08:50
@ns6089
ns6089force-pushed the yuv444in420 branch 2 times, most recently from 2cc6a6a to a4ffe24CompareAugust 1, 2024 07:52
@ns6089ns6089 changed the title Support recombined YUV 4:4:4 encoding (Prototype, Windows-only for now) feat(win/video): add support for recombined YUV444 encodingAug 22, 2024
@sonarqubecloud

Copy link
Copy Markdown

@ns6089

Copy link
Copy Markdown
ContributorAuthor

The code in this pull request is Not a Contribution under LizardByte Individual Contributor License Agreement.
The code in this pull request is shared under GNU GENERAL PUBLIC LICENSE Version 3.

The feature itself is completed on sunshine side.

danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 21, 2026
Convert absolute mouse input to relative deltas, accumulating fractional
deltas so slow (sub-pixel) motion still moves the cursor on both axes.
Moonlight TV clients send absolute positions; emitting relative deltas makes
the COSMIC magnifier follow the cursor (PointerMotionAbsolute never updates
the zoom focal point — pop-os/cosmic-comp LizardByte#2760).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (6ad5f97):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Dead reckoning is the only
thing in the smooth motion path (1:1); the estimate snaps to the real
cursor position (published by the KMS cursor plane via platf::kms_cursor_*
atomics) while the client is idle; a raw-saturated axis targets the
matching host edge, with fresh feedback authoritative on that axis —
phantom walls cannot persist. History: v1 walls, v2 partial, v3 jumps,
v4 pseudo-acceleration, v5 zero-slide latch — see CACHYOS.md.
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-9 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 27, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@ns6089@mirh@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(win/video): add support for recombined YUV444 encoding - #2760

Closed
ns6089 wants to merge 1 commit into
LizardByte:masterfrom
ns6089:yuv444in420
Closed

feat(win/video): add support for recombined YUV444 encoding#2760
ns6089 wants to merge 1 commit into
LizardByte:masterfrom
ns6089:yuv444in420

Conversation

@ns6089

@ns6089ns6089 commented Jun 26, 2024

Copy link
Copy Markdown
Contributor

Description

The continuation of #2533. It's possible to emulate YUV 4:4:4 on gpus that don't support it natively by doubling the YUV 4:2:0 pixel count and running custom recombination shaders on both encoding and decoding side. Like Microsoft did it in MS-RDPEGFX.

Prototype stage. Requires changes on moonlight's side: I currently have custom libplacebo mpv shader implemented for plvk backend, in the future it should be possible to add Direct3D11 and OpenGL shaders.

https://github.com/ns6089/Sunshine/compare/yuv444..yuv444in420

moonlight-common-c pull request: TBD
moonlight-qt pull request: TBD, testing branch https://github.com/ns6089/moonlight-qt/tree/yuv444in420

What works and what doesn't

  1. First prototype, left half of U_src and V_src planes in Y_out. Good DCT, bad motion compensation.
  2. Second prototype. U_src in Y_out. V_src is spread across U_out and V_out in a pattern that is spatially consistent with Y_out. Good motion compensation, relatively fat DCT on U_out and V_out due to high frequencies.
  3. Third prototype, dropped. Maybe can slightly improve the DCT by running 1/4 of V through averaging low pass filter.

To Do

  • decide what to do with resolutions not divisible by 2
  • decide in which part of the protocol dimension doubling will be taking place, e.g. will the client request the doubled dimension or will it be done implicitly

Screenshot

before
after

Issues Fixed or Closed

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

Awesomely crazy.
Could there be anything worth doing with an emulated 4:2:2 stream then? Like, I don't know, slightly lower recombination overhead, or lower bandwidth requirements?

Or perhaps not hitting encoding limits at higher resolutions. Like, is my understanding correct that this pixel doubling would not allow for 1440p on (say) older VCE versions that max out at 4K?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

I don't think anyone but Intel supports 4:2:2. About 4K limit, 1440p might still work depending on how exactly said limit is implemented, the overall pixel count stays within 4K range.


amd

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

I don't think anyone but Intel supports 4:2:2.

To be honest, I was more thinking of TVs than computers here. It's a mixed bag even there, but still it's not so rare.
But now that you mention pcs, decoding is much lighter on the cpu than encoding. I don't think that would usually be a deal breaker. Or nevertheless, couldn't the client-side recombination just fake to be 4:4:4 then? Or would whatever empty padding you add ruin the image more than the results you could get with just plain 4:2:0?

About 4K limit, 1440p might still work depending on how exactly said limit is implemented, the overall pixel count stays within 4K range.

You mean if the limit is actually implemented like 4096x2160 (usual old amd) vs 4096x4096 (usual old nvidia)?
Or can you really call it a day just as long as the supported total pixel count, whatever the "shape", is 7.372.800 (2560x1440x2) or more?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

You can't encode 4:2:2 on nvidia gpus, implementing a path exclusively for intel will be too expensive.

Or can you really call it a day just as long as the supported total pixel count, whatever the "shape", is 7.372.800 (2560x1440x2) or more?

I'm already calling it a day 😎
Doubling one dimension allows to minimize discontinuities in motion estimation, in contrast to tiling. Current half-naive implementation for example has single motion estimation vertical "seam" in U and V planes.

// Y U V// +-------+ +---+ +---+// | | | | | |// | Y | |UR | |VR |// | | | | | |// +---+---+ +---+ +---+// | | |// |UL |VL |// | | |// +---+---+

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

You can't encode 4:2:2 on nvidia gpus, implementing a path exclusively for intel will be too expensive.

You can't encode 4:4:4 on amd gpus either, and yet this is what this PR is about isn't it?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

Personally, I don't see a point in supporting recombination into 4:2:2
It will still have visible artifacts while having computational overhead close to 4:4:4 and significant amount of additional development time. And this development time will be multiplied by the amount of distinct clients,

@mirh

mirh commented Jun 28, 2024

Copy link
Copy Markdown

I mean, sure, of course this is already miraculous.
I was just trying to think outside the box (4:2:2 is still subpar, but even the worst case scenario starts to be bearable instead).

If any I guess the improvement isn't that clear cut, because unlike with a direct cable connection it's not like there aren't already compression artifacts anyway. So if 4:4:4 couldn't fit in some whatever doubled 4:2:0 4K scenario, just lowering the resolution could also be a possible (and if not any easily immediate) alternative?

@ns6089
ns6089force-pushed the yuv444in420 branch 4 times, most recently from fc48f22 to 3a1115dCompareJune 30, 2024 08:50
@ns6089
ns6089force-pushed the yuv444in420 branch 2 times, most recently from 2cc6a6a to a4ffe24CompareAugust 1, 2024 07:52
@ns6089ns6089 changed the title Support recombined YUV 4:4:4 encoding (Prototype, Windows-only for now) feat(win/video): add support for recombined YUV444 encodingAug 22, 2024
@sonarqubecloud

Copy link
Copy Markdown

@ns6089

Copy link
Copy Markdown
ContributorAuthor

The code in this pull request is Not a Contribution under LizardByte Individual Contributor License Agreement.
The code in this pull request is shared under GNU GENERAL PUBLIC LICENSE Version 3.

The feature itself is completed on sunshine side.

danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 21, 2026
Convert absolute mouse input to relative deltas, accumulating fractional
deltas so slow (sub-pixel) motion still moves the cursor on both axes.
Moonlight TV clients send absolute positions; emitting relative deltas makes
the COSMIC magnifier follow the cursor (PointerMotionAbsolute never updates
the zoom focal point — pop-os/cosmic-comp LizardByte#2760).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (6ad5f97):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Dead reckoning is the only
thing in the smooth motion path (1:1); the estimate snaps to the real
cursor position (published by the KMS cursor plane via platf::kms_cursor_*
atomics) while the client is idle; a raw-saturated axis targets the
matching host edge, with fresh feedback authoritative on that axis —
phantom walls cannot persist. History: v1 walls, v2 partial, v3 jumps,
v4 pseudo-acceleration, v5 zero-slide latch — see CACHYOS.md.
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-9 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 27, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@ns6089@mirh@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat(win/video): add support for recombined YUV444 encoding - #2760

Closed
ns6089 wants to merge 1 commit into
LizardByte:masterfrom
ns6089:yuv444in420
Closed

feat(win/video): add support for recombined YUV444 encoding#2760
ns6089 wants to merge 1 commit into
LizardByte:masterfrom
ns6089:yuv444in420

Conversation

@ns6089

@ns6089ns6089 commented Jun 26, 2024

Copy link
Copy Markdown
Contributor

Description

The continuation of #2533. It's possible to emulate YUV 4:4:4 on gpus that don't support it natively by doubling the YUV 4:2:0 pixel count and running custom recombination shaders on both encoding and decoding side. Like Microsoft did it in MS-RDPEGFX.

Prototype stage. Requires changes on moonlight's side: I currently have custom libplacebo mpv shader implemented for plvk backend, in the future it should be possible to add Direct3D11 and OpenGL shaders.

https://github.com/ns6089/Sunshine/compare/yuv444..yuv444in420

moonlight-common-c pull request: TBD
moonlight-qt pull request: TBD, testing branch https://github.com/ns6089/moonlight-qt/tree/yuv444in420

What works and what doesn't

  1. First prototype, left half of U_src and V_src planes in Y_out. Good DCT, bad motion compensation.
  2. Second prototype. U_src in Y_out. V_src is spread across U_out and V_out in a pattern that is spatially consistent with Y_out. Good motion compensation, relatively fat DCT on U_out and V_out due to high frequencies.
  3. Third prototype, dropped. Maybe can slightly improve the DCT by running 1/4 of V through averaging low pass filter.

To Do

  • decide what to do with resolutions not divisible by 2
  • decide in which part of the protocol dimension doubling will be taking place, e.g. will the client request the doubled dimension or will it be done implicitly

Screenshot

before
after

Issues Fixed or Closed

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

Awesomely crazy.
Could there be anything worth doing with an emulated 4:2:2 stream then? Like, I don't know, slightly lower recombination overhead, or lower bandwidth requirements?

Or perhaps not hitting encoding limits at higher resolutions. Like, is my understanding correct that this pixel doubling would not allow for 1440p on (say) older VCE versions that max out at 4K?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

I don't think anyone but Intel supports 4:2:2. About 4K limit, 1440p might still work depending on how exactly said limit is implemented, the overall pixel count stays within 4K range.


amd

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

I don't think anyone but Intel supports 4:2:2.

To be honest, I was more thinking of TVs than computers here. It's a mixed bag even there, but still it's not so rare.
But now that you mention pcs, decoding is much lighter on the cpu than encoding. I don't think that would usually be a deal breaker. Or nevertheless, couldn't the client-side recombination just fake to be 4:4:4 then? Or would whatever empty padding you add ruin the image more than the results you could get with just plain 4:2:0?

About 4K limit, 1440p might still work depending on how exactly said limit is implemented, the overall pixel count stays within 4K range.

You mean if the limit is actually implemented like 4096x2160 (usual old amd) vs 4096x4096 (usual old nvidia)?
Or can you really call it a day just as long as the supported total pixel count, whatever the "shape", is 7.372.800 (2560x1440x2) or more?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

You can't encode 4:2:2 on nvidia gpus, implementing a path exclusively for intel will be too expensive.

Or can you really call it a day just as long as the supported total pixel count, whatever the "shape", is 7.372.800 (2560x1440x2) or more?

I'm already calling it a day 😎
Doubling one dimension allows to minimize discontinuities in motion estimation, in contrast to tiling. Current half-naive implementation for example has single motion estimation vertical "seam" in U and V planes.

// Y U V// +-------+ +---+ +---+// | | | | | |// | Y | |UR | |VR |// | | | | | |// +---+---+ +---+ +---+// | | |// |UL |VL |// | | |// +---+---+

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

You can't encode 4:2:2 on nvidia gpus, implementing a path exclusively for intel will be too expensive.

You can't encode 4:4:4 on amd gpus either, and yet this is what this PR is about isn't it?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

Personally, I don't see a point in supporting recombination into 4:2:2
It will still have visible artifacts while having computational overhead close to 4:4:4 and significant amount of additional development time. And this development time will be multiplied by the amount of distinct clients,

@mirh

mirh commented Jun 28, 2024

Copy link
Copy Markdown

I mean, sure, of course this is already miraculous.
I was just trying to think outside the box (4:2:2 is still subpar, but even the worst case scenario starts to be bearable instead).

If any I guess the improvement isn't that clear cut, because unlike with a direct cable connection it's not like there aren't already compression artifacts anyway. So if 4:4:4 couldn't fit in some whatever doubled 4:2:0 4K scenario, just lowering the resolution could also be a possible (and if not any easily immediate) alternative?

@ns6089
ns6089force-pushed the yuv444in420 branch 4 times, most recently from fc48f22 to 3a1115dCompareJune 30, 2024 08:50
@ns6089
ns6089force-pushed the yuv444in420 branch 2 times, most recently from 2cc6a6a to a4ffe24CompareAugust 1, 2024 07:52
@ns6089ns6089 changed the title Support recombined YUV 4:4:4 encoding (Prototype, Windows-only for now) feat(win/video): add support for recombined YUV444 encodingAug 22, 2024
@sonarqubecloud

Copy link
Copy Markdown

@ns6089

Copy link
Copy Markdown
ContributorAuthor

The code in this pull request is Not a Contribution under LizardByte Individual Contributor License Agreement.
The code in this pull request is shared under GNU GENERAL PUBLIC LICENSE Version 3.

The feature itself is completed on sunshine side.

danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 21, 2026
Convert absolute mouse input to relative deltas, accumulating fractional
deltas so slow (sub-pixel) motion still moves the cursor on both axes.
Moonlight TV clients send absolute positions; emitting relative deltas makes
the COSMIC magnifier follow the cursor (PointerMotionAbsolute never updates
the zoom focal point — pop-os/cosmic-comp LizardByte#2760).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (6ad5f97):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Dead reckoning is the only
thing in the smooth motion path (1:1); the estimate snaps to the real
cursor position (published by the KMS cursor plane via platf::kms_cursor_*
atomics) while the client is idle; a raw-saturated axis targets the
matching host edge, with fresh feedback authoritative on that axis —
phantom walls cannot persist. History: v1 walls, v2 partial, v3 jumps,
v4 pseudo-acceleration, v5 zero-slide latch — see CACHYOS.md.
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-9 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 27, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@ns6089@mirh@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat(win/video): add support for recombined YUV444 encoding - #2760

Closed
ns6089 wants to merge 1 commit into
LizardByte:masterfrom
ns6089:yuv444in420
Closed

feat(win/video): add support for recombined YUV444 encoding#2760
ns6089 wants to merge 1 commit into
LizardByte:masterfrom
ns6089:yuv444in420

Conversation

@ns6089

@ns6089ns6089 commented Jun 26, 2024

Copy link
Copy Markdown
Contributor

Description

The continuation of #2533. It's possible to emulate YUV 4:4:4 on gpus that don't support it natively by doubling the YUV 4:2:0 pixel count and running custom recombination shaders on both encoding and decoding side. Like Microsoft did it in MS-RDPEGFX.

Prototype stage. Requires changes on moonlight's side: I currently have custom libplacebo mpv shader implemented for plvk backend, in the future it should be possible to add Direct3D11 and OpenGL shaders.

https://github.com/ns6089/Sunshine/compare/yuv444..yuv444in420

moonlight-common-c pull request: TBD
moonlight-qt pull request: TBD, testing branch https://github.com/ns6089/moonlight-qt/tree/yuv444in420

What works and what doesn't

  1. First prototype, left half of U_src and V_src planes in Y_out. Good DCT, bad motion compensation.
  2. Second prototype. U_src in Y_out. V_src is spread across U_out and V_out in a pattern that is spatially consistent with Y_out. Good motion compensation, relatively fat DCT on U_out and V_out due to high frequencies.
  3. Third prototype, dropped. Maybe can slightly improve the DCT by running 1/4 of V through averaging low pass filter.

To Do

  • decide what to do with resolutions not divisible by 2
  • decide in which part of the protocol dimension doubling will be taking place, e.g. will the client request the doubled dimension or will it be done implicitly

Screenshot

before
after

Issues Fixed or Closed

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

Awesomely crazy.
Could there be anything worth doing with an emulated 4:2:2 stream then? Like, I don't know, slightly lower recombination overhead, or lower bandwidth requirements?

Or perhaps not hitting encoding limits at higher resolutions. Like, is my understanding correct that this pixel doubling would not allow for 1440p on (say) older VCE versions that max out at 4K?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

I don't think anyone but Intel supports 4:2:2. About 4K limit, 1440p might still work depending on how exactly said limit is implemented, the overall pixel count stays within 4K range.


amd

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

I don't think anyone but Intel supports 4:2:2.

To be honest, I was more thinking of TVs than computers here. It's a mixed bag even there, but still it's not so rare.
But now that you mention pcs, decoding is much lighter on the cpu than encoding. I don't think that would usually be a deal breaker. Or nevertheless, couldn't the client-side recombination just fake to be 4:4:4 then? Or would whatever empty padding you add ruin the image more than the results you could get with just plain 4:2:0?

About 4K limit, 1440p might still work depending on how exactly said limit is implemented, the overall pixel count stays within 4K range.

You mean if the limit is actually implemented like 4096x2160 (usual old amd) vs 4096x4096 (usual old nvidia)?
Or can you really call it a day just as long as the supported total pixel count, whatever the "shape", is 7.372.800 (2560x1440x2) or more?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

You can't encode 4:2:2 on nvidia gpus, implementing a path exclusively for intel will be too expensive.

Or can you really call it a day just as long as the supported total pixel count, whatever the "shape", is 7.372.800 (2560x1440x2) or more?

I'm already calling it a day 😎
Doubling one dimension allows to minimize discontinuities in motion estimation, in contrast to tiling. Current half-naive implementation for example has single motion estimation vertical "seam" in U and V planes.

// Y U V// +-------+ +---+ +---+// | | | | | |// | Y | |UR | |VR |// | | | | | |// +---+---+ +---+ +---+// | | |// |UL |VL |// | | |// +---+---+

@mirh

mirh commented Jun 27, 2024

Copy link
Copy Markdown

You can't encode 4:2:2 on nvidia gpus, implementing a path exclusively for intel will be too expensive.

You can't encode 4:4:4 on amd gpus either, and yet this is what this PR is about isn't it?

@ns6089

Copy link
Copy Markdown
ContributorAuthor

Personally, I don't see a point in supporting recombination into 4:2:2
It will still have visible artifacts while having computational overhead close to 4:4:4 and significant amount of additional development time. And this development time will be multiplied by the amount of distinct clients,

@mirh

mirh commented Jun 28, 2024

Copy link
Copy Markdown

I mean, sure, of course this is already miraculous.
I was just trying to think outside the box (4:2:2 is still subpar, but even the worst case scenario starts to be bearable instead).

If any I guess the improvement isn't that clear cut, because unlike with a direct cable connection it's not like there aren't already compression artifacts anyway. So if 4:4:4 couldn't fit in some whatever doubled 4:2:0 4K scenario, just lowering the resolution could also be a possible (and if not any easily immediate) alternative?

@ns6089
ns6089force-pushed the yuv444in420 branch 4 times, most recently from fc48f22 to 3a1115dCompareJune 30, 2024 08:50
@ns6089
ns6089force-pushed the yuv444in420 branch 2 times, most recently from 2cc6a6a to a4ffe24CompareAugust 1, 2024 07:52
@ns6089ns6089 changed the title Support recombined YUV 4:4:4 encoding (Prototype, Windows-only for now) feat(win/video): add support for recombined YUV444 encodingAug 22, 2024
@sonarqubecloud

Copy link
Copy Markdown

@ns6089

Copy link
Copy Markdown
ContributorAuthor

The code in this pull request is Not a Contribution under LizardByte Individual Contributor License Agreement.
The code in this pull request is shared under GNU GENERAL PUBLIC LICENSE Version 3.

The feature itself is completed on sunshine side.

danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 21, 2026
Convert absolute mouse input to relative deltas, accumulating fractional
deltas so slow (sub-pixel) motion still moves the cursor on both axes.
Moonlight TV clients send absolute positions; emitting relative deltas makes
the COSMIC magnifier follow the cursor (PointerMotionAbsolute never updates
the zoom focal point — pop-os/cosmic-comp LizardByte#2760).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (6ad5f97):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Dead reckoning is the only
thing in the smooth motion path (1:1); the estimate snaps to the real
cursor position (published by the KMS cursor plane via platf::kms_cursor_*
atomics) while the client is idle; a raw-saturated axis targets the
matching host edge, with fresh feedback authoritative on that axis —
phantom walls cannot persist. History: v1 walls, v2 partial, v3 jumps,
v4 pseudo-acceleration, v5 zero-slide latch — see CACHYOS.md.
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-9 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 26, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
danalec pushed a commit to danalec/Sunshine that referenced this pull request Aug 27, 2026
…, perf patches
Single squashed commit collecting the CachyOS-specific patches on top of
upstream master (f273ce8):
- abs-mouse-relative v6: convert absolute mouse to relative deltas so the
COSMIC screen magnifier tracks the cursor (PointerMotionAbsolute does not
move its focal point; pop-os/cosmic-comp LizardByte#2760). Opt-in via the
absolute_mouse_as_relative config flag. Dead reckoning is the only motion
path (1:1); the estimate snaps to the real cursor position (published by
the KMS cursor plane via platf::kms_cursor_feedback()) while the client is
idle; a raw-saturated axis targets the matching host edge. Refactored to
pass the upstream SonarQube gate (no globals, extracted functions,
static_cast, seq_cst atomics).
- copy-framebuffer-reuse: no per-frame framebuffer allocation.
- kms-prop-cache: cache DRM property IDs (~30-40 fewer ioctls per frame).
- tools/set-accel-flat: flat pointer-acceleration profile helper.
Shipped as sunshine 2026.818.1-12 (PKGBUILD lives in danalec/cachyos-rebuilds).
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants

@ns6089@mirh@ReenigneArcher