Add standalone NVENC encoder - #1427

Merged
cgutman merged 4 commits into
LizardByte:nightlyfrom
ns6089:nvenc
Aug 13, 2023
Merged

Add standalone NVENC encoder#1427
cgutman merged 4 commits into
LizardByte:nightlyfrom
ns6089:nvenc

Conversation

@ns6089

@ns6089ns6089 commented Jul 7, 2023

Copy link
Copy Markdown
Contributor

Description

Add standalone NVENC encoder for reference frames invalidation right now. And possibly for VFR-like bitrate adjustment and more somewhere down the line. Windows version is fully functional . Linux cuda (and possibly opengl) support is out of scope for this PR, but can easily be done later (at least the encoder side),

In later PRs

  • Investigate and test the viability of "increased vbv" option. Ideally, P-frames should not steal bitrate budget from future frames, maybe this can be achieved with vbv offset in encoder.
  • Investigate the viability of using multiple ref frames (L0 > 1) since we apply strict limits on the vbv. Maybe nvenc can intelligently pick blocks from previous frames if they have higher qp? Can't imagine this being free though.
  • Run VMAF benchmark for const-qp mode (document the process in .md file), and pick default values for min-qp
  • Check GFE default ref frame buffer values, particularly for h264 (level4 vs level5)
  • New configuration page and documentation

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

Things to implement before merge (may be expanded)

  • Refactor colospace selection logic, it has too much duplication right now
  • Decide how to handle encoder ref frames buffer size, currently it's at 16 frames (dynamically based on resolution and framerate? through configuration?)
    Result: use 5 ref frames buffer, which corresponds to DPB=5 in h264 terms, and DPB=6 in HEVC. Both are required minimums that must be supported by decoder profiles for given resolution.
    Update: h265 level 4 only supports 4 ref frames, but in this case the client can set the limit, this level is very outdated nowadays
  • Investigate if NVENC can be told explicitly which ref frame to use, this can allow wider decoder support for ref frames invalidation.
    Verdict: when L0 (forward prediction) list size is 1, it should always use last frame as reference,
  • Try to patch num_ref_frames in SPS header
    Verdict: may be possible for h264, close to impossible for HEVC, either way too much hassle
  • Test encoder caps
    NV_ENC_CAPS_SUPPORT_CABAC
    NV_ENC_CAPS_WIDTH_MAX
    NV_ENC_CAPS_HEIGHT_MAX
    NV_ENC_CAPS_SUPPORT_CUSTOM_VBV_BUF_SIZE
    NV_ENC_CAPS_SUPPORT_REF_PIC_INVALIDATION
    NV_ENC_CAPS_SUPPORT_YUV444_ENCODE
    NV_ENC_CAPS_SUPPORT_10BIT_ENCODE
    NV_ENC_CAPS_SUPPORT_MULTIPLE_REF_FRAMES
    NV_ENC_CAPS_SUPPORT_QPELMV
  • Switch to nv-codec-headers
  • Update Linux and MacOS platforms to new structures
  • Use Peak-Signal-to-Noise-Ratio (Y-PSNR), Structural Similarity Index (Y-SSIM), and Video Multimethod Assessment Fusion (VMAF) for default min qp values Not in this PR
  • Add configuration page and documentation Not in this PR

From review

  • Look into entropyCodingMode encoding parameter
  • Proper cleanup in create_encoder()
  • Look into last_encoder_probe_supported_invalidate_ref_frames, if needs to be changed in multiple places
  • Lock h264 into High profile

Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/platform/windows/display_vram.cpp Outdated
Comment threadsrc/platform/windows/display_vram.cpp
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/video.cpp Outdated
@ns6089
ns6089force-pushed the nvenc branch 2 times, most recently from 0d3cee2 to 6836c6eCompareJuly 9, 2023 15:20
@ns6089ns6089 self-assigned this Jul 11, 2023
@ns6089

Copy link
Copy Markdown
ContributorAuthor

@cgutman, need your opinion on colorspace refactoring a2f34ab. When you have time, it's not blocking me.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

@cgutman, need your opinion on colorspace refactoring a2f34ab. When you have time, it's not blocking me.

Or better wait until I finish the Linux part, it approaches things slightly differently and might need adjustments.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

New options page (not yet pushed to this PR). Some things might still be slightly reworded, and need to run a few tests for better default QP values. But conceptually it should be done, and I think I will be able to backport it to ffmpeg nvenc backend so we don't end up with two pages.
2023-07-18 15_21_35-Sunshine — Mozilla Firefox

Comment threadsrc/platform/linux/vaapi.cpp
@cgutman

Copy link
Copy Markdown
Collaborator

LGTM. Let's try to get this refactoring in soon if possible. I've got some AV1 and client-side cursor changes that I'd like to work on, but I don't want to stomp all over your work.

If you can back out any debugging changes or WIP test changes (the scheduling priority revert?), I can test this across all my systems and we can get it merged.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

LGTM. Let's try to get this refactoring in soon if possible. I've got some AV1 and client-side cursor changes that I'd like to work on, but I don't want to stomp all over your work.

If you can back out any debugging changes or WIP test changes (the scheduling priority revert?), I can test this across all my systems and we can get it merged.

Understandable. Wired the encoder to use the existing configuration page, since some new options pretty much depend on whether I can make it work with realtime priority (and priority selection itself will add another option).

Haven't tested native HDR (only SDR in BT.2020) and CUDA path on Linux (didn't make drastic changes, but there's always a chance). Other than that, should be alright to merge in its current state. The rest of the features shouldn't produce conflicts, and can be done in later PRs.

@ns6089
ns6089 marked this pull request as ready for review August 4, 2023 12:31
Comment threadthird-party/nv-codec-headers
@ns6089
ns6089force-pushed the nvenc branch 2 times, most recently from 3ea8f10 to a318da3CompareAugust 5, 2023 13:59
Comment threadsrc/nvenc/nvenc_base.cpp
Comment threadsrc/nvenc/nvenc_utils.cpp
Comment threadsrc/video_colorspace.cpp Outdated
@cgutman

Copy link
Copy Markdown
Collaborator

The latest changes look good. Once you squash the fixup commits and rebase, I'll do a final testing pass on my local machines and we can get this in.

Shouldn't matter on x64 since everything is fastcall here, but cdecl is
the correct declaration.
@ns6089

Copy link
Copy Markdown
ContributorAuthor

Squashed and rebased. Also added small "fix" for nvapi, it doesn't affect anything since we're strictly x64, so no point in opening full pull request for it.

@cgutman

Copy link
Copy Markdown
Collaborator

Everything looks good in my tests:

  • HDR with and without NVENC
  • NvFBC on Linux
  • All valid colorspace and color range combos

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

No open projects
Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@ns6089@cgutman@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Add standalone NVENC encoder - #1427

Merged
cgutman merged 4 commits into
LizardByte:nightlyfrom
ns6089:nvenc
Aug 13, 2023
Merged

Add standalone NVENC encoder#1427
cgutman merged 4 commits into
LizardByte:nightlyfrom
ns6089:nvenc

Conversation

@ns6089

@ns6089ns6089 commented Jul 7, 2023

Copy link
Copy Markdown
Contributor

Description

Add standalone NVENC encoder for reference frames invalidation right now. And possibly for VFR-like bitrate adjustment and more somewhere down the line. Windows version is fully functional . Linux cuda (and possibly opengl) support is out of scope for this PR, but can easily be done later (at least the encoder side),

In later PRs

  • Investigate and test the viability of "increased vbv" option. Ideally, P-frames should not steal bitrate budget from future frames, maybe this can be achieved with vbv offset in encoder.
  • Investigate the viability of using multiple ref frames (L0 > 1) since we apply strict limits on the vbv. Maybe nvenc can intelligently pick blocks from previous frames if they have higher qp? Can't imagine this being free though.
  • Run VMAF benchmark for const-qp mode (document the process in .md file), and pick default values for min-qp
  • Check GFE default ref frame buffer values, particularly for h264 (level4 vs level5)
  • New configuration page and documentation

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

Things to implement before merge (may be expanded)

  • Refactor colospace selection logic, it has too much duplication right now
  • Decide how to handle encoder ref frames buffer size, currently it's at 16 frames (dynamically based on resolution and framerate? through configuration?)
    Result: use 5 ref frames buffer, which corresponds to DPB=5 in h264 terms, and DPB=6 in HEVC. Both are required minimums that must be supported by decoder profiles for given resolution.
    Update: h265 level 4 only supports 4 ref frames, but in this case the client can set the limit, this level is very outdated nowadays
  • Investigate if NVENC can be told explicitly which ref frame to use, this can allow wider decoder support for ref frames invalidation.
    Verdict: when L0 (forward prediction) list size is 1, it should always use last frame as reference,
  • Try to patch num_ref_frames in SPS header
    Verdict: may be possible for h264, close to impossible for HEVC, either way too much hassle
  • Test encoder caps
    NV_ENC_CAPS_SUPPORT_CABAC
    NV_ENC_CAPS_WIDTH_MAX
    NV_ENC_CAPS_HEIGHT_MAX
    NV_ENC_CAPS_SUPPORT_CUSTOM_VBV_BUF_SIZE
    NV_ENC_CAPS_SUPPORT_REF_PIC_INVALIDATION
    NV_ENC_CAPS_SUPPORT_YUV444_ENCODE
    NV_ENC_CAPS_SUPPORT_10BIT_ENCODE
    NV_ENC_CAPS_SUPPORT_MULTIPLE_REF_FRAMES
    NV_ENC_CAPS_SUPPORT_QPELMV
  • Switch to nv-codec-headers
  • Update Linux and MacOS platforms to new structures
  • Use Peak-Signal-to-Noise-Ratio (Y-PSNR), Structural Similarity Index (Y-SSIM), and Video Multimethod Assessment Fusion (VMAF) for default min qp values Not in this PR
  • Add configuration page and documentation Not in this PR

From review

  • Look into entropyCodingMode encoding parameter
  • Proper cleanup in create_encoder()
  • Look into last_encoder_probe_supported_invalidate_ref_frames, if needs to be changed in multiple places
  • Lock h264 into High profile

Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/platform/windows/display_vram.cpp Outdated
Comment threadsrc/platform/windows/display_vram.cpp
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/video.cpp Outdated
@ns6089
ns6089force-pushed the nvenc branch 2 times, most recently from 0d3cee2 to 6836c6eCompareJuly 9, 2023 15:20
@ns6089ns6089 self-assigned this Jul 11, 2023
@ns6089

Copy link
Copy Markdown
ContributorAuthor

@cgutman, need your opinion on colorspace refactoring a2f34ab. When you have time, it's not blocking me.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

@cgutman, need your opinion on colorspace refactoring a2f34ab. When you have time, it's not blocking me.

Or better wait until I finish the Linux part, it approaches things slightly differently and might need adjustments.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

New options page (not yet pushed to this PR). Some things might still be slightly reworded, and need to run a few tests for better default QP values. But conceptually it should be done, and I think I will be able to backport it to ffmpeg nvenc backend so we don't end up with two pages.
2023-07-18 15_21_35-Sunshine — Mozilla Firefox

Comment threadsrc/platform/linux/vaapi.cpp
@cgutman

Copy link
Copy Markdown
Collaborator

LGTM. Let's try to get this refactoring in soon if possible. I've got some AV1 and client-side cursor changes that I'd like to work on, but I don't want to stomp all over your work.

If you can back out any debugging changes or WIP test changes (the scheduling priority revert?), I can test this across all my systems and we can get it merged.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

LGTM. Let's try to get this refactoring in soon if possible. I've got some AV1 and client-side cursor changes that I'd like to work on, but I don't want to stomp all over your work.

If you can back out any debugging changes or WIP test changes (the scheduling priority revert?), I can test this across all my systems and we can get it merged.

Understandable. Wired the encoder to use the existing configuration page, since some new options pretty much depend on whether I can make it work with realtime priority (and priority selection itself will add another option).

Haven't tested native HDR (only SDR in BT.2020) and CUDA path on Linux (didn't make drastic changes, but there's always a chance). Other than that, should be alright to merge in its current state. The rest of the features shouldn't produce conflicts, and can be done in later PRs.

@ns6089
ns6089 marked this pull request as ready for review August 4, 2023 12:31
Comment threadthird-party/nv-codec-headers
@ns6089
ns6089force-pushed the nvenc branch 2 times, most recently from 3ea8f10 to a318da3CompareAugust 5, 2023 13:59
Comment threadsrc/nvenc/nvenc_base.cpp
Comment threadsrc/nvenc/nvenc_utils.cpp
Comment threadsrc/video_colorspace.cpp Outdated
@cgutman

Copy link
Copy Markdown
Collaborator

The latest changes look good. Once you squash the fixup commits and rebase, I'll do a final testing pass on my local machines and we can get this in.

Shouldn't matter on x64 since everything is fastcall here, but cdecl is
the correct declaration.
@ns6089

Copy link
Copy Markdown
ContributorAuthor

Squashed and rebased. Also added small "fix" for nvapi, it doesn't affect anything since we're strictly x64, so no point in opening full pull request for it.

@cgutman

Copy link
Copy Markdown
Collaborator

Everything looks good in my tests:

  • HDR with and without NVENC
  • NvFBC on Linux
  • All valid colorspace and color range combos

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

No open projects
Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@ns6089@cgutman@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add standalone NVENC encoder - #1427

Merged
cgutman merged 4 commits into
LizardByte:nightlyfrom
ns6089:nvenc
Aug 13, 2023
Merged

Add standalone NVENC encoder#1427
cgutman merged 4 commits into
LizardByte:nightlyfrom
ns6089:nvenc

Conversation

@ns6089

@ns6089ns6089 commented Jul 7, 2023

Copy link
Copy Markdown
Contributor

Description

Add standalone NVENC encoder for reference frames invalidation right now. And possibly for VFR-like bitrate adjustment and more somewhere down the line. Windows version is fully functional . Linux cuda (and possibly opengl) support is out of scope for this PR, but can easily be done later (at least the encoder side),

In later PRs

  • Investigate and test the viability of "increased vbv" option. Ideally, P-frames should not steal bitrate budget from future frames, maybe this can be achieved with vbv offset in encoder.
  • Investigate the viability of using multiple ref frames (L0 > 1) since we apply strict limits on the vbv. Maybe nvenc can intelligently pick blocks from previous frames if they have higher qp? Can't imagine this being free though.
  • Run VMAF benchmark for const-qp mode (document the process in .md file), and pick default values for min-qp
  • Check GFE default ref frame buffer values, particularly for h264 (level4 vs level5)
  • New configuration page and documentation

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

Things to implement before merge (may be expanded)

  • Refactor colospace selection logic, it has too much duplication right now
  • Decide how to handle encoder ref frames buffer size, currently it's at 16 frames (dynamically based on resolution and framerate? through configuration?)
    Result: use 5 ref frames buffer, which corresponds to DPB=5 in h264 terms, and DPB=6 in HEVC. Both are required minimums that must be supported by decoder profiles for given resolution.
    Update: h265 level 4 only supports 4 ref frames, but in this case the client can set the limit, this level is very outdated nowadays
  • Investigate if NVENC can be told explicitly which ref frame to use, this can allow wider decoder support for ref frames invalidation.
    Verdict: when L0 (forward prediction) list size is 1, it should always use last frame as reference,
  • Try to patch num_ref_frames in SPS header
    Verdict: may be possible for h264, close to impossible for HEVC, either way too much hassle
  • Test encoder caps
    NV_ENC_CAPS_SUPPORT_CABAC
    NV_ENC_CAPS_WIDTH_MAX
    NV_ENC_CAPS_HEIGHT_MAX
    NV_ENC_CAPS_SUPPORT_CUSTOM_VBV_BUF_SIZE
    NV_ENC_CAPS_SUPPORT_REF_PIC_INVALIDATION
    NV_ENC_CAPS_SUPPORT_YUV444_ENCODE
    NV_ENC_CAPS_SUPPORT_10BIT_ENCODE
    NV_ENC_CAPS_SUPPORT_MULTIPLE_REF_FRAMES
    NV_ENC_CAPS_SUPPORT_QPELMV
  • Switch to nv-codec-headers
  • Update Linux and MacOS platforms to new structures
  • Use Peak-Signal-to-Noise-Ratio (Y-PSNR), Structural Similarity Index (Y-SSIM), and Video Multimethod Assessment Fusion (VMAF) for default min qp values Not in this PR
  • Add configuration page and documentation Not in this PR

From review

  • Look into entropyCodingMode encoding parameter
  • Proper cleanup in create_encoder()
  • Look into last_encoder_probe_supported_invalidate_ref_frames, if needs to be changed in multiple places
  • Lock h264 into High profile

Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/platform/windows/display_vram.cpp Outdated
Comment threadsrc/platform/windows/display_vram.cpp
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/video.cpp Outdated
@ns6089
ns6089force-pushed the nvenc branch 2 times, most recently from 0d3cee2 to 6836c6eCompareJuly 9, 2023 15:20
@ns6089ns6089 self-assigned this Jul 11, 2023
@ns6089

Copy link
Copy Markdown
ContributorAuthor

@cgutman, need your opinion on colorspace refactoring a2f34ab. When you have time, it's not blocking me.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

@cgutman, need your opinion on colorspace refactoring a2f34ab. When you have time, it's not blocking me.

Or better wait until I finish the Linux part, it approaches things slightly differently and might need adjustments.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

New options page (not yet pushed to this PR). Some things might still be slightly reworded, and need to run a few tests for better default QP values. But conceptually it should be done, and I think I will be able to backport it to ffmpeg nvenc backend so we don't end up with two pages.
2023-07-18 15_21_35-Sunshine — Mozilla Firefox

Comment threadsrc/platform/linux/vaapi.cpp
@cgutman

Copy link
Copy Markdown
Collaborator

LGTM. Let's try to get this refactoring in soon if possible. I've got some AV1 and client-side cursor changes that I'd like to work on, but I don't want to stomp all over your work.

If you can back out any debugging changes or WIP test changes (the scheduling priority revert?), I can test this across all my systems and we can get it merged.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

LGTM. Let's try to get this refactoring in soon if possible. I've got some AV1 and client-side cursor changes that I'd like to work on, but I don't want to stomp all over your work.

If you can back out any debugging changes or WIP test changes (the scheduling priority revert?), I can test this across all my systems and we can get it merged.

Understandable. Wired the encoder to use the existing configuration page, since some new options pretty much depend on whether I can make it work with realtime priority (and priority selection itself will add another option).

Haven't tested native HDR (only SDR in BT.2020) and CUDA path on Linux (didn't make drastic changes, but there's always a chance). Other than that, should be alright to merge in its current state. The rest of the features shouldn't produce conflicts, and can be done in later PRs.

@ns6089
ns6089 marked this pull request as ready for review August 4, 2023 12:31
Comment threadthird-party/nv-codec-headers
@ns6089
ns6089force-pushed the nvenc branch 2 times, most recently from 3ea8f10 to a318da3CompareAugust 5, 2023 13:59
Comment threadsrc/nvenc/nvenc_base.cpp
Comment threadsrc/nvenc/nvenc_utils.cpp
Comment threadsrc/video_colorspace.cpp Outdated
@cgutman

Copy link
Copy Markdown
Collaborator

The latest changes look good. Once you squash the fixup commits and rebase, I'll do a final testing pass on my local machines and we can get this in.

Shouldn't matter on x64 since everything is fastcall here, but cdecl is
the correct declaration.
@ns6089

Copy link
Copy Markdown
ContributorAuthor

Squashed and rebased. Also added small "fix" for nvapi, it doesn't affect anything since we're strictly x64, so no point in opening full pull request for it.

@cgutman

Copy link
Copy Markdown
Collaborator

Everything looks good in my tests:

  • HDR with and without NVENC
  • NvFBC on Linux
  • All valid colorspace and color range combos

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

No open projects
Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@ns6089@cgutman@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add standalone NVENC encoder - #1427

Merged
cgutman merged 4 commits into
LizardByte:nightlyfrom
ns6089:nvenc
Aug 13, 2023
Merged

Add standalone NVENC encoder#1427
cgutman merged 4 commits into
LizardByte:nightlyfrom
ns6089:nvenc

Conversation

@ns6089

@ns6089ns6089 commented Jul 7, 2023

Copy link
Copy Markdown
Contributor

Description

Add standalone NVENC encoder for reference frames invalidation right now. And possibly for VFR-like bitrate adjustment and more somewhere down the line. Windows version is fully functional . Linux cuda (and possibly opengl) support is out of scope for this PR, but can easily be done later (at least the encoder side),

In later PRs

  • Investigate and test the viability of "increased vbv" option. Ideally, P-frames should not steal bitrate budget from future frames, maybe this can be achieved with vbv offset in encoder.
  • Investigate the viability of using multiple ref frames (L0 > 1) since we apply strict limits on the vbv. Maybe nvenc can intelligently pick blocks from previous frames if they have higher qp? Can't imagine this being free though.
  • Run VMAF benchmark for const-qp mode (document the process in .md file), and pick default values for min-qp
  • Check GFE default ref frame buffer values, particularly for h264 (level4 vs level5)
  • New configuration page and documentation

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

Things to implement before merge (may be expanded)

  • Refactor colospace selection logic, it has too much duplication right now
  • Decide how to handle encoder ref frames buffer size, currently it's at 16 frames (dynamically based on resolution and framerate? through configuration?)
    Result: use 5 ref frames buffer, which corresponds to DPB=5 in h264 terms, and DPB=6 in HEVC. Both are required minimums that must be supported by decoder profiles for given resolution.
    Update: h265 level 4 only supports 4 ref frames, but in this case the client can set the limit, this level is very outdated nowadays
  • Investigate if NVENC can be told explicitly which ref frame to use, this can allow wider decoder support for ref frames invalidation.
    Verdict: when L0 (forward prediction) list size is 1, it should always use last frame as reference,
  • Try to patch num_ref_frames in SPS header
    Verdict: may be possible for h264, close to impossible for HEVC, either way too much hassle
  • Test encoder caps
    NV_ENC_CAPS_SUPPORT_CABAC
    NV_ENC_CAPS_WIDTH_MAX
    NV_ENC_CAPS_HEIGHT_MAX
    NV_ENC_CAPS_SUPPORT_CUSTOM_VBV_BUF_SIZE
    NV_ENC_CAPS_SUPPORT_REF_PIC_INVALIDATION
    NV_ENC_CAPS_SUPPORT_YUV444_ENCODE
    NV_ENC_CAPS_SUPPORT_10BIT_ENCODE
    NV_ENC_CAPS_SUPPORT_MULTIPLE_REF_FRAMES
    NV_ENC_CAPS_SUPPORT_QPELMV
  • Switch to nv-codec-headers
  • Update Linux and MacOS platforms to new structures
  • Use Peak-Signal-to-Noise-Ratio (Y-PSNR), Structural Similarity Index (Y-SSIM), and Video Multimethod Assessment Fusion (VMAF) for default min qp values Not in this PR
  • Add configuration page and documentation Not in this PR

From review

  • Look into entropyCodingMode encoding parameter
  • Proper cleanup in create_encoder()
  • Look into last_encoder_probe_supported_invalidate_ref_frames, if needs to be changed in multiple places
  • Lock h264 into High profile

Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/platform/windows/display_vram.cpp Outdated
Comment threadsrc/platform/windows/display_vram.cpp
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/video.cpp Outdated
@ns6089
ns6089force-pushed the nvenc branch 2 times, most recently from 0d3cee2 to 6836c6eCompareJuly 9, 2023 15:20
@ns6089ns6089 self-assigned this Jul 11, 2023
@ns6089

Copy link
Copy Markdown
ContributorAuthor

@cgutman, need your opinion on colorspace refactoring a2f34ab. When you have time, it's not blocking me.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

@cgutman, need your opinion on colorspace refactoring a2f34ab. When you have time, it's not blocking me.

Or better wait until I finish the Linux part, it approaches things slightly differently and might need adjustments.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

New options page (not yet pushed to this PR). Some things might still be slightly reworded, and need to run a few tests for better default QP values. But conceptually it should be done, and I think I will be able to backport it to ffmpeg nvenc backend so we don't end up with two pages.
2023-07-18 15_21_35-Sunshine — Mozilla Firefox

Comment threadsrc/platform/linux/vaapi.cpp
@cgutman

Copy link
Copy Markdown
Collaborator

LGTM. Let's try to get this refactoring in soon if possible. I've got some AV1 and client-side cursor changes that I'd like to work on, but I don't want to stomp all over your work.

If you can back out any debugging changes or WIP test changes (the scheduling priority revert?), I can test this across all my systems and we can get it merged.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

LGTM. Let's try to get this refactoring in soon if possible. I've got some AV1 and client-side cursor changes that I'd like to work on, but I don't want to stomp all over your work.

If you can back out any debugging changes or WIP test changes (the scheduling priority revert?), I can test this across all my systems and we can get it merged.

Understandable. Wired the encoder to use the existing configuration page, since some new options pretty much depend on whether I can make it work with realtime priority (and priority selection itself will add another option).

Haven't tested native HDR (only SDR in BT.2020) and CUDA path on Linux (didn't make drastic changes, but there's always a chance). Other than that, should be alright to merge in its current state. The rest of the features shouldn't produce conflicts, and can be done in later PRs.

@ns6089
ns6089 marked this pull request as ready for review August 4, 2023 12:31
Comment threadthird-party/nv-codec-headers
@ns6089
ns6089force-pushed the nvenc branch 2 times, most recently from 3ea8f10 to a318da3CompareAugust 5, 2023 13:59
Comment threadsrc/nvenc/nvenc_base.cpp
Comment threadsrc/nvenc/nvenc_utils.cpp
Comment threadsrc/video_colorspace.cpp Outdated
@cgutman

Copy link
Copy Markdown
Collaborator

The latest changes look good. Once you squash the fixup commits and rebase, I'll do a final testing pass on my local machines and we can get this in.

Shouldn't matter on x64 since everything is fastcall here, but cdecl is
the correct declaration.
@ns6089

Copy link
Copy Markdown
ContributorAuthor

Squashed and rebased. Also added small "fix" for nvapi, it doesn't affect anything since we're strictly x64, so no point in opening full pull request for it.

@cgutman

Copy link
Copy Markdown
Collaborator

Everything looks good in my tests:

  • HDR with and without NVENC
  • NvFBC on Linux
  • All valid colorspace and color range combos

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

No open projects
Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@ns6089@cgutman@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Add standalone NVENC encoder - #1427

Merged
cgutman merged 4 commits into
LizardByte:nightlyfrom
ns6089:nvenc
Aug 13, 2023
Merged

Add standalone NVENC encoder#1427
cgutman merged 4 commits into
LizardByte:nightlyfrom
ns6089:nvenc

Conversation

@ns6089

@ns6089ns6089 commented Jul 7, 2023

Copy link
Copy Markdown
Contributor

Description

Add standalone NVENC encoder for reference frames invalidation right now. And possibly for VFR-like bitrate adjustment and more somewhere down the line. Windows version is fully functional . Linux cuda (and possibly opengl) support is out of scope for this PR, but can easily be done later (at least the encoder side),

In later PRs

  • Investigate and test the viability of "increased vbv" option. Ideally, P-frames should not steal bitrate budget from future frames, maybe this can be achieved with vbv offset in encoder.
  • Investigate the viability of using multiple ref frames (L0 > 1) since we apply strict limits on the vbv. Maybe nvenc can intelligently pick blocks from previous frames if they have higher qp? Can't imagine this being free though.
  • Run VMAF benchmark for const-qp mode (document the process in .md file), and pick default values for min-qp
  • Check GFE default ref frame buffer values, particularly for h264 (level4 vs level5)
  • New configuration page and documentation

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

Things to implement before merge (may be expanded)

  • Refactor colospace selection logic, it has too much duplication right now
  • Decide how to handle encoder ref frames buffer size, currently it's at 16 frames (dynamically based on resolution and framerate? through configuration?)
    Result: use 5 ref frames buffer, which corresponds to DPB=5 in h264 terms, and DPB=6 in HEVC. Both are required minimums that must be supported by decoder profiles for given resolution.
    Update: h265 level 4 only supports 4 ref frames, but in this case the client can set the limit, this level is very outdated nowadays
  • Investigate if NVENC can be told explicitly which ref frame to use, this can allow wider decoder support for ref frames invalidation.
    Verdict: when L0 (forward prediction) list size is 1, it should always use last frame as reference,
  • Try to patch num_ref_frames in SPS header
    Verdict: may be possible for h264, close to impossible for HEVC, either way too much hassle
  • Test encoder caps
    NV_ENC_CAPS_SUPPORT_CABAC
    NV_ENC_CAPS_WIDTH_MAX
    NV_ENC_CAPS_HEIGHT_MAX
    NV_ENC_CAPS_SUPPORT_CUSTOM_VBV_BUF_SIZE
    NV_ENC_CAPS_SUPPORT_REF_PIC_INVALIDATION
    NV_ENC_CAPS_SUPPORT_YUV444_ENCODE
    NV_ENC_CAPS_SUPPORT_10BIT_ENCODE
    NV_ENC_CAPS_SUPPORT_MULTIPLE_REF_FRAMES
    NV_ENC_CAPS_SUPPORT_QPELMV
  • Switch to nv-codec-headers
  • Update Linux and MacOS platforms to new structures
  • Use Peak-Signal-to-Noise-Ratio (Y-PSNR), Structural Similarity Index (Y-SSIM), and Video Multimethod Assessment Fusion (VMAF) for default min qp values Not in this PR
  • Add configuration page and documentation Not in this PR

From review

  • Look into entropyCodingMode encoding parameter
  • Proper cleanup in create_encoder()
  • Look into last_encoder_probe_supported_invalidate_ref_frames, if needs to be changed in multiple places
  • Lock h264 into High profile

Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/platform/windows/display_vram.cpp Outdated
Comment threadsrc/platform/windows/display_vram.cpp
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/video.cpp Outdated
@ns6089
ns6089force-pushed the nvenc branch 2 times, most recently from 0d3cee2 to 6836c6eCompareJuly 9, 2023 15:20
@ns6089ns6089 self-assigned this Jul 11, 2023
@ns6089

Copy link
Copy Markdown
ContributorAuthor

@cgutman, need your opinion on colorspace refactoring a2f34ab. When you have time, it's not blocking me.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

@cgutman, need your opinion on colorspace refactoring a2f34ab. When you have time, it's not blocking me.

Or better wait until I finish the Linux part, it approaches things slightly differently and might need adjustments.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

New options page (not yet pushed to this PR). Some things might still be slightly reworded, and need to run a few tests for better default QP values. But conceptually it should be done, and I think I will be able to backport it to ffmpeg nvenc backend so we don't end up with two pages.
2023-07-18 15_21_35-Sunshine — Mozilla Firefox

Comment threadsrc/platform/linux/vaapi.cpp
@cgutman

Copy link
Copy Markdown
Collaborator

LGTM. Let's try to get this refactoring in soon if possible. I've got some AV1 and client-side cursor changes that I'd like to work on, but I don't want to stomp all over your work.

If you can back out any debugging changes or WIP test changes (the scheduling priority revert?), I can test this across all my systems and we can get it merged.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

LGTM. Let's try to get this refactoring in soon if possible. I've got some AV1 and client-side cursor changes that I'd like to work on, but I don't want to stomp all over your work.

If you can back out any debugging changes or WIP test changes (the scheduling priority revert?), I can test this across all my systems and we can get it merged.

Understandable. Wired the encoder to use the existing configuration page, since some new options pretty much depend on whether I can make it work with realtime priority (and priority selection itself will add another option).

Haven't tested native HDR (only SDR in BT.2020) and CUDA path on Linux (didn't make drastic changes, but there's always a chance). Other than that, should be alright to merge in its current state. The rest of the features shouldn't produce conflicts, and can be done in later PRs.

@ns6089
ns6089 marked this pull request as ready for review August 4, 2023 12:31
Comment threadthird-party/nv-codec-headers
@ns6089
ns6089force-pushed the nvenc branch 2 times, most recently from 3ea8f10 to a318da3CompareAugust 5, 2023 13:59
Comment threadsrc/nvenc/nvenc_base.cpp
Comment threadsrc/nvenc/nvenc_utils.cpp
Comment threadsrc/video_colorspace.cpp Outdated
@cgutman

Copy link
Copy Markdown
Collaborator

The latest changes look good. Once you squash the fixup commits and rebase, I'll do a final testing pass on my local machines and we can get this in.

Shouldn't matter on x64 since everything is fastcall here, but cdecl is
the correct declaration.
@ns6089

Copy link
Copy Markdown
ContributorAuthor

Squashed and rebased. Also added small "fix" for nvapi, it doesn't affect anything since we're strictly x64, so no point in opening full pull request for it.

@cgutman

Copy link
Copy Markdown
Collaborator

Everything looks good in my tests:

  • HDR with and without NVENC
  • NvFBC on Linux
  • All valid colorspace and color range combos

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

No open projects
Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@ns6089@cgutman@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add standalone NVENC encoder - #1427

Merged
cgutman merged 4 commits into
LizardByte:nightlyfrom
ns6089:nvenc
Aug 13, 2023
Merged

Add standalone NVENC encoder#1427
cgutman merged 4 commits into
LizardByte:nightlyfrom
ns6089:nvenc

Conversation

@ns6089

@ns6089ns6089 commented Jul 7, 2023

Copy link
Copy Markdown
Contributor

Description

Add standalone NVENC encoder for reference frames invalidation right now. And possibly for VFR-like bitrate adjustment and more somewhere down the line. Windows version is fully functional . Linux cuda (and possibly opengl) support is out of scope for this PR, but can easily be done later (at least the encoder side),

In later PRs

  • Investigate and test the viability of "increased vbv" option. Ideally, P-frames should not steal bitrate budget from future frames, maybe this can be achieved with vbv offset in encoder.
  • Investigate the viability of using multiple ref frames (L0 > 1) since we apply strict limits on the vbv. Maybe nvenc can intelligently pick blocks from previous frames if they have higher qp? Can't imagine this being free though.
  • Run VMAF benchmark for const-qp mode (document the process in .md file), and pick default values for min-qp
  • Check GFE default ref frame buffer values, particularly for h264 (level4 vs level5)
  • New configuration page and documentation

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

Things to implement before merge (may be expanded)

  • Refactor colospace selection logic, it has too much duplication right now
  • Decide how to handle encoder ref frames buffer size, currently it's at 16 frames (dynamically based on resolution and framerate? through configuration?)
    Result: use 5 ref frames buffer, which corresponds to DPB=5 in h264 terms, and DPB=6 in HEVC. Both are required minimums that must be supported by decoder profiles for given resolution.
    Update: h265 level 4 only supports 4 ref frames, but in this case the client can set the limit, this level is very outdated nowadays
  • Investigate if NVENC can be told explicitly which ref frame to use, this can allow wider decoder support for ref frames invalidation.
    Verdict: when L0 (forward prediction) list size is 1, it should always use last frame as reference,
  • Try to patch num_ref_frames in SPS header
    Verdict: may be possible for h264, close to impossible for HEVC, either way too much hassle
  • Test encoder caps
    NV_ENC_CAPS_SUPPORT_CABAC
    NV_ENC_CAPS_WIDTH_MAX
    NV_ENC_CAPS_HEIGHT_MAX
    NV_ENC_CAPS_SUPPORT_CUSTOM_VBV_BUF_SIZE
    NV_ENC_CAPS_SUPPORT_REF_PIC_INVALIDATION
    NV_ENC_CAPS_SUPPORT_YUV444_ENCODE
    NV_ENC_CAPS_SUPPORT_10BIT_ENCODE
    NV_ENC_CAPS_SUPPORT_MULTIPLE_REF_FRAMES
    NV_ENC_CAPS_SUPPORT_QPELMV
  • Switch to nv-codec-headers
  • Update Linux and MacOS platforms to new structures
  • Use Peak-Signal-to-Noise-Ratio (Y-PSNR), Structural Similarity Index (Y-SSIM), and Video Multimethod Assessment Fusion (VMAF) for default min qp values Not in this PR
  • Add configuration page and documentation Not in this PR

From review

  • Look into entropyCodingMode encoding parameter
  • Proper cleanup in create_encoder()
  • Look into last_encoder_probe_supported_invalidate_ref_frames, if needs to be changed in multiple places
  • Lock h264 into High profile

Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/platform/windows/display_vram.cpp Outdated
Comment threadsrc/platform/windows/display_vram.cpp
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/video.cpp Outdated
@ns6089
ns6089force-pushed the nvenc branch 2 times, most recently from 0d3cee2 to 6836c6eCompareJuly 9, 2023 15:20
@ns6089ns6089 self-assigned this Jul 11, 2023
@ns6089

Copy link
Copy Markdown
ContributorAuthor

@cgutman, need your opinion on colorspace refactoring a2f34ab. When you have time, it's not blocking me.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

@cgutman, need your opinion on colorspace refactoring a2f34ab. When you have time, it's not blocking me.

Or better wait until I finish the Linux part, it approaches things slightly differently and might need adjustments.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

New options page (not yet pushed to this PR). Some things might still be slightly reworded, and need to run a few tests for better default QP values. But conceptually it should be done, and I think I will be able to backport it to ffmpeg nvenc backend so we don't end up with two pages.
2023-07-18 15_21_35-Sunshine — Mozilla Firefox

Comment threadsrc/platform/linux/vaapi.cpp
@cgutman

Copy link
Copy Markdown
Collaborator

LGTM. Let's try to get this refactoring in soon if possible. I've got some AV1 and client-side cursor changes that I'd like to work on, but I don't want to stomp all over your work.

If you can back out any debugging changes or WIP test changes (the scheduling priority revert?), I can test this across all my systems and we can get it merged.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

LGTM. Let's try to get this refactoring in soon if possible. I've got some AV1 and client-side cursor changes that I'd like to work on, but I don't want to stomp all over your work.

If you can back out any debugging changes or WIP test changes (the scheduling priority revert?), I can test this across all my systems and we can get it merged.

Understandable. Wired the encoder to use the existing configuration page, since some new options pretty much depend on whether I can make it work with realtime priority (and priority selection itself will add another option).

Haven't tested native HDR (only SDR in BT.2020) and CUDA path on Linux (didn't make drastic changes, but there's always a chance). Other than that, should be alright to merge in its current state. The rest of the features shouldn't produce conflicts, and can be done in later PRs.

@ns6089
ns6089 marked this pull request as ready for review August 4, 2023 12:31
Comment threadthird-party/nv-codec-headers
@ns6089
ns6089force-pushed the nvenc branch 2 times, most recently from 3ea8f10 to a318da3CompareAugust 5, 2023 13:59
Comment threadsrc/nvenc/nvenc_base.cpp
Comment threadsrc/nvenc/nvenc_utils.cpp
Comment threadsrc/video_colorspace.cpp Outdated
@cgutman

Copy link
Copy Markdown
Collaborator

The latest changes look good. Once you squash the fixup commits and rebase, I'll do a final testing pass on my local machines and we can get this in.

Shouldn't matter on x64 since everything is fastcall here, but cdecl is
the correct declaration.
@ns6089

Copy link
Copy Markdown
ContributorAuthor

Squashed and rebased. Also added small "fix" for nvapi, it doesn't affect anything since we're strictly x64, so no point in opening full pull request for it.

@cgutman

Copy link
Copy Markdown
Collaborator

Everything looks good in my tests:

  • HDR with and without NVENC
  • NvFBC on Linux
  • All valid colorspace and color range combos

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

No open projects
Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@ns6089@cgutman@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Add standalone NVENC encoder - #1427

Merged
cgutman merged 4 commits into
LizardByte:nightlyfrom
ns6089:nvenc
Aug 13, 2023
Merged

Add standalone NVENC encoder#1427
cgutman merged 4 commits into
LizardByte:nightlyfrom
ns6089:nvenc

Conversation

@ns6089

@ns6089ns6089 commented Jul 7, 2023

Copy link
Copy Markdown
Contributor

Description

Add standalone NVENC encoder for reference frames invalidation right now. And possibly for VFR-like bitrate adjustment and more somewhere down the line. Windows version is fully functional . Linux cuda (and possibly opengl) support is out of scope for this PR, but can easily be done later (at least the encoder side),

In later PRs

  • Investigate and test the viability of "increased vbv" option. Ideally, P-frames should not steal bitrate budget from future frames, maybe this can be achieved with vbv offset in encoder.
  • Investigate the viability of using multiple ref frames (L0 > 1) since we apply strict limits on the vbv. Maybe nvenc can intelligently pick blocks from previous frames if they have higher qp? Can't imagine this being free though.
  • Run VMAF benchmark for const-qp mode (document the process in .md file), and pick default values for min-qp
  • Check GFE default ref frame buffer values, particularly for h264 (level4 vs level5)
  • New configuration page and documentation

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

Things to implement before merge (may be expanded)

  • Refactor colospace selection logic, it has too much duplication right now
  • Decide how to handle encoder ref frames buffer size, currently it's at 16 frames (dynamically based on resolution and framerate? through configuration?)
    Result: use 5 ref frames buffer, which corresponds to DPB=5 in h264 terms, and DPB=6 in HEVC. Both are required minimums that must be supported by decoder profiles for given resolution.
    Update: h265 level 4 only supports 4 ref frames, but in this case the client can set the limit, this level is very outdated nowadays
  • Investigate if NVENC can be told explicitly which ref frame to use, this can allow wider decoder support for ref frames invalidation.
    Verdict: when L0 (forward prediction) list size is 1, it should always use last frame as reference,
  • Try to patch num_ref_frames in SPS header
    Verdict: may be possible for h264, close to impossible for HEVC, either way too much hassle
  • Test encoder caps
    NV_ENC_CAPS_SUPPORT_CABAC
    NV_ENC_CAPS_WIDTH_MAX
    NV_ENC_CAPS_HEIGHT_MAX
    NV_ENC_CAPS_SUPPORT_CUSTOM_VBV_BUF_SIZE
    NV_ENC_CAPS_SUPPORT_REF_PIC_INVALIDATION
    NV_ENC_CAPS_SUPPORT_YUV444_ENCODE
    NV_ENC_CAPS_SUPPORT_10BIT_ENCODE
    NV_ENC_CAPS_SUPPORT_MULTIPLE_REF_FRAMES
    NV_ENC_CAPS_SUPPORT_QPELMV
  • Switch to nv-codec-headers
  • Update Linux and MacOS platforms to new structures
  • Use Peak-Signal-to-Noise-Ratio (Y-PSNR), Structural Similarity Index (Y-SSIM), and Video Multimethod Assessment Fusion (VMAF) for default min qp values Not in this PR
  • Add configuration page and documentation Not in this PR

From review

  • Look into entropyCodingMode encoding parameter
  • Proper cleanup in create_encoder()
  • Look into last_encoder_probe_supported_invalidate_ref_frames, if needs to be changed in multiple places
  • Lock h264 into High profile

Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/platform/windows/display_vram.cpp Outdated
Comment threadsrc/platform/windows/display_vram.cpp
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/video.cpp Outdated
@ns6089
ns6089force-pushed the nvenc branch 2 times, most recently from 0d3cee2 to 6836c6eCompareJuly 9, 2023 15:20
@ns6089ns6089 self-assigned this Jul 11, 2023
@ns6089

Copy link
Copy Markdown
ContributorAuthor

@cgutman, need your opinion on colorspace refactoring a2f34ab. When you have time, it's not blocking me.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

@cgutman, need your opinion on colorspace refactoring a2f34ab. When you have time, it's not blocking me.

Or better wait until I finish the Linux part, it approaches things slightly differently and might need adjustments.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

New options page (not yet pushed to this PR). Some things might still be slightly reworded, and need to run a few tests for better default QP values. But conceptually it should be done, and I think I will be able to backport it to ffmpeg nvenc backend so we don't end up with two pages.
2023-07-18 15_21_35-Sunshine — Mozilla Firefox

Comment threadsrc/platform/linux/vaapi.cpp
@cgutman

Copy link
Copy Markdown
Collaborator

LGTM. Let's try to get this refactoring in soon if possible. I've got some AV1 and client-side cursor changes that I'd like to work on, but I don't want to stomp all over your work.

If you can back out any debugging changes or WIP test changes (the scheduling priority revert?), I can test this across all my systems and we can get it merged.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

LGTM. Let's try to get this refactoring in soon if possible. I've got some AV1 and client-side cursor changes that I'd like to work on, but I don't want to stomp all over your work.

If you can back out any debugging changes or WIP test changes (the scheduling priority revert?), I can test this across all my systems and we can get it merged.

Understandable. Wired the encoder to use the existing configuration page, since some new options pretty much depend on whether I can make it work with realtime priority (and priority selection itself will add another option).

Haven't tested native HDR (only SDR in BT.2020) and CUDA path on Linux (didn't make drastic changes, but there's always a chance). Other than that, should be alright to merge in its current state. The rest of the features shouldn't produce conflicts, and can be done in later PRs.

@ns6089
ns6089 marked this pull request as ready for review August 4, 2023 12:31
Comment threadthird-party/nv-codec-headers
@ns6089
ns6089force-pushed the nvenc branch 2 times, most recently from 3ea8f10 to a318da3CompareAugust 5, 2023 13:59
Comment threadsrc/nvenc/nvenc_base.cpp
Comment threadsrc/nvenc/nvenc_utils.cpp
Comment threadsrc/video_colorspace.cpp Outdated
@cgutman

Copy link
Copy Markdown
Collaborator

The latest changes look good. Once you squash the fixup commits and rebase, I'll do a final testing pass on my local machines and we can get this in.

Shouldn't matter on x64 since everything is fastcall here, but cdecl is
the correct declaration.
@ns6089

Copy link
Copy Markdown
ContributorAuthor

Squashed and rebased. Also added small "fix" for nvapi, it doesn't affect anything since we're strictly x64, so no point in opening full pull request for it.

@cgutman

Copy link
Copy Markdown
Collaborator

Everything looks good in my tests:

  • HDR with and without NVENC
  • NvFBC on Linux
  • All valid colorspace and color range combos

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

No open projects
Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@ns6089@cgutman@ReenigneArcher
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Add standalone NVENC encoder - #1427

Merged
cgutman merged 4 commits into
LizardByte:nightlyfrom
ns6089:nvenc
Aug 13, 2023
Merged

Add standalone NVENC encoder#1427
cgutman merged 4 commits into
LizardByte:nightlyfrom
ns6089:nvenc

Conversation

@ns6089

@ns6089ns6089 commented Jul 7, 2023

Copy link
Copy Markdown
Contributor

Description

Add standalone NVENC encoder for reference frames invalidation right now. And possibly for VFR-like bitrate adjustment and more somewhere down the line. Windows version is fully functional . Linux cuda (and possibly opengl) support is out of scope for this PR, but can easily be done later (at least the encoder side),

In later PRs

  • Investigate and test the viability of "increased vbv" option. Ideally, P-frames should not steal bitrate budget from future frames, maybe this can be achieved with vbv offset in encoder.
  • Investigate the viability of using multiple ref frames (L0 > 1) since we apply strict limits on the vbv. Maybe nvenc can intelligently pick blocks from previous frames if they have higher qp? Can't imagine this being free though.
  • Run VMAF benchmark for const-qp mode (document the process in .md file), and pick default values for min-qp
  • Check GFE default ref frame buffer values, particularly for h264 (level4 vs level5)
  • New configuration page and documentation

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Dependency update (updates to dependencies)
  • Documentation update (changes to documentation)
  • Repository update (changes to repository files, e.g. .github/...)

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code, particularly in hard-to-understand areas
  • I have added or updated the in code docstring/documentation-blocks for new or existing methods/components

Branch Updates

LizardByte requires that branches be up-to-date before merging. This means that after any PR is merged, this branch
must be updated before it can be merged. You must also
Allow edits from maintainers.

  • I want maintainers to keep my branch updated

Things to implement before merge (may be expanded)

  • Refactor colospace selection logic, it has too much duplication right now
  • Decide how to handle encoder ref frames buffer size, currently it's at 16 frames (dynamically based on resolution and framerate? through configuration?)
    Result: use 5 ref frames buffer, which corresponds to DPB=5 in h264 terms, and DPB=6 in HEVC. Both are required minimums that must be supported by decoder profiles for given resolution.
    Update: h265 level 4 only supports 4 ref frames, but in this case the client can set the limit, this level is very outdated nowadays
  • Investigate if NVENC can be told explicitly which ref frame to use, this can allow wider decoder support for ref frames invalidation.
    Verdict: when L0 (forward prediction) list size is 1, it should always use last frame as reference,
  • Try to patch num_ref_frames in SPS header
    Verdict: may be possible for h264, close to impossible for HEVC, either way too much hassle
  • Test encoder caps
    NV_ENC_CAPS_SUPPORT_CABAC
    NV_ENC_CAPS_WIDTH_MAX
    NV_ENC_CAPS_HEIGHT_MAX
    NV_ENC_CAPS_SUPPORT_CUSTOM_VBV_BUF_SIZE
    NV_ENC_CAPS_SUPPORT_REF_PIC_INVALIDATION
    NV_ENC_CAPS_SUPPORT_YUV444_ENCODE
    NV_ENC_CAPS_SUPPORT_10BIT_ENCODE
    NV_ENC_CAPS_SUPPORT_MULTIPLE_REF_FRAMES
    NV_ENC_CAPS_SUPPORT_QPELMV
  • Switch to nv-codec-headers
  • Update Linux and MacOS platforms to new structures
  • Use Peak-Signal-to-Noise-Ratio (Y-PSNR), Structural Similarity Index (Y-SSIM), and Video Multimethod Assessment Fusion (VMAF) for default min qp values Not in this PR
  • Add configuration page and documentation Not in this PR

From review

  • Look into entropyCodingMode encoding parameter
  • Proper cleanup in create_encoder()
  • Look into last_encoder_probe_supported_invalidate_ref_frames, if needs to be changed in multiple places
  • Lock h264 into High profile

Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/platform/windows/display_vram.cpp Outdated
Comment threadsrc/platform/windows/display_vram.cpp
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/nvenc/nvenc_base.cpp Outdated
Comment threadsrc/video.cpp Outdated
@ns6089
ns6089force-pushed the nvenc branch 2 times, most recently from 0d3cee2 to 6836c6eCompareJuly 9, 2023 15:20
@ns6089ns6089 self-assigned this Jul 11, 2023
@ns6089

Copy link
Copy Markdown
ContributorAuthor

@cgutman, need your opinion on colorspace refactoring a2f34ab. When you have time, it's not blocking me.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

@cgutman, need your opinion on colorspace refactoring a2f34ab. When you have time, it's not blocking me.

Or better wait until I finish the Linux part, it approaches things slightly differently and might need adjustments.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

New options page (not yet pushed to this PR). Some things might still be slightly reworded, and need to run a few tests for better default QP values. But conceptually it should be done, and I think I will be able to backport it to ffmpeg nvenc backend so we don't end up with two pages.
2023-07-18 15_21_35-Sunshine — Mozilla Firefox

Comment threadsrc/platform/linux/vaapi.cpp
@cgutman

Copy link
Copy Markdown
Collaborator

LGTM. Let's try to get this refactoring in soon if possible. I've got some AV1 and client-side cursor changes that I'd like to work on, but I don't want to stomp all over your work.

If you can back out any debugging changes or WIP test changes (the scheduling priority revert?), I can test this across all my systems and we can get it merged.

@ns6089

Copy link
Copy Markdown
ContributorAuthor

LGTM. Let's try to get this refactoring in soon if possible. I've got some AV1 and client-side cursor changes that I'd like to work on, but I don't want to stomp all over your work.

If you can back out any debugging changes or WIP test changes (the scheduling priority revert?), I can test this across all my systems and we can get it merged.

Understandable. Wired the encoder to use the existing configuration page, since some new options pretty much depend on whether I can make it work with realtime priority (and priority selection itself will add another option).

Haven't tested native HDR (only SDR in BT.2020) and CUDA path on Linux (didn't make drastic changes, but there's always a chance). Other than that, should be alright to merge in its current state. The rest of the features shouldn't produce conflicts, and can be done in later PRs.

@ns6089
ns6089 marked this pull request as ready for review August 4, 2023 12:31
Comment threadthird-party/nv-codec-headers
@ns6089
ns6089force-pushed the nvenc branch 2 times, most recently from 3ea8f10 to a318da3CompareAugust 5, 2023 13:59
Comment threadsrc/nvenc/nvenc_base.cpp
Comment threadsrc/nvenc/nvenc_utils.cpp
Comment threadsrc/video_colorspace.cpp Outdated
@cgutman

Copy link
Copy Markdown
Collaborator

The latest changes look good. Once you squash the fixup commits and rebase, I'll do a final testing pass on my local machines and we can get this in.

Shouldn't matter on x64 since everything is fastcall here, but cdecl is
the correct declaration.
@ns6089

Copy link
Copy Markdown
ContributorAuthor

Squashed and rebased. Also added small "fix" for nvapi, it doesn't affect anything since we're strictly x64, so no point in opening full pull request for it.

@cgutman

Copy link
Copy Markdown
Collaborator

Everything looks good in my tests:

  • HDR with and without NVENC
  • NvFBC on Linux
  • All valid colorspace and color range combos

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

No open projects
Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants

@ns6089@cgutman@ReenigneArcher