Support Base 64 URL #1658

Description

@commonsensesoftware

Updated by @MihaZupan on 2024-02-27

Proposed API

namespaceSystem.Buffers.Text;publicstaticclassBase64Url{// Encode bytes => utf8publicstaticOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticOperationStatusEncodeToUtf8InPlace(Span<byte>buffer,intdataLength,outintbytesWritten);// Decode from utf8 => bytespublicstaticOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticOperationStatusDecodeFromUtf8InPlace(Span<byte>buffer,outintbytesWritten);// Max length APIspublicstaticintGetMaxDecodedFromUtf8Length(intlength);publicstaticintGetMaxEncodedToUtf8Length(intlength);// IsValidpublicstaticboolIsValid(ReadOnlySpan<char>base64UrlText);publicstaticboolIsValid(ReadOnlySpan<char>base64UrlText,outintdecodedLength);publicstaticboolIsValid(ReadOnlySpan<byte>base64UrlTextUtf8);publicstaticboolIsValid(ReadOnlySpan<byte>base64UrlTextUtf8,outintdecodedLength);// Up to this point, this is a mirror of System.Buffers.Text.Base64// Below are more helpers that bring over functionality similar to Convert.*Base64*// Encode to / decode from charspublicstaticboolTryEncodeToChars(ReadOnlySpan<byte>bytes,Span<char>chars,outintcharsWritten){}publicstaticboolTryDecodeFromChars(ReadOnlySpan<char>chars,Span<byte>bytes,outintbytesWritten){}// These are just accelerator methods.// Should be efficiently implementable on top of the other ones in just a few lines.// Encode to stringpublicstaticstringEncodeToString(ReadOnlySpan<char>chars,Encodingencoding){}publicstaticstringEncodeToString(ReadOnlySpan<byte>bytes){}// Decode from chars => string// Decode from chars => byte[]// The names could also just be "Decode" without naming the return typepublicstaticstringDecodeToString(ReadOnlySpan<char>chars,Encodingencoding){}publicstaticbyte[]DecodeToByteArray(ReadOnlySpan<char>chars){}}

Original issue

The Base64 implementation in System.Memory provides excellent low-level optimizations for RFC 4648, but it currently uses a fixed alphabet. Minor refactoring would add additional use cases using the existing implementation.

Rationale and Usage

The most obvious use case for this change is to support the encoding variant known as Base 64 URL also described in RFC 4648 §4.

This excerpt from the RFC describes the difference between the standard Base 64 alphabet and the Base 64 URL alphabet.

This encoding is technically identical to the previous one, except
for the 62:nd and 63:rd alphabet character...

I have already been able to verify the encoding and decoding in a copy of the current implementation by merely changing the 2 relevant characters in the alphabet mappings at:

Today, this logic is suboptimally implemented in at least:

Furthermore, the encoding is generic and has use cases outside of ASP.NET (e.g. you shouldn't have to reference ASP.NET to use it).

Proposed API Change

The existing Base64.EncodeToUtf8 and Base64.DecodeFromUtf8 should each add a new method overload that allows one of the following:

  1. A new enumeration of allowed alphabets (ex: Base64Alphabet) which internally maps to well-known alphabet spans

  2. Allow any custom alphabet to be supplied as ReadOnlySpan<sbyte> that must have an exact length of 64 for encoding and 256 for decoding

    a. One or more new types could be provided with static properties for the well-known alphabet spans

Details

Option 1 requires less validation, but has less flexibility. Option 2 has more flexibility, including scenarios not described here, but requires more validation before using the alphabet.

Although the standard Base 64 and Base 64 URL are technically the same encoding with different alphabets, Base 64 URL typically does not include padding. RFC 4648 §3.2 indicates this is allowed (as it's explicitly stated), but there is no need to make that concession in this API. Including the = character for padding in Base 64 URL encoding is still correct.

Both approaches would have a similar looking API:

publicstaticunsafeOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,Base64Alphabetalphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticunsafeOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,Base64Alphabetalphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);

Figure 1: Supply an alphabet enumeration

publicstaticunsafeOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,ReadOnlySpan<byte>alphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticunsafeOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,ReadOnlySpan<byte>alphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);

Figure 2: Supply a custom alphabet

The support for trimming off and re-adding padding should be implemented separately. This could be in a separate Base64Url class or, perhaps more appropriately, added as a new type in System.Text.Encodings.Web.

Padding is generally very cheap to deal with compared to other implementation methods. Trimming involves walking the tail end of the span while there are padding characters and then slicing it off. Re-padding fills the end of the span buffer before decoding. Neither operation requires additional allocations.

Open Questions

  • Does Option 1 or Option 2 make more sense?
  • To complete the cycle, it feels like there should be a type that fully handles the encoding and decoding with trimmed padding. Does that make sense to be in System.Memory, System.Text.Encodings.Web, or perhaps some other place?

Metadata

Metadata

Assignees

Labels

api-approvedAPI was approved in API review, it can be implementedarea-System.Netin-prThere is an active PR which will close this issue when it is merged

Type

No type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions

    , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
     blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
    }
    } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
    })();
    (function(){
    try {
    var __m = "github.com";
    var __re = new RegExp('^' + "github\\.com" + '
    
    Skip to content

    Support Base 64 URL #1658

    Description

    @commonsensesoftware

    Updated by @MihaZupan on 2024-02-27

    Proposed API

    namespaceSystem.Buffers.Text;publicstaticclassBase64Url{// Encode bytes => utf8publicstaticOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticOperationStatusEncodeToUtf8InPlace(Span<byte>buffer,intdataLength,outintbytesWritten);// Decode from utf8 => bytespublicstaticOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticOperationStatusDecodeFromUtf8InPlace(Span<byte>buffer,outintbytesWritten);// Max length APIspublicstaticintGetMaxDecodedFromUtf8Length(intlength);publicstaticintGetMaxEncodedToUtf8Length(intlength);// IsValidpublicstaticboolIsValid(ReadOnlySpan<char>base64UrlText);publicstaticboolIsValid(ReadOnlySpan<char>base64UrlText,outintdecodedLength);publicstaticboolIsValid(ReadOnlySpan<byte>base64UrlTextUtf8);publicstaticboolIsValid(ReadOnlySpan<byte>base64UrlTextUtf8,outintdecodedLength);// Up to this point, this is a mirror of System.Buffers.Text.Base64// Below are more helpers that bring over functionality similar to Convert.*Base64*// Encode to / decode from charspublicstaticboolTryEncodeToChars(ReadOnlySpan<byte>bytes,Span<char>chars,outintcharsWritten){}publicstaticboolTryDecodeFromChars(ReadOnlySpan<char>chars,Span<byte>bytes,outintbytesWritten){}// These are just accelerator methods.// Should be efficiently implementable on top of the other ones in just a few lines.// Encode to stringpublicstaticstringEncodeToString(ReadOnlySpan<char>chars,Encodingencoding){}publicstaticstringEncodeToString(ReadOnlySpan<byte>bytes){}// Decode from chars => string// Decode from chars => byte[]// The names could also just be "Decode" without naming the return typepublicstaticstringDecodeToString(ReadOnlySpan<char>chars,Encodingencoding){}publicstaticbyte[]DecodeToByteArray(ReadOnlySpan<char>chars){}}

    Original issue

    The Base64 implementation in System.Memory provides excellent low-level optimizations for RFC 4648, but it currently uses a fixed alphabet. Minor refactoring would add additional use cases using the existing implementation.

    Rationale and Usage

    The most obvious use case for this change is to support the encoding variant known as Base 64 URL also described in RFC 4648 §4.

    This excerpt from the RFC describes the difference between the standard Base 64 alphabet and the Base 64 URL alphabet.

    This encoding is technically identical to the previous one, except
    for the 62:nd and 63:rd alphabet character...

    I have already been able to verify the encoding and decoding in a copy of the current implementation by merely changing the 2 relevant characters in the alphabet mappings at:

    Today, this logic is suboptimally implemented in at least:

    Furthermore, the encoding is generic and has use cases outside of ASP.NET (e.g. you shouldn't have to reference ASP.NET to use it).

    Proposed API Change

    The existing Base64.EncodeToUtf8 and Base64.DecodeFromUtf8 should each add a new method overload that allows one of the following:

    1. A new enumeration of allowed alphabets (ex: Base64Alphabet) which internally maps to well-known alphabet spans

    2. Allow any custom alphabet to be supplied as ReadOnlySpan<sbyte> that must have an exact length of 64 for encoding and 256 for decoding

      a. One or more new types could be provided with static properties for the well-known alphabet spans

    Details

    Option 1 requires less validation, but has less flexibility. Option 2 has more flexibility, including scenarios not described here, but requires more validation before using the alphabet.

    Although the standard Base 64 and Base 64 URL are technically the same encoding with different alphabets, Base 64 URL typically does not include padding. RFC 4648 §3.2 indicates this is allowed (as it's explicitly stated), but there is no need to make that concession in this API. Including the = character for padding in Base 64 URL encoding is still correct.

    Both approaches would have a similar looking API:

    publicstaticunsafeOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,Base64Alphabetalphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticunsafeOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,Base64Alphabetalphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);

    Figure 1: Supply an alphabet enumeration

    publicstaticunsafeOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,ReadOnlySpan<byte>alphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticunsafeOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,ReadOnlySpan<byte>alphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);

    Figure 2: Supply a custom alphabet

    The support for trimming off and re-adding padding should be implemented separately. This could be in a separate Base64Url class or, perhaps more appropriately, added as a new type in System.Text.Encodings.Web.

    Padding is generally very cheap to deal with compared to other implementation methods. Trimming involves walking the tail end of the span while there are padding characters and then slicing it off. Re-padding fills the end of the span buffer before decoding. Neither operation requires additional allocations.

    Open Questions

    • Does Option 1 or Option 2 make more sense?
    • To complete the cycle, it feels like there should be a type that fully handles the encoding and decoding with trimmed padding. Does that make sense to be in System.Memory, System.Text.Encodings.Web, or perhaps some other place?

    Metadata

    Metadata

    Assignees

    Labels

    api-approvedAPI was approved in API review, it can be implementedarea-System.Netin-prThere is an active PR which will close this issue when it is merged

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
      Skip to content

      Support Base 64 URL #1658

      Description

      @commonsensesoftware

      Updated by @MihaZupan on 2024-02-27

      Proposed API

      namespaceSystem.Buffers.Text;publicstaticclassBase64Url{// Encode bytes => utf8publicstaticOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticOperationStatusEncodeToUtf8InPlace(Span<byte>buffer,intdataLength,outintbytesWritten);// Decode from utf8 => bytespublicstaticOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticOperationStatusDecodeFromUtf8InPlace(Span<byte>buffer,outintbytesWritten);// Max length APIspublicstaticintGetMaxDecodedFromUtf8Length(intlength);publicstaticintGetMaxEncodedToUtf8Length(intlength);// IsValidpublicstaticboolIsValid(ReadOnlySpan<char>base64UrlText);publicstaticboolIsValid(ReadOnlySpan<char>base64UrlText,outintdecodedLength);publicstaticboolIsValid(ReadOnlySpan<byte>base64UrlTextUtf8);publicstaticboolIsValid(ReadOnlySpan<byte>base64UrlTextUtf8,outintdecodedLength);// Up to this point, this is a mirror of System.Buffers.Text.Base64// Below are more helpers that bring over functionality similar to Convert.*Base64*// Encode to / decode from charspublicstaticboolTryEncodeToChars(ReadOnlySpan<byte>bytes,Span<char>chars,outintcharsWritten){}publicstaticboolTryDecodeFromChars(ReadOnlySpan<char>chars,Span<byte>bytes,outintbytesWritten){}// These are just accelerator methods.// Should be efficiently implementable on top of the other ones in just a few lines.// Encode to stringpublicstaticstringEncodeToString(ReadOnlySpan<char>chars,Encodingencoding){}publicstaticstringEncodeToString(ReadOnlySpan<byte>bytes){}// Decode from chars => string// Decode from chars => byte[]// The names could also just be "Decode" without naming the return typepublicstaticstringDecodeToString(ReadOnlySpan<char>chars,Encodingencoding){}publicstaticbyte[]DecodeToByteArray(ReadOnlySpan<char>chars){}}

      Original issue

      The Base64 implementation in System.Memory provides excellent low-level optimizations for RFC 4648, but it currently uses a fixed alphabet. Minor refactoring would add additional use cases using the existing implementation.

      Rationale and Usage

      The most obvious use case for this change is to support the encoding variant known as Base 64 URL also described in RFC 4648 §4.

      This excerpt from the RFC describes the difference between the standard Base 64 alphabet and the Base 64 URL alphabet.

      This encoding is technically identical to the previous one, except
      for the 62:nd and 63:rd alphabet character...

      I have already been able to verify the encoding and decoding in a copy of the current implementation by merely changing the 2 relevant characters in the alphabet mappings at:

      Today, this logic is suboptimally implemented in at least:

      Furthermore, the encoding is generic and has use cases outside of ASP.NET (e.g. you shouldn't have to reference ASP.NET to use it).

      Proposed API Change

      The existing Base64.EncodeToUtf8 and Base64.DecodeFromUtf8 should each add a new method overload that allows one of the following:

      1. A new enumeration of allowed alphabets (ex: Base64Alphabet) which internally maps to well-known alphabet spans

      2. Allow any custom alphabet to be supplied as ReadOnlySpan<sbyte> that must have an exact length of 64 for encoding and 256 for decoding

        a. One or more new types could be provided with static properties for the well-known alphabet spans

      Details

      Option 1 requires less validation, but has less flexibility. Option 2 has more flexibility, including scenarios not described here, but requires more validation before using the alphabet.

      Although the standard Base 64 and Base 64 URL are technically the same encoding with different alphabets, Base 64 URL typically does not include padding. RFC 4648 §3.2 indicates this is allowed (as it's explicitly stated), but there is no need to make that concession in this API. Including the = character for padding in Base 64 URL encoding is still correct.

      Both approaches would have a similar looking API:

      publicstaticunsafeOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,Base64Alphabetalphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticunsafeOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,Base64Alphabetalphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);

      Figure 1: Supply an alphabet enumeration

      publicstaticunsafeOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,ReadOnlySpan<byte>alphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticunsafeOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,ReadOnlySpan<byte>alphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);

      Figure 2: Supply a custom alphabet

      The support for trimming off and re-adding padding should be implemented separately. This could be in a separate Base64Url class or, perhaps more appropriately, added as a new type in System.Text.Encodings.Web.

      Padding is generally very cheap to deal with compared to other implementation methods. Trimming involves walking the tail end of the span while there are padding characters and then slicing it off. Re-padding fills the end of the span buffer before decoding. Neither operation requires additional allocations.

      Open Questions

      • Does Option 1 or Option 2 make more sense?
      • To complete the cycle, it feels like there should be a type that fully handles the encoding and decoding with trimmed padding. Does that make sense to be in System.Memory, System.Text.Encodings.Web, or perhaps some other place?

      Metadata

      Metadata

      Assignees

      Labels

      api-approvedAPI was approved in API review, it can be implementedarea-System.Netin-prThere is an active PR which will close this issue when it is merged

      Type

      No type

      Projects

      No projects

        Milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions

        , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
        Skip to content

        Support Base 64 URL #1658

        Description

        @commonsensesoftware

        Updated by @MihaZupan on 2024-02-27

        Proposed API

        namespaceSystem.Buffers.Text;publicstaticclassBase64Url{// Encode bytes => utf8publicstaticOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticOperationStatusEncodeToUtf8InPlace(Span<byte>buffer,intdataLength,outintbytesWritten);// Decode from utf8 => bytespublicstaticOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticOperationStatusDecodeFromUtf8InPlace(Span<byte>buffer,outintbytesWritten);// Max length APIspublicstaticintGetMaxDecodedFromUtf8Length(intlength);publicstaticintGetMaxEncodedToUtf8Length(intlength);// IsValidpublicstaticboolIsValid(ReadOnlySpan<char>base64UrlText);publicstaticboolIsValid(ReadOnlySpan<char>base64UrlText,outintdecodedLength);publicstaticboolIsValid(ReadOnlySpan<byte>base64UrlTextUtf8);publicstaticboolIsValid(ReadOnlySpan<byte>base64UrlTextUtf8,outintdecodedLength);// Up to this point, this is a mirror of System.Buffers.Text.Base64// Below are more helpers that bring over functionality similar to Convert.*Base64*// Encode to / decode from charspublicstaticboolTryEncodeToChars(ReadOnlySpan<byte>bytes,Span<char>chars,outintcharsWritten){}publicstaticboolTryDecodeFromChars(ReadOnlySpan<char>chars,Span<byte>bytes,outintbytesWritten){}// These are just accelerator methods.// Should be efficiently implementable on top of the other ones in just a few lines.// Encode to stringpublicstaticstringEncodeToString(ReadOnlySpan<char>chars,Encodingencoding){}publicstaticstringEncodeToString(ReadOnlySpan<byte>bytes){}// Decode from chars => string// Decode from chars => byte[]// The names could also just be "Decode" without naming the return typepublicstaticstringDecodeToString(ReadOnlySpan<char>chars,Encodingencoding){}publicstaticbyte[]DecodeToByteArray(ReadOnlySpan<char>chars){}}

        Original issue

        The Base64 implementation in System.Memory provides excellent low-level optimizations for RFC 4648, but it currently uses a fixed alphabet. Minor refactoring would add additional use cases using the existing implementation.

        Rationale and Usage

        The most obvious use case for this change is to support the encoding variant known as Base 64 URL also described in RFC 4648 §4.

        This excerpt from the RFC describes the difference between the standard Base 64 alphabet and the Base 64 URL alphabet.

        This encoding is technically identical to the previous one, except
        for the 62:nd and 63:rd alphabet character...

        I have already been able to verify the encoding and decoding in a copy of the current implementation by merely changing the 2 relevant characters in the alphabet mappings at:

        Today, this logic is suboptimally implemented in at least:

        Furthermore, the encoding is generic and has use cases outside of ASP.NET (e.g. you shouldn't have to reference ASP.NET to use it).

        Proposed API Change

        The existing Base64.EncodeToUtf8 and Base64.DecodeFromUtf8 should each add a new method overload that allows one of the following:

        1. A new enumeration of allowed alphabets (ex: Base64Alphabet) which internally maps to well-known alphabet spans

        2. Allow any custom alphabet to be supplied as ReadOnlySpan<sbyte> that must have an exact length of 64 for encoding and 256 for decoding

          a. One or more new types could be provided with static properties for the well-known alphabet spans

        Details

        Option 1 requires less validation, but has less flexibility. Option 2 has more flexibility, including scenarios not described here, but requires more validation before using the alphabet.

        Although the standard Base 64 and Base 64 URL are technically the same encoding with different alphabets, Base 64 URL typically does not include padding. RFC 4648 §3.2 indicates this is allowed (as it's explicitly stated), but there is no need to make that concession in this API. Including the = character for padding in Base 64 URL encoding is still correct.

        Both approaches would have a similar looking API:

        publicstaticunsafeOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,Base64Alphabetalphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticunsafeOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,Base64Alphabetalphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);

        Figure 1: Supply an alphabet enumeration

        publicstaticunsafeOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,ReadOnlySpan<byte>alphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticunsafeOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,ReadOnlySpan<byte>alphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);

        Figure 2: Supply a custom alphabet

        The support for trimming off and re-adding padding should be implemented separately. This could be in a separate Base64Url class or, perhaps more appropriately, added as a new type in System.Text.Encodings.Web.

        Padding is generally very cheap to deal with compared to other implementation methods. Trimming involves walking the tail end of the span while there are padding characters and then slicing it off. Re-padding fills the end of the span buffer before decoding. Neither operation requires additional allocations.

        Open Questions

        • Does Option 1 or Option 2 make more sense?
        • To complete the cycle, it feels like there should be a type that fully handles the encoding and decoding with trimmed padding. Does that make sense to be in System.Memory, System.Text.Encodings.Web, or perhaps some other place?

        Metadata

        Metadata

        Assignees

        Labels

        api-approvedAPI was approved in API review, it can be implementedarea-System.Netin-prThere is an active PR which will close this issue when it is merged

        Type

        No type

        Projects

        No projects

          Milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
          Skip to content

          Support Base 64 URL #1658

          Description

          @commonsensesoftware

          Updated by @MihaZupan on 2024-02-27

          Proposed API

          namespaceSystem.Buffers.Text;publicstaticclassBase64Url{// Encode bytes => utf8publicstaticOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticOperationStatusEncodeToUtf8InPlace(Span<byte>buffer,intdataLength,outintbytesWritten);// Decode from utf8 => bytespublicstaticOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticOperationStatusDecodeFromUtf8InPlace(Span<byte>buffer,outintbytesWritten);// Max length APIspublicstaticintGetMaxDecodedFromUtf8Length(intlength);publicstaticintGetMaxEncodedToUtf8Length(intlength);// IsValidpublicstaticboolIsValid(ReadOnlySpan<char>base64UrlText);publicstaticboolIsValid(ReadOnlySpan<char>base64UrlText,outintdecodedLength);publicstaticboolIsValid(ReadOnlySpan<byte>base64UrlTextUtf8);publicstaticboolIsValid(ReadOnlySpan<byte>base64UrlTextUtf8,outintdecodedLength);// Up to this point, this is a mirror of System.Buffers.Text.Base64// Below are more helpers that bring over functionality similar to Convert.*Base64*// Encode to / decode from charspublicstaticboolTryEncodeToChars(ReadOnlySpan<byte>bytes,Span<char>chars,outintcharsWritten){}publicstaticboolTryDecodeFromChars(ReadOnlySpan<char>chars,Span<byte>bytes,outintbytesWritten){}// These are just accelerator methods.// Should be efficiently implementable on top of the other ones in just a few lines.// Encode to stringpublicstaticstringEncodeToString(ReadOnlySpan<char>chars,Encodingencoding){}publicstaticstringEncodeToString(ReadOnlySpan<byte>bytes){}// Decode from chars => string// Decode from chars => byte[]// The names could also just be "Decode" without naming the return typepublicstaticstringDecodeToString(ReadOnlySpan<char>chars,Encodingencoding){}publicstaticbyte[]DecodeToByteArray(ReadOnlySpan<char>chars){}}

          Original issue

          The Base64 implementation in System.Memory provides excellent low-level optimizations for RFC 4648, but it currently uses a fixed alphabet. Minor refactoring would add additional use cases using the existing implementation.

          Rationale and Usage

          The most obvious use case for this change is to support the encoding variant known as Base 64 URL also described in RFC 4648 §4.

          This excerpt from the RFC describes the difference between the standard Base 64 alphabet and the Base 64 URL alphabet.

          This encoding is technically identical to the previous one, except
          for the 62:nd and 63:rd alphabet character...

          I have already been able to verify the encoding and decoding in a copy of the current implementation by merely changing the 2 relevant characters in the alphabet mappings at:

          Today, this logic is suboptimally implemented in at least:

          Furthermore, the encoding is generic and has use cases outside of ASP.NET (e.g. you shouldn't have to reference ASP.NET to use it).

          Proposed API Change

          The existing Base64.EncodeToUtf8 and Base64.DecodeFromUtf8 should each add a new method overload that allows one of the following:

          1. A new enumeration of allowed alphabets (ex: Base64Alphabet) which internally maps to well-known alphabet spans

          2. Allow any custom alphabet to be supplied as ReadOnlySpan<sbyte> that must have an exact length of 64 for encoding and 256 for decoding

            a. One or more new types could be provided with static properties for the well-known alphabet spans

          Details

          Option 1 requires less validation, but has less flexibility. Option 2 has more flexibility, including scenarios not described here, but requires more validation before using the alphabet.

          Although the standard Base 64 and Base 64 URL are technically the same encoding with different alphabets, Base 64 URL typically does not include padding. RFC 4648 §3.2 indicates this is allowed (as it's explicitly stated), but there is no need to make that concession in this API. Including the = character for padding in Base 64 URL encoding is still correct.

          Both approaches would have a similar looking API:

          publicstaticunsafeOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,Base64Alphabetalphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticunsafeOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,Base64Alphabetalphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);

          Figure 1: Supply an alphabet enumeration

          publicstaticunsafeOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,ReadOnlySpan<byte>alphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticunsafeOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,ReadOnlySpan<byte>alphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);

          Figure 2: Supply a custom alphabet

          The support for trimming off and re-adding padding should be implemented separately. This could be in a separate Base64Url class or, perhaps more appropriately, added as a new type in System.Text.Encodings.Web.

          Padding is generally very cheap to deal with compared to other implementation methods. Trimming involves walking the tail end of the span while there are padding characters and then slicing it off. Re-padding fills the end of the span buffer before decoding. Neither operation requires additional allocations.

          Open Questions

          • Does Option 1 or Option 2 make more sense?
          • To complete the cycle, it feels like there should be a type that fully handles the encoding and decoding with trimmed padding. Does that make sense to be in System.Memory, System.Text.Encodings.Web, or perhaps some other place?

          Metadata

          Metadata

          Assignees

          Labels

          api-approvedAPI was approved in API review, it can be implementedarea-System.Netin-prThere is an active PR which will close this issue when it is merged

          Type

          No type

          Projects

          No projects

            Milestone

            Relationships

            None yet

            Development

            No branches or pull requests

            Issue actions

            , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
            Skip to content

            Support Base 64 URL #1658

            Description

            @commonsensesoftware

            Updated by @MihaZupan on 2024-02-27

            Proposed API

            namespaceSystem.Buffers.Text;publicstaticclassBase64Url{// Encode bytes => utf8publicstaticOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticOperationStatusEncodeToUtf8InPlace(Span<byte>buffer,intdataLength,outintbytesWritten);// Decode from utf8 => bytespublicstaticOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticOperationStatusDecodeFromUtf8InPlace(Span<byte>buffer,outintbytesWritten);// Max length APIspublicstaticintGetMaxDecodedFromUtf8Length(intlength);publicstaticintGetMaxEncodedToUtf8Length(intlength);// IsValidpublicstaticboolIsValid(ReadOnlySpan<char>base64UrlText);publicstaticboolIsValid(ReadOnlySpan<char>base64UrlText,outintdecodedLength);publicstaticboolIsValid(ReadOnlySpan<byte>base64UrlTextUtf8);publicstaticboolIsValid(ReadOnlySpan<byte>base64UrlTextUtf8,outintdecodedLength);// Up to this point, this is a mirror of System.Buffers.Text.Base64// Below are more helpers that bring over functionality similar to Convert.*Base64*// Encode to / decode from charspublicstaticboolTryEncodeToChars(ReadOnlySpan<byte>bytes,Span<char>chars,outintcharsWritten){}publicstaticboolTryDecodeFromChars(ReadOnlySpan<char>chars,Span<byte>bytes,outintbytesWritten){}// These are just accelerator methods.// Should be efficiently implementable on top of the other ones in just a few lines.// Encode to stringpublicstaticstringEncodeToString(ReadOnlySpan<char>chars,Encodingencoding){}publicstaticstringEncodeToString(ReadOnlySpan<byte>bytes){}// Decode from chars => string// Decode from chars => byte[]// The names could also just be "Decode" without naming the return typepublicstaticstringDecodeToString(ReadOnlySpan<char>chars,Encodingencoding){}publicstaticbyte[]DecodeToByteArray(ReadOnlySpan<char>chars){}}

            Original issue

            The Base64 implementation in System.Memory provides excellent low-level optimizations for RFC 4648, but it currently uses a fixed alphabet. Minor refactoring would add additional use cases using the existing implementation.

            Rationale and Usage

            The most obvious use case for this change is to support the encoding variant known as Base 64 URL also described in RFC 4648 §4.

            This excerpt from the RFC describes the difference between the standard Base 64 alphabet and the Base 64 URL alphabet.

            This encoding is technically identical to the previous one, except
            for the 62:nd and 63:rd alphabet character...

            I have already been able to verify the encoding and decoding in a copy of the current implementation by merely changing the 2 relevant characters in the alphabet mappings at:

            Today, this logic is suboptimally implemented in at least:

            Furthermore, the encoding is generic and has use cases outside of ASP.NET (e.g. you shouldn't have to reference ASP.NET to use it).

            Proposed API Change

            The existing Base64.EncodeToUtf8 and Base64.DecodeFromUtf8 should each add a new method overload that allows one of the following:

            1. A new enumeration of allowed alphabets (ex: Base64Alphabet) which internally maps to well-known alphabet spans

            2. Allow any custom alphabet to be supplied as ReadOnlySpan<sbyte> that must have an exact length of 64 for encoding and 256 for decoding

              a. One or more new types could be provided with static properties for the well-known alphabet spans

            Details

            Option 1 requires less validation, but has less flexibility. Option 2 has more flexibility, including scenarios not described here, but requires more validation before using the alphabet.

            Although the standard Base 64 and Base 64 URL are technically the same encoding with different alphabets, Base 64 URL typically does not include padding. RFC 4648 §3.2 indicates this is allowed (as it's explicitly stated), but there is no need to make that concession in this API. Including the = character for padding in Base 64 URL encoding is still correct.

            Both approaches would have a similar looking API:

            publicstaticunsafeOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,Base64Alphabetalphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticunsafeOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,Base64Alphabetalphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);

            Figure 1: Supply an alphabet enumeration

            publicstaticunsafeOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,ReadOnlySpan<byte>alphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticunsafeOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,ReadOnlySpan<byte>alphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);

            Figure 2: Supply a custom alphabet

            The support for trimming off and re-adding padding should be implemented separately. This could be in a separate Base64Url class or, perhaps more appropriately, added as a new type in System.Text.Encodings.Web.

            Padding is generally very cheap to deal with compared to other implementation methods. Trimming involves walking the tail end of the span while there are padding characters and then slicing it off. Re-padding fills the end of the span buffer before decoding. Neither operation requires additional allocations.

            Open Questions

            • Does Option 1 or Option 2 make more sense?
            • To complete the cycle, it feels like there should be a type that fully handles the encoding and decoding with trimmed padding. Does that make sense to be in System.Memory, System.Text.Encodings.Web, or perhaps some other place?

            Metadata

            Metadata

            Assignees

            Labels

            api-approvedAPI was approved in API review, it can be implementedarea-System.Netin-prThere is an active PR which will close this issue when it is merged

            Type

            No type

            Projects

            No projects

              Milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              Support Base 64 URL #1658

              Description

              @commonsensesoftware

              Updated by @MihaZupan on 2024-02-27

              Proposed API

              namespaceSystem.Buffers.Text;publicstaticclassBase64Url{// Encode bytes => utf8publicstaticOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticOperationStatusEncodeToUtf8InPlace(Span<byte>buffer,intdataLength,outintbytesWritten);// Decode from utf8 => bytespublicstaticOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticOperationStatusDecodeFromUtf8InPlace(Span<byte>buffer,outintbytesWritten);// Max length APIspublicstaticintGetMaxDecodedFromUtf8Length(intlength);publicstaticintGetMaxEncodedToUtf8Length(intlength);// IsValidpublicstaticboolIsValid(ReadOnlySpan<char>base64UrlText);publicstaticboolIsValid(ReadOnlySpan<char>base64UrlText,outintdecodedLength);publicstaticboolIsValid(ReadOnlySpan<byte>base64UrlTextUtf8);publicstaticboolIsValid(ReadOnlySpan<byte>base64UrlTextUtf8,outintdecodedLength);// Up to this point, this is a mirror of System.Buffers.Text.Base64// Below are more helpers that bring over functionality similar to Convert.*Base64*// Encode to / decode from charspublicstaticboolTryEncodeToChars(ReadOnlySpan<byte>bytes,Span<char>chars,outintcharsWritten){}publicstaticboolTryDecodeFromChars(ReadOnlySpan<char>chars,Span<byte>bytes,outintbytesWritten){}// These are just accelerator methods.// Should be efficiently implementable on top of the other ones in just a few lines.// Encode to stringpublicstaticstringEncodeToString(ReadOnlySpan<char>chars,Encodingencoding){}publicstaticstringEncodeToString(ReadOnlySpan<byte>bytes){}// Decode from chars => string// Decode from chars => byte[]// The names could also just be "Decode" without naming the return typepublicstaticstringDecodeToString(ReadOnlySpan<char>chars,Encodingencoding){}publicstaticbyte[]DecodeToByteArray(ReadOnlySpan<char>chars){}}

              Original issue

              The Base64 implementation in System.Memory provides excellent low-level optimizations for RFC 4648, but it currently uses a fixed alphabet. Minor refactoring would add additional use cases using the existing implementation.

              Rationale and Usage

              The most obvious use case for this change is to support the encoding variant known as Base 64 URL also described in RFC 4648 §4.

              This excerpt from the RFC describes the difference between the standard Base 64 alphabet and the Base 64 URL alphabet.

              This encoding is technically identical to the previous one, except
              for the 62:nd and 63:rd alphabet character...

              I have already been able to verify the encoding and decoding in a copy of the current implementation by merely changing the 2 relevant characters in the alphabet mappings at:

              Today, this logic is suboptimally implemented in at least:

              Furthermore, the encoding is generic and has use cases outside of ASP.NET (e.g. you shouldn't have to reference ASP.NET to use it).

              Proposed API Change

              The existing Base64.EncodeToUtf8 and Base64.DecodeFromUtf8 should each add a new method overload that allows one of the following:

              1. A new enumeration of allowed alphabets (ex: Base64Alphabet) which internally maps to well-known alphabet spans

              2. Allow any custom alphabet to be supplied as ReadOnlySpan<sbyte> that must have an exact length of 64 for encoding and 256 for decoding

                a. One or more new types could be provided with static properties for the well-known alphabet spans

              Details

              Option 1 requires less validation, but has less flexibility. Option 2 has more flexibility, including scenarios not described here, but requires more validation before using the alphabet.

              Although the standard Base 64 and Base 64 URL are technically the same encoding with different alphabets, Base 64 URL typically does not include padding. RFC 4648 §3.2 indicates this is allowed (as it's explicitly stated), but there is no need to make that concession in this API. Including the = character for padding in Base 64 URL encoding is still correct.

              Both approaches would have a similar looking API:

              publicstaticunsafeOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,Base64Alphabetalphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticunsafeOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,Base64Alphabetalphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);

              Figure 1: Supply an alphabet enumeration

              publicstaticunsafeOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,ReadOnlySpan<byte>alphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticunsafeOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,ReadOnlySpan<byte>alphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);

              Figure 2: Supply a custom alphabet

              The support for trimming off and re-adding padding should be implemented separately. This could be in a separate Base64Url class or, perhaps more appropriately, added as a new type in System.Text.Encodings.Web.

              Padding is generally very cheap to deal with compared to other implementation methods. Trimming involves walking the tail end of the span while there are padding characters and then slicing it off. Re-padding fills the end of the span buffer before decoding. Neither operation requires additional allocations.

              Open Questions

              • Does Option 1 or Option 2 make more sense?
              • To complete the cycle, it feels like there should be a type that fully handles the encoding and decoding with trimmed padding. Does that make sense to be in System.Memory, System.Text.Encodings.Web, or perhaps some other place?

              Metadata

              Metadata

              Assignees

              Labels

              api-approvedAPI was approved in API review, it can be implementedarea-System.Netin-prThere is an active PR which will close this issue when it is merged

              Type

              No type

              Projects

              No projects

                Milestone

                Relationships

                None yet

                Development

                No branches or pull requests

                Issue actions

                , 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                Skip to content

                Support Base 64 URL #1658

                Description

                @commonsensesoftware

                Updated by @MihaZupan on 2024-02-27

                Proposed API

                namespaceSystem.Buffers.Text;publicstaticclassBase64Url{// Encode bytes => utf8publicstaticOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticOperationStatusEncodeToUtf8InPlace(Span<byte>buffer,intdataLength,outintbytesWritten);// Decode from utf8 => bytespublicstaticOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticOperationStatusDecodeFromUtf8InPlace(Span<byte>buffer,outintbytesWritten);// Max length APIspublicstaticintGetMaxDecodedFromUtf8Length(intlength);publicstaticintGetMaxEncodedToUtf8Length(intlength);// IsValidpublicstaticboolIsValid(ReadOnlySpan<char>base64UrlText);publicstaticboolIsValid(ReadOnlySpan<char>base64UrlText,outintdecodedLength);publicstaticboolIsValid(ReadOnlySpan<byte>base64UrlTextUtf8);publicstaticboolIsValid(ReadOnlySpan<byte>base64UrlTextUtf8,outintdecodedLength);// Up to this point, this is a mirror of System.Buffers.Text.Base64// Below are more helpers that bring over functionality similar to Convert.*Base64*// Encode to / decode from charspublicstaticboolTryEncodeToChars(ReadOnlySpan<byte>bytes,Span<char>chars,outintcharsWritten){}publicstaticboolTryDecodeFromChars(ReadOnlySpan<char>chars,Span<byte>bytes,outintbytesWritten){}// These are just accelerator methods.// Should be efficiently implementable on top of the other ones in just a few lines.// Encode to stringpublicstaticstringEncodeToString(ReadOnlySpan<char>chars,Encodingencoding){}publicstaticstringEncodeToString(ReadOnlySpan<byte>bytes){}// Decode from chars => string// Decode from chars => byte[]// The names could also just be "Decode" without naming the return typepublicstaticstringDecodeToString(ReadOnlySpan<char>chars,Encodingencoding){}publicstaticbyte[]DecodeToByteArray(ReadOnlySpan<char>chars){}}

                Original issue

                The Base64 implementation in System.Memory provides excellent low-level optimizations for RFC 4648, but it currently uses a fixed alphabet. Minor refactoring would add additional use cases using the existing implementation.

                Rationale and Usage

                The most obvious use case for this change is to support the encoding variant known as Base 64 URL also described in RFC 4648 §4.

                This excerpt from the RFC describes the difference between the standard Base 64 alphabet and the Base 64 URL alphabet.

                This encoding is technically identical to the previous one, except
                for the 62:nd and 63:rd alphabet character...

                I have already been able to verify the encoding and decoding in a copy of the current implementation by merely changing the 2 relevant characters in the alphabet mappings at:

                Today, this logic is suboptimally implemented in at least:

                Furthermore, the encoding is generic and has use cases outside of ASP.NET (e.g. you shouldn't have to reference ASP.NET to use it).

                Proposed API Change

                The existing Base64.EncodeToUtf8 and Base64.DecodeFromUtf8 should each add a new method overload that allows one of the following:

                1. A new enumeration of allowed alphabets (ex: Base64Alphabet) which internally maps to well-known alphabet spans

                2. Allow any custom alphabet to be supplied as ReadOnlySpan<sbyte> that must have an exact length of 64 for encoding and 256 for decoding

                  a. One or more new types could be provided with static properties for the well-known alphabet spans

                Details

                Option 1 requires less validation, but has less flexibility. Option 2 has more flexibility, including scenarios not described here, but requires more validation before using the alphabet.

                Although the standard Base 64 and Base 64 URL are technically the same encoding with different alphabets, Base 64 URL typically does not include padding. RFC 4648 §3.2 indicates this is allowed (as it's explicitly stated), but there is no need to make that concession in this API. Including the = character for padding in Base 64 URL encoding is still correct.

                Both approaches would have a similar looking API:

                publicstaticunsafeOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,Base64Alphabetalphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticunsafeOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,Base64Alphabetalphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);

                Figure 1: Supply an alphabet enumeration

                publicstaticunsafeOperationStatusEncodeToUtf8(ReadOnlySpan<byte>bytes,Span<byte>utf8,ReadOnlySpan<byte>alphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);publicstaticunsafeOperationStatusDecodeFromUtf8(ReadOnlySpan<byte>utf8,Span<byte>bytes,ReadOnlySpan<byte>alphabet,outintbytesConsumed,outintbytesWritten,boolisFinalBlock=true);

                Figure 2: Supply a custom alphabet

                The support for trimming off and re-adding padding should be implemented separately. This could be in a separate Base64Url class or, perhaps more appropriately, added as a new type in System.Text.Encodings.Web.

                Padding is generally very cheap to deal with compared to other implementation methods. Trimming involves walking the tail end of the span while there are padding characters and then slicing it off. Re-padding fills the end of the span buffer before decoding. Neither operation requires additional allocations.

                Open Questions

                • Does Option 1 or Option 2 make more sense?
                • To complete the cycle, it feels like there should be a type that fully handles the encoding and decoding with trimmed padding. Does that make sense to be in System.Memory, System.Text.Encodings.Web, or perhaps some other place?

                Metadata

                Metadata

                Assignees

                Labels

                api-approvedAPI was approved in API review, it can be implementedarea-System.Netin-prThere is an active PR which will close this issue when it is merged

                Type

                No type

                Projects

                No projects

                  Milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions