Skip to content

Decode JWS payloads as UTF-8 instead of the JVM default charset - #266

Merged
alexanderjordanbaker merged 1 commit into
apple:mainfrom
matteobaccan:fix/utf8-jws-payload-decoding
Aug 27, 2026
Merged

Decode JWS payloads as UTF-8 instead of the JVM default charset#266
alexanderjordanbaker merged 1 commit into
apple:mainfrom
matteobaccan:fix/utf8-jws-payload-decoding

Conversation

@matteobaccan

Copy link
Copy Markdown

The problem

SignedDataVerifier.parseJWTPayload converts the base64url-decoded payload bytes into a String without specifying a charset:

String payload = new String(Base64.getUrlDecoder().decode(jwt.getPayload()));

new String(byte[]) uses the JVM default charset. A JWT Claims Set is always a UTF-8 encoded JSON object (RFC 7519 §3, RFC 8259 §8.1), so the two only agree when the JVM happens to run with UTF-8 as its default.

That is not guaranteed on the Java versions this library supports. JEP 400 made UTF-8 the default only in Java 18; on Java 11 through 17 the default charset comes from the platform locale, so a server running on Windows or with a non-UTF-8 LANG gets windows-1252, ISO-8859-1, Shift_JIS, and so on.

When that happens, every non-ASCII character in a decoded payload is corrupted silently. The signature is verified against the raw JWS, so verification still succeeds and no exception is raised — the caller simply receives mojibake. This affects every entry point that goes through decodeSignedObject: verifyAndDecodeTransaction, verifyAndDecodeRenewalInfo, verifyAndDecodeNotification, verifyAndDecodeAppTransaction and verifyAndDecodeRealtimeRequest.

It is easy to reach in practice. The Advanced Commerce descriptors and items carry merchant-supplied free text (displayName, description), so any app that is not English-only can hit it.

Reproduction

On JDK 17 with windows-1252 as the default charset, decoding a signed transaction whose advancedCommerceInfo.descriptors.description is Abonnement Café — 5,99 € par mois returns:

Abonnement Café — 5,99 € par mois

The existing test suite cannot catch this: all of its payload fixtures are ASCII-only, and CI sets JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8, which hides the difference.

The fix

new String(bytes, StandardCharsets.UTF_8) — one argument.

Source encoding

The build does not set options.encoding, so javac reads the sources with the platform charset too. 25 files under src/main contain non-ASCII characters (typographic apostrophes in the javadoc), which means those are corrupted in locally built javadoc on a non-UTF-8 machine. It also makes a regression test for this bug impossible to write in the natural way: the expected string literal would be corrupted by the compiler in exactly the same way as the actual value, and the assertion would pass.

This PR sets options.encoding = 'UTF-8' on the JavaCompile tasks and on javadoc. Note that ci-prb.yml currently compensates for the missing setting with JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8, while ci-release-javadocs.yml does not set it at all.

The test

SignedDataVerifierTest.testNonAsciiDataDecodingIsIndependentOfTheDefaultCharset decodes a new fixture containing French and Japanese text. Its expected values are written as \uXXXX escape sequences, so the test source is pure ASCII and the assertion does not depend on the encoding used to compile it.

Verification

On JDK 17, default charset windows-1252:

  • without the production change the new test fails with
    expected: <Abonnement Café — 5,99 € par mois> but was: <Abonnement Café — 5,99 € par mois>
  • with it, ./gradlew clean test and ./gradlew javadoc both pass

Not included

Two related items I left out to keep the diff focused, happy to add them if you would prefer them here:

  • the JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8 line in ci-prb.yml is now redundant for compilation and could be dropped
  • CI runs on a UTF-8 locale, so it would not catch a regression of this bug; a matrix entry running the tests with a non-UTF-8 default charset would lock the behavior in

I also did not touch CHANGELOG.md, since it appears to be updated by maintainers in the release commits.

The payload of a JWS is always UTF-8, but parseJWTPayload converted the
decoded bytes with new String(byte[]), which uses the JVM default charset.
On the Java versions this library supports that charset is platform
dependent, since JEP 400 only made UTF-8 the default in Java 18. On a JVM
running with, for example, windows-1252 every non-ASCII character in a
signed transaction, renewal info, notification or AppTransaction was
silently corrupted: the signature still verified, so no error was raised
and the caller received mojibake.
Free-text fields such as the Advanced Commerce descriptors and the item
displayName and description make this reachable for any non-English app.
Also set the source encoding of the compile and javadoc tasks to UTF-8.
Without it javac reads the sources with the platform charset, which
corrupts the non-ASCII characters present in the javadoc of 25 files and
prevents the new regression test from expressing its expected values.

@alexanderjordanbakeralexanderjordanbaker left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thank you! Will be filing equivalent PRs on the other languages shortly

@alexanderjordanbaker
alexanderjordanbaker merged commit f0ddedd into apple:mainAug 27, 2026
4 checks passed
alexanderjordanbaker added a commit to apple/app-store-server-library-node that referenced this pull request Aug 28, 2026
alexanderjordanbaker added a commit to apple/app-store-server-library-python that referenced this pull request Aug 28, 2026
alexanderjordanbaker added a commit to apple/app-store-server-library-swift that referenced this pull request Aug 28, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@matteobaccan@alexanderjordanbaker
, 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
 blocks
(function() {
function addCopyButtons() {
document.querySelectorAll('pre code').forEach(function(codeBlock) {
if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
codeBlock.parentElement.setAttribute('data-copy-added', 'true');
var btn = document.createElement('button');
btn.textContent = 'Copy';
btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
btn.onmouseover = function() { this.style.opacity = '1'; };
btn.onmouseout = function() { this.style.opacity = '0.7'; };
btn.onclick = function() {
navigator.clipboard.writeText(codeBlock.textContent).then(function() {
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
});
};
codeBlock.parentElement.style.position = 'relative';
codeBlock.parentElement.appendChild(btn);
});
}
addCopyButtons();
// Re-run on dynamic content
var observer = new MutationObserver(addCopyButtons);
observer.observe(document.body, { childList: true, subtree: true });
})();
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Decode JWS payloads as UTF-8 instead of the JVM default charset by matteobaccan · Pull Request #266 · apple/app-store-server-library-java · GitHub
Skip to content

Decode JWS payloads as UTF-8 instead of the JVM default charset - #266

Merged
alexanderjordanbaker merged 1 commit into
apple:mainfrom
matteobaccan:fix/utf8-jws-payload-decoding
Aug 27, 2026
Merged

Decode JWS payloads as UTF-8 instead of the JVM default charset#266
alexanderjordanbaker merged 1 commit into
apple:mainfrom
matteobaccan:fix/utf8-jws-payload-decoding

Conversation

@matteobaccan

Copy link
Copy Markdown

The problem

SignedDataVerifier.parseJWTPayload converts the base64url-decoded payload bytes into a String without specifying a charset:

String payload = new String(Base64.getUrlDecoder().decode(jwt.getPayload()));

new String(byte[]) uses the JVM default charset. A JWT Claims Set is always a UTF-8 encoded JSON object (RFC 7519 §3, RFC 8259 §8.1), so the two only agree when the JVM happens to run with UTF-8 as its default.

That is not guaranteed on the Java versions this library supports. JEP 400 made UTF-8 the default only in Java 18; on Java 11 through 17 the default charset comes from the platform locale, so a server running on Windows or with a non-UTF-8 LANG gets windows-1252, ISO-8859-1, Shift_JIS, and so on.

When that happens, every non-ASCII character in a decoded payload is corrupted silently. The signature is verified against the raw JWS, so verification still succeeds and no exception is raised — the caller simply receives mojibake. This affects every entry point that goes through decodeSignedObject: verifyAndDecodeTransaction, verifyAndDecodeRenewalInfo, verifyAndDecodeNotification, verifyAndDecodeAppTransaction and verifyAndDecodeRealtimeRequest.

It is easy to reach in practice. The Advanced Commerce descriptors and items carry merchant-supplied free text (displayName, description), so any app that is not English-only can hit it.

Reproduction

On JDK 17 with windows-1252 as the default charset, decoding a signed transaction whose advancedCommerceInfo.descriptors.description is Abonnement Café — 5,99 € par mois returns:

Abonnement Café — 5,99 € par mois

The existing test suite cannot catch this: all of its payload fixtures are ASCII-only, and CI sets JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8, which hides the difference.

The fix

new String(bytes, StandardCharsets.UTF_8) — one argument.

Source encoding

The build does not set options.encoding, so javac reads the sources with the platform charset too. 25 files under src/main contain non-ASCII characters (typographic apostrophes in the javadoc), which means those are corrupted in locally built javadoc on a non-UTF-8 machine. It also makes a regression test for this bug impossible to write in the natural way: the expected string literal would be corrupted by the compiler in exactly the same way as the actual value, and the assertion would pass.

This PR sets options.encoding = 'UTF-8' on the JavaCompile tasks and on javadoc. Note that ci-prb.yml currently compensates for the missing setting with JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8, while ci-release-javadocs.yml does not set it at all.

The test

SignedDataVerifierTest.testNonAsciiDataDecodingIsIndependentOfTheDefaultCharset decodes a new fixture containing French and Japanese text. Its expected values are written as \uXXXX escape sequences, so the test source is pure ASCII and the assertion does not depend on the encoding used to compile it.

Verification

On JDK 17, default charset windows-1252:

  • without the production change the new test fails with
    expected: <Abonnement Café — 5,99 € par mois> but was: <Abonnement Café — 5,99 € par mois>
  • with it, ./gradlew clean test and ./gradlew javadoc both pass

Not included

Two related items I left out to keep the diff focused, happy to add them if you would prefer them here:

  • the JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8 line in ci-prb.yml is now redundant for compilation and could be dropped
  • CI runs on a UTF-8 locale, so it would not catch a regression of this bug; a matrix entry running the tests with a non-UTF-8 default charset would lock the behavior in

I also did not touch CHANGELOG.md, since it appears to be updated by maintainers in the release commits.

The payload of a JWS is always UTF-8, but parseJWTPayload converted the
decoded bytes with new String(byte[]), which uses the JVM default charset.
On the Java versions this library supports that charset is platform
dependent, since JEP 400 only made UTF-8 the default in Java 18. On a JVM
running with, for example, windows-1252 every non-ASCII character in a
signed transaction, renewal info, notification or AppTransaction was
silently corrupted: the signature still verified, so no error was raised
and the caller received mojibake.
Free-text fields such as the Advanced Commerce descriptors and the item
displayName and description make this reachable for any non-English app.
Also set the source encoding of the compile and javadoc tasks to UTF-8.
Without it javac reads the sources with the platform charset, which
corrupts the non-ASCII characters present in the javadoc of 25 files and
prevents the new regression test from expressing its expected values.

@alexanderjordanbakeralexanderjordanbaker left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thank you! Will be filing equivalent PRs on the other languages shortly

@alexanderjordanbaker
alexanderjordanbaker merged commit f0ddedd into apple:mainAug 27, 2026
4 checks passed
alexanderjordanbaker added a commit to apple/app-store-server-library-node that referenced this pull request Aug 28, 2026
alexanderjordanbaker added a commit to apple/app-store-server-library-python that referenced this pull request Aug 28, 2026
alexanderjordanbaker added a commit to apple/app-store-server-library-swift that referenced this pull request Aug 28, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@matteobaccan@alexanderjordanbaker
, 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Decode JWS payloads as UTF-8 instead of the JVM default charset by matteobaccan · Pull Request #266 · apple/app-store-server-library-java · GitHub
Skip to content

Decode JWS payloads as UTF-8 instead of the JVM default charset - #266

Merged
alexanderjordanbaker merged 1 commit into
apple:mainfrom
matteobaccan:fix/utf8-jws-payload-decoding
Aug 27, 2026
Merged

Decode JWS payloads as UTF-8 instead of the JVM default charset#266
alexanderjordanbaker merged 1 commit into
apple:mainfrom
matteobaccan:fix/utf8-jws-payload-decoding

Conversation

@matteobaccan

Copy link
Copy Markdown

The problem

SignedDataVerifier.parseJWTPayload converts the base64url-decoded payload bytes into a String without specifying a charset:

String payload = new String(Base64.getUrlDecoder().decode(jwt.getPayload()));

new String(byte[]) uses the JVM default charset. A JWT Claims Set is always a UTF-8 encoded JSON object (RFC 7519 §3, RFC 8259 §8.1), so the two only agree when the JVM happens to run with UTF-8 as its default.

That is not guaranteed on the Java versions this library supports. JEP 400 made UTF-8 the default only in Java 18; on Java 11 through 17 the default charset comes from the platform locale, so a server running on Windows or with a non-UTF-8 LANG gets windows-1252, ISO-8859-1, Shift_JIS, and so on.

When that happens, every non-ASCII character in a decoded payload is corrupted silently. The signature is verified against the raw JWS, so verification still succeeds and no exception is raised — the caller simply receives mojibake. This affects every entry point that goes through decodeSignedObject: verifyAndDecodeTransaction, verifyAndDecodeRenewalInfo, verifyAndDecodeNotification, verifyAndDecodeAppTransaction and verifyAndDecodeRealtimeRequest.

It is easy to reach in practice. The Advanced Commerce descriptors and items carry merchant-supplied free text (displayName, description), so any app that is not English-only can hit it.

Reproduction

On JDK 17 with windows-1252 as the default charset, decoding a signed transaction whose advancedCommerceInfo.descriptors.description is Abonnement Café — 5,99 € par mois returns:

Abonnement Café — 5,99 € par mois

The existing test suite cannot catch this: all of its payload fixtures are ASCII-only, and CI sets JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8, which hides the difference.

The fix

new String(bytes, StandardCharsets.UTF_8) — one argument.

Source encoding

The build does not set options.encoding, so javac reads the sources with the platform charset too. 25 files under src/main contain non-ASCII characters (typographic apostrophes in the javadoc), which means those are corrupted in locally built javadoc on a non-UTF-8 machine. It also makes a regression test for this bug impossible to write in the natural way: the expected string literal would be corrupted by the compiler in exactly the same way as the actual value, and the assertion would pass.

This PR sets options.encoding = 'UTF-8' on the JavaCompile tasks and on javadoc. Note that ci-prb.yml currently compensates for the missing setting with JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8, while ci-release-javadocs.yml does not set it at all.

The test

SignedDataVerifierTest.testNonAsciiDataDecodingIsIndependentOfTheDefaultCharset decodes a new fixture containing French and Japanese text. Its expected values are written as \uXXXX escape sequences, so the test source is pure ASCII and the assertion does not depend on the encoding used to compile it.

Verification

On JDK 17, default charset windows-1252:

  • without the production change the new test fails with
    expected: <Abonnement Café — 5,99 € par mois> but was: <Abonnement Café — 5,99 € par mois>
  • with it, ./gradlew clean test and ./gradlew javadoc both pass

Not included

Two related items I left out to keep the diff focused, happy to add them if you would prefer them here:

  • the JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8 line in ci-prb.yml is now redundant for compilation and could be dropped
  • CI runs on a UTF-8 locale, so it would not catch a regression of this bug; a matrix entry running the tests with a non-UTF-8 default charset would lock the behavior in

I also did not touch CHANGELOG.md, since it appears to be updated by maintainers in the release commits.

The payload of a JWS is always UTF-8, but parseJWTPayload converted the
decoded bytes with new String(byte[]), which uses the JVM default charset.
On the Java versions this library supports that charset is platform
dependent, since JEP 400 only made UTF-8 the default in Java 18. On a JVM
running with, for example, windows-1252 every non-ASCII character in a
signed transaction, renewal info, notification or AppTransaction was
silently corrupted: the signature still verified, so no error was raised
and the caller received mojibake.
Free-text fields such as the Advanced Commerce descriptors and the item
displayName and description make this reachable for any non-English app.
Also set the source encoding of the compile and javadoc tasks to UTF-8.
Without it javac reads the sources with the platform charset, which
corrupts the non-ASCII characters present in the javadoc of 25 files and
prevents the new regression test from expressing its expected values.

@alexanderjordanbakeralexanderjordanbaker left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thank you! Will be filing equivalent PRs on the other languages shortly

@alexanderjordanbaker
alexanderjordanbaker merged commit f0ddedd into apple:mainAug 27, 2026
4 checks passed
alexanderjordanbaker added a commit to apple/app-store-server-library-node that referenced this pull request Aug 28, 2026
alexanderjordanbaker added a commit to apple/app-store-server-library-python that referenced this pull request Aug 28, 2026
alexanderjordanbaker added a commit to apple/app-store-server-library-swift that referenced this pull request Aug 28, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@matteobaccan@alexanderjordanbaker
, 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Decode JWS payloads as UTF-8 instead of the JVM default charset by matteobaccan · Pull Request #266 · apple/app-store-server-library-java · GitHub
Skip to content

Decode JWS payloads as UTF-8 instead of the JVM default charset - #266

Merged
alexanderjordanbaker merged 1 commit into
apple:mainfrom
matteobaccan:fix/utf8-jws-payload-decoding
Aug 27, 2026
Merged

Decode JWS payloads as UTF-8 instead of the JVM default charset#266
alexanderjordanbaker merged 1 commit into
apple:mainfrom
matteobaccan:fix/utf8-jws-payload-decoding

Conversation

@matteobaccan

Copy link
Copy Markdown

The problem

SignedDataVerifier.parseJWTPayload converts the base64url-decoded payload bytes into a String without specifying a charset:

String payload = new String(Base64.getUrlDecoder().decode(jwt.getPayload()));

new String(byte[]) uses the JVM default charset. A JWT Claims Set is always a UTF-8 encoded JSON object (RFC 7519 §3, RFC 8259 §8.1), so the two only agree when the JVM happens to run with UTF-8 as its default.

That is not guaranteed on the Java versions this library supports. JEP 400 made UTF-8 the default only in Java 18; on Java 11 through 17 the default charset comes from the platform locale, so a server running on Windows or with a non-UTF-8 LANG gets windows-1252, ISO-8859-1, Shift_JIS, and so on.

When that happens, every non-ASCII character in a decoded payload is corrupted silently. The signature is verified against the raw JWS, so verification still succeeds and no exception is raised — the caller simply receives mojibake. This affects every entry point that goes through decodeSignedObject: verifyAndDecodeTransaction, verifyAndDecodeRenewalInfo, verifyAndDecodeNotification, verifyAndDecodeAppTransaction and verifyAndDecodeRealtimeRequest.

It is easy to reach in practice. The Advanced Commerce descriptors and items carry merchant-supplied free text (displayName, description), so any app that is not English-only can hit it.

Reproduction

On JDK 17 with windows-1252 as the default charset, decoding a signed transaction whose advancedCommerceInfo.descriptors.description is Abonnement Café — 5,99 € par mois returns:

Abonnement Café — 5,99 € par mois

The existing test suite cannot catch this: all of its payload fixtures are ASCII-only, and CI sets JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8, which hides the difference.

The fix

new String(bytes, StandardCharsets.UTF_8) — one argument.

Source encoding

The build does not set options.encoding, so javac reads the sources with the platform charset too. 25 files under src/main contain non-ASCII characters (typographic apostrophes in the javadoc), which means those are corrupted in locally built javadoc on a non-UTF-8 machine. It also makes a regression test for this bug impossible to write in the natural way: the expected string literal would be corrupted by the compiler in exactly the same way as the actual value, and the assertion would pass.

This PR sets options.encoding = 'UTF-8' on the JavaCompile tasks and on javadoc. Note that ci-prb.yml currently compensates for the missing setting with JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8, while ci-release-javadocs.yml does not set it at all.

The test

SignedDataVerifierTest.testNonAsciiDataDecodingIsIndependentOfTheDefaultCharset decodes a new fixture containing French and Japanese text. Its expected values are written as \uXXXX escape sequences, so the test source is pure ASCII and the assertion does not depend on the encoding used to compile it.

Verification

On JDK 17, default charset windows-1252:

  • without the production change the new test fails with
    expected: <Abonnement Café — 5,99 € par mois> but was: <Abonnement Café — 5,99 € par mois>
  • with it, ./gradlew clean test and ./gradlew javadoc both pass

Not included

Two related items I left out to keep the diff focused, happy to add them if you would prefer them here:

  • the JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8 line in ci-prb.yml is now redundant for compilation and could be dropped
  • CI runs on a UTF-8 locale, so it would not catch a regression of this bug; a matrix entry running the tests with a non-UTF-8 default charset would lock the behavior in

I also did not touch CHANGELOG.md, since it appears to be updated by maintainers in the release commits.

The payload of a JWS is always UTF-8, but parseJWTPayload converted the
decoded bytes with new String(byte[]), which uses the JVM default charset.
On the Java versions this library supports that charset is platform
dependent, since JEP 400 only made UTF-8 the default in Java 18. On a JVM
running with, for example, windows-1252 every non-ASCII character in a
signed transaction, renewal info, notification or AppTransaction was
silently corrupted: the signature still verified, so no error was raised
and the caller received mojibake.
Free-text fields such as the Advanced Commerce descriptors and the item
displayName and description make this reachable for any non-English app.
Also set the source encoding of the compile and javadoc tasks to UTF-8.
Without it javac reads the sources with the platform charset, which
corrupts the non-ASCII characters present in the javadoc of 25 files and
prevents the new regression test from expressing its expected values.

@alexanderjordanbakeralexanderjordanbaker left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thank you! Will be filing equivalent PRs on the other languages shortly

@alexanderjordanbaker
alexanderjordanbaker merged commit f0ddedd into apple:mainAug 27, 2026
4 checks passed
alexanderjordanbaker added a commit to apple/app-store-server-library-node that referenced this pull request Aug 28, 2026
alexanderjordanbaker added a commit to apple/app-store-server-library-python that referenced this pull request Aug 28, 2026
alexanderjordanbaker added a commit to apple/app-store-server-library-swift that referenced this pull request Aug 28, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@matteobaccan@alexanderjordanbaker
, 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' Decode JWS payloads as UTF-8 instead of the JVM default charset by matteobaccan · Pull Request #266 · apple/app-store-server-library-java · GitHub
Skip to content

Decode JWS payloads as UTF-8 instead of the JVM default charset - #266

Merged
alexanderjordanbaker merged 1 commit into
apple:mainfrom
matteobaccan:fix/utf8-jws-payload-decoding
Aug 27, 2026
Merged

Decode JWS payloads as UTF-8 instead of the JVM default charset#266
alexanderjordanbaker merged 1 commit into
apple:mainfrom
matteobaccan:fix/utf8-jws-payload-decoding

Conversation

@matteobaccan

Copy link
Copy Markdown

The problem

SignedDataVerifier.parseJWTPayload converts the base64url-decoded payload bytes into a String without specifying a charset:

String payload = new String(Base64.getUrlDecoder().decode(jwt.getPayload()));

new String(byte[]) uses the JVM default charset. A JWT Claims Set is always a UTF-8 encoded JSON object (RFC 7519 §3, RFC 8259 §8.1), so the two only agree when the JVM happens to run with UTF-8 as its default.

That is not guaranteed on the Java versions this library supports. JEP 400 made UTF-8 the default only in Java 18; on Java 11 through 17 the default charset comes from the platform locale, so a server running on Windows or with a non-UTF-8 LANG gets windows-1252, ISO-8859-1, Shift_JIS, and so on.

When that happens, every non-ASCII character in a decoded payload is corrupted silently. The signature is verified against the raw JWS, so verification still succeeds and no exception is raised — the caller simply receives mojibake. This affects every entry point that goes through decodeSignedObject: verifyAndDecodeTransaction, verifyAndDecodeRenewalInfo, verifyAndDecodeNotification, verifyAndDecodeAppTransaction and verifyAndDecodeRealtimeRequest.

It is easy to reach in practice. The Advanced Commerce descriptors and items carry merchant-supplied free text (displayName, description), so any app that is not English-only can hit it.

Reproduction

On JDK 17 with windows-1252 as the default charset, decoding a signed transaction whose advancedCommerceInfo.descriptors.description is Abonnement Café — 5,99 € par mois returns:

Abonnement Café — 5,99 € par mois

The existing test suite cannot catch this: all of its payload fixtures are ASCII-only, and CI sets JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8, which hides the difference.

The fix

new String(bytes, StandardCharsets.UTF_8) — one argument.

Source encoding

The build does not set options.encoding, so javac reads the sources with the platform charset too. 25 files under src/main contain non-ASCII characters (typographic apostrophes in the javadoc), which means those are corrupted in locally built javadoc on a non-UTF-8 machine. It also makes a regression test for this bug impossible to write in the natural way: the expected string literal would be corrupted by the compiler in exactly the same way as the actual value, and the assertion would pass.

This PR sets options.encoding = 'UTF-8' on the JavaCompile tasks and on javadoc. Note that ci-prb.yml currently compensates for the missing setting with JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8, while ci-release-javadocs.yml does not set it at all.

The test

SignedDataVerifierTest.testNonAsciiDataDecodingIsIndependentOfTheDefaultCharset decodes a new fixture containing French and Japanese text. Its expected values are written as \uXXXX escape sequences, so the test source is pure ASCII and the assertion does not depend on the encoding used to compile it.

Verification

On JDK 17, default charset windows-1252:

  • without the production change the new test fails with
    expected: <Abonnement Café — 5,99 € par mois> but was: <Abonnement Café — 5,99 € par mois>
  • with it, ./gradlew clean test and ./gradlew javadoc both pass

Not included

Two related items I left out to keep the diff focused, happy to add them if you would prefer them here:

  • the JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8 line in ci-prb.yml is now redundant for compilation and could be dropped
  • CI runs on a UTF-8 locale, so it would not catch a regression of this bug; a matrix entry running the tests with a non-UTF-8 default charset would lock the behavior in

I also did not touch CHANGELOG.md, since it appears to be updated by maintainers in the release commits.

The payload of a JWS is always UTF-8, but parseJWTPayload converted the
decoded bytes with new String(byte[]), which uses the JVM default charset.
On the Java versions this library supports that charset is platform
dependent, since JEP 400 only made UTF-8 the default in Java 18. On a JVM
running with, for example, windows-1252 every non-ASCII character in a
signed transaction, renewal info, notification or AppTransaction was
silently corrupted: the signature still verified, so no error was raised
and the caller received mojibake.
Free-text fields such as the Advanced Commerce descriptors and the item
displayName and description make this reachable for any non-English app.
Also set the source encoding of the compile and javadoc tasks to UTF-8.
Without it javac reads the sources with the platform charset, which
corrupts the non-ASCII characters present in the javadoc of 25 files and
prevents the new regression test from expressing its expected values.

@alexanderjordanbakeralexanderjordanbaker left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thank you! Will be filing equivalent PRs on the other languages shortly

@alexanderjordanbaker
alexanderjordanbaker merged commit f0ddedd into apple:mainAug 27, 2026
4 checks passed
alexanderjordanbaker added a commit to apple/app-store-server-library-node that referenced this pull request Aug 28, 2026
alexanderjordanbaker added a commit to apple/app-store-server-library-python that referenced this pull request Aug 28, 2026
alexanderjordanbaker added a commit to apple/app-store-server-library-swift that referenced this pull request Aug 28, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@matteobaccan@alexanderjordanbaker
, 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Decode JWS payloads as UTF-8 instead of the JVM default charset by matteobaccan · Pull Request #266 · apple/app-store-server-library-java · GitHub
Skip to content

Decode JWS payloads as UTF-8 instead of the JVM default charset - #266

Merged
alexanderjordanbaker merged 1 commit into
apple:mainfrom
matteobaccan:fix/utf8-jws-payload-decoding
Aug 27, 2026
Merged

Decode JWS payloads as UTF-8 instead of the JVM default charset#266
alexanderjordanbaker merged 1 commit into
apple:mainfrom
matteobaccan:fix/utf8-jws-payload-decoding

Conversation

@matteobaccan

Copy link
Copy Markdown

The problem

SignedDataVerifier.parseJWTPayload converts the base64url-decoded payload bytes into a String without specifying a charset:

String payload = new String(Base64.getUrlDecoder().decode(jwt.getPayload()));

new String(byte[]) uses the JVM default charset. A JWT Claims Set is always a UTF-8 encoded JSON object (RFC 7519 §3, RFC 8259 §8.1), so the two only agree when the JVM happens to run with UTF-8 as its default.

That is not guaranteed on the Java versions this library supports. JEP 400 made UTF-8 the default only in Java 18; on Java 11 through 17 the default charset comes from the platform locale, so a server running on Windows or with a non-UTF-8 LANG gets windows-1252, ISO-8859-1, Shift_JIS, and so on.

When that happens, every non-ASCII character in a decoded payload is corrupted silently. The signature is verified against the raw JWS, so verification still succeeds and no exception is raised — the caller simply receives mojibake. This affects every entry point that goes through decodeSignedObject: verifyAndDecodeTransaction, verifyAndDecodeRenewalInfo, verifyAndDecodeNotification, verifyAndDecodeAppTransaction and verifyAndDecodeRealtimeRequest.

It is easy to reach in practice. The Advanced Commerce descriptors and items carry merchant-supplied free text (displayName, description), so any app that is not English-only can hit it.

Reproduction

On JDK 17 with windows-1252 as the default charset, decoding a signed transaction whose advancedCommerceInfo.descriptors.description is Abonnement Café — 5,99 € par mois returns:

Abonnement Café — 5,99 € par mois

The existing test suite cannot catch this: all of its payload fixtures are ASCII-only, and CI sets JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8, which hides the difference.

The fix

new String(bytes, StandardCharsets.UTF_8) — one argument.

Source encoding

The build does not set options.encoding, so javac reads the sources with the platform charset too. 25 files under src/main contain non-ASCII characters (typographic apostrophes in the javadoc), which means those are corrupted in locally built javadoc on a non-UTF-8 machine. It also makes a regression test for this bug impossible to write in the natural way: the expected string literal would be corrupted by the compiler in exactly the same way as the actual value, and the assertion would pass.

This PR sets options.encoding = 'UTF-8' on the JavaCompile tasks and on javadoc. Note that ci-prb.yml currently compensates for the missing setting with JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8, while ci-release-javadocs.yml does not set it at all.

The test

SignedDataVerifierTest.testNonAsciiDataDecodingIsIndependentOfTheDefaultCharset decodes a new fixture containing French and Japanese text. Its expected values are written as \uXXXX escape sequences, so the test source is pure ASCII and the assertion does not depend on the encoding used to compile it.

Verification

On JDK 17, default charset windows-1252:

  • without the production change the new test fails with
    expected: <Abonnement Café — 5,99 € par mois> but was: <Abonnement Café — 5,99 € par mois>
  • with it, ./gradlew clean test and ./gradlew javadoc both pass

Not included

Two related items I left out to keep the diff focused, happy to add them if you would prefer them here:

  • the JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8 line in ci-prb.yml is now redundant for compilation and could be dropped
  • CI runs on a UTF-8 locale, so it would not catch a regression of this bug; a matrix entry running the tests with a non-UTF-8 default charset would lock the behavior in

I also did not touch CHANGELOG.md, since it appears to be updated by maintainers in the release commits.

The payload of a JWS is always UTF-8, but parseJWTPayload converted the
decoded bytes with new String(byte[]), which uses the JVM default charset.
On the Java versions this library supports that charset is platform
dependent, since JEP 400 only made UTF-8 the default in Java 18. On a JVM
running with, for example, windows-1252 every non-ASCII character in a
signed transaction, renewal info, notification or AppTransaction was
silently corrupted: the signature still verified, so no error was raised
and the caller received mojibake.
Free-text fields such as the Advanced Commerce descriptors and the item
displayName and description make this reachable for any non-English app.
Also set the source encoding of the compile and javadoc tasks to UTF-8.
Without it javac reads the sources with the platform charset, which
corrupts the non-ASCII characters present in the javadoc of 25 files and
prevents the new regression test from expressing its expected values.

@alexanderjordanbakeralexanderjordanbaker left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thank you! Will be filing equivalent PRs on the other languages shortly

@alexanderjordanbaker
alexanderjordanbaker merged commit f0ddedd into apple:mainAug 27, 2026
4 checks passed
alexanderjordanbaker added a commit to apple/app-store-server-library-node that referenced this pull request Aug 28, 2026
alexanderjordanbaker added a commit to apple/app-store-server-library-python that referenced this pull request Aug 28, 2026
alexanderjordanbaker added a commit to apple/app-store-server-library-swift that referenced this pull request Aug 28, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@matteobaccan@alexanderjordanbaker
, 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' Decode JWS payloads as UTF-8 instead of the JVM default charset by matteobaccan · Pull Request #266 · apple/app-store-server-library-java · GitHub
Skip to content

Decode JWS payloads as UTF-8 instead of the JVM default charset - #266

Merged
alexanderjordanbaker merged 1 commit into
apple:mainfrom
matteobaccan:fix/utf8-jws-payload-decoding
Aug 27, 2026
Merged

Decode JWS payloads as UTF-8 instead of the JVM default charset#266
alexanderjordanbaker merged 1 commit into
apple:mainfrom
matteobaccan:fix/utf8-jws-payload-decoding

Conversation

@matteobaccan

Copy link
Copy Markdown

The problem

SignedDataVerifier.parseJWTPayload converts the base64url-decoded payload bytes into a String without specifying a charset:

String payload = new String(Base64.getUrlDecoder().decode(jwt.getPayload()));

new String(byte[]) uses the JVM default charset. A JWT Claims Set is always a UTF-8 encoded JSON object (RFC 7519 §3, RFC 8259 §8.1), so the two only agree when the JVM happens to run with UTF-8 as its default.

That is not guaranteed on the Java versions this library supports. JEP 400 made UTF-8 the default only in Java 18; on Java 11 through 17 the default charset comes from the platform locale, so a server running on Windows or with a non-UTF-8 LANG gets windows-1252, ISO-8859-1, Shift_JIS, and so on.

When that happens, every non-ASCII character in a decoded payload is corrupted silently. The signature is verified against the raw JWS, so verification still succeeds and no exception is raised — the caller simply receives mojibake. This affects every entry point that goes through decodeSignedObject: verifyAndDecodeTransaction, verifyAndDecodeRenewalInfo, verifyAndDecodeNotification, verifyAndDecodeAppTransaction and verifyAndDecodeRealtimeRequest.

It is easy to reach in practice. The Advanced Commerce descriptors and items carry merchant-supplied free text (displayName, description), so any app that is not English-only can hit it.

Reproduction

On JDK 17 with windows-1252 as the default charset, decoding a signed transaction whose advancedCommerceInfo.descriptors.description is Abonnement Café — 5,99 € par mois returns:

Abonnement Café — 5,99 € par mois

The existing test suite cannot catch this: all of its payload fixtures are ASCII-only, and CI sets JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8, which hides the difference.

The fix

new String(bytes, StandardCharsets.UTF_8) — one argument.

Source encoding

The build does not set options.encoding, so javac reads the sources with the platform charset too. 25 files under src/main contain non-ASCII characters (typographic apostrophes in the javadoc), which means those are corrupted in locally built javadoc on a non-UTF-8 machine. It also makes a regression test for this bug impossible to write in the natural way: the expected string literal would be corrupted by the compiler in exactly the same way as the actual value, and the assertion would pass.

This PR sets options.encoding = 'UTF-8' on the JavaCompile tasks and on javadoc. Note that ci-prb.yml currently compensates for the missing setting with JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8, while ci-release-javadocs.yml does not set it at all.

The test

SignedDataVerifierTest.testNonAsciiDataDecodingIsIndependentOfTheDefaultCharset decodes a new fixture containing French and Japanese text. Its expected values are written as \uXXXX escape sequences, so the test source is pure ASCII and the assertion does not depend on the encoding used to compile it.

Verification

On JDK 17, default charset windows-1252:

  • without the production change the new test fails with
    expected: <Abonnement Café — 5,99 € par mois> but was: <Abonnement Café — 5,99 € par mois>
  • with it, ./gradlew clean test and ./gradlew javadoc both pass

Not included

Two related items I left out to keep the diff focused, happy to add them if you would prefer them here:

  • the JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8 line in ci-prb.yml is now redundant for compilation and could be dropped
  • CI runs on a UTF-8 locale, so it would not catch a regression of this bug; a matrix entry running the tests with a non-UTF-8 default charset would lock the behavior in

I also did not touch CHANGELOG.md, since it appears to be updated by maintainers in the release commits.

The payload of a JWS is always UTF-8, but parseJWTPayload converted the
decoded bytes with new String(byte[]), which uses the JVM default charset.
On the Java versions this library supports that charset is platform
dependent, since JEP 400 only made UTF-8 the default in Java 18. On a JVM
running with, for example, windows-1252 every non-ASCII character in a
signed transaction, renewal info, notification or AppTransaction was
silently corrupted: the signature still verified, so no error was raised
and the caller received mojibake.
Free-text fields such as the Advanced Commerce descriptors and the item
displayName and description make this reachable for any non-English app.
Also set the source encoding of the compile and javadoc tasks to UTF-8.
Without it javac reads the sources with the platform charset, which
corrupts the non-ASCII characters present in the javadoc of 25 files and
prevents the new regression test from expressing its expected values.

@alexanderjordanbakeralexanderjordanbaker left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thank you! Will be filing equivalent PRs on the other languages shortly

@alexanderjordanbaker
alexanderjordanbaker merged commit f0ddedd into apple:mainAug 27, 2026
4 checks passed
alexanderjordanbaker added a commit to apple/app-store-server-library-node that referenced this pull request Aug 28, 2026
alexanderjordanbaker added a commit to apple/app-store-server-library-python that referenced this pull request Aug 28, 2026
alexanderjordanbaker added a commit to apple/app-store-server-library-swift that referenced this pull request Aug 28, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@matteobaccan@alexanderjordanbaker
, 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); Decode JWS payloads as UTF-8 instead of the JVM default charset by matteobaccan · Pull Request #266 · apple/app-store-server-library-java · GitHub
Skip to content

Decode JWS payloads as UTF-8 instead of the JVM default charset - #266

Merged
alexanderjordanbaker merged 1 commit into
apple:mainfrom
matteobaccan:fix/utf8-jws-payload-decoding
Aug 27, 2026
Merged

Decode JWS payloads as UTF-8 instead of the JVM default charset#266
alexanderjordanbaker merged 1 commit into
apple:mainfrom
matteobaccan:fix/utf8-jws-payload-decoding

Conversation

@matteobaccan

Copy link
Copy Markdown

The problem

SignedDataVerifier.parseJWTPayload converts the base64url-decoded payload bytes into a String without specifying a charset:

String payload = new String(Base64.getUrlDecoder().decode(jwt.getPayload()));

new String(byte[]) uses the JVM default charset. A JWT Claims Set is always a UTF-8 encoded JSON object (RFC 7519 §3, RFC 8259 §8.1), so the two only agree when the JVM happens to run with UTF-8 as its default.

That is not guaranteed on the Java versions this library supports. JEP 400 made UTF-8 the default only in Java 18; on Java 11 through 17 the default charset comes from the platform locale, so a server running on Windows or with a non-UTF-8 LANG gets windows-1252, ISO-8859-1, Shift_JIS, and so on.

When that happens, every non-ASCII character in a decoded payload is corrupted silently. The signature is verified against the raw JWS, so verification still succeeds and no exception is raised — the caller simply receives mojibake. This affects every entry point that goes through decodeSignedObject: verifyAndDecodeTransaction, verifyAndDecodeRenewalInfo, verifyAndDecodeNotification, verifyAndDecodeAppTransaction and verifyAndDecodeRealtimeRequest.

It is easy to reach in practice. The Advanced Commerce descriptors and items carry merchant-supplied free text (displayName, description), so any app that is not English-only can hit it.

Reproduction

On JDK 17 with windows-1252 as the default charset, decoding a signed transaction whose advancedCommerceInfo.descriptors.description is Abonnement Café — 5,99 € par mois returns:

Abonnement Café — 5,99 € par mois

The existing test suite cannot catch this: all of its payload fixtures are ASCII-only, and CI sets JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8, which hides the difference.

The fix

new String(bytes, StandardCharsets.UTF_8) — one argument.

Source encoding

The build does not set options.encoding, so javac reads the sources with the platform charset too. 25 files under src/main contain non-ASCII characters (typographic apostrophes in the javadoc), which means those are corrupted in locally built javadoc on a non-UTF-8 machine. It also makes a regression test for this bug impossible to write in the natural way: the expected string literal would be corrupted by the compiler in exactly the same way as the actual value, and the assertion would pass.

This PR sets options.encoding = 'UTF-8' on the JavaCompile tasks and on javadoc. Note that ci-prb.yml currently compensates for the missing setting with JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8, while ci-release-javadocs.yml does not set it at all.

The test

SignedDataVerifierTest.testNonAsciiDataDecodingIsIndependentOfTheDefaultCharset decodes a new fixture containing French and Japanese text. Its expected values are written as \uXXXX escape sequences, so the test source is pure ASCII and the assertion does not depend on the encoding used to compile it.

Verification

On JDK 17, default charset windows-1252:

  • without the production change the new test fails with
    expected: <Abonnement Café — 5,99 € par mois> but was: <Abonnement Café — 5,99 € par mois>
  • with it, ./gradlew clean test and ./gradlew javadoc both pass

Not included

Two related items I left out to keep the diff focused, happy to add them if you would prefer them here:

  • the JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8 line in ci-prb.yml is now redundant for compilation and could be dropped
  • CI runs on a UTF-8 locale, so it would not catch a regression of this bug; a matrix entry running the tests with a non-UTF-8 default charset would lock the behavior in

I also did not touch CHANGELOG.md, since it appears to be updated by maintainers in the release commits.

The payload of a JWS is always UTF-8, but parseJWTPayload converted the
decoded bytes with new String(byte[]), which uses the JVM default charset.
On the Java versions this library supports that charset is platform
dependent, since JEP 400 only made UTF-8 the default in Java 18. On a JVM
running with, for example, windows-1252 every non-ASCII character in a
signed transaction, renewal info, notification or AppTransaction was
silently corrupted: the signature still verified, so no error was raised
and the caller received mojibake.
Free-text fields such as the Advanced Commerce descriptors and the item
displayName and description make this reachable for any non-English app.
Also set the source encoding of the compile and javadoc tasks to UTF-8.
Without it javac reads the sources with the platform charset, which
corrupts the non-ASCII characters present in the javadoc of 25 files and
prevents the new regression test from expressing its expected values.

@alexanderjordanbakeralexanderjordanbaker left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thank you! Will be filing equivalent PRs on the other languages shortly

@alexanderjordanbaker
alexanderjordanbaker merged commit f0ddedd into apple:mainAug 27, 2026
4 checks passed
alexanderjordanbaker added a commit to apple/app-store-server-library-node that referenced this pull request Aug 28, 2026
alexanderjordanbaker added a commit to apple/app-store-server-library-python that referenced this pull request Aug 28, 2026
alexanderjordanbaker added a commit to apple/app-store-server-library-swift that referenced this pull request Aug 28, 2026
Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants

@matteobaccan@alexanderjordanbaker